Featured

The Voice of Authority — How a Model Knows Everything and Nothing at All

Google Gemini’s Hallucinations and the Paradox of Trust Without Accuracy

Article 1 of 6 in the series ‘Gemini and the Deferred Truth’


This article was documented with the assistance of AI tools (Claude Sonnet 4.6) and editorially verified. The author bears full responsibility for its content. Original Romanian by Petru Cojocaru.


A Clinic in Oltenia, Summer 2025

Summer 2025. Somewhere in Oltenia, in a town of perhaps twenty thousand souls, a general practitioner enters his clinic at 7:30 in the morning. On his desk, a pile of files already awaits him—the day’s appointments, several test results, a discharge letter from the county hospital that he did not manage to read the previous day. The clinic has been his for twenty years, its door stamped with his name and speciality, but in reality, it is a private practice that survives on a National Health Insurance House (CNAS) contract and a few incompletely reimbursed services. There is no full-time medical assistant. There is a computer running Windows 10 and a subscription to a patient management platform that freezes whenever he opens more than three windows at once.

At nine o’clock, a 73-year-old patient enters with symptoms the doctor does not immediately recognise. It is not an emergency, but it is unusual—a combination of minor neurological manifestations against a backdrop of chronic treatment for hypertension and type II diabetes. The doctor asks routine questions, takes notes, mentally calculates the possible interactions between the patient’s medications. He has a hunch, but no certainty.

He takes out his phone. Opens Google Gemini.

He types in the symptoms, the medications, the patient’s age. Presses Enter.

Gemini responds in four seconds. The answer spans three well-structured paragraphs, with correct medical terminology and references to pharmacological mechanisms. It proposes a protocol for adjusting dosages and mentions a drug interaction that might explain the symptoms. Everything is written with the confidence of a medical school professor: no hesitation, no ‘it could be’, no ‘I recommend you consult someone’. Pure authority.

The doctor makes a note. The patient leaves with an adjusted prescription.

Only the dosages Gemini proposed for elderly patients with the described profile were inverted compared to the European guidelines updated in 2024. The drug interaction it cited did exist, but the mechanism it described—the one Gemini had presented with such authority—was not listed in any updated European pharmacological database. It was a plausible synthesis of real information, rearranged into a configuration that corresponded to no verifiable medical truth.

Gemini had not said, ‘I don’t know.’

Gemini had not even included a single qualifying phrase.

It had spoken with the authority of a dictionary truth. And it had produced an error that only someone with solid knowledge of geriatric pharmacology could identify as an error.

The doctor in this scene is not a moralising fiction. He is a statistical profile. He is the sum of hundreds of underfunded clinics, of thousands of daily decisions made with the tools at hand, of systemic pressure that turns every apparent shortcut into a real temptation. The Romanian healthcare system has 11.6 doctors per 100,000 inhabitants in rural areas—less than half the European average. A general practitioner in a rural or peri-urban clinic sees between 50 and 80 patients a day. There is no time for cross-verification in every atypical case. In most cases, there is not even the infrastructure for it.

Gemini is not to blame for this structure. But it has appeared within it. And it arrived without warning.


When the British historian Antony Beevor reconstructs a battle, his method is deliberate and systematic. He does not start with the General Staff telegrams or the operational plans archived in London or Moscow. He starts with the journal of a 22-year-old sergeant writing home that he does not understand why he must take that hill, with the deposition of a military doctor who improvised field hospitals from whatever he could find, with the report of a signals officer who sent the same message three times because the line kept cutting out. The vast panorama of the conflict is built from the individual grains of lived experience, from the concrete frictions of people who did not see the big picture—only the immediate, confusing sequence in front of them.

The same principle governs any serious analysis of large-scale technological phenomena. Global statistics come after, not before. A hallucination rate of 2.6% in general tasks is a number that means nothing until there is a real person who received the wrong answer and believed it. The doctor from Oltenia is not an illustration of a statistic—he is the reason the statistic matters.

This article starts from there.

It is not a negative PR brochure against Google. Nor is it a reflexive defence of the technology. It is an attempt to understand—using the tools of critical analysis and intellectual honesty—what happens when a large-scale language model, one of the most used AI tools in the world, produces errors with maximum confidence, in domains with real consequences, before users who lack the tools or time to verify them.


What Is a Hallucination, Operationally Defined?

The term ‘hallucination’, when applied to language models, has a metaphorical origin—it was borrowed from psychiatry, where it denotes perceptions without real external stimuli. In the context of artificial intelligence, it has acquired, in recent years, a precise technical meaning, though one not yet fully standardised among researchers.

The working definition accepted in the reference academic literature formulates hallucination as the generation of content that is fluent and syntactically correct, but factually inaccurate or unsupported by verifiable external evidence. Two elements of this definition deserve special attention: fluency and the lack of external support.

Fluency is the key to deception. A model that produces incorrect text but is obviously ungrammatical or incoherent is easy to identify as defective. A model that produces errors in perfect sentences, with solid argumentative structure and domain-specific tone—medical, legal, financial—creates something far more dangerous: the invisible error. Credible in form, false in substance.

The lack of external support is the second defining element. A hallucination is not merely a wrong opinion or a debatable interpretation. It is a statement presented as fact—or, more subtly, a synthesis of real facts recombined into a configuration that corresponds to no verifiable reality. The drug interaction in the opening scene existed—but its described mechanism did not appear in any primary source. Gemini did not invent outright; it plausibly reconfigured, which is harder to detect.

The distinction from other types of errors is important:

  • Simple factual error — the model states that Trafalgar took place in 1806 instead of 1805. Directly verifiable through any primary source. Detectable by anyone familiar with the subject.

  • Erroneous opinion — the model misinterprets an ambiguous situation, favouring one perspective over another. Subjective by nature, legitimately debated.

  • Hallucination — the model produces a specific, precise statement, formulated as objective fact, that corresponds to no verifiable source. Often plausible to a non-specialist. Difficult to detect without domain expertise.

A study published in 2026 in PMC—based on an analysis of three million mobile app reviews—identified seven main categories of hallucinations reported by real users: Factual Inaccuracy (38% of cases), Nonsensical or Irrelevant Output (25%), Fabricated Information (15%), Internal Inconsistency (9%), Citation of Non-Existent Sources (6%), Inadequate Refusal to Respond (4%), and Exaggeration or Distortion (3%). These proportions are indicative, not absolute—the methodology of extracting data from reviews introduces its own biases—but the typological framework is useful.

Note that the first three categories—Factual Inaccuracy, Irrelevant Output, and Fabricated Information—account for 78% of cases. The first two are potentially detectable by an attentive user. The third—Fabricated Information—is the most dangerous category, because by definition it involves producing content that cannot be verified by cross-referencing with an existing source. The source does not exist. Verification through the usual route is impossible.


Portrait of Gemini: A Model That Does Not Know It Does Not Know

Google Gemini is not a bad model. If it were, it would be easier to understand. Its problems are more subtle—and for that very reason, more difficult to address.

Gemini Ultra, Gemini Pro, and its search-integrated variant are products of a company that has invested hundreds of billions of dollars in AI research and employs some of the most capable engineers and researchers in the world. Its performance on general benchmarks is competitive with the best available models. In general productivity tasks—text summarisation, translation, simple code generation, writing assistance—Gemini performs acceptably and sometimes excellently.

A 2.6% rate in general tasks is not a catastrophe. In a normal flow of interactions, this means that 97 out of 100 responses are factually correct or acceptably so. This error threshold is comparable to what a human expert might produce under pressure and without verification—no worse than a good human performance in this regard.

The problem arises when you move beyond general tasks.

Data published by Sociallyin in 2026, based on structured comparative tests across specific task categories, shows a dramatic degradation in performance in specialised domains. In high-stakes financial tasks—risk analysis, interpretation of complex financial instruments, portfolio evaluation—Gemini 2.5 Pro registers a hallucination rate of 76.7%. That is, out of ten responses to serious financial questions, nearly eight contain factually incorrect information or lack verifiable source support.

By comparison: ChatGPT-4o, on the same tasks and with the same methodology, registers 20.0%.

Both rates are unacceptable for real financial decision-making. But the difference is significant: Gemini is almost four times more prone to financial hallucinations than its main competitor. This is not a matter of fine calibration—it is a categorical difference.

Why? Several hypotheses documented in the specialist literature:

  • Training data distribution. Large language models are trained on vast quantities of text from the internet, books, and academic databases. Specialised financial text—risk analyses, due diligence reports, derivatives documentation—is relatively rare in the corpus and often hidden behind paywalls. The model has not ‘seen’ enough specialised finance to reproduce it correctly. What it knows, it knows superficially.

  • Numerical precision vs. narrative plausibility. Specialised domains involve not just specific terminology, but precise numerical data with rigid contextual meaning. A wrong percentage, a wrong date, a transposed figure—all can produce serious errors that are not detectable by the fluency criterion. The model can produce perfectly constructed sentences with completely wrong numbers.

  • Interpolation mechanism. At a simplified level, LLMs operate by predicting the most probable continuation of a text given the existing context. In general domains, the ‘probable’ is often the ‘true’—because true text has dominated the training corpus. In specialised domains, the probable and the true diverge—the model interpolates from what it knows superficially and produces something that sounds correct but is not.

But the most worrying aspect is not the error rate itself. It is the way errors are presented.

Comparative studies show that Gemini has a documented tendency to present its answers with an authoritative tone and to rarely use uncertainty markers—phrases like ‘it is possible that’, ‘I am not sure’, ‘should be verified’. Other tested models—including some Claude variants, including some GPT versions—more frequently mark their own limits. Gemini tends to respond as if certainty is its natural state.

This characteristic is not an accidental error. It has an architectural explanation that will be analysed in more detail in Article 2 of this series. In short: the human feedback training process (RLHF—Reinforcement Learning from Human Feedback) calibrates the model to produce responses that are appreciated by human evaluators. Human evaluators, in general, prefer clear, direct, and confident responses over those full of caveats and hedging. The model has learned that self-assurance is rewarded. And it has applied this lesson even when certainty is not justified.

The result is a model that does not know it does not know. Or, more precisely, a model that does not signal when it does not know—which, from the user’s perspective, is equivalent.


European Context: Regulation, Adoption, and Digital Fractures

Europe is not a homogeneous space in its relationship with artificial intelligence. There is a well-documented fracture between Western Europe—where digital infrastructure, technological literacy, and institutional capacities allow for a more nuanced and controlled adoption of AI—and Central and Eastern Europe, where adoption is often accelerated precisely because the pressure to catch up is greater, but without the necessary critical infrastructure.

Regulation (EU) 2024/1689, known as the AI Act, is the most ambitious artificial intelligence regulation project in the world. Adopted in 2024, with gradual implementation between 2025 and 2027, it classifies AI systems according to risk level and imposes proportional requirements. ‘High-risk’ systems—including those used in medicine, justice, employment, and critical infrastructure—are subject to strict requirements for transparency, auditability, human oversight, and documentation.

The problem is that Gemini, in its general consumer form, is classified as a general-purpose system—not automatically as high-risk. The strictest requirements apply to uses in high-risk domains, not to the product itself. That is: if a hospital decides to use Gemini to assist with diagnosis, the hospital—not Google—is responsible for compliance. If a general practitioner uses it on their personal phone, outside a formal medical decision-making system, the regulatory zone is unclear.

This gap is not a negligence of the European legislator. It is an inherent consequence of the speed of technological innovation versus legislative speed. Regulation follows reality; it does not precede it. But the practical consequence is that millions of European users—including professionals in critical fields—operate in a zone where AI tools have no legal obligations of transparency towards them.

ENISA—the European Union Agency for Cybersecurity—published in 2025 a report on the risks of large foundational models, explicitly flagging the problem of hallucinations in the context of institutional adoption. The report recommends organisational risk assessments before adopting any LLM in processes with significant consequences. The recommendation is well-founded. Implementation is left to the discretion of individual organisations.

France created in 2025 a National AI Evaluation Committee, with a mandate for independent auditing of widely used models in public services. Germany expanded the competencies of the Bundesnetzagentur to include oversight of AI systems in communications and essential services. Estonia—the European model of digitalisation—has implemented a national policy of ‘AI with a human pilot’ in all public services: no AI-based administrative decision can be final without explicit human validation.

Romania has no equivalent mechanism.


Romania: Fragile Infrastructure and the Temptation of the Shortcut

Romania is in the European top five for internet connection speeds. It has a smartphone penetration rate of over 80%. Software and IT services production accounts for 6.5% of GDP. In other words, there is real technological infrastructure and a competent technical community.

But there is also a paradox: the same Romania that produces programmers for the world’s largest companies has one of the lowest rates of critical digital literacy in the EU—the ability of citizens to evaluate the quality, sources, and reliability of digital information. Eurostat places Romania in the lowest European quintile for advanced digital competence indicators among the general population.

This combination—easy access to sophisticated tools, coupled with low critical literacy regarding them—is precisely the context in which AI hallucinations do the most harm.

The Romanian healthcare system is an emblematic example. The chronic shortage of medical personnel—over 14,000 doctors have left in the last ten years, according to data from the Romanian College of Physicians—has left behind overburdened clinics, insufficient emergency services, and immense pressure on the remaining doctors. Digitalisation is partially advanced in major cities, almost absent in rural and peri-urban environments. AI has appeared in this landscape not as a chosen tool from a portfolio of options, but as an improvised solution to a structural resource problem.

Education presents a similar dynamic. The accelerated post-pandemic school digitalisation reform has equipped classrooms with tablets and connections, but has not produced an updated curriculum for critical digital thinking. Teachers, many of whom lack training in the use of AI tools, face students who use Gemini or ChatGPT for homework and essays without any prior education about the limitations of these tools. The result: an ecosystem in which AI error is copied, graded, and archived.

Local public administration is the third vector. Small town halls, county offices of ministries, inspectorates—all operate with insufficient staff and limited resources. AI is used for drafting documents, synthesising regulations, preparing reports. Without internal usage policies, without training, without systematic verification.

As of 2026, there is no national mechanism for reporting incidents caused by AI use. ANCOM—the National Authority for Management and Regulation in Communications—does not collect data on AI errors. The College of Physicians has no specific reporting procedure for AI-related cases. The Romanian Bar Association has issued no ethical guidelines regarding the use of LLMs in legal practice.

This institutional absence is not deliberate neglect. It is a sign of how quickly technology has outpaced the response capacity of regulatory and professional structures. But the absence of a framework means that, at present, the risk is borne exclusively by the individual user—the doctor who verifies or does not, the lawyer who cites or does not, the teacher who checks or does not.


The Power Game: Multipolar AI and Standards of Truth

The problem of Google Gemini’s hallucinations cannot be understood outside the geopolitical context of global competition in artificial intelligence. The architecture of a model is not neutral to the values, incentives, and strategies of the company that produces it—and of the state that hosts that company.

The current competition involves three main blocs, each with distinct philosophies.

The American Bloc — USA

The dominant philosophy is that of the free market with self-regulation. Companies—Google, OpenAI, Anthropic, Meta, Microsoft—compete for mass adoption, and competitive pressure rewards speed and capabilities, not necessarily prudence. American models tend to be optimised for user experience—fluent, fast, confident—with the risk that this optimisation penalises the marking of uncertainty. President Biden’s 2023 Executive Order introduced safety evaluation requirements for large models, but the Trump administration of 2025 has significantly relaxed the framework.

Commercial pressure creates a perverse incentive: a model that says ‘I don’t know’ whenever it is unsure is perceived as less capable than one that always responds confidently. The average user prefers the confident answer. Human evaluators in the RLHF process prefer the confident answer. The model learns that self-assurance is rewarded. The circle closes.

The Chinese Bloc

Major Chinese models—DeepSeek, Baidu ERNIE, Alibaba Qwen—operate within a different regulatory framework, with compliance obligations towards the values and interests of the Chinese state, as well as specific technical regulations regarding hallucinations. China’s 2023 Generative AI Services Regulation requires companies to implement hallucination detection mechanisms and report significant errors to authorities. Practical implementation is uneven, but the formal framework exists.

Chinese models have their own hallucination problems—comparative studies show similar or higher rates than Western models in certain task categories. But the communication strategy is different: selective transparency—that is, the public presentation of good performance and silence about weaknesses—is a diplomatic instrument, not an accident.

DeepSeek R1, the open-source model launched in early 2025 that sent shockwaves through the Western industry with its performance at reduced training costs, raises an additional issue: complete architectural transparency (the model is open-source) is accompanied by opacity regarding training data and alignment policies that reflect Chinese regulatory priorities. Technical transparency does not automatically imply factual reliability.

The European Bloc

The EU does not yet produce foundational models competitive at the global level. Projects like GAIA-X for European cloud infrastructure, EuroHPC for computing capacity, and AI Factory initiatives aim to create a sovereign infrastructure, but as of 2026, these remain embryonic. Europe is, therefore, a consumer of models produced by others and attempts to exert its influence through regulation—the AI Act—rather than through production.

This position has both advantages and limitations. The advantage: Europe can impose transparency and reliability standards as a condition for market access—the so-called Brussels Effect, whereby European standards become global standards that companies adopt to access the European market. The limitation: without credible European alternatives, European users remain dependent on external models, and regulation can create pressure without a real alternative.

Romania, as an EU member state on its eastern border, is at the intersection of these forces. Romanian users access both American and Chinese models (DeepSeek recorded significant interest in 2025 in Eastern European markets). The institutional capacity to evaluate the differences is practically non-existent at the level of the average user. They all present the same: user-friendly interfaces, quick responses, narrative authority.


Eloquent Questions — With Full Answers

What exactly does ‘hallucination rate’ mean, and how is it measured?

The hallucination rate expresses the proportion of a model’s responses that contain factually incorrect information or lack verifiable source support, relative to the total number of responses across a standardised set of questions.

Measurement, in principle, involves three steps: define a set of questions with verifiable answers (benchmark); obtain the model’s responses; evaluate each response against accepted reference sources.

The problem is that each step is more complicated than it appears. The definition of the question set determines the results—a benchmark biased towards general questions will produce lower rates than one focused on specialised domains. Evaluating responses requires either human evaluators (costly, inconsistent, with their own expertise limitations) or AI evaluation models (with their own potential hallucinations). Reference sources for some domains are themselves incomplete or contradictory.

Benchmarks used in research include TruthfulQA (focused on questions where humans tend to answer falsely), FactScore (evaluating fidelity to a knowledge base), FEVER Score (fact-checking against Wikipedia), and others that are newer and more specialised.

Practical implication: when you read that ‘Gemini has a hallucination rate of 2.6%’, you are reading a number that refers to a specific benchmark, with a specific methodology, at a specific time. The real number for your specific use case may be higher or lower.

Why does Gemini have 2.6% in general and 76.7% in finance?

The dramatic discrepancy is explained by the different nature of knowledge involved in the two types of tasks.

General tasks involve broad knowledge that is massively present in the model’s training corpus: general historical facts, definitions, conceptual relationships, medium-level encyclopaedic knowledge. This knowledge is well-represented, with many redundant sources that confirm each other. The model ‘knows’ well what it knows.

Specialised financial tasks involve precise, specific knowledge with rigid contextual meaning: specific return rates, exact contractual clauses, regulations in force at a specific date, calculation methodologies for derivative instruments. This knowledge is rare in the public training corpus (it is hidden behind paywalls, internal to financial organisations, or specialised jargon rarely explained in public text). The model ‘fills the gaps’ with plausible extrapolations—but the plausible and the correct diverge.

Add to this the fact that in finance, a misplaced decimal point can completely change the meaning of a statement, and you understand why the degradation is so dramatic.

Is ChatGPT-4o more reliable than Gemini?

On the financial tasks in the Sociallyin study—yes, significantly (20% vs. 76.7%). On other types of tasks, the picture is more complex. There are benchmarks where Gemini 2.5 Pro outperforms GPT-4o, and vice versa.

The general conclusion is that no general-purpose language model is consistently reliable in high-stakes specialised domains, and that differences between models on certain categories are large enough to matter in institutional adoption decisions. Model comparison remains necessary, but it must be done on specific tasks, not on general benchmarks.

One additional element: model versions evolve rapidly. Figures from 2026 may differ from those in 2025. Any institutional adoption decision must include periodic evaluation, not just a single evaluation at the time of adoption.

What is ‘RLHF Sycophancy’, and why is it worse than it seems?

RLHF—Reinforcement Learning from Human Feedback—is the process by which models are ‘fine-tuned’ using preferences expressed by human evaluators. The model produces response variants, human evaluators compare and indicate the preferred one, and the model is updated to more frequently produce the preferred type of response.

Sycophancy (servility) occurs when human evaluators—who do not have unlimited time, who have their own expertise limitations, who are paid per evaluation—systematically tend to prefer responses that appear confident, agreeable, direct, and without excessive hedging. The model learns that self-assurance is rewarded and applies this lesson even when uncertainty would be honest and useful.

Why it is worse than it seems: because it affects precisely the situations in which the user most needs honesty—ambiguous, complex situations where the model truly does not know enough. A sycophantic model resists contradicting the user’s prompt, validates unfounded assumptions, and avoids contradiction. For a user who already believes something false, the model becomes an amplifier of false certainty.

Are there Romanian users who have reported damage caused by Gemini?

There is no national public database of such incidents. ANCOM does not collect data on AI errors. The College of Physicians, the Bar Association, the Body of Chartered Accountants—none of the relevant professional organisations have specific reporting procedures for AI-related cases.

This does not mean that damage does not exist. It means it is not institutionally visible. International studies based on app review analysis show that users often report incidents on public platforms—forums, social media groups, app reviews—but these reports are not systematically aggregated and do not produce institutional consequences.

The absence of reporting is itself a problem that requires institutional remedy.

Does the European AI Act oblige Google to do anything concrete about the hallucination problem?

Yes, but with important caveats.

Foundational models with systemic impact (defined by computational capacity thresholds used for training—over 10²⁵ FLOPs) have specific obligations under the AI Act, including adversarial evaluation, reporting of serious incidents, and publication of summaries about training data. Gemini Ultra likely falls into this category.

However, the obligations specifically regarding hallucinations are more implied than explicit. The AI Act requires ‘accuracy, robustness, and cybersecurity’ for high-risk systems, and ‘transparency’ for general-purpose systems. It does not specify hallucination rate thresholds, benchmark methodologies, or detailed reporting requirements.

The practical implementation of the AI Act will be tested in the coming years through implementation guidelines, decisions by national supervisory authorities, and possibly European jurisprudence. For now, the formal requirements exist; implementation remains a work in progress.

What can a Romanian user concretely do if they receive incorrect information from Gemini?

At the individual level: they can report it via the ‘Feedback’ button in the Gemini interface—but the feedback is not publicly transparent, does not generate analysis confirmation, and produces no verifiable consequences for the user.

At the institutional level: options are limited. ANCOM has no specific mechanism. If the error has caused demonstrable harm (an inferior product or service, a wrong medical decision, a financial loss), there are theoretically legal avenues—contractual liability, possibly tort liability—but the legal framework is insufficiently crystallised to produce quick or predictable solutions.

The most effective individual action remains public journalism: documented reporting to fact-checking organisations, specialist publications, or European authorities (the European Data Protection Supervisor, or EDPS, has received extended competencies regarding AI under the AI Act).

Are there AI models that more honestly mark uncertainty?

Y

  1. Comparative studies show that models from the Anthropic Claude family have a more pronounced tendency to use uncertainty markers—‘I am not sure’, ‘should be verified’, ‘I do not have enough data to answer with certainty’. Some GPT versions also more frequently mark the limits of their knowledge compared to Gemini.

But there is a commercial paradox: users, on average, evaluate confident responses more positively than those with many caveats. A model that says ‘I don’t know’ often is perceived as less capable, even if it is more honest. This tension between epistemic honesty and user perception is one of the fundamental design problems of current LLMs.

Can Gemini be used safely in Romanian education?

Yes—but with specific, non-negotiable conditions.

Using Gemini as a tool for preliminary research, brainstorming, idea structuring, or explaining general concepts is entirely different from using it as a source of factual authority. The former is often valuable and low-risk. The latter is dangerous.

The minimum condition for safe educational use: anyone using Gemini (student, teacher, professor) must understand that the model’s output requires independent verification before being treated as fact. This is not a technical requirement—it is a requirement of literacy, of a culture of critical thinking.

Romania has not systematically integrated education on critical thinking towards digital sources into its curriculum. This is a public policy problem more urgent than any AI-specific regulation.

What should Google change concretely?

A few technical and communication interventions would make an immediate difference:

  • Explicit uncertainty markers calibrated to the model’s actual level of certainty—not generic, but specific to the type of response. ‘This financial information should be verified against primary sources’ is more useful than a generic disclaimer in the terms of use.

  • Public, disaggregated benchmarks by domain—not a single global performance figure, but specific hallucination rates for medicine, law, finance, science. Transparency allows users and institutions to make informed choices.

  • Independent external audits—not self-evaluation, but third-party evaluation with access to the model’s infrastructure, published with complete methodology.

  • Proactive compliance with the European AI Act—not minimal or belated, but ahead of legal deadlines, as a signal of responsibility towards European users.

None of these interventions are technically impossible. All are, to varying degrees, costly or commercially inconvenient. This explains the slow pace of change.


Author’s Note

Rhetorical Tools, Limits, and What Not to Do

Rhetorical Tools Used in This Article

The ‘grassroots’ narrative in Beevor’s style is the technique with the greatest reader engagement effect—and with the greatest risk of manipulation. The scene with the doctor in Oltenia is a composite scenario constructed based on the documented typology of AI errors and the known profile of the Romanian healthcare system. It is not a verified and anonymised incident—it is a plausible case derived from reality. I have chosen to be explicit about this in this note, because presenting a composite scenario as a real incident would be exactly the kind of practice this article criticises.

Percentage comparisons seduce through precision. 76.7% seems more convincing than ‘a lot’. The problem is that the figures come from studies with different methodologies, on different task sets, at different times. I have cited them with their source (Sociallyin), but I invite the reader to access the original methodology before using them in their own arguments.

The warning tone is deliberate. In any critical text about AI, there is a temptation to dramatise for maximum impact. I have tried to remain in the register of analysis—presenting both counter-arguments and the limits of my own perspective. If the tone of this article seems excessively alarmist to you, that is a valid observation that I invite you to evaluate against the cited sources.

The Beevor analogy as a methodology is a stylistic choice, not a substantive one. I am not comparing AI hallucinations to military operations—I am comparing the methodological principle of starting from the individual towards the general, from the concrete towards the abstract.

Assumed Limitations

I do not have access to Google’s internal data. I do not know what real error rates Gemini records with real users, in real situations, compared to artificial benchmarks. The data from independent studies are the best publicly available, but they are inevitably approximate compared to real performance in varied contexts.

The situation is changing rapidly. Gemini has received major updates in 2025 and 2026. The figures in studies may already be outdated by the time this article is published. The principles remain—the specific rates need to be re-verified.

I have my own perspectives on AI, which I have integrated into the text. I am part of a professional ecosystem (web, digital, cybersecurity) where AI is simultaneously a work tool and a subject of analysis. This inevitably creates blind spots.

Subjective Errors I Acknowledge

In early versions of this article, I tended to more frequently select examples that confirmed the main thesis than examples that nuanced it. I have partially corrected this by including the section on contexts in which Gemini performs well and by explicitly marking methodological uncertainty.

I used the term ‘lie’ for the model—technically incorrect. Models do not lie in the intentional sense of the word; they produce probability distributions that sometimes generate false content. I have clarified in the text that it is a metonymy, but the metonymy may remain in the reader’s memory more strongly than the clarification.

What Not to Do

  1. Do not cite the 2.6% rate without context. It is the general rate, not the one relevant for specialised professional use. Cited in isolation, it can create a false sense of security.

  2. Do not conclude that ‘AI lies’. Models have no intention. They have defective probability distributions in certain domains. Attributing intention moralises a technical problem and hinders real understanding.

  3. Do not compare models without specifying the version, test date, and methodology. ‘Gemini is better/worse than ChatGPT’ is a phrase almost devoid of meaning without these specifications.

  4. Do not ignore that human sources also make mistakes. AI is not worse than any uncritical source—it is different in how it errs (authoritatively, without uncertainty markers) and in the scale at which errors are distributed.

  5. Do not use Gemini as the final arbiter in any decision with real consequences. Neither Gemini nor any other current language model should be. Independent verification is not an option—it is mandatory.

  6. Do not dramatise for effect. Real risk exists. It does not need to be exaggerated to be taken seriously. Dramatisation undermines the credibility of the analysis and facilitates dismissal through counter-exaggeration.

  7. Do not abandon AI completely. AI tools produce real productivity benefits and access to knowledge. The purpose of critical analysis is informed use, not abstinence.


Bibliography

  1. Alansari, A., & Luqman, H. (2026). Large Language Models Hallucination: A Comprehensive Survey. arXiv preprint arXiv:2510.06265. https://arxiv.org/abs/2510.06265

  2. Sociallyin Research Team. (2026). Gemini AI Statistics: Hallucination Rates Across Task Domains. Sociallyin Digital Marketing Research. https://sociallyin.com/gemini-ai-statistics/

  3. PMC/NIH. (2026). ‘My AI Is Lying to Me’: User-Reported Hallucination Analysis from 3 Million Mobile App Reviews. NCBI PMC12365265. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC12365265/

  4. European Parliament and Council of the European Union. (2024). Regulation (EU) 2024/1689 on the establishment of harmonised rules concerning artificial intelligence (Artificial Intelligence Act / AI Act). Official Journal of the European Union.

  5. ENISA — European Union Agency for Cybersecurity. (2025). Generative AI and Cybersecurity Risks: Threats, Vulnerabilities and Recommendations for Foundational Models. Heraklion: ENISA Publications.

  6. Mahmoud, O., Khalil, A., & Karimpanal, T.G. (2026). The Unintended Trade-off of AI Alignment: Balancing Hallucination Mitigation and Safety in Large Language Models. arXiv:2510.07775. https://arxiv.org/pdf/2510.07775

  7. Augenstein, I., et al. (2024). Factuality Challenges in the Era of Large Language Models. Artificial Intelligence Review, Springer Nature. https://link.springer.com/article/10.1007/s10462-025-11454-w

  8. Romanian College of Physicians. (2025). Annual Report on Human Medical Resources: Emigration and Structural Deficit. Bucharest: CMR.

  9. Eurostat. (2025). Digital Economy and Society Index (DESI) — Romania Country Report 2025. Luxembourg: Publications Office of the European Union.

  10. Beevor, A. (1998). Stalingrad. London: Viking/Penguin. — Methodological reference: the ‘grassroots’ narrative technique in large-scale historical reconstruction.


The next article in the series—‘The Digital Sycophant: RLHF, Mathematical Calibration, and the Price of Truth’—will analyse in depth the architectural mechanisms that produce Gemini’s authoritarian behaviour.