⏱️ Reading time: 11 min
A developer searched Google for a nine-word phrase about a 2014 basketball meme, and the search engine responded as if he had just ended a relationship. Instead of links to old forum posts or tweets, Google Search returned an AI-generated summary consoling him over a breakup that never happened.
📑 En este artículo
- TL;DR
- What are AI hallucinations?
- Why it matters when a search engine hallucinates
- How it works: from search to generated text
- Practical examples: when ambiguity breaks search
- Where else algorithmic confabulation shows up
- Common mistakes and best practices
- Comparison: classic search, generative summary, and RAG with citations
- Going deeper: why no patch fully solves this
- Frequently Asked Questions
- Why does Google AI Overviews sometimes respond with emotional support instead of a link?
- Are the model’s made-up answers a programming error?
- How do I spot algorithmic confabulation in a search summary?
- Does a RAG system with citations like Perplexity or Bing Copilot fabricate less than a search engine with an automatic summary?
- Can I turn off Google’s generative summaries?
- References
The case, recounted on Sancho Panza’s personal blog, is anecdotal, but it exposes something structural. This is what the industry calls AI hallucinations, a phenomenon worth understanding whether you code, do research, or simply search for information every day.
TL;DR
- Language models generate the statistically most probable word, not the most truthful one.
- Google AI Overviews combines retrieval (RAG) with generation: if it retrieves the wrong context, the model writes it up with total confidence.
- No current architecture verifies its own claims before showing them to the user.
- A proper name without context (like ‘Dario’ with no last name) can trigger an empathetic response instead of search results.
- Perplexity and Bing Copilot show numbered citations; a summary without links can’t be audited.
What are AI hallucinations?
AI hallucinations are responses generated by a language model that present false, made-up, or unsupported claims as facts, without backing from any real source. They happen because the model doesn’t verify data: it predicts the statistically most probable sequence of words, with the same confidence whether or not there’s evidence behind it.
The term became popular starting in 2022, when ChatGPT and other assistants began inventing academic citations, legal cases, or programming functions that don’t exist. Natural language processing research documents them as a structural problem, not as an occasional error that a one-off patch will fix.
Researchers usually distinguish two types. Intrinsic hallucination directly contradicts the context the model received; extrinsic hallucination adds information that context never even mentioned. In an AI-powered search engine, both coexist: the model can misrepresent the retrieved passage or simply fill in something that came from no passage at all.
Why it matters when a search engine hallucinates
A traditional search engine shows ten links and lets the user decide which one to trust. An AI-generated summary makes that decision for you: it synthesizes several sources into a paragraph and presents it with the same confident tone, whether or not the data is solid.
Google integrated these generative summaries, known as AI Overviews, into its main search engine starting in 2024, according to its own official Search blog. For millions of daily queries, that means AI hallucinations no longer show up only in a separate chatbot: they appear above the ten blue links, in the position that gets the most clicks.
For a developer the risk is concrete. If you search for how to use a library flag or an API’s behavior and the summary gives you made-up information with the same confidence as correct information, you end up copying code that compiles but fails in production, or citing a fact that doesn’t exist in any real documentation.
How it works: from search to generated text
When you type a query, the system doesn’t hand your raw question to the language model. It first converts it into a numerical vector (an embedding) and searches, among millions of indexed pages, for the passages whose vector is closest to your query’s. This technique is called retrieval-augmented generation (RAG) and was formalized in a 2020 paper that combined a classic search engine with a generative model.
flowchart TD
A["User query"] --> B["Generate embedding"]
B --> C["Search for similar passages in the index"]
C --> D["Build context from retrieved passages"]
D --> E["Language model writes the answer"]
E --> F["Generative summary shown to the user"]
The problem shows up when the search retrieves the wrong context. If your query is ambiguous (a proper name like “Dario” with no last name or basketball context), the system might pull passages about romantic relationships instead of an NBA player. The model doesn’t know the context is wrong: it writes the best possible answer with what it was given, with the same confidence as if the data were correct.
The model doesn’t search for the truth: it searches for the most probable continuation given the context it received. Under the hood, it generates the answer token by token, calculating a probability distribution over thousands of possible words at each step and picking one of the most probable ones, not the most truthful one. That mechanism works perfectly for writing fluent text, and it’s the same reason it can sound confident while saying something false.
💭 Key point: A language model doesn’t have an “I don’t know” button: it always calculates a probability distribution over words, even when there’s no real context behind it.
Practical examples: when ambiguity breaks search
Sancho Panza’s case illustrates the pattern better than any abstract explanation. In 2014 the NBA’s Philadelphia 76ers drafted Dario Saric, a Croatian power forward who was playing professionally in Turkey at the time. Saric announced he would finish his contract there before joining the team, and some fans started joking that he was “never coming over,” an inside joke that circulated on forums and social media for years.
A decade later, Sancho wanted to dig up those old tweets and literally searched “hes never coming over dario.” Google, which has no way of knowing that “Dario” was a basketball player and not a real person in the user’s life, generated an empathetic response about how to process the pain of someone cutting off contact. The search engine acted as if it had to console him instead of helping him find information.
The links he was actually looking for (old posts about the meme) were on the page, but several hundred pixels down, after the generated block. That visual hierarchy, the AI summary first and the verifiable results after, is what turns a made-up answer into the user’s first, and sometimes only, contact with the information.
Where else algorithmic confabulation shows up
The Dario pattern isn’t an isolated basketball case. The same ambiguous-retrieval mechanics show up in other technical and everyday contexts.
- Proper name ambiguity: without a last name, team, or domain, the model assumes the meaning most common in its training data, not the one you had in mind.
- Fabricated citations and sources: ask an assistant for an academic reference on a very specific topic and it will likely invent a title, author, and year that sound plausible.
- Outdated information presented as current: the model doesn’t always know when a price, a software version, or an executive position changed.
- Nonexistent API functions or parameters: when completing code, an LLM can invent a flag or method that doesn’t exist in the real library, with perfect syntax.
Common mistakes and best practices
The most common mistake is treating the first answer that appears as the most trustworthy, when it’s actually the one that went through the least verification. Before copying a fact from an AI-generated summary, it’s worth applying a few simple checks.
- Scroll down to the original source: a generative summary almost always cites or links to pages below it; read that page before trusting the summary.
- Be wary of proper names without context: if your search has an ambiguous name, add the last name, domain, or a keyword from the real context.
- Ask for the explicit source: if you use a chatbot for research, ask it to cite the specific document or link, not just the claim.
- Verify generated code before running it: a made-up flag or function compiles silently until it fails in production; check the library’s official documentation.
💡 Tip: on Google you can force the classic list of links without the generative summary by adding the parameter &udm=14 to the end of the search URL.
Comparison: classic search, generative summary, and RAG with citations
Not all AI-powered search systems carry the same risk of generative fabrication. The key difference is whether the system shows where each piece of information came from.
| System | How it gets the answer | Risk of algorithmic confabulation | How to verify |
|---|---|---|---|
| Classic search engine (10 links) | Indexes and ranks pages by relevance, without generating new text | Low: the user decides which link to trust | Open the page and read the source |
| Generative summary (AI Overviews) | RAG with retrieved context and automatic writing | High if retrieval pulls the wrong context | Scroll down to the links the summary cites, if it shows any |
| RAG with verifiable citations (Perplexity, Bing Copilot) | RAG with numbered links next to each claim | Medium: still depends on retrieval, but it’s auditable | Open the numbered citation tied to the specific fact |
| Pure chatbot without retrieval | Only the knowledge learned during training | High for recent or very specific topics | Ask for the source and look it up yourself independently |
Going deeper: why no patch fully solves this
AI hallucination isn’t an isolated bug that a patch will fix, but the expected result of optimizing a model for fluency instead of truthfulness. Reducing AI hallucinations to zero would require the model to measure its own uncertainty, something current architectures don’t do reliably.
sequenceDiagram
participant U as User
participant B as Search engine
participant M as Language model
U->>B: "hes never coming over dario"
B->>B: retrieves passages about romantic relationships
B->>M: passes along ambiguous context
M-->>U: empathetic response about a breakup
Note over U,M: the retrieved context never had anything to do with basketball
Recent improvements target specific parts of the problem, not the whole problem. A better reranker reduces the chance of pulling the wrong passage; reinforcement learning from human feedback (RLHF) pushes the model to sound less confident when the context is weak; forcing citations requires every claim to be tied to a specific passage. None of the three eliminates RAG system errors when retrieval fails from the start.
There’s no direct way to verify from the outside whether a specific AI Overviews answer comes from an LLM invention or a solid source, because Google doesn’t expose the model’s confidence score or the full list of retrieved passages. The only way to infer it is by checking whether the summary cites a source and whether that source, read directly, actually backs up the claim.
Your next step: the next time a generative summary gives you a specific fact (a figure, a date, a name), open the source it cites before repeating it or pasting it into your code.
Frequently Asked Questions
Why does Google AI Overviews sometimes respond with emotional support instead of a link?
Because the system retrieves passages based on semantic similarity, not the user’s actual intent. If the query uses an ambiguous proper name, retrieval can pull context about personal relationships, and the model writes based on that context without knowing it’s wrong.
Are the model’s made-up answers a programming error?
Not exactly. They’re a consequence of the design: the model optimizes for generating fluent, probable text, not for checking facts against a reliable database in real time.
How do I spot algorithmic confabulation in a search summary?
Scroll down to the links the summary cites and read the original page. If the summary doesn’t show any source for a specific claim, treat it as unverified.
Does a RAG system with citations like Perplexity or Bing Copilot fabricate less than a search engine with an automatic summary?
The risk drops because every claim is tied to a numbered source you can open, but it doesn’t disappear: if retrieval pulls the wrong passage, the citation will be wrong too.
Can I turn off Google’s generative summaries?
Yes, adding the parameter &udm=14 to the end of the search URL forces the classic list of links, without the AI block up top.
References
- Sancho Panza’s Thoughts: the original account of the “hes never coming over dario” case that inspired this article.
- Wikipedia: academic definition and classification of the hallucination phenomenon in AI systems.
- arXiv: “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” the 2020 paper that formalized the RAG architecture.
- Google Blog: Search: official Google announcements about integrating generative summaries into its search engine.
- Wikipedia: general explanation of retrieval-augmented generation and its variants.
📱 Do you like this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.
Featured image: Foto de Markus Spiske en Unsplash
Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.
Leave a comment
0 Comments