When a bot invents a return window, the post-mortem usually starts with the model. That is the wrong default. A hallucination is fluent, confident, wrong text — and the most common reason it appears in support is that nothing trustworthy was available to say instead.
Models complete patterns. If your site never states the rule, or states it three ways, or buries it on a PDF the crawler never saw, “helpful” generation fills the gap. Swapping vendors without fixing the gap just relocates the invention.
Separate three different failures
- Missing content — the fact was never published in a retrievable place.
- Retrieval miss — the fact exists, but the wrong passage (or none) was supplied.
- Ungrounded generation — material was thin or absent, and the system answered anyway.
Customers experience all three as “the bot made something up.” Only the third is hallucination in the strict sense. The first is a content gap. The second is a search problem. Collapsing them into one complaint makes every evaluation noisy and every RFP response sound the same.
How content creates “model” errors
Ambiguity is enough. “Returns within 30 days” on one page and “exchanges only on sale items” on another, with no link between them, invites a blended answer that matches neither policy. Duplicate stubs across locales do the same. So do marketing pages that sound like policy and help pages that sound like suggestions.
So do unpublished truths: the rule everyone on the team knows, enforced in the warehouse, never written down. Humans paper over that gap in chat. An assistant cannot — unless you allow it to invent, which many demos quietly do by rewarding answer rate.
An evaluation that blames the right layer
Take twenty real tickets from last month. For each, ask: is the correct answer on a public page today? If no, mark a content gap — inventing here is both a product-behaviour failure and a publishing failure. If yes, ask the live assistant. If it declines, your refusal path works and you have a retrieval or ranking bug to fix. If it answers, open the citation and check the claim.
| Observation | Call it | Next move |
|---|---|---|
| No page states the fact | Content gap | Write the page; keep refusing until live |
| Page exists; bot declines | Retrieval / ranking | Fix index, titles, chunks |
| Page exists; bot invents anyway | Missing grounding / refusal | Fail the vendor test |
| Right page; incomplete clause | Chunking / structure | Rewrite the section |
Why answer rate makes this worse
Answer rate rises when the system guesses. Teams under pressure to “automate more” loosen refusal, and the dashboard improves while trust decays. The useful companion metric is the refusal list ranked by frequency — each row is either a page to write or a retrieve to fix. That is the loop in find content gaps.
What “good” looks like on a gap question
A good assistant does not perform omniscience. It says it does not have that information, points at the closest real page if one exists without overstating it, and offers a human. That behaviour is refusal working as designed — and it is how you tell a grounded product from a demo that never leaves the happy path.
If you are mid-evaluation, run the outside-content question from how to test an AI support tool on every shortlist candidate. The ones that invent will also be the ones that look strongest in a scripted demo. That is not a coincidence.
The shortest version
- Most “hallucinations” start as missing, conflicting, or unindexed content.
- Split gaps, retrieval misses, and true ungrounded generation before you judge a model.
- Citations decide the fork; answer rate hides it.
- Refuse until the page exists — then re-test the same question.



