Every AI support tool demos well. That is not a conspiracy. A demo is assembled from questions the knowledge base already covers, asked in language close to the page titles, by someone who knows which answers look impressive. Under those conditions almost any retrieval system looks competent. The prediction problem is that those conditions are almost never the ones that decide whether the product helps or hurts once it is live.
So treat the demo as a product tour, not an evaluation. It can show you the interface, the install path, and whether the vendor has thought about escalation. It cannot tell you how the system behaves when a visitor asks something your site does not say, or asks the same fact three ways, or wants a person. Those are the behaviours that matter, and they are the ones a scripted call is designed to avoid.
What a demo is optimised for
A good sales engineer is not trying to trick you. They are trying to show the product under conditions where it works. That means curated content, known questions, and a path that ends in a citation that resolves. You leave with a feeling of confidence that is real — and poorly calibrated to the week after launch, when the first genuinely awkward question arrives.
- Questions are chosen because the pages exist, not because they are representative of your inbox.
- Phrasing tends to match the vocabulary on the page, which flatters keyword-shaped retrieval.
- Edge cases (policy gaps, clinical nuance, billing conditions) are deferred to “we can tune that later”.
- The human handoff, if shown at all, is usually the happy path with a prepared transcript.
Where demos systematically miss
The expensive failures in this category are fluent and wrong, not blank and honest. Demos almost never surface them, because inventing an answer requires a question the content does not cover — and nobody puts that on the slide deck. The same goes for paraphrase inconsistency: if you only ask in the page’s own words, a keyword system and a retrieval system look identical.
| Behaviour | Usually in the demo? | Why it matters live |
|---|---|---|
| Answers a covered FAQ | Yes | Baseline competence, not differentiation |
| Declines outside the content | Rarely | This is where hallucination shows up |
| Same fact, three phrasings | Almost never | Separates search-with-chat from real retrieval |
| Citation contains the claim | Sometimes | Decorative sources lend false confidence |
| “I want a person” in one turn | Optional | The public complaint vector if it fails |
Run the evaluation off the call
The antidote is not a longer demo. It is a short, adversarial session on a live widget: theirs, and yours once you have a trial, using questions you already know the answers to. We wrote the four-question version as how to test an AI support tool; the point here is why that has to happen without the salesperson in the room.
- Use your content, or insist on a trial seeded with a slice of it. A vendor’s showcase site only proves they can answer their own pages.
- Write the questions before you open the widget. Improvising on the call pulls you back into the happy path.
- Record the exact answer and the citation URL. Memory softens “fluent and wrong” into “pretty good”.
- Include one request for a human. Count the turns. One is correct.
If a vendor will not put a working assistant on a public page or a time-boxed trial with your material, that is already a result. Evaluation that only happens inside a controlled walkthrough is not evaluation.
What the demo is still good for
None of this means you should skip the call. It means you should ask different questions on it. Use the time for things a solo widget session cannot show: how crawl and recrawl work, what happens when a page changes, how escalations land in your inbox or helpdesk, where spend caps sit, and who owns the knowledge when something goes wrong.
- Install and domain constraints — including what breaks outside localhost. See domain locks.
- Escalation payload: does the human get the transcript, or only a name and email?
- Content ops: who updates what when the refusal list grows? The content gap loop is the product over time.
- Failure modes the vendor will admit to. A team that only has success stories has not operated the tool in anger.
A buying filter that survives the pitch
Rank vendors on behaviours you can reproduce, not on the feeling left by the deck. Prefer the system that declines cleanly on your gap question over the one that answered every prepared FAQ with more polish. Prefer the citation you can open and verify over the one that points at a homepage. Prefer the handoff that takes one turn over the one that “tries to help a bit more first”.
“A demo predicts that the product can look good. It does not predict that it will stay honest when the question is slightly wrong.”
Run the same off-call tests on Matter Chat. If we invent an answer on your content, or bury the handoff, or cite a page that does not contain the claim, you should walk away — from us as readily as from anyone else. The demo is a bad predictor for every vendor in the category. The twenty-minute edge-case session is not.



