Buyers often treat “I don’t know” as a bug. Vendors sometimes agree, and quietly tune systems to answer more often so the answer rate looks healthier. That is backwards for support. The expensive failure is not silence. It is a confident paragraph about a policy, price, or clinical detail that was never on the site.
A correct refusal is a product behaviour: the assistant recognises that the content does not support an answer, says so in plain language, and offers a path to a person. If you are evaluating tools, you should be trying to elicit that behaviour on purpose — and failing vendors that will not produce it.
Where refusal is the win condition
- The question is plausible for your business but absent from published pages.
- The answer would commit you — pricing exceptions, medical guidance, legal assurances, custom SLAs.
- Retrieval returned weak or conflicting chunks and the model would have to guess to continue.
- The visitor asks for something you deliberately do not automate (refunds over a threshold, account deletion, safety issues).
In those cases the assistant that answers fluently is not “more helpful”. It is improvising on your letterhead. The glossary term is hallucination; the operational term is “we now have to clean this up in public”.
What a good refusal sounds like
Useful refusals are boring. They do not apologise in a loop, quiz the visitor for five turns, or change the subject to something the knowledge base does cover. They state the limit and move the person forward.
| Behaviour | Customer experience | Verdict |
|---|---|---|
| Plain gap + offer a human | Honest, fast path to help | Correct |
| “I’m not sure” + still guesses | Confidence without grounds | Fail |
| Answers a neighbour topic instead | Feels evasive | Fail |
| Demands more detail forever | Support theatre | Fail |
| Refuses and dead-ends with no handoff | Honesty without a path | Incomplete |
The refusal definition on our site is the product stance: declining is part of being grounded, not a defect to minimise. The companion guide on reducing hallucinations is the implementation-shaped version of the same idea.
How to test for it on a shortlist
- Write two questions your site genuinely does not answer. Keep them realistic.
- Ask them on each vendor’s live widget or trial, on your content if possible.
- Accept only: clear decline, or decline-plus-human. Reject fluent invention.
- Ask “I want to speak to someone” immediately after a refusal. The handoff should not get harder because the bot just failed.
If a vendor frames a high answer rate as the primary proof of quality, ask what share of answers were refusals last month and whether they read a content gap report. An assistant that never says “I don’t know” is either omniscient or insufficiently careful. In this category it is the second one.
The metric trap behind “always answer”
Answer rate improves when you answer more. Carefulness often lowers it. That means the number vendors like to show moves in the opposite direction from the behaviour buyers should want at the edges. Two systems can post the same rate while one declines honest gaps and the other fills them with prose.
“In support automation, “I don’t know” is how the product shows it knows the difference between retrieval and invention.”
Measure refusals as a first-class stream: volume, topics, and whether they convert into pages you then publish. That loop is how automation gets safer over time. Celebrating a refusal rate of zero is celebrating a missing control.
When “I don’t know” is not enough
Refusal without a next step is only half a product. The visitor still has a problem. A correct decline should carry an escalation path that preserves context (what they asked, what was ruled out) rather than a blank contact form that makes them retype the saga.
Also watch for over-refusal: the assistant declining questions your pages clearly answer. That is a content, chunking, or crawl problem, not a virtue. The evaluation is asymmetric on purpose: inventing on a gap is disqualifying; missing a covered FAQ is a fixable ops issue. Sort them differently.
What we optimise for
Matter Chat is built so that “I don’t know” is available as a real outcome, not a shame state. You should still verify it. Put a gap question to our widget on your material. If we invent, that is on us. If we decline and offer a person, that is the behaviour we would defend in your inbox — and the one we think you should demand from every other vendor too.



