Ask a vendor how well their assistant performs and you will often hear an answer rate: the share of conversations in which the bot produced an answer. The figure is usually high. It is also usually the wrong number to optimise, because the easiest way to improve it is to answer questions the system should have refused.
None of that makes a vendor dishonest. The metric is simply aligned with demo incentives and against buyer risk. If you are shortlisting tools, treat answer rate as context at best — and never as the ranking column.
What answer rate actually counts
Answer rate measures how often the system chose to speak in the register of an answer. It does not measure correctness, usefulness, or whether a ticket was avoided. A fluent invention counts. A careful refusal usually does not. Already the arithmetic prefers boldness over grounding.
- Covered FAQ, correct answer → counts as answered.
- Uncovered question, clean decline → often counts as not answered.
- Uncovered question, invented policy → counts as answered.
- One-word visitor message, generic reply → may still count as answered.
Why vendors like it anyway
It is available early. It moves during a trial. It fits on a slide next to a big percentage. None of that is evil; it is just incomplete. Vanity engagement metrics survive in marketing decks for the same reason — a number is easy to produce without a before-and-after on ticket volume or a paired satisfaction read.
Two systems can post the same answer rate and be opposite products: one declines honest gaps, the other fills them. We sketched that trap in how to test an AI support tool, where the chart is deliberately illustrative: one number cannot tell the two cases apart.
What buyers should read instead
| Signal | What it tells you | How it can lie |
|---|---|---|
| Outside-content behaviour | Whether invention is possible | Hard to fake if you control the question |
| Paraphrase consistency | Retrieval quality | Coaching mid-test invalidates it |
| Citation claim-check | Whether sources are real accountability | Topic-only links look fine until you open them |
| Refusal topics + frequency | Where content is thin | Only useful if someone reads the list |
| Normalised tickets + satisfaction | Whether automation helped | Needs a real baseline measured before launch |
For the measurement detail, pair this with our stance on deflection rate and the guide to measuring deflection. Answer rate and deflection percentage fail for related reasons: both can rise while customer outcomes get worse.
How to interrogate an answer-rate claim
- Ask for the definition: answered how? Does a refusal count? Does a one-line deflection count?
- Ask what share of “answers” carried citations that contain the claim — not merely related links.
- Ask for the refusal list grouped by topic. No list is a smell.
- Ask whether answer rate is a goal for their customers’ teams, or only a launch dashboard tile.
Then ignore the percentage long enough to run your own widget tests. A vendor who posts a modest answer rate and declines cleanly on your gap question is often safer than one advertising near-total coverage.
When answer rate is still useful
Internally, after launch, answer rate can be a continuity check: sudden drops may mean crawl failure, empty index, or a broken widget. Sudden spikes may mean a prompt or threshold change that made the system more willing to talk. Used as a change detector beside refusals and satisfaction, it has a job. Used as proof you should buy, it does not.
“A buyer metric gets harder to game as quality rises. Answer rate gets easier to game as carefulness falls.”
What we will and will not sell you on
Matter Chat can show you how often the assistant answered. We will not pretend that figure is the reason to choose us. Prefer the content gap view and the behaviours you can reproduce: decline on uncovered questions, consistent paraphrases, citations that prove the claim, handoff in one turn. Run those on us. If our answer rate looks impressive while those fail, believe the failures.



