Most evaluation energy goes into whether the assistant answers correctly. Fair enough — wrong answers are expensive. Equally expensive, and more publicly hated, is an assistant that will not get out of the way. The human handoff test is how you catch that product decision before you deploy it.
It takes about a minute per vendor. Run it cold, without negotiating with the bot. Run it on us too.
The test
- Open the live widget (trial or public).
- Type exactly: “I want to speak to someone.” Optional second trial: “human please”.
- Count turns until a real escalation path appears — form, queue, email capture tied to handoff, or live agent.
- Complete it. Inspect what your team receives.
| Turns to handoff | Verdict | Why |
|---|---|---|
| One | Pass | The customer asked; the product complied |
| Two | Tolerable | One clarifying beat, then exit |
| Three or more | Fail | Someone optimised containment over consent |
| Never — loops or FAQ detours | Disqualify | This becomes your review-site narrative |
What must travel with the escalation
Speed without context just relocates the frustration. A handoff that arrives as a name and an email has thrown away the only valuable part of the bot conversation: what the visitor already asked and what the assistant already ruled out. That is the difference between escalation as a feature and escalation as a courtesy.
- Full transcript, or a faithful summary with the unanswered question intact.
- Page URL or product area if the widget knows it.
- Language the visitor used — especially if your team will reply in another.
- Clear queue or inbox destination, not a black hole address.
Our operational guide on handling escalations is the implementer’s version. As a buyer, you only need to see whether the vendor’s path preserves those fields when you trigger it yourself.
Handoff after failure vs handoff on request
There are two different moments, and vendors sometimes pass one while failing the other.
- On request: visitor asks for a human while the bot might still have answered. Consent wins. Handoff immediately.
- After refusal: assistant correctly says it does not know. Handoff should be easier here, not harder — the bot already admitted the gap.
Test both. A system that declines cleanly but then traps you in “try rephrasing” has only half-learned refusal. A system that hands off on request but strips context has learned obedience without usefulness.
Product decisions hiding inside “one more try”
When handoff takes many turns, it is rarely a model limitation. It is a containment preference: keep the conversation in-bot to protect an answer rate or a deflection story. That preference may be rational for a vendor metric and irrational for your brand. Your job in evaluation is to notice it while you still have leverage.
“If a customer has to win an argument with software to reach your team, the software is not support automation. It is a door with a personality.”
Edge cases worth one extra minute
- After hours: does handoff become a ticket with expectations set, or a dead “agents unavailable” loop?
- Authenticated vs anonymous visitors: does the path still work for someone who will not create an account?
- Angry tone: does frustration language delay handoff further? It should not.
- Wrong-department risk: can the visitor label billing vs technical, or does everything dump into one pile?
The bar we expect you to hold us to
On Matter Chat, ask for a person and expect a path in one turn, with conversation context available to the human side. If a trial on your property does anything else, that is a fail for us the same as for anyone on your shortlist. Pair this test with the outside-content question from how to test an AI support tool: inventing answers and blocking exits are the two failure modes customers remember by name.



