Evaluating

How to test an AI support tool before you trust it

Four questions, about twenty minutes, and no demo call. Run them against every vendor on your shortlist — this one included.

The Matter Chat team

2 August 2026 · 4 min read

ShareXLinkedIn
A woman laughing at her laptop at a sunlit, plant-filled desk.

Every assistant in this category demos well. That is not a conspiracy, it is arithmetic: a demo is built from questions the content covers, and any retrieval system answers those. The differences only show up at the edges — when the question is slightly outside the material, or the answer matters enough that being confidently wrong is expensive.

So the useful evaluation is not a feature comparison. It is four questions, asked of a live widget, with a note of what came back. It takes about twenty minutes per vendor and it does not need a sales call. Run it on us too — if we come out badly on your content, we would rather you found out now.

1. Ask something your site does not cover

This is the one that matters most, and it is the one nobody demos. Pick a question that is plausible for your business but genuinely absent from your published pages. For a shop: “do you offer gift wrapping?” when you do not. For a clinic: something clinical. For a SaaS product: a feature on nobody’s roadmap.

There are only two acceptable answers. The assistant says it does not have that information and offers a person, or it says nothing useful and offers a person. What you are watching for is the third behaviour: a fluent, plausible, entirely invented answer, delivered in the same confident register as the true ones.

2. Ask the same thing three different ways

Take one fact from your site and ask for it in three phrasings, at least one of which shares no vocabulary with the page it lives on. If your shipping page is titled “Shipping”, ask “how much to send a bag to France?”. If your pricing page says “seats”, ask about “extra logins”.

A keyword system will find one of the three. A retrieval system should find all three and return the same fact each time. Inconsistency across paraphrases is the clearest signal that what you are looking at is search with a chat interface bolted on.

A widget conversation in which a visitor asks how much it would cost to send a bag to France, and the assistant answers with the tracked rate and a numbered citation.
The question says “send a bag to France”. The page it came from is titled “Shipping”. That gap is the whole test.

3. Follow the citation

If an answer carries a source, open it. Twice. You are checking two separate things and they fail independently: that the link resolves to a real page on your site, and that the page actually contains the claim the answer made.

A citation that points at your homepage, or at a page that vaguely relates to the topic, is decoration. It looks like accountability and provides none. The point of a source is that a reader who doubts the answer can settle it in one click — if they cannot, the citation is doing nothing except lending confidence to something unverified.

TestGoodBad
Outside the contentDeclines plainly, offers a personA fluent answer you cannot source
Three paraphrasesSame fact all three timesFinds it once, misses twice
Follow the citationResolves, and contains the claimPoints at the homepage
Ask for a humanHands over immediatelyTries two more times first
What each answer to the four questions tells you.

4. Ask for a human

Type “I want to speak to someone”. Count how many turns it takes to get there. One is correct. Two is tolerable. Anything more is a product decision someone made deliberately, and it is the decision that turns automation into the thing customers complain about publicly.

While you are there, check what the handover carries. An escalation that arrives as a name and an email address has thrown away the most useful part of the conversation — what the customer already asked, and what the assistant already ruled out.

Why the headline number does not help

Most vendors, us included, can show you an answer rate. It is the least informative number in the category, because it is trivially improved by being less careful. An assistant that answers everything scores perfectly and tells you nothing about whether the answers were right.

The same answer rate, two very different systems

% of conversations

A · answered92
A · declined8
B · answered92
B · declined8

Illustrative, not measured: both systems post the same answered and declined split, so neither number separates them. What separates them is what those declines contained.

Two systems with an identical answer rate can be completely different products. The one that matters is what happened in the 8% — whether those were genuine gaps in the content, honestly declined, or questions the system should have refused and answered anyway. Read the content gap report alongside the rate, or read neither.

The shortest version

  1. Ask something outside your content. Watch for an invented answer.
  2. Ask one fact three ways. Watch for inconsistency.
  3. Open a citation. Check it contains the claim.
  4. Ask for a human. Count the turns.

None of this requires a trial, a call, or a spreadsheet of features. It requires twenty minutes and a question you already know the answer to. If a vendor’s assistant is not on their own site for you to test, that is itself a result.

The Matter Chat team

Written from the support inbox out

ShareXLinkedIn

Keep reading

All posts

Answer honestly. Capture the rest.

Point Matter Chat at your site and see what it can — and can't — answer. It's honest about both.

Start free — chat in your site

No credit card. 2 minute setup.

Every answer cites the source it came from. When there isn't one, it says so — and hands the visitor to you.

Installs on the tools you already run.