The default enterprise buying motion for software is a feature matrix: forty rows, five vendors, traffic lights. For AI support tools that process is unusually bad. The rows that are easy to fill — “citations”, “multilingual”, “integrations”, “analytics” — are almost always yes for everyone you would seriously consider. The rows that decide whether the product helps your customers are behaviours, and behaviours do not fit in a checkbox.
You can still shortlist rigorously. You just shortlist on tests, not on claims. The output is a small table of outcomes you observed, not a grid of marketing nouns.
Why matrices mislead in this category
- Booleans hide quality. “Has citations” does not mean the citation contains the claim.
- Vendors optimise for the matrix. Anything that can be renamed into a row will be.
- The costly failures — fluent invention, buried handoffs — rarely appear as features at all.
- Your content and inbox shape outcomes more than a generic capability list does.
The four behaviours that rank a shortlist
These are the same tests we recommend in how to test an AI support tool, framed as a buying filter. Run them on a live widget with your content when you can. Score pass / partial / fail — not stars.
| Behaviour | Pass | Fail |
|---|---|---|
| Outside-content question | Declines, offers a person | Fluent invented answer |
| Paraphrase triad | Same fact three ways | Hits once, misses twice |
| Citation check | Link contains the claim | Homepage or topic-only link |
| Ask for a human | Handoff in one turn, with context | Loops, or drops the transcript |
Anything that fails the first or the fourth is usually out, regardless of how pretty the dashboard is. Invention and trapped customers are not polish issues.
Add two constraints of your own
Behaviour tests are universal. Your constraints are not. Pick two that actually bind, not ten that sound strategic, and apply them after the scorecard.
- Ops reality: who will update content when the refusal list grows? If the answer is “nobody”, prefer tools that make gaps obvious in analytics.
- Channel reality: website widget only, or Slack / helpdesk too? Do not buy a platform for a surface you will not staff.
- Risk reality: clinics and regulated answers need stricter refusal than a brochure site. Weight the outside-content test accordingly.
- Build reality: if you were about to build, read build vs buy before you shortlist vendors as a formality.
How many vendors to keep
Three is enough for a behaviour bake-off. Five is the point where people rebuild the matrix to feel in control. Start from a long list of six or seven names gathered from peers and search, eliminate anyone without a testable assistant, then run the four behaviours on the remainder in one afternoon.
Invite each finalist to a call only after the widget tests. Use the call for crawl ops, escalation routing, spend caps, and failure stories — not for another scripted FAQ tour. The demo is a bad predictor; do not let it re-enter as the decider.
What to ignore on purpose
- Answer rate as a primary ranking signal — see why it is a vendor metric.
- Deflection percentages without a before/after ticket baseline.
- Persona demos and mascot energy. Tone is tunable; grounding is not a coat of paint.
- Competitor pricing tables copied from blogs. They age in weeks and we will not pretend otherwise.
“If a row cannot be falsified in a live conversation, it should not decide the purchase.”
A one-page shortlist template
- Vendor name and URL of the live widget you tested.
- Pass/fail on the four behaviours, with one quoted answer each.
- Handoff destination and whether context survived.
- Your two constraints: met or not.
- Open question you still need answered on a call — one sentence max.
That page is enough to brief a founder or a finance partner without importing a spreadsheet religion. Include Matter Chat on the same page with the same rules. If we lose on refusal or handoff against your content, that is the shortlist working — not a relationship problem.



