An RFP response is a genre. It is written to survive a scoring matrix, not a furious customer. That is fine for commodity software. It is dangerous for AI support, where the failure mode is fluent, public, and hard to unwind. Reading these responses well means hunting for places where confidence substitutes for a testable behaviour.
We answer RFPs too. Mark the same red flags in ours. If a response cannot survive the questions below, it should not survive procurement.
Vanity metrics offered as proof
Treat big answer rate and deflection rate claims as marketing until the method is spelled out. If the response cannot say whether refusals count, whether tickets were baselined, and whether satisfaction travelled with the volume change, the number is decorative.
Ask for the buyer metrics instead: outside-content behaviour, paraphrase consistency, citation claim-checks, refusal topics, normalised ticket volume with satisfaction. If those do not appear, the vendor is still selling the demo.
Grounding claimed, citation unspecified
“RAG-powered” and “grounded in your knowledge base” are table stakes language now. The red flag is the absence of a falsifiable citation standard. Does every answer include a source? Only when confident? What happens when retrieval is weak — refuse, or answer anyway with a related link?
- No mention of refusal as a first-class outcome → assume answer-rate pressure wins.
- Citations described as “references” or “related articles” → topic match, not claim proof.
- No discussion of stale crawl / recrawl → your deleted policy page may still answer for weeks.
- Hallucination framed only as a model problem → they will not help you fix content gaps.
Handoff fog
Search the response for escalation. If it says “seamless handoff” without turn count, context payload, and after-hours behaviour, mark it. Seamless is not a protocol. “I want to speak to someone” in one turn, with transcript attached, is a protocol — and you can test it before signature.
| What they wrote | What you ask next |
|---|---|
| “Escalates when needed” | Who decides ‘needed’ — visitor or model? |
| “Integrates with our helpdesk” | Which fields arrive? Transcript yes/no? |
| “Reduces unnecessary escalations” | Show the prompt or rule that delays handoff |
| “Warm transfer” | Demo it on your queue, not a slide |
Security theatre without product consequences
SOC paperwork matters. It does not replace product questions: domain locks, spend caps, prompt-injection posture on a public widget, and what staff can export. A response that leads with certifications and never mentions abuse cases on an embeddable chat is incomplete. Point them at concrete surfaces. For us that includes domain locks and security and spend caps; ask for their equivalents you can configure, not only audit.
Implementation plans that end at go-live
Red flag: a timeline that climaxes at launch week with no owner for the refusal list, no recrawl cadence, and no first-30-days analytics plan. Support assistants decay when pages change. If the RFP response treats content as a one-time ingest, you are buying a temporary FAQ with a chat bubble.
- Who reviews content gaps weekly?
- What happens when staging URLs enter the index?
- How do you re-evaluate after a major docs rewrite?
- Which metric gets a kill switch if satisfaction drops with tickets?
What a strong response looks like instead
It invites tests. It defines refusal. It admits failure modes. It separates answer rate from outcomes. It tells you how to falsify citations. It describes handoff with fields and timing. It is slightly less exciting than the polished alternative — which is the point.
“In this category, the reassuring RFP response is the one that makes it easier for you to prove them wrong before you buy.”
Attach the four behavioural tests from how to test an AI support tool as a mandatory appendix in your RFP. Score the live results higher than the prose. If Matter Chat’s written response ever conflicts with what the widget does on your content, believe the widget — and make every other vendor live by the same rule.



