Notes on making support automation honest
How to evaluate a tool in this category, what the headline numbers actually measure, and where automation stops being useful. Written so you can run the tests against us as readily as against anyone else.

More in the blog
Evaluating6 minThe best AI chatbot for a Webflow site, compared honestly (2026)Matter Chat, Social Intents, Ultimo Bots, Chatling, WeblyChat and Chatbase on a Webflow site: Apps marketplace or custom code, the paid-site-plan gate, CMS coverage, citations, handoff and pricing. Checked on the vendors' own pages.BuyingComparisonRead
Evaluating6 minChatbase alternatives for teams that need cited answers (2026)Matter Chat, Chatbase, Intercom Fin, Tidio Lyro, DocsBot and Chatling compared on citations, refusal behaviour, handoff and what the entry price meters. Checked on the vendors' own pages, with the test to run on all six.BuyingComparisonRead
Measuring3 minDeflection rate is the most overstated number in support automationCounting bot conversations as deflected tickets overstates the result. The honest version is a before-and-after on ticket volume, read next to satisfaction.MetricsDeflectionRead
Evaluating4 minHow to test an AI support tool before you trust itEvery AI support tool demos well, because demos ask questions the content covers. Four questions that separate them, and what a good answer looks like.EvaluationGroundingRead
Building4 minWhat to index first when your site is a messStart with the pages that already resolve real support questions. Noise, archives, and unfinished docs can wait — indexing them first makes answers worse.KnowledgeIndexingRead
Evaluating5 minThe demo is a bad predictor of AI support qualityA polished demo tells you the product can answer questions it was prepared for. The useful evaluation happens at the edges — outside the script, on your content.EvaluationDemosRead
Building3 minStaging domains will poison your answers — fix that firstStaging, preview, and leftover demo hosts quietly win retrieval. Remove them from the index before you tune prompts or blame the model.KnowledgeIndexingRead
Measuring3 minResolution rate vs deflection: stop mixing the twoResolution and deflection answer different questions. Mixing them inflates the result and hides whether customers actually got what they needed.MetricsDeflectionRead
Evaluating4 minWhat a citation should prove (and what most of them don’t)Citations look like grounding. Many are just related links. Here is the two-check test — resolves, and contains the claim — and what fails it.CitationsGroundingRead
Building3 minWriting help pages an assistant can actually retrieveAnswer early, name headings like questions, keep one topic per page, and put facts in text. Retrieval fails on structure more often than on “AI”.KnowledgeWritingRead
Measuring3 minHow to baseline ticket volume before you launch a botRecord normalised ticket volume for several weeks before launch. Skip the baseline and you will never know whether the bot displaced work or only looked busy.MetricsDeflectionRead
Evaluating4 minParaphrase testing: the five-minute retrieval checkA five-minute paraphrase test separates keyword matching from real retrieval. Same fact, three phrasings, one expected answer — including on Matter Chat.EvaluationRetrievalRead
Building3 minDesigning refusals that still feel helpfulRefuse plainly, say what you can help with instead, and offer a person or a form. Helpful refusals protect trust without inventing answers.PersonaEscalationRead
Measuring3 minSatisfaction must travel with every automation numberNever report deflection, containment, or answer rate alone. Pair every automation figure with the same satisfaction instrument you used before launch.MetricsCSATRead
Evaluating4 minWhen “I don’t know” is the correct product behaviourAssistants that always answer look helpful in demos and dangerous in production. When refusal is the right behaviour — and how to evaluate it.RefusalGroundingRead
Building3 minEscalation that keeps the conversation contextHand over on refusal and on request, send the full transcript, and be honest about office hours. Context is the whole point of escalation.EscalationSupport opsRead
Measuring3 minWhat a refusal list is for (and how to rank it)What a refusal list is for: turning honest “I don’t know” moments into a ranked content backlog that drives support automation forward.MetricsRefusalsRead
Evaluating3 minHow to shortlist AI support tools without a feature matrixSkip the 40-row comparison spreadsheet. A behaviour-first shortlist — refusal, paraphrase, citation, handoff — ranks vendors on what customers feel.BuyingEvaluationRead
Building3 minPersona tuning without turning the bot into a mascotSet a specific greeting, match your brand’s register, and write hard boundaries. Skip the forced personality that makes refusals feel like a bit.PersonaBrandRead
Measuring3 minContainment rate, explained without the sales deckContainment rate without the spin: what it counts, how it differs from resolution and deflection, and when a high number is a warning.MetricsContainmentRead
Evaluating4 minBuild vs buy for website support (without the usual trap)Build vs buy for a website support assistant is really a question about who owns crawl, refusal, handoff, and the content loop — not who can wrap an API first.Build vs buyBuyingRead
Building2 minLaunch checklist: five questions before you go liveBefore launch: confirm sources, citations, refusals, human handoff, and cost controls. If any fail, the widget is not ready — the model is not the issue.LaunchEvaluationRead
Measuring3 minMeasuring escalation quality, not just escalation countEscalation volume alone rewards obstruction. Measure time-to-human, context completeness, and repeat-asking so handoffs prove automation helped.MetricsEscalationRead
Evaluating3 minWhy answer rate is a vendor metric, not a buyer metricAnswer rate is trivially improved by being less careful. Here is what the number actually measures — and what to read instead when you evaluate vendors.MetricsAnswer rateRead
Building4 minRe-crawling after a docs rewriteDocs rewrites feel finished when the pages go live. For a retrieval-based assistant, they are finished when the crawl catches up — and you have checked what the bot still cites.KnowledgeCrawlingRead
Measuring3 minThe first 30 days of analytics after a widget launchWhat to measure in the first 30 days after launching a support widget: baselines, sampling, refusal themes, and what to ignore until the crawl settles.MetricsLaunchRead
Evaluating3 minThe human handoff test every vendor should passHuman handoff is where automation either stays honest or becomes the thing customers complain about. A simple vendor test — including for Matter Chat.EscalationHandoffRead
Building4 minSpend caps as a product decision, not just billingSpend caps belong in product design: they define refusal behaviour, protect visitors from runaway loops, and force explicit choices about what support automation is worth.ProductBillingRead
Measuring3 minWhen falling tickets means attrition, not successFalling support tickets after launching a bot can mean deflection or attrition. Here is how to tell which — before you report savings.MetricsDeflectionRead
Evaluating4 minEvaluating multilingual AI support without a language you speakHow to evaluate multilingual AI support when you cannot read every reply: locked facts, paraphrase pairs, citation checks, and handoff — without trusting the demo language switch.MultilingualEvaluationRead
Building4 minWhen a contact form should stay (and the bot should defer)Chat widgets do not replace every contact path. Keep forms where attachment, identity, or sensitivity matter — and teach the bot to step aside without a fight.EscalationUXRead
Measuring3 minA monthly support-automation report a CFO will acceptHow to write a monthly support automation report finance will trust: baselines, sampled quality, refusal backlog, and honest limits on savings claims.MetricsReportingRead
Evaluating3 minRed flags in an AI support RFP responseRed flags in AI support RFP responses: vanity metrics, citation theatre, handoff fog, and demos that never leave the happy path — and what to demand instead.RFPBuyingRead
Building4 minMultilingual support when your source pages are English-onlyYou can serve multilingual visitors without translated help centres — if you are explicit about limits, avoid invented policy in translation, and test refusals in every language you enable.MultilingualKnowledgeRead
Measuring3 minContent gaps as a metric, not a side projectTreat content gaps as a core metric: define them, count refusals and wrong answers that reveal them, and report gap closure rate monthly.MetricsContent gapsRead
Building4 minDomain locks and why “works on localhost” isn’t enoughDomain allowlists exist so your bot answers only where you installed it. Testing on localhost proves the script loads; it does not prove production embeds are locked down.SecurityEmbedRead
Building3 minGrounding without the mystiqueStrip the jargon and grounding is simple — supply the passages, require the model to use them, refuse when they are missing. Here is how to check it is actually happening.GroundingRAGRead
Building3 minChunking mistakes that look like model failuresSplit a policy table from its header, bury the exception in a footer, or index a nav-heavy page as one blob — and the model will look broken while retrieval did the damage.ChunkingRAGRead
Evaluating3 minHallucination is usually a content problemBefore you swap vendors over “hallucinations,” audit whether the fact exists, retrieves, and is allowed to be refused. Most stacks fail that audit before they fail the model.HallucinationContent gapsRead
Evaluating3 minRAG vs fine-tuning for support: the decision that sticks“Train it on our docs” sounds right and usually isn’t. For website support, RAG wins on currency, citations, and cost; fine-tuning earns its keep for behaviour, not policies.RAGFine-tuningRead
Building4 minCitations customers will actually clickCitations earn clicks when they are specific, stable, and visibly tied to the claim. Generic homepage links and broken anchors teach visitors to ignore your sources.CitationsTrustRead
Building4 minKnowledge ops for a two-person teamA support bot on a tiny team survives on refusal lists, ten-minute recrawls, and docs edits tied to real tickets — not on a KM platform nobody maintains.KnowledgeOperationsRead
Building3 minPrompt injection on a public support widgetPublic widgets get jailbreak theatre and quieter injection attempts. Here is what actually matters for support: grounding, tool limits, domain locks, and not following instructions in the user message.SecurityPrompt injectionRead
Building4 minWhy your bot answers from the wrong pageWrong-page answers are usually indexing and structure problems dressed as intelligence. Fix duplicates, noise, and chunk boundaries before you touch the model.RetrievalIndexingRead
Building4 minEcommerce: returns questions that invent policyEcommerce bots fail returns questions by inventing windows, restocking fees, and exceptions. Ground them in one canonical policy or refuse until you have one.EcommerceReturnsRead
Building4 minSaaS: billing questions the pricing page doesn’t answerSaaS support bots fail on billing when only the pricing page is indexed. Close gaps on invoices, seats, trials, and downgrades — or refuse and route to finance.SaaSBillingRead
Building1 minClinics: when fluent answers are unacceptableFor clinics and health-adjacent sites, AI support must refuse outside published admin facts. Fluency without grounding is unacceptable.IndustryClinicsRead
Building1 minAgencies: one bot, many client knowledge basesHow agencies should run AI support across clients: separate knowledge bases, domain locks, and evaluation per property.IndustryAgenciesRead
Building1 minLocal services: hours, areas, and the lies maps createFor local service businesses, ground AI support on your published hours and service areas — not on whatever maps currently invent.IndustryLocal servicesRead
Want to go deeper? Pick a topic
The long-form material lives in the resources. These are the entries people reach for most often.
Answer honestly. Capture the rest.
Point Matter Chat at your site and see what it can — and can't — answer. It's honest about both.
Start free — chat in your site
No credit card. 2 minute setup.
Every answer cites the source it came from. When there isn't one, it says so — and hands the visitor to you.
Installs on the tools you already run.
