Multilingual demos are easy to stage. The vendor switches the widget to Spanish, asks a prepared question, and a fluent answer appears. If you do not speak Spanish, the fluency itself becomes the evaluation — which is how you buy a translation costume over a retrieval system. You can do better without becoming a linguist.
The method is to lock the facts in a language you do control, then check whether other languages return the same facts, decline the same gaps, and hand off as quickly. Run it on every shortlisted vendor, including Matter Chat.
Separate three different claims
| Claim | What to verify | How it fails |
|---|---|---|
| UI languages | Widget chrome translates | Answers still English-only |
| Answer languages | Replies in the visitor’s language | Fluent tone, wrong or invented facts |
| Source languages | Index includes non-English pages | Answers from English only, silently |
Vendors often demonstrate the first and imply the third. Your evaluation should pin which one you are buying. The setup guide for our own product is multilingual setup; use it as a map of the moving parts, not as proof we pass your test.
Build a fact card before you open any widget
- Pick two facts from your site you can verify in English (or your strongest language): a number, a window, a region constraint.
- Write the canonical answer in that language in a notepad — the exact claim, not a vibe.
- Write one question your site does not answer at all (the gap question).
- Get machine translations of the three questions into the target language, then have a human fix anything obviously broken if you can — a colleague, contractor, or native speaker for ten minutes.
The checks you can score without fluency
You may not be able to judge style. You can still judge structure.
- Numbers and proper nouns: does the reply contain the same price, SLA hours, SKU, or city names as your fact card?
- Citation target: open the source. Does it resolve to the page that holds the claim — ideally the localised page if you have one?
- Gap behaviour: on the uncovered question, do you get a decline-shaped answer and a human path, or a long fluent paragraph with no verifiable numbers?
- Handoff: type the local equivalent of “I want to speak to someone” (use a known phrase). Count turns. One is still correct.
If citations point at English pages while the answer is in another language, that can be acceptable — if the claim is truly on that page and your customers can use it. What is not acceptable is a localised-sounding answer whose source does not contain the claim at all. Same citation bar as English.
English-only sources are a special case
Many teams index English help content and still want French or German replies. That can work when the system retrieves the English passage and answers in the visitor’s language without inventing. It fails when the model “helps” by filling in what the English page never stated.
Your evaluation should include at least one question whose answer is a precise exception buried in English — and one gap question. If the exception survives translation and the gap becomes a refusal, you are looking at something operable. If both become smooth paragraphs, you are looking at multilingual hallucination.
Use a bilingual spot-check, not a full review
If you have any access to a speaker of the target language, even briefly, do not ask them “does this sound good?”. Ask them to mark whether the reply asserts your fact card, asserts something else, or declines. That is a fifteen-minute pass, not a localisation project.
- Give them your canonical fact and the transcript only.
- Ask for a three-way label: match / mismatch / refuse.
- If mismatch, have them paste the wrong claim in English. That becomes vendor feedback.
What to ignore in the multilingual demo
- Perfect idioms in a scripted FAQ the vendor chose.
- Language counts on a marketing page (“40+ languages”) without retrieval behaviour attached.
- Auto-detect theatre that still answers from the wrong locale’s prices.
“Fluency in a language you do not speak is not evidence. Consistency with a fact you do know is.”
Hold Matter Chat to the same fact-card method. If our non-English reply invents on your gap question, or cites a page that does not carry the claim, treat it as a failed evaluation — not as a translation nit. Multilingual support is still support. The honesty bar does not lower when the alphabet changes.



