Building

What to index first when your site is a mess

You do not need a clean knowledge architecture to launch. You need the ten pages that already answer the questions people ask.

The Matter Chat team

30 July 2026 · 4 min read

ShareXLinkedIn
A woman holding a fan of tabbed cards above a table spread with pale paper cards, choosing which to place first.

Most teams arrive at indexing with a mess: a marketing site, a half-migrated help centre, a Notion dump, three PDFs nobody owns, and a staging subdomain that still has last year’s prices. The instinct is to crawl everything so nothing is missed. That instinct is usually wrong.

An assistant that retrieves from everything will answer from whatever matches closest — including tag archives, draft pages, and the blog post that contradicts your returns policy. Coverage is not the same as usefulness. The first crawl should be a short, deliberate list of pages that already resolve the questions your inbox already knows.

Start from the inbox, not the sitemap

Open the last month of tickets, chat transcripts, or the folder where “quick replies” live. Rank the topics by volume. Shipping, refunds, pricing, account access, and “where is my order” usually occupy the top of the pile for consumer businesses; seats, billing, SSO, and export formats do the same for SaaS.

For each repeated question, name the one page that should answer it. If that page does not exist, write it before you crawl anything else — indexing a gap does not close it. The guide to crawling your first site assumes you already know which URLs matter; this post is how you decide that list when the site will not tell you.

A practical first ten

Ten pages is enough to launch a useful bot on most sites. Prefer canonical, maintained pages over blog explainers that happen to mention the same topic.

  • Returns / refunds / cancellation policy
  • Shipping / delivery / service area
  • Pricing (the page with the numbers, not the landing pitch)
  • Account / login / password reset
  • Billing / invoices / plan changes
  • Product FAQ that your team already trusts
  • Hours, locations, or contact rules if humans still matter
  • One “how it works” page for the thing people misunderstand most
  • Warranty, SLA, or guarantee — whichever you actually honour
  • Anything legal or safety-related you must never invent around

What to leave out on purpose

Pagination, tag archives, author pages, search-result URLs, and “related posts” hubs dilute retrieval. So do marketing campaigns that reuse product names without stating facts. So does staging — that gets its own post, because it is the poison that looks like progress.

Internal jargon pages aimed at your sales team are another trap. If a visitor would never land there, and the wording does not match how customers ask, the page will either never retrieve or retrieve for the wrong reason. Prefer the public help wording, even if it feels less precise to you.

When the answer lives in three places

Shipping mentioned on the product page, the FAQ, and a 2019 blog post is not three sources of truth — it is one fact fighting itself. Pick the page you will keep updating. Index that. Deprecate or rewrite the others so they do not win the retrieval lottery with an outdated sentence.

If you cannot pick a winner in five minutes, the content problem is upstream of the bot. Fix the contradiction in writing first; otherwise every confident answer is a coin flip between versions.

Check the citations before you celebrate coverage

After the first index, ask five questions you already know the answers to. Open every citation. You are not checking eloquence — you are checking that the retrieved page is the one you intended and that the claim is still on it.

A right answer from the wrong page is a temporary win. The next paraphrase will miss. Treat that as an indexing or structure problem, not a model problem. The longer walkthrough of structuring pages for retrieval is in writing docs your bot can answer from.

Expand only after the refusal list is boring

Once the first ten are solid, grow from refusals (real questions the assistant declined) rather than from the sitemap. Content gaps ranked by frequency tell you which eleventh page is worth the crawl. Everything else can wait.

  1. List the questions your inbox already answers on repeat.
  2. Index the pages that answer them, and only those.
  3. Remove noise and duplicates that fight the canonical pages.
  4. Verify citations in the playground.
  5. Add the next page only when refusals show demand for it.

A messy site does not require a messy index. It requires the discipline to start small enough that you can still tell which page is speaking.

The Matter Chat team

Written from the support inbox out

ShareXLinkedIn

Keep reading

All posts

Answer honestly. Capture the rest.

Point Matter Chat at your site and see what it can — and can't — answer. It's honest about both.

Start free — chat in your site

No credit card. 2 minute setup.

Every answer cites the source it came from. When there isn't one, it says so — and hands the visitor to you.

Installs on the tools you already run.