The first thirty days after a widget launch are the most misread month in support automation. Traffic spikes from the announcement. The team watches answer rate hourly. Marketing asks for a win story. Finance asks for savings. The crawl is still finding pages you forgot existed. Treating day twelve like day three hundred is how you publish a heroic deck and reopen it quietly at day ninety.
Month one is for baselines and instrumentation — not for proving ROI. You are learning what visitors actually ask, how often the bot should refuse, and whether escalation lands with context. The goal is an honest picture, not a peak.
Tell stakeholders explicitly: week four is orientation, not verdict. That single sentence prevents a premature case study that has to be walked back when seasonality normalises.
- Week 1 owner: engineering — crawl scope and domain locks.
- Week 2–3 owner: content — refusal themes and gap ranking.
- Week 4 owner: support lead — sample audit and escalation review.
Before day one: lock what you can
If you skipped baselining ticket volume, start there anyway, even imperfectly. Note four weeks of ticket counts by channel, seasonality you already know, and any campaign scheduled this month. You will need it to interpret anything that looks like deflection later.
- Turn on conversation logging and refusal capture if your product supports it.
- Define escalation the same way you will in month six.
- Pick a weekly manual sample size you can actually sustain — twenty threads is fine.
- Agree internally: no ROI slide until day forty-five unless labelled preliminary.
Week one: crawl and configuration noise
Expect odd answers early. Staging domains left indexed, PDFs missing, persona too chatty — week one catches configuration mistakes, not model quality. Priorities:
- Verify indexed URL list against what should be public.
- Run the launch checklist questions on production, not staging.
- Log every wrong answer that cites a real page — chunking or retrieval issue.
- Log every invented answer — grounding or persona issue.
Weeks two and three: theme the refusals
By week two, refusals stabilise into themes. That list is the product roadmap. Rank it as in what a refusal list is for: volume, severity, writeability.
Start the weekly sample audit: mark each pulled thread correct, wrong, honest refusal, bad escalation. Track those four counts, not just containment. If wrong answers cluster on one page, fix the page or the chunk before you write new FAQs elsewhere.
Compare widget question themes to ticket themes. Divergence early is normal; visitors experiment. Convergence on refusals means the bot is surfacing real gaps; convergence on wrong answers means crawl or grounding debt.
Week four: what you can say out loud
| Claim | Safe at day 30? | Why |
|---|---|---|
| Top question themes | Yes | Grounds content plan |
| Refusal backlog ranked | Yes | Operational output |
| Sampled correctness trend | If sample honest | Still wide confidence interval |
| Ticket volume down | Only with baseline | Campaigns confound |
| Dollar savings | Rarely | Counterfactual too noisy |
| Beat vendor demo metrics | No | Different denominator |
The four charts worth keeping
- Conversations started per day — usage, not success.
- Refusal themes by weekly count — content pipeline.
- Sampled correctness and wrong-answer rate — quality.
- Escalations with median turns to human — handoff health.
“Day thirty is when you know what to build next — not when you know how much you saved.”
Schedule the day forty-five review now: ticket volume vs baseline, satisfaction split by bot-only vs human, first monthly automation report draft. Month one builds the instruments. Month two is when numbers start to mean what the deck promises.
Archive week-one wrong answers with citations. They are regression fixtures when someone proposes “just loosen refusal a bit” in month two.



