Most teams install the widget, wait two weeks, and then ask whether tickets went down. By then the only numbers available are post-launch noise and a memory of what “felt busy” before. That is not a baseline. A baseline is a recorded series, taken on purpose, while nothing interesting is happening.
If you want an honest deflection story later, the work starts before anyone chats with the bot. The good news: the work is boring and finite.
What you are trying to freeze in place
You need a picture of inbound load that would have continued if you had launched nothing. Raw ticket count is almost never enough on its own, because businesses grow, seasons move, and a product launch next door can swamp a support experiment.
Pick a normaliser that moves with demand for help: active customers, orders, sessions, or seats — whichever is closest to “people who might write in.” Then track tickets per that unit, weekly. Absolute volume can still be noted; the ratio is what you will compare.
How long is long enough
Several weeks at minimum. Weekly series are noisy; a single quiet week will flatter whatever comes next. Four weeks is a common floor for small teams. Longer is better if your volume is spiky or if you already know a promotional calendar is about to move the numbers.
- Exclude known anomalies (outage week, billing incident, one-off PR spike) or mark them explicitly.
- Keep the same helpdesk filters you will use after launch.
- Capture satisfaction on the same instrument you will keep — baseline quality matters as much as baseline volume.
- Note staffing changes: a new hire answering faster is not deflection.
What to write down before go-live
| Field | Why it matters |
|---|---|
| Date range and timezone | So “before” cannot quietly expand later |
| Normaliser definition | So growth is not mistaken for failure or success |
| Weekly tickets / normaliser | The series you will actually compare |
| Channel mix | So chat displacing email is visible |
| Satisfaction method + score | So volume drops can be checked for attrition |
| Known exclusions | So nobody re-litigates the outage week later |
Store it somewhere boring: a sheet, a note in the project channel, a screenshot of the helpdesk report with the filters visible. The format matters less than the fact that it existed before the widget went live.
What not to count as the before
A vendor pilot on a staging site is not a baseline for production tickets. Internal dogfooding is not a baseline. The week you soft-launched to 5% of traffic is already contaminated. If you missed the quiet period, say so — a partial baseline with a caveat is still more honest than inventing a pre-period from memory.
tickets / 1k sessions (index)
Illustrative shape only: a two-week snapshot can sit above or below the true run-rate; a longer window dampens that luck.
After you launch, leave it alone
The first fortnight after a widget appears is distorted by novelty: people click because it is new, agents change behaviour because they are watching, and you will be tempted to retune daily. Let the measurement window run. Then compare normalised volume after against the baseline you already froze, with satisfaction beside it.
That comparison is the honest deflection check. Conversation counts, answer rate and time-to-first-response in the bot can still be useful for product tuning. None of them substitutes for the baseline. The guide on measuring deflection spells out the arithmetic once you have both sides of the window.
The shortest version
- Choose a normaliser tied to demand.
- Record weekly tickets against it for several weeks with unchanged filters.
- Capture satisfaction the same way you will after launch.
- Launch, wait, then compare — do not rebuild the “before” from recollection.
If you skip this, you can still ship a useful assistant. You just cannot claim ticket displacement with a straight face. Better to know which claim you are allowed to make.



