Every chatbot change — new fragments, retuned weights, a different fallback threshold — risks breaking queries that used to work. An offline test set catches those regressions in seconds: a fixed list of queries with known-good answers that you re-run before every release.
What goes into the test set
Collect 100 to 300 queries covering the chatbot's real workload:
- Common questions sampled from actual user logs, paraphrases included.
- Edge cases: typos, follow-ups, topic switches, greetings, and out-of-scope queries.
- Regression cases: every bug report becomes a permanent test so fixed bugs stay fixed.
Pair each query with the expected topic or entry id, not the full expected text. Topics are stable; wording evolves. For out-of-scope queries the expected result is the fallback action itself.
Grading answers automatically
For retrieval-based bots, grading is mechanical: run the query, check whether the expected entry appears in the top 1 (or top 3) results. Track top-1 accuracy, top-3 accuracy, and the fallback rate separately. A change that lifts top-1 accuracy but doubles false fallbacks is not a clear win.
function grade(results) {
let top1 = 0, top3 = 0;
for (const r of results) {
if (r.ranked[0] === r.expected) top1++;
if (r.ranked.slice(0, 3).includes(r.expected)) top3++;
}
return { top1: top1 / results.length, top3: top3 / results.length };
}
For response quality beyond ranking, add BLEU/ROUGE spot checks or human review on a rotating sample — automation covers ranking, humans cover wording.
Keeping the test set honest
A test set you train against stops measuring generalization. Protect it:
- Never add paraphrases from the test set into the knowledge base or query expansion lists.
- Split by topic like grouped validation: test topics stay disjoint from anything tuned during development.
- Freeze the set per release and version it, so accuracy numbers stay comparable over time.
If scores climb because the test leaked into training data, you have theater, not progress.
Running tests in CI
The full suite should run headless in under a minute. ReLU.chat's release checks run conversation regressions alongside unit tests on every push. Structure yours the same way: a single command that loads the knowledge base, grades every query, and fails the build when accuracy drops below the frozen baseline.
Growing the set over time
Review production logs weekly. Every misunderstood query is either a knowledge gap (add content) or a retrieval gap (fix ranking) — and either way, a new test case. A test set that grows with real failures compounds in value: after a year, it encodes hundreds of lessons no single engineer remembers.