Accuracy alone hides where a classifier fails. A confusion matrix shows the full picture: rows for true intents, columns for predicted intents, and counts in every cell. The diagonal holds correct predictions; every off-diagonal cell is a specific, fixable mistake.
Reading the matrix
Suppose your chatbot classifies queries into question, follow-up, greeting, and out-of-scope. A confusion matrix might show:
- question: 82 correct, 9 predicted as follow-up, 4 as out-of-scope
- follow-up: 45 correct, 12 predicted as question
- greeting: 48 correct, 2 predicted as question
- out-of-scope: 30 correct, 8 predicted as question
The pattern jumps out: the classifier over-predicts 'question'. That single insight directs your effort — tighten the question class boundary instead of tuning everything blindly.
Building one in code
Construction is a few lines over labeled test queries:
function confusionMatrix(items, labels) {
const idx = new Map(labels.map((l, i) => [l, i]));
const m = labels.map(() => labels.map(() => 0));
for (const { actual, predicted } of items) {
m[idx.get(actual)][idx.get(predicted)] += 1;
}
return m;
}
Normalize rows to percentages for readability when classes are imbalanced. Raw counts matter for impact ('12 errors' beats '3%'), percentages matter for comparison across classes.
From matrix to per-class metrics
Each class reads as a one-versus-rest problem. For class C:
- True positives: cell (C, C).
- False positives: column C excluding the diagonal.
- False negatives: row C excluding the diagonal.
From these come precision and recall per class. Report macro-averaged F1 (average of per-class F1) when classes matter equally, and weighted F1 when frequent classes matter more. For intent prototypes, per-class recall directly shows which prototypes need more examples.
Acting on what you see
Common patterns and their fixes:
- Symmetric confusion between two classes: their definitions overlap. Merge them or add distinguishing training examples.
- One column absorbs everything: the model defaults to that class under uncertainty. Raise its decision threshold.
- Sparse row, scattered errors: too few test examples. Collect more before concluding anything.
Rebuild the matrix after every fix. If the off-diagonal mass moves rather than shrinks, you are trading errors between classes — check whether the trade matches your priorities. Not all mistakes cost the same: misclassifying out-of-scope as question produces wrong answers, while the reverse produces only an unnecessary fallback.