With thousands of examples you can spare a fixed test split. With a few hundred — typical for chatbot routing labels, intent examples, or fragment quality ratings — a single split wastes data and yields noisy estimates. Cross-validation spends every example on both training and testing.
How k-fold validation works
Split the data into k equal folds. Train on k-1 folds, test on the remaining one, and repeat until each fold has served as the test set. Average the k scores. With k=5, every example trains four models and tests one, and the averaged score varies far less than any single split.
function kFoldIndices(n, k) {
const folds = Array.from({ length: k }, () => []);
for (let i = 0; i < n; i++) folds[i % k].push(i);
return folds.map((test, f) => ({
test,
train: folds.flatMap((fold, g) => (g === f ? [] : fold)),
}));
}
Shuffle before assigning folds so that ordering artifacts (all of one class first) do not pile into a single fold.
Stratification matters
With imbalanced classes, plain random folds can starve a fold of a rare class. Stratified folds preserve each class's overall proportion in every fold. Implement it by dealing each class's indices round-robin across folds instead of dealing the whole dataset at once. For small datasets this one detail often changes results more than the choice of k.
Grouped folds prevent leakage
When examples come in groups — multiple queries per topic, multiple messages per session — random folds leak: near-duplicates land on both sides of the split and inflate scores. Assign whole groups to folds, never splitting a group. This is the grouped validation discipline, and cross-validation without it overstates real accuracy.
Choosing k
- k=5 is the default: each model trains on 80% of the data, and five training runs stay cheap for small models.
- k=10 reduces bias slightly (90% training data) at double the compute.
- Leave-one-out (k=n) is unbiased but expensive and high-variance; reserve it for tiny datasets under a hundred examples.
Report the mean and standard deviation across folds. A mean of 91% with a standard deviation of 6% tells a different story than the same mean with 1% — the first model is gambling on the split.
Validation versus test
Cross-validation tunes decisions: which features, which threshold, which k. Once those decisions are made, evaluate once on a held-out test set that never influenced any choice. Tuning on the test set — even informally, by stopping when test scores look good — leaks information and inflates the final number. Keep a frozen offline set for that final check.