Data science

Cross-Validation

Cross-validation evaluates a learning procedure on held-out data. In k-fold CV, each fold takes a turn as the validation set while the model trains on the remaining folds.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

kk-fold CV trains kk models on k−1k\frac{k-1}{k} of the data each, giving kk validation scores averaged into a robust estimate.

Leave-one-out CV uses k=nk = n folds and trains on n−1n-1 observations each time. Its test-error estimate often has low bias but can have high variance, and fitting nn models can be expensive.

Intuition

CV estimates how the model will perform on unseen data, accounting for the fact that a single train-test split might be unlucky.

Scores can vary across folds because of model instability or differences in the held-out observations. The folds share training data, so their scores are not independent.

Worked example

5-fold CV on 1000 samples: 5 models each trained on 800 samples and tested on 200, giving 5 accuracy scores averaged.

Stratified KFold preserves class proportions in each fold, important for imbalanced classification.

The math

Expected generalization error estimate: E^=1k∑i=1kL^i\hat{E} = \frac{1}{k}\sum_{i=1}^k \hat{\mathcal{L}}_i, where L^i\hat{\mathcal{L}}_i is the validation loss on fold ii.

Nested CV separates hyperparameter selection in the inner loop from evaluation in the outer loop, reducing the optimistic bias from selecting and evaluating on the same validation scores.

In practice

Use CV to compare models, tune hyperparameters (grid search over CC, kk, etc.), and select features without overfitting to the validation set.

For time series, use forward-chaining (expanding window) or blocked CV to avoid lookahead bias.

Sources

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.