Data science

Train/Test Split

Splitting data into training and test sets evaluates how a model generalizes to unseen data.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

A common split is 80/20 or 70/30; the test set should mirror the real-world distribution the model will face.

Stratified splitting preserves class proportions in classification problems: pd.get_dummiesummies(sample(frac=0.8, stratify=y)).

Intuition

The test set is a simulation of future data — if it's leaky (contains future information), your evaluation is optimistic and misleading.

Never tune your model on the test set; it's only for final honest evaluation after all tuning is done.

Worked example

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) from sklearn.

If your dataset has a time column, split chronologically rather than randomly to simulate a real forecasting scenario.

The math

Training error R^train\hat{R}_{train} is biased low; test error R^test\hat{R}_{test} estimates the true generalization error R∗R^*.

The train-test gap R^test−R^train\hat{R}_{test} - \hat{R}_{train} diagnoses overfitting: large gaps indicate the model memorized noise.

In practice

For time-series data, use a temporal split (train on older data, test on newer) rather than random shuffling.

Multiple train-test splits (cross-validation) give more stable error estimates than a single split.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.