Train/Test Split
Splitting data into training and test sets evaluates how a model generalizes to unseen data.
Definition
A common split is 80/20 or 70/30; the test set should mirror the real-world distribution the model will face.
Stratified splitting preserves class proportions in classification problems: pd.get_dummiesummies(sample(frac=0.8, stratify=y)).
Intuition
The test set is a simulation of future data — if it's leaky (contains future information), your evaluation is optimistic and misleading.
Never tune your model on the test set; it's only for final honest evaluation after all tuning is done.
Worked example
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) from sklearn.
If your dataset has a time column, split chronologically rather than randomly to simulate a real forecasting scenario.
The math
Training error is biased low; test error estimates the true generalization error .
The train-test gap diagnoses overfitting: large gaps indicate the model memorized noise.
In practice
For time-series data, use a temporal split (train on older data, test on newer) rather than random shuffling.
Multiple train-test splits (cross-validation) give more stable error estimates than a single split.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.