Data science

Bias-Variance Tradeoff

Bias-variance decomposition separates a model's error into bias (systematic error from wrong assumptions) and variance (sensitivity to training data).

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

Bias: error from overly simplistic assumptions — the model consistently gets it wrong in the same direction.

Variance: error from sensitivity to small fluctuations in training data — different training sets produce very different models.

Intuition

High bias = underfitting (model is too simple to capture the pattern). High variance = overfitting (model is too complex and memorizes noise).

The total error is Bias2+Variance+Irreducible Noise\text{Bias}^2 + \text{Variance} + \text{Irreducible Noise}. You can reduce one only at the cost of increasing the other.

Worked example

Linear regression on a truly linear relationship: low bias, low variance. On a non-linear relationship: high bias (systematically misses the curve).

A depth-20 decision tree on small data: low bias (fits complex patterns) but high variance (different trees on different data).

The math

E[(y−f^(x))2]=(E[f^(x)]−f(x))2+E[(f^(x)−E[f^(x)])2]+σ2\mathbb{E}[(y - \hat{f}(x))^2] = (\mathbb{E}[\hat{f}(x)] - f(x))^2 + \mathbb{E}[(\hat{f}(x) - \mathbb{E}[\hat{f}(x)])^2] + \sigma^2.

Bagging (bootstrap aggregating) reduces variance by averaging predictions from models trained on different bootstrapped samples.

In practice

Use bias-variance thinking to choose models: simple models (linear regression) for noisy/simple problems; complex models (trees, neural nets) for signal-rich problems.

Ensemble methods like random forests and gradient boosting trade bias for variance reduction to improve overall prediction.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.