Data science

Regularization

Regularization adds a penalty to the loss function to discourage overly complex models, reducing overfitting.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

L2 regularization (Ridge): adds λ∑jwj2\lambda\sum_j w_j^2 to the loss, shrinking all weights toward zero proportionally.

L1 regularization (Lasso): adds λ∑j∣wj∣\lambda\sum_j |w_j|, which drives some weights exactly to zero (feature selection).

Intuition

Regularization trades a small increase in training error for a larger decrease in test error by preventing the model from fitting noise.

L1 sparse solutions are interpretable — only the most important features survive. L2 solutions are dense but more stable when features are correlated.

Worked example

Ridge regression: min⁡w∑(y−Xw)2+λ∣∣w∣∣22\min_w \sum(y - Xw)^2 + \lambda||w||_2^2. As λ→∞\lambda \to \infty, all w→0w \to 0 (predictions approach the mean).

Lasso on a dataset with 1000 features but only 50 relevant: sets 950 coefficients to exactly 0.

The math

Elastic Net combines L1 and L2: λ(ρ∑∣wj∣+(1−ρ)∑wj2/2)\lambda(\rho\sum|w_j| + (1-\rho)\sum w_j^2/2), with ρ∈[0,1]\rho \in [0,1] controlling the mix.

The regularization path traces how each coefficient changes as λ\lambda varies from 0 (OLS) to ∞\infty (all zeros).

In practice

Always standardize features before regularization — regularization penalizes the magnitude of coefficients, which should be on comparable scales.

Use cross-validation to choose λ\lambda: too small lets overfitting through; too large underfits by suppressing all signal.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.