Regularization
Regularization adds a penalty to the loss function to discourage overly complex models, reducing overfitting.
Definition
L2 regularization (Ridge): adds to the loss, shrinking all weights toward zero proportionally.
L1 regularization (Lasso): adds , which drives some weights exactly to zero (feature selection).
Intuition
Regularization trades a small increase in training error for a larger decrease in test error by preventing the model from fitting noise.
L1 sparse solutions are interpretable — only the most important features survive. L2 solutions are dense but more stable when features are correlated.
Worked example
Ridge regression: . As , all (predictions approach the mean).
Lasso on a dataset with 1000 features but only 50 relevant: sets 950 coefficients to exactly 0.
The math
Elastic Net combines L1 and L2: , with controlling the mix.
The regularization path traces how each coefficient changes as varies from 0 (OLS) to (all zeros).
In practice
Always standardize features before regularization — regularization penalizes the magnitude of coefficients, which should be on comparable scales.
Use cross-validation to choose : too small lets overfitting through; too large underfits by suppressing all signal.
Go deeper
- InteractiveGradient Descent Lab
- InteractiveDecision Tree Explorer
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.