Ensemble Methods
Ensemble methods combine multiple models into one prediction, typically reducing variance (bagging) or bias (boosting).
Definition
Bagging (Bootstrap Aggregating): train many models on bootstrapped datasets, average predictions. Random forests extend this by also randomizing feature subsets at each split.
Boosting: train models sequentially; each new model focuses on the residuals (errors) of the ensemble so far. Popular variants: AdaBoost, Gradient Boosting, XGBoost, LightGBM.
Intuition
The "wisdom of the crowd" effect: individual models make different errors, and averaging cancels out uncorrelated mistakes.
Boosting reduces both bias and variance but is more prone to overfitting than bagging — careful tuning of the number of rounds is needed.
Worked example
A random forest with 500 trees is more stable and accurate than any single decision tree drawn from the same data.
XGBoost on a tabular dataset often beats deep learning for structured data with fewer than ~10,000 features.
The math
Random forest variance reduction: each tree sees unique data points per bootstrap; averaging trees reduces variance by a factor of .
Gradient boosting: fit to pseudo-residuals ; update .
In practice
Ensemble methods are the workhorses of Kaggle competitions and production ML for tabular data.
Stacking (a meta-learner on top of base model predictions) can further improve but adds complexity.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.