Data science

Decision Trees

A decision tree splits data recursively on feature values, creating a flowchart-like model that's easy to interpret.

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

At each node, the tree chooses the feature and threshold that best separates the target classes (using Gini impurity or entropy).

Leaves hold the final prediction — a class label for classification or a real value for regression.

Intuition

Trees naturally model non-linear relationships and interactions without explicit feature engineering.

A shallow tree is interpretable: you can trace any prediction back through a few "if-then" rules.

Worked example

A tree for loan approval might split first on "income > 50k", then "debt < 20k", then "employment length".

A depth-3 tree can capture "X > 5 AND Y < 3" as a path of three splits.

The math

Gini impurity: G=1−∑kpk2G = 1 - \sum_k p_k^2; entropy: H=−∑kpklog⁡2pkH = -\sum_k p_k \log_2 p_k; information gain = Hparent−∑childrennchildnparentHchildH_{parent} - \sum_{children} \frac{n_{child}}{n_{parent}} H_{child}.

Pruning (pre- or post-) reduces overfitting by cutting branches that add more complexity than predictive power.

In practice

Good for exploratory analysis and as building blocks (random forests, gradient boosting) that combine many trees for better accuracy.

Can handle mixed numeric and categorical features and are robust to outliers after splitting.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.