Data science

Outlier Detection

Outliers are data points that deviate significantly from the overall pattern; they may indicate errors, rare events, or genuine extreme behavior.

Ask the Data science assistant 1 min read · Updated September 9, 2026

Definition

Detection methods: Z-score ( ∣z∣>3|z| > 3 flags outliers), IQR rule ( <Q1−1.5⋅IQR< Q1 - 1.5\cdot IQR or >Q3+1.5⋅IQR> Q3 + 1.5\cdot IQR ), and DBSCAN (points labeled as noise).

Context matters: an outlier in one domain may be normal in another (e.g., a $10M house price is extreme but real).

Intuition

Not all outliers are errors — some represent the most interesting cases (fraud, rare diseases, high-value customers).

Outliers disproportionately affect mean-based statistics and models assuming normality — median and robust methods are less affected.

Worked example

Age = 250 in a customer dataset is likely a data entry error (replace or drop).

A single transaction of $1M might be an outlier but also a genuine signal worth investigating.

The math

Winsorizing caps extreme values at a percentile (e.g., all values above 99th percentile set to the 99th percentile value) instead of dropping them.

Robust statistics: median absolute deviation (MAD) replaces std for location estimation: MAD=median(∣xi−x~∣)\text{MAD} = \text{median}(|x_i - \tilde{x}|).

In practice

Always investigate outliers before deciding what to do with them — understanding their cause prevents accidentally discarding signal.

Some models (decision trees, random forests) are robust to outliers; linear regression and logistic regression are not.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.