Outlier Detection
Outliers are data points that deviate significantly from the overall pattern; they may indicate errors, rare events, or genuine extreme behavior.
Definition
Detection methods: Z-score ( flags outliers), IQR rule ( or ), and DBSCAN (points labeled as noise).
Context matters: an outlier in one domain may be normal in another (e.g., a $10M house price is extreme but real).
Intuition
Not all outliers are errors — some represent the most interesting cases (fraud, rare diseases, high-value customers).
Outliers disproportionately affect mean-based statistics and models assuming normality — median and robust methods are less affected.
Worked example
Age = 250 in a customer dataset is likely a data entry error (replace or drop).
A single transaction of $1M might be an outlier but also a genuine signal worth investigating.
The math
Winsorizing caps extreme values at a percentile (e.g., all values above 99th percentile set to the 99th percentile value) instead of dropping them.
Robust statistics: median absolute deviation (MAD) replaces std for location estimation: .
In practice
Always investigate outliers before deciding what to do with them — understanding their cause prevents accidentally discarding signal.
Some models (decision trees, random forests) are robust to outliers; linear regression and logistic regression are not.
Go deeper
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.