Data science

Feature Scaling

Feature scaling standardizes the range or distribution of features so that no single feature dominates distance calculations.

Ask the Data science assistant1 min read · Updated September 9, 2026

Definition

Standardization (z-score): x′=(x−μ)/σx' = (x - \mu) / \sigma, giving features mean 0 and std 1.

Min-max scaling: x′=(x−xmin⁡)/(xmax⁡−xmin⁡)x' = (x - x_{\min}) / (x_{\max} - x_{\min}), mapping to [0, 1].

Intuition

Distance-based models (k-NN, SVM, k-means, neural nets) are sensitive to feature scale — a feature with large values dominates Euclidean distance.

Linear models, tree-based models, and probability-based models are generally invariant to scaling — but it doesn't hurt them.

Worked example

Height (cm) and weight (kg): height ~[150, 200], weight ~[40, 120] — weight has higher variance but not necessarily more predictive power.

Image pixel values [0, 255] are often standardized to roughly [-1, 1] for neural network training.

The math

RobustScaler uses median and IQR instead of mean and std — appropriate when data has outliers.

For highly skewed data, log transform or Box-Cox transform can make distributions more symmetric before scaling.

In practice

Always fit scaling on the training set only, then apply the same transformation to test/production data.

In a cross-validation pipeline, scaling must be done inside each fold to avoid leakage.

Go deeper

More in Data science

Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.