Feature Scaling
Feature scaling standardizes the range or distribution of features so that no single feature dominates distance calculations.
Definition
Standardization (z-score): , giving features mean 0 and std 1.
Min-max scaling: , mapping to [0, 1].
Intuition
Distance-based models (k-NN, SVM, k-means, neural nets) are sensitive to feature scale — a feature with large values dominates Euclidean distance.
Linear models, tree-based models, and probability-based models are generally invariant to scaling — but it doesn't hurt them.
Worked example
Height (cm) and weight (kg): height ~[150, 200], weight ~[40, 120] — weight has higher variance but not necessarily more predictive power.
Image pixel values [0, 255] are often standardized to roughly [-1, 1] for neural network training.
The math
RobustScaler uses median and IQR instead of mean and std — appropriate when data has outliers.
For highly skewed data, log transform or Box-Cox transform can make distributions more symmetric before scaling.
In practice
Always fit scaling on the training set only, then apply the same transformation to test/production data.
In a cross-validation pipeline, scaling must be done inside each fold to avoid leakage.
Go deeper
- InteractiveK-Means Clustering Playground
More in Data science
Assembled from the ReLU.chat curated knowledge base. These explanations are concise on purpose; check the sources for anything important.