Machine learning models learn faster and generalize better when input features share a similar scale. A feature ranging 0 to 1,000,000 beside one ranging 0 to 1 distorts gradients, distances, and regularization alike. Normalization fixes the imbalance in one preprocessing step.
Why scale matters
Consider gradient descent on two features: house size in square millimeters and room count. The loss surface stretches into a narrow valley — steep along one axis, flat along the other — forcing tiny learning rates and slow convergence. Distance-based methods suffer directly: in cosine or Euclidean space, the large feature dominates every comparison.
Standardization: zero mean, unit variance
Standardization rescales each feature using the training set's mean and standard deviation: z = (x - mean) / std. Resulting features center on zero with spread near one, which suits neural networks and linear models.
function standardize(values, mean, std) {
return values.map(v => (v - mean) / (std || 1));
}
Compute mean and std on the training split only, then apply the same numbers to validation and production inputs. Recomputing statistics on each new batch leaks information and shifts the model's operating point — the same leakage discipline as any train/test separation.
Min-max scaling: fixed ranges
Min-max scaling maps each feature into [0, 1] via (x - min) / (max - min). It suits algorithms that expect bounded inputs and features with meaningful endpoints, like pixel intensities. Its weakness is outliers: one extreme value compresses everything else into a sliver. Prefer standardization for heavy-tailed data.
Normalizing text features
Text pipelines normalize differently. TF-IDF vectors get L2-normalized so long documents do not outscore short ones by magnitude. Embedding vectors are normalized before cosine similarity so the dot product equals the cosine. When mixing text scores with dense features in one ranker, standardize each signal's historical range first or the widest-ranging signal wins by default.
Folding normalization into the model
At inference time, explicit normalization is an extra step — and an extra place for training/serving skew. For linear first layers, fold the scaling into the weights: divide each weight by the feature's std and adjust the bias by the mean. The exported model then accepts raw features directly. ReLU.chat's policy export uses exactly this trick, folding training feature scaling into the first layer so the browser runs raw inputs with zero preprocessing mismatch.