Small networks trained on small datasets memorize fast. Regularization trades a little training accuracy for better behavior on unseen inputs. Dropout is the most famous technique, but for tiny MLPs the full toolkit — weight decay, early stopping, and noise — matters more than any single trick.
How dropout works
During training, dropout randomly zeroes each hidden activation with probability p (typically 0.1 to 0.5) and scales survivors by 1 / (1 - p) to preserve expected magnitudes. Each update trains a different random sub-network; at inference, dropout turns off and the full network predicts.
function dropoutForward(activations, p, training) {
if (!training || p <= 0) return activations;
const scale = 1 / (1 - p);
return activations.map(a => (Math.random() < p ? 0 : a * scale));
}
The effect: no neuron can rely on any other neuron always being present, so representations spread across the layer instead of concentrating in a few units.
When dropout hurts tiny networks
Dropout was designed for large networks with redundancy to spare. A 128-unit hidden layer tolerates p=0.3; a 16-unit layer may not — dropping a third of 16 units destroys too much signal per update. For tiny MLPs (tens of units), use p=0.1 or skip dropout in favor of stronger alternatives.
Weight decay and early stopping
- Weight decay (L2): add
lambda * sum(w^2)to the loss, penalizing large weights. Start near 1e-4 and tune on validation. It costs one line and helps almost every small network. - Early stopping: keep the checkpoint with the best validation score and stop when validation stalls for N epochs. It is free regularization that also saves training time.
Both techniques suit CPU training budgets perfectly: no extra compute, no architecture changes, and early stopping shortens runs.
Noise as regularization
Small input noise — jittering normalized features by 1% during training — smooths decision boundaries similarly to dropout but without touching the architecture. For policy networks trained with reward signals, action-space noise (epsilon-greedy exploration) doubles as both regularizer and data collector.
A regularization recipe
For a tiny MLP on hundreds of examples: standardize inputs, add weight decay at 1e-4, train with a decaying learning rate, and early-stop on a grouped validation split. Add light dropout (p=0.1) only if validation still trails training by a wide margin. Measure every addition with cross-validation — regularization you cannot measure is superstition.