Small networks trained on small datasets memorize fast. Regularization trades a little training accuracy for better behavior on unseen inputs. Dropout is the most famous technique, but for tiny MLPs the full toolkit — weight decay, early stopping, and noise — matters more than any single trick.

How dropout works

During training, dropout randomly zeroes each hidden activation with probability p (typically 0.1 to 0.5) and scales survivors by 1 / (1 - p) to preserve expected magnitudes. Each update trains a different random sub-network; at inference, dropout turns off and the full network predicts.

function dropoutForward(activations, p, training) {
  if (!training || p <= 0) return activations;
  const scale = 1 / (1 - p);
  return activations.map(a => (Math.random() < p ? 0 : a * scale));
}

The effect: no neuron can rely on any other neuron always being present, so representations spread across the layer instead of concentrating in a few units.

When dropout hurts tiny networks

Dropout was designed for large networks with redundancy to spare. A 128-unit hidden layer tolerates p=0.3; a 16-unit layer may not — dropping a third of 16 units destroys too much signal per update. For tiny MLPs (tens of units), use p=0.1 or skip dropout in favor of stronger alternatives.

Weight decay and early stopping

  • Weight decay (L2): add lambda * sum(w^2) to the loss, penalizing large weights. Start near 1e-4 and tune on validation. It costs one line and helps almost every small network.
  • Early stopping: keep the checkpoint with the best validation score and stop when validation stalls for N epochs. It is free regularization that also saves training time.

Both techniques suit CPU training budgets perfectly: no extra compute, no architecture changes, and early stopping shortens runs.

Noise as regularization

Small input noise — jittering normalized features by 1% during training — smooths decision boundaries similarly to dropout but without touching the architecture. For policy networks trained with reward signals, action-space noise (epsilon-greedy exploration) doubles as both regularizer and data collector.

A regularization recipe

For a tiny MLP on hundreds of examples: standardize inputs, add weight decay at 1e-4, train with a decaying learning rate, and early-stop on a grouped validation split. Add light dropout (p=0.1) only if validation still trails training by a wide margin. Measure every addition with cross-validation — regularization you cannot measure is superstition.