The learning rate is the step size of gradient descent: too large and training diverges, too small and it crawls. A schedule varies the rate over training — starting bold, finishing careful — and it is one of the cheapest ways to improve small-network results.
Why constant rates struggle
Early in training, weights sit far from good values and large steps make fast progress. Late in training, the same large steps bounce around the minimum instead of settling into it. A constant rate compromises between these phases and serves neither well. Schedules give each phase the step size it wants.
Schedules that work
- Step decay: halve the rate every N epochs. Simple, predictable, easy to reason about.
- Cosine annealing: decay along a cosine curve from the initial rate to near zero. Smooth and strong for fixed training budgets.
- Linear warmup then decay: ramp up over the first few epochs (stabilizing early gradients), then decay. Essential for transformers, helpful everywhere.
function cosineSchedule(epoch, totalEpochs, baseLR, minLR = 0) {
const t = Math.min(epoch / totalEpochs, 1);
return minLR + 0.5 * (baseLR - minLR) * (1 + Math.cos(Math.PI * t));
}
Picking the base rate
Find the base rate with a quick range test: train for a few epochs each at rates spanning 1e-4 to 1e-1 on a log scale, and pick the largest rate whose loss decreases smoothly. Small MLPs on normalized features typically land near 1e-3 with Adam or 1e-2 with SGD.
Scale the rate with batch size: doubling the batch roughly justifies doubling the rate (the linear scaling rule), since averaged gradients are less noisy. For full-batch training on tiny datasets, stay conservative.
Schedules on tiny CPU budgets
When training finishes in seconds on a laptop, elaborate schedules are theater — there are not enough epochs for phases to matter. ReLU.chat's CPU policy training runs a short supervised warm-up followed by brief policy updates; a simple decay across the run suffices. Match schedule complexity to training length: step decay for dozens of epochs, cosine for hundreds, warmup-plus-decay for thousands.
Debugging with the loss curve
Plot training loss against the scheduled rate. Healthy training shows loss falling fast early, then settling as the rate decays. Loss that plateaus early suggests the rate decayed too fast; loss that oscillates late suggests the floor is too high. One glance at this curve diagnoses most schedule mistakes faster than any grid search.