Neural networks output raw scores (logits) that can be negative, huge, or uncalibrated. Softmax turns them into probabilities that sum to one. Temperature tunes how sharp or flat that distribution is. Together they control every sampling decision a model makes.

From logits to probabilities

Softmax exponentiates each logit and divides by the total: p(i) = exp(z(i)) / sum(exp(z)). The largest logit becomes the most probable class, but every class keeps some probability. Subtract the maximum logit first for numerical stability — the math is identical and the exponentials never overflow:

function softmax(logits, temperature = 1) {
  const max = Math.max(...logits);
  const exps = logits.map(z => Math.exp((z - max) / temperature));
  const sum = exps.reduce((a, b) => a + b, 0);
  return exps.map(e => e / sum);
}

What temperature does

Temperature divides the logits before softmax. Low temperature (0.3) sharpens the distribution toward the top choice — nearly greedy. High temperature (1.5) flattens it, giving unlikely options real sampling chances. At temperature 1 you get the model's native distribution.

  • Classification: use temperature 1 and take the argmax. Temperature only matters when sampling.
  • Creative generation: raise temperature for variety, accepting occasional odd choices.
  • Factual answers: lower temperature toward greedy for consistency, or skip sampling entirely with fragment composition.

Calibration: when probabilities lie

Softmax outputs sum to one, but that does not make them calibrated: a model can assign 90% probability and be right 60% of the time. Modern networks tend toward overconfidence. Temperature scaling on a held-out set — fitting one temperature value that minimizes validation cross-entropy — is the cheapest calibration fix and changes no predictions, only their stated confidence.

Softmax in small policies

Policy networks like ReLU.chat's action heads use softmax over a handful of actions: answer directly, ask a clarifying question, show examples, and so on. With few actions, inspect the full distribution during debugging instead of just the top action — a flat distribution signals genuine uncertainty where the heuristic fallback may decide better than the learned head.

Common pitfalls

Never apply softmax twice (once in the model, once in post-processing) — the second application distorts already-normalized values. And remember that argmax discards temperature entirely: if your pipeline always takes the top class, temperature settings change nothing. Temperature is a sampling parameter, not a quality dial.