A language model outputs a probability distribution per token; decoding turns those distributions into a finished sequence. Greedy decoding picks the best token at each step. Beam search keeps several promising prefixes alive. The choice shapes output quality, diversity, and latency.

Greedy decoding

Greedy decoding takes the argmax at every step and feeds it back as context. It is fast (one forward pass per token), deterministic, and optimal when each step's best choice leads to the best sequence — which is often true for short, constrained outputs like classifications or single-phrase answers.

Its failure mode is myopia: an early token that looks best locally can trap the sequence. 'The bank of the...' commits to a riverbank before the model sees whether the sentence needed a financial bank. No backtracking, no recovery.

Beam search maintains k candidate prefixes (the beam), expands each with the top next tokens, and keeps the k highest-scoring results. With k=4, the riverbank and the financial bank both survive until later context disambiguates:

function beamStep(beams, logProbs, k) {
  const candidates = [];
  for (const b of beams)
    for (let t = 0; t < logProbs.length; t++)
      candidates.push({ seq: [...b.seq, t], score: b.score + logProbs[t] });
  candidates.sort((a, b) => b.score - a.score);
  return candidates.slice(0, k);
}

Normalize scores by length or the beam favors short sequences (fewer added log-probs, which are negative). A length penalty of score / len^0.7 is a common default.

Cost versus quality

Beam search costs roughly k times greedy's compute and memory. Quality gains concentrate in open-ended generation: translations, summaries, and long answers. For narrow tasks — intent labels, entity spans, temperature-1 sampling with argmax — beams add latency without visible wins.

Diminishing returns hit early: k=4 captures most of the gain over greedy, and k beyond 8 rarely helps while slowing interactive latency past acceptable budgets.

Sampling alternatives

Deterministic decoding repeats itself: the same prompt always yields the same output. Stochastic methods trade determinism for diversity — top-k sampling restricts draws to the k most likely tokens, nucleus (top-p) sampling takes the smallest set covering probability mass p. Chatbots that want varied phrasings sample; chatbots that want consistent answers decode greedily or skip generation for fragment composition.

Choosing for your system

  • Greedy: classifications, short answers, latency-critical paths.
  • Small beam (k=2-4): translations, summaries, anything where fluency matters and users wait anyway.
  • Sampling: creative or varied output where determinism reads as robotic.

Evaluate candidates with BLEU/ROUGE on reference outputs, but confirm with human reads: decoding changes are exactly the kind humans notice and metrics underweight.