When a system writes text — summaries, translations, composed answers — how do you score it without reading every output? BLEU and ROUGE compare generated text against human references using n-gram overlap. They are imperfect but fast, standard, and good enough to rank systems during development.

BLEU: precision over n-grams

BLEU measures how much of the generated text appears in the reference. It computes precision for unigrams through 4-grams, takes a geometric mean, and multiplies by a brevity penalty that punishes outputs shorter than the reference:

function ngramPrecision(candidate, reference, n) {
  const refCounts = countNgrams(reference, n);
  const candCounts = countNgrams(candidate, n);
  let clipped = 0, total = 0;
  for (const [g, c] of candCounts) {
    clipped += Math.min(c, refCounts.get(g) || 0);
    total += c;
  }
  return total ? clipped / total : 0;
}

Clipping caps each n-gram's credit at its reference count, so repeating 'the the the' cannot game the score. The brevity penalty is exp(1 - refLen / candLen) when the candidate is shorter, else 1.

ROUGE: recall over n-grams

ROUGE flips the emphasis: how much of the reference appears in the generated text. ROUGE-1 and ROUGE-2 measure unigram and bigram recall; ROUGE-L measures the longest common subsequence, rewarding fluent word order even with gaps. Summarization research reports ROUGE because missing reference content is the failure mode that matters — a summary that omits the key point fails no matter how fluent it is.

For extractive summarization, ROUGE recall is naturally high since outputs copy source sentences. Compare systems with ROUGE F1 to balance coverage against verbosity.

What the scores mean

Neither metric has an absolute scale. BLEU in the 30s is solid for translation; ROUGE-1 F1 near 0.4 is respectable for summarization. Use them comparatively: same references, same tokenizer, different systems. A two-point BLEU gap between your own checkpoints is meaningful; comparing your BLEU against another paper's number usually is not, because tokenization and reference counts differ.

Known blind spots

Both metrics match exact n-grams, so fluent paraphrases score poorly while awkward near-copies score well. They ignore factuality entirely — a summary can score highly while negating the source. For chatbot answers, pair overlap metrics with retrieval accuracy: first verify the system found the right knowledge, then check the wording resembles a good answer.

Practical workflow

Keep 50 to 100 reference outputs for representative queries, freeze them, and compute BLEU/ROUGE on every change alongside your offline ranking tests. When overlap scores and human spot-checks agree, ship. When they disagree, trust the humans and investigate — the metric is a proxy, not the goal.