When retrieved knowledge runs long, users benefit from a compressed view: the three sentences that matter most. Extractive summarization picks those sentences from the source text instead of generating new ones, which guarantees every word in the summary actually appeared in the documents.

Sentence scoring with word frequencies

The simplest extractor scores each sentence by the sum of its words' frequencies, normalized by sentence length. Frequent content words indicate the document's main topics, so sentences dense in them are likely representative:

function summarize(sentences, topK = 3) {
  const freq = new Map();
  for (const s of sentences)
    for (const w of contentWords(s))
      freq.set(w, (freq.get(w) || 0) + 1);
  const scored = sentences.map((s, i) => {
    const words = contentWords(s);
    const sum = words.reduce((a, w) => a + (freq.get(w) || 0), 0);
    return { s, i, score: words.length ? sum / words.length : 0 };
  });
  return scored.sort((a, b) => b.score - a.score).slice(0, topK)
    .sort((a, b) => a.i - b.i).map(x => x.s);
}

Note the final re-sort into original order: summaries read far better when sentences keep their source sequence.

Position and title bonuses

Pure frequency scoring misses structural cues. Two cheap bonuses help a lot: boost sentences near the start of each section (topic sentences live there), and boost sentences sharing words with the title or the user's query. Query-biased scoring turns a generic summary into an answer-focused one — exactly what a chatbot needs.

Removing redundancy

Top-scoring sentences often say the same thing twice. After picking the best sentence, penalize remaining candidates by their word overlap with already-chosen ones (maximal marginal relevance in miniature). A Jaccard overlap penalty of 50% or more keeps summaries diverse without any embeddings.

Extractive versus generative

Generative summarizers rewrite content fluently but can hallucinate details — dangerous when answers cite sources. Extractive summaries are choppier but faithful by construction: each sentence links back to its fragment. For a fragment-composed chatbot, extraction fits naturally: select the best sentences from the top-ranked fragments and join them with linguistic connectors.

Evaluation and limits

Measure summary quality with ROUGE, which counts n-gram overlap against reference summaries. Expect extractive methods to score well on ROUGE recall — they copy source phrases verbatim — while trailing humans on coherence.

Extraction fails when the answer needs synthesis across documents ('compare X and Y') rather than selection within them. Detect comparison and how-to intents with your intent classifier and fall back to structured templates for those cases instead of forcing extraction.