Search indexes, tag clouds, and related-article links all need the same thing: the handful of phrases that capture what a document is about. RAKE and TextRank extract those phrases with no training data and no model, using only word statistics and graph structure.

RAKE: statistics over stop-word gaps

RAKE (Rapid Automatic Keyword Extraction) splits text on stop words and punctuation, treating each surviving word run as a candidate phrase. It scores words by how often they appear and how often they co-occur with other content words, then sums word scores into phrase scores.

function rakeCandidates(tokens, stopWords) {
  const phrases = [];
  let current = [];
  for (const t of tokens) {
    if (stopWords.has(t)) {
      if (current.length) phrases.push(current);
      current = [];
    } else {
      current.push(t);
    }
  }
  if (current.length) phrases.push(current);
  return phrases.filter(p => p.length <= 4);
}

The word score is degree / frequency: total co-occurrences with other content words divided by occurrences. Words that appear mostly inside longer phrases score higher than words that appear alone. This simple ratio is RAKE's whole insight, and it works surprisingly well.

TextRank: PageRank for words

TextRank builds a graph where words are nodes and edges connect words that co-occur within a small window. It then runs PageRank: words linked to many high-scoring words become high-scoring themselves. Top-ranked words are merged into phrases from the original text order.

PageRank iteration is a loop of weighted sums — twenty lines of code and fast convergence on document-sized graphs. The window size (typically 2 to 4 words) controls how local the relationships are.

Which to choose

  • RAKE is simpler, faster, and better at multi-word phrases out of the box. It needs a good stop-word list since stop words define phrase boundaries.
  • TextRank handles documents where stop-word splitting produces poor candidates, and its graph formulation extends to sentence extraction for summarization.

For auto-tagging knowledge-base fragments, start with RAKE. For keyphrase highlighting inside long answers, TextRank's word graph gives finer control.

Practical improvements

Both methods benefit from the same preprocessing used in retrieval indexing: lowercase, stem for scoring but display original forms, and filter candidates by length (one to four words) and frequency (seen at least twice, or once in short texts).

Deduplicate aggressively: if 'neural network' and 'neural networks' both rank highly, keep one. Substring and stem-overlap filters remove most duplicates without semantic analysis.

Browser use cases

Extracted keywords power several on-device features: tag pills on articles, 'related topics' links between fragments, query expansion candidates, and search-result highlighting. All of it runs at index time in milliseconds per document, with results small enough to store beside each fragment.