Before a search engine scores documents, it normalizes their text. Two of the oldest normalization steps are stop-word removal and stemming: dropping filler words and reducing inflected forms to a common root. On small indexes they cut index size, speed up queries, and improve matches — if you apply them carefully.
What stop words are
Stop words are high-frequency terms that carry little meaning on their own: 'the', 'is', 'and', 'of', 'to'. In a TF-IDF pipeline they earn near-zero weights anyway, so removing them up front mostly saves index space and query time rather than changing rankings.
A typical English stop list has 100 to 200 entries. Keep the list short for small knowledge bases: every removed word is a word users can never search for. If your content includes phrases like 'to be or not to be', aggressive stop-word removal destroys the query.
What stemming does
Stemming reduces related word forms to a shared stem so that 'running', 'runs', and 'ran' all match. The Porter stemmer is the classic choice: a sequence of suffix-stripping rules with no dictionary. It is fast, dependency-free, and good enough for retrieval in most cases.
// Porter-style suffix rules (simplified excerpt)
function stem(word) {
if (word.endsWith('ies') && word.length > 4) return word.slice(0, -3) + 'i';
if (word.endsWith('es') && word.length > 3) return word.slice(0, -2);
if (word.endsWith('s') && word.length > 3) return word.slice(0, -1);
if (word.endsWith('ing') && word.length > 5) return word.slice(0, -3);
if (word.endsWith('ed') && word.length > 4) return word.slice(0, -2);
return word;
}
Apply the identical function to documents at index time and to queries at search time. A stemmer applied on only one side silently breaks matching.
Stemming versus lemmatization
Lemmatization maps words to dictionary forms ('better' becomes 'good') using a vocabulary and part-of-speech tags. It is more accurate but needs language data and more code. For browser search over a few hundred documents, stemming wins: a Porter implementation is roughly a hundred lines with zero data files, while lemmatization ships a lexicon.
When normalization hurts
Normalization is lossy, and the losses matter in specific cases:
- Negation: removing 'not' and 'no' as stop words turns 'not working' into 'working'.
- Code and identifiers: stemming 'classes' to 'class' is fine, but mangling API names breaks exact lookup.
- Short queries: with one or two meaningful words, each dropped stop word removes a large share of the signal.
A pragmatic rule: keep 'not', 'no', and 'never' out of your stop list, and skip stemming for fields that hold identifiers or titles matched with field-weighted BM25.
A minimal normalization pipeline
For small on-device indexes, this order works well:
- Lowercase and strip punctuation.
- Split on whitespace (subword tokenization is overkill here).
- Drop stop words, except negation terms.
- Stem the survivors.
The whole pipeline is a few dozen lines and runs in microseconds per query. Combined with an inverted index, it gives you a complete keyword search stack with no network calls and no model weights.
Measure before and after
Normalization changes rankings, so evaluate with precision and recall on a small set of real queries before shipping. If stemming helps long-tail queries but hurts exact-phrase ones, index both the stemmed and unstemmed forms and weight exact matches higher — the best of both worlds for a few kilobytes of index.