A chatbot that knows which language the user is writing in can route to the right knowledge base, pick the right stop-word list, and avoid answering English questions with another language's fragments. Language detection sounds like a job for a large model, but short-text classifiers built on character n-grams run in kilobytes and answer in microseconds.

Why character n-grams work

Every language has a fingerprint in its character sequences. The trigram 'the' dominates English, 'sch' marks German, 'eau' marks French. A detector compares the n-gram profile of the input against stored profiles for each supported language and picks the closest match. No dictionary, no grammar — just counting.

function trigrams(text) {
  const counts = new Map();
  const clean = text.toLowerCase().replace(/[^a-zà-ÿ]/g, ' ');
  for (let i = 0; i + 3 <= clean.length; i++) {
    const g = clean.slice(i, i + 3);
    if (!g.includes(' ')) counts.set(g, (counts.get(g) || 0) + 1);
  }
  return counts;
}

Compare profiles with cosine similarity over the shared n-gram space — the same cosine similarity used in vector search, applied to count vectors instead of embeddings.

Building language profiles

For each supported language, collect a few paragraphs of representative text, count trigrams, and keep the top few hundred by frequency. Serialize the profiles as JSON. Ten languages at 300 trigrams each compress to roughly 30 KB — trivial to ship with the page or cache in IndexedDB.

Normalize aggressively before counting: lowercase, strip digits and punctuation, and collapse whitespace. The profile should capture letter patterns, not formatting quirks of your sample text.

Handling short messages

Chat messages are short, and short text means sparse n-grams. A three-word message may share no trigrams with any profile. Three techniques help:

  • Back off to bigrams and unigrams when trigram overlap is too thin.
  • Require a confidence margin: if the top two languages score within a small delta, report 'unknown' instead of guessing.
  • Carry session context: once a user writes two messages in Turkish, default the third ambiguous message to Turkish. This is the same session memory idea applied to language.

Mixed-language input

Multilingual users mix languages mid-sentence. Rather than forcing one label, score each sentence separately and route on the majority language, or detect per-message and let retrieval search both languages' indexes. For keyword search, mixed input mostly works if stop-word removal runs per detected segment instead of once globally.

Where detection fits

Run detection first in the pipeline, before tokenization and normalization, since those steps are language-specific. Cache the result per message and per session. The whole operation — profile lookup plus a few hundred arithmetic ops — finishes in well under a millisecond, keeping you inside tight latency budgets while making every downstream step smarter.