Tachyontachyon
Concepts

Tokenization

How Tachyon turns text fields into the tokens its inverted index is built from.

Every text field is broken into tokens before it's indexed — the inverted index is built entirely on tokens, never on raw field values. What counts as a token, and how it's normalized, directly determines what a query will and won't match.

The rules

  • Alphanumeric runs — a token is (roughly) a contiguous run of letters and digits. Punctuation and whitespace split tokens; they aren't indexed themselves.
  • Unicode normalization (NFD) — text is normalized to NFD (canonical decomposition) before tokenizing, so visually identical characters that could otherwise be encoded differently are treated the same.
  • Combining marks stripped — after NFD decomposition, combining marks (accents, diacritics) are dropped, so "café" and "cafe" tokenize to the same term.
  • Lowercased — all tokens are lowercased; search is case-insensitive by construction, not as an opt-in query flag.
  • Truncated at 64 characters — a single token longer than that is cut off, which in practice only matters for pathological input (a URL or hash dropped into a text field), not real words.
  • Han ideographs segmented individually — CJK text without natural whitespace boundaries is split into standalone single-character tokens rather than left as one long unsegmented run.

What's deliberately not done

  • No stemming. "running" and "run" are different tokens today. Searching one will not match documents containing only the other.
  • No stop-word removal. Common words ("the", "and", "of") are indexed and searched like any other token.

Both are listed on the Roadmap as planned query-surface extensions — they're absent because they haven't shipped, not because they're considered unnecessary.

Why this matters for query behavior

Tokenization rules apply identically to indexed documents and incoming queries — the same function turns "Café" in a document and "cafe" in a query into the matching token "cafe". Understanding the rules above is mostly useful for predicting edge cases: why a hyphenated term splits into two tokens, why an accented and unaccented spelling match each other, or why a long identifier in a text field doesn't tokenize the way you might expect past 64 characters.

Typo tolerance operates on top of tokenization, not as a replacement for it — see Typo Tolerance for how a query token that doesn't exactly match an indexed token can still find it within a length-scaled edit distance.

On this page