Tokenization
How Tachyon turns text fields into the tokens its inverted index is built from.
Every text field is broken into tokens before it's indexed — the
inverted index is built entirely on tokens, never on raw field values.
What counts as a token, and how it's normalized, directly determines what
a query will and won't match.
The rules
- Alphanumeric runs — a token is (roughly) a contiguous run of letters and digits. Punctuation and whitespace split tokens; they aren't indexed themselves.
- Unicode normalization (NFD) — text is normalized to NFD (canonical decomposition) before tokenizing, so visually identical characters that could otherwise be encoded differently are treated the same.
- Combining marks stripped — after NFD decomposition, combining marks
(accents, diacritics) are dropped, so
"café"and"cafe"tokenize to the same term. - Lowercased — all tokens are lowercased; search is case-insensitive by construction, not as an opt-in query flag.
- Truncated at 64 characters — a single token longer than that is cut off, which in practice only matters for pathological input (a URL or hash dropped into a text field), not real words.
- Han ideographs segmented individually — CJK text without natural whitespace boundaries is split into standalone single-character tokens rather than left as one long unsegmented run.
What's deliberately not done
- No stemming.
"running"and"run"are different tokens today. Searching one will not match documents containing only the other. - No stop-word removal. Common words (
"the","and","of") are indexed and searched like any other token.
Both are listed on the Roadmap as planned query-surface extensions — they're absent because they haven't shipped, not because they're considered unnecessary.
Why this matters for query behavior
Tokenization rules apply identically to indexed documents and incoming
queries — the same function turns "Café" in a document and "cafe" in a
query into the matching token "cafe". Understanding the rules above is
mostly useful for predicting edge cases: why a hyphenated term splits into
two tokens, why an accented and unaccented spelling match each other, or why
a long identifier in a text field doesn't tokenize the way you might
expect past 64 characters.
Typo tolerance operates on top of tokenization, not as a replacement for it — see Typo Tolerance for how a query token that doesn't exactly match an indexed token can still find it within a length-scaled edit distance.