# Tokenization
URL: /docs/concepts/tokenization

How Tachyon turns text fields into the tokens its inverted index is built from.



Every `text` field is broken into tokens before it's indexed — the
inverted index is built entirely on tokens, never on raw field values.
What counts as a token, and how it's normalized, directly determines what
a query will and won't match.

## The rules [#the-rules]

* **Alphanumeric runs** — a token is (roughly) a contiguous run of letters
  and digits. Punctuation and whitespace split tokens; they aren't indexed
  themselves.
* **Unicode normalization (NFD)** — text is normalized to NFD
  (canonical decomposition) before tokenizing, so visually identical
  characters that could otherwise be encoded differently are treated the
  same.
* **Combining marks stripped** — after NFD decomposition, combining marks
  (accents, diacritics) are dropped, so `"café"` and `"cafe"` tokenize to
  the same term.
* **Lowercased** — all tokens are lowercased; search is case-insensitive by
  construction, not as an opt-in query flag.
* **Truncated at 64 characters** — a single token longer than that is cut
  off, which in practice only matters for pathological input (a URL or hash
  dropped into a text field), not real words.
* **Han ideographs segmented individually** — CJK text without natural
  whitespace boundaries is split into standalone single-character tokens
  rather than left as one long unsegmented run.

## What's deliberately not done [#whats-deliberately-not-done]

* **No stemming.** `"running"` and `"run"` are different tokens today.
  Searching one will not match documents containing only the other.
* **No stop-word removal.** Common words (`"the"`, `"and"`, `"of"`) are
  indexed and searched like any other token.

Both are listed on the [Roadmap](/roadmap) as planned query-surface
extensions — they're absent because they haven't shipped, not because
they're considered unnecessary.

## Why this matters for query behavior [#why-this-matters-for-query-behavior]

Tokenization rules apply identically to indexed documents and incoming
queries — the same function turns `"Café"` in a document and `"cafe"` in a
query into the matching token `"cafe"`. Understanding the rules above is
mostly useful for predicting edge cases: why a hyphenated term splits into
two tokens, why an accented and unaccented spelling match each other, or why
a long identifier in a `text` field doesn't tokenize the way you might
expect past 64 characters.

Typo tolerance operates on top of tokenization, not as a replacement for
it — see [Typo Tolerance](/docs/typo-tolerance) for how a query token that
doesn't exactly match an indexed token can still find it within a
length-scaled edit distance.
