Tachyontachyon
Concepts

Segment Merging

Why segment count grows over time, why that alone slows queries down, and how Tachyon's background tiered merge bounds it.

Every flush from the in-memory memtable creates a new, immutable segment on disk. Left unchecked, that means segment count grows without bound as a collection accumulates writes — and segment count, not corpus size directly, is what drives query latency up over time.

Why segment count matters

A search fans out across the current memtable plus every committed segment, merging results from each. More segments means more places a single query has to look, even if total document count hasn't changed. A collection that received the same documents in one large flush versus a hundred small ones holds identical data, but the hundred-segment version is slower to query — purely because of fragmentation, not data volume.

The merge trigger

Tachyon bounds segment count with a background tiered merge. Once a collection holds more than --merge-trigger-segments/ TACHYON_MERGE_TRIGGER_SEGMENTS (default 8) segments, a merge kicks off automatically, right after the flush that crossed the threshold.

What gets merged

A merge folds the smallest --merge-fan-in/TACHYON_MERGE_FAN_IN (default 4) segments — by document count — into one new segment. Picking the smallest segments keeps merge cost proportional to the data actually being folded together, rather than repeatedly re-merging the collection's largest segments.

A merge holds everything it's folding together in memory at once while it runs. --merge-fan-in is the knob that trades merge memory usage against how many segments accumulate between merges — see Persistence for how this interacts with the other memory-bounding flag, --max-memtable-docs.

Concurrency during a merge

Segments are immutable and shared via reference counting. A query that has already taken its snapshot of the segment list keeps reading that exact snapshot consistently, even while a merge swaps in the new merged segment and retires the old ones underneath it for the next query. Readers are never blocked by a merge in progress, and a merge never sees a half-written segment — see Architecture → Concurrency.

The tradeoff

Merging is background work that competes for I/O and CPU with foreground query and ingest traffic while it runs, and it temporarily needs disk space for both the old segments being merged and the new one being written. In exchange, it's what keeps query latency from degrading indefinitely as a collection accumulates writes over its lifetime — the alternative is segment count growing forever and every query getting incrementally slower with no ceiling.

On this page