Benchmarking Tachyon against Typesense: the numbers, and their limits
A single, fully-disclosed run comparing Tachyon and Typesense 30.2 at 100K, 1M, and 5M documents — what it found, and exactly what it doesn't prove yet.
We ran a single, informal comparison of Tachyon against Typesense 30.2 at 100K, 1M, and 5M documents — ingest time, memory, disk footprint, and search latency, all fully disclosed on the Benchmarks page along with hardware, dataset generation, and query mix. This post is the narrative version: what the numbers actually show, and — just as important — what this run doesn't prove.
This is one run, on a laptop, against default configuration for both engines, with no repeated trials and no concurrency sweep. Treat every number below as directional until it's superseded by the reproducible benchmark suite both engines deserve to be measured with.
The setup, briefly
MacBook Air (M4, 24 GB), both engines run one at a time inside Docker
Desktop, adikeshri/tachyon:latest vs. typesense/typesense:30.2, both on
default configuration. 500 fixed-vocabulary queries per run, sequential,
single client, first 5 discarded as warmup. Full detail, including the
exact caveats this run doesn't satisfy against our own published
methodology, is on Benchmarks — this post
doesn't repeat it in full.
Memory: the largest gap, and the most explainable one
At every corpus size tested, Tachyon's memory footprint was substantially smaller:
| Tachyon (steady) | Typesense (steady) | |
|---|---|---|
| 100K docs | 24.2 MiB | 343.3 MiB |
| 1M docs | 108.3 MiB | 462.9 MiB |
| 5M docs | 426.3 MiB | 944.7 MiB |
Part of the 100K gap is a fixed cost, not a scaling one: Typesense's single-node consensus layer carries roughly 200–300 MiB of baseline overhead before any data is loaded at all, which dominates the comparison at the smallest corpus size. But the gap doesn't close as the corpus grows — by 5M documents, Tachyon was using well under half the memory Typesense was. That's consistent with the architectural difference described in Persistence: Tachyon's segments are memory-mapped and decoded lazily, so working-set memory tracks what a query actually touches rather than total corpus size, while Typesense is built to hold its index resident in RAM.
Ingest: faster throughout, by a widening margin
| Tachyon | Typesense | Tachyon throughput | |
|---|---|---|---|
| 100K docs | 0.99s | 1.57s | 101.3K docs/s |
| 1M docs | 11.73s | 18.25s | 85.3K docs/s |
| 5M docs | 79.74s | 105.41s | 62.7K docs/s |
Tachyon ingested faster at every size, and the absolute gap widened with corpus size even as both engines' throughput dropped somewhat as their indexes grew. One run isn't enough to call this a scaling trend with confidence — it's a data point worth confirming with repeated trials.
Search latency: closer, and size-dependent
This is where the picture is more mixed. At 100K and 1M documents, Tachyon was consistently faster across best-case, p95, p99, and worst-case latency. At 5M documents, best-case latency essentially tied (15.6ms vs. 15.28ms), while Tachyon held a clear lead at the tail — p95 was 38.98ms vs. 89.12ms, and worst-case was 106.39ms vs. 280.02ms.
Tail latency separating out more than best-case latency, and doing so more as the corpus grows, is the kind of signal that's genuinely interesting — and exactly the kind of signal that needs a proper concurrency sweep and repeated trials to trust, not a single sequential run. It's on the list for the real suite, not a conclusion we're drawing from this one.
Where Typesense came out ahead
On-disk footprint: Tachyon used more disk than Typesense at every size tested (2875.6 MiB vs. 1764.6 MiB at 5M documents) — the flip side of keeping less in memory. That's a real tradeoff, not an oversight: more of the index lives on disk in a form the OS page cache manages, rather than being held resident. Worth knowing if disk, not memory, is your tighter constraint.
What this run doesn't establish
Against our own published methodology, this run
has real gaps: it isn't pinned to a specific Tachyon commit (just the
latest image tag), it has no separate numbers for filtered, faceted, or
sorted queries — every query was plain text — and it has no concurrency
sweep at all, just a single sequential client. Those are exactly the gaps
the reproducible benchmark suite on the Roadmap is meant to
close. Until then, this is one fully-disclosed data point, not a verdict.
See the full Benchmarks page for every number, the complete methodology, and the dataset generation details.