Tachyontachyon

Why we shipped a search engine as a single binary

The case against a distributed system as the default starting point for application search — and what a single-node design actually costs you.

· Tachyon

Most search engines you'd reach for today — Elasticsearch, OpenSearch, Solr — are distributed systems first and search engines second. That's a reasonable design if your baseline assumption is "this needs to scale past what one machine can hold." It's a strange default if your actual workload is a product catalog, a docs site, or a support queue that fits comfortably on a single node with room to spare.

Tachyon starts from the opposite assumption: most application search is not distributed-systems-shaped, and the tooling shouldn't force you to operate one anyway.

What "distributed by default" actually costs

A minimal production Elasticsearch deployment isn't minimal. It's a JVM to size and tune garbage collection for, a cluster protocol to reason about, shard and replica counts to plan, and master-eligible nodes to keep healthy. None of that is optional complexity you can opt out of later — it's baked into the architecture from the first node.

For a workload that will never need to shard past one machine, all of that is pure overhead: operational surface with no corresponding benefit. You pay the complexity tax whether or not you ever collect on the distributed-scale payoff.

What single-node buys you back

Tachyon is one binary. docker run it, point it at a data directory, and you have a running search engine — no cluster to bootstrap, no JVM heap to size, no shard allocation to reason about. See Getting Started for the entire setup.

The tradeoff is explicit, not hidden: a single Tachyon node has a ceiling. Persistence explains exactly how memory and disk usage scale as your corpus grows, so you can see that ceiling coming rather than discover it in production. Distributed search — sharding, replication — is on the Roadmap, for teams that outgrow one node. It isn't the starting assumption baked into every deployment from day one.

The actual bet

The bet isn't "distributed search is bad." It's that most teams building application search never need it, and the ones that do can tell in advance that they will. Building the single-node case well, and being honest about where its ceiling is, serves the common case better than defaulting every deployment into cluster operations it doesn't need yet.

If your corpus fits on one machine — and for most product catalogs, docs sites, and support systems, it does — that's the whole cost-benefit calculation. See Architecture for how the single-node design is actually put together.

On this page