System design explainer
The interview question behind every search box: you type three letters and suggestions appear instantly. How would you build that for millions of people typing at once?
All the suggestions live in a trie sitting entirely in RAM, rebuilt offline from query logs. Nothing touches a database on the hot path. A keystroke walks down the trie, grabs the top suggestions stored under that node, and you're done in microseconds. Everything else — sharding, ranking, cache layers — is plumbing around that one fact. If you remember one sentence in the interview, make it that one.
Think of the Google search box. The analogy that carries the whole design: autocomplete is a library card catalog where every drawer is already open. You don't search the library per keystroke — you walk to the drawer labeled with your prefix and read the most popular cards inside.
"Autocomplete queries a database on every keystroke." It doesn't. A database lookup per keystroke — at tens of thousands of QPS with a 50ms budget — would melt. The entire index is a trie in memory; a lookup is just following one pointer per character you typed. The database only shows up offline, when the trie is rebuilt.
State your assumptions out loud, then do arithmetic the interviewer can follow:
| Assumption | Value |
|---|---|
| Search queries per day | 100M (a large but not Google-scale engine) |
| Keystrokes per query | ~10 → 1B suggestion requests/day |
| Average QPS | 1B ÷ 86,400 ≈ 11,600 QPS |
| Peak QPS (3× for evenings) | ≈ 35,000 QPS |
| Distinct queries worth indexing | 10M (the head of the popularity curve) |
| Trie size | 10M × 20 chars avg, prefix-sharing ≈ 3× compression → ~70M nodes × ~48 bytes ≈ ~3.4 GB per full copy — fits in RAM |
| Bandwidth | 35k QPS × ~1 KB response ≈ 35 MB/s ≈ 280 Mbps — trivial |
| Edge cache | Top 1M prefixes cached, ~90% hit rate → only ~3.5k QPS reaches the trie tier |
The punchline to say out loud: "The whole index fits in memory on a single beefy box — so this is a sharding-for-QPS-and-reliability problem, not a storage problem."
Query popularity follows a power law — a small head of queries accounts for most traffic. Indexing the full long tail (hundreds of millions of distinct queries, many typed once ever) would multiply memory for almost zero hit-rate gain. Production systems index the head aggressively and let the tail fall through to "no suggestions" or a slower path. The cutoff threshold is itself a tunable: lower threshold = more memory, marginally better coverage.
Two halves that barely talk to each other: a serving path (fast, in-memory, per keystroke) and an offline pipeline (slow, batch, rebuilds the index daily). They meet only when a fresh trie snapshot is pushed to the servers.
flowchart TB
WEB["User typing in search box"]
CDN["CDN / edge cache
hot prefixes, ~90% hit rate"]
AG["API gateway
rate limiting + routing"]
TS["Typeahead service
fan-out + rerank"]
SH1["Trie shard A
in-memory"]
SH2["Trie shard B
in-memory"]
SHN["Trie shard N
in-memory"]
AGG["Aggregator
merge + personalize"]
LOGS["Query logs"]
FILT["Filter bots, PII,
unsafe queries"]
COUNT["Count frequencies"]
BUILD["Build trie per shard"]
PUSH["Rolling push
shard by shard"]
WEB --> CDN
CDN -->|"cache miss"| AG
CDN -->|"cache hit"| WEB
AG --> TS
TS --> SH1
TS --> SH2
TS --> SHN
SH1 --> AGG
SH2 --> AGG
SHN --> AGG
AGG --> WEB
LOGS --> FILT --> COUNT --> BUILD --> PUSH
PUSH -. "fresh snapshot" .-> SH1
PUSH -. "fresh snapshot" .-> SH2
PUSH -. "fresh snapshot" .-> SHN
Sharding key insight: shard by hash of the prefix, so all completions for one prefix live on one shard — no cross-shard merging per keystroke. The aggregator only merges when personalization pulls in extra candidates.
sequenceDiagram
autonumber
participant C as Client
participant E as Edge cache
participant G as API gateway
participant T as Typeahead service
participant S as Trie shard
C->>E: GET /suggest?q=tesl&k=10
alt hot prefix — cache hit
E-->>C: 200 suggestions (~5ms)
else cache miss
E->>G: forward
G->>T: route by hash(prefix)
T->>S: walk t-e-s-l, read top-k
S-->>T: candidates + frequencies
T->>T: rerank (recency, personal history)
T-->>E: JSON, cache 60s
E-->>C: 200 suggestions (~40ms)
end
Each node is tiny: the character, a map of child pointers, and — if a query ends here — its frequency count. Walking the prefix tesl is 4 pointer hops. The top-k suggestions are found by a bounded DFS under that node, guided by stored frequencies. Total work per keystroke: O(prefix length + k log k).
flowchart LR
A["Raw query logs"] --> B["Strip bots + PII"]
B --> C["Count query frequencies"]
C --> D["Drop queries below threshold"]
D --> E["Blocklist unsafe terms"]
E --> F["Build one trie per shard"]
F --> G["Serialize + checksum snapshot"]
G --> H["Rolling deploy, shard by shard"]
H --> I["Health-check: spot-check
known prefixes before cutover"]
Keep a small per-user structure — recent searches in Redis, ~100 entries. The typeahead service fetches it in parallel with the trie lookup and boosts candidates the user has typed before. Because it's a fixed-size fetch done concurrently, p99 barely moves. Full learning-to-rank (dozens of features, ML model) comes later and only if the interviewer asks — name it, don't build it unprompted.
A naive trie of 70M nodes with hash-map children is memory-hungry. Production tricks: store children in sorted arrays (binary search is fine at tiny fan-out), collapse single-child chains (like a radix tree), or compile to a DAWG / finite-state transducer. These cut memory 3–10× at the cost of a more complex builder — which is fine, because the builder runs offline.
GET /v1/suggest?q=tesl&k=10&lang=en
→ 200 OK
[
{"text": "tesla", "score": 0.98},
{"text": "tesla stock", "score": 0.87},
{"text": "tesla model y", "score": 0.81}
]
POST /v1/feedback # client reports which suggestion was clicked
{"q": "tesl", "chosen": "tesla model y"} # trains future ranking
Data model: TrieNode { children: Map<char, TrieNode>, freq: int, is_word: bool }, serialized as an immutable snapshot per shard (versioned, checksummed). Mutable per-user recents live separately in Redis with a TTL — never mixed into the shared trie.
| Decision | Why | Cost |
|---|---|---|
| Trie in RAM vs database index | Microsecond lookups, O(prefix) walk | GBs of RAM per replica; rebuild pipeline to maintain |
| Daily batch rebuild vs live updates | Simple, testable, atomic cutover | Trending queries lag by hours — acceptable for v1 |
| Edge cache in front | 90% of traffic never reaches origin | Stale suggestions for up to TTL; cache invalidation on push |
| Prefix-hash sharding | One shard per lookup, no fan-out merge | Hot prefixes can unbalance shards — mitigate with replicas + cache |
| No fuzzy matching on hot path | Keeps p99 tiny | Typos get no help — handled client-side or by a separate service |
Serving: Go service, trie built at startup from a snapshot file, one process per shard on Kubernetes, consistent-hash routing by prefix. Cache: Cloudflare in front with 60s TTL on suggestion responses. Personalization: Redis, per-user recent queries, fetched concurrently. Offline: nightly Spark job over query logs → per-shard snapshot → rolling deploy. Observability: p50/p99 per prefix-length bucket, cache hit rate, snapshot age. That's a one-team, two-week v1 — and it handles the math above with room to spare.
This is a real trie built in JavaScript from words, living in your browser's memory — the same idea as the production design, just smaller. Type a prefix and watch the animated traversal, then see the top completions ranked by frequency.