System design explainer

Design an autocomplete system

The interview question behind every search box: you type three letters and suggestions appear instantly. How would you build that for millions of people typing at once?

The takeaway, up front

All the suggestions live in a trie sitting entirely in RAM, rebuilt offline from query logs. Nothing touches a database on the hot path. A keystroke walks down the trie, grabs the top suggestions stored under that node, and you're done in microseconds. Everything else — sharding, ranking, cache layers — is plumbing around that one fact. If you remember one sentence in the interview, make it that one.

1 · Requirements

What are we actually building?

Think of the Google search box. The analogy that carries the whole design: autocomplete is a library card catalog where every drawer is already open. You don't search the library per keystroke — you walk to the drawer labeled with your prefix and read the most popular cards inside.

Functional

  • User types a prefix → return top 10 suggestions
  • Suggestions ranked by popularity, recency, and personal history
  • Results update as the user keeps typing (per keystroke)
  • Support multiple languages and regions
  • Filter out offensive / unsafe suggestions

Non-functional

  • Latency: p99 under 100ms, ideally under 50ms — slower than that and it feels broken
  • Availability: the search box can never be down; degrade gracefully
  • Scale: tens of thousands of requests per second at peak
  • Freshness: new trending queries appear within hours, not weeks
  • Consistency: eventual is fine — two users may briefly see different orderings

🚫 Common misconception

"Autocomplete queries a database on every keystroke." It doesn't. A database lookup per keystroke — at tens of thousands of QPS with a 50ms budget — would melt. The entire index is a trie in memory; a lookup is just following one pointer per character you typed. The database only shows up offline, when the trie is rebuilt.

2 · Back-of-the-envelope

Capacity math

State your assumptions out loud, then do arithmetic the interviewer can follow:

AssumptionValue
Search queries per day100M (a large but not Google-scale engine)
Keystrokes per query~10 → 1B suggestion requests/day
Average QPS1B ÷ 86,400 ≈ 11,600 QPS
Peak QPS (3× for evenings)≈ 35,000 QPS
Distinct queries worth indexing10M (the head of the popularity curve)
Trie size10M × 20 chars avg, prefix-sharing ≈ 3× compression → ~70M nodes × ~48 bytes ≈ ~3.4 GB per full copy — fits in RAM
Bandwidth35k QPS × ~1 KB response ≈ 35 MB/s ≈ 280 Mbps — trivial
Edge cacheTop 1M prefixes cached, ~90% hit rate → only ~3.5k QPS reaches the trie tier

The punchline to say out loud: "The whole index fits in memory on a single beefy box — so this is a sharding-for-QPS-and-reliability problem, not a storage problem."

Go deeper: why 10M queries and not all of them?

Query popularity follows a power law — a small head of queries accounts for most traffic. Indexing the full long tail (hundreds of millions of distinct queries, many typed once ever) would multiply memory for almost zero hit-rate gain. Production systems index the head aggressively and let the tail fall through to "no suggestions" or a slower path. The cutoff threshold is itself a tunable: lower threshold = more memory, marginally better coverage.

3 · Architecture

The system, end to end

Two halves that barely talk to each other: a serving path (fast, in-memory, per keystroke) and an offline pipeline (slow, batch, rebuilds the index daily). They meet only when a fresh trie snapshot is pushed to the servers.

flowchart TB
    WEB["User typing in search box"]
    CDN["CDN / edge cache
hot prefixes, ~90% hit rate"] AG["API gateway
rate limiting + routing"] TS["Typeahead service
fan-out + rerank"] SH1["Trie shard A
in-memory"] SH2["Trie shard B
in-memory"] SHN["Trie shard N
in-memory"] AGG["Aggregator
merge + personalize"] LOGS["Query logs"] FILT["Filter bots, PII,
unsafe queries"] COUNT["Count frequencies"] BUILD["Build trie per shard"] PUSH["Rolling push
shard by shard"] WEB --> CDN CDN -->|"cache miss"| AG CDN -->|"cache hit"| WEB AG --> TS TS --> SH1 TS --> SH2 TS --> SHN SH1 --> AGG SH2 --> AGG SHN --> AGG AGG --> WEB LOGS --> FILT --> COUNT --> BUILD --> PUSH PUSH -. "fresh snapshot" .-> SH1 PUSH -. "fresh snapshot" .-> SH2 PUSH -. "fresh snapshot" .-> SHN

Sharding key insight: shard by hash of the prefix, so all completions for one prefix live on one shard — no cross-shard merging per keystroke. The aggregator only merges when personalization pulls in extra candidates.

4 · Component deep-dives

What happens on a single keystroke

sequenceDiagram
    autonumber
    participant C as Client
    participant E as Edge cache
    participant G as API gateway
    participant T as Typeahead service
    participant S as Trie shard
    C->>E: GET /suggest?q=tesl&k=10
    alt hot prefix — cache hit
        E-->>C: 200 suggestions (~5ms)
    else cache miss
        E->>G: forward
        G->>T: route by hash(prefix)
        T->>S: walk t-e-s-l, read top-k
        S-->>T: candidates + frequencies
        T->>T: rerank (recency, personal history)
        T-->>E: JSON, cache 60s
        E-->>C: 200 suggestions (~40ms)
    end

Inside a trie node

Each node is tiny: the character, a map of child pointers, and — if a query ends here — its frequency count. Walking the prefix tesl is 4 pointer hops. The top-k suggestions are found by a bounded DFS under that node, guided by stored frequencies. Total work per keystroke: O(prefix length + k log k).

flowchart LR
    A["Raw query logs"] --> B["Strip bots + PII"]
    B --> C["Count query frequencies"]
    C --> D["Drop queries below threshold"]
    D --> E["Blocklist unsafe terms"]
    E --> F["Build one trie per shard"]
    F --> G["Serialize + checksum snapshot"]
    G --> H["Rolling deploy, shard by shard"]
    H --> I["Health-check: spot-check
known prefixes before cutover"]
Go deeper: personalization without slowing the hot path

Keep a small per-user structure — recent searches in Redis, ~100 entries. The typeahead service fetches it in parallel with the trie lookup and boosts candidates the user has typed before. Because it's a fixed-size fetch done concurrently, p99 barely moves. Full learning-to-rank (dozens of features, ML model) comes later and only if the interviewer asks — name it, don't build it unprompted.

Go deeper: compressing the trie

A naive trie of 70M nodes with hash-map children is memory-hungry. Production tricks: store children in sorted arrays (binary search is fine at tiny fan-out), collapse single-child chains (like a radix tree), or compile to a DAWG / finite-state transducer. These cut memory 3–10× at the cost of a more complex builder — which is fine, because the builder runs offline.

5 · API + data model

The contract

GET /v1/suggest?q=tesl&k=10&lang=en
→ 200 OK
[
  {"text": "tesla",        "score": 0.98},
  {"text": "tesla stock",  "score": 0.87},
  {"text": "tesla model y", "score": 0.81}
]

POST /v1/feedback   # client reports which suggestion was clicked
{"q": "tesl", "chosen": "tesla model y"}   # trains future ranking

Data model: TrieNode { children: Map<char, TrieNode>, freq: int, is_word: bool }, serialized as an immutable snapshot per shard (versioned, checksummed). Mutable per-user recents live separately in Redis with a TTL — never mixed into the shared trie.

6 · Trade-offs

What you give up, on purpose

DecisionWhyCost
Trie in RAM vs database indexMicrosecond lookups, O(prefix) walkGBs of RAM per replica; rebuild pipeline to maintain
Daily batch rebuild vs live updatesSimple, testable, atomic cutoverTrending queries lag by hours — acceptable for v1
Edge cache in front90% of traffic never reaches originStale suggestions for up to TTL; cache invalidation on push
Prefix-hash shardingOne shard per lookup, no fan-out mergeHot prefixes can unbalance shards — mitigate with replicas + cache
No fuzzy matching on hot pathKeeps p99 tinyTypos get no help — handled client-side or by a separate service
7 · Failure modes

What breaks, and what saves you

8 · What I'd actually build

Opinionated, concrete, shippable

Serving: Go service, trie built at startup from a snapshot file, one process per shard on Kubernetes, consistent-hash routing by prefix. Cache: Cloudflare in front with 60s TTL on suggestion responses. Personalization: Redis, per-user recent queries, fetched concurrently. Offline: nightly Spark job over query logs → per-shard snapshot → rolling deploy. Observability: p50/p99 per prefix-length bucket, cache hit rate, snapshot age. That's a one-team, two-week v1 — and it handles the math above with room to spare.

9 · Interview tips

How to run the room

  1. Say "trie in memory" in the first two minutes. It's the answer; everything else is elaboration.
  2. Do the math out loud — QPS from keystrokes, GBs of RAM. Interviewers promote candidates who quantify.
  3. Draw the two halves (serving vs offline) before any deep dive. It shows you separate hot path from batch.
  4. Name personalization and ML ranking as follow-ups, not v1. Scope discipline scores points.
  5. When asked "how do updates work?", answer "we don't update live — we rebuild." Then discuss the freshness trade-off.
  6. Have one scaling story ready: "at 10× traffic we add shards and lean harder on the edge cache; the trie itself doesn't change."
10 · Interactive widget

Trie visualizer — type and watch it walk

This is a real trie built in JavaScript from words, living in your browser's memory — the same idea as the production design, just smaller. Type a prefix and watch the animated traversal, then see the top completions ranked by frequency.

Nodes: Words: Built in: Idle — type above

The traversal you just watched is the entire hot path: one pointer hop per character, then a bounded walk under the final node. No database, no index scan, no network — that's why it answers in microseconds.

v2026.10.03-01