How to study

Step-by-step system design interview framework

Use this checklist on every system design question. The shape is the same whether the prompt is a URL shortener, a rate limiter, or a full news feed.

Step 1: Clarify requirements and scope

Split requirements into functional (what must the system do, i.e. the core use cases) and non-functional (scale, latency, availability targets). Ask about read/write ratio and rough traffic numbers before drawing a single box. Confirm what is explicitly out of scope; interviewers expect you to negotiate scope, not design everything.

Step 2: Estimate scale

Back-of-envelope QPS, storage, and bandwidth from the numbers in Step 1 (daily active users, requests per user, average payload size). You do not need precision; round to the nearest order of magnitude and state your assumptions out loud. These numbers drive later decisions, since a single Postgres instance is fine at 100 QPS but not at 100k.

Step 3: Define the core entities and API

List the nouns the system needs to model (for a URL shortener: ShortURL, User, ClickEvent) and sketch the API surface, the endpoints or function signatures a client would call. This forces you to pin down the contract before the architecture.

Step 4: High-level design

Draw boxes and arrows: client, API layer, application servers, data stores, cache, queue, whatever the problem needs. Walk through one request end to end (write path, then read path) before adding anything else. Keep this version simple; you will deepen it in the next step.

Step 5: Deep dive on the bottleneck

Pick the one or two components where the interesting trade-offs live, usually the database choice, a caching layer, or a hot-path service, and go deep: SQL vs. NoSQL and why, cache invalidation strategy, how you shard or partition, how you keep a counter or ranking consistent under concurrent writes. This is where most of the interview signal comes from; do not spend all your time on Step 4's boxes.

Step 6: Scale, fail, and trade off

Address what breaks first as load grows (single point of failure, hot shard, cache stampede) and how you would fix it: replication, load balancing, horizontal partitioning, rate limiting at the edge. State the trade-off explicitly (consistency vs. availability, cost vs. latency); naming the trade-off matters more than picking a "correct" side.

Step 7: Wrap up

Summarize the design in two sentences, note what you would monitor in production (error rate, p99 latency, queue depth), and flag anything you would revisit with more time. A clean close signals seniority as much as the design itself.

What you're actually being graded on

Interviewers are scoring four things underneath the seven steps. Problem navigation is whether you scoped and sized the problem before designing anything (Steps 1-2); skip it and everything downstream is a confident answer to the wrong question. Solution design is whether the boxes in Step 4 satisfy the requirements you gathered, not whether they resemble a familiar diagram. Technical depth is almost entirely Step 5: going past naming a component into the mechanism, why this eviction policy or this shard key and not another. Communication is graded the whole time: thinking out loud and naming trade-offs explicitly instead of presenting one design as the only option. A candidate strong on the first three who stays silent while drawing usually scores lower than the diagram deserves, since the interviewer has no reasoning to grade.

Core system design concepts you need cold

Every deep dive in Step 5 eventually touches the same handful of ideas, so know them well enough to reach for the right one without stalling. CAP theorem says a distributed system facing a network partition can guarantee consistency or availability, not both at once: choose consistency and a node that cannot confirm the latest write refuses the request, choose availability and it serves a possibly stale one. Most real systems pick availability for reads and consistency for the specific writes that actually need it, a bank balance, rather than applying one answer everywhere. Horizontal scaling adds more machines and spreads the load across them, which is how you handle traffic that outgrows any single box; vertical scaling just buys a bigger box, which is simpler but hits a hard ceiling and a single point of failure. Interviewers expect you to default to horizontal for anything at real scale and explain why.

Load balancing decides which of those horizontally scaled machines handles a given request. Round robin is the simplest option and works fine when requests cost roughly the same to serve; least-connections routes to whichever server is least busy right now, which matters when request cost varies a lot; consistent hashing keeps the same client or key landing on the same server as machines get added or removed, which is what you want in front of a cache or a sharded data store so you are not invalidating everything on every scaling event. Latency, throughput, and availability get used loosely in casual conversation but mean distinct things in an interview: latency is how long one request takes, throughput is how many requests the system handles per second, and availability is the fraction of time the system successfully responds at all. A system can have low latency and low availability at the same time, fast when it works, down a lot, so name which one you are actually optimizing for in the deep dive instead of treating speed as a single goal.

Sharding and partitioning both split data across multiple machines, but the distinction matters when you say it out loud: partitioning is the general idea of splitting a dataset, sharding specifically means splitting it across separate database instances rather than within one instance. Pick a shard key that matches your access pattern, since a bad one creates a hot shard that no amount of horizontal scaling fixes. SQL gives you transactions, joins, and a fixed schema in exchange for coordination cost at scale; NoSQL trades that structure for horizontal scale and schema flexibility, and the honest answer to "SQL or NoSQL" is almost always "depends which of those two things this specific entity needs", not a blanket preference. Caching sits in front of whichever store you pick: cache-aside, where the application checks the cache and falls back to the database on a miss, is the default because it only caches what actually gets read, while write-through keeps the cache updated on every write at the cost of extra write latency. Whatever the strategy, name your invalidation approach explicitly, a short TTL, an explicit invalidation on write, or accepting brief staleness, because caching without an invalidation story is the fastest way to lose points in a deep dive. Consistency patterns follow the same logic as CAP: strong consistency means every read sees the latest write, eventual consistency means it will eventually, and most systems mix the two on purpose, strong for the data where staleness is unacceptable, eventual for everything else where it buys availability and speed. A CDN pushes static and slow-changing content physically closer to the user by caching it at edge locations, which is why it comes up in almost every deep dive that touches media or a global user base, and leader election exists to make sure exactly one node coordinates a task like handling writes to a shard, so two nodes never make conflicting decisions at the same time.

Specific technologies worth knowing by name

Naming a real tool instead of a generic noun ("a message queue" vs. "Kafka") signals you've operated something, not just read about it. Redis covers the cache-aside and rate-limiter conversations: an in-memory store fast enough for sub-millisecond reads, with optional persistence for durability. Kafka is the default answer whenever a design needs to decouple a producer from a slower consumer or fan out one event to several systems, the ad-click-aggregator and notification prompts both lean on it. Cassandra and DynamoDB are the wide-column/key-value stores worth naming for horizontal write scale with a simple access pattern; reach for PostgreSQL instead the moment the design needs transactions or joins, which is more often than candidates assume. Elasticsearch is what you name for full-text search or log aggregation, not a general-purpose database. ZooKeeper (or etcd) answers "how do you do leader election or distributed configuration," a question most candidates leave blank. An API gateway sits in front of the application servers in Step 4 to handle auth, rate limiting, and routing in one place instead of duplicating that logic per service. You don't need production experience with all seven, just enough to say which problem each solves and drop it into the right box in Step 4.

What Are the Most Common System Design Interview Questions?

Most prompts are a variation on a handful of underlying problems. Run each one through the 7-step framework above instead of memorizing a fixed architecture; the interesting part is almost always the deep dive in Step 5, not the boxes in Step 4.

Social and feed systems

  • Design a news feed (Facebook, LinkedIn) — the real difficulty is fan-out on write vs. fan-out on read for users with millions of followers.
  • Design Instagram or a photo-sharing app — object storage and CDN placement for media, plus a feed ranking service on top.
  • Design a notification system — fan-out to push, email, and SMS with per-channel retry and de-duplication.

Messaging and real-time

  • Design WhatsApp or a chat system — message ordering, delivery receipts, and connection state at scale (long-lived WebSocket or persistent connections per user).
  • Design a live comments or reaction feed — high write volume against a single hot key (the post), and how you shard or batch around it.
  • Design a collaborative document editor (Google Docs) — operational transforms or CRDTs for concurrent edits, not a database schema question.

Infra and platform

  • Design a rate limiter — token bucket vs. sliding window, and where the limiter sits (edge, gateway, or per-service). Worked through step by step.
  • Design a URL shortener — the canonical warm-up question: ID generation strategy, redirect latency, and read-heavy caching. Worked through step by step.
  • Design a web crawler — politeness/rate limits per domain, dedup at scale, and a distributed frontier queue.
  • Design a distributed cache or key-value store — consistent hashing, replication, and how you handle a hot key.
  • Design a job scheduler or distributed task queue — exactly-once vs. at-least-once delivery, and retry/backoff design.
  • Design an ad click aggregator or metrics pipeline — high-throughput ingestion with a streaming aggregation layer (Kafka-style), not a single database write path.

Storage and retrieval

  • Design a search autocomplete or type-ahead service — a trie or prefix index kept in memory, refreshed from a slower batch job.
  • Design a distributed file storage system (Dropbox, Google Drive) — chunking, deduplication, and sync conflict resolution.
  • Design a ticketing or seat-reservation system (Ticketmaster) — preventing double-booking under concurrent writes, usually with optimistic locking or a short-lived hold.
  • Design a video streaming service (Netflix, YouTube) — encoding into multiple bitrates, chunked delivery, and CDN placement so playback adapts to the viewer's connection instead of buffering.

Commerce and marketplace

  • Design a ride-sharing dispatch system (Uber, Lyft) — geospatial indexing (geohashing or a quadtree) to match riders and drivers nearby in real time.
  • Design a payment or checkout system — idempotency keys so a retried request never double-charges, plus a reconciliation path for failures.
  • Design a recommendation system — matching users to items from a behavior signal, usually a candidate-generation step followed by a ranking step, not a single database query.
  • Design a bill-splitting app (Splitwise) — modeling debts as a graph and simplifying who owes whom into the fewest transactions.

None of these need a memorized textbook diagram. The prompt is a stand-in for testing whether you can scope it (Step 1), size it (Step 2), and go deep on the one or two components where the real trade-off lives (Step 5) — the same reasoning the parking lot exercises on a single-process design instead of a distributed one.

FAQ

How long should a system design interview answer take?

Most onsite loops give you 45 minutes. Roughly: 5 min requirements, 5 min estimation, 10 min high-level design, 15 min deep dive, 10 min wrap-up and trade-offs. The exact split moves depending on the interviewer's questions, but deep dive should always be your longest block. A shallow deep dive is the most common reason a design interview underperforms.

Do I need real production experience to answer these well?

It helps but is not required. What interviewers actually grade is structured reasoning: can you break scope into requirements, defend a design decision with a trade-off, and reason about what breaks at scale. Working through curated questions with the algorithms and placement decisions spelled out builds that reasoning even without prior infrastructure experience.

What's the difference between a system design interview and an OOD interview?

System design is about distributed systems at scale: multiple servers, databases, caches, network boundaries. OOD interviews are about class structure and responsibilities in a single process, like a parking lot or an elevator system. Some prompts (a rate limiter, for example) get asked both ways: as a distributed-systems problem with Redis and sharding, or as a single-class design problem. Confirm which one you're being asked before you start.

Do I need to memorize specific architectures?

No. Memorizing "the Twitter architecture" fails the moment the interviewer changes one requirement. Memorize the framework above instead; it generates a reasonable architecture for almost any prompt because the estimation and trade-off steps adapt to whatever numbers and constraints you're given.

What are the most important system design concepts to know before the interview?

CAP theorem, horizontal versus vertical scaling, caching strategies and invalidation, consistency patterns, and sharding versus partitioning come up in nearly every deep dive, regardless of which prompt you draw. They're covered above, in the concept primer between the framework and the full question list, each with the trade-off interviewers actually want to hear instead of a textbook definition.

How many weeks should I spend preparing for system design interviews?

Three to four weeks of consistent practice for most candidates, less if you already work with distributed systems day to day, more if this is your first design round. The framework and concept primer above take one sitting to read, but recognizing the right deep-dive angle under time pressure is a rehearsal skill, not a reading skill: budget most of those weeks to talking through 8-12 problems out loud against a timer, not to reading more explanations. Candidates who spend three weeks reading and one weekend practicing consistently underperform candidates who read once and practice for three weeks, because the graded skill is retrieval and narration under pressure, not familiarity with the material.

Each question below includes its own step-by-step walkthrough. Read the framework first, then try the problem before unlocking the full solution.

System Design

System Design Interview

System design interview questions — distributed systems, API design, and scalability — from real interview loops, plus the most common prompts (feed systems, rate limiters, chat, URL shorteners) and the framework to answer any of them.

LiveCodingInterview.com prepares you for the system design interview with questions asked at FAANG and other top tech companies.

Showing 1 of 1 questions