backenddrills

System design and architecture

System design: workload, consistency and recovery before diagrams

By the BackendDrills editorial team · Published and checked October 6, 2026 · 8-minute read

Before you start: API, database and failure-handling fundamentals. Find this reading in a study path →

Start a design with measurable requirements and a write owner. A diagram becomes useful when its boundaries explain capacity, consistency and recovery. These numbers are an original exercise, not a production benchmark.

Translate the workload into downstream demand

A dashboard expects 120 page requests per second. Each page requests six uncached summaries. That is up to 720 summary calls per second before retries and background refresh. If a summary call occupies a connection for an average 80 milliseconds, the simple steady-state estimate is about 58 concurrently occupied connections. This calculation assumes stable throughput and duration; tails, bursts and uneven tenants require additional investigation.

downstream calls/s = page requests/s × calls/page
average in-flight work ≈ throughput × average duration
Test separately: bursts, slow dependencies, retries and tenant skew.

Allocate demand across all replicas, not just one instance. State the queue bound, admission policy and customer behavior when capacity is exhausted. An unbounded queue converts rejected demand into growing latency and memory use. A bounded test should report completed business operations, latency distributions and correctness alongside resource saturation.

Choose where stale data is acceptable

A reporting view may tolerate a short delay after a write. A purchase confirmation may need to reflect the authoritative operation immediately. CQRS can separate read and write models, but introduces synchronization and consistency choices when those models diverge. Define the expected delay, what the client sees during it and what detects a stuck projection. Avoid routing a correctness-sensitive decision through a stale read merely because it is fast.

Make recovery part of the proposal

Imagine restoring the primary database while a consumer checkpoint remains ahead of the restored data. Replaying only from that checkpoint could leave missing derived records. Your design needs a restore and reconciliation procedure that relates source versions, consumers and external effects. A successful backup job is only one observation; restore a fixture and verify useful business state.

Write a one-page decision record

  1. State workload ranges, latency target and what must remain correct.
  2. Identify write ownership, authoritative state and permitted read staleness.
  3. Compare two plausible designs, including operational cost.
  4. Name overload behavior, failure domains and a rollback boundary.
  5. Define a capacity experiment and a restore acceptance test.
  6. Record which changed assumption would trigger a redesign.

For a lead or architect, review the proposal by challenging one assumption rather than adding more boxes. For an engineer, implement the smallest fixture that could disprove that assumption. Both roles should be able to explain the same invariant and recovery gate.

Check the underlying behavior

Original illustrative examples, prepared with AI assistance and checked against the linked primary documentation. No customer incident or vendor endorsement is claimed. Our editorial approach.

Practice the next decision

Try a complete free backend drill: inspect evidence, make three decisions and review the reasoning. No account or card required.

Try a free incident drill →

Explore 1008 scenarios · $28 one-time

Selected practice from the study paths

Free readings need no account. Full edition drills require verified access; opening a paid link does not expose its answers.

Recommended next readings