SystemDesign.io
Track 1 · Module 1.3 · Architectural Foundations

Capacity Estimation &
Back-of-the-Envelope Math

Learn the 5 simple formulas that let you estimate how much traffic, storage, memory, and bandwidth any system needs — the most asked skill in system design interviews.

⏱️ Estimated: 15 min
📊 Difficulty: Beginner
🎯 You'll learn: 5 Sizing Formulas + Live Calculator
Your Progress5 formulas

Why Capacity Estimation Matters

🎯 The Interview Scenario

Imagine you're in a system design interview and the interviewer says:

"Design Twitter. How many servers would you need? How much storage? How fast should the network be?"

That's capacity estimation — a quick, rough calculation done on a whiteboard without a calculator. You don't need to be perfectly accurate. You just need to be in the right ballpark (within 2-3× of the real answer).

💡
Beginner Tip: Think of it like estimating grocery costs before checkout. You don't calculate to the penny — you round to the nearest dollar. Same idea here, but with servers and gigabytes!

By the end of this module, you'll know 5 simple formulas that cover every capacity question:

① QPS② Peak QPS③ Storage④ RAM Cache⑤ Bandwidth
📋 Quick Reference Cheatsheet (expand when you need it)

The 86,400 Seconds Shortcut

Daily Volume÷ 86,400 ≈ QPS
1 Million / day~10 QPS
10 Million / day~100 QPS
100 Million / day~1,000 QPS
1 Billion / day~10,000 QPS

Trick: 86,400 ≈ 100,000 (10⁵). Just drop 5 zeros from daily volume!

Powers of Two & Storage Units

PowerApprox.Unit
2¹⁰~1,0001 KB
2²⁰~1 Million1 MB
2³⁰~1 Billion1 GB
2⁴⁰~1 Trillion1 TB
2⁵⁰~1 Quadrillion1 PB

Latency Numbers to Know

⚡ CPU & RAM
0.5 ns — 100 ns

L1 Cache ref = 0.5 ns · Main memory = 100 ns

💿 SSD
100 µs — 1 ms

SSD random read ≈ 100 µs

🌐 Disk & Network
10 ms — 150 ms

HDD seek = 10 ms · Cross-region = 150 ms


Formula 1 — Queries Per Second (QPS)

🚦 How Much Traffic Can Your System Handle?

QPS tells you how many requests your servers process every single second. It's the most fundamental number in system design — like knowing a highway's lane capacity before building on-ramps.

QPS = Daily Requests ÷ 86,400 seconds// 86,400 = seconds in a day (24 × 60 × 60)
💡
Memory trick: 86,400 is close to 100,000. So in your head, just drop 5 zeros from the daily number.
100 Million requests/day → drop 5 zeros → ~1,000 QPS. Done!

Real Example: Twitter's Read Traffic

Step 1 · Volume300M users × 50 reads/day
15 Billion reads/day
Step 2 · QPS Math15,000,000,000 reads ÷ 86,400 sec
≈ 173,600 Read QPS
✅ Quick Check
An app gets 432 Million requests per day. What's the approximate QPS?

Formula 2 — Peak QPS

📈 Planning for the Traffic Surge

Average QPS is like average daily traffic on a road. But what about rush hour? On Black Friday, a flash sale, or when a celebrity tweets — traffic can spike 2× to 5× above average. Your system must survive the peak, not just the average.

Peak QPS = Average QPS × Peak Multiplier (2× to 5×)// 2× for normal apps · 3× for social media · 5× for flash sales/breaking news
💡
Rule of thumb: When in doubt, use as your peak multiplier. It's the industry standard for most social media and e-commerce apps. Mention "we'll design for 2× peak" in interviews — interviewers love that.

Real Example: Twitter Peak

Peak Surge173,600 Average QPS × Peak Multiplier
≈ 347,200 Peak QPS

This means your auto-scaling group must handle ~350K requests/second during the busiest moment of the day.


Formula 3 — Storage Capacity

💾 How Much Disk Space Do You Need?

Every time a user posts a tweet, uploads a photo, or sends a message, data gets written to disk. The question is: how much disk do you need over 5 years? (5 years is the standard retention window interviewers expect.)

Daily Storage = Daily Writes × Average Payload Size
5-Year Storage = Daily Storage × 365 × 5 × Replication Factor// Replication Factor = 3× (standard) — data is stored on 3 separate machines for safety
💡
Why 3× replication? If one server's disk dies, the other 2 copies keep your data safe. It's like keeping 3 copies of your important files — one on your laptop, one on Google Drive, one on a USB. So your raw storage triples!

Real Example: Twitter Storage

What's in a tweet? ~200 bytes text + ~100 bytes metadata (user ID, timestamp) = ~300 bytes. About 20% of tweets have an image (avg ~200 KB each).

Step 1 · Text100M tweets/day × 300 Bytes
30 GB / day
Step 2 · Media20M images/day (20%) × 200 KB
4 TB / day
Step 3 · 5-Yr Storage4.03 TB/day × 365 × 5 yrs × 3× replicas
≈ 22 Petabytes
✅ Quick Check
A service writes 1 TB of new data daily. With 3× replication and 5-year retention, how much raw disk is needed?

Formula 4 — Memory (RAM) Caching

🧠 The 80/20 Rule for Caching

Here's a powerful insight: 20% of your data gets 80% of the traffic. Think about it — on Twitter, most people view the same trending tweets, not random tweets from 2019.

So if you cache just the hottest 20% of daily reads in RAM (using Redis or Memcached), you'll serve ~80% of all requests from super-fast memory instead of slow disk. This is called the Pareto Principle or 80/20 rule.

Daily Read Volume = Daily Reads × Average Read Size
RAM Cache Needed = Daily Read Volume × 20%// This 20% working set satisfies ~80% of all read requests from memory!

Real Example: Twitter Cache

Step 1 · Daily Reads15 Billion reads/day × 300 Bytes
4.5 TB / day
Step 2 · 80/20 Cache4.5 TB × 20% hot working set
900 GB RAM

900 GB fits in just 4 Redis nodes (256 GB RAM each). That's a tiny cluster to serve 80% of Twitter's reads!

✅ Quick Check
Users generate 2 TB of read traffic per day. Using the 80/20 rule, how much RAM do you need for the cache?

Formula 5 — Network Bandwidth

🌐 How Fat Should Your Network Pipe Be?

There's one common gotcha here: storage is measured in Bytes, but network speed is measured in Bits. Since 1 Byte = 8 Bits, you need to multiply by 8 when converting.

Bandwidth = QPS × Payload Size × 8 bits/Byte// Ingress = incoming data (writes) · Egress = outgoing data (reads)
💡
Don't forget the ×8! This is the #1 mistake in bandwidth calculations. Your ISP says "1 Gbps" — that's Giga-bits, not Giga-bytes. Actual file transfer speed is ~125 MB/s.

Real Example: Twitter Egress

Step 1 · Byte Speed173,600 read QPS × 300 Bytes
~52.1 MB/s
Step 2 · Bits Conversion52.1 MB/s × 8 bits/Byte
≈ 417 Mbps Egress

If we include media (images/videos), the actual bandwidth is much higher — but the formula stays the same. Just use a larger payload size.

✅ Quick Check
A system handles 10,000 QPS with 50 KB average response payload. What's the outbound bandwidth?

🧮 Interactive System Sizer — Try It Yourself!

Live Capacity Estimator

Adjust the sliders or pick a preset to see real-time results.

Daily Active Users (DAU)300 Million
Reads / Views per User / Day50 reads
Writes / Posts per User / Day1 write
Avg Write Payload Size30 KB
Avg Read Payload Size10 KB
Average QPS
R: — · W: —
Peak QPS
Design target for scaling
Daily Storage Ingress
5-Year Retention
With 3× replication
RAM (80/20 Cache)
20% daily read working set
Outbound Bandwidth
Network egress pipe
Estimated Infrastructure
— Web Servers(At 5,000 Peak QPS / Node)
— Redis Nodes(At 256 GB RAM / Node)

🏗️ Practice: Right-Size a Scaled Web Application

Scenario: You just calculated the capacity for an application handling 50,000 Peak QPS, 20% hot data in RAM, and 5-year persistent storage.
Wire the architecture: Connect incoming Users to the Load Balancer, distribute traffic to the App Server Fleet, connect servers to Redis Cache (for fast RAM reads), and persist writes to the Database.

Interactive CanvasMode: Freeform Drag & Connect

Task: Right-Sizing Infrastructure Tiers

🎯
Scenario Prompt: Connect Users ➔ Load Balancer ➔ App Servers, then wire App Servers to both Redis Cache (for 80/20 RAM) and the Database (for 5-Year storage).
Available Components (Click to place on canvas):
🖱️Canvas is emptyClick components above to place them onto the canvas.
Drag components to arrange them freely, and click two nodes to connect them.
💡 Drag to move • Click Node A ➔ Node B to connect • Click ✕ to delete