Reference · quick sheet
The Scaling Ladder
Take the cheapest rung that clears your runway requirement. Every rung you climb permanently removes something you used to get for free.
| # | Move | Buys | Costs you, permanently |
|---|---|---|---|
| 1 | Index it; fix the query | Often an order of magnitude | Write cost per index |
| 2 | Bigger machine | Everything, briefly | Money, and a ceiling you will hit |
| 3 | Read replicas | Read capacity, survivability | Read-your-writes — see staleness |
| 4 | Caching | Absorbs repeated reads | Invalidation, herds, stale sets — see cache misses |
| 5 | Vertical partitioning | Runway, quickly and cheaply | Joins and transactions across partitions |
| 6 | Horizontal sharding | Write capacity, no ceiling | Months of work, a routing layer you own, and your database's guarantees |
Vertical vs horizontal
Vertical moves whole tables onto separate databases — each table still lives in one place, and SQL inside a partition is untouched. Horizontal splits one table's rows across machines, which is what takes joins, transactions, unique constraints and global ordering away from you.
When each rung is exhausted
- Replicas: when writes are the bottleneck. Replication copies; it never divides write load.
- Caching: when the working set is not repeatedly read — check whether the OS page cache is already holding it for free, as Discord found.
- Vertical partitioning: when a single table no longer fits one machine. Objective, not a feeling. Figma's largest were several terabytes and billions of rows.
The timing question
Figma deferred and bought runway deliberately. Notion's first lesson learned was "Shard earlier", because
they arrived in an emergency with VACUUM stalling. Both are right:
defer the migration, never defer the decision. Ask what will force this, and how far away
is it? — not should we shard yet?