Reference · quick sheet

The Scaling Ladder

Take the cheapest rung that clears your runway requirement. Every rung you climb permanently removes something you used to get for free.

#MoveBuysCosts you, permanently
1Index it; fix the queryOften an order of magnitudeWrite cost per index
2Bigger machineEverything, brieflyMoney, and a ceiling you will hit
3Read replicasRead capacity, survivability Read-your-writes — see staleness
4CachingAbsorbs repeated reads Invalidation, herds, stale sets — see cache misses
5Vertical partitioningRunway, quickly and cheaply Joins and transactions across partitions
6Horizontal shardingWrite capacity, no ceiling Months of work, a routing layer you own, and your database's guarantees

Vertical vs horizontal

Vertical moves whole tables onto separate databases — each table still lives in one place, and SQL inside a partition is untouched. Horizontal splits one table's rows across machines, which is what takes joins, transactions, unique constraints and global ordering away from you.

When each rung is exhausted

  • Replicas: when writes are the bottleneck. Replication copies; it never divides write load.
  • Caching: when the working set is not repeatedly read — check whether the OS page cache is already holding it for free, as Discord found.
  • Vertical partitioning: when a single table no longer fits one machine. Objective, not a feeling. Figma's largest were several terabytes and billions of rows.

The timing question

Figma deferred and bought runway deliberately. Notion's first lesson learned was "Shard earlier", because they arrived in an emergency with VACUUM stalling. Both are right: defer the migration, never defer the decision. Ask what will force this, and how far away is it? — not should we shard yet?