Connect physical design to workload and reliability
Storage layout, distribution, and replication solve different problems. This chapter helps choose among them using access patterns, capacity, consistency, and recovery requirements.
Suggested route: Start with the decision map, then compare the engine-specific meanings of clustering. Study sharding and replicas as separate architectural choices.
By the end: Explain which bottleneck a change addresses, what it cannot fix, and which operational trade-offs it adds.
Decision reference
| Requirement | Investigate |
|---|---|
| Selective row access | Indexes |
| Time pruning and lifecycle | Table partitioning |
| Related values stored together | Engine-specific clustering |
| Distributed capacity | Sharding |
| Read isolation or failover | Replicas and their consistency contract |
In this chapter
- Choose the right storage or scaling changeSeparate query access, physical layout, capacity, and availability requirements.
- Clustering means different thingsDistinguish BigQuery storage blocks, clustered indexes, PostgreSQL CLUSTER, and database clusters.
- Sharding and distribution keysChoose where data lives while accounting for routing, hot keys, joins, and growth.
- Replicas, freshness, and failoverSeparate read capacity from availability and define what a user may observe after a write.