Caching strategies
Cache-aside, read-through, write-through, write-behind. Pick one, run traffic through it, and see the read path light up - and the database load drop.
9 min read
A cache is a small, fast store in front of a slow one. The idea takes one sentence. Everything interesting about caching lives in the second-order effects: what happens on a miss, what happens on a write, and what happens the moment the cache is wrong or gone.
Going from 90% to 95% hit rate halves your database load
The hit rate is the share of requests the cache answers. It is the number that decides whether the database behind it is comfortable or on fire, and it does not behave the way the percentage suggests. What reaches the database is the miss rate, and going from 10% misses to 5% is a halving.
Improvements at the top of the range are worth far more than they look on a dashboard: 99% to 99.5% halves the backend load again. It runs the other way too, which is why a small dip in hit rate during a deploy can double database load with nothing else changing.
That makes a cache load-bearing in the literal sense. The database is provisioned for the miss traffic, not the real traffic, so the day the cache disappears is the day you find out what you were actually running.
Four strategies, four different things to be wrong about
Every strategy reads roughly the same way: try the cache, fall back to the database on a miss. They differ in who fetches on a miss, and above all in what happens on a write.
Cache-aside is the default because it is explicit, and the application stays in control. Read-through moves the miss handling into the cache. Write-through keeps the cache and database in lockstep at the cost of slower writes.
Write-behind is the one to choose deliberately. Its risk is not stale reads, because reads go to the cache, which has the newest value. The risk is that you have told the user “saved” while the only copy lives in one process's memory. That makes it excellent for view counters and terrible for payments.
Three ways a cache takes down what it protects
Caches fail in ways specific to caches, and all three failures share a shape: the database suddenly receives traffic it was never sized for.
A stampede is a hot key expiring while thousands of requests want it. Nothing coordinates them, so they all miss together. Penetration is requests for keys that do not exist, which are never cached, so each one costs a query; a bloom filter (see the bloom filter lab) turns them into free rejections. An avalanche is many keys expiring at once, or a cache node restarting empty.
Jittered TTLs fix avalanches for the same reason jitter fixes retry storms in the retries lab: synchronised clients are the real problem in both, and randomness breaks the synchrony. What changes when there is more than one cache is the subject of distributed caching.
The short version
- The database sees the miss rate, so 90% to 95% hits halves its load.
- Strategies differ mainly on writes: explicit, inline, in lockstep, or deferred.
- Write-behind is fast and can lose acknowledged writes.
- Stampedes, penetration and avalanches all jump the miss rate to 100%; locks, bloom filters and jitter stop them.