Tutorials › Core Cloud Architecture › Caching

Core Cloud Architecture · Part 7 of 13

Caching

Every cache trades some staleness for speed. The design question is how much, and where it's tolerable.

A cache keeps a copy of frequently accessed data somewhere faster to read from than its original source, usually in memory instead of a database's disk-backed storage. The speed gain is large and cheap to get. The cost, easy to underweight while designing one, is that a cache is a second copy of the truth, and two copies of the truth can disagree.

Cache-aside

In the cache-aside pattern (also called lazy loading), the application checks the cache first; on a hit, it uses the cached value directly. On a miss, it reads from the database, then writes that value into the cache before returning it, so the next request for the same data is a hit. This is the most common pattern because it's simple to reason about and only caches data that's being requested; nothing is pre-loaded speculatively.

Write-through

In write-through caching, every write goes to the cache and the database together, as one operation, so the cache is never behind the database for data that's been written this way. It trades a small amount of write latency (every write now touches two systems instead of one) for a much stronger consistency guarantee than cache-aside provides on its own.

TTL and invalidation

A TTL (time to live) is an expiration set on a cached value; once it elapses, the cache treats the entry as gone, forcing the next read to go back to the source and refresh it. It's the simplest way to bound how stale a cache-aside value can get, without ever explicitly telling the cache "this changed." Invalidation is the more precise, and more failure-prone, alternative: explicitly removing or updating a cache entry the moment its underlying data changes. It keeps the cache fresher than a TTL alone, but only if every code path that changes the data remembers to invalidate the right key. Miss one, and that entry is silently wrong until its TTL, if it has one at all, catches up.

Hot keys

A hot key is a single cache entry receiving a disproportionate share of all traffic: a trending post, a widely shared link, a piece of global configuration read on every request. Even a cache built to handle enormous aggregate load can struggle when nearly all of that load lands on one key, since most caching systems shard data across multiple nodes by key, and a hot key means one shard doing far more work than the rest. Mitigations include replicating the hot value across multiple cache nodes deliberately, or adding a low-TTL layer of local (in-process) caching in front of the shared cache for the handful of keys that need it.

Compare: managed caching services

AWSGCPAzure
ElastiCache (Valkey / Redis OSS / Memcached)Memorystore (Valkey / Redis / Memcached)Azure Managed Redis (replacing Azure Cache for Redis)

What caching costs: consistency risk

Caching's latency benefit is close to automatic: reading from memory is reliably faster than reading from a database, almost regardless of how the cache is designed. The cost that needs design attention is consistency: the moment a cache exists, there are two answers to "what is this value right now," and they can diverge.

A stale read happens when the underlying data has changed but the cache hasn't caught up: a user updates their profile, reloads the page, and briefly sees the old value, because the read hit a cache entry that hasn't expired or been invalidated. Cache/DB divergence is the more serious version of the same problem: a bug in invalidation logic, a race between a write and a cache refresh, or a cache node that fell out of sync means the cache keeps serving a wrong value indefinitely.

Cache what tolerates being briefly wrong. A product listing's view count can be a few seconds stale with no consequence. An account balance or an inventory count being briefly wrong can mean someone is shown a number they act on incorrectly. Before caching a piece of data, ask what happens if a reader sees a slightly outdated version of it. The answer should be "nothing that matters." Caching makes almost anything faster, so speed alone doesn't tell you what to cache.