Some answers are expensive to produce: a complex database query, a call to a slow third-party API, a rendered page. If the same answer is requested over and over, recomputing it every time is wasteful — and it puts the slow component under load it didn't need.
A cache keeps a copy of the answer somewhere fast (usually memory) so that the next time someone asks the same question, you can hand back the saved copy instead of doing the work again.
The problem: doing the same expensive work twice
Imagine a product page that runs a heavy database query on every view. The data barely changes minute to minute, yet each visitor triggers the full query. As traffic grows, the database becomes the bottleneck — even though almost every request is asking for the same thing. A database that comfortably handles a few hundred queries a second will drown at a thousand page views a second, and no amount of tuning changes the fact that it's answering the same question over and over.
The insight: reads usually outnumber writes, and the same few items tend to be requested far more than the rest (hot keys). That repetition is exactly what a cache exploits.
How it works
A cache sits between the application and the slow source (a database, an API). On each read, the app checks the cache first:
- Cache hit — the answer is already there. Return it immediately. The database is never touched.
- Cache miss — the answer isn't there. Fetch it from the source, store it in the cache, then return it. The next identical request will be a hit.
When the app does this itself, as in the code below, the pattern is called cache-aside. (Read-through is the same flow handled by the cache layer, so the app only ever talks to the cache.) The first request for an item pays the full cost; everyone after rides for nearly free, until the entry expires.
Step through the first two requests for one product. Before the second one runs, predict whether it hits and what it costs. Then, at the busy hour, flip to No cache to replay the same traffic without one.
Why it's so much faster: an in-memory cache like Redis or Memcached usually answers in well under a millisecond, while a query that joins tables or touches disk can take tens of milliseconds. (A cache inside the app's own process is faster still: microseconds.) If 95% of reads are hits, you've removed 95% of the read load from your database and made most requests dramatically faster at the same time.
Your database tops out at about 400 queries a second. You put a cache in front of it, traffic is 2,000 reads a second, and 90% of reads are hits. What load does the database see?
Keeping the cache fresh
A cached copy is a snapshot. The moment the underlying data changes, the cache can be stale — serving an old answer. Two mechanisms keep this under control:
- TTL (time to live) — each entry is stamped with an expiry. After the TTL elapses, the entry is dropped and the next read is a miss, refreshing it from the source. A short TTL means fresher data but more misses; a long TTL means fewer misses but more staleness.
- Eviction — memory is finite, so when the cache fills up it must drop something. LRU (least-recently-used) is the common choice: evict whatever hasn't been touched in the longest time, keeping the hot keys resident.
Step through a price change on a cached product. You'll predict what a shopper sees ten seconds after the price drops, then flip to Delete on write to see invalidation close the stale window. The last step shows a problem neither approach solves on its own: a hot key that vanishes from the cache.
Cache invalidation is famously hard. Deleting on every write sounds airtight, but there are gaps: a write path that forgets to delete, a second service that writes to the same table, or a slow read that loads the old value just before the write and saves it to the cache just after the delete. Keep a sensible TTL as a safety net even when you invalidate explicitly, so any stale entry eventually ages out. As the saying goes, there are only two hard things in computer science: cache invalidation and naming things.
The trade-offs
Caching is not free:
- Staleness — you accept that reads may be slightly out of date. Fine for a product description; dangerous for an account balance.
- Cold starts — an empty cache (after a deploy or restart) sends a burst of misses straight to the database. If your database only survives because of the cache, a cold start can take it down.
- Stampedes — when a hot key expires or is deleted, every request in the reload gap misses at once. Common defences: let one request reload while the others wait for it (often called request coalescing or single flight), keep serving the old value while one request refreshes it in the background, and add a little random jitter to TTLs so many keys don't expire in the same second.
- Complexity — another moving part to run, monitor, and reason about, plus the invalidation logic.
- Memory cost — fast storage isn't free; you cache the valuable subset, not everything.
Every 60 seconds, right as a very popular product's cache entry expires, your database CPU spikes. What's the best fix?
When to reach for it
Caching pays off when reads vastly outnumber writes, when the same items are requested repeatedly, and when slightly stale data is acceptable. It's one of the highest-leverage performance tools available — often a few lines in front of a hot query.
It pairs naturally with load balancing: the load balancer spreads requests across servers, and a shared cache keeps each of those servers from re-doing the same expensive work.