Why does a processor keep a tiny memory right next to itself?
From 1986 to 2000, processors got 55 percent faster every year. Memory improved by only 10. Something had to give.
▶ Start the storyMoore's law kept packing more transistors onto chips, so processors got faster almost every year. Main memory did not keep up. From 1986 to 2000, CPU speed improved at about 55 percent a year, while the response time of memory sitting off the chip improved at only about 10 percent. That growing gap is called the memory wall, and it means a processor can spend much of its time simply waiting for memory to answer.
CPU speed vs. memory speed, annual growth (1986-2000)
% per year
| Annual improvement | |
|---|---|
| CPU speed | 55 % per year |
| Off-chip memory speed | 10 % per year |
The fix is a cache: a small, fast memory placed right next to the processor core, holding copies of data the processor has used recently. Checking the cache first means the processor often avoids a trip to main memory, which can be tens to hundreds of times slower to reach. A modern chip usually layers several of these caches: L1, closest to the core and fastest; L2, a bit farther away and slower but bigger; and often L3, shared across multiple cores.
Caches work because of a pattern called locality of reference: a processor tends to use the same memory locations again soon (temporal locality), and it tends to use memory locations near each other (spatial locality). Guess right about what's about to be needed, and the cache already has it ready.
It isn't a free upgrade. Cache memory is built from SRAM, which needs four or six transistors to hold a single bit; the DRAM of main memory needs just one transistor and a capacitor. Back in the 386 era, main memory could have latencies up to 120 nanoseconds, while SRAM cache ran at roughly 10 to 25 nanoseconds. That speed costs chip space, which is exactly why caches stay small while memory stays big.
Quiz me
0/3
Recap
Caches work because of locality of reference: processors tend to reuse the same memory locations soon (temporal locality) and ones nearby (spatial locality), so a small fast cache can usually guess right about what's needed next.
Surprising fact · From 1986 to 2000, CPU speed improved about 55 percent a year while off-chip memory improved only about 10 percent a year.
Connects to
- 🗄️ Why do programs live in the same memory as the data they work on?
- 📈 Why did computer chips double in power every two years, and can it last?
- 🔁 What does a processor actually do billions of times every second?
- 👻 How could a processor's habit of guessing ahead leak your secrets?
- 🧮 Why can you only keep a few things in mind at once?
Sources (5)
No source, no claim. Every fact in this lesson (14 claims) cites at least one of these.