The main idea is that L1 and L2 have different jobs, so they are intentionally designed differently.
differently.
L1: optimize for speed
L2: optimize for capacity + lower miss rate
That one idea explains most of the differences in your slide.
| Design aspect | Typical L1 design | Typical L2 design | Why? |
|---|
| Size | Small | Larger | L1 must be extremely fast; L2 can trade speed for capacity |
| Number of sets | Fewer | More | Larger cache naturally has more sets |
| Associativity | Lower | Higher | L1 wants fast lookup; L2 wants fewer conflict misses |
| Block size | Small/moderate | Same or sometimes larger | Larger blocks exploit spatial locality but increase miss penalty |
| Split vs unified | Usually split I-cache + D-cache | Usually unified | L1 needs parallel instruction/data access; L2 benefits from flexible capacity |
| Ports | Often multiple/banked | May be fewer/slower | L1 bandwidth matters more |
| Replacement policy | Simple | More sophisticated | L2 can afford extra logic |
| Write policy | Usually D-cache write-back | Usually write-back | Reduces traffic to lower levels |
| Inclusion | Depends on architecture | Inclusive/exclusive/non-inclusive | Determines whether the same block can exist at multiple levels |
1. Cache Size
L1 is normally small because it must be accessed very quickly.
Example:
L1=32KB
L2=512KB
Small cache→faster
Large cache→lower miss rate, but slower
2. Number of Sets
Remember:
#Sets=BlockSize×AssociativityCacheSize
Example:
L1:
32KB, 64B, 4-way
Sets=64×432768=128
L2:
512KB, 64B, 8-way
Sets=64×8524288=1024
So L2 usually has many more sets.
The number of sets affects the address:
IndexBits=log2(Sets)
3. Associativity
Associativity tells us how many cache lines are inside each set.
- Direct mapped = 1-way
- 2-way = 2 cache lines per set
- 4-way = 4 cache lines per set
- 8-way = 8 cache lines per set
Higher associativity usually reduces conflict misses because a memory block has more possible locations inside its set.
However, higher associativity also means:
- more tag comparisons
- more hardware
- potentially longer hit time
So L1 often uses relatively modest associativity because hit time is very important.
L2 can often use higher associativity because reducing misses is more important there.
4. Block Size
A block is the amount of data transferred between memory levels at one time.
Typical examples:
- 32B
- 64B
- 128B
A larger block can help because of spatial locality.
If the CPU accesses one address, it may soon access nearby addresses, so bringing nearby data together can reduce compulsory misses.
However, a larger block also has disadvantages:
- larger miss penalty
- more transfer time
- fewer cache lines for the same cache capacity
- possibly more capacity or conflict misses if the block becomes too large
So block size is a trade-off.
5. Split Cache vs Unified Cache
L1 is commonly split into:
- L1 instruction cache
- L1 data cache
The instruction cache stores instructions.
The data cache stores data used by loads and stores.
The reason for splitting them is bandwidth.
The CPU may need to do both of these in the same cycle:
- fetch an instruction
- access data
With separate caches, both accesses can happen at the same time.
So:
- instruction fetch → L1 I-cache
- load/store → L1 D-cache
L2 is commonly unified.
That means instructions and data share the same L2 capacity.
This allows the space to be used more flexibly.
In general:
- L1 → often split
- L2 → usually unified
- L3 → usually unified
6. Multiple Ports
A cache port is an access path into the cache.
A single-port cache can usually handle only one access at a time.
If a unified cache must handle both:
- instruction fetch
- data access
in the same cycle, one access may have to wait.
This is why some problems say a load/store requires one extra cycle in a single-port unified cache.
A split L1 cache naturally reduces this conflict because instruction and data accesses use different caches.
7. Which Usually Has a Higher Miss Rate: I-cache or D-cache?
Usually:
D-cache miss rate>I-cache miss rate
Instruction accesses often have strong spatial and temporal locality.
For example, instructions are frequently executed sequentially:
- instruction 100
- instruction 101
- instruction 102
- instruction 103
Loops also repeatedly execute the same instructions.
Data accesses can be much more irregular.
Examples:
- pointer chasing
- linked lists
- hash tables
- random array accesses
So D-cache often has a higher miss rate.
This is workload-dependent, though.
8. Write-Through vs Write-Back
This mainly matters for data caches.
Write-Through
When the CPU writes data, both the cache and the next memory level are updated immediately.
Advantages:
- simpler consistency
- memory is always up to date
Disadvantages:
- more memory traffic
- potentially slower writes
Write-Back
The CPU writes only to the cache first.
The cache line is marked dirty.
When that line is later evicted, the modified data is written to the next memory level.
Advantages:
- less lower-level memory traffic
- usually better performance
Disadvantages:
- more complex
- requires dirty-bit management
Modern high-performance data caches commonly use write-back.
9. Replacement Policy
In a set-associative cache, a set may become full.
Suppose a 4-way set already contains four blocks.
If a new block must enter that set, one old block must be removed.
The replacement policy decides which one.
Common policies include:
- LRU
- Pseudo-LRU
- FIFO
- Random
L1 often prefers simpler policies because access speed is critical.
L2 may use more sophisticated policies because reducing misses is more important.
10. Inclusive Cache
In an inclusive hierarchy, every block in L1 must also exist in L2.
Conceptually:
L1⊆L2
Example:
L1 contains:
- A
- B
- C
L2 contains:
- A
- B
- C
- D
- E
Advantages:
- can simplify some coherence and tracking mechanisms
Disadvantages:
- duplicate copies use cache capacity
11. Exclusive Cache
In an exclusive hierarchy, the same block is generally kept in only one cache level at a time.
Example:
L1 contains:
- A
- B
- C
L2 contains:
- D
- E
- F
- G
Advantage:
- better effective total cache capacity
Disadvantage:
- more complicated movement of blocks between levels
12. Non-Inclusive / Non-Exclusive
Some processors use neither a strictly inclusive nor strictly exclusive policy.
A block may exist in multiple cache levels, but it is not required to.
This gives the cache hierarchy more flexibility.
13. L3 Cache
Modern processors often also include L3 cache.
The hierarchy becomes:
CPU → L1 → L2 → L3 → DRAM
Typical trend:
- L1: smallest and fastest
- L2: larger and slower
- L3: much larger and slower
- DRAM: largest but much slower
The main purpose of L3 is to reduce expensive DRAM accesses.
In multicore processors, L3 is often shared by multiple cores.
L1 vs L2 Summary
| Feature | L1 | L2 |
|---|
| Size | Small | Larger |
| Speed | Fastest | Slower |
| Main goal | Low hit time | Lower miss rate and larger capacity |
| Associativity | Usually lower | Usually higher |
| Split / Unified | Often split | Usually unified |
| Ports / bandwidth | Very important | Less latency-critical than L1 |
| Replacement policy | Usually simpler | Can be more sophisticated |
| Write policy | D-cache often write-back | Often write-back |
| Miss rate | Usually higher than L2 | Lower than L1 |
| Latency | Very low | Higher |
The main design principle is:
L1: optimize for speed
L2/L3: optimize more for capacity and miss-rate reduction