Design L1, L2 caches

Seungyun Lee·4일 전

Computer Arch (Memory)

목록 보기
12/16

The main idea is that L1 and L2 have different jobs, so they are intentionally designed differently.

differently.
L1: optimize for speed\boxed{\text{L1: optimize for speed}}

L2: optimize for capacity + lower miss rate\boxed{\text{L2: optimize for capacity + lower miss rate}}

That one idea explains most of the differences in your slide.

Design aspectTypical L1 designTypical L2 designWhy?
SizeSmallLargerL1 must be extremely fast; L2 can trade speed for capacity
Number of setsFewerMoreLarger cache naturally has more sets
AssociativityLowerHigherL1 wants fast lookup; L2 wants fewer conflict misses
Block sizeSmall/moderateSame or sometimes largerLarger blocks exploit spatial locality but increase miss penalty
Split vs unifiedUsually split I-cache + D-cacheUsually unifiedL1 needs parallel instruction/data access; L2 benefits from flexible capacity
PortsOften multiple/bankedMay be fewer/slowerL1 bandwidth matters more
Replacement policySimpleMore sophisticatedL2 can afford extra logic
Write policyUsually D-cache write-backUsually write-backReduces traffic to lower levels
InclusionDepends on architectureInclusive/exclusive/non-inclusiveDetermines whether the same block can exist at multiple levels

1. Cache Size

L1 is normally small because it must be accessed very quickly.
Example:
L1=32KBL1=32KB
L2=512KBL2=512KB

Small cache→faster\text{Small cache} \rightarrow \text{faster}

Large cache→lower miss rate, but slower\text{Large cache} \rightarrow \text{lower miss rate, but slower}

2. Number of Sets

Remember:

#Sets=CacheSizeBlockSize×Associativity\boxed{\#Sets= \frac{CacheSize}{BlockSize\times Associativity}}

Example:
L1:

32KB, 64B, 4-way32KB,\ 64B,\ 4\text{-way}

Sets=3276864×4=128Sets= \frac{32768}{64\times4} = 128

L2:

512KB, 64B, 8-way512KB,\ 64B,\ 8\text{-way}

Sets=52428864×8=1024Sets= \frac{524288}{64\times8} = 1024

So L2 usually has many more sets.
The number of sets affects the address:

IndexBits=log⁡2(Sets)\boxed{IndexBits=\log_2(Sets)}

3. Associativity

Associativity tells us how many cache lines are inside each set.

  • Direct mapped = 1-way
  • 2-way = 2 cache lines per set
  • 4-way = 4 cache lines per set
  • 8-way = 8 cache lines per set

Higher associativity usually reduces conflict misses because a memory block has more possible locations inside its set.

However, higher associativity also means:

  • more tag comparisons
  • more hardware
  • potentially longer hit time
    So L1 often uses relatively modest associativity because hit time is very important.
    L2 can often use higher associativity because reducing misses is more important there.

4. Block Size

A block is the amount of data transferred between memory levels at one time.
Typical examples:

  • 32B
  • 64B
  • 128B
    A larger block can help because of spatial locality.
    If the CPU accesses one address, it may soon access nearby addresses, so bringing nearby data together can reduce compulsory misses.
    However, a larger block also has disadvantages:
  • larger miss penalty
  • more transfer time
  • fewer cache lines for the same cache capacity
  • possibly more capacity or conflict misses if the block becomes too large
    So block size is a trade-off.

5. Split Cache vs Unified Cache

L1 is commonly split into:

  • L1 instruction cache
  • L1 data cache

The instruction cache stores instructions.
The data cache stores data used by loads and stores.
The reason for splitting them is bandwidth.

The CPU may need to do both of these in the same cycle:

  • fetch an instruction
  • access data
    With separate caches, both accesses can happen at the same time.
    So:
  • instruction fetch → L1 I-cache
  • load/store → L1 D-cache
    L2 is commonly unified.
    That means instructions and data share the same L2 capacity.
    This allows the space to be used more flexibly.
    In general:
  • L1 → often split
  • L2 → usually unified
  • L3 → usually unified

6. Multiple Ports

A cache port is an access path into the cache.
A single-port cache can usually handle only one access at a time.
If a unified cache must handle both:

  • instruction fetch
  • data access
    in the same cycle, one access may have to wait.
    This is why some problems say a load/store requires one extra cycle in a single-port unified cache.
    A split L1 cache naturally reduces this conflict because instruction and data accesses use different caches.

7. Which Usually Has a Higher Miss Rate: I-cache or D-cache?

Usually:

D-cache miss rate>I-cache miss rate\text{D-cache miss rate} > \text{I-cache miss rate}

Instruction accesses often have strong spatial and temporal locality.
For example, instructions are frequently executed sequentially:

  • instruction 100
  • instruction 101
  • instruction 102
  • instruction 103
    Loops also repeatedly execute the same instructions.
    Data accesses can be much more irregular.
    Examples:
  • pointer chasing
  • linked lists
  • hash tables
  • random array accesses
    So D-cache often has a higher miss rate.
    This is workload-dependent, though.

8. Write-Through vs Write-Back

This mainly matters for data caches.
Write-Through
When the CPU writes data, both the cache and the next memory level are updated immediately.
Advantages:

  • simpler consistency
  • memory is always up to date
    Disadvantages:
  • more memory traffic
  • potentially slower writes
    Write-Back
    The CPU writes only to the cache first.
    The cache line is marked dirty.
    When that line is later evicted, the modified data is written to the next memory level.
    Advantages:
  • less lower-level memory traffic
  • usually better performance
    Disadvantages:
  • more complex
  • requires dirty-bit management
    Modern high-performance data caches commonly use write-back.

9. Replacement Policy

In a set-associative cache, a set may become full.
Suppose a 4-way set already contains four blocks.
If a new block must enter that set, one old block must be removed.
The replacement policy decides which one.
Common policies include:

  • LRU
  • Pseudo-LRU
  • FIFO
  • Random
    L1 often prefers simpler policies because access speed is critical.
    L2 may use more sophisticated policies because reducing misses is more important.

10. Inclusive Cache

In an inclusive hierarchy, every block in L1 must also exist in L2.
Conceptually:

L1⊆L2L1 \subseteq L2

Example:
L1 contains:
- A
- B
- C
L2 contains:
- A
- B
- C
- D
- E

Advantages:

  • can simplify some coherence and tracking mechanisms
    Disadvantages:
  • duplicate copies use cache capacity

11. Exclusive Cache

In an exclusive hierarchy, the same block is generally kept in only one cache level at a time.

Example:
L1 contains:
- A
- B
- C
L2 contains:
- D
- E
- F
- G

Advantage:

  • better effective total cache capacity
    Disadvantage:
  • more complicated movement of blocks between levels

12. Non-Inclusive / Non-Exclusive

Some processors use neither a strictly inclusive nor strictly exclusive policy.
A block may exist in multiple cache levels, but it is not required to.
This gives the cache hierarchy more flexibility.

13. L3 Cache

Modern processors often also include L3 cache.
The hierarchy becomes:
CPU → L1 → L2 → L3 → DRAM
Typical trend:

  • L1: smallest and fastest
  • L2: larger and slower
  • L3: much larger and slower
  • DRAM: largest but much slower
    The main purpose of L3 is to reduce expensive DRAM accesses.
    In multicore processors, L3 is often shared by multiple cores.

L1 vs L2 Summary

FeatureL1L2
SizeSmallLarger
SpeedFastestSlower
Main goalLow hit timeLower miss rate and larger capacity
AssociativityUsually lowerUsually higher
Split / UnifiedOften splitUsually unified
Ports / bandwidthVery importantLess latency-critical than L1
Replacement policyUsually simplerCan be more sophisticated
Write policyD-cache often write-backOften write-back
Miss rateUsually higher than L2Lower than L1
LatencyVery lowHigher

The main design principle is:

L1: optimize for speed\text{L1: optimize for speed}

L2/L3: optimize more for capacity and miss-rate reduction\text{L2/L3: optimize more for capacity and miss-rate reduction}

profile
Design Verification engineer

0개의 댓글