There is a trade-off between the size of the cache and the access time
Small caches are faster but cannot hold much data, therefore having larger miss rate
Large caches can hold enough data but are slower, thus increasing hit time
We can add another, smaller cache (L1) between the current cache (L2) and the CPU
Since everything that is in L1 is also in L2, it only makes sense to add L2 if it is much bigger than L1
The CPU checks L1 first. If L1 misses, it checks L2. If L2 also misses, it accesses main memory.




We now have local and global miss rate
Local miss rate is number of misses in the given cache level divided by number of accesses to this level and it is equal to Miss_RateL1 for L1 and Miss_RateL2 for L2
Global miss rate is the number of misses in the given cache level divided by total number of accesses generated by the CPU. It is equal to Miss_RateL1 for L1, but it is Miss_RateL1*Miss_RateL2 for L2
Assume that in 1000 memory references there are 40 misses in L1 cache and 20 misses in L2 cache. Assume the miss penalty for L2 is 200 clock cycles, hit time of L2 is 10 clock cycles, hit time of L1 is 1 clock cycle and there are 1.5 memory references per instruction. What are the average memory access time and the average memory stalls per instruction?
Given the data below, what is the impact of second-level cache associativity on its miss
penalty: