You are building a computer system around a processor with in-order execution that runs at 1 GHz and has a CPI of 1, excluding memory accesses. The only instructions that read or write data from/to memory are loads (20% of all instructions) and stores (5% of all instructions).
The memory system for this computer has a split L1 cache. Both the I-cache and the D-cache are direct mapped and hold 32 KB each. The I-cache has a 2% miss rate and 64 byte blocks, and the D-cache is a write-through, no-write-allocate cache with a 5% miss rate and 64 byte blocks. The hit time for both the Icache and the D-cache is 1ns. The L1 cache has a write buffer. 95% of writes to L1 find a free entry in the write buffer immediately. The other 5% of the writes have to wait until an entry frees up in the write buffer (assume that such writes arrive just as the write buffer initiates a request to L2 to free up its entry and the entry is not freed up until the L2 is done with the request). The processor is stalled on a write until a free write buffer entry is available.
The L2 cache is a unified write-back, write-allocate cache with a total size of 512 KB and a block size of 64-bytes. The hit time of the L2 cache is 15ns. Note that this is also the time taken to write a word to the L2 cache. The local hit rate of the L2 cache is 80%. Also, 50% of all L2 cache blocks replaced are dirty. The 64-bit wide main memory has an access latency of 20ns (including the time for the request to reach from the L2 cache to the main memory, you can consider this is the initial setup time for a memory access), after which any number of bus words may be transferred at the rate of one bus word (64-bit) per bus cycle on the 64-bit wide 100 MHz main memory bus. Assume inclusion between the L1 and L2 caches and assume there is no write-back buffer at the L2 cache. Assume a write-back takes the same amount of time as an L2 read miss of the same size.
While calculating any time values (such as hit time, miss penalty, AMAT), please use ns (nanoseconds) as the unit of time. For miss rates below, give the local miss rate for that cache. By miss penaltyL2, we mean the time from the miss request issued by the L2 cache up to the time the data comes back to the L2 cache from main memory.
only asked instruction accesss so doesn't need to care D-cache

1 + 0.02 (15+0.2 x 150) = 1.9ns
step 1
Read the new block from memory (always).
Write back the old block that is being replaced, if it is dirty (sometimes).
step 2
Access latency = 20
Bus is 64 bits wide -> 8 bytes
Bus runs at 100 MHz -> 10ns CCT
How many transfers for one block?
Block size ÷ bus width = 64 bytes ÷ 8 bytes = 8 bus words
How long do those 8 transfers take?
8 bus words × 10 ns = 80 ns
Total to read one block:
Read time = latency + transfer time
= 20 ns + 80 ns
= 100 ns

MP L2 = [latency + (block size ÷ bus width) × bus cycle time] × (1 + fraction dirty)
= [20 + (64 ÷ 8) × 10] × (1 + 0.5)
= 100 × 1.5
= 150 ns
AMAT = hit time L1 + miss rate L1 × (hit time L2 + miss rate L2 × miss penalty L2)
AMAT = hit time L1 + miss rate L1 × (hit time L2 + miss rate L2 × miss penalty L2)
miss rate L1 = 0.05 (5%), the D-cache's miss rate.
AMAT(data read) = 1 + 0.05 × (15 + 0.20 × 150)
= 1 + 0.05 × 45
= 1 + 2.25
= 3.25 ns
First, look at the path a write takes. The D-cache is write-through and no-write-allocate, so every store, hit or miss, is placed in the write buffer. The buffer then writes it to L2 in the background. The 5% D-cache miss rate therefore does not affect writes: a write never waits for a block to be brought into L1.
A write reaching L2 that misses is handled the same way as a read miss. L2 is write-allocate, so it fetches the whole block from memory, and it is write-back, so the replaced block is dirty 50% of the time and must be written out first, with no write-back buffer.
miss penalty L2 (writes) = 150 ns, the same as for reads.
This is the time for one write-buffer entry to be written into L2, including the L2 misses:
write time L2Buff = hit time L2 + miss rate L2 × miss penalty L2
= 15 + 0.20 × 150
= 15 + 30
= 45 ns
write time L2Buff = 45 ns (15 ns is the given time to write a word into L2).
There are two cases:
Case 1: buffer has a free entry (95% of writes).
The store goes into the buffer immediately and the processor continues. Stall = 0 ns.
Case 2: buffer is full (5% of writes).
The problem says the write arrives just as the buffer starts sending an entry to L2, and that entry isn't free until L2 finishes. The processor waits the entire L2 write. Stall = 45 ns.
AMAT(data write) = 0.95 × 0 + 0.05 × 45
= 2.25 ns
AMAT for data writes = 2.25 ns
(Some instructors also count the 1 ns L1 write access in every case. That gives 0.95 × 1 + 0.05 × (1 + 45) = 3.25 ns. The problem says to include only stall time, so 2.25 ns is the intended answer.)