HW3

Seungyun Lee·어제

Computer Arch (Memory)

목록 보기
16/16

You are building a computer system around a processor with in-order execution that runs at 1 GHz and has a CPI of 1, excluding memory accesses. The only instructions that read or write data from/to memory are loads (20% of all instructions) and stores (5% of all instructions).

The memory system for this computer has a split L1 cache. Both the I-cache and the D-cache are direct mapped and hold 32 KB each. The I-cache has a 2% miss rate and 64 byte blocks, and the D-cache is a write-through, no-write-allocate cache with a 5% miss rate and 64 byte blocks. The hit time for both the Icache and the D-cache is 1ns. The L1 cache has a write buffer. 95% of writes to L1 find a free entry in the write buffer immediately. The other 5% of the writes have to wait until an entry frees up in the write buffer (assume that such writes arrive just as the write buffer initiates a request to L2 to free up its entry and the entry is not freed up until the L2 is done with the request). The processor is stalled on a write until a free write buffer entry is available.

The L2 cache is a unified write-back, write-allocate cache with a total size of 512 KB and a block size of 64-bytes. The hit time of the L2 cache is 15ns. Note that this is also the time taken to write a word to the L2 cache. The local hit rate of the L2 cache is 80%. Also, 50% of all L2 cache blocks replaced are dirty. The 64-bit wide main memory has an access latency of 20ns (including the time for the request to reach from the L2 cache to the main memory, you can consider this is the initial setup time for a memory access), after which any number of bus words may be transferred at the rate of one bus word (64-bit) per bus cycle on the 64-bit wide 100 MHz main memory bus. Assume inclusion between the L1 and L2 caches and assume there is no write-back buffer at the L2 cache. Assume a write-back takes the same amount of time as an L2 read miss of the same size.

While calculating any time values (such as hit time, miss penalty, AMAT), please use ns (nanoseconds) as the unit of time. For miss rates below, give the local miss rate for that cache. By miss penaltyL2, we mean the time from the miss request issued by the L2 cache up to the time the data comes back to the L2 cache from main memory.


1. Computing the AMAT (average memory access time) for instruction accesses.

a. Give the values of the following terms for instruction accesses. hit timeL1, miss rateL1, hit timeL2, miss rateL2

only asked instruction accesss so doesn't need to care D-cache

b. Give the formula for calculating miss penaltyL2, and compute the value of miss penaltyL2.

1 + 0.02 (15+0.2 x 150) = 1.9ns 
step 1
Read the new block from memory (always).
Write back the old block that is being replaced, if it is dirty (sometimes).

step 2
Access latency = 20
Bus is 64 bits wide -> 8 bytes
Bus runs at 100 MHz -> 10ns CCT

How many transfers for one block?
Block size ÷ bus width = 64 bytes ÷ 8 bytes = 8 bus words

How long do those 8 transfers take?
8 bus words × 10 ns = 80 ns

Total to read one block:

Read time = latency + transfer time
          = 20 ns + 80 ns
          = 100 ns

MP L2 = [latency + (block size ÷ bus width) × bus cycle time] × (1 + fraction dirty)
                = [20 + (64 ÷ 8) × 10] × (1 + 0.5)
                = 100 × 1.5
                = 150 ns

c. Give the formula for calculating the AMAT for this system using the five terms whose values you computed above and any other values you need.

AMAT = hit time L1 + miss rate L1 × (hit time L2 + miss rate L2 × miss penalty L2)

d. Plug in the values into the AMAT formula above, and compute a numerical value for AMAT for instruction accesses.

AMAT = hit time L1 + miss rate L1 × (hit time L2 + miss rate L2 × miss penalty L2)

2. Computing the AMAT for data reads

a. Give the value of miss rate L1 for data reads.

miss rate L1 = 0.05 (5%), the D-cache's miss rate.

b. Calculate the value of the AMAT for data reads using the above value, and other values you need.

AMAT(data read) = 1 + 0.05 × (15 + 0.20 × 150)
                = 1 + 0.05 × 45
                = 1 + 2.25
                = 3.25 ns

3. Computing the AMAT for data writes.

First, look at the path a write takes. The D-cache is write-through and no-write-allocate, so every store, hit or miss, is placed in the write buffer. The buffer then writes it to L2 in the background. The 5% D-cache miss rate therefore does not affect writes: a write never waits for a block to be brought into L1.

a. Give the value of miss penaltyL2 for data writes.

A write reaching L2 that misses is handled the same way as a read miss. L2 is write-allocate, so it fetches the whole block from memory, and it is write-back, so the replaced block is dirty 50% of the time and must be written out first, with no write-back buffer.

miss penalty L2 (writes) = 150 ns, the same as for reads.

b. Give the value of write timeL2Buff for a write buffer entry being written to the L2 cache.

This is the time for one write-buffer entry to be written into L2, including the L2 misses:

write time L2Buff = hit time L2 + miss rate L2 × miss penalty L2
                  = 15 + 0.20 × 150
                  = 15 + 30
                  = 45 ns

write time L2Buff = 45 ns (15 ns is the given time to write a word into L2).

c. Calculate the value of the AMAT for data writes using the above two values, and any other values that you need. Only include the time that the processor will be stalled. Hint: There are two cases to be considered here depending upon whether the write buffer is full or not.

There are two cases:

Case 1: buffer has a free entry (95% of writes).
The store goes into the buffer immediately and the processor continues. Stall = 0 ns.

Case 2: buffer is full (5% of writes).
The problem says the write arrives just as the buffer starts sending an entry to L2, and that entry isn't free until L2 finishes. The processor waits the entire L2 write. Stall = 45 ns.

AMAT(data write) = 0.95 × 0 + 0.05 × 45
                 = 2.25 ns

AMAT for data writes = 2.25 ns

(Some instructors also count the 1 ns L1 write access in every case. That gives 0.95 × 1 + 0.05 × (1 + 45) = 3.25 ns. The problem says to include only stall time, so 2.25 ns is the intended answer.)

profile
Design Verification engineer

0개의 댓글