Increasing Cache BandWidth

Seungyun Lee·2일 전

Computer Arch (Memory)

목록 보기
15/16

1. Pipelined caches

A processor's clock cycle can be no shorter than its slowest single step (the critical path). A cache access is one of the slowest steps: decode the index, read the arrays, compare tags, select the data. If the whole access has to fit in one cycle, it limits the clock speed of the entire CPU.

The idea: split the cache access into stages

Pipelining the cache works like an assembly line. The access is cut into stages, for example index → read → tag compare → data out, with latches between them. Each stage is short, so the clock can be faster. A new access can enter the cache every cycle, while earlier ones are still in later stages.

Two things change, and it's important to keep them apart:

  • Bandwidth goes up: more accesses finish per nanosecond, one per (shorter) cycle.
  • Latency does not improve: each access still takes about the same real time, a little more because of latch overhead. Measured in cycles, a hit now takes k cycles instead of 1.

1stage

2stages

3stages

4stages

What to notice as you go from 1 to 4 stages:

The clock cycle shrinks, from about 2.1 ns to 0.6 ns, so the whole CPU can run faster.
Bandwidth more than triples, because a new load starts every cycle.
Hit latency in real time rises slightly, from 2.1 ns to 2.4 ns, because each latch adds overhead. Pipelining never makes a single access faster.
Hit latency in cycles grows from 1 to 4, which is where the costs come from.

Each step allowed a much higher clock rate. The Pentium 4 was designed for very high GHz, which is why it pipelined the cache (and everything else) so deeply.

The two costs in your notes

  1. More cycles between issuing a load and using its data (load-use delay)
    With a 4-cycle hit, an instruction that needs the loaded value must wait extra cycles, the yellow "wait" boxes in the widget:
lw   r1, 0(r2)     # load, takes k cycles
add  r3, r1, r4    # needs r1, so it stalls k-1 cycles

The processor hides this by doing independent work in the gap. The compiler schedules other instructions between the load and its use, or out-of-order execution (which the Pentium Pro introduced) runs later instructions that don't depend on the load. If there's nothing independent to run, the stall shows up.

  1. Greater penalty on mispredicted branches
    The instruction cache is pipelined too, so fetching instructions takes more stages, and the whole CPU pipeline gets longer. When a branch is mispredicted, every instruction fetched after it is wrong and must be thrown away, and the pipeline must refill from the correct address. More stages mean more wasted cycles per mistake. On a very deep pipeline like the Pentium 4's, this penalty is large, so good branch prediction becomes essential.

Pipelining the cache trades latency in cycles for higher clock rate and bandwidth. It pays off when the processor can keep the pipeline full with independent instructions and correct branch predictions, and it hurts when it can't.

2. Multibanked caches

3. Nonblocking caches

profile
Design Verification engineer

0개의 댓글