A processor's clock cycle can be no shorter than its slowest single step (the critical path). A cache access is one of the slowest steps: decode the index, read the arrays, compare tags, select the data. If the whole access has to fit in one cycle, it limits the clock speed of the entire CPU.
The idea: split the cache access into stages
Pipelining the cache works like an assembly line. The access is cut into stages, for example index → read → tag compare → data out, with latches between them. Each stage is short, so the clock can be faster. A new access can enter the cache every cycle, while earlier ones are still in later stages.
Two things change, and it's important to keep them apart:
1stage

2stages

3stages

4stages

What to notice as you go from 1 to 4 stages:
The clock cycle shrinks, from about 2.1 ns to 0.6 ns, so the whole CPU can run faster.
Bandwidth more than triples, because a new load starts every cycle.
Hit latency in real time rises slightly, from 2.1 ns to 2.4 ns, because each latch adds overhead. Pipelining never makes a single access faster.
Hit latency in cycles grows from 1 to 4, which is where the costs come from.

Each step allowed a much higher clock rate. The Pentium 4 was designed for very high GHz, which is why it pipelined the cache (and everything else) so deeply.
The two costs in your notes
lw r1, 0(r2) # load, takes k cycles
add r3, r1, r4 # needs r1, so it stalls k-1 cycles
The processor hides this by doing independent work in the gap. The compiler schedules other instructions between the load and its use, or out-of-order execution (which the Pentium Pro introduced) runs later instructions that don't depend on the load. If there's nothing independent to run, the stall shows up.
Pipelining the cache trades latency in cycles for higher clock rate and bandwidth. It pays off when the processor can keep the pipeline full with independent instructions and correct branch predictions, and it hurts when it can't.