Performance

Seungyun Lee·2026년 8월 26일

Computer Arch (Memory)

목록 보기
1/16

Response Time and Throughput

  • Response time (also called execution
    time)
    • How long it takes to do a task
  • Throughput
    • Total work done per unit time
    • e.g., tasks/transactions/… per hour

Realative Performance

PerformancexPerformancey=Execution timeyExecution timex=n\frac {Performance_x} {Performance_y} = \frac {Execution\ time_y} {Execution\ time_x} = n

x,y 분자 분모 뒤바뀌는거 주의
x가 n만큼 y보다 빠르다 or 느리다

Measuring Execution Time

  • Elapsed time ( total excution time, wall time)
    - Total response time, Including all aspects

    • $ time ls: measure time under the file, ls(list files in derectory)
  • CPU Time
    - Time spent processing a given job

CPU Time

CPUTime=CPU Clock CyclesClock Rate=CPU Clock Cycle×CCTCPU_{Time} = \frac {CPU\ Clock\ Cycles} {Clock\ Rate} = CPU\ Clock\ Cycle \times CCT

거속시 관계처럼 유연하게 생각하기

Time=CCRate=CC×CCTTime = \frac{CC} {Rate} = CC \times CCT

Perfomance improved by

  • reducing number of clock cycle
  • increasing clock rate

Example

Computer A: 2GHz clock, 10s CPU time
• Designing Computer B
– Aim for 6s CPU time
– Can do faster clock, but causes 1.2 × clock cycles

• How fast must Computer B clock be?

CPI (Cycle per Insturction)

Instruction Count for a program

  • Determined by program, ISA and compiler
ClockCycle=Instruction Count×Cycle per InstructionClock Cycle = Instruction\ Count \times Cycle\ per\ Instruction
CPUTime=IC×CPIClock Rate=IC×CPI×CCTCPU_{Time} = \frac {IC \times CPI} {Clock\ Rate} = IC \times CPI \times CCT

ISA (instruction set Architecture)

  • defines types of instuctions

  • defines addressing mode

    	ALU: ADD R1, R2, R3 -> Register addressing mode
    	BNE R1, R2, L1
    	LD R, 8(R2) -> Indirect addressing mode

bracket means no data -> need to access memory

RISC (Reduced Inst Set Computing)

  • ADD R1 R2 R3: add R2, R3

CISC (Complex Inst Set Computing)

  • ADD R1, (R2), (R3): bring R2, R3 data from memory and add -> store to R1
  • include multiple instructions
  • inefficient
 j = b(i) + c(i)
 
 RISC - MIPS, ARM
 LD R4, (R2)
 LD R5, (R3)
 ADD R1, R4, R5
 
 CISC
 ADD R1, (R2), (R3)

Average cycles per instruction

  • Determined by CPU hardware
  • If different instructions have different CPI

CPU Performance Equation

이 공식은 "Iron Law of Processor Performance"라고 불릴 정도로 중요합니다. CPU 시간을 세 가지 요소로 분해합니다.

① 기본 공식 유도

CPU 시간은 단순히 (총 클럭 사이클 수) ×\times (한 사이클의 길이)입니다.

CPU Time=CPU Clock Cycles×Clock Cycle TimeCPU~Time = CPU~Clock~Cycles \times Clock~Cycle~Time

또는 클럭 속도(Rate, Frequency)를 이용하면:

CPU Time=CPU Clock CyclesClock RateCPU~Time = \frac{CPU~Clock~Cycles}{Clock~Rate}

참고: Clock Rate=1Clock Cycle TimeClock~Rate = \frac{1}{Clock~Cycle~Time} (예: 1GHz = 1ns 주기)

② 심화 공식 (Instruction Count & CPI 도입)

하지만 '총 사이클 수'만으로는 분석이 어렵습니다. 그래서 명령어 개수(Instruction Count)와 CPI(Cycles Per Instruction) 개념을 도입합니다.

  • Instruction Count (IC): 프로그램이 실행한 기계어 명령어의 총개수.

  • CPI (Cycles Per Instruction): 명령어 하나를 실행하는 데 평균적으로 드는 사이클 수.

CPI=CPU Clock CyclesInstruction CountCPI = \frac{CPU~Clock~Cycles}{Instruction~Count}

이걸 위 공식에 대입하면 그 유명한 3변수 공식이 나옵니다.

CPU Time=Instruction Count×CPI×Clock Cycle TimeCPU~Time = Instruction~Count \times CPI \times Clock~Cycle~Time

=IC×CPI×CCT= IC \times CPI \times CCT

=Instruction Count×CPIClock Rate= \frac{Instruction~Count \times CPI}{Clock~Rate}

IPC (Instructions Per Cycle) 란?
정의: 한 사이클(Clock) 동안 처리할 수 있는 명령어의 개수.
수식:
IPC=1CPI=Instruction CountCPU Clock CyclesIPC = \frac{1}{CPI} = \frac{\text{Instruction Count}}{\text{CPU Clock Cycles}}
의미: "우리 CPU는 한 번 쿵! 할 때마다 명령어를 몇 개씩 해치우나?" (높을수록 좋음)

CPI (낮을수록 좋음): "이번 신형 CPU는 CPI가 0.8에서 0.5로 줄었어!" -> 뭔가 줄었다니 좋은 건가? (헷갈림)
IPC (높을수록 좋음): "이번 신형 CPU는 IPC가 1.2에서 2.0으로 늘었어!" -> 와! 성능이 두 배 가까이 뛰었네! (직관적)

CPU Time=Instruction CountClock Rate×IPCCPU \text{ Time} = \frac{\text{Instruction Count}}{\text{Clock Rate} \times \text{IPC}}

1. 기본 개념:

가중 평균 CPI전체 프로그램의 실행 시간(Clock Cycles)은 각 명령어 클래스의 CPI와 발생 빈도(Relative Frequency)를 곱한 값들의 합으로 결정됩니다.

CPI=∑i=1n(CPIi×InstructionCountiInstructionCountTotal)CPI = \sum_{i=1}^{n} \left( CPI_i \times \frac{Instruction Count_i}{Instruction Count_{Total}} \right)

2. 현재 시스템 분석

(Original CPI)시스템의 현재 상태는 다음과 같습니다.

  • FP (부동소수점 연산): 전체 명령어의 25% 차지, CPI는 4. (이 25% 안에는 2% 빈도로 발생하는 FPSQR이 포함되어 있음)
  • Others (기타 연산): 나머지 75% 차지, CPI는 1.33.
  • Original CPI: (4×25%)+(1.33×75%)=2.0(4 \times 25\%) + (1.33 \times 75\%) = 2.0.

3. 두 가지 하드웨어 개선안 비교

설계자는 다음 두 가지 개선안 중 하나를 선택해야 합니다.

  • Design 1 (특정 악성 병목 해결): FPSQR 명령어는 CPI가 20으로 매우 느립니다. 이 연산 유닛을 대폭 뜯어고쳐 CPI를 2로 10배 개선합니다.
    • 개선된 CPI = 기존 CPI - (FPSQR 빈도 ×\times 감소한 CPI)
    • CPI=2.0−2%×(20−2)=1.64CPI = 2.0 - 2\% \times (20 - 2) = 1.64
  • Design 2 (범용적 성능 개선): 특정한 하나가 아닌, FP 유닛 전체의 파이프라인을 개선하여 FP 전체 CPI를 4에서 2.5로 낮춥니다.
    • CPI=(75%×1.33)+(25%×2.5)=1.625CPI = (75\% \times 1.33) + (25\% \times 2.5) = 1.625

4. 현실적인 결론

단일 명령어(FPSQR)의 성능을 무려 10배나 향상시킨 Design 1보다, FP 유닛 전체의 성능을 1.6배 정도 향상시킨 Design 2의 최종 CPI(1.625)가 더 낮아 전체 성능이 우수합니다. 이는 아무리 극적인 성능 향상(CPI 20 →\rightarrow 2)을 이루더라도, 자주 쓰이지 않는 연산(2%)이라면 전체 시스템에 미치는 영향은 미미하다는 암달의 법칙(Amdahl's Law)을 여실히 보여주는 예시입니다.

Performance Summary

Performance depends on
– Algorithm: affects IC, possibly CPI
– Programming language: affects IC, CPI
– Compiler: affects IC, CPI
– Instruction set architecture: affects IC, CPI, Tc

Example

A given application written in Java runs 15 seconds on a desktop processor. A new Java compiler is released that requires only 0.6 as
many instructions as the old compiler. Unfortunately, it increases the CPI by 1.1. How fast can we expect the application to run using this new compiler?

Answer: 150.61.1=9.9seconds
CPU time = IC CPI CCT

profile
Design Verification engineer

0개의 댓글