26O01e

QK·약 21시간 전

A — B300 Multi-Model Serving Architecture

1. 목적

본 문서는 NVIDIA B300 8-GPU 단일 노드 환경에서 여러 LLM 모델을 동시에 서비스하기 위한 Multi-Model Inference Serving Architecture를 정의한다.

대상 환경은 다음과 같다.

  • NVIDIA B300 8-GPU 단일 서버
  • Kubernetes 기반 운영
  • NVIDIA GPU Operator / Device Plugin
  • vLLM 기반 LLM inference
  • 여러 모델의 동시 serving
  • 여러 사용자 및 팀(tenant)의 공동 사용
  • Keycloak / Active Directory 기반 인증
  • LiteLLM 기반 LLM Gateway 및 quota 관리
  • Redis 기반 distributed rate limiting / state
  • Prometheus / Grafana / DCGM 기반 observability
  • Cilium Gateway API / Envoy 기반 Kubernetes network ingress
  • 향후 여러 B300 노드 및 ClusterMesh 환경으로 확장 가능

핵심 설계 원칙은 다음과 같다.

Kubernetes가 GPU 자원을 할당하고, vLLM이 GPU에서 inference를 수행하며, LiteLLM이 사용자/팀/모델 단위의 API 정책과 quota를 관리한다.


2. Architecture Overview

전체 논리 구조는 다음과 같다.

flowchart TB

    U[Users / Applications]

    IDP[Keycloak / Active Directory]
    GW[Cilium Gateway API / Envoy]

    LL[LiteLLM Proxy]
    R[(Redis)]
    DB[(PostgreSQL)]

    subgraph B300["B300 8-GPU Kubernetes Node"]
        V70["vLLM 70B<br/>TP=4"]
        V32["vLLM 32B<br/>TP=2"]
        V8["vLLM 8B<br/>TP=1"]
        VE["Embedding / Reranker<br/>TP=1"]

        G0["GPU 0-3"]
        G1["GPU 4-5"]
        G2["GPU 6"]
        G3["GPU 7"]

        V70 --> G0
        V32 --> G1
        V8 --> G2
        VE --> G3
    end

    PROM[Prometheus]
    DCGM[DCGM Exporter]
    GRAF[Grafana]

    U -->|OIDC / JWT| IDP
    U -->|API Request| GW
    IDP -->|Identity / Claims| GW
    GW --> LL

    LL --> R
    LL --> DB

    LL -->|model=70b| V70
    LL -->|model=32b| V32
    LL -->|model=8b| V8
    LL -->|embedding/reranker| VE

    DCGM --> PROM
    V70 --> PROM
    V32 --> PROM
    V8 --> PROM
    VE --> PROM
    PROM --> GRAF

이 구조에서 각 계층의 책임을 명확하게 분리한다.

계층주요 컴포넌트책임
IdentityKeycloak / AD사용자 및 서비스 인증
Network GatewayCilium Gateway API / EnvoyTLS, JWT 검증, network/L7 policy
AI GatewayLiteLLMModel routing, tenant quota, RPM/TPM, concurrency
StateRedisDistributed rate-limit 및 shared state
PersistencePostgreSQL사용자/팀/API key/budget 등의 영속 데이터
InferencevLLMModel execution, batching, KV Cache, TP
GPU ResourceNVIDIA GPU OperatorGPU allocation 및 device management
Hardware MonitoringDCGM ExporterGPU utilization, HBM, power, temperature, NVLink
Platform MonitoringPrometheus/Grafana통합 모니터링 및 dashboard

3. B300 GPU Resource Partitioning

3.1 기본 원칙

8-GPU B300 노드는 여러 inference instance가 동시에 동작할 수 있는 하나의 GPU resource pool이다.

예시적인 초기 partition은 다음과 같다.

B300 Node
│
├── GPU 0 ─┐
├── GPU 1  │
├── GPU 2  ├── vLLM 70B / TP=4
├── GPU 3 ─┘
│
├── GPU 4 ─┐
├── GPU 5 ─┘── vLLM 32B / TP=2
│
├── GPU 6 ──── vLLM 8B / TP=1
│
└── GPU 7 ──── Embedding / Reranker / Spare

이는 초기 serving profile의 예시이며 고정된 최적 배치로 간주하지 않는다.

실제 배치는 다음 항목의 benchmark 결과를 기준으로 결정한다.

  • GPU memory usage
  • KV Cache capacity
  • TTFT
  • ITL
  • output token throughput
  • request concurrency
  • GPU utilization
  • NVLink traffic
  • CPU/NUMA locality
  • model size
  • context length
  • workload mix

3.2 Tensor Parallelism

후보 TP 구성은 다음과 같다.

Model class후보 TPGPU 수용도
Lightweight modelTP=117B~14B, embedding, reranker
Mid-size LLMTP=2232B급
Large LLMTP=4470B급
Very large modelTP=88대형 MoE / full-node serving

TP 크기를 1/2/4/8과 같이 구성하면 운영 및 benchmark matrix를 단순화할 수 있다.

다만,

TP는 반드시 2의 거듭제곱이어야 한다는 절대적인 규칙은 아니다.

실제 B300/NVLink topology와 모델의 communication pattern에 따라 TP=2/4/8 각각을 benchmark하여 결정한다.


4. Kubernetes GPU Allocation

GPU 자원 할당의 책임은 LiteLLM이나 vLLM이 아니라 Kubernetes + NVIDIA Device Plugin에 둔다.

예:

resources:
  requests:
    nvidia.com/gpu: "4"
  limits:
    nvidia.com/gpu: "4"

70B TP4 instance:

Kubernetes
    │
    │ nvidia.com/gpu = 4
    ▼
vLLM Pod
    │
    └── tensor-parallel-size = 4

32B:

resources:
  requests:
    nvidia.com/gpu: "2"
  limits:
    nvidia.com/gpu: "2"

8B:

resources:
  requests:
    nvidia.com/gpu: "1"
  limits:
    nvidia.com/gpu: "1"

중요한 운영 원칙

GPU 번호를 운영자가 직접 다음과 같이 고정하는 방식은 기본 설계로 사용하지 않는다.

CUDA_VISIBLE_DEVICES=0,1,2,3

대신 Kubernetes GPU resource allocation을 기준으로 한다.

CUDA_VISIBLE_DEVICES는 컨테이너 내부에서 NVIDIA runtime이 할당된 GPU를 노출하기 위한 실행 환경의 결과로 취급한다.


5. vLLM Instance Architecture

각 모델은 독립적인 vLLM deployment/service로 관리한다.

flowchart LR

    subgraph K8s["Kubernetes"]
        S70["Service<br/>vllm-70b"]
        S32["Service<br/>vllm-32b"]
        S8["Service<br/>vllm-8b"]

        V70["vLLM 70B<br/>TP=4"]
        V32["vLLM 32B<br/>TP=2"]
        V8["vLLM 8B<br/>TP=1"]

        S70 --> V70
        S32 --> V32
        S8 --> V8
    end

    G70["GPU x4"]
    G32["GPU x2"]
    G8["GPU x1"]

    V70 --> G70
    V32 --> G32
    V8 --> G8

각 vLLM instance는 다음을 독립적으로 관리한다.

  • Model weights
  • CUDA context
  • Tensor Parallelism
  • KV Cache
  • Continuous batching
  • Request scheduling
  • Maximum context
  • Maximum concurrent sequences
  • Prefix Cache
  • Prometheus metrics

6. vLLM Resource Configuration

예시:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: vllm-70b
  namespace: llm
spec:
  replicas: 1

  selector:
    matchLabels:
      app: vllm-70b

  template:
    metadata:
      labels:
        app: vllm-70b
        model: llama-70b

    spec:
      nodeSelector:
        accelerator: b300

      containers:
      - name: vllm
        image: vllm/vllm-openai:<validated-version>

        args:
        - "meta-llama/Llama-3.1-70B-Instruct"
        - "--tensor-parallel-size"
        - "4"
        - "--gpu-memory-utilization"
        - "0.90"
        - "--max-model-len"
        - "32768"

        ports:
        - containerPort: 8000

        resources:
          requests:
            cpu: "16"
            memory: "64Gi"
            nvidia.com/gpu: "4"

          limits:
            cpu: "32"
            memory: "128Gi"
            nvidia.com/gpu: "4"

gpu-memory-utilization

0.85~0.90은 초기 tuning 범위로 사용할 수 있지만 모든 모델에 적용되는 고정 권장값은 아니다.

최종값은 다음 결과를 기반으로 결정한다.

Model
  +
Context Length
  +
Concurrency
  +
KV Cache
  +
Prefix Cache
       ↓
GPU Memory Requirement
       ↓
OOM / latency / throughput benchmark

7. CPU / NUMA Affinity

GPU inference는 GPU memory만 고려해서는 안 된다.

다음 자원의 locality를 함께 고려한다.

GPU
 │
 ├── PCIe topology
 ├── NVLink topology
 ├── CPU socket
 ├── NUMA node
 ├── NIC
 └── memory

예시:

CPU NUMA 0
 ├── CPU cores
 ├── Memory
 └── GPU 0-3

CPU NUMA 1
 ├── CPU cores
 ├── Memory
 └── GPU 4-7

Kubernetes에서는 다음 기능을 함께 검토한다.

  • CPU Manager
  • Topology Manager
  • NVIDIA Device Plugin
  • Node Feature Discovery
  • Guaranteed QoS
  • CPU pinning
  • NUMA alignment

예를 들어 특정 serving workload에서는:

resources:
  requests:
    cpu: "16"
    memory: "64Gi"
    nvidia.com/gpu: "4"
  limits:
    cpu: "16"
    memory: "64Gi"
    nvidia.com/gpu: "4"

처럼 CPU request/limit을 동일하게 설정하여 Guaranteed QoS를 사용하는 방식을 검토할 수 있다.


8. Model Serving Gateway

모든 vLLM Service를 외부에 직접 노출하지 않는다.

권장 구조:

Client
  │
  ▼
Cilium Gateway API / Envoy
  │
  ▼
LiteLLM
  │
  ├── llama-70b ──> vllm-70b
  ├── qwen-32b  ──> vllm-32b
  ├── llama-8b  ──> vllm-8b
  └── embedding ──> embedding service

클라이언트는 backend의 IP/port를 알 필요가 없다.

예:

client = OpenAI(
    base_url="https://llm.example.internal/v1",
    api_key="..."
)

response = client.chat.completions.create(
    model="llama-70b",
    messages=[
        {"role": "user", "content": "Hello"}
    ]
)

9. LiteLLM as AI Gateway

LiteLLM은 GPU scheduler가 아니다.

역할을 다음과 같이 정의한다.

LiteLLM
│
├── Model routing
├── Deployment routing
├── User / Team identification
├── API key management
├── RPM
├── TPM
├── Budget
├── Max concurrent requests
├── Model access control
├── Fallback
└── Usage accounting

반면 vLLM은:

vLLM
│
├── Model execution
├── Tensor Parallelism
├── Continuous batching
├── KV Cache
├── Prefix Cache
└── GPU scheduling inside inference engine

을 담당한다.


10. LiteLLM Model Deployment

예:

model_list:

  - model_name: llama-70b
    litellm_params:
      model: openai/llama-70b
      api_base: http://vllm-70b.llm.svc.cluster.local:8000/v1
      api_key: dummy
      max_parallel_requests: 8

  - model_name: qwen-32b
    litellm_params:
      model: openai/qwen-32b
      api_base: http://vllm-32b.llm.svc.cluster.local:8000/v1
      api_key: dummy
      max_parallel_requests: 16

  - model_name: llama-8b
    litellm_params:
      model: openai/llama-8b
      api_base: http://vllm-8b.llm.svc.cluster.local:8000/v1
      api_key: dummy
      max_parallel_requests: 32

여기서 max_parallel_requests는 GPU 메모리를 직접 제한하는 값이 아니다.

예:

LiteLLM
  max_parallel_requests = 8
          │
          ▼
vLLM
  request scheduler
          │
          ▼
GPU
  KV Cache / batching

따라서 LiteLLM concurrency와 vLLM의 sequence/concurrency 설정은 별도로 benchmark하여 결정한다.


11. Multi-Deployment Routing

향후 동일 모델의 replica가 증가하면 하나의 logical model에 여러 deployment를 연결할 수 있다.

                    model = llama-70b
                           │
                     LiteLLM Router
                      /           \
                     /             \
                    ▼               ▼
             vLLM-70B-A       vLLM-70B-B
                TP4                TP4
              GPU 0-3            GPU 4-7

이 구조에서는 LiteLLM이 다음과 같은 routing을 수행할 수 있다.

  • Load balancing
  • Least-loaded deployment 선택
  • RPM/TPM-aware routing
  • Retry
  • Fallback
  • Deployment health 기반 제외

따라서 향후 B300 여러 대로 확장할 때도 동일한 논리 구조를 유지할 수 있다.


12. Authentication Architecture

기업 환경에서는 사용자 인증의 중심을 Keycloak / AD에 둔다.

권장 흐름:

sequenceDiagram

    participant U as User / Application
    participant K as Keycloak / AD
    participant E as Envoy Gateway
    participant L as LiteLLM
    participant V as vLLM

    U->>K: OIDC authentication
    K-->>U: Access Token (JWT)

    U->>E: API request + JWT
    E->>E: JWT signature / issuer / audience validation

    E->>L: Authenticated request
    L->>L: User / Team / Model policy
    L->>L: RPM / TPM / concurrency check

    L->>V: OpenAI-compatible request
    V-->>L: Inference response
    L-->>E: Response
    E-->>U: Response

13. Keycloak / AD Claims

JWT에는 다음과 같은 정보를 사용할 수 있다.

{
  "sub": "user-1234",
  "preferred_username": "user-a",
  "groups": [
    "ai-research",
    "data-platform"
  ],
  "roles": [
    "llm-user"
  ]
}

이를 다음과 같이 policy에 연결한다.

JWT
 │
 ├── sub
 │      ↓
 │    User
 │
 ├── groups
 │      ↓
 │    Team
 │
 └── roles
        ↓
      Permission

예:

ai-research
 ├── llama-70b
 ├── qwen-32b
 └── llama-8b

general-users
 ├── qwen-32b
 └── llama-8b

14. Tenant / User / Model Quota

사용자에게 GPU를 직접 할당하기보다는 API capacity를 할당한다.

권장 policy hierarchy:

Organization
     │
     ▼
   Team
     │
     ▼
   User
     │
     ▼
   API Key / Client
     │
     ▼
   Model

예:

Team70B TPM32B TPM8B TPMMax Concurrent
AI Research30K50K100K10
AI Engineering10K30K50K5
General-10K30K3

이러한 quota는 GPU의 물리적 할당과 독립적이다.


15. Rate Limit

LLM workload에서는 RPM만으로는 충분하지 않다.

최소한 다음을 고려한다.

항목의미목적
RPMRequests/minuteRequest burst 제어
TPMTokens/minute실제 inference load 제어
TPDTokens/day장기 사용량 제어
Budget비용/사용량팀별 resource accounting
Concurrency동시 requestBackend queue 보호
Context limit요청당 최대 contextKV Cache 보호
Output limit최대 생성 토큰장시간 generation 방지

16. Noisy Neighbor Protection

멀티테넌트 환경에서 가장 중요한 보호 장치 중 하나다.

예:

User A
  │
  └── 200K token request
           │
           ▼
      KV Cache pressure
           │
           ▼
      User B latency ↑

따라서 다음과 같은 다단계 제한을 적용한다.

             Client
                │
                ▼
       ┌────────────────┐
       │ Gateway        │
       │ JWT / Auth     │
       └───────┬────────┘
               │
       ┌───────▼────────┐
       │ LiteLLM        │
       │ RPM / TPM      │
       │ Team quota     │
       │ concurrency    │
       └───────┬────────┘
               │
       ┌───────▼────────┐
       │ vLLM           │
       │ context limit  │
       │ sequence limit │
       │ KV Cache       │
       └───────┬────────┘
               │
               ▼
              GPU

특히 --max-model-len은 단일 요청이 과도한 KV Cache를 사용하는 것을 방지하기 위한 중요한 보호 장치다.


17. Prefix Caching

동일한 system prompt 또는 반복되는 prefix가 많은 workload에서는 vLLM Prefix Caching을 별도 benchmark 대상으로 관리한다.

Request A
[System Prompt][Context A][Answer A]

Request B
[System Prompt][Context B][Answer B]

Request C
[System Prompt][Context C][Answer C]
       │
       ▼
  Prefix Cache

다만 Prefix Caching 효과를 고정된 비율로 가정하지 않는다.

실제 효과는 다음에 따라 달라진다.

  • Prefix length
  • Prefix reuse rate
  • Context length
  • Concurrency
  • Cache residency
  • Request arrival pattern
  • Model architecture

따라서 실제 B300 환경에서는 Prefix Cache ON/OFF A/B benchmark로 효과를 측정한다.

주요 측정값:

  • TTFT
  • Prefill throughput
  • Total throughput
  • ITL
  • GPU utilization
  • KV Cache utilization
  • Request latency

18. Observability Architecture

관측성은 세 계층으로 구성한다.

flowchart LR

    subgraph GPU["GPU / Hardware"]
        DCGM[DCGM Exporter]
        NV[NVLink]
        HBM[HBM]
        PWR[Power / Temperature]
    end

    subgraph INF["Inference"]
        V1[vLLM 70B]
        V2[vLLM 32B]
        V3[vLLM 8B]
    end

    subgraph NET["Network"]
        HUBBLE[Hubble]
        CILIUM[Cilium Metrics]
        NODE[Node Exporter]
    end

    PROM[Prometheus]
    GRAF[Grafana]

    DCGM --> PROM
    NV --> DCGM
    HBM --> DCGM
    PWR --> DCGM

    V1 --> PROM
    V2 --> PROM
    V3 --> PROM

    HUBBLE --> PROM
    CILIUM --> PROM
    NODE --> PROM

    PROM --> GRAF

19. GPU Metrics

DCGM에서 최소한 다음 항목을 수집한다.

Metric category목적
GPU utilizationCompute utilization
HBM utilizationMemory bandwidth / pressure
HBM usageModel + KV Cache memory
PowerPower saturation
TemperatureThermal condition
ECCHardware error
XIDGPU driver/hardware error
NVLink trafficTP communication
PCIe trafficHost↔GPU traffic

20. vLLM Metrics

vLLM의 실제 metric 이름은 사용 중인 vLLM version에서 /metrics 및 해당 version documentation을 기준으로 검증한다.

주요 관측 항목은 다음과 같다.

Request
 ├── request count
 ├── request latency
 ├── TTFT
 ├── ITL
 └── output throughput

Scheduler
 ├── running requests
 ├── waiting requests
 └── queue time

KV Cache
 ├── utilization
 └── available capacity

Engine
 ├── prompt tokens
 ├── generation tokens
 └── batch behavior

특히 다음 correlation이 중요하다.

KV Cache ↑
    +
Waiting Requests ↑
    +
TTFT ↑
    +
GPU Utilization ≈ 100%
        │
        ▼
Inference saturation

21. Kubernetes / Network Observability

B300 inference workload에서는 GPU metric만으로 원인을 판단하지 않는다.

함께 수집한다.

Kubernetes

  • Pod CPU
  • Pod memory
  • CPU throttling
  • OOMKilled
  • Pod restart
  • scheduling latency
  • GPU allocation
  • NUMA placement

Cilium

  • Hubble flows
  • dropped packets
  • retransmission
  • TCP RTT
  • connection errors
  • endpoint identity
  • NetworkPolicy drops

Node

  • CPU utilization
  • IRQ
  • softirq
  • NIC throughput
  • RX/TX drops
  • PCIe status
  • memory pressure

22. Failure Isolation

모델별 serving instance를 독립 Deployment로 운영하면 장애 격리가 쉬워진다.

vLLM-70B OOM
      │
      ▼
70B Service unavailable

        X

vLLM-32B ─────────── 정상
vLLM-8B  ─────────── 정상
Embedding ────────── 정상

반대로 하나의 process에서 여러 모델을 함께 실행하면 한 모델의 OOM이나 CUDA failure가 다른 모델까지 영향을 줄 가능성이 증가한다.

따라서 기본 원칙은:

One Model / One Inference Deployment

로 한다.


23. Health Check

LLM 모델은 일반적인 HTTP service보다 startup 시간이 길다.

따라서 다음을 구분한다.

Pod created
    │
    ▼
Container started
    │
    ▼
Model loading
    │
    ▼
CUDA initialization
    │
    ▼
KV Cache initialization
    │
    ▼
Inference engine ready
    │
    ▼
Readiness = TRUE

livenessProbe가 model loading 중인 Pod를 재시작하지 않도록 충분한 startup/readiness 정책을 사용한다.

특히 대형 model의 경우 initialDelaySeconds를 단순히 짧게 설정하지 않고 실제 load time을 benchmark하여 결정한다.


24. GPU Sharing Strategy

본 architecture에서는 다음 우선순위를 사용한다.

기본

Dedicated GPU set
      ↓
Dedicated vLLM instance
      ↓
Multiple users share the vLLM instance

즉:

GPU 0-3
   │
   └── vLLM 70B
          │
          ├── User A
          ├── User B
          ├── User C
          └── User D

사용자별로 GPU를 직접 분할하지 않는다.

이유

vLLM은 하나의 model instance 안에서 여러 request를:

  • continuous batching
  • KV Cache
  • request scheduling

을 이용해 처리할 수 있기 때문이다.

따라서 여러 사용자가 하나의 GPU set을 공유하는 가장 자연스러운 계층은 vLLM request scheduler다.


25. MIG / GPU Time-Slicing

MIG나 GPU time-slicing은 기본 serving architecture와 분리해서 평가한다.

기본 Multi-Model Serving

GPU set
  ↓
vLLM instance
  ↓
multiple tenants

GPU Time-Slicing

GPU
 ├── Pod A
 ├── Pod B
 └── Pod C

MIG

Physical GPU
 ├── MIG instance A
 ├── MIG instance B
 └── MIG instance C

MIG/time-slicing은 GPU isolation 요구사항이나 소형 workload에서 검토할 수 있지만, TP 기반 대형 LLM serving과는 별도의 architecture/benchmark 항목으로 관리한다.


26. Recommended Initial B300 Profile

초기 POC는 다음 profile로 시작할 수 있다.

ResourceInitial profile
GPU 0-370B TP4
GPU 4-532B TP2
GPU 68B TP1
GPU 7Embedding/Reranker 또는 spare
LiteLLM2 replicas
RedisHA 구성 검토
PostgreSQLLiteLLM persistent DB
KeycloakExisting enterprise IdP
GatewayCilium Gateway API / Envoy
MonitoringPrometheus + Grafana
GPU MonitoringDCGM Exporter
NetworkCilium + Hubble
AuthenticationOIDC/JWT
AuthorizationTeam/User/Model policy

이 profile은 production 최종값이 아니라 baseline으로 사용한다.


27. Request Flow

전체 request flow는 다음과 같다.

sequenceDiagram

    participant User as User / Application
    participant KC as Keycloak / AD
    participant GW as Cilium Gateway / Envoy
    participant LL as LiteLLM
    participant R as Redis
    participant V as vLLM
    participant GPU as B300 GPU

    User->>KC: Login / Client Credentials
    KC-->>User: JWT Access Token

    User->>GW: POST /v1/chat/completions
    Note over User,GW: Authorization: Bearer JWT

    GW->>GW: JWT validation
    GW->>LL: Authenticated request

    LL->>LL: Identify User / Team
    LL->>R: Check RPM / TPM / Concurrency

    alt Quota available
        R-->>LL: Allow
        LL->>V: Route by model
        V->>GPU: Prefill / Decode
        GPU-->>V: Tokens
        V-->>LL: Response
        LL-->>GW: Response
        GW-->>User: Response
    else Quota exceeded
        R-->>LL: Reject
        LL-->>GW: HTTP 429
        GW-->>User: HTTP 429 / Retry-After
    end

28. Capacity Control Model

전체 system의 capacity는 하나의 값이 아니라 여러 단계에서 제한된다.

                    Global Capacity
                          │
                 ┌────────┴────────┐
                 │                 │
             Team quota        Model quota
                 │                 │
                 └────────┬────────┘
                          │
                   Concurrent Requests
                          │
                     vLLM Queue
                          │
                    KV Cache
                          │
                       GPU

따라서 운영자는 다음 질문을 각각 구분해야 한다.

Q1. 사용자가 너무 많이 요청하는가?

→ LiteLLM quota

Q2. 특정 모델이 포화되었는가?

→ vLLM queue / concurrency / KV Cache

Q3. GPU가 포화되었는가?

→ DCGM

Q4. GPU는 한가한데 inference가 느린가?

→ CPU / NUMA / PCIe / network / storage / scheduler 확인

Q5. 여러 GPU를 사용하는 TP model이 느린가?

→ NVLink / topology / NCCL / CPU affinity 확인


29. Scaling Strategy

현재는 single B300 node지만 향후 다음 구조로 확장할 수 있다.

flowchart TB

    Client[Users / Applications]

    GW[Cilium Gateway]
    LL[LiteLLM Cluster]
    R[(Redis)]
    
    subgraph C1["Inference Cluster A"]
        A1["B300 Node 1<br/>70B TP4"]
        A2["B300 Node 2<br/>70B TP4"]
        A3["B300 Node 3<br/>32B / 8B"]
    end

    subgraph C2["Inference Cluster B"]
        B1["B300 Node 4<br/>70B"]
        B2["B300 Node 5<br/>32B"]
    end

    Client --> GW
    GW --> LL
    LL --> R

    LL --> A1
    LL --> A2
    LL --> A3

    LL --> B1
    LL --> B2

이 경우 LiteLLM은 logical model → 여러 physical deployment 구조로 확장할 수 있다.

Kubernetes/Cilium ClusterMesh를 사용하는 경우에는 cluster-level service discovery와 network policy를 별도로 설계한다.


30. Security Architecture

보안 책임을 계층별로 분리한다.

영역책임
IdentityKeycloak / AD
AuthenticationEnvoy / Gateway
AuthorizationGateway + LiteLLM
API quotaLiteLLM
Network isolationCilium NetworkPolicy
Workload identityKubernetes ServiceAccount
Service identitymTLS 검토
GPU isolationKubernetes/NVIDIA
SecretKubernetes Secret / Vault
AuditGateway + LiteLLM + vLLM logs

특히 Kubernetes ServiceAccount와 사용자 인증을 동일하게 취급하지 않는다.

Human
  ↓
Keycloak / AD
  ↓
OIDC / JWT

Kubernetes Workload
  ↓
ServiceAccount / Workload Identity

두 identity domain은 목적이 다르다.


31. Network Policy

LLM namespace에서는 필요한 통신만 허용한다.

예:

Gateway
   │
   ▼
LiteLLM
   │
   ├── vLLM-70B
   ├── vLLM-32B
   └── vLLM-8B

LiteLLM
   │
   ├── Redis
   └── PostgreSQL

Cilium NetworkPolicy 관점에서는 다음과 같이 최소 권한 정책을 구성한다.

Gateway
   └── allow → LiteLLM

LiteLLM
   ├── allow → vLLM
   ├── allow → Redis
   └── allow → PostgreSQL

User Pod
   └── deny → vLLM direct access

즉, 사용자가 vLLM Service를 직접 호출하여 LiteLLM quota를 우회하지 못하도록 한다.


32. Direct Backend Access Prevention

운영 환경에서는 다음을 방지한다.

                  ┌──> LiteLLM ──> vLLM
Client ── Gateway ┤
                  └──> vLLM  ← DENY

vLLM Service는 ClusterIP 기반 internal service로 유지하고 외부에는 직접 노출하지 않는다.

외부 요청의 표준 경로는:

Client
 ↓
Cilium Gateway
 ↓
LiteLLM
 ↓
vLLM

으로 고정한다.


33. Operational Dashboard

최소한 다음 dashboard를 구성한다.

Dashboard 1 — GPU

GPU Utilization
HBM Usage
HBM Bandwidth
Power
Temperature
ECC
XID
NVLink Traffic

Dashboard 2 — Model

Request Rate
Running Requests
Waiting Requests
TTFT
ITL
Input Tokens/sec
Output Tokens/sec
KV Cache Utilization
Error Rate

Dashboard 3 — Tenant

Requests/user
Tokens/user
Tokens/team
RPM
TPM
429 count
Model usage
Quota utilization

Dashboard 4 — Kubernetes

Pod CPU
CPU throttling
Memory
Restart
OOM
GPU allocation
Node pressure
NUMA placement

Dashboard 5 — Network

Network throughput
RTT
Packet loss
Retransmission
Cilium drops
Hubble flows
NIC drops

34. Alerting

예시 alert policy:

Alert조건 예시의미
GPU SaturationGPU util 지속 highGPU capacity 부족
KV Cache PressureKV cache 지속 highContext/concurrency 과다
Queue Growthwaiting requests 증가Backend saturation
TTFT SpikeTTFT 급증Prefill/queue 문제
GPU XIDXID 발생GPU/driver 문제
ECC ErrorECC 증가HBM/GPU hardware
Pod Restartrestart 증가Engine failure
OOMKilledOOM 발생Memory sizing 문제
429 Spike429 급증Quota 또는 capacity 부족
Network DropCilium/Hubble drop 증가Network 문제

35. Benchmark Requirements

Production serving configuration을 확정하기 전에 다음 benchmark를 수행한다.

Model-level

Model
 ×
TP
 ×
Precision
 ×
Context
 ×
Concurrency

System-level

Single model
      ↓
Multi model
      ↓
Multi tenant
      ↓
Mixed workload
      ↓
Saturation
      ↓
Soak

주요 측정값

TTFT
ITL
E2E latency
Input tok/s
Output tok/s
Request throughput
GPU utilization
HBM utilization
KV Cache utilization
Power
Temperature
NVLink traffic
CPU utilization
Network utilization
Error rate
429 rate

36. Recommended Stage 3 Benchmark Mapping

현재 B300 validation 계획과 연결하면 다음과 같이 매핑할 수 있다.

TestArchitecture validation
Stage 3a-1Basic model serving
Stage 3a-2Engine / TP tuning
Stage 3a-3Mixed context
Stage 3a-4Long-duration / soak
Stage 3a-5Quantization
Stage 3a-6Concurrency / saturation
Stage 3a-7GPU scaling
Stage 3a-8Prefix Cache ON/OFF
Stage 3a-9RAG workload
Stage 3a-10External KV / MemKV
K-seriesKubernetes / network / resilience

특히 Multi-Model Serving Architecture를 확정하기 위해서는 Stage 3a-6 + 3a-7 + 3a-8의 결과가 중요하다.


37. Recommended Initial Configuration

최초 production-like POC에서는 다음과 같이 시작한다.

B300 8 GPU
│
├── GPU 0-3
│    └── 70B / TP4
│
├── GPU 4-5
│    └── 32B / TP2
│
├── GPU 6
│    └── 8B / TP1
│
└── GPU 7
     └── Embedding / Reranker / Spare

                    │
                    ▼

             Kubernetes Services

                    │
                    ▼

                LiteLLM x2

                    │
             ┌──────┴──────┐
             │             │
           Redis       PostgreSQL

                    │
                    ▼

          Cilium Gateway / Envoy

                    │
                    ▼

             Keycloak / AD

단, 실제 네트워크 방향은 일반적으로 다음과 같이 생각하는 것이 더 정확하다.

                    Keycloak / AD
                         ▲
                         │ OIDC
                         │
User ───────────────> Gateway
                         │
                         ▼
                      LiteLLM
                         │
                ┌────────┼────────┐
                ▼        ▼        ▼
              vLLM      vLLM     vLLM
               70B       32B       8B
                │        │        │
                ▼        ▼        ▼
              GPU0-3    GPU4-5    GPU6

GPU 7은 workload 특성에 따라 embedding/reranker 또는 spare capacity로 사용할 수 있다.


38. Design Principles

최종적으로 본 architecture의 핵심 원칙은 다음과 같다.

Principle 1 — GPU와 User를 직접 연결하지 않는다

User → GPU

가 아니라:

User → LiteLLM → vLLM → GPU

로 구성한다.

Principle 2 — GPU allocation은 Kubernetes가 담당한다

nvidia.com/gpu

를 기준으로 GPU 자원을 관리한다.

Principle 3 — Model instance는 독립시킨다

1 Model
   ↓
1 vLLM Deployment

를 기본으로 한다.

Principle 4 — Tenant quota는 API capacity로 관리한다

TPM
RPM
Concurrency
Budget
Model access

를 LiteLLM에서 관리한다.

Principle 5 — GPU capacity와 API quota를 분리한다

GPU resource
≠
User quota

이다.

Principle 6 — Authentication과 Authorization을 분리한다

Keycloak / AD
     ↓
Authentication

Gateway / LiteLLM
     ↓
Authorization / Quota

Principle 7 — GPU뿐 아니라 전체 path를 관측한다

User
 ↓
Gateway
 ↓
LiteLLM
 ↓
vLLM
 ↓
CPU / NUMA
 ↓
GPU / NVLink

각 계층의 latency와 resource usage를 timestamp 기준으로 correlation한다.


39. Target Architecture

최종적으로 B300 Multi-Model Serving Platform은 다음 구조를 목표로 한다.

flowchart TB

    USERS["Users / Applications"]

    IDP["Keycloak / Active Directory"]

    GW["Cilium Gateway API / Envoy"]

    LL["LiteLLM<br/>AI Gateway"]

    REDIS["Redis<br/>Rate Limit / Shared State"]

    DB["PostgreSQL<br/>Persistent Metadata"]

    subgraph K8S["Kubernetes Inference Platform"]

        subgraph B300["NVIDIA B300 8-GPU Node"]

            V70["vLLM 70B<br/>TP4"]
            V32["vLLM 32B<br/>TP2"]
            V8["vLLM 8B<br/>TP1"]
            EMB["Embedding / Reranker"]

            GPU70["GPU 0-3"]
            GPU32["GPU 4-5"]
            GPU8["GPU 6"]
            GPUEMB["GPU 7"]

            V70 --> GPU70
            V32 --> GPU32
            V8 --> GPU8
            EMB --> GPUEMB
        end

        NP["NetworkPolicy / Cilium"]
        SA["ServiceAccount / Workload Identity"]
    end

    PROM["Prometheus"]
    DCGM["DCGM Exporter"]
    HUBBLE["Hubble"]
    GRAF["Grafana"]

    USERS -->|"OIDC / JWT"| IDP
    USERS -->|"HTTPS / OpenAI API"| GW

    IDP -->|"Identity / Claims"| GW
    GW --> LL

    LL --> REDIS
    LL --> DB

    LL -->|"llama-70b"| V70
    LL -->|"qwen-32b"| V32
    LL -->|"llama-8b"| V8
    LL -->|"embedding"| EMB

    DCGM --> PROM
    V70 --> PROM
    V32 --> PROM
    V8 --> PROM
    EMB --> PROM
    HUBBLE --> PROM

    PROM --> GRAF

    NP -.-> GW
    NP -.-> LL
    NP -.-> V70
    NP -.-> V32
    NP -.-> V8

이 구조를 기준으로 하면 향후 B300이 여러 대로 증가하더라도:

Single B300
     ↓
Multiple B300 Nodes
     ↓
Multiple Kubernetes Nodes
     ↓
Multiple Inference Clusters
     ↓
Cilium ClusterMesh

로 확장할 수 있으며, 사용자에게 노출되는 OpenAI-compatible API와 Team/User/Model 정책은 최대한 동일하게 유지할 수 있다.


40. Summary

이 아키텍처의 가장 중요한 설계 경계는 다음과 같다.

┌───────────────────────────────────────────────┐
│              User / Tenant Layer              │
│                                               │
│ Keycloak / AD                                │
│ Team / User / Model Permission               │
│ RPM / TPM / Budget                            │
└──────────────────────┬────────────────────────┘
                       │
┌──────────────────────▼────────────────────────┐
│                AI Gateway Layer                │
│                                               │
│ Cilium Gateway / Envoy                        │
│ LiteLLM                                      │
│ Routing / Quota / Concurrency / Accounting    │
└──────────────────────┬────────────────────────┘
                       │
┌──────────────────────▼────────────────────────┐
│              Inference Layer                   │
│                                               │
│ vLLM                                          │
│ TP / Batching / KV Cache / Prefix Cache       │
└──────────────────────┬────────────────────────┘
                       │
┌──────────────────────▼────────────────────────┐
│                GPU Layer                       │
│                                               │
│ NVIDIA GPU Operator                           │
│ B300 / NVLink / HBM                           │
│ GPU allocation / NUMA / topology              │
└───────────────────────────────────────────────┘

핵심적으로 Kubernetes → vLLM → LiteLLM의 역할을 뒤섞지 않는 것이 이 architecture의 가장 중요한 원칙이다.

  • Kubernetes/NVIDIA: GPU 자원 할당
  • vLLM: GPU에서 실제 inference
  • LiteLLM: 사용자/팀/모델/API capacity 관리
  • Envoy/Cilium: network ingress 및 security boundary
  • Keycloak/AD: identity
  • Redis/PostgreSQL: shared state 및 persistence
  • DCGM/Prometheus/Grafana/Hubble: end-to-end observability
profile
engineer

0개의 댓글