26O05d2

QK·6일 전

네. 여기서 제일 중요한 점은 4개를 각각 독립 제품으로 설치하는 구조가 아니라는 것입니다.

LiteLLM과 vLLM은 실제 애플리케이션이고, InferencePool은 Kubernetes CRD, EPP는 Inference Extension이 제공하는 routing component입니다.

전체 요청 흐름

사용자 / Jupyter / Spark / Application
                 │
                 │ HTTPS
                 ▼
┌──────────────────────────────────────┐
│ ① LiteLLM                           │
│                                      │
│ 사용자/팀 인증                       │
│ Virtual Key                          │
│ Model ACL                            │
│ RPM / TPM / 동시성 제한              │
│ Usage / Budget                       │
│ Logical Model Alias                  │
│                                      │
│ "discovery" → deepseek-v4.1-flash   │
└──────────────────┬───────────────────┘
                   │
                   │ OpenAI API
                   ▼
┌──────────────────────────────────────┐
│ ② K8s Gateway API                   │
│                                      │
│ HTTPRoute                            │
│              │                       │
│              ▼                       │
│ ③ InferencePool                     │
│                                      │
│ deepseek-v4.1-flash endpoint 집합    │
└──────────────────┬───────────────────┘
                   │
                   ▼
┌──────────────────────────────────────┐
│ ④ EPP - Endpoint Picker             │
│                                      │
│ Pod A? Pod B?                       │
│ queue / load / KV locality 등        │
└─────────┬─────────────────┬──────────┘
          │                 │
          ▼                 ▼
┌─────────────────┐ ┌─────────────────┐
│ ⑤ vLLM A       │ │ ⑤ vLLM B       │
│ Node #1 TP4    │ │ Node #2 TP4    │
│ GPU 0-3        │ │ GPU 0-3        │
└─────────────────┘ └─────────────────┘

Gateway API Inference Extension은 이 중 InferencePool + EPP를 inference-aware routing 계층으로 정의합니다.


① LiteLLM

역할

LiteLLM은 GPU scheduler가 아니라 AI API Management 계층으로 두겠습니다.

예를 들어 사용자는 실제 모델 이름을 몰라도 됩니다.

Application

model = "discovery"

LiteLLM에서:

discovery
      ↓
deepseek-v4.1-flash

로 매핑합니다.

다른 팀은:

ontology
    ↓
glm-5.3

를 사용합니다.

그리고 여기서 사용자/팀별로:

Discovery Team
  models:
    - discovery
    - qwen

  TPM: 2,000,000
  RPM: 100
  max_parallel_requests: 20


Wizard Team
  models:
    - qwen
    - document

  TPM: 500,000
  RPM: 50
  max_parallel_requests: 8

같은 정책을 관리합니다.

LiteLLM은 virtual keys, team/model access, RPM/TPM 및 max parallel requests 같은 관리 기능을 제공합니다.

설치

Kubernetes에 Deployment로 올립니다.

개념적으로:

namespace: ai-gateway

LiteLLM
 ├─ Deployment
 ├─ Service
 ├─ ConfigMap/Secret
 └─ PostgreSQL

형태입니다.

LiteLLM의 key/team/budget 같은 관리 기능을 사용할 경우 PostgreSQL을 같이 사용하는 구성이 적절합니다.

실제 production에서는 대략:

                 :443
                   │
                   ▼
          Cilium Gateway
                   │
                   ▼
             LiteLLM Service
                   │
          ┌────────┴────────┐
          │                 │
      LiteLLM #1        LiteLLM #2
          │                 │
          └────────┬────────┘
                   │
              PostgreSQL

로 만들 수 있습니다.

GPU는 필요 없습니다.


② Gateway API

여기서 약간 헷갈릴 수 있습니다.

Gateway API는 프로그램 이름이 아니라 Kubernetes API/CRD입니다.

예:

apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
  name: inference-gateway
  namespace: ai-serving

같은 리소스를 선언합니다.

실제 packet을 처리하는 Gateway implementation이 따로 필요합니다.

현재 환경에서는 이미 Cilium을 사용하므로:

Gateway API
      ↓
Cilium Gateway implementation
      ↓
Envoy dataplane

를 우선 검토할 수 있습니다. Cilium은 Gateway API를 지원합니다.

다만 정확히 사용할 Cilium 버전과 Gateway API Inference Extension 버전의 호환성은 구축 전에 확인해야 합니다.


③ InferencePool

InferencePool은 프로그램이 아닙니다.

Kubernetes CRD입니다.

예를 들어 6개월 뒤 B300 node가 2대가 되었다고 합시다.

Node #1
deepseek Pod A
TP4

Node #2
deepseek Pod B
TP4

둘은 동일한 모델을 제공합니다.

그럼:

InferencePool
name = deepseek-pool

        │
        ├── deepseek Pod A
        │
        └── deepseek Pod B

라고 정의합니다.

개념적인 manifest는 이런 형태입니다.

apiVersion: inference.networking.x-k8s.io/v1
kind: InferencePool

metadata:
  name: deepseek-pool

spec:
  selector:
    matchLabels:
      app: vllm
      model: deepseek-v4-1-flash

즉 Kubernetes에게:

이 label을 가진 Pod들은 같은 inference endpoint pool이다.

라고 알려주는 것입니다.

공식 Inference Extension에서도 InferencePool은 inference workload를 제공하는 endpoint 집합을 나타내는 API입니다.


④ EPP

EPP가 실제로 재미있는 부분입니다.

EPP = Endpoint Picker

이건 실제 component입니다.

예를 들어:

InferencePool

Pod A
running=24
waiting=10
KV=88%

Pod B
running=8
waiting=1
KV=45%

일 때 EPP가:

             EPP
              │
     ┌────────┴────────┐
     │                 │
 Pod A score       Pod B score
    BAD               GOOD
                         │
                         ▼
                      Pod B

처럼 endpoint를 선택합니다.

따라서 EPP의 질문은 딱 하나입니다.

"이 요청을 어느 replica가 처리해야 하는가?"

Gateway API Inference Extension은 EPP를 inference-specific routing decision을 담당하는 component로 사용합니다.

설치

EPP를 직접 Python으로 만드는 것부터 시작하지 않습니다.

Inference Extension의 배포 구성을 사용합니다.

Kubernetes에서는 대략:

namespace: ai-serving

EPP
 ├─ Deployment
 ├─ Service
 ├─ ServiceAccount
 ├─ RBAC
 └─ InferencePool 연동

으로 보입니다.


⑤ vLLM

여기가 실제 GPU inference engine입니다.

vLLM Pod
     │
     ├── model
     ├── tokenizer
     ├── KV cache
     ├── batching
     ├── scheduling
     └── generation
           │
           ▼
        B300 GPU

예를 들어 v12 Stage7 결과가:

DeepSeek

TP2  FAIL
TP4  PASS
TP8  PASS

TP4
SLO capacity C48
required C39

→ TP4 production

라면 Deployment가 대략:

apiVersion: apps/v1
kind: Deployment

metadata:
  name: deepseek-vllm

spec:
  template:
    metadata:
      labels:
        app: vllm
        model: deepseek-v4-1-flash

    spec:
      containers:
        - name: vllm

          resources:
            limits:
              nvidia.com/gpu: 4

이고 command는 production profile에서 생성합니다.

vllm serve /models/deepseek-v4.1-flash

--tensor-parallel-size 4
--max-model-len ...
--max-num-seqs ...
--gpu-memory-utilization ...
--enable-prefix-caching

즉 v12 Stage7의 역할은 바로 여기입니다.


이제 실제 요청 하나를 따라가 보겠습니다

Discovery 팀의 Jupyter에서:

client.chat.completions.create(
    model="discovery",
    messages=[...]
)

를 호출했다고 합시다.

STEP 1 — LiteLLM

LiteLLM이 key를 확인합니다.

API Key
   ↓
Team = discovery

Allowed models?
   ↓
discovery ✓

RPM?
   ↓
PASS

TPM?
   ↓
PASS

Concurrent requests?
   ↓
12 / 20

PASS

그리고:

logical model
discovery
     ↓
physical model
deepseek-v4.1-flash

로 보냅니다.


STEP 2 — Gateway / HTTPRoute

Gateway가 요청을 받습니다.

논리적으로:

deepseek-v4.1-flash
       │
       ▼
deepseek-pool

로 route합니다.


STEP 3 — InferencePool

현재:

deepseek-pool

├─ Pod A
└─ Pod B

가 등록되어 있습니다.


STEP 4 — EPP

EPP가 상태를 봅니다.

              Pod A        Pod B

running         22            8
waiting          8            1
KV              84%          47%
prefix hit      HIGH         LOW

그리고 routing policy에 따라 Pod를 선택합니다.

예를 들어 Pod B.


STEP 5 — vLLM

Pod B
  ↓
continuous batching
  ↓
KV allocation
  ↓
B300
  ↓
token generation

결과가 다시:

vLLM
 ↓
Gateway
 ↓
LiteLLM
 ↓
Jupyter

로 돌아갑니다.


현재 GPU 1-node에서는 어떻게 되나

현재는 사실 이렇게 됩니다.

LiteLLM
    │
    ▼
Gateway
    │
    ▼
InferencePool
    │
    ▼
EPP
    │
    ▼
DeepSeek Pod A

EPP가 할 일이 별로 없습니다.

endpoint가 하나뿐이니까요.

그런데 구조는 유지합니다.

6개월 뒤 Node #2가 들어오면:

                 LiteLLM
                    │
                    ▼
                  Gateway
                    │
                    ▼
              InferencePool
                    │
                   EPP
                    │
             ┌──────┴──────┐
             │             │
             ▼             ▼
        Node #1         Node #2

        Pod A           Pod B
        TP4             TP4

가 됩니다.

Application은 아무것도 바꿀 필요가 없습니다.


모델별로 Pool을 하나씩 만드는 것이 이해하기 쉽습니다

예를 들면:

LiteLLM logical models

discovery
   │
   ▼
deepseek-pool
   │
   ├── DeepSeek A
   └── DeepSeek B


fast-discovery
   │
   ▼
glm-flash-pool
   │
   ├── GLM Flash A
   └── GLM Flash B


qwen
   │
   ▼
qwen-pool
   │
   ├── Qwen A
   └── Qwen B

이런 구조입니다.


그런데 Document 계열은 별도로 보는 게 좋습니다

PaddleOCR/Docling까지 억지로 vLLM InferencePool 구조에 넣을 필요는 없습니다.

LiteLLM
 │
 ├─ LLM API
 │      │
 │      ▼
 │  InferencePool
 │      ↓
 │     EPP
 │      ↓
 │    vLLM
 │
 └─ X Document pipeline

보다는 document는:

Document API Gateway

      ↓

Document Orchestrator

 ┌────┴─────────────┐
 ▼                  ▼

PaddleOCR        Docling
 GPU worker      CPU/GPU worker

형태가 더 자연스럽습니다.

LiteLLM은 기본적으로 LLM API gateway로 사용하는 것이 좋습니다.


Kubernetes namespace로 내려보면

저라면:

ai-gateway
│
├─ LiteLLM
├─ Secret
└─ config


ai-serving-system
│
├─ Gateway
├─ EPP
├─ RBAC
└─ common policies


ai-model-discovery
│
├─ InferencePool
└─ DeepSeek vLLM Deployment


ai-model-glm
│
├─ InferencePool
└─ GLM vLLM Deployment


ai-model-qwen
│
├─ InferencePool
└─ Qwen vLLM Deployment


ai-observability
│
├─ Prometheus
├─ DCGM
├─ OTel
└─ Grafana

정도로 분리하겠습니다.


설치 순서도 중요합니다

다음 순서가 좋습니다.

1 Kubernetes
       ↓
2 NVIDIA GPU Operator
       ↓
3 Cilium
       ↓
4 Gateway API CRDs
       ↓
5 Gateway API Inference Extension
       ↓
6 Gateway implementation 확인
       ↓
7 EPP
       ↓
8 vLLM Deployment
       ↓
9 InferencePool
       ↓
10 HTTPRoute
       ↓
11 LiteLLM + PostgreSQL
       ↓
12 Prometheus/DCGM/Hubble
       ↓
13 Stage7D validation

처음부터 LiteLLM까지 올려서 문제를 디버깅하지 않는 게 좋습니다.

먼저:

Gateway → Pool → EPP → vLLM

이 정상인지 확인합니다.

그 다음:

LiteLLM → Gateway → Pool → EPP → vLLM

로 확장합니다.


그리고 여기서 v12를 한 단계 더 발전시킬 수 있습니다

현재 v12가 생성하는 것이:

Stage7
  ↓
model-production-profile.yaml
  ↓
helm-values.yaml

인데 다음 버전에서는:

Stage7
 ↓
TP / capacity / SLO
 ↓
model-production-profile.yaml
 │
 ├─→ vLLM Helm values
 │
 ├─→ InferencePool
 │
 ├─→ HTTPRoute
 │
 └─→ LiteLLM model/team capacity config

까지 생성하게 할 수 있습니다.

예를 들어 Stage7 결과가:

DeepSeek
TP4
production C30
max SLO C48

이면 자동으로:

vLLM
GPU=4
TP=4

InferencePool
deepseek-pool

LiteLLM
discovery
max_parallel_requests=30

까지 연결하는 방식입니다.

지금 구축하려는 환경에서는 이게 꽤 중요한 다음 단계입니다. Stage7에서 측정한 실제 B300 capacity가 LiteLLM의 quota와 K8s serving configuration 양쪽에 동일한 source-of-truth로 반영되기 때문에, 사람이 세 군데에 서로 다른 값을 넣는 문제를 피할 수 있습니다.

LiteLLM Proxy 공식 문서
Gateway API Inference Extension 공식 문서
InferencePool 공식 설명
Cilium Gateway API 공식 문서


수정했습니다. v12 기반을 유지하면서 v13으로 확장했습니다.

이제 흐름은 다음과 같습니다.

Stage7
 ↓
TP / Capacity / SLO / Headroom
 ↓
model-production-profile.yaml
 │
 ├─→ vllm-helm-values.yaml
 ├─→ inferencepool.yaml
 ├─→ httproute.yaml
 └─→ litellm-model-team-capacity.yaml

추가로 v12에서 중요했던 문제도 같이 보완했습니다. Concurrency + RPS headroom을 TP 선택에 반영할 수 있게 했고, model_selection_summary.json이 REVIEW_REQUIRED인데 generator가 임의로 다른 TP를 골라 PASS시키는 fallback을 막았습니다. max_error_rate가 설정됐는데 error metric 자체가 없을 때도 이제 SLO PASS로 보지 않습니다.

LiteLLM 출력은 특정 LiteLLM 버전의 Admin API payload를 섣불리 고정하지 않고, Stage7에서 측정된 max_parallel_requests / sustainable_rps / team allocation을 담는 version-neutral capacity source-of-truth로 생성합니다. 실제 배포 단계에서 사용하는 LiteLLM 버전의 config/Admin API 형식으로 변환하도록 분리했습니다.

새 통합 실행은:

./stage7_generate_full_production_bundle.sh \
  deepseek_v4_1_flash \
  /var/log/b300_validation/stage7/online/deepseek_v4_1_flash/<RUN> \
  ./stage7_production_evidence_example.yaml \
  /tmp/deepseek-prod

입니다. Python compile, shell syntax, YAML parse와 InferencePool / HTTPRoute / LiteLLM capacity synthetic PASS 생성까지 검증했습니다.

b300_validation_common_env_v13.tar.gz전체 v13 패키지 다운로드

:chatgpt-content-reference{index="1"}새 Stage7 Serving Platform 가이드

:chatgpt-content-reference{index="2"}CHANGELOG v13

:chatgpt-content-reference{index="3"}Production policy v13

다음 단계에서는 이 v13을 기준으로 실제 Cilium Gateway API + Gateway API Inference Extension/EPP + LiteLLM Helm 배포 구조까지 포함한 ai-serving Kubernetes manifest/Helm chart를 붙이면, Stage7 qualification 결과에서 실제 GPU 서비스 배포까지 거의 일관된 GitOps 파이프라인으로 연결할 수 있습니다.


네. 현재 v13 기준 현장 인수검증용 Excel 결과서를 만들었습니다.

:chatgpt-content-reference{index="0"}B300 v13 Acceptance Test Report.xlsx 다운로드

요청하신 6개 시트로 구성했습니다.

  • 01_Overview — 프로젝트/노드/OS/vLLM/K8s/검증자/최종 승인 정보
  • 02_HW_Acceptance — B300 8장, PCIe, NVLink/NVSwitch, ECC/XID, DCGM, NCCL, FP8/FP4, thermal, NIC, NVMe
  • 03_vLLM_Function — vLLM 0.30 compatibility, readiness/warm-up, BF16/FP8, prefix cache, KV tier, metrics, K8s 재검증
  • 04_Performance — TP1/2/4/8 × concurrency, TTFT/ITL P50/P99, Req/s, Output tok/s, Error Rate, KV tier 및 Stage7 capacity/SLO/headroom
  • 05_Soak — HW burn, vLLM 8h, KV 안정성, K8s restart, production profile soak
  • 06_Result_Summary — 각 시트의 PASS / FAIL / N/A / NOT RUN 자동 집계 및 Overall 판정

각 검사항목에는 PASS/FAIL/N/A/NOT RUN 드롭다운, Actual Result, Evidence Path, 비고/조치 입력란을 넣었습니다. 초기 상태는 전부 NOT RUN으로 두어서 검증하지 않은 항목이 실수로 PASS 처리되지 않습니다.

특히 04_Performance에는 현재 v13 방향에 맞춰 Expected Peak Concurrency/RPS + Headroom → Required Capacity → Stage7 Selected TP를 기록할 수 있는 영역도 포함했습니다. 수식 오류 검사도 완료했습니다.

profile
engineer

0개의 댓글