네. 여기서 제일 중요한 점은 4개를 각각 독립 제품으로 설치하는 구조가 아니라는 것입니다.
LiteLLM과 vLLM은 실제 애플리케이션이고, InferencePool은 Kubernetes CRD, EPP는 Inference Extension이 제공하는 routing component입니다.
사용자 / Jupyter / Spark / Application
│
│ HTTPS
▼
┌──────────────────────────────────────┐
│ ① LiteLLM │
│ │
│ 사용자/팀 인증 │
│ Virtual Key │
│ Model ACL │
│ RPM / TPM / 동시성 제한 │
│ Usage / Budget │
│ Logical Model Alias │
│ │
│ "discovery" → deepseek-v4.1-flash │
└──────────────────┬───────────────────┘
│
│ OpenAI API
▼
┌──────────────────────────────────────┐
│ ② K8s Gateway API │
│ │
│ HTTPRoute │
│ │ │
│ ▼ │
│ ③ InferencePool │
│ │
│ deepseek-v4.1-flash endpoint 집합 │
└──────────────────┬───────────────────┘
│
▼
┌──────────────────────────────────────┐
│ ④ EPP - Endpoint Picker │
│ │
│ Pod A? Pod B? │
│ queue / load / KV locality 등 │
└─────────┬─────────────────┬──────────┘
│ │
▼ ▼
┌─────────────────┐ ┌─────────────────┐
│ ⑤ vLLM A │ │ ⑤ vLLM B │
│ Node #1 TP4 │ │ Node #2 TP4 │
│ GPU 0-3 │ │ GPU 0-3 │
└─────────────────┘ └─────────────────┘
Gateway API Inference Extension은 이 중 InferencePool + EPP를 inference-aware routing 계층으로 정의합니다.
LiteLLM은 GPU scheduler가 아니라 AI API Management 계층으로 두겠습니다.
예를 들어 사용자는 실제 모델 이름을 몰라도 됩니다.
Application
model = "discovery"
LiteLLM에서:
discovery
↓
deepseek-v4.1-flash
로 매핑합니다.
다른 팀은:
ontology
↓
glm-5.3
를 사용합니다.
그리고 여기서 사용자/팀별로:
Discovery Team
models:
- discovery
- qwen
TPM: 2,000,000
RPM: 100
max_parallel_requests: 20
Wizard Team
models:
- qwen
- document
TPM: 500,000
RPM: 50
max_parallel_requests: 8
같은 정책을 관리합니다.
LiteLLM은 virtual keys, team/model access, RPM/TPM 및 max parallel requests 같은 관리 기능을 제공합니다.
Kubernetes에 Deployment로 올립니다.
개념적으로:
namespace: ai-gateway
LiteLLM
├─ Deployment
├─ Service
├─ ConfigMap/Secret
└─ PostgreSQL
형태입니다.
LiteLLM의 key/team/budget 같은 관리 기능을 사용할 경우 PostgreSQL을 같이 사용하는 구성이 적절합니다.
실제 production에서는 대략:
:443
│
▼
Cilium Gateway
│
▼
LiteLLM Service
│
┌────────┴────────┐
│ │
LiteLLM #1 LiteLLM #2
│ │
└────────┬────────┘
│
PostgreSQL
로 만들 수 있습니다.
GPU는 필요 없습니다.
여기서 약간 헷갈릴 수 있습니다.
Gateway API는 프로그램 이름이 아니라 Kubernetes API/CRD입니다.
예:
apiVersion: gateway.networking.k8s.io/v1
kind: Gateway
metadata:
name: inference-gateway
namespace: ai-serving
같은 리소스를 선언합니다.
실제 packet을 처리하는 Gateway implementation이 따로 필요합니다.
현재 환경에서는 이미 Cilium을 사용하므로:
Gateway API
↓
Cilium Gateway implementation
↓
Envoy dataplane
를 우선 검토할 수 있습니다. Cilium은 Gateway API를 지원합니다.
다만 정확히 사용할 Cilium 버전과 Gateway API Inference Extension 버전의 호환성은 구축 전에 확인해야 합니다.
InferencePool은 프로그램이 아닙니다.
Kubernetes CRD입니다.
예를 들어 6개월 뒤 B300 node가 2대가 되었다고 합시다.
Node #1
deepseek Pod A
TP4
Node #2
deepseek Pod B
TP4
둘은 동일한 모델을 제공합니다.
그럼:
InferencePool
name = deepseek-pool
│
├── deepseek Pod A
│
└── deepseek Pod B
라고 정의합니다.
개념적인 manifest는 이런 형태입니다.
apiVersion: inference.networking.x-k8s.io/v1
kind: InferencePool
metadata:
name: deepseek-pool
spec:
selector:
matchLabels:
app: vllm
model: deepseek-v4-1-flash
즉 Kubernetes에게:
이 label을 가진 Pod들은 같은 inference endpoint pool이다.
라고 알려주는 것입니다.
공식 Inference Extension에서도 InferencePool은 inference workload를 제공하는 endpoint 집합을 나타내는 API입니다.
EPP가 실제로 재미있는 부분입니다.
EPP = Endpoint Picker
이건 실제 component입니다.
예를 들어:
InferencePool
Pod A
running=24
waiting=10
KV=88%
Pod B
running=8
waiting=1
KV=45%
일 때 EPP가:
EPP
│
┌────────┴────────┐
│ │
Pod A score Pod B score
BAD GOOD
│
▼
Pod B
처럼 endpoint를 선택합니다.
따라서 EPP의 질문은 딱 하나입니다.
"이 요청을 어느 replica가 처리해야 하는가?"
Gateway API Inference Extension은 EPP를 inference-specific routing decision을 담당하는 component로 사용합니다.
EPP를 직접 Python으로 만드는 것부터 시작하지 않습니다.
Inference Extension의 배포 구성을 사용합니다.
Kubernetes에서는 대략:
namespace: ai-serving
EPP
├─ Deployment
├─ Service
├─ ServiceAccount
├─ RBAC
└─ InferencePool 연동
으로 보입니다.
여기가 실제 GPU inference engine입니다.
vLLM Pod
│
├── model
├── tokenizer
├── KV cache
├── batching
├── scheduling
└── generation
│
▼
B300 GPU
예를 들어 v12 Stage7 결과가:
DeepSeek
TP2 FAIL
TP4 PASS
TP8 PASS
TP4
SLO capacity C48
required C39
→ TP4 production
라면 Deployment가 대략:
apiVersion: apps/v1
kind: Deployment
metadata:
name: deepseek-vllm
spec:
template:
metadata:
labels:
app: vllm
model: deepseek-v4-1-flash
spec:
containers:
- name: vllm
resources:
limits:
nvidia.com/gpu: 4
이고 command는 production profile에서 생성합니다.
vllm serve /models/deepseek-v4.1-flash
--tensor-parallel-size 4
--max-model-len ...
--max-num-seqs ...
--gpu-memory-utilization ...
--enable-prefix-caching
즉 v12 Stage7의 역할은 바로 여기입니다.
Discovery 팀의 Jupyter에서:
client.chat.completions.create(
model="discovery",
messages=[...]
)
를 호출했다고 합시다.
LiteLLM이 key를 확인합니다.
API Key
↓
Team = discovery
Allowed models?
↓
discovery ✓
RPM?
↓
PASS
TPM?
↓
PASS
Concurrent requests?
↓
12 / 20
PASS
그리고:
logical model
discovery
↓
physical model
deepseek-v4.1-flash
로 보냅니다.
Gateway가 요청을 받습니다.
논리적으로:
deepseek-v4.1-flash
│
▼
deepseek-pool
로 route합니다.
현재:
deepseek-pool
├─ Pod A
└─ Pod B
가 등록되어 있습니다.
EPP가 상태를 봅니다.
Pod A Pod B
running 22 8
waiting 8 1
KV 84% 47%
prefix hit HIGH LOW
그리고 routing policy에 따라 Pod를 선택합니다.
예를 들어 Pod B.
Pod B
↓
continuous batching
↓
KV allocation
↓
B300
↓
token generation
결과가 다시:
vLLM
↓
Gateway
↓
LiteLLM
↓
Jupyter
로 돌아갑니다.
현재는 사실 이렇게 됩니다.
LiteLLM
│
▼
Gateway
│
▼
InferencePool
│
▼
EPP
│
▼
DeepSeek Pod A
EPP가 할 일이 별로 없습니다.
endpoint가 하나뿐이니까요.
그런데 구조는 유지합니다.
6개월 뒤 Node #2가 들어오면:
LiteLLM
│
▼
Gateway
│
▼
InferencePool
│
EPP
│
┌──────┴──────┐
│ │
▼ ▼
Node #1 Node #2
Pod A Pod B
TP4 TP4
가 됩니다.
Application은 아무것도 바꿀 필요가 없습니다.
예를 들면:
LiteLLM logical models
discovery
│
▼
deepseek-pool
│
├── DeepSeek A
└── DeepSeek B
fast-discovery
│
▼
glm-flash-pool
│
├── GLM Flash A
└── GLM Flash B
qwen
│
▼
qwen-pool
│
├── Qwen A
└── Qwen B
이런 구조입니다.
PaddleOCR/Docling까지 억지로 vLLM InferencePool 구조에 넣을 필요는 없습니다.
LiteLLM
│
├─ LLM API
│ │
│ ▼
│ InferencePool
│ ↓
│ EPP
│ ↓
│ vLLM
│
└─ X Document pipeline
보다는 document는:
Document API Gateway
↓
Document Orchestrator
┌────┴─────────────┐
▼ ▼
PaddleOCR Docling
GPU worker CPU/GPU worker
형태가 더 자연스럽습니다.
LiteLLM은 기본적으로 LLM API gateway로 사용하는 것이 좋습니다.
저라면:
ai-gateway
│
├─ LiteLLM
├─ Secret
└─ config
ai-serving-system
│
├─ Gateway
├─ EPP
├─ RBAC
└─ common policies
ai-model-discovery
│
├─ InferencePool
└─ DeepSeek vLLM Deployment
ai-model-glm
│
├─ InferencePool
└─ GLM vLLM Deployment
ai-model-qwen
│
├─ InferencePool
└─ Qwen vLLM Deployment
ai-observability
│
├─ Prometheus
├─ DCGM
├─ OTel
└─ Grafana
정도로 분리하겠습니다.
다음 순서가 좋습니다.
1 Kubernetes
↓
2 NVIDIA GPU Operator
↓
3 Cilium
↓
4 Gateway API CRDs
↓
5 Gateway API Inference Extension
↓
6 Gateway implementation 확인
↓
7 EPP
↓
8 vLLM Deployment
↓
9 InferencePool
↓
10 HTTPRoute
↓
11 LiteLLM + PostgreSQL
↓
12 Prometheus/DCGM/Hubble
↓
13 Stage7D validation
처음부터 LiteLLM까지 올려서 문제를 디버깅하지 않는 게 좋습니다.
먼저:
Gateway → Pool → EPP → vLLM
이 정상인지 확인합니다.
그 다음:
LiteLLM → Gateway → Pool → EPP → vLLM
로 확장합니다.
현재 v12가 생성하는 것이:
Stage7
↓
model-production-profile.yaml
↓
helm-values.yaml
인데 다음 버전에서는:
Stage7
↓
TP / capacity / SLO
↓
model-production-profile.yaml
│
├─→ vLLM Helm values
│
├─→ InferencePool
│
├─→ HTTPRoute
│
└─→ LiteLLM model/team capacity config
까지 생성하게 할 수 있습니다.
예를 들어 Stage7 결과가:
DeepSeek
TP4
production C30
max SLO C48
이면 자동으로:
vLLM
GPU=4
TP=4
InferencePool
deepseek-pool
LiteLLM
discovery
max_parallel_requests=30
까지 연결하는 방식입니다.
지금 구축하려는 환경에서는 이게 꽤 중요한 다음 단계입니다. Stage7에서 측정한 실제 B300 capacity가 LiteLLM의 quota와 K8s serving configuration 양쪽에 동일한 source-of-truth로 반영되기 때문에, 사람이 세 군데에 서로 다른 값을 넣는 문제를 피할 수 있습니다.
LiteLLM Proxy 공식 문서
Gateway API Inference Extension 공식 문서
InferencePool 공식 설명
Cilium Gateway API 공식 문서
수정했습니다. v12 기반을 유지하면서 v13으로 확장했습니다.
이제 흐름은 다음과 같습니다.
Stage7
↓
TP / Capacity / SLO / Headroom
↓
model-production-profile.yaml
│
├─→ vllm-helm-values.yaml
├─→ inferencepool.yaml
├─→ httproute.yaml
└─→ litellm-model-team-capacity.yaml
추가로 v12에서 중요했던 문제도 같이 보완했습니다. Concurrency + RPS headroom을 TP 선택에 반영할 수 있게 했고, model_selection_summary.json이 REVIEW_REQUIRED인데 generator가 임의로 다른 TP를 골라 PASS시키는 fallback을 막았습니다. max_error_rate가 설정됐는데 error metric 자체가 없을 때도 이제 SLO PASS로 보지 않습니다.
LiteLLM 출력은 특정 LiteLLM 버전의 Admin API payload를 섣불리 고정하지 않고, Stage7에서 측정된 max_parallel_requests / sustainable_rps / team allocation을 담는 version-neutral capacity source-of-truth로 생성합니다. 실제 배포 단계에서 사용하는 LiteLLM 버전의 config/Admin API 형식으로 변환하도록 분리했습니다.
새 통합 실행은:
./stage7_generate_full_production_bundle.sh \
deepseek_v4_1_flash \
/var/log/b300_validation/stage7/online/deepseek_v4_1_flash/<RUN> \
./stage7_production_evidence_example.yaml \
/tmp/deepseek-prod
입니다. Python compile, shell syntax, YAML parse와 InferencePool / HTTPRoute / LiteLLM capacity synthetic PASS 생성까지 검증했습니다.
b300_validation_common_env_v13.tar.gz전체 v13 패키지 다운로드
:chatgpt-content-reference{index="1"}새 Stage7 Serving Platform 가이드
:chatgpt-content-reference{index="2"}CHANGELOG v13
:chatgpt-content-reference{index="3"}Production policy v13
다음 단계에서는 이 v13을 기준으로 실제 Cilium Gateway API + Gateway API Inference Extension/EPP + LiteLLM Helm 배포 구조까지 포함한 ai-serving Kubernetes manifest/Helm chart를 붙이면, Stage7 qualification 결과에서 실제 GPU 서비스 배포까지 거의 일관된 GitOps 파이프라인으로 연결할 수 있습니다.
네. 현재 v13 기준 현장 인수검증용 Excel 결과서를 만들었습니다.
:chatgpt-content-reference{index="0"}B300 v13 Acceptance Test Report.xlsx 다운로드
요청하신 6개 시트로 구성했습니다.
01_Overview — 프로젝트/노드/OS/vLLM/K8s/검증자/최종 승인 정보02_HW_Acceptance — B300 8장, PCIe, NVLink/NVSwitch, ECC/XID, DCGM, NCCL, FP8/FP4, thermal, NIC, NVMe03_vLLM_Function — vLLM 0.30 compatibility, readiness/warm-up, BF16/FP8, prefix cache, KV tier, metrics, K8s 재검증04_Performance — TP1/2/4/8 × concurrency, TTFT/ITL P50/P99, Req/s, Output tok/s, Error Rate, KV tier 및 Stage7 capacity/SLO/headroom05_Soak — HW burn, vLLM 8h, KV 안정성, K8s restart, production profile soak06_Result_Summary — 각 시트의 PASS / FAIL / N/A / NOT RUN 자동 집계 및 Overall 판정각 검사항목에는 PASS/FAIL/N/A/NOT RUN 드롭다운, Actual Result, Evidence Path, 비고/조치 입력란을 넣었습니다. 초기 상태는 전부 NOT RUN으로 두어서 검증하지 않은 항목이 실수로 PASS 처리되지 않습니다.
특히 04_Performance에는 현재 v13 방향에 맞춰 Expected Peak Concurrency/RPS + Headroom → Required Capacity → Stage7 Selected TP를 기록할 수 있는 영역도 포함했습니다. 수식 오류 검사도 완료했습니다.