B300 8장(총 8×192GB or 8×288GB HBM3e)이 장착된 단일 노드는 연산력과 메모리 대역폭이 매우 압도적인 머신입니다.
위의 문서 전처리 스크립트는 수많은 청크를 동시에 밀어 넣는 비동기 배치(High Throughput) 작업이므로, 모델 크기 및 운영 방식에 맞춰 아래 2가지 구성 전략 중 하나를 선택하는 것이 가장 효율적입니다.
Qwen/Qwen2.5-32B-Instruct 또는 14B, Llama-3.1-8B/70BTP=2, Port 8001)TP=2, Port 8002)TP=2, Port 8003)TP=2, Port 8004)Llama-3.1-70B-Instruct or Qwen2.5-72B-Instruct (또는 추후 DeepSeek-V2/V3 계열)tensor-parallel-size 8로 단일 vLLM 서버(Port 8000)를 띄웁니다.nvidia-smi
# B300 인식 확인 (Driver 550+ / 560+ 권장)
# NVIDIA Container Toolkit 동작 확인
docker run --rm --gpus all nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi
가장 간단하게 단일 포트(8000)로 8장을 모두 활용하는 실행 방법입니다.
docker run -d --name vllm-server \
--runtime nvidia \
--gpus all \
--ipc=host \
-p 8000:8000 \
-v /data/models:/data/models \
vllm/vllm-openai:latest \
--model /data/models/Qwen2.5-32B-Instruct \
--tensor-parallel-size 8 \
--gpu-memory-utilization 0.90 \
--max-model-len 8192 \
--max-num-seqs 256 \
--port 8000
--tensor-parallel-size 8: 8장의 GPU로 모델 가중치를 분산.--max-num-seqs 256: 배치 내에서 동시에 동적 스케줄링할 최대 요청 수(스크립트의 --concurrency 48 이상을 거뜬히 처리).--gpu-memory-utilization 0.90: KV Cache에 VRAM 90% 할당 (대용량 컨텍스트 및 다중 배치 수용).--ipc=host: PyTorch 분산 통신(NCCL) 공유 메모리 부족 방지 (필수).8장의 GPU를 2장씩 4개 인스턴스로 분할 실행하면 스크립트의 수천 개 청크 처리 속도가 극대화됩니다.
# Instance 1 (GPU 0, 1) -> Port 8001
docker run -d --name vllm-01 --runtime nvidia --gpus '"device=0,1"' --ipc=host \
-p 8001:8000 -v /data/models:/data/models vllm/vllm-openai:latest \
--model /data/models/Qwen2.5-32B-Instruct \
--tensor-parallel-size 2 --max-model-len 8192 --max-num-seqs 128
# Instance 2 (GPU 2, 3) -> Port 8002
docker run -d --name vllm-02 --runtime nvidia --gpus '"device=2,3"' --ipc=host \
-p 8002:8000 -v /data/models:/data/models vllm/vllm-openai:latest \
--model /data/models/Qwen2.5-32B-Instruct \
--tensor-parallel-size 2 --max-model-len 8192 --max-num-seqs 128
# Instance 3 (GPU 4, 5) -> Port 8003
docker run -d --name vllm-03 --runtime nvidia --gpus '"device=4,5"' --ipc=host \
-p 8003:8000 -v /data/models:/data/models vllm/vllm-openai:latest \
--model /data/models/Qwen2.5-32B-Instruct \
--tensor-parallel-size 2 --max-model-len 8192 --max-num-seqs 128
# Instance 4 (GPU 6, 7) -> Port 8004
docker run -d --name vllm-04 --runtime nvidia --gpus '"device=6,7"' --ipc=host \
-p 8004:8000 -v /data/models:/data/models vllm/vllm-openai:latest \
--model /data/models/Qwen2.5-32B-Instruct \
--tensor-parallel-size 2 --max-model-len 8192 --max-num-seqs 128
앞단에 Nginx를 띄워 8000번 포트로 들어오는 요청을 8001~8004로 라운드로빈 로드밸런싱하거나, 단순하게 Nginx 없이 단일 인스턴스(방법 1)로 시작하셔도 B300 스펙상 충분히 강력합니다.
서버가 정상 구동되었는지 curl로 헬스체크 및 OpenAI 호환 엔드포인트를 호출해 봅니다.
# 1. 모델 목록 확인
curl http://localhost:8000/v1/models
# 2. 간단한 추론 테스트 (JSON 모드 지원 확인)
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "/data/models/Qwen2.5-32B-Instruct",
"messages": [{"role": "user", "content": "Ping test"}],
"max_tokens": 10
}'
batch_enrich_markdown.py) 호출vLLM이 베어메탈에서 실행되면, 작성해 둔 스크립트에서 --vllm-url과 --model 인자를 지정하여 실행합니다.
python batch_enrich_markdown.py \
--input-dir ./docs_markdown \
--output-dir ./vectordb_dataset \
--final-json bulk_chunks_for_vectordb.json \
--component "directpv" \
--vllm-url "http://localhost:8000/v1" \
--model "/data/models/Qwen2.5-32B-Instruct" \
--concurrency 64
성능 팁: B300 환경에서는
--concurrency를64~128수준으로 올려도 vLLM의 PagedAttention과 Continuous Batching 덕분에 GPU 연산 유닛이 풀로 돌면서 수천 개 청크 처리가 순식간에 끝납니다.
==
vllm/vllm-openai:latest는 레지스트리 도메인이 생략되어 있으므로, Docker의 기본 레지스트리인 Docker Hub(docker.io)에서 가져옵니다.
전체 정규 주소(Full Image Name)는 다음과 같습니다:
docker.io/vllm/vllm-openai:latest (또는 특정 버전 태그 권장: vllm/vllm-openai:v0.6.x 등)ghcr.io)에도 빌드 이미지를 배포합니다.ghcr.io/vllm-project/vllm:latest# 외부망에서 다운로드 및 저장
docker pull vllm/vllm-openai:latest
docker save -o vllm-openai-latest.tar vllm/vllm-openai:latest
# 폐쇄망 반입 후 로드
docker load -i vllm-openai-latest.tar
# 사내 레지스트리에 등록 시
docker tag vllm/vllm-openai:latest harbor.internal.net/ai-platform/vllm-openai:latest
docker push harbor.internal.net/ai-platform/vllm-openai:latest
latest보다는 Blackwell(Compute Capability 10.0/12.0 계열)과 CUDA 12.4+ 환경이 정상 패키징된 공식 릴리스 태그를 확인하고 다운로드하는 것을 권장합니다.===
RCA(근본 원인 분석)에 특화된 5가지 확장 메타데이터(doc_category, symptoms_and_errors, related_metrics, target_resources, is_destructive_action)를 vLLM에서 고속 비동기로 추출하고 청크 헤더와 JSON에 반영하도록 보강한 전체 스크립트 파일을 생성하겠습니다.
python batch_enrich_markdown.py \
--input-dir ./docs_markdown \
--output-dir ./vectordb_dataset \
--final-json bulk_chunks_for_vectordb.json \
--component "directpv" \
--vllm-url "http://localhost:8000/v1" \
--model "/data/models/Qwen2.5-32B-Instruct" \
--concurrency 48
bulk_chunks_for_vectordb.json의 각 청크는 아래와 같이 메타데이터 필터링(doc_category, related_metrics, is_destructive_action 등)에 완벽히 대응할 수 있는 구조로 저장됩니다.{
"id": "directpv_troubleshooting_drives.md_0001",
"page_content": "[Context: Troubleshooting > Drive Discovery]\n[Category: TROUBLESHOOTING] [Target Resources: DirectPVDrive, Node]\n[Summary: DirectPV 디스크 자동 탐색 실패 시 해결 절차]\n[Symptoms/Errors: Drive formatted status stuck in pending, Discovery timeout]\n- Q: directpv에서 신규 드라이브가 인식되지 않을 때 확인 커맨드는?\n- Q: kubectl directpv drives discover 타임아웃 해결 방법\n\n## Drive Discovery Failure\nDirectPV discovers drives...",
"metadata": {
"component": "directpv",
"doc_category": "TROUBLESHOOTING",
"breadcrumb": "Troubleshooting > Drive Discovery",
"summary": "DirectPV 디스크 자동 탐색 실패 시 해결 절차",
"questions": [
"directpv에서 신규 드라이브가 인식되지 않을 때 확인 커맨드는?",
"kubectl directpv drives discover 타임아웃 해결 방법"
],
"symptoms_and_errors": [
"Drive formatted status stuck in pending",
"Discovery timeout"
],
"related_metrics": [
"minio_node_drive_total_duration_microseconds"
],
"target_resources": [
"DirectPVDrive",
"Node"
],
"commands": [
"kubectl directpv drives discover"
],
"is_destructive_action": false,
"h1": "Troubleshooting",
"h2": "Drive Discovery"
}
}
===
사내 Confluence에 저장된 장애 이력(Post-mortem) 및 SOP 문서를 REST API(단일 Page ID 또는 특정 Space 내 전체 검색 CQL)로 일괄 조회하여 Markdown으로 변환하고, 앞서 설계한 B300 vLLM을 활용해 RCA 특화 메타데이터와 가상 질문을 보강한 뒤 단일 JSON으로 묶어주는 통합 파이프라인 스크립트 파일을 생성하겠습니다.
root_cause_hypotheses (알려진 근본 원인) 자동 추출:--page-id) 및 Space 일괄(--space) 지원:--space SRE_STORAGE)를 한 번에 긁어모아 변환할 수도 있습니다.raw_confluence_md/ 폴더에 깔끔하게 변환된 .md 파일들을 남겨두어, 사람이 직접 열람하거나 다른 용도로 활용할 수 있습니다.# 1. 특정 장애 보고서/SOP 단일 페이지 변환 및 vLLM 보강
python confluence_to_vectordb_pipeline.py \
--url "https://confluence.mycorp.internal" \
--page-id "98765432" \
--token "your-confluence-pat-token" \
--no-verify-ssl \
--component "minio-aistor" \
--vllm-url "http://localhost:8000/v1" \
--model "/data/models/Qwen2.5-32B-Instruct" \
--concurrency 48
# 2. 특정 Space의 장애 매뉴얼 전체 일괄 변환
python confluence_to_vectordb_pipeline.py \
--url "https://confluence.mycorp.internal" \
--space "STORAGE_SOP" \
--token "your-confluence-pat-token" \
--no-verify-ssl \
--component "minio-aistor" \
--output-dir ./storage_sop_dataset \
--final-json bulk_sop_chunks.json