질문하신 접근 방식은 대규모 Vector DB 적재 시 가장 권장되는 표준 베스트 프랙티스(Best Practice)입니다.
List[Dict])을 500~1,000개 단위로 배치 임베딩 및 업서트하는 것이 속도와 자원 효율 측면에서 수십 배 빠릅니다.지정한 디렉터리(--input-dir) 내의 모든 .md 파일을 재귀 탐색하여 청킹 및 vLLM 가상 질문/메타데이터 주입을 수행하고, 최종적으로 Vector DB 대량 적재용 단일 JSON 파일(bulk_chunks_for_vectordb.json)로 통합해 주는 전체 파이프라인 스크립트 파일을 생성하겠습니다.
수집해 둔 Markdown 폴더(docs_markdown/)를 지정하여 실행합니다.
python batch_enrich_markdown.py \
--input-dir ./docs_markdown \
--output-dir ./vectordb_dataset \
--final-json bulk_chunks_for_vectordb.json \
--component "directpv" \
--vllm-url "http://localhost:8000/v1" \
--model "Qwen/Qwen2.5-14B-Instruct" \
--concurrency 48
최종 산출물인 bulk_chunks_for_vectordb.json은 모든 문서의 청크들이 하나의 배열로 묶여 있어, 사내 Vector DB 적재 스크립트에서 단 3~4줄로 즉시 대량 적재(Batch Upsert)가 가능합니다.
[
{
"id": "directpv_troubleshooting_drives.md_0001",
"page_content": "[Context: Troubleshooting > Drive Discovery]\n[Summary: DirectPV 디스크 자동 탐색 실패 시 해결 절차]\n- Q: directpv에서 신규 드라이브가 인식되지 않을 때 확인 커맨드는?\n- Q: kubectl directpv drives discover 타임아웃 해결 방법\n\n## Drive Discovery Failure\nDirectPV discovers drives...",
"metadata": {
"component": "directpv",
"source_url": "https://docs.min.io/directpv/troubleshooting/drives/",
"breadcrumb": "Troubleshooting > Drive Discovery",
"summary": "DirectPV 디스크 자동 탐색 실패 시 해결 절차",
"questions": [
"directpv에서 신규 드라이브가 인식되지 않을 때 확인 커맨드는?",
"kubectl directpv drives discover 타임아웃 해결 방법"
],
"error_keywords": ["DriveDiscoveryFailed", "Timeout"],
"commands": ["kubectl directpv drives discover"],
"h1": "Troubleshooting",
"h2": "Drive Discovery"
}
}
]
import json
# 1. 파일 한번에 로드
with open("./vectordb_dataset/bulk_chunks_for_vectordb.json", "r", encoding="utf-8") as f:
chunks = json.load(f)
# 2. 500개씩 묶어 Vector DB에 Bulk Insert (임베딩 및 업서트)
batch_size = 500
for i in range(0, len(chunks), batch_size):
batch = chunks[i:i + batch_size]
# vector_db.upsert(batch)
===
TEST
B300(NVIDIA Blackwell 아키텍처 기반의 고성능 GPU) 환경에 vLLM을 올려두었다면, 처리 성능과 추론 속도가 매우 압도적이므로 해당 작업을 수행하기에 최적의 환경입니다.
문서 전처리 파이프라인은 "결정론적 텍스트 분할(청킹/규칙 기반 코드 보호)"을 파이썬으로 먼저 수행한 뒤, 분할된 각 청크를 "vLLM을 통한 비동기 고속 배치 추론(가상 질문 생성 + 메타데이터 추출)"으로 보강하는 구조로 설계하는 것이 가장 효율적이고 비용/시간이 절약됩니다.
구체적인 설계 원리와 파이프라인 구현 방법을 안내해 드립니다.
Markdown을 #, ##, ### 단위로 쪼갤 때 bash ... 같은 코드 블록이 중간에 잘려 나가는 것을 방지하는 기법입니다.
MarkdownHeaderTextSplitter를 사용해 상위 계층 구조(Breadcrumb)를 메타데이터로 추출.RecursiveCharacterTextSplitter의 구분자(separators)를 조정하여 코드 블록 닫힘 기호(\n```\n)나 문단(\n\n)을 최우선 분할점으로 설정.keep_separator=True를 적용하고, 각 청크의 상단에 계층형 헤더 경로(예: Path: MinIO > DirectPV > Troubleshooting)를 강제 주입.청크 텍스트를 vLLM에 입력하여 JSON Schema 형태로 아래 항목을 한 번에 뽑아냅니다.
hypothetical_questions: "이 문서를 검색할 운영자가 던질 법한 가상 질문" (HyDE 효과 → 검색 유사도 극대화)component: 대상 소프트웨어 컴포넌트error_keywords: 에러 코드, 상태 메시지 (예: DriveNotReady, Timeout, 503)cli_commands: 본문에 포함된 명령어 목록 (kubectl directpv ...)preprocess_chunks_vllm.py)로컬 B300 서버에서 vLLM 인라인 엔진을 직접 로드하거나, 별도로 띄워둔 vLLM OpenAI-Compatible API 서버를 호출하여 비동기 처리하는 스크립트입니다.
import os
import re
import json
import asyncio
from typing import List, Dict, Any
from pydantic import BaseModel, Field
from openai import AsyncOpenAI
from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter
# ==========================================
# 0. 설정
# ==========================================
VLLM_API_BASE = os.getenv("VLLM_API_BASE", "http://localhost:8000/v1")
VLLM_API_KEY = "EMPTY" # vLLM 기본값
MODEL_NAME = "Qwen/Qwen2.5-14B-Instruct" # 또는 Llama-3.1-8B-Instruct 등
client = AsyncOpenAI(base_url=VLLM_API_BASE, api_key=VLLM_API_KEY)
# ==========================================
# 1. 코드 블록 보호형 계층적 청킹
# ==========================================
def split_markdown_preserving_code(md_text: str, default_component: str = "directpv") -> List[Dict[str, Any]]:
"""
1) Markdown 헤더 계층 보존
2) 코드 블록이 중간에 잘리지 않도록 보호하며 청킹
"""
headers_to_split = [
("#", "h1"),
("##", "h2"),
("###", "h3"),
("####", "h4"),
]
header_splitter = MarkdownHeaderTextSplitter(
headers_to_split_on=headers_to_split,
strip_headers=False
)
section_docs = header_splitter.split_text(md_text)
# 코드 블록/문단 우선 분할기
# 코드 블록 닫는 ``` 기호나 빈 줄 단위로 우선 분할
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=1200,
chunk_overlap=150,
separators=[
"\n```\n", # 코드 블록 경계 최우선 분할
"\n```bash\n",
"\n```yaml\n",
"\n\n", # 문단 단위
"\n", # 줄 단위
" "
],
keep_separator=True
)
final_chunks = []
chunk_index = 0
for section in section_docs:
sub_docs = text_splitter.split_documents([section])
# 상위 헤더 경로 구성 (Breadcrumbs)
headers = [section.metadata.get(h) for h in ["h1", "h2", "h3", "h4"] if section.metadata.get(h)]
breadcrumb_path = " > ".join(headers) if headers else "Root"
for doc in sub_docs:
chunk_index += 1
final_chunks.append({
"chunk_id": f"{default_component}_chunk_{chunk_index:04d}",
"component": default_component,
"breadcrumb": breadcrumb_path,
"content": doc.page_content,
"headers": section.metadata
})
return final_chunks
# ==========================================
# 2. vLLM 비동기 가상 질문 & 메타데이터 추출
# ==========================================
PROMPT_TEMPLATE = """당신은 대규모 인프라 및 쿠버네티스 SRE/DevOps 전문가입니다.
주어진 기술 문서 청크를 분석하여, 운영자가 문제 해결 시 검색할 법한 가상 질문과 핵심 메타데이터를 JSON 형식으로 추출하십시오.
[문서 헤더 경로]
{breadcrumb}
[문서 내용]
{content}
[추출 규칙]
1. questions: 장애 조치, 에러 해결, 설정 질문 위주로 2~3개 한국어/영어 혼용 생성
2. error_keywords: 본문에 언급된 에러 코드, 상태(예: DriveNotReady, 503, Offline 등)
3. commands: 본문에 포함된 CLI 명령어 프리픽스 (예: `kubectl directpv`, `mc admin` 등)
4. summary: 이 청크의 1줄 요약
반드시 유효한 JSON 문자열만 출력하십시오. 마크다운 코드블록이나 불필요한 설명은 금지합니다.
{{
"summary": "1줄 요약",
"questions": ["가상 질문 1", "가상 질문 2"],
"error_keywords": ["키워드1", "키워드2"],
"commands": ["명령어1"]
}}"""
async def enrich_chunk_with_vllm(chunk: Dict[str, Any], semaphore: asyncio.Semaphore) -> Dict[str, Any]:
"""vLLM 비동기 API 호출로 청크 단위 메타데이터 보강"""
async with semaphore:
prompt = PROMPT_TEMPLATE.format(
breadcrumb=chunk["breadcrumb"],
content=chunk["content"]
)
try:
response = await client.chat.completions.create(
model=MODEL_NAME,
messages=[{"role": "user", "content": prompt}],
temperature=0.1,
max_tokens=512,
# JSON 모드 강제 지원 시 사용
response_format={"type": "json_object"}
)
raw_json = response.choices[0].message.content
enrichment_data = json.loads(raw_json)
except Exception as e:
# LLM 파싱 실패 시 기본 Fallback
enrichment_data = {
"summary": chunk["breadcrumb"],
"questions": [],
"error_keywords": [],
"commands": []
}
# 1) Chunk Header에 가상 질문 및 요약 삽입
questions_block = "\n".join([f"- Q: {q}" for q in enrichment_data.get("questions", [])])
enriched_content = f"""[Context: {chunk['breadcrumb']}]
[Summary: {enrichment_data.get('summary', '')}]
{questions_block}
{chunk['content']}
"""
# 2) 메타데이터 보강
return {
"chunk_id": chunk["chunk_id"],
"page_content": enriched_content.strip(),
"metadata": {
"component": chunk["component"],
"breadcrumb": chunk["breadcrumb"],
"summary": enrichment_data.get("summary", ""),
"questions": enrichment_data.get("questions", []),
"error_keywords": enrichment_data.get("error_keywords", []),
"commands": enrichment_data.get("commands", []),
**chunk["headers"]
}
}
async def process_all_chunks(chunks: List[Dict[str, Any]], concurrency: int = 32):
# B300 환경에서는 동시 요청(concurrency) 32~64 이상 충분히 소화 가능
semaphore = asyncio.Semaphore(concurrency)
tasks = [enrich_chunk_with_vllm(c, semaphore) for c in chunks]
return await asyncio.gather(*tasks)
# ==========================================
# 3. 메인 실행 흐름
# ==========================================
def main():
sample_md_path = "docs_markdown/directpv/troubleshooting_drives.md"
if not os.path.exists(sample_md_path):
print(f"[!] 파일을 찾을 수 없습니다: {sample_md_path}")
return
with open(sample_md_path, "r", encoding="utf-8") as f:
md_text = f.read()
print("[1] 코드 보존형 계층 청킹 진행 중...")
initial_chunks = split_markdown_preserving_code(md_text, default_component="directpv")
print(f"[*] 총 {len(initial_chunks)}개 청크 생성 완료.")
print(f"[2] B300 vLLM을 이용한 비동기 Enrichment (가상 질문 + 메타데이터) 시작...")
loop = asyncio.get_event_loop()
enriched_results = loop.run_until_complete(process_all_chunks(initial_chunks, concurrency=32))
# 결과 검증 및 저장
output_json = "enriched_chunks_for_vectordb.json"
with open(output_json, "w", encoding="utf-8") as f:
json.dump(enriched_results, f, ensure_ascii=False, indent=2)
print(f"[✓] 전처리 완료! {output_json} 파일에 저장되었습니다.")
print("\n--- [샘플 청크 1번 결과 확인] ---")
print("Content:\n", enriched_results[0]["page_content"][:400], "...\n")
print("Metadata:\n", json.dumps(enriched_results[0]["metadata"], ensure_ascii=False, indent=2))
if __name__ == "__main__":
main()
이렇게 전처리된 청크를 적재하면 Vector DB 검색 단계에서 놀라운 효과를 보입니다.
"directpv 드라이브 포맷 중 timeout 날 때 어떻게 해?"라고 구어체로 질문해도, 청크 상단에 주입된 - Q: directpv 포맷 실패 및 타임아웃 발생 시 조치 방법과 벡터 공간에서 즉각 매칭됩니다.DriveNotReady를 감지했을 때, Vector DB 검색 쿼리에 filter={"error_keywords": {"$contains": "DriveNotReady"}}를 걸 수 있어 엉뚱한 설치 문서가 검색되는 문제를 100% 방어합니다.