26S26f1

QK·2026년 9월 27일

질문하신 접근 방식은 대규모 Vector DB 적재 시 가장 권장되는 표준 베스트 프랙티스(Best Practice)입니다.

왜 이 방식이 좋은가?

  1. Vector DB 일괄 적재(Bulk Insert / Batch Upsert) 최적화:
  • 청크 파일마다 수천 번 네트워크 I/O 및 트랜잭션을 날리는 것보다, 모든 청크가 담긴 단일 JSON 배열(List[Dict])을 500~1,000개 단위로 배치 임베딩 및 업서트하는 것이 속도와 자원 효율 측면에서 수십 배 빠릅니다.
  1. 단일 실패 격리 및 재시도(Checkpointing):
  • 문서 수백 개를 변환하는 도중 특정 파일 파싱이나 LLM 요청에서 에러가 발생하더라도, 파일별 개별 임시 JSON을 남겨두면 작업 중단 시 처음부터 다시 돌릴 필요 없이 이어서 병합할 수 있습니다.

지정한 디렉터리(--input-dir) 내의 모든 .md 파일을 재귀 탐색하여 청킹 및 vLLM 가상 질문/메타데이터 주입을 수행하고, 최종적으로 Vector DB 대량 적재용 단일 JSON 파일(bulk_chunks_for_vectordb.json)로 통합해 주는 전체 파이프라인 스크립트 파일을 생성하겠습니다.


실행 방법 및 활용 안내

1. 실행 예시

수집해 둔 Markdown 폴더(docs_markdown/)를 지정하여 실행합니다.

python batch_enrich_markdown.py \
  --input-dir ./docs_markdown \
  --output-dir ./vectordb_dataset \
  --final-json bulk_chunks_for_vectordb.json \
  --component "directpv" \
  --vllm-url "http://localhost:8000/v1" \
  --model "Qwen/Qwen2.5-14B-Instruct" \
  --concurrency 48

2. 최종 생성되는 단일 통합 JSON의 구조

최종 산출물인 bulk_chunks_for_vectordb.json은 모든 문서의 청크들이 하나의 배열로 묶여 있어, 사내 Vector DB 적재 스크립트에서 단 3~4줄로 즉시 대량 적재(Batch Upsert)가 가능합니다.

[
  {
    "id": "directpv_troubleshooting_drives.md_0001",
    "page_content": "[Context: Troubleshooting > Drive Discovery]\n[Summary: DirectPV 디스크 자동 탐색 실패 시 해결 절차]\n- Q: directpv에서 신규 드라이브가 인식되지 않을 때 확인 커맨드는?\n- Q: kubectl directpv drives discover 타임아웃 해결 방법\n\n## Drive Discovery Failure\nDirectPV discovers drives...",
    "metadata": {
      "component": "directpv",
      "source_url": "https://docs.min.io/directpv/troubleshooting/drives/",
      "breadcrumb": "Troubleshooting > Drive Discovery",
      "summary": "DirectPV 디스크 자동 탐색 실패 시 해결 절차",
      "questions": [
        "directpv에서 신규 드라이브가 인식되지 않을 때 확인 커맨드는?",
        "kubectl directpv drives discover 타임아웃 해결 방법"
      ],
      "error_keywords": ["DriveDiscoveryFailed", "Timeout"],
      "commands": ["kubectl directpv drives discover"],
      "h1": "Troubleshooting",
      "h2": "Drive Discovery"
    }
  }
]

3. Vector DB 적재 시 활용 예시 (Python 팁)

import json
# 1. 파일 한번에 로드
with open("./vectordb_dataset/bulk_chunks_for_vectordb.json", "r", encoding="utf-8") as f:
    chunks = json.load(f)

# 2. 500개씩 묶어 Vector DB에 Bulk Insert (임베딩 및 업서트)
batch_size = 500
for i in range(0, len(chunks), batch_size):
    batch = chunks[i:i + batch_size]
    # vector_db.upsert(batch)

===

TEST

B300(NVIDIA Blackwell 아키텍처 기반의 고성능 GPU) 환경에 vLLM을 올려두었다면, 처리 성능과 추론 속도가 매우 압도적이므로 해당 작업을 수행하기에 최적의 환경입니다.

문서 전처리 파이프라인은 "결정론적 텍스트 분할(청킹/규칙 기반 코드 보호)"을 파이썬으로 먼저 수행한 뒤, 분할된 각 청크를 "vLLM을 통한 비동기 고속 배치 추론(가상 질문 생성 + 메타데이터 추출)"으로 보강하는 구조로 설계하는 것이 가장 효율적이고 비용/시간이 절약됩니다.

구체적인 설계 원리와 파이프라인 구현 방법을 안내해 드립니다.


1. B300 + vLLM으로 이 작업이 가능한 이유 & 최적화 포인트

  • B300의 대용량 HBM3e 메모리 및 연산 성능: 수백~수천 개의 청크에 대해 LLM 추론을 돌릴 때 단건 호출을 하면 시간이 걸리지만, vLLM의 Continuous Batching(연속 배치 처리)과 Async 엔진을 활용하면 수천 개 청크의 가상 질문/메타데이터 생성을 수 분 내에 완료할 수 있습니다.
  • 권장 모델: 폐쇄망 로컬 환경이라면 코딩 및 다국어(한국어/영어)에 강하고 Function Calling/JSON 출력이 안정적인 Qwen2.5-14B/32B-Instruct 또는 Llama-3.1-8B/70B-Instruct를 FP8 또는 BF16으로 띄우면 최고의 가성비와 정확도를 얻을 수 있습니다.
  • 구조적 분리 원칙:
  • 3) 헤더 청킹 및 코드 블록 보호: LLM에 맡기지 않고 Python 정규식/AST 파서로 처리 (비용 0, 토큰 누락 0%, 100% 결정론적 동작).
  • 1) 가상 질문 삽입 & 2) 메타데이터 강화: Python으로 쪼개진 각 청크를 vLLM 비동기 배치로 전달해 병렬 추출.

2. 단계별 핵심 로직 설계

(1) 코드 블록 & CLI 보호형 헤더 청킹 (Chunking Strategy)

Markdown을 #, ##, ### 단위로 쪼갤 때 bash ... 같은 코드 블록이 중간에 잘려 나가는 것을 방지하는 기법입니다.

  1. MarkdownHeaderTextSplitter를 사용해 상위 계층 구조(Breadcrumb)를 메타데이터로 추출.
  2. 2차 분할 시 RecursiveCharacterTextSplitter의 구분자(separators)를 조정하여 코드 블록 닫힘 기호(\n```\n)나 문단(\n\n)을 최우선 분할점으로 설정.
  3. 청크의 시작이 코드 블록 내부에서 끊기지 않도록 keep_separator=True를 적용하고, 각 청크의 상단에 계층형 헤더 경로(예: Path: MinIO > DirectPV > Troubleshooting)를 강제 주입.

(2) vLLM Structured Output을 이용한 가상 질문 및 메타데이터 추출

청크 텍스트를 vLLM에 입력하여 JSON Schema 형태로 아래 항목을 한 번에 뽑아냅니다.

  • hypothetical_questions: "이 문서를 검색할 운영자가 던질 법한 가상 질문" (HyDE 효과 → 검색 유사도 극대화)
  • component: 대상 소프트웨어 컴포넌트
  • error_keywords: 에러 코드, 상태 메시지 (예: DriveNotReady, Timeout, 503)
  • cli_commands: 본문에 포함된 명령어 목록 (kubectl directpv ...)

3. 엔드투엔드 파이프라인 구현 코드 (preprocess_chunks_vllm.py)

로컬 B300 서버에서 vLLM 인라인 엔진을 직접 로드하거나, 별도로 띄워둔 vLLM OpenAI-Compatible API 서버를 호출하여 비동기 처리하는 스크립트입니다.

import os
import re
import json
import asyncio
from typing import List, Dict, Any
from pydantic import BaseModel, Field
from openai import AsyncOpenAI
from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter

# ==========================================
# 0. 설정
# ==========================================
VLLM_API_BASE = os.getenv("VLLM_API_BASE", "http://localhost:8000/v1")
VLLM_API_KEY = "EMPTY"  # vLLM 기본값
MODEL_NAME = "Qwen/Qwen2.5-14B-Instruct"  # 또는 Llama-3.1-8B-Instruct 등

client = AsyncOpenAI(base_url=VLLM_API_BASE, api_key=VLLM_API_KEY)


# ==========================================
# 1. 코드 블록 보호형 계층적 청킹
# ==========================================
def split_markdown_preserving_code(md_text: str, default_component: str = "directpv") -> List[Dict[str, Any]]:
    """
    1) Markdown 헤더 계층 보존
    2) 코드 블록이 중간에 잘리지 않도록 보호하며 청킹
    """
    headers_to_split = [
        ("#", "h1"),
        ("##", "h2"),
        ("###", "h3"),
        ("####", "h4"),
    ]
    header_splitter = MarkdownHeaderTextSplitter(
        headers_to_split_on=headers_to_split, 
        strip_headers=False
    )
    section_docs = header_splitter.split_text(md_text)

    # 코드 블록/문단 우선 분할기
    # 코드 블록 닫는 ``` 기호나 빈 줄 단위로 우선 분할
    text_splitter = RecursiveCharacterTextSplitter(
        chunk_size=1200,
        chunk_overlap=150,
        separators=[
            "\n```\n",      # 코드 블록 경계 최우선 분할
            "\n```bash\n",
            "\n```yaml\n",
            "\n\n",         # 문단 단위
            "\n",           # 줄 단위
            " "
        ],
        keep_separator=True
    )

    final_chunks = []
    chunk_index = 0

    for section in section_docs:
        sub_docs = text_splitter.split_documents([section])
        
        # 상위 헤더 경로 구성 (Breadcrumbs)
        headers = [section.metadata.get(h) for h in ["h1", "h2", "h3", "h4"] if section.metadata.get(h)]
        breadcrumb_path = " > ".join(headers) if headers else "Root"

        for doc in sub_docs:
            chunk_index += 1
            final_chunks.append({
                "chunk_id": f"{default_component}_chunk_{chunk_index:04d}",
                "component": default_component,
                "breadcrumb": breadcrumb_path,
                "content": doc.page_content,
                "headers": section.metadata
            })

    return final_chunks


# ==========================================
# 2. vLLM 비동기 가상 질문 & 메타데이터 추출
# ==========================================
PROMPT_TEMPLATE = """당신은 대규모 인프라 및 쿠버네티스 SRE/DevOps 전문가입니다.
주어진 기술 문서 청크를 분석하여, 운영자가 문제 해결 시 검색할 법한 가상 질문과 핵심 메타데이터를 JSON 형식으로 추출하십시오.

[문서 헤더 경로]
{breadcrumb}

[문서 내용]
{content}

[추출 규칙]
1. questions: 장애 조치, 에러 해결, 설정 질문 위주로 2~3개 한국어/영어 혼용 생성
2. error_keywords: 본문에 언급된 에러 코드, 상태(예: DriveNotReady, 503, Offline 등)
3. commands: 본문에 포함된 CLI 명령어 프리픽스 (예: `kubectl directpv`, `mc admin` 등)
4. summary: 이 청크의 1줄 요약

반드시 유효한 JSON 문자열만 출력하십시오. 마크다운 코드블록이나 불필요한 설명은 금지합니다.
{{
  "summary": "1줄 요약",
  "questions": ["가상 질문 1", "가상 질문 2"],
  "error_keywords": ["키워드1", "키워드2"],
  "commands": ["명령어1"]
}}"""

async def enrich_chunk_with_vllm(chunk: Dict[str, Any], semaphore: asyncio.Semaphore) -> Dict[str, Any]:
    """vLLM 비동기 API 호출로 청크 단위 메타데이터 보강"""
    async with semaphore:
        prompt = PROMPT_TEMPLATE.format(
            breadcrumb=chunk["breadcrumb"],
            content=chunk["content"]
        )
        try:
            response = await client.chat.completions.create(
                model=MODEL_NAME,
                messages=[{"role": "user", "content": prompt}],
                temperature=0.1,
                max_tokens=512,
                # JSON 모드 강제 지원 시 사용
                response_format={"type": "json_object"}
            )
            raw_json = response.choices[0].message.content
            enrichment_data = json.loads(raw_json)
        except Exception as e:
            # LLM 파싱 실패 시 기본 Fallback
            enrichment_data = {
                "summary": chunk["breadcrumb"],
                "questions": [],
                "error_keywords": [],
                "commands": []
            }

        # 1) Chunk Header에 가상 질문 및 요약 삽입
        questions_block = "\n".join([f"- Q: {q}" for q in enrichment_data.get("questions", [])])
        enriched_content = f"""[Context: {chunk['breadcrumb']}]
[Summary: {enrichment_data.get('summary', '')}]
{questions_block}

{chunk['content']}
"""

        # 2) 메타데이터 보강
        return {
            "chunk_id": chunk["chunk_id"],
            "page_content": enriched_content.strip(),
            "metadata": {
                "component": chunk["component"],
                "breadcrumb": chunk["breadcrumb"],
                "summary": enrichment_data.get("summary", ""),
                "questions": enrichment_data.get("questions", []),
                "error_keywords": enrichment_data.get("error_keywords", []),
                "commands": enrichment_data.get("commands", []),
                **chunk["headers"]
            }
        }

async def process_all_chunks(chunks: List[Dict[str, Any]], concurrency: int = 32):
    # B300 환경에서는 동시 요청(concurrency) 32~64 이상 충분히 소화 가능
    semaphore = asyncio.Semaphore(concurrency)
    tasks = [enrich_chunk_with_vllm(c, semaphore) for c in chunks]
    return await asyncio.gather(*tasks)


# ==========================================
# 3. 메인 실행 흐름
# ==========================================
def main():
    sample_md_path = "docs_markdown/directpv/troubleshooting_drives.md"
    if not os.path.exists(sample_md_path):
        print(f"[!] 파일을 찾을 수 없습니다: {sample_md_path}")
        return

    with open(sample_md_path, "r", encoding="utf-8") as f:
        md_text = f.read()

    print("[1] 코드 보존형 계층 청킹 진행 중...")
    initial_chunks = split_markdown_preserving_code(md_text, default_component="directpv")
    print(f"[*] 총 {len(initial_chunks)}개 청크 생성 완료.")

    print(f"[2] B300 vLLM을 이용한 비동기 Enrichment (가상 질문 + 메타데이터) 시작...")
    loop = asyncio.get_event_loop()
    enriched_results = loop.run_until_complete(process_all_chunks(initial_chunks, concurrency=32))

    # 결과 검증 및 저장
    output_json = "enriched_chunks_for_vectordb.json"
    with open(output_json, "w", encoding="utf-8") as f:
        json.dump(enriched_results, f, ensure_ascii=False, indent=2)

    print(f"[✓] 전처리 완료! {output_json} 파일에 저장되었습니다.")
    print("\n--- [샘플 청크 1번 결과 확인] ---")
    print("Content:\n", enriched_results[0]["page_content"][:400], "...\n")
    print("Metadata:\n", json.dumps(enriched_results[0]["metadata"], ensure_ascii=False, indent=2))

if __name__ == "__main__":
    main()

4. Vector DB 적재 시 얻게 되는 이점

이렇게 전처리된 청크를 적재하면 Vector DB 검색 단계에서 놀라운 효과를 보입니다.

  1. HyDE 효과 (가상 질문 매칭):
  • 운영자가 "directpv 드라이브 포맷 중 timeout 날 때 어떻게 해?"라고 구어체로 질문해도, 청크 상단에 주입된 - Q: directpv 포맷 실패 및 타임아웃 발생 시 조치 방법과 벡터 공간에서 즉각 매칭됩니다.
  1. Hybrid Search (키워드 하드 필터링):
  • Agent가 에러 로그에서 DriveNotReady를 감지했을 때, Vector DB 검색 쿼리에 filter={"error_keywords": {"$contains": "DriveNotReady"}}를 걸 수 있어 엉뚱한 설치 문서가 검색되는 문제를 100% 방어합니다.
  1. 코드/명령어 보존:
  • 청크 분할 시 코드 블록 문법이 깨지지 않으므로, RCA Agent가 최종 리포트를 쓸 때 실제 실행 가능한 정확한 복구 명령어를 제시할 수 있습니다.
profile
engineer

0개의 댓글