26S26i

QK·6일 전

사내 자체 워크플로우 빌더 도구(GUI 기반 노드 배치 또는 YAML/JSON 기반 파이프라인 정의)에 그대로 이식할 수 있도록, 코드에 의존하지 않고 워크플로우의 순서, 노드별 역할/입출력 명세, 내부 프롬프트 및 판단 로직을 중심으로 상세히 설계했습니다.


Incident RCA Agent 워크플로우 아키텍처

[Fake/Real Incident JSON] 
         │
         ▼
[Step 1. Symptom Normalizer & Entity Extractor] (LLM)
         │
         ▼
[Step 2. Knowledge Base Query Planner & Retriever] (Search/DB Tool)
         │
         ▼
[Step 3. RCA Hypothesis Generator] (LLM)
         │
         ▼
[Step 4. Evidence Cross-Checker & Verifier] (LLM)
         │
         ├──────────────────────────┐
         ▼ (모든 가설 기각 & 루프 < 2)  ▼ (원인 확정 OR 루프 종료)
[Step 4-A. Hypothesis Refiner]       [Step 5. RCA Report & Runbook Generator] (LLM)
         │                                  │
         └──────────────────────────────────┤
                                            ▼
                                     [최종 결과 리포트]

Step 1. 증상 정규화 및 이상 엔티티 추출 노드 (Symptom Normalizer)

  • 노드 유형: LLM Node (JSON Mode / 구조화 출력 권장)
  • 목적: 원시 Incident JSON(수백 줄의 로그, 시계열 메트릭 스냅샷, K8s Event)의 노이즈를 걷어내고, 이상(Anomaly)을 보이는 인프라 리소스와 주요 현상을 정규화된 키-값으로 압축합니다.
  • 입력 데이터 (Input):
  • incident_payload: 인시던트 JSON 원문
  • 프롬프트 내용 (Prompt):

    역할: 대규모 분산 스토리지/K8s 전문 SRE 엔지니어
    지시사항:

    1. 주어진 Incident 데이터에서 명백한 이상 징후(임계치 초과 메트릭, ERROR/WARN 로그)를 감지하십시오.
    2. 문제가 발생한 특정 타깃 엔티티(노드명, 파드명, 드라이브 마운트 경로, 에러 코드)를 빠짐없이 추출하십시오.
    3. 과거 장애 보고서 및 SOP를 검색할 때 사용할 핵심 검색 키워드(영문/한글) 3개를 생성하십시오.
  • 출력 데이터 (Output State):
  • target_component: 컴포넌트 명 (예: minio-aistor, directpv)
  • anomaly_entities: [{"type": "node", "name": "storage-worker-node-04"}, {"type": "drive", "path": "/mnt/directpv/nvme03"}, {"type": "error", "code": "503 Slow Down"}]
  • symptom_summary: 발생 증상 1~2줄 요약 (예: "특정 NVMe I/O 지연으로 인한 클라이언트 503 반환")
  • search_queries: ["MinIO drive slow down timeout", "DirectPV NVMe I/O delay latency", "503 Slow Down PutObject"]

Step 2. 지식 검색 플래너 및 리트리버 노드 (Knowledge Retriever)

  • 노드 유형: Tool / DB Integration Node (사내 Vector DB 연동 노드)
  • 목적: 앞선 스텝에서 추출한 메타데이터와 검색어를 이용해 앞서 적재해 둔 bulk_chunks_for_vectordb.json 지식 베이스에서 가장 관련성 높은 SOP, 과거 포스트모템, 트러블슈팅 가이드를 추출합니다.
  • 동작 방식 (Logic):
  1. 1차 메타데이터 하드 필터링:
  • component == target_component (예: minio-aistor 또는 directpv)
  • doc_category가 TROUBLESHOOTING 또는 SOP인 문서를 우선 순위에 배치
  1. 2차 시맨틱/키워드 검색 (Top-k=4):
  • search_queries와 임베딩 유사도가 높은 청크 4개 선별
  • 출력 데이터 (Output State):
  • retrieved_knowledge: 검색된 청크 목록 (가상 질문, 요약, 조치 명령어, 과거 근본 원인 포함)

Step 3. 가설 수립 노드 (Hypothesis Generator)

  • 노드 유형: LLM Node
  • 목적: 관측된 증상과 검색된 과거 장애 사례/SOP를 종합하여, 발생 가능한 원인 가설을 상호 배타적(MECE)으로 최대 3개 수립합니다.
  • 입력 데이터 (Input):
  • symptom_summary & anomaly_entities (Step 1 결과)
  • retrieved_knowledge (Step 2 결과)
  • 프롬프트 내용 (Prompt):

    지시사항:
    제공된 증상과 지식 베이스(SOP/과거 장애)를 토대로, 이 장애를 유발했을 가능성이 있는 잠재 원인 가설을 2~3개 도출하십시오.
    [중요 - 검증 규칙(Verification Rule) 명시]:
    각 가설마다 "이 가설이 진실이기 위해 Incident 데이터에서 반드시 확인되어야 하는 조건(필요충분조건)"을 함께 정의하십시오.
    예시:

    • 가설: "스토리지 드라이브 하드웨어 I/O Hang으로 인한 분산 락 지연"
    • 검증 규칙: "특정 단일 드라이브의 I/O Duration 메트릭만 비정상 스파이크를 보여야 하며, 피어 노드에서는 타임아웃 로그가 발생해야 함"
  • 출력 데이터 (Output State):
  • hypotheses: [{"id": "H1", "title": "...", "rule": "..."}, {"id": "H2", "title": "...", "rule": "..."}]

Step 4. 가설 교차 검증 노드 (Evidence Cross-Checker)

  • 노드 유형: LLM Node (엄격 감사 모드)
  • 목적: Step 3에서 세운 가설의 검증 규칙과 Step 1의 원시 로그/메트릭 데이터를 1:1로 대조하여 가설의 참/거짓을 판정합니다. (환각 억제 핵심 단계)
  • 입력 데이터 (Input):
  • hypotheses (Step 3 결과)
  • raw_metrics_summary & raw_logs_summary (최초 Incident JSON)
  • 프롬프트 내용 (Prompt):

    역할: 엄격하고 보수적인 장애 감사관
    지시사항:
    각 가설의 검증 규칙을 실제 메트릭과 로그 데이터와 비교 검증하십시오.
    데이터에 명시적인 근거가 없는 내용은 절대 참으로 인정하지 마십시오.
    판정 기준:

    • CONFIRMED: 데이터(수치, 로그)가 규칙을 100% 뒷받침하고 반증이 없음.
    • DISPROVED: 데이터와 모순되는 명백한 반증 지표가 있음 (예: 디스크 문제라 했는데 I/O latency가 정상).
    • INCONCLUSIVE: 제공된 데이터만으로는 확인이 불가능함.

    각 가설별로 status, confidence (0.0~1.0), supporting_evidence(인용된 로그/메트릭), refuting_evidence를 명시하십시오.

  • 출력 데이터 (Output State):
  • verified_hypotheses: 가설 검증 결과 리스트

Step 4-B. 조건부 분기 (Router / Condition)

  • 노드 유형: Condition / Branch Node (사내 툴의 조건 분기 블록)
  • 분기 조건:
  • 조건 1: verified_hypotheses 중 CONFIRMED 상태인 가설이 1개 이상 존재할 때 ➔ Step 5로 이동
  • 조건 2: 모두 DISPROVED 또는 INCONCLUSIVE이고, 재시도 횟수 < 1회 일 때 ➔ Step 4-A (Hypothesis Refiner)로 이동
  • 조건 3: 재시도 횟수 초과 ➔ Step 5로 이동 (미해결 리포트 모드로 출력)

Step 4-A. 가설 보완 노드 (Hypothesis Refiner - 재시도 루프)

  • 노드 유형: LLM Node
  • 목적: 1차 가설이 모두 기각되었을 때, 기각 사유를 바탕으로 놓친 제3의 원인(예: 애플리케이션 버그성 무한 리스팅, 네트워크 비대칭 단절 등)을 재탐색하여 가설을 업데이트한 뒤 다시 Step 4 검증으로 보냅니다.

Step 5. 최종 RCA 보고서 & 조치 Runbook 생성 노드 (Report Generator)

  • 노드 유형: LLM Node
  • 목적: 검증된 최종 원인과 Step 2에서 확보한 SOP 매뉴얼의 조치 절차를 결합하여, 운영자가 즉시 실행할 수 있는 표준화된 RCA 리포트를 발행합니다.
  • 입력 데이터 (Input):
  • verified_hypotheses (Step 4 결과)
  • retrieved_knowledge (Step 2의 SOP 및 조치 커맨드)
  • incident_metadata (장애 ID, 대상 리소스)
  • 프롬프트 내용 (Prompt):

    검증된 근본 원인과 첨부된 SOP 문서를 기반으로, 아래 4가지 섹션을 갖춘 표준 RCA 리포트를 작성하십시오.

    1. 인시던트 요약: 장애 ID, 발생 시간, 영향도, 관련 타깃 리소스
    2. 근본 원인 분석(RCA):
    • 확정된 원인 (단정적 서술)
    • 결정적 근거 (메트릭 지표 수치, 결정적 에러 로그 라인 명시)
    • 기각된 다른 가설과 기각 사유 (운영 신뢰도 확보용)
    1. 단계별 조치 런북 (Action Runbook):
    • 1단계: 즉시 긴급 조치: SOP에 명시된 CLI 커맨드 제시 (예: mc admin drive offline ... 또는 kubectl directpv ...)
    • 2단계: 완화 확인: 복구 여부를 판별할 확인 커맨드 및 메트릭
    • 3단계: 재발 방지 권고사항: 하드웨어 교체, 타임아웃 값 조정 등
    1. 분석 확신도 및 데이터 한계: 분석 확신도(HIGH/MEDIUM/LOW) 및 추가 수집 필요 지표
  • 출력 데이터:
  • 완성된 Markdown 포맷의 최종 RCA 리포트.

사내 워크플로우 툴 구성 팁

  1. 상태값(Context) 공유: 사내 워크플로우 툴의 글로벌 세션 객체(State/Context Dict)에 incident_payload, entities, retrieved_docs, hypotheses를 누적하여 각 노드가 필요한 필드만 참조할 수 있도록 구성합니다.
  2. Deterministic 검색 유지: 지식 검색(Step 2)은 LLM의 임의 추론에 맡기지 말고 사내 Vector DB 플러그인 노드를 사용해 정확한 필터(component) + 쿼리 매칭으로 동작시켜 검색 안정성을 확보합니다.
  3. 가설 검증의 분리: 가설 생성(Step 3)과 검증(Step 4)을 반드시 별개 노드로 분리해야 단일 LLM 호출 시 발생하는 자기 확증 편향(자기가 만든 가설을 스스로 무조건 맞다고 우기는 현상)을 완벽히 방지할 수 있습니다.

==

네. 이 경우에는 “Incident → Context 수집 → 증거 정규화 → 증상 분류 → RAG 검색 → 가설 생성 → 증거 검증 → 원인 추론 → 영향/범위 판단 → RCA 보고서 → 후속 조치” 형태로 설계하는 것이 좋습니다.

특히 지금 가지고 있는 데이터가

  • RAG Data: AIStor/DirectPV/Cilium/Kubernetes/운영 문서 및 장애 지식
  • Fake Incident Data: K8s Event + Pod Log + Prometheus metric + Incident metadata
  • 향후 실제 데이터: Prometheus/OpenSearch/K8s API/AIStor 로그 등

이라는 점을 고려하면, 단순히 LLM에게 "이 장애 원인이 뭐야?"라고 묻는 구조보다는 각 단계에서 증거를 구조화하고 다음 단계로 넘기는 workflow가 훨씬 좋습니다.


1. 전체 Workflow

제가 권장하는 전체 구조는 다음입니다.

┌───────────────────────┐
│  01 Incident Trigger  │
│  Fake / Real Incident │
└───────────┬───────────┘
            │
            ▼
┌────────────────────────────┐
│ 02 Incident Normalizer     │
│ Incident Context 구조화    │
└────────────┬───────────────┘
             │
             ▼
┌────────────────────────────┐
│ 03 Evidence Extractor      │
│ Log/Event/Metric 분리      │
└────────────┬───────────────┘
             │
             ▼
┌────────────────────────────┐
│ 04 Symptom Analyzer        │
│ "무슨 현상이 발생했나?"    │
└────────────┬───────────────┘
             │
             ├───────────────┐
             ▼               ▼
   ┌────────────────┐  ┌────────────────┐
   │ 05 RAG Search  │  │ 06 Metric      │
   │ Knowledge      │  │ Correlation    │
   └───────┬────────┘  └───────┬────────┘
           │                   │
           └─────────┬─────────┘
                     ▼
          ┌─────────────────────┐
          │ 07 Hypothesis      │
          │ Generator          │
          └──────────┬──────────┘
                     │
                     ▼
          ┌─────────────────────┐
          │ 08 Evidence        │
          │ Validator           │
          └──────────┬──────────┘
                     │
               ┌─────┴─────┐
               │           │
          Evidence OK   Evidence 부족
               │           │
               │           └──────► RAG Search / Context
               │                     Expansion
               ▼
       ┌─────────────────────┐
       │ 09 Causal Analyzer │
       │ Root Cause Graph   │
       └──────────┬──────────┘
                  │
                  ▼
       ┌─────────────────────┐
       │ 10 Blast Radius    │
       │ / Impact Analyzer  │
       └──────────┬──────────┘
                  │
                  ▼
       ┌─────────────────────┐
       │ 11 Confidence      │
       │ Assessment         │
       └──────────┬──────────┘
                  │
                  ▼
       ┌─────────────────────┐
       │ 12 RCA Report      │
       │ Generator          │
       └──────────┬──────────┘
                  │
                  ▼
       ┌─────────────────────┐
       │ 13 Action Planner  │
       │ Verify / Mitigate  │
       └─────────────────────┘

핵심은 RAG를 처음부터 끝까지 한 번만 호출하지 않는 것입니다.

RAG는 최소한 다음 세 단계에서 사용하는 것을 권장합니다.

Initial RAG
     ↓
Hypothesis 생성
     ↓
Targeted RAG
     ↓
Evidence 검증
     ↓
Final RCA

2. 먼저 Workflow에서 사용하는 공통 State

회사 workflow tool이 어떤 형태든 각 node 사이에 전달되는 공통 JSON State를 하나 정의하는 것이 좋습니다.

예를 들어:

{
  "incident": {},
  "environment": {},
  "symptoms": [],
  "events": [],
  "logs": [],
  "metrics": [],
  "rag_results": [],
  "hypotheses": [],
  "evidence": [],
  "causal_chain": [],
  "impact": {},
  "confidence": {},
  "actions": [],
  "final_rca": {}
}

이것을 workflow의 공통 Context Object로 생각하면 됩니다.


3. Node 01 — Incident Trigger

역할

RCA의 시작점입니다.

Fake Incident JSON을 넣으면 됩니다.

예:

scenario:
slow_drive

cluster:
prod-kr01-k8s

namespace:
storage

target_node:
kr01-aistor-dn03

target_pod:
minio-aistor-pool-0-2

mount_drive:
/data/nvme02

service_tier:
P1-CRITICAL

Input

Fake Incident JSON

Output

{
  "incident_id": "...",
  "scenario": "slow_drive",
  "cluster": "prod-kr01-k8s",
  "namespace": "storage",
  "target_node": "kr01-aistor-dn03",
  "target_pod": "minio-aistor-pool-0-2",
  "severity": "P1-CRITICAL"
}

여기서는 LLM을 굳이 사용할 필요가 없습니다.

단순 JSON parsing node로 만드는 것이 좋습니다.


4. Node 02 — Incident Normalizer

이 노드는 상당히 중요합니다.

Fake Incident가 조금씩 다른 형태로 들어와도 RCA Agent 내부에서는 동일한 schema로 만들어야 합니다.

Prompt

You are an incident normalization agent.

Convert the incoming incident payload into a normalized incident context.

Extract:

- incident id
- timestamp
- cluster
- namespace
- node
- pod
- service
- component
- severity
- affected resource
- suspected infrastructure layer
- alert name
- scenario

Do not infer root cause.

Only extract facts explicitly present in the incident payload.

Output

예:

{
  "incident": {
    "id": "INC-20260928-001",
    "severity": "P1-CRITICAL",
    "timestamp": "...",
    "alert": "MinIODriveLatencyHigh"
  },

  "target": {
    "cluster": "prod-kr01-k8s",
    "namespace": "storage",
    "node": "kr01-aistor-dn03",
    "pod": "minio-aistor-pool-0-2",
    "drive": "/data/nvme02"
  },

  "service": {
    "name": "AIStor",
    "component": "storage"
  }
}

중요: 여기서는 원인을 판단하지 않습니다.


5. Node 03 — Evidence Extractor

Incident JSON 안에 들어있는 데이터를 유형별로 분리합니다.

                    Incident JSON
                         │
            ┌────────────┼────────────┐
            ▼            ▼            ▼
          Events        Logs        Metrics

Output:

{
  "events": [...],
  "logs": [...],
  "metrics": [...]
}

그리고 각각에 metadata를 붙입니다.

예:

{
  "type": "metric",
  "name": "node_disk_io_time_seconds_total",
  "source": "prometheus",
  "timestamp": "...",
  "entity": "kr01-aistor-dn03",
  "value": 0.87
}

이 단계도 가능한 한 deterministic하게 만드는 것이 좋습니다.


6. Node 04 — Symptom Analyzer

여기부터 LLM을 사용합니다.

질문은:

"무슨 일이 발생했는가?"

이지

"왜 발생했는가?"

가 아닙니다.

Prompt

Analyze the incident evidence.

Your task is to identify observable symptoms only.

Separate:

1. Application symptoms
2. Storage symptoms
3. Network symptoms
4. Kubernetes symptoms
5. Resource symptoms
6. Performance symptoms

Do not identify root cause.

Every symptom must reference supporting evidence.

예를 들어 slow_drive라면:

{
  "symptoms": [
    {
      "category": "storage",
      "symptom": "drive latency increased",
      "evidence": [
        "w_await increased",
        "wareq_sz increased"
      ]
    },
    {
      "category": "application",
      "symptom": "S3 PutObject requests timeout",
      "evidence": [
        "HTTP 503 Slow Down"
      ]
    },
    {
      "category": "application",
      "symptom": "goroutine count increased",
      "evidence": [
        "goroutine metric spike"
      ]
    }
  ]
}

이 단계에서 Root Cause를 말하면 안 됩니다.


7. Node 05 — Initial RAG Search

이제 RAG를 사용합니다.

그런데 검색 query를 단순히:

MinIO drive latency

라고 하면 안 됩니다.

앞 Node의 증상들을 기반으로 검색 query를 생성합니다.

Query Generator

예:

AIStor drive latency w_await wareq_sz
AIStor taking drive offline unable to write read 1m
AIStor disk latency PutObject timeout
MinIO Slow Down drive latency goroutine
DirectPV disk latency troubleshooting

그리고 RAG DB를 검색합니다.


8. RAG 결과를 그냥 LLM에 통째로 넣지 않는 것이 중요

검색 결과에는 다음 metadata가 있어야 합니다.

{
  "document": "aistor-troubleshooting.md",
  "section": "Drive Health",
  "content": "...",
  "source": "MinIO AIStor documentation",
  "relevance": 0.91
}

그리고 RAG 결과를 다음처럼 분류합니다.

RAG Evidence
│
├── Known Cause
├── Known Symptom
├── Troubleshooting Procedure
├── Configuration
├── Operational Guidance
└── Unrelated

이것을 RAG Evidence Classifier node로 만드는 것을 추천합니다.


9. Node 06 — Metric Correlation Analyzer

이 노드는 별도로 두는 것이 좋습니다.

RAG는 "일반적인 지식"이고 Prometheus는 "이번 장애에서 실제 발생한 사실"이기 때문입니다.

예:

Drive latency
       │
       ├── w_await ↑
       ├── wareq_sz ↑
       ├── disk I/O time ↑
       │
       ▼
MinIO request latency ↑
       │
       ▼
PutObject timeout
       │
       ▼
HTTP 503

이런 temporal relationship을 분석합니다.

Prompt

Analyze the temporal and causal relationships between metrics.

Determine:

- what changed first
- what changed next
- what changed last
- which metrics are correlated
- which metrics are merely coincidental

Do not declare root cause.

Output:

{
  "timeline": [
    {
      "order": 1,
      "event": "disk latency increased"
    },
    {
      "order": 2,
      "event": "request waiting increased"
    },
    {
      "order": 3,
      "event": "goroutines increased"
    },
    {
      "order": 4,
      "event": "PutObject timeout"
    }
  ]
}

이게 RCA에서 상당히 중요합니다.


10. Node 07 — Hypothesis Generator

이제 처음으로 "원인이 무엇일 가능성이 있는가?"를 생성합니다.

다만 하나의 원인만 만들지 않는 것이 중요합니다.

예:

H1: Physical disk degradation
H2: NVMe/controller issue
H3: DirectPV/storage layer issue
H4: Node CPU/IRQ contention
H5: Network latency
H6: MinIO software issue

Output:

{
  "hypotheses": [
    {
      "id": "H1",
      "hypothesis": "physical drive degradation",
      "layer": "hardware",
      "reason": [
        "drive latency increased",
        "I/O wait increased"
      ]
    },
    {
      "id": "H2",
      "hypothesis": "storage controller problem",
      "layer": "hardware"
    },
    {
      "id": "H3",
      "hypothesis": "network problem",
      "layer": "network"
    }
  ]
}

11. 중요한 설계 — Hypothesis는 계층별로 생성

제가 특히 추천하는 부분입니다.

RCA Agent가 다음 계층을 항상 검사하게 합니다.

Application
    ↓
AIStor
    ↓
Storage / DirectPV
    ↓
Filesystem
    ↓
Block Device
    ↓
Node Kernel
    ↓
Hardware

Network는 별도의 branch:

Client
  ↓
Pod
  ↓
Cilium
  ↓
Node NIC
  ↓
Switch
  ↓
Peer Node

Kubernetes도:

Pod
 ↓
Node
 ↓
Kubelet
 ↓
CNI
 ↓
Control Plane

즉 RCA Agent가 특정 기술에 편향되지 않게 합니다.


12. Node 08 — Evidence Validator

이 단계가 RCA Agent의 핵심입니다.

각 hypothesis에 대해:

Hypothesis
     │
     ├── Supporting Evidence
     ├── Contradicting Evidence
     └── Missing Evidence

를 찾습니다.

예:

H1 Physical Drive Problem

Supporting:

+ w_await increased
+ wareq_sz increased
+ drive-specific error
+ same drive repeatedly affected

Contradicting:

- other drives normal

Missing:

? SMART
? nvme error log
? kernel I/O error

이렇게 만듭니다.


13. Evidence Matrix를 만들어라

workflow에서 이 구조를 적극적으로 사용하는 것을 추천합니다.

HypothesisEvidence ForEvidence AgainstMissing Evidence
Physical diskw_await ↑다른 disk 정상SMART
Networktimeoutnetwork metrics 정상packet loss
CPU contentiongoroutine ↑CPU normalIRQ
AIStor bug503disk latency 선행version-specific issue
DirectPVdrive offline일부 증거 부족DirectPV logs

이 표가 만들어지면 RCA 품질이 크게 올라갑니다.


14. Node 09 — Targeted RAG Search

첫 번째 RAG 검색과 다른 점입니다.

첫 RAG:

"이런 증상이 어떤 문제와 관련 있는가?"

두 번째 RAG:

"현재 생성된 hypothesis를 검증할 문서가 있는가?"

예:

H1 = NVMe drive latency

검색:
"AIStor drive offline unable to write read 1m"
"DirectPV drive latency"
"NVMe timeout kernel"
"AIStor disk health"

H2:

"AIStor network timeout"
"Cilium packet loss"
"MinIO peer unreachable"

즉 Hypothesis마다 별도의 targeted retrieval을 수행합니다.


15. Node 10 — Causal Analyzer

여기에서 최종적인 causal chain을 만듭니다.

예를 들어 slow_drive:

Physical / Device
       │
       ▼
NVMe I/O latency
       │
       ▼
Filesystem I/O wait
       │
       ▼
AIStor disk operation latency
       │
       ▼
Request queue buildup
       │
       ▼
Goroutine increase
       │
       ▼
S3 PutObject latency
       │
       ▼
HTTP 503 / timeout
       │
       ▼
User-visible incident

중요한 것은 Root Cause와 Symptoms를 분리하는 것입니다.


16. Network Partition 시나리오는 다른 graph

예를 들어 fake incident의 network_partition이라면:

Cilium/BGP issue
       │
       ▼
Route/interface disruption
       │
       ▼
AIStor peer unreachable
       │
       ▼
Node communication failure
       │
       ▼
Distributed synchronization failure
       │
       ▼
Quorum degradation
       │
       ▼
S3 request failure

이렇게 scenario에 따라 causal graph가 달라집니다.


17. Node 11 — Blast Radius Analyzer

RCA Agent가 원인만 찾으면 실제 운영에서는 부족합니다.

반드시:

"얼마나 영향을 받았는가?"

를 분석해야 합니다.

분석 대상

Cluster
 ├── Node
 ├── Pod
 ├── Service
 ├── Namespace
 ├── Storage pool
 ├── Disk
 ├── Tenant
 └── User request

예:

{
  "scope": {
    "cluster": "prod-kr01-k8s",
    "nodes": 1,
    "pods": 1,
    "drives": 1
  },
  "service_impact": {
    "aistor": "degraded",
    "s3": "partial failure"
  }
}

18. Node 12 — Confidence Analyzer

RCA Agent에서 반드시 넣는 것을 추천합니다.

하지만 단순히:

Confidence: 95%

라고 LLM이 마음대로 숫자를 만들게 하면 안 됩니다.

다음 요소를 기반으로 계산하도록 합니다.

Evidence strength
+
Temporal consistency
+
RAG support
+
Contradicting evidence
+
Missing evidence

예:

{
  "root_cause": "physical drive degradation",
  "confidence": "HIGH",

  "basis": [
    "drive-specific latency",
    "I/O metrics correlate",
    "application symptoms follow disk latency",
    "other drives unaffected"
  ],

  "missing": [
    "SMART data",
    "kernel NVMe error log"
  ]
}

저라면 처음에는 숫자보다는:

HIGH
MEDIUM
LOW
UNKNOWN

으로 시작하겠습니다.


19. Node 13 — RCA Report Generator

최종 결과는 사람이 읽을 수 있어야 합니다.

저라면 다음 형태로 고정 schema를 만듭니다.

# Incident RCA

## 1. Incident Summary

## 2. Impact

## 3. Timeline

## 4. Observed Symptoms

## 5. Evidence

## 6. Candidate Root Causes

## 7. Root Cause

## 8. Causal Chain

## 9. Supporting Evidence

## 10. Contradicting Evidence

## 11. Missing Evidence

## 12. Blast Radius

## 13. Confidence

## 14. Recommended Verification

## 15. Immediate Mitigation

## 16. Long-term Prevention

20. 여기서 아주 중요한 안전장치

RCA Agent가 이런 말을 하면 안 됩니다.

"NVMe가 고장났다."

증거가 없다면:

"NVMe drive degradation이 가장 일관된 hypothesis이며, 현재 evidence는 drive-specific latency 증가를 지원한다. 그러나 SMART/NVMe error log가 없어 physical failure는 아직 확정할 수 없다."

라고 해야 합니다.

즉:

FACT
  ↓
OBSERVATION
  ↓
CORRELATION
  ↓
HYPOTHESIS
  ↓
VALIDATION
  ↓
ROOT CAUSE

를 workflow에서 강제하는 것이 중요합니다.


21. 제가 추천하는 실제 Node 구성

회사 Workflow Tool에서는 아래 정도로 만들면 좋습니다.

#NodeTypeLLM
01Incident TriggerInput❌
02Incident NormalizerTransform❌/LLM
03Evidence ExtractorTransform❌
04Symptom AnalyzerAgent✅
05Query GeneratorAgent✅
06RAG RetrieverRetrieval❌
07RAG Evidence ClassifierAgent✅
08Metric CorrelationAgent✅
09Hypothesis GeneratorAgent✅
10Evidence Matrix BuilderAgent✅
11Targeted RAGRetrieval❌
12Evidence ValidatorAgent✅
13Causal AnalyzerAgent✅
14Blast RadiusAgent✅
15Confidence AnalyzerAgent✅
16RCA Report GeneratorAgent✅
17Action PlannerAgent✅
18Final OutputOutput❌

22. Workflow에 Loop를 하나 넣는 것이 좋음

사실 가장 중요한 부분 중 하나입니다.

단순히:

RAG → RCA

로 끝내지 말고:

                 ┌──────────────┐
                 │ Hypothesis   │
                 │ Generator    │
                 └──────┬───────┘
                        ↓
                 ┌──────────────┐
                 │ Evidence     │
                 │ Validator    │
                 └──────┬───────┘
                        ↓
                 Evidence sufficient?
                    /          \
                  NO            YES
                  │              │
                  ▼              ▼
           Targeted RAG       Causal
                  │            Analysis
                  │              │
                  └──────►───────┘

이 loop를 넣으면 훨씬 "Agent"다운 구조가 됩니다.


23. 그리고 제가 강하게 추천하는 별도 Node: "Contradiction Detector"

이것을 하나 넣는 것이 좋습니다.

예를 들어:

Log:
network timeout

Metric:
packet loss = 0
BGP = stable

RAG:
network partition can cause timeout

LLM이 RAG만 보고:

Network partition

이라고 판단할 위험이 있습니다.

Contradiction Detector가:

Network hypothesis

Supporting:
+ timeout

Contradicting:
- packet loss absent
- BGP stable

Status:
WEAK

로 만들어야 합니다.


24. 최종적으로 RCA Agent의 사고 구조

전체적으로 보면 다음 구조가 됩니다.

                 INCIDENT
                    │
                    ▼
             ┌────────────┐
             │ Normalize  │
             └─────┬──────┘
                   │
                   ▼
             ┌────────────┐
             │ Evidence   │
             │ Extraction │
             └─────┬──────┘
                   │
                   ▼
             ┌────────────┐
             │  Symptoms  │
             └─────┬──────┘
                   │
          ┌────────┴────────┐
          ▼                 ▼
       Metrics             RAG
          │                 │
          └────────┬────────┘
                   ▼
             ┌────────────┐
             │Hypotheses  │
             └─────┬──────┘
                   │
                   ▼
             ┌────────────┐
             │ Evidence   │
             │ Matrix     │
             └─────┬──────┘
                   │
                   ▼
             ┌────────────┐
             │ Targeted   │
             │ RAG        │
             └─────┬──────┘
                   │
                   ▼
             ┌────────────┐
             │Validator   │
             └─────┬──────┘
                   │
              ┌────┴────┐
              │         │
             weak      strong
              │         │
              ▼         ▼
          More RAG    Causal Graph
              │         │
              └────┐ ┌──┘
                   ▼ ▼
              Blast Radius
                   │
                   ▼
              Confidence
                   │
                   ▼
                RCA
                   │
                   ▼
              Action Plan

25. 현재 가지고 있는 Fake Incident로 테스트한다면

첫 번째는 반드시:

slow_drive

를 넣으세요.

Agent가 다음을 만들어내는지 보면 됩니다.

Symptom
 ├─ disk latency
 ├─ request waiting
 ├─ goroutine increase
 └─ PutObject timeout

Hypothesis
 ├─ disk degradation
 ├─ controller issue
 ├─ DirectPV issue
 ├─ node resource contention
 └─ network issue

Evidence Matrix
 └─ disk degradation에 가장 많은 직접 증거

Causal Chain
 └─ disk latency
       ↓
   I/O wait
       ↓
   request queue
       ↓
   goroutine
       ↓
   S3 latency
       ↓
   503

Missing Evidence
 ├─ SMART
 ├─ nvme error log
 └─ kernel log

Conclusion
 └─ drive/storage-path problem
    (physical failure는 추가 검증 필요)

그 다음:

network_partition

을 넣어서 Agent가 disk hypothesis에 끌려가지 않고 network/BGP/Cilium/peer/quorum 계층으로 reasoning을 전환하는지 확인합니다.


26. 한 단계 더 발전시키면

지금 사용자 환경에서는 최종적으로 다음 형태로 만드는 것이 상당히 좋습니다.

                  ┌─────────────────────┐
                  │    RCA Agent        │
                  └──────────┬──────────┘
                             │
       ┌─────────────────────┼──────────────────────┐
       ▼                     ▼                      ▼
   Incident DB           RAG Knowledge          Live Data
       │                     │                      │
       │                ┌────┴────┐          ┌────┴─────┐
       │                │          │          │          │
       │             Docs       Past RCA   Prometheus OpenSearch
       │
       └─────────────────────┬───────────────────────┘
                             ▼
                       Evidence Store
                             │
                             ▼
                       RCA Reasoning
                             │
                  ┌──────────┼──────────┐
                  ▼          ▼          ▼
                Cause      Impact     Action
                  │          │          │
                  └──────────┼──────────┘
                             ▼
                       RCA Knowledge
                             │
                             ▼
                       RAG에 재저장

특히 Past RCA를 RAG에 다시 넣는 Feedback Loop가 중요합니다.

처음에는 Fake Incident로:

Incident → RCA

를 만들고,

운영하면서:

Real Incident
     ↓
RCA Agent
     ↓
Engineer Verification
     ↓
Corrected RCA
     ↓
RCA Knowledge
     ↓
RAG
     ↓
다음 Incident

로 발전시키면 됩니다.

이 구조라면 나중에 단순한 RAG 기반 RCA chatbot이 아니라, Incident Evidence → Hypothesis → Verification → Causal Reasoning → Organizational RCA Knowledge로 발전하는 실제 운영형 RCA Agent가 됩니다.

profile
engineer

0개의 댓글