사내 자체 워크플로우 빌더 도구(GUI 기반 노드 배치 또는 YAML/JSON 기반 파이프라인 정의)에 그대로 이식할 수 있도록, 코드에 의존하지 않고 워크플로우의 순서, 노드별 역할/입출력 명세, 내부 프롬프트 및 판단 로직을 중심으로 상세히 설계했습니다.
[Fake/Real Incident JSON]
│
▼
[Step 1. Symptom Normalizer & Entity Extractor] (LLM)
│
▼
[Step 2. Knowledge Base Query Planner & Retriever] (Search/DB Tool)
│
▼
[Step 3. RCA Hypothesis Generator] (LLM)
│
▼
[Step 4. Evidence Cross-Checker & Verifier] (LLM)
│
├──────────────────────────┐
▼ (모든 가설 기각 & 루프 < 2) ▼ (원인 확정 OR 루프 종료)
[Step 4-A. Hypothesis Refiner] [Step 5. RCA Report & Runbook Generator] (LLM)
│ │
└──────────────────────────────────┤
▼
[최종 결과 리포트]
incident_payload: 인시던트 JSON 원문역할: 대규모 분산 스토리지/K8s 전문 SRE 엔지니어
지시사항:
- 주어진 Incident 데이터에서 명백한 이상 징후(임계치 초과 메트릭, ERROR/WARN 로그)를 감지하십시오.
- 문제가 발생한 특정 타깃 엔티티(노드명, 파드명, 드라이브 마운트 경로, 에러 코드)를 빠짐없이 추출하십시오.
- 과거 장애 보고서 및 SOP를 검색할 때 사용할 핵심 검색 키워드(영문/한글) 3개를 생성하십시오.
target_component: 컴포넌트 명 (예: minio-aistor, directpv)anomaly_entities: [{"type": "node", "name": "storage-worker-node-04"}, {"type": "drive", "path": "/mnt/directpv/nvme03"}, {"type": "error", "code": "503 Slow Down"}]symptom_summary: 발생 증상 1~2줄 요약 (예: "특정 NVMe I/O 지연으로 인한 클라이언트 503 반환")search_queries: ["MinIO drive slow down timeout", "DirectPV NVMe I/O delay latency", "503 Slow Down PutObject"]bulk_chunks_for_vectordb.json 지식 베이스에서 가장 관련성 높은 SOP, 과거 포스트모템, 트러블슈팅 가이드를 추출합니다.component == target_component (예: minio-aistor 또는 directpv)doc_category가 TROUBLESHOOTING 또는 SOP인 문서를 우선 순위에 배치search_queries와 임베딩 유사도가 높은 청크 4개 선별retrieved_knowledge: 검색된 청크 목록 (가상 질문, 요약, 조치 명령어, 과거 근본 원인 포함)symptom_summary & anomaly_entities (Step 1 결과)retrieved_knowledge (Step 2 결과)지시사항:
제공된 증상과 지식 베이스(SOP/과거 장애)를 토대로, 이 장애를 유발했을 가능성이 있는 잠재 원인 가설을 2~3개 도출하십시오.
[중요 - 검증 규칙(Verification Rule) 명시]:
각 가설마다 "이 가설이 진실이기 위해 Incident 데이터에서 반드시 확인되어야 하는 조건(필요충분조건)"을 함께 정의하십시오.
예시:
- 가설: "스토리지 드라이브 하드웨어 I/O Hang으로 인한 분산 락 지연"
- 검증 규칙: "특정 단일 드라이브의 I/O Duration 메트릭만 비정상 스파이크를 보여야 하며, 피어 노드에서는 타임아웃 로그가 발생해야 함"
hypotheses: [{"id": "H1", "title": "...", "rule": "..."}, {"id": "H2", "title": "...", "rule": "..."}]검증 규칙과 Step 1의 원시 로그/메트릭 데이터를 1:1로 대조하여 가설의 참/거짓을 판정합니다. (환각 억제 핵심 단계)hypotheses (Step 3 결과)raw_metrics_summary & raw_logs_summary (최초 Incident JSON)역할: 엄격하고 보수적인 장애 감사관
지시사항:
각 가설의검증 규칙을 실제 메트릭과 로그 데이터와 비교 검증하십시오.
데이터에 명시적인 근거가 없는 내용은 절대 참으로 인정하지 마십시오.
판정 기준:
CONFIRMED: 데이터(수치, 로그)가 규칙을 100% 뒷받침하고 반증이 없음.DISPROVED: 데이터와 모순되는 명백한 반증 지표가 있음 (예: 디스크 문제라 했는데 I/O latency가 정상).INCONCLUSIVE: 제공된 데이터만으로는 확인이 불가능함.각 가설별로
status,confidence (0.0~1.0),supporting_evidence(인용된 로그/메트릭),refuting_evidence를 명시하십시오.
verified_hypotheses: 가설 검증 결과 리스트verified_hypotheses 중 CONFIRMED 상태인 가설이 1개 이상 존재할 때 ➔ Step 5로 이동DISPROVED 또는 INCONCLUSIVE이고, 재시도 횟수 < 1회 일 때 ➔ Step 4-A (Hypothesis Refiner)로 이동verified_hypotheses (Step 4 결과)retrieved_knowledge (Step 2의 SOP 및 조치 커맨드)incident_metadata (장애 ID, 대상 리소스)검증된 근본 원인과 첨부된 SOP 문서를 기반으로, 아래 4가지 섹션을 갖춘 표준 RCA 리포트를 작성하십시오.
- 인시던트 요약: 장애 ID, 발생 시간, 영향도, 관련 타깃 리소스
- 근본 원인 분석(RCA):
- 확정된 원인 (단정적 서술)
- 결정적 근거 (메트릭 지표 수치, 결정적 에러 로그 라인 명시)
- 기각된 다른 가설과 기각 사유 (운영 신뢰도 확보용)
- 단계별 조치 런북 (Action Runbook):
- 1단계: 즉시 긴급 조치: SOP에 명시된 CLI 커맨드 제시 (예:
mc admin drive offline ...또는kubectl directpv ...)- 2단계: 완화 확인: 복구 여부를 판별할 확인 커맨드 및 메트릭
- 3단계: 재발 방지 권고사항: 하드웨어 교체, 타임아웃 값 조정 등
- 분석 확신도 및 데이터 한계: 분석 확신도(
HIGH/MEDIUM/LOW) 및 추가 수집 필요 지표
incident_payload, entities, retrieved_docs, hypotheses를 누적하여 각 노드가 필요한 필드만 참조할 수 있도록 구성합니다.component) + 쿼리 매칭으로 동작시켜 검색 안정성을 확보합니다.==
네. 이 경우에는 “Incident → Context 수집 → 증거 정규화 → 증상 분류 → RAG 검색 → 가설 생성 → 증거 검증 → 원인 추론 → 영향/범위 판단 → RCA 보고서 → 후속 조치” 형태로 설계하는 것이 좋습니다.
특히 지금 가지고 있는 데이터가
이라는 점을 고려하면, 단순히 LLM에게 "이 장애 원인이 뭐야?"라고 묻는 구조보다는 각 단계에서 증거를 구조화하고 다음 단계로 넘기는 workflow가 훨씬 좋습니다.
제가 권장하는 전체 구조는 다음입니다.
┌───────────────────────┐
│ 01 Incident Trigger │
│ Fake / Real Incident │
└───────────┬───────────┘
│
▼
┌────────────────────────────┐
│ 02 Incident Normalizer │
│ Incident Context 구조화 │
└────────────┬───────────────┘
│
▼
┌────────────────────────────┐
│ 03 Evidence Extractor │
│ Log/Event/Metric 분리 │
└────────────┬───────────────┘
│
▼
┌────────────────────────────┐
│ 04 Symptom Analyzer │
│ "무슨 현상이 발생했나?" │
└────────────┬───────────────┘
│
├───────────────┐
▼ ▼
┌────────────────┐ ┌────────────────┐
│ 05 RAG Search │ │ 06 Metric │
│ Knowledge │ │ Correlation │
└───────┬────────┘ └───────┬────────┘
│ │
└─────────┬─────────┘
▼
┌─────────────────────┐
│ 07 Hypothesis │
│ Generator │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ 08 Evidence │
│ Validator │
└──────────┬──────────┘
│
┌─────┴─────┐
│ │
Evidence OK Evidence 부족
│ │
│ └──────► RAG Search / Context
│ Expansion
▼
┌─────────────────────┐
│ 09 Causal Analyzer │
│ Root Cause Graph │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ 10 Blast Radius │
│ / Impact Analyzer │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ 11 Confidence │
│ Assessment │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ 12 RCA Report │
│ Generator │
└──────────┬──────────┘
│
▼
┌─────────────────────┐
│ 13 Action Planner │
│ Verify / Mitigate │
└─────────────────────┘
핵심은 RAG를 처음부터 끝까지 한 번만 호출하지 않는 것입니다.
RAG는 최소한 다음 세 단계에서 사용하는 것을 권장합니다.
Initial RAG
↓
Hypothesis 생성
↓
Targeted RAG
↓
Evidence 검증
↓
Final RCA
회사 workflow tool이 어떤 형태든 각 node 사이에 전달되는 공통 JSON State를 하나 정의하는 것이 좋습니다.
예를 들어:
{
"incident": {},
"environment": {},
"symptoms": [],
"events": [],
"logs": [],
"metrics": [],
"rag_results": [],
"hypotheses": [],
"evidence": [],
"causal_chain": [],
"impact": {},
"confidence": {},
"actions": [],
"final_rca": {}
}
이것을 workflow의 공통 Context Object로 생각하면 됩니다.
RCA의 시작점입니다.
Fake Incident JSON을 넣으면 됩니다.
예:
scenario:
slow_drive
cluster:
prod-kr01-k8s
namespace:
storage
target_node:
kr01-aistor-dn03
target_pod:
minio-aistor-pool-0-2
mount_drive:
/data/nvme02
service_tier:
P1-CRITICAL
Fake Incident JSON
{
"incident_id": "...",
"scenario": "slow_drive",
"cluster": "prod-kr01-k8s",
"namespace": "storage",
"target_node": "kr01-aistor-dn03",
"target_pod": "minio-aistor-pool-0-2",
"severity": "P1-CRITICAL"
}
여기서는 LLM을 굳이 사용할 필요가 없습니다.
단순 JSON parsing node로 만드는 것이 좋습니다.
이 노드는 상당히 중요합니다.
Fake Incident가 조금씩 다른 형태로 들어와도 RCA Agent 내부에서는 동일한 schema로 만들어야 합니다.
You are an incident normalization agent.
Convert the incoming incident payload into a normalized incident context.
Extract:
- incident id
- timestamp
- cluster
- namespace
- node
- pod
- service
- component
- severity
- affected resource
- suspected infrastructure layer
- alert name
- scenario
Do not infer root cause.
Only extract facts explicitly present in the incident payload.
예:
{
"incident": {
"id": "INC-20260928-001",
"severity": "P1-CRITICAL",
"timestamp": "...",
"alert": "MinIODriveLatencyHigh"
},
"target": {
"cluster": "prod-kr01-k8s",
"namespace": "storage",
"node": "kr01-aistor-dn03",
"pod": "minio-aistor-pool-0-2",
"drive": "/data/nvme02"
},
"service": {
"name": "AIStor",
"component": "storage"
}
}
중요: 여기서는 원인을 판단하지 않습니다.
Incident JSON 안에 들어있는 데이터를 유형별로 분리합니다.
Incident JSON
│
┌────────────┼────────────┐
▼ ▼ ▼
Events Logs Metrics
Output:
{
"events": [...],
"logs": [...],
"metrics": [...]
}
그리고 각각에 metadata를 붙입니다.
예:
{
"type": "metric",
"name": "node_disk_io_time_seconds_total",
"source": "prometheus",
"timestamp": "...",
"entity": "kr01-aistor-dn03",
"value": 0.87
}
이 단계도 가능한 한 deterministic하게 만드는 것이 좋습니다.
여기부터 LLM을 사용합니다.
질문은:
"무슨 일이 발생했는가?"
이지
"왜 발생했는가?"
가 아닙니다.
Analyze the incident evidence.
Your task is to identify observable symptoms only.
Separate:
1. Application symptoms
2. Storage symptoms
3. Network symptoms
4. Kubernetes symptoms
5. Resource symptoms
6. Performance symptoms
Do not identify root cause.
Every symptom must reference supporting evidence.
예를 들어 slow_drive라면:
{
"symptoms": [
{
"category": "storage",
"symptom": "drive latency increased",
"evidence": [
"w_await increased",
"wareq_sz increased"
]
},
{
"category": "application",
"symptom": "S3 PutObject requests timeout",
"evidence": [
"HTTP 503 Slow Down"
]
},
{
"category": "application",
"symptom": "goroutine count increased",
"evidence": [
"goroutine metric spike"
]
}
]
}
이 단계에서 Root Cause를 말하면 안 됩니다.
이제 RAG를 사용합니다.
그런데 검색 query를 단순히:
MinIO drive latency
라고 하면 안 됩니다.
앞 Node의 증상들을 기반으로 검색 query를 생성합니다.
예:
AIStor drive latency w_await wareq_sz
AIStor taking drive offline unable to write read 1m
AIStor disk latency PutObject timeout
MinIO Slow Down drive latency goroutine
DirectPV disk latency troubleshooting
그리고 RAG DB를 검색합니다.
검색 결과에는 다음 metadata가 있어야 합니다.
{
"document": "aistor-troubleshooting.md",
"section": "Drive Health",
"content": "...",
"source": "MinIO AIStor documentation",
"relevance": 0.91
}
그리고 RAG 결과를 다음처럼 분류합니다.
RAG Evidence
│
├── Known Cause
├── Known Symptom
├── Troubleshooting Procedure
├── Configuration
├── Operational Guidance
└── Unrelated
이것을 RAG Evidence Classifier node로 만드는 것을 추천합니다.
이 노드는 별도로 두는 것이 좋습니다.
RAG는 "일반적인 지식"이고 Prometheus는 "이번 장애에서 실제 발생한 사실"이기 때문입니다.
예:
Drive latency
│
├── w_await ↑
├── wareq_sz ↑
├── disk I/O time ↑
│
▼
MinIO request latency ↑
│
▼
PutObject timeout
│
▼
HTTP 503
이런 temporal relationship을 분석합니다.
Analyze the temporal and causal relationships between metrics.
Determine:
- what changed first
- what changed next
- what changed last
- which metrics are correlated
- which metrics are merely coincidental
Do not declare root cause.
Output:
{
"timeline": [
{
"order": 1,
"event": "disk latency increased"
},
{
"order": 2,
"event": "request waiting increased"
},
{
"order": 3,
"event": "goroutines increased"
},
{
"order": 4,
"event": "PutObject timeout"
}
]
}
이게 RCA에서 상당히 중요합니다.
이제 처음으로 "원인이 무엇일 가능성이 있는가?"를 생성합니다.
다만 하나의 원인만 만들지 않는 것이 중요합니다.
예:
H1: Physical disk degradation
H2: NVMe/controller issue
H3: DirectPV/storage layer issue
H4: Node CPU/IRQ contention
H5: Network latency
H6: MinIO software issue
Output:
{
"hypotheses": [
{
"id": "H1",
"hypothesis": "physical drive degradation",
"layer": "hardware",
"reason": [
"drive latency increased",
"I/O wait increased"
]
},
{
"id": "H2",
"hypothesis": "storage controller problem",
"layer": "hardware"
},
{
"id": "H3",
"hypothesis": "network problem",
"layer": "network"
}
]
}
제가 특히 추천하는 부분입니다.
RCA Agent가 다음 계층을 항상 검사하게 합니다.
Application
↓
AIStor
↓
Storage / DirectPV
↓
Filesystem
↓
Block Device
↓
Node Kernel
↓
Hardware
Network는 별도의 branch:
Client
↓
Pod
↓
Cilium
↓
Node NIC
↓
Switch
↓
Peer Node
Kubernetes도:
Pod
↓
Node
↓
Kubelet
↓
CNI
↓
Control Plane
즉 RCA Agent가 특정 기술에 편향되지 않게 합니다.
이 단계가 RCA Agent의 핵심입니다.
각 hypothesis에 대해:
Hypothesis
│
├── Supporting Evidence
├── Contradicting Evidence
└── Missing Evidence
를 찾습니다.
예:
Supporting:
+ w_await increased
+ wareq_sz increased
+ drive-specific error
+ same drive repeatedly affected
Contradicting:
- other drives normal
Missing:
? SMART
? nvme error log
? kernel I/O error
이렇게 만듭니다.
workflow에서 이 구조를 적극적으로 사용하는 것을 추천합니다.
| Hypothesis | Evidence For | Evidence Against | Missing Evidence |
|---|---|---|---|
| Physical disk | w_await ↑ | 다른 disk 정상 | SMART |
| Network | timeout | network metrics 정상 | packet loss |
| CPU contention | goroutine ↑ | CPU normal | IRQ |
| AIStor bug | 503 | disk latency 선행 | version-specific issue |
| DirectPV | drive offline | 일부 증거 부족 | DirectPV logs |
이 표가 만들어지면 RCA 품질이 크게 올라갑니다.
첫 번째 RAG 검색과 다른 점입니다.
첫 RAG:
"이런 증상이 어떤 문제와 관련 있는가?"
두 번째 RAG:
"현재 생성된 hypothesis를 검증할 문서가 있는가?"
예:
H1 = NVMe drive latency
검색:
"AIStor drive offline unable to write read 1m"
"DirectPV drive latency"
"NVMe timeout kernel"
"AIStor disk health"
H2:
"AIStor network timeout"
"Cilium packet loss"
"MinIO peer unreachable"
즉 Hypothesis마다 별도의 targeted retrieval을 수행합니다.
여기에서 최종적인 causal chain을 만듭니다.
예를 들어 slow_drive:
Physical / Device
│
▼
NVMe I/O latency
│
▼
Filesystem I/O wait
│
▼
AIStor disk operation latency
│
▼
Request queue buildup
│
▼
Goroutine increase
│
▼
S3 PutObject latency
│
▼
HTTP 503 / timeout
│
▼
User-visible incident
중요한 것은 Root Cause와 Symptoms를 분리하는 것입니다.
예를 들어 fake incident의 network_partition이라면:
Cilium/BGP issue
│
▼
Route/interface disruption
│
▼
AIStor peer unreachable
│
▼
Node communication failure
│
▼
Distributed synchronization failure
│
▼
Quorum degradation
│
▼
S3 request failure
이렇게 scenario에 따라 causal graph가 달라집니다.
RCA Agent가 원인만 찾으면 실제 운영에서는 부족합니다.
반드시:
"얼마나 영향을 받았는가?"
를 분석해야 합니다.
Cluster
├── Node
├── Pod
├── Service
├── Namespace
├── Storage pool
├── Disk
├── Tenant
└── User request
예:
{
"scope": {
"cluster": "prod-kr01-k8s",
"nodes": 1,
"pods": 1,
"drives": 1
},
"service_impact": {
"aistor": "degraded",
"s3": "partial failure"
}
}
RCA Agent에서 반드시 넣는 것을 추천합니다.
하지만 단순히:
Confidence: 95%
라고 LLM이 마음대로 숫자를 만들게 하면 안 됩니다.
다음 요소를 기반으로 계산하도록 합니다.
Evidence strength
+
Temporal consistency
+
RAG support
+
Contradicting evidence
+
Missing evidence
예:
{
"root_cause": "physical drive degradation",
"confidence": "HIGH",
"basis": [
"drive-specific latency",
"I/O metrics correlate",
"application symptoms follow disk latency",
"other drives unaffected"
],
"missing": [
"SMART data",
"kernel NVMe error log"
]
}
저라면 처음에는 숫자보다는:
HIGH
MEDIUM
LOW
UNKNOWN
으로 시작하겠습니다.
최종 결과는 사람이 읽을 수 있어야 합니다.
저라면 다음 형태로 고정 schema를 만듭니다.
# Incident RCA
## 1. Incident Summary
## 2. Impact
## 3. Timeline
## 4. Observed Symptoms
## 5. Evidence
## 6. Candidate Root Causes
## 7. Root Cause
## 8. Causal Chain
## 9. Supporting Evidence
## 10. Contradicting Evidence
## 11. Missing Evidence
## 12. Blast Radius
## 13. Confidence
## 14. Recommended Verification
## 15. Immediate Mitigation
## 16. Long-term Prevention
RCA Agent가 이런 말을 하면 안 됩니다.
"NVMe가 고장났다."
증거가 없다면:
"NVMe drive degradation이 가장 일관된 hypothesis이며, 현재 evidence는 drive-specific latency 증가를 지원한다. 그러나 SMART/NVMe error log가 없어 physical failure는 아직 확정할 수 없다."
라고 해야 합니다.
즉:
FACT
↓
OBSERVATION
↓
CORRELATION
↓
HYPOTHESIS
↓
VALIDATION
↓
ROOT CAUSE
를 workflow에서 강제하는 것이 중요합니다.
회사 Workflow Tool에서는 아래 정도로 만들면 좋습니다.
| # | Node | Type | LLM |
|---|---|---|---|
| 01 | Incident Trigger | Input | ❌ |
| 02 | Incident Normalizer | Transform | ❌/LLM |
| 03 | Evidence Extractor | Transform | ❌ |
| 04 | Symptom Analyzer | Agent | ✅ |
| 05 | Query Generator | Agent | ✅ |
| 06 | RAG Retriever | Retrieval | ❌ |
| 07 | RAG Evidence Classifier | Agent | ✅ |
| 08 | Metric Correlation | Agent | ✅ |
| 09 | Hypothesis Generator | Agent | ✅ |
| 10 | Evidence Matrix Builder | Agent | ✅ |
| 11 | Targeted RAG | Retrieval | ❌ |
| 12 | Evidence Validator | Agent | ✅ |
| 13 | Causal Analyzer | Agent | ✅ |
| 14 | Blast Radius | Agent | ✅ |
| 15 | Confidence Analyzer | Agent | ✅ |
| 16 | RCA Report Generator | Agent | ✅ |
| 17 | Action Planner | Agent | ✅ |
| 18 | Final Output | Output | ❌ |
사실 가장 중요한 부분 중 하나입니다.
단순히:
RAG → RCA
로 끝내지 말고:
┌──────────────┐
│ Hypothesis │
│ Generator │
└──────┬───────┘
↓
┌──────────────┐
│ Evidence │
│ Validator │
└──────┬───────┘
↓
Evidence sufficient?
/ \
NO YES
│ │
▼ ▼
Targeted RAG Causal
│ Analysis
│ │
└──────►───────┘
이 loop를 넣으면 훨씬 "Agent"다운 구조가 됩니다.
이것을 하나 넣는 것이 좋습니다.
예를 들어:
Log:
network timeout
Metric:
packet loss = 0
BGP = stable
RAG:
network partition can cause timeout
LLM이 RAG만 보고:
Network partition
이라고 판단할 위험이 있습니다.
Contradiction Detector가:
Network hypothesis
Supporting:
+ timeout
Contradicting:
- packet loss absent
- BGP stable
Status:
WEAK
로 만들어야 합니다.
전체적으로 보면 다음 구조가 됩니다.
INCIDENT
│
▼
┌────────────┐
│ Normalize │
└─────┬──────┘
│
▼
┌────────────┐
│ Evidence │
│ Extraction │
└─────┬──────┘
│
▼
┌────────────┐
│ Symptoms │
└─────┬──────┘
│
┌────────┴────────┐
▼ ▼
Metrics RAG
│ │
└────────┬────────┘
▼
┌────────────┐
│Hypotheses │
└─────┬──────┘
│
▼
┌────────────┐
│ Evidence │
│ Matrix │
└─────┬──────┘
│
▼
┌────────────┐
│ Targeted │
│ RAG │
└─────┬──────┘
│
▼
┌────────────┐
│Validator │
└─────┬──────┘
│
┌────┴────┐
│ │
weak strong
│ │
▼ ▼
More RAG Causal Graph
│ │
└────┐ ┌──┘
▼ ▼
Blast Radius
│
▼
Confidence
│
▼
RCA
│
▼
Action Plan
첫 번째는 반드시:
slow_drive
를 넣으세요.
Agent가 다음을 만들어내는지 보면 됩니다.
Symptom
├─ disk latency
├─ request waiting
├─ goroutine increase
└─ PutObject timeout
Hypothesis
├─ disk degradation
├─ controller issue
├─ DirectPV issue
├─ node resource contention
└─ network issue
Evidence Matrix
└─ disk degradation에 가장 많은 직접 증거
Causal Chain
└─ disk latency
↓
I/O wait
↓
request queue
↓
goroutine
↓
S3 latency
↓
503
Missing Evidence
├─ SMART
├─ nvme error log
└─ kernel log
Conclusion
└─ drive/storage-path problem
(physical failure는 추가 검증 필요)
그 다음:
network_partition
을 넣어서 Agent가 disk hypothesis에 끌려가지 않고 network/BGP/Cilium/peer/quorum 계층으로 reasoning을 전환하는지 확인합니다.
지금 사용자 환경에서는 최종적으로 다음 형태로 만드는 것이 상당히 좋습니다.
┌─────────────────────┐
│ RCA Agent │
└──────────┬──────────┘
│
┌─────────────────────┼──────────────────────┐
▼ ▼ ▼
Incident DB RAG Knowledge Live Data
│ │ │
│ ┌────┴────┐ ┌────┴─────┐
│ │ │ │ │
│ Docs Past RCA Prometheus OpenSearch
│
└─────────────────────┬───────────────────────┘
▼
Evidence Store
│
▼
RCA Reasoning
│
┌──────────┼──────────┐
▼ ▼ ▼
Cause Impact Action
│ │ │
└──────────┼──────────┘
▼
RCA Knowledge
│
▼
RAG에 재저장
특히 Past RCA를 RAG에 다시 넣는 Feedback Loop가 중요합니다.
처음에는 Fake Incident로:
Incident → RCA
를 만들고,
운영하면서:
Real Incident
↓
RCA Agent
↓
Engineer Verification
↓
Corrected RCA
↓
RCA Knowledge
↓
RAG
↓
다음 Incident
로 발전시키면 됩니다.
이 구조라면 나중에 단순한 RAG 기반 RCA chatbot이 아니라, Incident Evidence → Hypothesis → Verification → Causal Reasoning → Organizational RCA Knowledge로 발전하는 실제 운영형 RCA Agent가 됩니다.