맞습니다. 이 workflow tool에서는 각 노드를 단순히 Prompt 하나로 생각하기보다,
Input → State → Processing/Prompt → Output → Next State
를 하나의 단위로 설계하는 것이 좋습니다.
특히 RCA Agent에서는 모든 노드가 자유로운 텍스트를 주고받게 하면 나중에 연결이 굉장히 어려워지므로, 공통 State Schema를 먼저 정의하고 각 Node가 그 State의 일부를 읽고 일부를 갱신하는 구조를 추천합니다.
제가 권장하는 형태는 다음입니다.
┌────────────────────┐
│ Incident Input │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ Incident State │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ Normalize Node │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ Evidence State │
└─────────┬──────────┘
│
┌────────────────┼────────────────┐
▼ ▼ ▼
Events Logs Metrics
│ │ │
└────────────────┼────────────────┘
▼
┌────────────────────┐
│ Symptom Analyzer │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ RAG Search │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ Hypothesis │
│ Generator │
└─────────┬──────────┘
│
▼
┌────────────────────┐
│ Evidence Validator │
└─────────┬──────────┘
│
┌──────┴──────┐
│ │
부족함 충분함
│ │
▼ ▼
Targeted RAG Causal Analyzer
│ │
└──────┬──────┘
▼
┌────────────────────┐
│ Impact Analyzer │
└─────────┬──────────┘
▼
┌────────────────────┐
│ Confidence │
└─────────┬──────────┘
▼
┌────────────────────┐
│ RCA Report │
└─────────┬──────────┘
▼
┌────────────────────┐
│ Action Planner │
└────────────────────┘
회사 workflow tool에서 State 개념을 지원한다면 가장 먼저 이 schema를 만드는 것을 추천합니다.
예를 들어:
{
"incident": {},
"environment": {},
"target": {},
"symptoms": [],
"timeline": [],
"events": [],
"logs": [],
"metrics": [],
"rag": {},
"hypotheses": [],
"evidence_matrix": [],
"causal_chain": [],
"impact": {},
"confidence": {},
"actions": [],
"rca": {}
}
그런데 실제 운영을 생각하면 조금 더 세분화하는 것이 좋습니다.
{
"incident": {
"id": "",
"timestamp": "",
"severity": "",
"alert_name": "",
"scenario": "",
"status": "",
"source": ""
}
}
장애 자체의 기본 정보입니다.
예:
{
"id": "INC-20260928-001",
"severity": "P1",
"alert_name": "MinIODriveLatencyHigh",
"scenario": "slow_drive",
"source": "synthetic"
}
{
"environment": {
"cluster": "",
"kubernetes_version": "",
"cni": "",
"storage": "",
"region": "",
"site": "",
"environment": ""
}
}
현재 사용자 환경이라면:
{
"cluster": "prod-kr01-k8s",
"kubernetes_version": "1.33.x",
"cni": "Cilium",
"storage": "AIStor",
"environment": "production"
}
이 정보가 RAG query에도 상당히 중요합니다.
장애가 발생한 대상을 분리합니다.
{
"target": {
"namespace": "",
"pod": "",
"node": "",
"service": "",
"component": "",
"drive": "",
"ip": "",
"resource": ""
}
}
예:
{
"namespace": "storage",
"pod": "minio-aistor-pool-0-2",
"node": "kr01-aistor-dn03",
"service": "AIStor",
"component": "storage",
"drive": "/data/nvme02"
}
이 부분이 중요합니다.
모든 Evidence에 공통 metadata를 붙이는 것을 추천합니다.
{
"evidence": {
"events": [],
"logs": [],
"metrics": []
}
}
각 evidence:
{
"id": "EV-001",
"type": "metric",
"source": "prometheus",
"timestamp": "2026-09-28T10:20:00",
"entity": "kr01-aistor-dn03",
"name": "node_disk_io_time",
"value": 0.91,
"unit": "",
"raw": "",
"reliability": "HIGH"
}
이렇게 해놓으면 나중에 모든 LLM node가:
"어떤 evidence를 근거로 했는가?"
를 추적할 수 있습니다.
각 Node를 다음 7개 항목으로 정의하면 됩니다.
1. Node Name
2. Input State
3. External Input
4. Processing / Prompt
5. Output Schema
6. State Update
7. Next Node
즉 workflow 문서를 다음 표준으로 만드는 것을 추천합니다.
외부에서 장애 데이터를 받습니다.
Incident JSON
없음
incident
environment
target
events
logs
metrics
JSON parsing.
초기 State:
{
"incident": {},
"environment": {},
"target": {},
"events": [],
"logs": [],
"metrics": []
}
incident
environment
target
{
"incident": {},
"environment": {},
"target": {}
}
Normalize the incident information.
Extract only facts explicitly present in the input.
Do not infer root cause.
{
"incident": {
"id": "...",
"severity": "...",
"alert_name": "...",
"timestamp": "..."
},
"environment": {},
"target": {}
}
incident
environment
target
raw incident
events
logs
metrics
{
"events": [],
"logs": [],
"metrics": []
}
여기서는 LLM보다 JSON transformation node가 좋습니다.
여기부터 Agent/LLM Node입니다.
incident
environment
target
events
logs
metrics
Incident
+
Target
+
Events
+
Logs
+
Metrics
Identify observable symptoms.
Do not determine root cause.
Every symptom must reference evidence IDs.
{
"symptoms": [
{
"id": "SYM-001",
"category": "storage",
"description": "Drive latency increased",
"severity": "HIGH",
"evidence_ids": [
"EV-001",
"EV-002"
]
}
]
}
symptoms
이 노드를 별도로 두는 것도 추천합니다.
events
logs
metrics
symptoms
{
"timeline": [
{
"timestamp": "...",
"event": "drive latency increased"
},
{
"timestamp": "...",
"event": "request queue increased"
},
{
"timestamp": "...",
"event": "S3 timeout"
}
]
}
timeline
RCA에서는 시간 순서가 매우 중요하기 때문에 별도 State로 유지하는 것이 좋습니다.
environment
target
symptoms
timeline
{
"rag": {
"queries": [
"AIStor drive latency",
"AIStor drive offline",
"DirectPV disk latency",
"NVMe I/O latency"
]
}
}
rag.queries
여기서는 아직 RAG 검색 결과를 넣지 않습니다.
이것은 Agent가 아니라 Retrieval Node입니다.
rag.queries
RAG Vector DB
+
BM25
+
Metadata filter
현재 환경에서는 예를 들어:
{
"product": "AIStor",
"component": "storage",
"version": "...",
"document_type": [
"documentation",
"troubleshooting",
"incident"
]
}
{
"rag": {
"results": [
{
"document_id": "DOC-001",
"title": "...",
"content": "...",
"score": 0.91,
"source": "...",
"section": "..."
}
]
}
}
RAG 결과를 그냥 context로 넣지 말고 분류합니다.
symptoms
rag.results
{
"rag": {
"evidence": [
{
"document_id": "DOC-001",
"classification": "KNOWN_SYMPTOM",
"relevance": "HIGH",
"supports": [
"SYM-001"
]
}
]
}
}
Classification:
KNOWN_SYMPTOM
KNOWN_CAUSE
TROUBLESHOOTING
CONFIGURATION
BACKGROUND
IRRELEVANT
incident
environment
target
symptoms
timeline
metrics
rag.evidence
{
"hypotheses": [
{
"id": "H-001",
"description": "Physical drive degradation",
"layer": "hardware",
"initial_confidence": "MEDIUM"
},
{
"id": "H-002",
"description": "Storage controller issue",
"layer": "hardware",
"initial_confidence": "LOW"
},
{
"id": "H-003",
"description": "Network latency",
"layer": "network",
"initial_confidence": "LOW"
}
]
}
여기서 중요한 점: root_cause라고 하지 않고 hypothesis라고 합니다.
이 노드는 RCA Agent의 핵심 State를 만듭니다.
hypotheses
symptoms
events
logs
metrics
rag.evidence
timeline
{
"evidence_matrix": [
{
"hypothesis_id": "H-001",
"supporting": [
"EV-001",
"EV-004"
],
"contradicting": [
"EV-010"
],
"missing": [
"SMART data",
"NVMe error log"
]
}
]
}
이렇게 만들어 놓으면 이후 Agent가 훨씬 안정적으로 판단합니다.
여기서는 Hypothesis별로 query를 생성합니다.
hypotheses
evidence_matrix
{
"rag": {
"targeted_queries": [
{
"hypothesis_id": "H-001",
"queries": [
"AIStor drive offline unable to write read",
"NVMe latency drive failure"
]
}
]
}
}
targeted_queries
rag.targeted_results
기존 RAG와 별도 field를 두는 것을 추천합니다.
{
"rag": {
"initial_results": [],
"targeted_results": []
}
}
왜냐하면 나중에:
최초 검색 결과가 뭐였고, 추가 검색 결과가 뭐였는가?
를 추적해야 하기 때문입니다.
hypotheses
evidence_matrix
rag.targeted_results
metrics
timeline
{
"hypotheses": [
{
"id": "H-001",
"status": "SUPPORTED",
"support_score": 0.86,
"supporting_evidence": [],
"contradicting_evidence": [],
"missing_evidence": []
}
]
}
여기서 중요한 것은 score를 LLM이 임의로 만드는 것이 아니라 workflow에서 정한 규칙을 사용하는 것입니다.
초기에는 차라리:
SUPPORTED
PARTIALLY_SUPPORTED
WEAK
CONTRADICTED
UNKNOWN
가 좋습니다.
이것은 LLM Node가 아니라 조건 Node입니다.
IF
hypothesis status == SUPPORTED
AND
critical missing evidence == 0
THEN
→ Causal Analyzer
ELSE
→ Additional Evidence Collection
이렇게 해야 Agent가 무한 loop에 빠지지 않습니다.
예:
max investigation loop = 3
validated hypotheses
symptoms
timeline
evidence
{
"causal_chain": [
{
"step": 1,
"layer": "hardware",
"cause": "drive I/O latency"
},
{
"step": 2,
"layer": "storage",
"cause": "storage operation latency"
},
{
"step": 3,
"layer": "application",
"cause": "request queue buildup"
},
{
"step": 4,
"layer": "service",
"cause": "S3 timeout"
}
]
}
incident
target
symptoms
events
metrics
causal_chain
{
"impact": {
"cluster": "partial",
"nodes": 1,
"pods": 1,
"drives": 1,
"service": "degraded",
"user_impact": "partial S3 failures"
}
}
validated hypotheses
evidence_matrix
causal_chain
impact
{
"confidence": {
"level": "HIGH",
"reason": [
"Multiple independent evidence sources",
"Temporal sequence is consistent",
"No strong contradictory evidence"
],
"missing": [
"SMART data"
]
}
}
여기서는 앞에서 만든 모든 State를 넣습니다.
하지만 전체 raw log를 다시 넣으면 안 됩니다.
다음만 넣는 것을 추천합니다.
incident
target
symptoms
timeline
validated hypotheses
evidence_matrix
causal_chain
impact
confidence
{
"rca": {
"summary": "",
"root_cause": "",
"causal_chain": [],
"evidence": [],
"impact": "",
"confidence": "",
"missing_evidence": []
}
}
rca
confidence
missing_evidence
impact
{
"actions": [
{
"priority": "P1",
"type": "verification",
"action": "Collect NVMe SMART/error log"
},
{
"priority": "P1",
"type": "mitigation",
"action": "Remove affected drive from service"
},
{
"priority": "P2",
"type": "prevention",
"action": "Add drive latency alert"
}
]
}
여기서 verification, mitigation, prevention을 구분하면 좋습니다.
제가 실제 구현한다면 최종적으로:
{
"incident": {},
"environment": {},
"target": {},
"events": [],
"logs": [],
"metrics": [],
"symptoms": [],
"timeline": [],
"rag": {
"queries": [],
"initial_results": [],
"evidence": [],
"targeted_queries": [],
"targeted_results": []
},
"hypotheses": [],
"evidence_matrix": [],
"causal_chain": [],
"impact": {},
"confidence": {},
"rca": {},
"actions": [],
"workflow": {
"iteration": 0,
"status": "",
"started_at": "",
"completed_at": ""
}
}
이 부분이 실제 workflow 구현에서 중요합니다.
Global State 전체를 매번 모든 Node에 넣지 않는 것을 권장합니다.
예를 들어 Symptom Analyzer가:
전체 State 20KB
를 받는 것보다:
incident
target
events
logs
metrics
만 받게 합니다.
즉:
Global State
│
├── Node 1 → 필요한 field만
│
├── Node 2 → 필요한 field만
│
├── Node 3 → 필요한 field만
│
└── Node 4 → 필요한 field만
이렇게 해야 token 비용도 줄고 hallucination도 줄어듭니다.
회사 workflow tool에서 node type을 다음처럼 정의하면 좋습니다.
외부 데이터를 받음.
Incident
Prometheus
K8s
OpenSearch
RAG
데이터를 구조화.
Parser
Normalizer
Extractor
Aggregator
Timeline
판단/추론.
Symptom Analyzer
Hypothesis Generator
Evidence Validator
Causal Analyzer
RCA Generator
조건 판단.
Evidence sufficient?
Hypothesis supported?
More investigation?
Severity?
이렇게 나누면 workflow가 상당히 명확해집니다.
RCA Agent를 운영형으로 만들 생각이라면 모든 데이터에 ID를 부여하세요.
예:
INC-001
SYM-001
EV-001
EV-002
HYP-001
HYP-002
RAG-001
RAG-002
CAUSE-001
ACTION-001
그러면 최종 RCA가:
Root Cause
↓
HYP-001
↓
EV-001
EV-004
RAG-007
RAG-012
↓
CAUSE-001
처럼 traceable해집니다.
이것이 단순 LLM chatbot과 RCA Agent를 구분하는 중요한 부분입니다.
각 Node의 표준 template을 하나 만들어두세요.
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
NODE SPEC
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Node ID:
NODE-XX
Node Name:
Symptom Analyzer
Type:
LLM / Transform / Retrieval / Router
Purpose:
무엇을 하기 위한 Node인가?
INPUT STATE:
- incident
- environment
- target
- events
- logs
- metrics
EXTERNAL INPUT:
없음
PROMPT:
...
OUTPUT SCHEMA:
{
...
}
STATE UPDATE:
- symptoms
ERROR HANDLING:
- LLM JSON parsing failure
- empty evidence
NEXT:
NODE-05
FAILURE NEXT:
NODE-ERROR
MAX RETRY:
2
━━━━━━━━━━━━━━━━━━━━━━━━━━━━
최종적으로 회사 Workflow Tool에 아래 19개 Node를 만들어 두는 것이 좋습니다.
[INPUT]
01 Incident Input
↓
02 Incident Normalizer
↓
03 Evidence Extractor
↓
04 Symptom Analyzer
↓
05 Timeline Analyzer
↓
06 RAG Query Generator
↓
07 RAG Retriever
↓
08 RAG Evidence Classifier
↓
09 Hypothesis Generator
↓
10 Evidence Matrix Builder
↓
11 Targeted RAG Query Generator
↓
12 Targeted RAG Retriever
↓
13 Evidence Validator
↓
14 Decision Router
↙ ↘
부족 충분
↓ ↓
[Loop] 15 Causal Analyzer
↓
16 Blast Radius
↓
17 Confidence
↓
18 RCA Generator
↓
19 Action Planner
↓
[END]
핵심은 Global State + Node별 Input/Output Schema + Evidence ID + Router 네 가지입니다.
그리고 지금 가지고 있는 slow_drive, network_partition fake incident 두 개를 이용하면 이 workflow를 만들면서 각 Node의 실제 Input JSON / Output JSON / Prompt / State Update까지 하나씩 확정할 수 있습니다. 이 방식으로 설계하면 나중에 Fake Incident를 실제 Prometheus/K8s/OpenSearch 데이터로 교체해도 workflow 자체는 거의 그대로 유지할 수 있습니다.