26S26j

QK·5일 전

맞습니다. 이 workflow tool에서는 각 노드를 단순히 Prompt 하나로 생각하기보다,

Input → State → Processing/Prompt → Output → Next State

를 하나의 단위로 설계하는 것이 좋습니다.

특히 RCA Agent에서는 모든 노드가 자유로운 텍스트를 주고받게 하면 나중에 연결이 굉장히 어려워지므로, 공통 State Schema를 먼저 정의하고 각 Node가 그 State의 일부를 읽고 일부를 갱신하는 구조를 추천합니다.


1. 전체 구조부터 정의

제가 권장하는 형태는 다음입니다.

                    ┌────────────────────┐
                    │  Incident Input    │
                    └─────────┬──────────┘
                              │
                              ▼
                    ┌────────────────────┐
                    │  Incident State    │
                    └─────────┬──────────┘
                              │
                              ▼
                    ┌────────────────────┐
                    │ Normalize Node     │
                    └─────────┬──────────┘
                              │
                              ▼
                    ┌────────────────────┐
                    │ Evidence State     │
                    └─────────┬──────────┘
                              │
             ┌────────────────┼────────────────┐
             ▼                ▼                ▼
          Events             Logs            Metrics
             │                │                │
             └────────────────┼────────────────┘
                              ▼
                    ┌────────────────────┐
                    │ Symptom Analyzer   │
                    └─────────┬──────────┘
                              │
                              ▼
                    ┌────────────────────┐
                    │ RAG Search         │
                    └─────────┬──────────┘
                              │
                              ▼
                    ┌────────────────────┐
                    │ Hypothesis         │
                    │ Generator          │
                    └─────────┬──────────┘
                              │
                              ▼
                    ┌────────────────────┐
                    │ Evidence Validator │
                    └─────────┬──────────┘
                              │
                       ┌──────┴──────┐
                       │             │
                    부족함          충분함
                       │             │
                       ▼             ▼
                  Targeted RAG   Causal Analyzer
                       │             │
                       └──────┬──────┘
                              ▼
                    ┌────────────────────┐
                    │ Impact Analyzer    │
                    └─────────┬──────────┘
                              ▼
                    ┌────────────────────┐
                    │ Confidence         │
                    └─────────┬──────────┘
                              ▼
                    ┌────────────────────┐
                    │ RCA Report         │
                    └─────────┬──────────┘
                              ▼
                    ┌────────────────────┐
                    │ Action Planner     │
                    └────────────────────┘

2. 가장 먼저 만들어야 하는 것: Global State

회사 workflow tool에서 State 개념을 지원한다면 가장 먼저 이 schema를 만드는 것을 추천합니다.

예를 들어:

{
  "incident": {},
  "environment": {},
  "target": {},
  "symptoms": [],
  "timeline": [],
  "events": [],
  "logs": [],
  "metrics": [],
  "rag": {},
  "hypotheses": [],
  "evidence_matrix": [],
  "causal_chain": [],
  "impact": {},
  "confidence": {},
  "actions": [],
  "rca": {}
}

그런데 실제 운영을 생각하면 조금 더 세분화하는 것이 좋습니다.


3. 제가 권장하는 최종 State

3.1 Incident

{
  "incident": {
    "id": "",
    "timestamp": "",
    "severity": "",
    "alert_name": "",
    "scenario": "",
    "status": "",
    "source": ""
  }
}

의미

장애 자체의 기본 정보입니다.

예:

{
  "id": "INC-20260928-001",
  "severity": "P1",
  "alert_name": "MinIODriveLatencyHigh",
  "scenario": "slow_drive",
  "source": "synthetic"
}

4. Environment State

{
  "environment": {
    "cluster": "",
    "kubernetes_version": "",
    "cni": "",
    "storage": "",
    "region": "",
    "site": "",
    "environment": ""
  }
}

현재 사용자 환경이라면:

{
  "cluster": "prod-kr01-k8s",
  "kubernetes_version": "1.33.x",
  "cni": "Cilium",
  "storage": "AIStor",
  "environment": "production"
}

이 정보가 RAG query에도 상당히 중요합니다.


5. Target State

장애가 발생한 대상을 분리합니다.

{
  "target": {
    "namespace": "",
    "pod": "",
    "node": "",
    "service": "",
    "component": "",
    "drive": "",
    "ip": "",
    "resource": ""
  }
}

예:

{
  "namespace": "storage",
  "pod": "minio-aistor-pool-0-2",
  "node": "kr01-aistor-dn03",
  "service": "AIStor",
  "component": "storage",
  "drive": "/data/nvme02"
}

6. Evidence State

이 부분이 중요합니다.

모든 Evidence에 공통 metadata를 붙이는 것을 추천합니다.

{
  "evidence": {
    "events": [],
    "logs": [],
    "metrics": []
  }
}

각 evidence:

{
  "id": "EV-001",
  "type": "metric",
  "source": "prometheus",
  "timestamp": "2026-09-28T10:20:00",
  "entity": "kr01-aistor-dn03",
  "name": "node_disk_io_time",
  "value": 0.91,
  "unit": "",
  "raw": "",
  "reliability": "HIGH"
}

이렇게 해놓으면 나중에 모든 LLM node가:

"어떤 evidence를 근거로 했는가?"

를 추적할 수 있습니다.


7. Node마다 무엇을 정의해야 하는가?

각 Node를 다음 7개 항목으로 정의하면 됩니다.

1. Node Name
2. Input State
3. External Input
4. Processing / Prompt
5. Output Schema
6. State Update
7. Next Node

즉 workflow 문서를 다음 표준으로 만드는 것을 추천합니다.


8. Node 01 — Incident Input

목적

외부에서 장애 데이터를 받습니다.

Input

Incident JSON

State Read

없음

State Write

incident
environment
target
events
logs
metrics

Processing

JSON parsing.

Output

초기 State:

{
  "incident": {},
  "environment": {},
  "target": {},
  "events": [],
  "logs": [],
  "metrics": []
}

9. Node 02 — Incident Normalizer

Input

incident
environment
target

Prompt에 전달할 State

{
  "incident": {},
  "environment": {},
  "target": {}
}

Prompt

Normalize the incident information.

Extract only facts explicitly present in the input.

Do not infer root cause.

Output

{
  "incident": {
    "id": "...",
    "severity": "...",
    "alert_name": "...",
    "timestamp": "..."
  },
  "environment": {},
  "target": {}
}

State Update

incident
environment
target

10. Node 03 — Evidence Extractor

Input

raw incident

Output

events
logs
metrics

State

{
  "events": [],
  "logs": [],
  "metrics": []
}

여기서는 LLM보다 JSON transformation node가 좋습니다.


11. Node 04 — Symptom Analyzer

여기부터 Agent/LLM Node입니다.

Input State

incident
environment
target
events
logs
metrics

Prompt에 전달

Incident
+
Target
+
Events
+
Logs
+
Metrics

Prompt 핵심

Identify observable symptoms.

Do not determine root cause.

Every symptom must reference evidence IDs.

Output

{
  "symptoms": [
    {
      "id": "SYM-001",
      "category": "storage",
      "description": "Drive latency increased",
      "severity": "HIGH",
      "evidence_ids": [
        "EV-001",
        "EV-002"
      ]
    }
  ]
}

State Update

symptoms

12. Node 05 — Timeline Analyzer

이 노드를 별도로 두는 것도 추천합니다.

Input

events
logs
metrics
symptoms

Output

{
  "timeline": [
    {
      "timestamp": "...",
      "event": "drive latency increased"
    },
    {
      "timestamp": "...",
      "event": "request queue increased"
    },
    {
      "timestamp": "...",
      "event": "S3 timeout"
    }
  ]
}

State Update

timeline

RCA에서는 시간 순서가 매우 중요하기 때문에 별도 State로 유지하는 것이 좋습니다.


13. Node 06 — RAG Query Generator

Input

environment
target
symptoms
timeline

Output

{
  "rag": {
    "queries": [
      "AIStor drive latency",
      "AIStor drive offline",
      "DirectPV disk latency",
      "NVMe I/O latency"
    ]
  }
}

State Update

rag.queries

여기서는 아직 RAG 검색 결과를 넣지 않습니다.


14. Node 07 — RAG Retriever

이것은 Agent가 아니라 Retrieval Node입니다.

Input

rag.queries

External resource

RAG Vector DB
+
BM25
+
Metadata filter

Metadata Filter

현재 환경에서는 예를 들어:

{
  "product": "AIStor",
  "component": "storage",
  "version": "...",
  "document_type": [
    "documentation",
    "troubleshooting",
    "incident"
  ]
}

Output

{
  "rag": {
    "results": [
      {
        "document_id": "DOC-001",
        "title": "...",
        "content": "...",
        "score": 0.91,
        "source": "...",
        "section": "..."
      }
    ]
  }
}

15. Node 08 — RAG Evidence Classifier

RAG 결과를 그냥 context로 넣지 말고 분류합니다.

Input

symptoms
rag.results

Output

{
  "rag": {
    "evidence": [
      {
        "document_id": "DOC-001",
        "classification": "KNOWN_SYMPTOM",
        "relevance": "HIGH",
        "supports": [
          "SYM-001"
        ]
      }
    ]
  }
}

Classification:

KNOWN_SYMPTOM
KNOWN_CAUSE
TROUBLESHOOTING
CONFIGURATION
BACKGROUND
IRRELEVANT

16. Node 09 — Hypothesis Generator

Input

incident
environment
target
symptoms
timeline
metrics
rag.evidence

Output

{
  "hypotheses": [
    {
      "id": "H-001",
      "description": "Physical drive degradation",
      "layer": "hardware",
      "initial_confidence": "MEDIUM"
    },
    {
      "id": "H-002",
      "description": "Storage controller issue",
      "layer": "hardware",
      "initial_confidence": "LOW"
    },
    {
      "id": "H-003",
      "description": "Network latency",
      "layer": "network",
      "initial_confidence": "LOW"
    }
  ]
}

여기서 중요한 점: root_cause라고 하지 않고 hypothesis라고 합니다.


17. Node 10 — Evidence Matrix Builder

이 노드는 RCA Agent의 핵심 State를 만듭니다.

Input

hypotheses
symptoms
events
logs
metrics
rag.evidence
timeline

Output

{
  "evidence_matrix": [
    {
      "hypothesis_id": "H-001",

      "supporting": [
        "EV-001",
        "EV-004"
      ],

      "contradicting": [
        "EV-010"
      ],

      "missing": [
        "SMART data",
        "NVMe error log"
      ]
    }
  ]
}

이렇게 만들어 놓으면 이후 Agent가 훨씬 안정적으로 판단합니다.


18. Node 11 — Targeted RAG Query Generator

여기서는 Hypothesis별로 query를 생성합니다.

Input

hypotheses
evidence_matrix

Output

{
  "rag": {
    "targeted_queries": [
      {
        "hypothesis_id": "H-001",
        "queries": [
          "AIStor drive offline unable to write read",
          "NVMe latency drive failure"
        ]
      }
    ]
  }
}

19. Node 12 — Targeted RAG Retriever

Input

targeted_queries

Output

rag.targeted_results

기존 RAG와 별도 field를 두는 것을 추천합니다.

{
  "rag": {
    "initial_results": [],
    "targeted_results": []
  }
}

왜냐하면 나중에:

최초 검색 결과가 뭐였고, 추가 검색 결과가 뭐였는가?

를 추적해야 하기 때문입니다.


20. Node 13 — Evidence Validator

Input

hypotheses
evidence_matrix
rag.targeted_results
metrics
timeline

Output

{
  "hypotheses": [
    {
      "id": "H-001",
      "status": "SUPPORTED",
      "support_score": 0.86,
      "supporting_evidence": [],
      "contradicting_evidence": [],
      "missing_evidence": []
    }
  ]
}

여기서 중요한 것은 score를 LLM이 임의로 만드는 것이 아니라 workflow에서 정한 규칙을 사용하는 것입니다.

초기에는 차라리:

SUPPORTED
PARTIALLY_SUPPORTED
WEAK
CONTRADICTED
UNKNOWN

가 좋습니다.


21. Node 14 — Decision Router

이것은 LLM Node가 아니라 조건 Node입니다.

IF
    hypothesis status == SUPPORTED
AND
    critical missing evidence == 0

THEN
    → Causal Analyzer

ELSE
    → Additional Evidence Collection

이렇게 해야 Agent가 무한 loop에 빠지지 않습니다.

예:

max investigation loop = 3

22. Node 15 — Causal Analyzer

Input

validated hypotheses
symptoms
timeline
evidence

Output

{
  "causal_chain": [
    {
      "step": 1,
      "layer": "hardware",
      "cause": "drive I/O latency"
    },
    {
      "step": 2,
      "layer": "storage",
      "cause": "storage operation latency"
    },
    {
      "step": 3,
      "layer": "application",
      "cause": "request queue buildup"
    },
    {
      "step": 4,
      "layer": "service",
      "cause": "S3 timeout"
    }
  ]
}

23. Node 16 — Blast Radius Analyzer

Input

incident
target
symptoms
events
metrics
causal_chain

Output

{
  "impact": {
    "cluster": "partial",
    "nodes": 1,
    "pods": 1,
    "drives": 1,
    "service": "degraded",
    "user_impact": "partial S3 failures"
  }
}

24. Node 17 — Confidence Analyzer

Input

validated hypotheses
evidence_matrix
causal_chain
impact

Output

{
  "confidence": {
    "level": "HIGH",
    "reason": [
      "Multiple independent evidence sources",
      "Temporal sequence is consistent",
      "No strong contradictory evidence"
    ],
    "missing": [
      "SMART data"
    ]
  }
}

25. Node 18 — RCA Report Generator

Input

여기서는 앞에서 만든 모든 State를 넣습니다.

하지만 전체 raw log를 다시 넣으면 안 됩니다.

다음만 넣는 것을 추천합니다.

incident
target
symptoms
timeline
validated hypotheses
evidence_matrix
causal_chain
impact
confidence

Output

{
  "rca": {
    "summary": "",
    "root_cause": "",
    "causal_chain": [],
    "evidence": [],
    "impact": "",
    "confidence": "",
    "missing_evidence": []
  }
}

26. Node 19 — Action Planner

Input

rca
confidence
missing_evidence
impact

Output

{
  "actions": [
    {
      "priority": "P1",
      "type": "verification",
      "action": "Collect NVMe SMART/error log"
    },
    {
      "priority": "P1",
      "type": "mitigation",
      "action": "Remove affected drive from service"
    },
    {
      "priority": "P2",
      "type": "prevention",
      "action": "Add drive latency alert"
    }
  ]
}

여기서 verification, mitigation, prevention을 구분하면 좋습니다.


27. 최종 State는 이렇게 만들면 됩니다

제가 실제 구현한다면 최종적으로:

{
  "incident": {},

  "environment": {},

  "target": {},

  "events": [],

  "logs": [],

  "metrics": [],

  "symptoms": [],

  "timeline": [],

  "rag": {
    "queries": [],
    "initial_results": [],
    "evidence": [],
    "targeted_queries": [],
    "targeted_results": []
  },

  "hypotheses": [],

  "evidence_matrix": [],

  "causal_chain": [],

  "impact": {},

  "confidence": {},

  "rca": {},

  "actions": [],

  "workflow": {
    "iteration": 0,
    "status": "",
    "started_at": "",
    "completed_at": ""
  }
}

28. 그리고 State와 Node Input을 구분해야 합니다

이 부분이 실제 workflow 구현에서 중요합니다.

Global State 전체를 매번 모든 Node에 넣지 않는 것을 권장합니다.

예를 들어 Symptom Analyzer가:

전체 State 20KB

를 받는 것보다:

incident
target
events
logs
metrics

만 받게 합니다.

즉:

Global State
     │
     ├── Node 1 → 필요한 field만
     │
     ├── Node 2 → 필요한 field만
     │
     ├── Node 3 → 필요한 field만
     │
     └── Node 4 → 필요한 field만

이렇게 해야 token 비용도 줄고 hallucination도 줄어듭니다.


29. Node를 크게 4종류로 나누면 설계가 쉬워집니다

회사 workflow tool에서 node type을 다음처럼 정의하면 좋습니다.

A. Input / Collector

외부 데이터를 받음.

Incident
Prometheus
K8s
OpenSearch
RAG

B. Transform

데이터를 구조화.

Parser
Normalizer
Extractor
Aggregator
Timeline

C. Agent

판단/추론.

Symptom Analyzer
Hypothesis Generator
Evidence Validator
Causal Analyzer
RCA Generator

D. Router

조건 판단.

Evidence sufficient?
Hypothesis supported?
More investigation?
Severity?

이렇게 나누면 workflow가 상당히 명확해집니다.


30. 실제로는 "State"보다 더 중요한 것이 Evidence ID

RCA Agent를 운영형으로 만들 생각이라면 모든 데이터에 ID를 부여하세요.

예:

INC-001
SYM-001
EV-001
EV-002
HYP-001
HYP-002
RAG-001
RAG-002
CAUSE-001
ACTION-001

그러면 최종 RCA가:

Root Cause
  ↓
HYP-001
  ↓
EV-001
EV-004
RAG-007
RAG-012
  ↓
CAUSE-001

처럼 traceable해집니다.

이것이 단순 LLM chatbot과 RCA Agent를 구분하는 중요한 부분입니다.


31. 최종적으로 Workflow Tool에서 Node를 이렇게 정의하면 됩니다

각 Node의 표준 template을 하나 만들어두세요.

━━━━━━━━━━━━━━━━━━━━━━━━━━━━
NODE SPEC
━━━━━━━━━━━━━━━━━━━━━━━━━━━━

Node ID:
NODE-XX

Node Name:
Symptom Analyzer

Type:
LLM / Transform / Retrieval / Router

Purpose:
무엇을 하기 위한 Node인가?

INPUT STATE:
- incident
- environment
- target
- events
- logs
- metrics

EXTERNAL INPUT:
없음

PROMPT:
...

OUTPUT SCHEMA:
{
   ...
}

STATE UPDATE:
- symptoms

ERROR HANDLING:
- LLM JSON parsing failure
- empty evidence

NEXT:
NODE-05

FAILURE NEXT:
NODE-ERROR

MAX RETRY:
2
━━━━━━━━━━━━━━━━━━━━━━━━━━━━

32. 제가 추천하는 실제 전체 구성

최종적으로 회사 Workflow Tool에 아래 19개 Node를 만들어 두는 것이 좋습니다.

[INPUT]
01 Incident Input
        ↓
02 Incident Normalizer
        ↓
03 Evidence Extractor
        ↓
04 Symptom Analyzer
        ↓
05 Timeline Analyzer
        ↓
06 RAG Query Generator
        ↓
07 RAG Retriever
        ↓
08 RAG Evidence Classifier
        ↓
09 Hypothesis Generator
        ↓
10 Evidence Matrix Builder
        ↓
11 Targeted RAG Query Generator
        ↓
12 Targeted RAG Retriever
        ↓
13 Evidence Validator
        ↓
14 Decision Router
       ↙ ↘
    부족   충분
     ↓       ↓
   [Loop] 15 Causal Analyzer
             ↓
        16 Blast Radius
             ↓
        17 Confidence
             ↓
        18 RCA Generator
             ↓
        19 Action Planner
             ↓
           [END]

핵심은 Global State + Node별 Input/Output Schema + Evidence ID + Router 네 가지입니다.

그리고 지금 가지고 있는 slow_drive, network_partition fake incident 두 개를 이용하면 이 workflow를 만들면서 각 Node의 실제 Input JSON / Output JSON / Prompt / State Update까지 하나씩 확정할 수 있습니다. 이 방식으로 설계하면 나중에 Fake Incident를 실제 Prometheus/K8s/OpenSearch 데이터로 교체해도 workflow 자체는 거의 그대로 유지할 수 있습니다.

profile
engineer

0개의 댓글