
https://github.com/prometheus/alertmanager.git
프로젝트에서 Discord Bot의 모니터링 시스템을 완성하기 위해 AlertManager를 도입했다. Prometheus로 메트릭을 수집하고, Grafana로 시각화한 다음, 마지막으로 AlertManager를 통해 실시간 알림 시스템을 구축했습니다.
Discord Bot이 24시간 안정적으로 돌아가는지 확인하려면 계속 모니터링해야 하는데, 사람이 24시간 대시보드를 지켜볼 수는 없다. 이 부분을 AlertManager를 통해 자동으로 알림이 오도록 설정했다.
전체적인 모니터링 시스템의 구조는 다음과 같다:
Flask 서버에서는 3개의 엔드포인트를 제공한다:
/metrics (8000 포트): Prometheus가 메트릭 수집/health: 헬스체크용/test-error: 테스트용 에러 발생Discord Bot에서 다음 메트릭들을 실시간으로 수집한다:
# 명령어 실행 통계 (성공/실패 구분)
discord_bot_commands_total{command="ping", status="success"} 45
discord_bot_commands_total{command="add", status="success"} 31
discord_bot_commands_total{command="roll", status="success"} 3
# 메시지 처리 시간
discord_bot_message_latency_seconds_bucket{le="0.1"} 42
discord_bot_message_latency_seconds_bucket{le="0.5"} 78
# 전송한 메시지 수
discord_bot_messages_sent_total 79
# 에러 발생 횟수 (타입별)
discord_bot_errors_total{error_type="command_error"} 2
discord_bot_errors_total{error_type="startup"} 0
# 봇 상태 확인 (하트비트)
discord_bot_heartbeat_timestamp_seconds 1717084523
# 서버 및 사용자 통계
discord_bot_active_guilds 1
discord_bot_active_users 5
모니터링 시스템을 테스트하기 위해 의도적으로 에러를 발생시키는 엔드포인트도 만들었습니다:
@app.route('/test-error')
def test_error():
"""테스트용 에러 발생"""
metrics.error_count.labels(error_type='test_error').inc()
logging.error("Test error triggered via /test-error endpoint")
return {"status": "error", "message": "Test error generated"}, 500
@app.route('/test-crash')
def test_crash():
"""테스트용 크래시 시뮬레이션"""
metrics.error_count.labels(error_type='crash_simulation').inc()
logging.critical("Crash simulation triggered")
raise Exception("Simulated crash for testing alerts")
처음에는 AlertManager 설정에 Slack webhook URL을 평문으로 넣었는데, 보안상 문제가 있어서 Kubernetes Secret으로 분리했습니다:
# slack-webhook-secret.yaml
apiVersion: v1
kind: Secret
metadata:
name: slack-webhook-secret
namespace: monitoring
type: Opaque
data:
webhook-url: aHR0cHM6Ly9ob29rcy5zbGFjay5jb20vc2VydmljZXMvVDA4UUJMTlVUQjQvQjA4VUZQUzVRVEQvVEdBSnBza3B6cXU2TUNPRVM4OGVKMklX
# alertmanager-configmap.yaml
global:
slack_api_url_file: '/etc/slack-secrets/webhook-url'
route:
receiver: 'default'
group_by: ['alertname', 'cluster', 'service']
group_wait: 10s
group_interval: 5m
repeat_interval: 12h
routes:
- match:
alertname: DiscordBotHighErrors
receiver: 'discord-bot-critical'
repeat_interval: 30m
- match:
alertname: DiscordBotDown
receiver: 'discord-bot-critical'
repeat_interval: 15m
receivers:
- name: 'discord-bot-critical'
slack_configs:
- channel: 'C08QCUMR0GL' # 실제 사용 중인 채널 ID
send_resolved: true
username: "Discord Bot Alert"
icon_emoji: ":robot_face:"
color: 'danger'
title: "🚨 Discord Bot Alert: {{ .GroupLabels.alertname }}"
text: |
{{ range .Alerts }}
*Summary:* {{ .Annotations.summary }}
*Description:* {{ .Annotations.description }}
*Started:* {{ .StartsAt.Format "2006-01-02 15:04:05" }}
{{ end }}
# prometheus-rules-configmap.yaml
groups:
- name: discord-bot-alerts
rules:
- alert: DiscordBotHighErrors
expr: discord_bot_errors_total > 5
for: 30s
labels:
severity: critical
service: discord-bot
annotations:
summary: "Discord Bot high error count"
description: "Error count is {{ $value }}"
- alert: DiscordBotDown
expr: up{job="discord-bot"} == 0
for: 1m
labels:
severity: critical
service: discord-bot
annotations:
summary: "Discord Bot is down"
description: "Discord Bot has been down for more than 1 minute"
# 실제 webhook URL을 Base64로 인코딩
echo -n "https://hooks.slack.com/services/YOUR/WEBHOOK/URL" | base64
# Secret 적용
kubectl apply -f k8s/monitoring/slack-bot/slack-webhook-secret.yaml
kubectl apply -f k8s/monitoring/alertmanager/alertmanager-configmap.yaml
kubectl apply -f k8s/monitoring/alertmanager/alertmanager-deployment.yaml
kubectl apply -f k8s/monitoring/alertmanager/alertmanager-service.yaml
kubectl apply -f k8s/monitoring/prometheus/prometheus-rules-configmap.yaml
kubectl rollout restart deployment/prometheus -n monitoring
# Discord Bot Pod 강제 삭제
kubectl delete pod -l app=discord-bot
# 1분 후 Slack에 도착한 알림:
실제 받은 Slack 알림:
Discord Bot Alert: DiscordBotDown Summary: Discord Bot is down Description: Discord Bot has been down for more than 1 minute Started: 2025-05-30 15:42:33
# 테스트 에러 엔드포인트 호출
curl http://localhost:8000/test-error
# 6번 호출하여 임계값(5) 초과
for i in {1..6}; do curl http://localhost:8000/test-error; done
# 30초 후 알림 도착!
실제 받은 Slack 알림

현재까지 실제로 받은 알림들:
ping: 45번 (가장 많이 사용) add: 31번 roll: 3번
ashcircle03/discord-bot:116 이미지로 안정적 실행C08QCUMR0GL로 알림 전송실제 운영 중인 메트릭들:
# 현재 수집되는 실제 메트릭 예시
discord_bot_commands_total{command="ping",status="success"} 45
discord_bot_messages_sent_total 79
discord_bot_errors_total{error_type="command_error"} 2
discord_bot_heartbeat_timestamp_seconds 1717084523
discord_bot_active_guilds 1
discord_bot_active_users 5
처음에는 단순히 "Discord Bot 만들어보자"였는데, 어느새 운영 환경 수준의 모니터링 시스템을 구축하게 되었다.