Erich Blume 549c57ab82 C2(deploy-infra-alerting): impl add first alert rule and runbook

- Add ServiceProbeFailure alert rule to Grafana alerting provisioning
  - Queries probe_success metric from Alloy blackbox exporter
  - Extracts service name from job label via label_replace
  - Fires after 2 minutes of failure, noDataState=Alerting
  - Annotations include summary with service name and runbook URL
- Add runbook at docs/how-to/alerts/runbook-service-probe-failure.md
  - Covers all 5 probed services (miniflux, kiwix, transmission, devpi, argocd)
  - Diagnostic steps, common causes, silencing instructions
- Add alerting section to observability.md reference doc

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

2026-03-22 10:57:23 -07:00

556 B

Raw Blame History

title

modified

Observability

Metrics, logs, traces, and dashboards for BlumeOps infrastructure.

Components

prometheus - Metrics storage and querying
loki - Log aggregation
tempo - Distributed tracing
alloy - Metrics, log, and trace collection
grafana - Dashboards and visualization

Alerting

deploy-infra-alerting - Alerting pipeline (Grafana Unified Alerting → ntfy)
runbook-service-probe-failure - Service health check failure runbook

556 B Raw Blame History

Observability

Components

Alerting

556 B

Raw Blame History