- Add ServiceProbeFailure alert rule to Grafana alerting provisioning - Queries probe_success metric from Alloy blackbox exporter - Extracts service name from job label via label_replace - Fires after 2 minutes of failure, noDataState=Alerting - Annotations include summary with service name and runbook URL - Add runbook at docs/how-to/alerts/runbook-service-probe-failure.md - Covers all 5 probed services (miniflux, kiwix, transmission, devpi, argocd) - Diagnostic steps, common causes, silencing instructions - Add alerting section to observability.md reference doc Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
556 B
556 B
| title | modified | tags | |
|---|---|---|---|
| Observability | 2026-03-22 |
|
Observability
Metrics, logs, traces, and dashboards for BlumeOps infrastructure.
Components
- prometheus - Metrics storage and querying
- loki - Log aggregation
- tempo - Distributed tracing
- alloy - Metrics, log, and trace collection
- grafana - Dashboards and visualization
Alerting
- deploy-infra-alerting - Alerting pipeline (Grafana Unified Alerting → ntfy)
- runbook-service-probe-failure - Service health check failure runbook