- Add ServiceProbeFailure alert rule to Grafana alerting provisioning
- Queries probe_success metric from Alloy blackbox exporter
- Extracts service name from job label via label_replace
- Fires after 2 minutes of failure, noDataState=Alerting
- Annotations include summary with service name and runbook URL
- Add runbook at docs/how-to/alerts/runbook-service-probe-failure.md
- Covers all 5 probed services (miniflux, kiwix, transmission, devpi, argocd)
- Diagnostic steps, common causes, silencing instructions
- Add alerting section to observability.md reference doc
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>