blumeops/docs/reference
Erich Blume 549c57ab82 C2(deploy-infra-alerting): impl add first alert rule and runbook
- Add ServiceProbeFailure alert rule to Grafana alerting provisioning
  - Queries probe_success metric from Alloy blackbox exporter
  - Extracts service name from job label via label_replace
  - Fires after 2 minutes of failure, noDataState=Alerting
  - Annotations include summary with service name and runbook URL
- Add runbook at docs/how-to/alerts/runbook-service-probe-failure.md
  - Covers all 5 probed services (miniflux, kiwix, transmission, devpi, argocd)
  - Diagnostic steps, common causes, silencing instructions
- Add alerting section to observability.md reference doc

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 10:57:23 -07:00
..
infrastructure Review power.md: add ringtail, mark reviewed 2026-03-18 07:37:31 -07:00
kubernetes Deploy Mealie recipe manager (#299) 2026-03-16 21:59:10 -07:00
operations C2(deploy-infra-alerting): impl add first alert rule and runbook 2026-03-22 10:57:23 -07:00
services Review jellyfin and automounter services 2026-03-17 13:06:23 -07:00
storage Review restore-1password-backup doc: fix offsite TBD, clarify archive name, add BorgBase to backups 2026-03-15 10:13:07 -07:00
tools Document ai-sources in AI guide, change process, and mise-tasks ref 2026-03-15 18:43:39 -07:00