blumeops/docs/how-to
Erich Blume 549c57ab82 C2(deploy-infra-alerting): impl add first alert rule and runbook
- Add ServiceProbeFailure alert rule to Grafana alerting provisioning
  - Queries probe_success metric from Alloy blackbox exporter
  - Extracts service name from job label via label_replace
  - Fires after 2 minutes of failure, noDataState=Alerting
  - Annotations include summary with service name and runbook URL
- Add runbook at docs/how-to/alerts/runbook-service-probe-failure.md
  - Covers all 5 probed services (miniflux, kiwix, transmission, devpi, argocd)
  - Diagnostic steps, common causes, silencing instructions
- Add alerting section to observability.md reference doc

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 10:57:23 -07:00
..
alerts C2(deploy-infra-alerting): impl add first alert rule and runbook 2026-03-22 10:57:23 -07:00
authentik Restructure docs: consolidate, recategorize, and extract 2026-03-15 19:55:59 -07:00
configuration Restructure docs: consolidate, recategorize, and extract 2026-03-15 19:55:59 -07:00
dagger Add how-to guide for upgrading Dagger 2026-03-06 20:31:30 -08:00
deployment Restructure docs: consolidate, recategorize, and extract 2026-03-15 19:55:59 -07:00
forgejo-runner Remove mikado frontmatter from closed chains, clarify finalization rules 2026-03-04 20:43:19 -08:00
grafana Restructure docs: consolidate, recategorize, and extract 2026-03-15 19:55:59 -07:00
jobsync Review deploy-jobsync doc: add missing env var, update tag example 2026-03-13 15:45:07 -07:00
knowledgebase Review build-jobsync-container, refine docs-preview tooling 2026-03-11 18:11:34 -07:00
mealie Fix plan-a-meal random recipe API queries 2026-03-17 11:10:48 -07:00
operations Review operations docs: add last-reviewed dates and improve troubleshooting 2026-03-16 07:38:02 -07:00
ringtail
zot Remove mikado frontmatter from closed chains, clarify finalization rules 2026-03-04 20:43:19 -08:00