blumeops/docs/reference
Erich Blume 8e6a803076 C2(deploy-infra-alerting): impl add probes and alert rules for services-check coverage
Extend Alloy blackbox probes:
- Add prometheus, loki, grafana, teslamate, immich, navidrome
- Now probing 11 services (was 5), covering most HTTP checks from
  services-check

Add alert rules:
- PostgresClusterUnhealthy: cnpg_collector_up < 1 for 3m (critical)
- PodNotReady: kube_pod_status_ready{condition="true"} == 0 for 5m

Add runbooks:
- runbook-postgres-unhealthy.md
- runbook-pod-not-ready.md

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
2026-03-22 12:11:12 -07:00
..
infrastructure Review power.md: add ringtail, mark reviewed 2026-03-18 07:37:31 -07:00
kubernetes Deploy Mealie recipe manager (#299) 2026-03-16 21:59:10 -07:00
operations C2(deploy-infra-alerting): impl add probes and alert rules for services-check coverage 2026-03-22 12:11:12 -07:00
services Review jellyfin and automounter services 2026-03-17 13:06:23 -07:00
storage Review restore-1password-backup doc: fix offsite TBD, clarify archive name, add BorgBase to backups 2026-03-15 10:13:07 -07:00
tools Document ai-sources in AI guide, change process, and mise-tasks ref 2026-03-15 18:43:39 -07:00