C2(deploy-infra-alerting): impl add textfile staleness and Frigate alerts
- TextfileStale: fires when a .prom textfile on indri hasn't been updated in 1 hour (node_textfile_mtime_seconds). Covers borgmatic, zot, minikube, jellyfin exporters. - FrigateCameraDown: fires when frigate_camera_fps drops to 0 for 5m. - Add runbooks for both alerts. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
This commit is contained in:
parent
b2b0d6efa7
commit
2fa536e547
4 changed files with 214 additions and 0 deletions
|
|
@ -23,3 +23,5 @@ Metrics, logs, traces, and dashboards for BlumeOps infrastructure.
|
|||
- [[runbook-service-probe-failure]] - Service health check failure runbook
|
||||
- [[runbook-postgres-unhealthy]] - PostgreSQL cluster health runbook
|
||||
- [[runbook-pod-not-ready]] - Pod not ready runbook
|
||||
- [[runbook-textfile-stale]] - Metrics textfile freshness runbook
|
||||
- [[runbook-frigate-camera-down]] - Frigate camera health runbook
|
||||
|
|
|
|||
Loading…
Add table
Add a link
Reference in a new issue