- Add ArgoCD metrics scrape target to Prometheus (argocd-metrics:8082) - Add ArgoCDAppOutOfSync alert: fires when argocd_app_info has sync_status != Synced for 30 minutes - Add runbook with diagnostic steps and common fixes Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
28 lines
871 B
Markdown
28 lines
871 B
Markdown
---
|
|
title: Observability
|
|
modified: 2026-03-22
|
|
tags:
|
|
- operations
|
|
---
|
|
|
|
# Observability
|
|
|
|
Metrics, logs, traces, and dashboards for BlumeOps infrastructure.
|
|
|
|
## Components
|
|
|
|
- [[prometheus]] - Metrics storage and querying
|
|
- [[loki]] - Log aggregation
|
|
- [[tempo]] - Distributed tracing
|
|
- [[alloy|Alloy]] - Metrics, log, and trace collection
|
|
- [[grafana]] - Dashboards and visualization
|
|
|
|
## Alerting
|
|
|
|
- [[deploy-infra-alerting]] - Alerting pipeline (Grafana Unified Alerting → ntfy)
|
|
- [[runbook-service-probe-failure]] - Service health check failure runbook
|
|
- [[runbook-postgres-unhealthy]] - PostgreSQL cluster health runbook
|
|
- [[runbook-pod-not-ready]] - Pod not ready runbook
|
|
- [[runbook-textfile-stale]] - Metrics textfile freshness runbook
|
|
- [[runbook-frigate-camera-down]] - Frigate camera health runbook
|
|
- [[runbook-argocd-out-of-sync]] - ArgoCD sync status runbook
|