blumeops/docs/reference/services/alloy.md
Erich Blume 813ce2ddaf Recurring review sweep: 4 doc cards + nvidia-device-plugin v0.19.2
Doc review (last-reviewed 2026-06-04):
- cluster.md: k8s v1.34.0→v1.35.0; ringtail workload list updated for
  the in-progress minikube→k3s migration
- ntfy/tempo/alloy: images are now locally-built registry.ops.eblu.me
  nix containers (v2.19.2 / v2.10.3 / v1.16.0); Fly alloy binary v1.16.1

Service review:
- nvidia-device-plugin v0.19.0→v0.19.2 (upstream patch, no breaking
  changes for our CDI + RuntimeClass manifests)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-04 13:23:03 -07:00

2.4 KiB

title modified last-reviewed tags
Alloy 2026-06-04 2026-06-04
service
observability

Grafana Alloy

Unified observability collector for metrics and logs with three deployments:

  1. Indri (host) - System metrics and service logs from macOS host
  2. Kubernetes (DaemonSet) - Automatic pod log collection and service health probes
  3. Fly.io proxy (embedded) - nginx access log metrics and log forwarding from flyio-proxy

Quick Reference

Property Value
Indri Binary ~/.local/bin/alloy
Indri Config ~/.config/grafana-alloy/config.alloy
K8s Namespace alloy
K8s Image registry.ops.eblu.me/blumeops/alloy:v1.16.0-9564435 (locally built)
ArgoCD App alloy-k8s
Fly.io Config fly/alloy.river
Fly.io Image grafana/alloy:v1.16.1 (binary copied into nginx container, sha-pinned)

Metrics Collected

From Indri

  • System metrics via prometheus.exporter.unix
  • Textfile collector: minikube.prom, borgmatic.prom, zot.prom, jellyfin.prom
  • Zot registry metrics from http://localhost:5050/metrics
  • Pushed to prometheus via remote_write

From Kubernetes

  • All pod logs via loki.source.kubernetes
  • Service health probes: miniflux, kiwix, transmission, devpi, argocd

From Fly.io Proxy

  • flyio_nginx_http_requests_total — request rate by status/method/host
  • flyio_nginx_http_request_duration_seconds — latency histogram
  • flyio_nginx_http_response_bytes_total — response bandwidth
  • flyio_nginx_cache_requests_total — cache HIT/MISS/EXPIRED counts
  • Pushed to prometheus via remote_write through caddy

Logs Collected

Brew services: forgejo, tailscale

mcquack LaunchAgents: alloy, borgmatic, zot, jellyfin

Logs pushed to loki at https://loki.tail8d86e.ts.net/loki/api/v1/push.

Fly.io proxy: nginx JSON access logs pushed to loki at https://loki.ops.eblu.me/loki/api/v1/push (via caddy).

Why Built from Source

The Homebrew bottle uses CGO_ENABLED=0, which breaks Tailscale MagicDNS. Building with CGO_ENABLED=1 uses the macOS native resolver.

Note: This may no longer be needed now that services use *.ops.eblu.me URLs (routed via Caddy) instead of *.tail8d86e.ts.net. Should be tested in the future.