Files
manage/docs/monitoring-logging-design.md
T
Developer d4f95b64d4 chore(observability): externalize stack from root compose files
Manage now connects to existing Grafana/Prometheus/Alertmanager instances
and never deploys its own stack.

- docker-compose.yml / docker-compose.dev.yml: removed prometheus, loki,
  alloy, grafana, alertmanager, node-exporter services, the monitoring
  network, and observability named volumes; they now ship only backend +
  frontend. Dev frontend now joins the web network so the Vite dev proxy
  can reach the backend.
- backend: alertmanager_url default is now empty; /api/monitoring/alerts
  and /alertmanager-status return graceful "not configured" responses
  when ALERTMANAGER_URL is unset. Added not-configured tests.
- docker-compose.observability.yml: kept as the optional standalone
  example; header clarifies Manage does not deploy it.
- Removed orphaned combined monitoring/prometheus/prometheus.yml
  (standalone stack uses prometheus.standalone.yml).
- Docs (README, REQUIREMENTS decision log, monitoring-logging-design,
  observability-runbooks, context.md, MIGRATION_PLAN, frontend/README,
  CHANGELOG) updated to the connect-to-existing model.

VITE_GRAFANA_URL / VITE_PROMETHEUS_URL remain as optional frontend
deep-link overrides. .env.example still needs a manual update (safety
policy blocks assistant edits): set ALERTMANAGER_URL empty/optional and
move standalone-only vars out of the root file.
2026-06-23 21:20:07 +00:00

461 lines
27 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Monitoring and Logging Design — Manage
> **Status (2026-06-23):** This document is the **reference architecture** for the
> self-hosted observability stack Manage integrates with. Manage itself **does
> not deploy** Grafana, Prometheus, Loki, Alertmanager, Alloy, or Node Exporter.
> The root `docker-compose.yml` / `docker-compose.dev.yml` ship **only** the
> backend and frontend; Manage connects to **existing** Grafana / Prometheus /
> Alertmanager instances. The stack described here is available as an optional,
> deploy-it-yourself example in `docker-compose.observability.yml` (see
> `docs/observability-runbooks.md`). The "Add services to `docker-compose.yml`"
> phase notes below describe the **original** in-compose deployment, which has
> since been extracted to the standalone compose file.
## Executive Summary
Manage currently uses ad-hoc observability: plain-text Python logs, a custom POSIX shell metrics collector on remote machines, and a background poller that stores snapshots in SQLite. This works for a single-instance homelab but becomes painful as the fleet grows and as users need faster incident response.
This document proposes a dedicated, self-hosted observability subsystem built on the standard Prometheus/Grafana stack:
- **Metrics**: Prometheus pulling from Node Exporter on machines and from application exporters in containers.
- **Logs**: Structured JSON logs shipped to **Grafana Loki** by **Promtail/Grafana Alloy**.
- **Dashboards**: Grafana for deep-dive dashboards, embedded in the Manage React UI via iframes.
- **Alerting**: Prometheus Alertmanager for routing and notifications (email first, webhooks later).
- **Auth**: Grafana authenticates through the existing OIDC/Authentik provider.
The existing POSIX remote collector will be removed, and the Python backup alert engine will be migrated to Alertmanager rules.
---
## Goals
1. **Fast query and alerting**: move from SQLite scan-based history to a real time-series database and indexed log store.
2. **Unified view**: monitor both local containers/apps and remote Linux machines from one place.
3. **Standard tooling**: use de-facto open-source tools so dashboards, exporters, and runbooks are reusable.
4. **Room to grow**: design supports adding traces, more notification channels, and longer retention later without re-architecture.
## Non-Goals
1. **Traces**: deferred to a later phase; the data flow and collector choice (Promtail/Alloy) will be trace-ready.
2. **Multi-tenant RBAC**: Manage is single-instance/homelab; Grafana teams are sufficient for now.
3. **SLA/SLO framework**: out of scope; we focus on metrics, logs, and alerts, not SLO budgeting.
4. **Cloud-hosted observability vendors**: all components run self-hosted in Docker Compose.
---
## Decisions
| Area | Decision | Rationale |
|------|----------|-----------|
| Coupling | Dedicated observability subsystem consumed by Manage | Keeps Manage fast and lets the observability stack evolve independently. |
| Metrics backend | Prometheus | Pull model, huge ecosystem, standard exporters, easy Grafana integration. |
| Machine metrics | Node Exporter | Rich OS metrics, reusable dashboards, no custom shell to maintain. |
| Log backend | Grafana Loki | Prometheus-style labels, low resource use, tight Grafana integration. |
| Log collection | Promtail / Grafana Alloy | Tails Docker logs and journald; no per-app network calls. |
| App logs | Structured JSON to stdout | Standard 12-factor pattern; collector handles routing. |
| Dashboards | Grafana + iframe embeds | Fast to implement, rich dashboards, Manage UI stays focused on summary. |
| Alerting | Prometheus Alertmanager | Mature routing, silencing, inhibition; single source of truth for infra alerts. |
| Auth | Grafana OAuth via Authentik | Reuses existing identity provider; consistent UX. |
| Retention | 30 days metrics, 30 days logs | Matches current retention policy; disk usage stays predictable. |
| Migration | Remove POSIX collector, migrate backup alerts | Eliminates duplicate alerting paths and custom remote code. |
---
## Current State
### Logging
- `backend/src/media_library_viewer_api/logging_utils.py` configures stdlib `logging` with a plain-text format.
- `main.py` has a `log_requests` middleware that emits method, path, client IP, status code, and elapsed time.
- Frontend uses standard `console.log` / browser dev tools; no server-side log aggregation.
### Metrics
> **Historical note (2026-06-17):** The legacy Manage-side `MonitoringPoller` that
> SSH-scraped `/proc` + `df` into a local SQLite table (`monitoring_machine_actions`)
> has been **decommissioned**. System metrics now live entirely in the external
> observability stack: `node_exporter` on each machine is scraped by **Prometheus**
> and visualised in **Grafana** (see the standalone `docker-compose.observability.yml`
> stack). Manage is a thin dashboard: it surfaces Alertmanager alerts + Prometheus
> target health + Grafana deep-links, and does not collect or store its own metrics.
- `main.py` has a `log_requests` middleware that emits method, path, client IP, status code, and elapsed time.
- Frontend uses standard `console.log` / browser dev tools; no server-side log aggregation.
### Alerting
- `backup_alert_engine.py` / `backup_poller.py` generate backup-related alerts (failure, anomaly, missed schedule) and store them in SQLite.
- No general infrastructure alerting (disk full, machine down, high CPU, etc.).
---
## Target Architecture
```
┌─────────────────────────────────────────────────────────────────────────────┐
│ Docker Compose Network │
│ │
│ ┌─────────────┐ scrape ┌──────────────┐ │
│ │ Prometheus │◄────────────────│ Node Exporter│◄── host / remote hosts │
│ │ (TSDB) │ └──────────────┘ │
│ └──────┬──────┘ │
│ │ query │
│ ▼ │
│ ┌─────────────┐ alert ┌─────────────┐ email ┌──────────┐ │
│ │ Grafana │──────────────►│ Alertmanager│──────────────►│ SMTP │ │
│ │ (OAuth) │ └─────────────┘ └──────────┘ │
│ └──────┬──────┘ │
│ │ embed (iframe) │
│ ▼ │
│ ┌─────────────────┐ │
│ │ Manage React │ │
│ │ (summary) │ │
│ └─────────────────┘ │
│ │
│ Logs: │
│ Manage / containers ──stdout──► Promtail/Alloy ──push──► Loki ◄──────┐ │
│ host / remote journald ───────► Promtail/Alloy ──push──► Loki │ │
│ │ │
│ Grafana queries Loki for logs ◄──────────────────────────────────────┘ │
└─────────────────────────────────────────────────────────────────────────────┘
```
### Component Responsibilities
| Component | Responsibility |
|-----------|----------------|
| **Prometheus** | Scrape and store metrics; evaluate alert rules. |
| **Node Exporter** | Expose host-level metrics (CPU, memory, disk, network, filesystem). |
| **Loki** | Store and index log streams by labels. |
| **Promtail / Alloy** | Discover log sources, parse labels, and push to Loki. |
| **Grafana** | Visualize metrics and logs; serve as the alert UI. |
| **Alertmanager** | Deduplicate, group, route, and deliver alerts. |
| **Manage backend** | Emit structured logs; expose `/metrics` for Prometheus; forward health/status to summary endpoints. |
| **Manage frontend** | Embed Grafana panels; show high-level status cards. |
---
## Instrumentation Changes
### Backend Logging
1. Switch to structured JSON logging via `python-json-logger` or `structlog`.
2. Include fields:
- `timestamp`, `level`, `logger`, `message`
- `request_id` (correlation ID propagated from frontend or generated)
- `method`, `path`, `status_code`, `elapsed_ms`
- `user_id`, `machine_id` where relevant
- `error` / `error_type` / `traceback` for exceptions
3. Keep emitting to stdout; Promtail/Alloy will parse JSON.
**Status**: implemented in `backend/src/media_library_viewer_api/logging_utils.py` with `LOG_FORMAT=json|text`, secret-safe `sanitize_log_extra`, and request logging in `main.py`.
### Request Middleware
- Extend `log_requests` to attach `request_id` to `request.state`.
- Include `request_id` in response headers (`X-Request-Id`) so the frontend can correlate.
- Log all outbound SSH commands with `machine_id`, `action`, `duration_ms`, and `request_id`.
**Status**: implemented in `main.py` and `observability.py`; `record_ssh_command` is called from `monitoring_actions.py` for every machine operation.
### Application Metrics Endpoint
- Add a `/metrics` endpoint using `prometheus-client`.
- Initial counters/gauges:
- `manage_api_requests_total` (method, path, status)
- `manage_api_request_duration_seconds` histogram
- `manage_ssh_commands_total` (machine_id, action, status)
- `manage_ssh_command_duration_seconds` histogram
- `manage_media_index_build_duration_seconds`
- `manage_backup_runs_total` (job_name, status)
- `manage_mail_queue_size`, `manage_mail_queue_failures_total`
**Status**: implemented in `backend/src/media_library_viewer_api/observability.py` and wired into `main.py`, `monitoring_actions.py`, `backups.py`, `media.py`, and `mail_queue.py`.
### Frontend Observability
- Keep first phase minimal: capture JS errors and send them to the backend as structured log events.
- Optional later: expose RUM-style metrics (page loads, API call latencies) via Prometheus client library or manual instrumentation.
---
## Node Exporter Deployment
### Local / Docker Host
- Add a `node-exporter` service to `docker-compose.yml` with host PID/network mounts.
- Prometheus scrapes it as `job="node"`.
### Remote Machines
- Add a managed task/template in `jobs.py` to install/upgrade Node Exporter via the package manager or a static binary.
- Manage exposes a settings flag per machine: `node_exporter_enabled`.
- For machines behind NAT, use one of:
1. Reverse SSH tunnel from machine to Manage host.
2. VPN/Wireguard already in place.
3. Prometheus federation or pushgateway for unreachable targets (later phase).
- If Node Exporter cannot be installed, temporarily keep the POSIX collector as a fallback until migration is complete.
---
## Log Shipping
### Docker Compose Services
- Add `logging` driver config or Promtail sidecar to each service.
- Preferred: run **Grafana Alloy** as a single daemon container with `docker_sd_config` to discover all Compose services automatically.
### Host Logs
- Alloy mounts `/var/log` and `/var/lib/docker/containers` (read-only).
- Alloy also tails journald where available.
### Remote Machines
- Option A: install Alloy on remote hosts and have it push logs to Loki.
- Option B: keep logs on remote hosts and use Node Exporter logs only; defer centralized remote logs.
- Recommendation: **Option A** for important machines, **Option B** for constrained ones.
---
## Dashboards
### Grafana
- Provision dashboards from YAML/JSON in version control:
- Node Exporter Full dashboard (import from Grafana.com).
- Manage API overview (request rate, latency, errors).
- Manage operations (SSH commands, media index builds, mail queue).
- Backup runs and alert history.
- Manage iframe embeds point to specific dashboard panels using Grafana's `panelId` and `kiosk` mode.
### Manage React UI
- Add an "Observability" page with:
- System health cards (Prometheus up, Loki up, Alertmanager up).
- Recent alerts summary from Alertmanager API.
- Iframe panels for key metrics (CPU/memory of selected machine, recent logs).
- Drill-down links open the full Grafana dashboard.
---
## Alerting
### Alertmanager Configuration
- Reuse existing SMTP settings for email notifications.
- Initial routing:
- `severity=critical` → email immediately.
- `severity=warning` → email with 5-minute group wait.
- `job=backup` → grouped by job name.
### Initial Alert Rules
- Infrastructure:
- Node down for > 5 minutes.
- Disk usage > 85% (warning), > 95% (critical).
- Memory usage > 90% for > 10 minutes.
- CPU iowait > 30% for > 10 minutes.
- Application:
- Manage API 5xx rate > 1% over 5 minutes.
- SSH command failure rate > 10% over 5 minutes.
- Mail queue growing or failures increasing.
- Backup:
- Backup job failed (`manage_backup_runs_total{status="failure"}`).
- Backup job missing for > 1.5× schedule interval.
- Backup run duration or size anomaly compared to rolling median.
### Backup Alert Migration
- Re-implement rules as Prometheus recording/alerting rules where possible.
- Keep historical comparison logic (median duration/size) as a small scheduled task that writes anomaly metrics to a Pushgateway or custom exporter, then Alertmanager consumes them.
- Preserve acknowledge/resolve workflow by storing Alertmanager webhook events in SQLite if needed, or by using Grafana alert annotations.
---
## Authentication
- Grafana configured with generic OAuth pointing at Authentik (same issuer as Manage).
- Grafana role mapping: default `Viewer`; admin group mapped to `Admin`.
- Traefik routes `grafana.${BACKEND_APP_HOST}` or a sub-path.
- Iframe embedding requires Grafana `allow_embedding = true` and matching cookie domain/samesite settings.
---
## Retention and Storage
| Store | Retention | Notes |
|-------|-----------|-------|
| Prometheus | 30 days | Default TSDB block compaction. |
| Loki | 30 days | Single-store boltdb-shipper or filesystem target. |
| Grafana | persistent SQLite/Postgres later | Dashboards and users are config, not runtime data. |
- Volumes: `prometheus-data`, `loki-data`, `grafana-data`.
- Backups: snapshot these volumes alongside existing `backend_cache`.
---
## Security
1. Network: all observability services on an internal Docker network; exposed only through Traefik where needed.
2. Node Exporter: bind to localhost on remote hosts and use a reverse tunnel, or firewall to Manage IP only.
3. Secrets: SMTP password, OIDC client secret, and any remote scrape credentials in environment variables or Docker secrets; never commit them.
4. Logs: sanitize tokens, passwords, and private keys before JSON serialization.
5. Alertmanager: disable unauthenticated UI if exposed publicly; rely on OIDC/Traefik.
---
## Implementation Plan
### Phase 0 — Foundation and Cleanup
- [x] Add `prometheus-client` and `python-json-logger` to `backend/pyproject.toml`.
- [x] Refactor `logging_utils.py` to emit JSON when `LOG_FORMAT=json`.
- [x] Add `request_id` propagation in `log_requests` middleware.
- [x] Add `/metrics` endpoint with initial counters/gauges.
- [x] Remove the POSIX remote collector code in `resources.py`; keep `disk_space` as a lightweight SSH/local helper in `monitoring_actions.py`.
- [x] Add `X-Request-Id` response header.
### Phase 1 — Local Observability Stack
- [x] Add services to `docker-compose.yml`: Prometheus, Loki, Grafana, Alertmanager, Grafana Alloy.
- [x] Add `node-exporter` service for the Docker host.
- [x] Configure Alloy to scrape all Docker container logs and ship to Loki.
- [x] Configure Prometheus to scrape `node-exporter` and Manage `/metrics`.
- [x] Provision Grafana datasources and a basic Manage API dashboard.
- [x] Wire Grafana OAuth to Authentik.
**Phase 1 files**:
- `monitoring/prometheus/prometheus.yml`
- `monitoring/prometheus/rules/backup_alerts.yml`
- `monitoring/loki/loki.yml`
- `monitoring/alloy/config.alloy`
- `monitoring/alertmanager/alertmanager.yml`
- `monitoring/grafana/grafana.ini`
- `monitoring/grafana/provisioning/datasources/datasources.yml`
- `monitoring/grafana/provisioning/dashboards-json/dashboards.yml`
- `monitoring/grafana/provisioning/dashboards-json/dashboards/manage-overview.json`
- `docker-compose.yml` and `docker-compose.dev.yml` updated with observability services.
- `frontend/vite.config.ts` updated with `/grafana` dev proxy.
- `.env.example` updated with Grafana OAuth and alerting variables.
### Phase 2 — Remote Machine Metrics
- [x] Add Node Exporter install/restart/status job templates in `jobs.py` (`install_node_exporter`, `restart_node_exporter`, `node_exporter_status`).
- [x] Add `node_exporter_enabled`, `node_exporter_port`, and `node_exporter_scrape_host` fields to `MonitoringMachineInput` and `SettingsStore`.
- [x] Implement `media_library_viewer_api.services.targets` to build Prometheus file-SD target lists for enabled SSH machines and write them to `PROMETHEUS_FILE_SD_DIR/node_exporter_targets.json`.
- [x] Regenerate file-SD targets on machine create/update/delete in `routers/settings.py`.
- [x] Add `/api/monitoring/prometheus-targets` endpoint returning live targets from the store.
- [x] Configure Prometheus `node-exporter-remote` job with `file_sd_configs` reading `/etc/prometheus/file-sd/node_exporter_targets.json`.
- [x] Mount the backend cache `prometheus-file-sd` directory into the Prometheus container as a read-only file-SD source.
- [x] Add `PROMETHEUS_FILE_SD_DIR` setting and `.env.example` entry.
- [x] Provision a minimal `Node Exporter Overview` Grafana dashboard (`monitoring/grafana/provisioning/dashboards-json/dashboards/node-exporter-overview.json`) covering CPU, memory, root disk, and network traffic.
- [x] Remove POSIX collector fallback. The legacy collector code in `backend/src/media_library_viewer_api/clients/resources.py` has been deleted, the collector control endpoints were removed from `routers/monitoring.py`, and `disk_space` was relocated to `services/monitoring_actions.py` as a lightweight SSH/local helper. Metrics are now sourced exclusively from Prometheus/Node Exporter.
**Phase 2 files**:
- `backend/src/media_library_viewer_api/jobs.py` (Node Exporter job templates).
- `backend/src/media_library_viewer_api/routers/settings.py` (machine input fields + target regeneration).
- `backend/src/media_library_viewer_api/services/settings_store.py` (machine persistence fields).
- `backend/src/media_library_viewer_api/services/targets.py` (file-SD target builder/writer).
- `backend/src/media_library_viewer_api/routers/monitoring.py` (`/prometheus-targets` endpoint).
- `backend/src/media_library_viewer_api/config.py` (`prometheus_file_sd_dir` setting).
- `backend/tests/test_targets.py` and `backend/tests/test_api.py` (target + endpoint tests).
- `monitoring/prometheus/prometheus.yml` (`node-exporter-remote` file SD job).
- `monitoring/grafana/provisioning/dashboards-json/dashboards/node-exporter-overview.json`.
- `docker-compose.yml` and `docker-compose.dev.yml` (file-SD volume mount + backend env var).
- `.env.example` (`PROMETHEUS_FILE_SD_DIR`).
### Phase 3 — Alerting
- [x] Define initial Prometheus alert rules for backup failures (infrastructure rules deferred to Phase 2/3).
- [x] Configure Alertmanager with email routing using existing SMTP settings; `monitoring/alertmanager/alertmanager.yml` uses env vars for SMTP and routing.
- [x] Migrate backup alert rules to Alertmanager:
- `BackupJobFailed` triggers on `increase(manage_backup_runs_total{status="failed"}[1h]) > 0`.
- `BackupJobStuck` triggers on `time() - manage_backup_runs_last_success_timestamp > 86400`.
- Added `manage_backup_runs_last_success_timestamp` gauge in `observability.py` and updated `routers/backups.py` to set it on successful runs.
- SQLite backup alerts (`backup_alert_engine.py` and `backup_poller.py`) are preserved for now alongside Alertmanager rules; the UI can consume either source during transition.
- [x] Add Alertmanager status summary endpoints in Manage backend:
- `GET /api/monitoring/alerts` proxies `/api/v1/alerts` and returns a UI-friendly summary (total, by_severity, alerts list).
- `GET /api/monitoring/alertmanager-status` proxies `/api/v2/status` and returns `up`, `version`, `uptime`, `peers`.
- [x] Added `alertmanager_url` setting to `config.py` (default `http://alertmanager:9093`) and `ALERTMANAGER_URL` env var in both compose files and `.env.example`.
- [x] Added tests for the Alertmanager endpoints and the backup success gauge.
**Phase 3 files**:
- `backend/src/media_library_viewer_api/routers/monitoring.py` (`/alerts` and `/alertmanager-status` endpoints).
- `backend/src/media_library_viewer_api/observability.py` (`BACKUP_RUNS_LAST_SUCCESS` gauge + updated `record_backup_run`).
- `backend/src/media_library_viewer_api/routers/backups.py` (pass `success=True` to `record_backup_run` on successful reports).
- `backend/src/media_library_viewer_api/config.py` (`alertmanager_url` setting).
- `monitoring/alertmanager/alertmanager.yml` (SMTP + routing config).
- `monitoring/prometheus/rules/backup_alerts.yml` (backup alert rules).
- `docker-compose.yml` / `docker-compose.dev.yml` (`ALERTMANAGER_URL` env var).
- `.env.example` (`ALERTMANAGER_URL`).
- `backend/tests/test_api.py` (`TestAlertmanager` tests).
- `backend/tests/test_observability.py` (backup metric tests).
### Phase 4 — Manage UI Integration
- [x] Add "Observability" page in React with summary cards (Alertmanager health, active alerts, Prometheus targets, machines) and Grafana iframe panels.
- [x] Add recent alerts list from Alertmanager API via `GET /api/monitoring/alerts`.
- [x] Add drill-down links to full Grafana dashboards for Node Exporter metrics and Loki logs.
- [x] Handle iframe sandbox attributes (`allow-scripts allow-same-origin allow-popups allow-forms`); CSP is delegated to the reverse proxy / Grafana `allow_embedding` configuration.
- [x] Add `useObservability` hook and API client wrappers for alerts, Alertmanager status, and Prometheus targets.
- [x] Add TypeScript types for Alertmanager summary/status and Prometheus targets.
- [x] Wire the new `/observability` route into `App.tsx` and the sidebar navigation.
**Phase 4 files**:
- `frontend/src/components/ObservabilityPage.tsx` (page component).
- `frontend/src/hooks/useObservability.ts` (React Query hooks).
- `frontend/src/api/client.ts` (API client functions).
- `frontend/src/types/index.ts` (new interfaces).
- `frontend/src/App.tsx` (route + nav item).
- `frontend/src/components/ui/{card,badge,alert,skeleton,select}.tsx` (shadcn/ui components).
- `frontend/vite.config.ts` already has `/grafana` dev proxy for iframe source.
### Phase 5 — Hardening and Future-Proofing
- [x] Add health checks and `deploy.resources` limits for Prometheus, Loki, Alloy, Grafana, Alertmanager, and Node Exporter in both compose files.
- [x] Use `depends_on` with `condition: service_healthy` for Alloy → Loki and Grafana → Prometheus/Loki.
- [x] Add Prometheus scrape jobs for Loki, Alertmanager, and Grafana so their `up` metrics are available for health alerts.
- [x] Add observability health alerting rules (`PrometheusTargetMissing`, `AlertmanagerDown`, `GrafanaDown`).
- [x] Add `ALERTMANAGER_WEBHOOK_URL` backend setting, `POST /api/monitoring/alertmanager-webhook` receiver, and Alertmanager `webhook` receiver config.
- [x] Document runbooks for common alerts.
- [x] Add volume backups for Prometheus/Loki/Grafana data. Backup/restore procedures for `prometheus_data`, `loki_data`, `grafana_data`, and `alertmanager_data` are documented in `docs/observability-runbooks.md`.
- [ ] Optional: add OpenTelemetry Collector as a translation layer for traces later.
**Phase 5 files**:
- `docker-compose.yml` and `docker-compose.dev.yml` (health checks, resource limits, `depends_on` conditions).
- `monitoring/prometheus/prometheus.yml` (additional scrape jobs for observability services).
- `monitoring/prometheus/rules/backup_alerts.yml` (renamed scope to include observability health alerts).
- `backend/src/media_library_viewer_api/config.py` (`alertmanager_webhook_url` setting).
- `backend/src/media_library_viewer_api/routers/monitoring.py` (`POST /api/monitoring/alertmanager-webhook`).
- `monitoring/alertmanager/alertmanager.yml` (`webhook` receiver).
- `backend/tests/test_api.py` (`TestAlertmanagerWebhook`).
- `docs/observability-runbooks.md` (new runbook documentation).
---
## Risks and Mitigations
| Risk | Mitigation |
|------|------------|
| Node Exporter hard to install on NAS/minimal hosts | Keep POSIX collector as opt-in fallback; document manual install steps. |
| Grafana iframe embedding blocked by CSP or cookies | Test early in Phase 4; use matching domains and `allow_embedding`. |
| Prometheus storage grows faster than expected | Start with 30-day retention; add compaction alerts. |
| Backup alert semantics lost in migration | Write tests comparing old Python alerts vs new Alertmanager rules. |
| OIDC configuration drift between Manage and Grafana | Use same env vars/Authentik application for both. |
| Remote scrape requires network path | Use reverse SSH tunnels or defer remote scraping until VPN is ready. |
---
## Open Questions
1. ~~What sub-domain or sub-path should Grafana use? (`grafana.${BACKEND_APP_HOST}` vs `${BACKEND_APP_HOST}/grafana`)~~ Decided: dedicated `GRAFANA_APP_HOST` subdomain in production; dev uses port 3000 and a `/grafana` proxy in Vite.
2. Should remote machines run Node Exporter as a systemd service or a container?
3. Do we need remote log aggregation immediately, or can it wait until after metrics alerting is stable?
4. Should the existing backup alert acknowledgement/resolve UI be rebuilt on top of Alertmanager, or replaced by Grafana alert annotations?