Files
manage/openspec/changes/archive/observability-service-registry/proposal.md
T
Developer 38b2de54ff chore: archive observability-service-registry, track pi-map artifacts
- Archive the completed observability-service-registry SDD change into
  openspec/changes/archive/ (delivered across 5 slices; only
  jellyfin-service-registry remains active).
- Stop ignoring .pi-map.md / .pi-map.index.md so the navigation maps are
  versioned alongside the code, and add the regenerated map pairs repo-wide.
2026-06-24 13:28:23 +00:00

119 lines
6.5 KiB
Markdown

# Proposal — Observability service registry
**Change:** `observability-service-registry`
**Phase:** proposal
**Date:** 2026-06-23
## Problem
Grafana and Prometheus are already modeled as **service-registry service types**
with dashboard widgets (`GrafanaLinkWidget`, `PrometheusMetricWidget`) whose
adapters resolve URLs from service records — not env vars. The Observability
page, however, still treats observability as an env-configured afterthought:
1. **Alertmanager is not a service type.** It has no `integrations/alertmanager.py`
definition. Its connection URL comes from the `ALERTMANAGER_URL` **environment
variable** (`config.py`), read directly by the monitoring router. This is the
only observability component still env-bound, and it cannot be configured,
named, multi-instanced, or surfaced as a widget like Grafana/Prometheus.
2. **The Observability page hard-codes `VITE_GRAFANA_URL`**
(`ObservabilityPage.tsx``GRAFANA_BASE_URL = import.meta.env.VITE_GRAFANA_URL`).
It ignores any Grafana service instances configured in the registry, so the
deep-links break unless the env var is also set.
3. **No Grafana/Prometheus health card.** The page shows Alertmanager health and
Prometheus *targets*, but never reports whether the configured Grafana or
Prometheus instance is actually reachable.
4. **Two env-bound relay knobs** with no service-registry equivalent:
`ALERTMANAGER_WEBHOOK_URL` (the router forwards received webhooks onward) and
the leftover `VITE_PROMETHEUS_URL` build arg.
The result is a split-brain model: Grafana/Prometheus are first-class registry
services, but Alertmanager is an env var and the page reads a different env var
for Grafana than the registry holds.
## Proposal
Make **all** observability integration (Alertmanager, Grafana, Prometheus) a
first-class, UI-configured part of the existing service registry, with a
dedicated Observability page and dashboard widgets — and **zero observability
service env vars**.
1. **Add `alertmanager` as a 6th service integration type**
(`integrations/alertmanager.py`): config (`base_url`, `timeout_seconds`),
optional secret (`api_key`/bearer token for auth'd instances), and an
`active_alerts` widget kind. Register it in `integrations/registry.py`,
following the exact pattern of `grafana.py`/`prometheus.py`.
2. **Rewire the monitoring router to resolve Alertmanager from a service record**
instead of `config.alertmanager_url`. `/alerts` and `/alertmanager-status`
accept an optional `?service_id=` and default to the first enabled
`alertmanager` instance. Graceful "not configured" responses are preserved
when none exists.
3. **Add health/status endpoints** for Grafana (`/api/monitoring/grafana-status`)
and Prometheus (`/api/monitoring/prometheus-status`) that resolve a service
record and probe its health endpoint, returning `{ up, version, service_id,
name, error? }`.
4. **Observability page discovers services** from the registry. Grafana deep-links
come from the discovered instance's `base_url`; `VITE_GRAFANA_URL` is removed.
Health cards for Alertmanager, Grafana, and Prometheus are shown.
5. **Add an `active_alerts` Alertmanager dashboard widget** (backend adapter +
frontend registry + component), mirroring the Grafana/Prometheus widgets.
6. **Delete observability service env vars** entirely: backend
`alertmanager_url`, `alertmanager_webhook_url`; frontend `VITE_GRAFANA_URL`,
`VITE_PROMETHEUS_URL`. Drop the inbound→outbound webhook relay (the receiver
stays, log-only).
## Non-goals
- Do **not** build an in-app metric explorer or charting library. Manage stays a
thin dashboard: deep-links/iframe to Grafana only.
- Do **not** introduce a "primary/default instance" flag on service records in
this change. Default selection is "first enabled instance of type" (see Open
questions). A default flag is a cross-cutting registry feature deferred to a
later change.
- Do **not** remove the Prometheus target **data** path: `build_node_exporter_targets`
and the `/prometheus-targets` HTTP endpoint stay. Only the **file writer**
(`write_prometheus_targets` + `prometheus_file_sd_dir`) is removed; external
Prometheus uses `http_sd_configs` against the endpoint instead.
- Do **not** persist Prometheus query history or alert acknowledgements in Manage.
- Do **not** migrate the existing `/machines` + Node Exporter target model; it is
orthogonal to service instances.
## Acceptance
- `alertmanager` appears in `GET /api/services/types` and can be
created/edited/deleted like the other service types.
- `ALERTMANAGER_URL`, `ALERTMANAGER_WEBHOOK_URL`, `VITE_GRAFANA_URL`, and
`VITE_PROMETHEUS_URL` have **no remaining references** in backend source,
frontend source, or the root compose files.
- `/alerts` and `/alertmanager-status` work against a configured alertmanager
service instance and return the existing "not configured" shape when none is
configured.
- `/api/monitoring/grafana-status` and `/prometheus-status` return sensible JSON
for configured / missing / unreachable services.
- The Observability page renders Alertmanager, Grafana, and Prometheus health
cards and builds Grafana deep-links from the discovered service instance.
- An `active_alerts` widget can be added to the dashboard bound to an
alertmanager service and renders firing alerts.
- Backend `pytest` + `ruff` green; frontend `npm run lint`, `npm run build`
(tsc -b + vite build), `npm run test` green.
- `docs/REQUIREMENTS.md` decision log, `CHANGELOG.md`, and `backend/README.md`
updated.
## Impact
- **Backend:** new integration module; monitoring router rewire + two new
endpoints; config.py loses three fields (`alertmanager_url`,
`alertmanager_webhook_url`, `prometheus_file_sd_dir`) and the file-writer call
sites; widget adapter + registry.
- **Frontend:** Observability page rewire (service discovery, new health cards,
Grafana URL source); new alertmanager widget component + registry entry; types;
removal of two `import.meta.env` reads.
- **Config/deploy:** root compose files and Dockerfile build args lose the
observability service env vars and the `prometheus-file-sd` volume mount;
`.env.example` needs a manual update (assistant-edit-blocked — tracked as a
task follow-up).
- **Compatibility:** `ALERTMANAGER_URL` users must re-create their Alertmanager
as a service instance after upgrading; admins using `file_sd_configs` against
the shared volume must switch to `http_sd_configs` against
`/api/monitoring/prometheus-targets`. Both breaking, called out in CHANGELOG.