Files
manage/openspec/changes/observability-service-registry/design.md
T
Developer c13e274ca4 docs(openspec): re-scope observability-service-registry change
Rename grafana-prometheus-polish -> observability-service-registry and
rewrite proposal/design/tasks for the approved vision: all observability
integration (alertmanager, grafana, prometheus) configured as service-
registry instances in the UI, surfaced on a dedicated page, with widgets
per service definition -- nothing in the env.

Key scope decisions captured:
- Add alertmanager as a 6th service type + active_alerts widget.
- Rewire /alerts + /alertmanager-status to resolve from service records
  (first-enabled-instance default; no primary flag in v1).
- Add /grafana-status + /prometheus-status health endpoints.
- Observability page discovers services; kill VITE_GRAFANA_URL /
  VITE_PROMETHEUS_URL deep-links.
- Webhook receiver stays log-only (drop the outbound forward).
- Remove PROMETHEUS_FILE_SD_DIR + the file-writer; external Prometheus
  uses http_sd_configs against GET /api/monitoring/prometheus-targets.
  build_node_exporter_targets + that endpoint stay.
- PROMETHEUS_ENABLED stays (Manage's own /metrics toggle).
- End state: zero observability *service* env vars.

Plan = 5 slices, each <=400 changed lines, green tests/lint/build,
commit per slice.
2026-06-23 21:48:58 +00:00

11 KiB

Design — Observability service registry

Change: observability-service-registry Phase: design Date: 2026-06-23

Current state

  • Service registry (integrations/registry.py): five types — grafana, prometheus, jellyfin, jellyseerr, nextcloud, ssh_tasks. Grafana (grafana.py: base_url, secret api_key, widget link) and Prometheus (prometheus.py: base_url, secret api_key, widget metric) are already service types. Their widget sources (widgets/sources.py: GrafanaWidgetSource, PrometheusWidgetSource) already resolve URLs from the ServiceRecord, not env vars.
  • Monitoring router (routers/monitoring.py): /alerts, /alertmanager-status, /alertmanager-webhook, /prometheus-targets, /machines. The alertmanager endpoints read settings.alertmanager_url (env). _alertmanager_client() returns (requests.Session(), url); _webhook_client() returns the forward target.
  • Config (config.py): prometheus_enabled (Manage's own /metrics toggle — stays), prometheus_file_sd_dir, alertmanager_url="", alertmanager_webhook_url="". The latter three are removed by this change. services/targets.py::write_prometheus_targets is the file-writer consuming prometheus_file_sd_dir; build_node_exporter_targets + the /prometheus-targets endpoint stay (used by external Prometheus via http_sd_configs).
  • Frontend Observability page (ObservabilityPage.tsx): hard-codes GRAFANA_BASE_URL = import.meta.env.VITE_GRAFANA_URL || "http://localhost:3000" and builds Node Exporter + Loki deep-links from it. Hooks (useObservability.ts) call /alerts, /alertmanager-status, /prometheus-targets, /machines.
  • Service resolution: settings_store.list_services(service_type) orders by name ASC; get_service(id). There is no is_default/primary flag.
  • VITE_PROMETHEUS_URL is referenced only in the Dockerfile/compose build args, not in any frontend source (verified). VITE_GRAFANA_URL is read in ObservabilityPage.tsx only.
  • Webhook relay: monitoring/alertmanager/alertmanager.yml (standalone stack) points webhook_configs at http://backend:8000/api/monitoring/alertmanager-webhook. The receiver then optionally forwards to ALERTMANAGER_WEBHOOK_URL.

Target state

1. New alertmanager integration type

integrations/alertmanager.py, mirroring prometheus.py:

class AlertmanagerConfig(ServiceConfigBase):
    base_url: str
    timeout_seconds: int = 5

class AlertmanagerAlertsWidgetConfig(WidgetConfigBase):
    severity_filter: str | None = None  # optional: "critical", "warning", ...

DEFINITION = ServiceDefinition(
    service_type="alertmanager",
    name="Alertmanager",
    description="Alert routing and firing-alert summaries.",
    config_model=AlertmanagerConfig,
    secret_fields=[SecretField(key="api_key", label="API key", helper="Optional bearer token")],
    widget_kinds=[widget_kind("active_alerts", "Active alerts",
                              "Firing-alert summary from this Alertmanager.",
                              model_cls=AlertmanagerAlertsWidgetConfig,
                              default_config={}, refresh_interval_ms=15_000)],
)

Register in registry.py (SERVICE_DEFINITIONS). Add AlertmanagerWidgetSource to SERVICE_ADAPTERS in widgets/sources.py.

2. Service resolution helper (shared)

Add a single helper in the monitoring router used by the alertmanager endpoints and the new grafana/prometheus status endpoints:

def _resolve_service_record(
    store: SettingsStore, service_type: str, service_id: str | None
) -> ServiceRecord | None:
    """Return the requested instance, else the first enabled instance of type."""
    if service_id:
        row = store.get_service(service_id)
        if row and row.get("service_type") == service_type and row.get("enabled", True):
            return build_service_record(store, row)
        return None
    for row in store.list_services(service_type):
        if row.get("enabled", True):
            return build_service_record(store, row)
    return None

This factors build_service_record (already in widgets/sources.py) and keeps secret decryption in one place. Default selection = first enabled instance of type (ordered by name ASC). No DB schema change.

3. Rewired alertmanager endpoints

/alerts?service_id=<optional> and /alertmanager-status?service_id=<optional> become Depends(get_settings_store) endpoints:

  • Resolve the service record via _resolve_service_record(store, "alertmanager", service_id).
  • If None → return the existing not-configured bodies ({"error": "alertmanager_not_configured"}, {"up": False, ...}).
  • Else build base_url from service.config["base_url"], attach Authorization: Bearer <api_key> if a secret is present, and call /api/v1/alerts / /api/v2/status.
  • On exception → existing unreachable bodies.
  • Response gains service_id + name on success so the UI can show which instance was used.

_alertmanager_client() and _webhook_client() (the env readers) are deleted.

4. New health endpoints

GET /api/monitoring/grafana-status?service_id= and GET /api/monitoring/prometheus-status?service_id=:

  • Resolve via _resolve_service_record(store, "grafana"|"prometheus", service_id).
  • Probe Grafana GET {base_url}/api/health (200 → {database: "ok"}), Prometheus GET {base_url}/-/healthy + GET {base_url}/api/v1/status/buildinfo (version).
  • Return { up: bool, version: str, service_id: str, name: str, error: str | None }.
  • None configured → {up: false, version: "", service_id: "", name: "", error: "no_service_configured"}.
  • Attach bearer token if a secret is present (some auth'd setups need it).

Decision — dedicated endpoints vs. widget-data path: dedicated endpoints. The Observability page needs page-level health for the default instance independent of any widget instance; widget data fetch is per-widget and keyed on a widget kind with a different shape. Reusing _resolve_service_record keeps the service-resolution logic single-sourced; only the probe differs per type.

5. active_alerts widget adapter

AlertmanagerWidgetSource.fetch(service, "active_alerts", config) reuses the existing _summary_from_alerts to shape /api/v1/alerts into the same summary the page endpoint returns. Optional severity_filter limits severities.

6. Webhook relay decision

Decision: drop the outbound relay (alertmanager_webhook_url env var and the forward POST). The POST /alertmanager-webhook receiver stays as log-only (it records received alerts at INFO for audit/debug and returns {"status": "received"}). Rationale: the inbound→outbound relay to a second webhook is a niche feature whose target is the env var we are removing; a log-only receiver remains useful and keeps the standalone Alertmanager config working unchanged.

7. PROMETHEUS_FILE_SD_DIR decision

Decision: remove it. It is a vestige of the in-project Prometheus model where Manage's container shared a volume with a Prometheus container and wrote node_exporter_targets.json into it for file_sd_configs. In the connect-to-existing model Prometheus is a separate deployment, and a shared volume across deployments is awkward (network share / rsync hack).

The same data is already served over HTTP by GET /api/monitoring/prometheus-targets, so an external Prometheus consumes it via standard http_sd_configs instead — pull-based, no shared volume, no env var:

# external Prometheus
- job_name: node-exporter-remote
  http_sd_configs:
    - url: https://manage.example.com/api/monitoring/prometheus-targets
      refresh_interval: 30s

Concretely in this change: delete the prometheus_file_sd_dir config field, the write_prometheus_targets() file-writer call sites (startup + machine create/update/delete hooks), and the PROMETHEUS_FILE_SD_DIR compose/Dockerfile mounts. Keep services/targets.py::build_node_exporter_targets (the data builder) and the /api/monitoring/prometheus-targets HTTP endpoint.

PROMETHEUS_ENABLED stays as an env var — it toggles Manage's own /metrics endpoint, which is app instrumentation (Role 1), not service integration. It is the single legitimate observability env survivor.

Frontend changes

  1. Observability page (ObservabilityPage.tsx):
    • Discover grafana services via useServiceInstances("grafana"); pick the first enabled; build GRAFANA_BASE_URL from its base_url. Remove the import.meta.env.VITE_GRAFANA_URL read. Empty-state when none configured (link to /services).
    • Add useGrafanaStatus() and usePrometheusStatus() hooks; add Grafana and Prometheus HealthCards alongside Alertmanager.
  2. Hooks/client (useObservability.ts, api/client.ts):
    • fetchGrafanaStatus(serviceId?), fetchPrometheusStatus(serviceId?), useGrafanaStatus, usePrometheusStatus.
    • Existing fetchAlertmanagerAlerts / fetchAlertmanagerStatus keep their signatures (no service id needed for the page's default instance).
  3. Widget registry (integrations/registry.ts): add alertmanager binding with the active_alerts kind → new AlertmanagerAlertsWidget component.
  4. Types (types/index.ts): GrafanaStatus, PrometheusStatus, AlertmanagerAlertsWidgetConfig; add service_id/name to status shapes.
  5. VITE_PROMETHEUS_URL is unused in source — only its removal from the Dockerfile/compose build args is needed.

Data model / compatibility

  • No DB schema change. Alertmanager instances live in the existing services table as service_type="alertmanager".
  • Breaking for ALERTMANAGER_URL users: after upgrade the env var is ignored; the Alertmanager instance must be (re)created on the Services page. Document in CHANGELOG. No data migration (it was an env var, not a DB row).
  • Type-contract cleanup (pre-existing drift): the down-branch /alertmanager-status bodies omit name/peers. The frontend AlertmanagerStatus type declares them required but never reads them. Fix by adding name: "", peers: [] to the not-configured/unreachable branches (and service_id) so the response is internally consistent.

Open questions

  1. Default-instance flag. v1 uses "first enabled instance of type" (by name). If multiple Grafana/Prometheus/Alertmanager instances exist, status reflects only the first. A per-type "primary" flag is deferred — confirm this is acceptable, or whether to add the flag in this change (adds a UI affordance + DB column; would grow scope past one slice).
  2. Service-id in page endpoints. Should /alerts?service_id= be surfaced in the Observability page UI (an instance selector), or kept backend-only for future use? v1 leaves the page on the default instance.
  3. .env.example is assistant-edit-blocked by safety policy. The four env vars must be removed manually by the user; tracked as a task follow-up.