Slice 5 (final) of observability-service-registry. Completes the move to service-registry-only observability config: no observability service env vars remain. - config.py: removed alertmanager_url + alertmanager_webhook_url fields. - docker-compose.yml / docker-compose.dev.yml: removed ALERTMANAGER_URL, ALERTMANAGER_WEBHOOK_URL (backend env), and VITE_GRAFANA_URL, VITE_PROMETHEUS_URL (frontend build args / dev env). - frontend/Dockerfile: removed the VITE_GRAFANA_URL / VITE_PROMETHEUS_URL ARG, build-stage ENV, and dev-stage ENV lines. - docs: REQUIREMENTS decision-log entry; CHANGELOG Added/Changed/BREAKING for the observability service registry; backend/README monitoring section (Observability page, services page config, http_sd_configs, new health endpoints, log-only webhook). The only observability env var remaining is PROMETHEUS_ENABLED (Manage's own /metrics toggle). Grep-gated: no live references to the removed vars/fields in backend src, frontend src, compose, or Dockerfile. ruff clean; 239 backend tests pass; frontend 0 lint errors, build clean, 72 tests pass. .env.example is assistant-edit-blocked; user follow-up noted in the SDD tasks: drop the removed vars there too.
5.6 KiB
Changelog
All notable changes to Manage. Breaking changes are marked with BREAKING.
[Unreleased]
Added — Observability service registry
- Alertmanager is now a service type. Configure Alertmanager, Grafana, and
Prometheus instances in the UI on the Services page; all three are first-class
service-registry entries with dashboard widgets (
active_alerts, Grafana link, Prometheus metric). - New monitoring endpoints resolve the configured service instance and probe its
health:
GET /api/monitoring/grafana-status,/prometheus-status. The/alertsand/alertmanager-statusendpoints now take an optionalservice_idand pick the first enabled alertmanager instance by default. - The Observability page discovers Grafana/Prometheus/Alertmanager from the
registry and renders health cards; the dashboard
active_alertswidget sums firing alerts by severity.
Changed — Observability is now external only
- Removed all observability services from
docker-compose.ymlanddocker-compose.dev.yml. They now deploy only the backend and frontend. Themonitoringnetwork and theprometheus/loki/alloy/grafana/alertmanager/node-exporterservices and their named volumes were deleted, and theGRAFANA_APP_HOSTTraefik rule was removed. - Manage now connects to existing Grafana/Prometheus/Alertmanager instances
and never ships its own stack. The previous in-compose stack is preserved as
an optional, deploy-it-yourself example in
docker-compose.observability.yml(config undermonitoring/, documented indocs/observability-runbooks.md). - Removed the now-orphaned combined
monitoring/prometheus/prometheus.yml; the standalone stack usesmonitoring/prometheus/prometheus.standalone.yml. - Removed the Prometheus file-SD bridge (
PROMETHEUS_FILE_SD_DIR+ thewrite_prometheus_targetsfile writer). External Prometheus instances now consume node-exporter targets viahttp_sd_configsagainstGET /api/monitoring/prometheus-targets. The webhook receiver is log-only.
BREAKING
- Observability is configured entirely via the service registry; the backend
alertmanager_url/alertmanager_webhook_urland frontendVITE_GRAFANA_URL/VITE_PROMETHEUS_URLenvironment variables, plusPROMETHEUS_FILE_SD_DIR, were removed. Re-create your Alertmanager / Grafana / Prometheus instances on the Services page after upgrading. The only observability env var remaining isPROMETHEUS_ENABLED(toggles Manage's own/metricsendpoint).
Added — Service registry
- Runtime service registry persisted in the backend SQLite database. External services (Grafana, Prometheus, Jellyfin, Nextcloud, SSH task runner) are now configured in the app instead of via environment variables.
- Services page (
/services) to create, list, and delete service instances. - Service detail pages (
/services/:serviceType/:serviceId) to edit name/enabled state, rotate secrets, and view the widgets a service provides. - Service definitions live as Pydantic modules in
backend/.../integrations/, each declaring its config schema, secret fields, and widget kinds. - Multi-instance support: multiple Grafana/Jellyfin/etc. instances per type.
- SSH task runner service records run history in a new
service_task_runstable, shown on the runner's service page.
Changed
- Dashboard widgets are now service-bound (reference a service instance + widget kind) or built-in (backups, static text). The "Add widget" flow is pick-service → pick-widget-kind → configure.
- Deleting a service cascade-deletes widgets that reference it.
Security
- Service secrets (API keys, tokens, passphrases) are encrypted at rest with Fernet.
BREAKING
-
Saved Actions (server tasks) now target
ssh_tasksservice instances instead of monitoring machines. Thedefault_machine_idfield on saved tasks was replaced withdefault_service_id; the legacysaved_task_runstable was dropped and run history now lives inservice_task_runs. Re-create SSH task runner services on the Services page and re-link saved actions after upgrading. -
MANAGE_ENCRYPTION_KEYis now required to start the backend. Generate one with:python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())" -
The
GRAFANA_URLandPROMETHEUS_URLbackend environment variables were removed; Grafana/Prometheus URLs now live on service records configured in the UI. Re-create them on the Services page after upgrading. -
The legacy widget/addon-pages model (
/addons/:addonId,/api/widgets/types,/api/widgets/sources) was removed in favor of the service registry. -
Default dashboard widget seeding was removed; a fresh install starts with an empty dashboard. Add widgets from the dashboard's edit dialog after configuring services.
Notes / follow-ups
- Machine-level Jellyfin/Jellyseerr app config still powers the Media/Users/Files
pages. Migrating those onto the service registry is a separate follow-up change
(see
openspec/changes/service-registry/design.md§12.5).
Follow-up #1 — remove dead machine Jellyfin/Jellyseerr fields
With Jellyfin/Jellyseerr now resolved from the service registry, the machine-level
Jellyfin/Jellyseerr fields are dead config. Removed from dependencies.py (dead
_jellyseerr_client_for; _resolve_machine simplified to SSH-only),
services/settings_store.py, routers/settings.py (MachineInput), frontend
types, the Settings.tsx form, and frontend test fixtures. Existing DB rows may
still carry these keys in config_json; they are inert and get dropped on the
next machine save. No data migration required.