Files
manage/CHANGELOG.md
T
Developer b1a66a1ab7 chore(observability): remove remaining observability env vars, docs
Slice 5 (final) of observability-service-registry. Completes the move to
service-registry-only observability config: no observability service env
vars remain.

- config.py: removed alertmanager_url + alertmanager_webhook_url fields.
- docker-compose.yml / docker-compose.dev.yml: removed ALERTMANAGER_URL,
  ALERTMANAGER_WEBHOOK_URL (backend env), and VITE_GRAFANA_URL,
  VITE_PROMETHEUS_URL (frontend build args / dev env).
- frontend/Dockerfile: removed the VITE_GRAFANA_URL / VITE_PROMETHEUS_URL
  ARG, build-stage ENV, and dev-stage ENV lines.
- docs: REQUIREMENTS decision-log entry; CHANGELOG Added/Changed/BREAKING
  for the observability service registry; backend/README monitoring
  section (Observability page, services page config, http_sd_configs,
  new health endpoints, log-only webhook).

The only observability env var remaining is PROMETHEUS_ENABLED (Manage's
own /metrics toggle). Grep-gated: no live references to the removed
vars/fields in backend src, frontend src, compose, or Dockerfile.

ruff clean; 239 backend tests pass; frontend 0 lint errors, build clean,
72 tests pass.

.env.example is assistant-edit-blocked; user follow-up noted in the SDD
tasks: drop the removed vars there too.
2026-06-24 08:49:53 +00:00

115 lines
5.6 KiB
Markdown

# Changelog
All notable changes to Manage. Breaking changes are marked with **BREAKING**.
## [Unreleased]
### Added — Observability service registry
- **Alertmanager is now a service type.** Configure Alertmanager, Grafana, and
Prometheus instances in the UI on the Services page; all three are first-class
service-registry entries with dashboard widgets (`active_alerts`, Grafana link,
Prometheus metric).
- New monitoring endpoints resolve the configured service instance and probe its
health: `GET /api/monitoring/grafana-status`, `/prometheus-status`. The
`/alerts` and `/alertmanager-status` endpoints now take an optional
`service_id` and pick the first enabled alertmanager instance by default.
- The Observability page discovers Grafana/Prometheus/Alertmanager from the
registry and renders health cards; the dashboard `active_alerts` widget sums
firing alerts by severity.
### Changed — Observability is now external only
- **Removed** all observability services from `docker-compose.yml` and
`docker-compose.dev.yml`. They now deploy **only** the backend and frontend.
The `monitoring` network and the `prometheus`/`loki`/`alloy`/`grafana`/
`alertmanager`/`node-exporter` services and their named volumes were deleted,
and the `GRAFANA_APP_HOST` Traefik rule was removed.
- Manage now connects to **existing** Grafana/Prometheus/Alertmanager instances
and never ships its own stack. The previous in-compose stack is preserved as
an optional, deploy-it-yourself example in `docker-compose.observability.yml`
(config under `monitoring/`, documented in `docs/observability-runbooks.md`).
- Removed the now-orphaned combined `monitoring/prometheus/prometheus.yml`; the
standalone stack uses `monitoring/prometheus/prometheus.standalone.yml`.
- Removed the Prometheus file-SD bridge (`PROMETHEUS_FILE_SD_DIR` + the
`write_prometheus_targets` file writer). External Prometheus instances now
consume node-exporter targets via `http_sd_configs` against
`GET /api/monitoring/prometheus-targets`. The webhook receiver is log-only.
### **BREAKING**
- Observability is configured entirely via the service registry; the backend
`alertmanager_url`/`alertmanager_webhook_url` and frontend
`VITE_GRAFANA_URL`/`VITE_PROMETHEUS_URL` environment variables, plus
`PROMETHEUS_FILE_SD_DIR`, were **removed**. Re-create your Alertmanager /
Grafana / Prometheus instances on the Services page after upgrading. The only
observability env var remaining is `PROMETHEUS_ENABLED` (toggles Manage's own
`/metrics` endpoint).
### Added — Service registry
- Runtime **service registry** persisted in the backend SQLite database. External
services (Grafana, Prometheus, Jellyfin, Nextcloud, SSH task runner) are now
configured in the app instead of via environment variables.
- Services page (`/services`) to create, list, and delete service instances.
- Service detail pages (`/services/:serviceType/:serviceId`) to edit name/enabled
state, rotate secrets, and view the widgets a service provides.
- Service definitions live as Pydantic modules in `backend/.../integrations/`,
each declaring its config schema, secret fields, and widget kinds.
- Multi-instance support: multiple Grafana/Jellyfin/etc. instances per type.
- SSH task runner service records run history in a new `service_task_runs`
table, shown on the runner's service page.
### Changed
- Dashboard widgets are now **service-bound** (reference a service instance +
widget kind) or **built-in** (backups, static text). The "Add widget" flow is
pick-service → pick-widget-kind → configure.
- Deleting a service cascade-deletes widgets that reference it.
### Security
- Service secrets (API keys, tokens, passphrases) are **encrypted at rest** with
Fernet.
### **BREAKING**
- Saved Actions (server tasks) now target `ssh_tasks` service instances instead
of monitoring machines. The `default_machine_id` field on saved tasks was
replaced with `default_service_id`; the legacy `saved_task_runs` table was
dropped and run history now lives in `service_task_runs`. Re-create SSH task
runner services on the Services page and re-link saved actions after
upgrading.
- **`MANAGE_ENCRYPTION_KEY` is now required** to start the backend. Generate one
with:
```bash
python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"
```
- The `GRAFANA_URL` and `PROMETHEUS_URL` backend environment variables were
removed; Grafana/Prometheus URLs now live on service records configured in the
UI. Re-create them on the Services page after upgrading.
- The legacy widget/addon-pages model (`/addons/:addonId`,
`/api/widgets/types`, `/api/widgets/sources`) was removed in favor of the
service registry.
- Default dashboard widget seeding was removed; a fresh install starts with an
empty dashboard. Add widgets from the dashboard's edit dialog after
configuring services.
### Notes / follow-ups
- Machine-level Jellyfin/Jellyseerr app config still powers the Media/Users/Files
pages. Migrating those onto the service registry is a separate follow-up change
(see `openspec/changes/service-registry/design.md` §12.5).
## Follow-up #1 — remove dead machine Jellyfin/Jellyseerr fields
With Jellyfin/Jellyseerr now resolved from the service registry, the machine-level
Jellyfin/Jellyseerr fields are dead config. Removed from `dependencies.py` (dead
`_jellyseerr_client_for`; `_resolve_machine` simplified to SSH-only),
`services/settings_store.py`, `routers/settings.py` (`MachineInput`), frontend
types, the `Settings.tsx` form, and frontend test fixtures. Existing DB rows may
still carry these keys in `config_json`; they are inert and get dropped on the
next machine save. No data migration required.