Route all prometheus widget queries through Grafana /api/ds/query instead of
direct Prom HTTP. PrometheusConfig: drop base_url, add grafana_url +
datasource_uid; secret grafana_api_key (required). PrometheusWidgetSource →
MetricSource with _gateway_query POST method. normalize_grafana_frames
recovered from 65bae95 + shared _dedup_label helper. Gateway-path status
check. Startup old-config validation. CHANGELOG migration note. All adapter
tests rewritten for POST /api/ds/query + Grafana frames mock. Backend: 331
pytest pass, ruff clean. Frontend: build green (unchanged in S1).
8.2 KiB
Changelog
All notable changes to Manage. Breaking changes are marked with BREAKING.
[Unreleased]
BREAKING — Prometheus queries now route through Grafana gateway
- The
prometheusservice config changed:base_urlis replaced bygrafana_url+datasource_uid, and theapi_keysecret is replaced bygrafana_api_key(a Grafana service account token or API key with read access to the Prometheus datasource). All metric widget queries (chart,gauge,mean,metric) now issuePOST {grafana_url}/api/ds/queryinstead of direct Prometheus HTTP calls. - Migration: Reconfigure existing
prometheusservices — replacebase_urlwithgrafana_url(your Grafana instance URL), add thegrafana_api_keysecret, and optionally setdatasource_uid(defaults to"prometheus").
Added — Direct Prometheus charting
- Prometheus is now the direct source for in-app charts. New widget kinds
on the
prometheusservice:chart(multi-series line chart via recharts, backed by/api/v1/query_range),gauge(instant scalar with configurable threshold bands), andmean(client-side average over a time window).
BREAKING — Grafana service type removed
- The
grafanaservice type, Grafana link widget, Grafana chart widget, andGET /api/monitoring/grafana-statusendpoint were removed. Manage now queries Prometheus directly for all chart data. - Migration: Delete any existing Grafana service instances and create
Prometheus service instances instead (pointing at your Prometheus URL). Any
configured
grafana/chartwidgets must be recreated asprometheus/chartwidgets. Grafana link widgets are gone — use Prometheus chart/metric widgets instead.
Added — Observability service registry
- Alertmanager is now a service type. Configure Alertmanager, Grafana, and
Prometheus instances in the UI on the Services page; all three are first-class
service-registry entries with dashboard widgets (
active_alerts, Grafana link, Prometheus metric). - New monitoring endpoints resolve the configured service instance and probe its
health:
GET /api/monitoring/grafana-status,/prometheus-status. The/alertsand/alertmanager-statusendpoints now take an optionalservice_idand pick the first enabled alertmanager instance by default. - The Observability page discovers Grafana/Prometheus/Alertmanager from the
registry and renders health cards; the dashboard
active_alertswidget sums firing alerts by severity.
Changed — Observability is now external only
- Removed all observability services from
docker-compose.ymlanddocker-compose.dev.yml. They now deploy only the backend and frontend. Themonitoringnetwork and theprometheus/loki/alloy/grafana/alertmanager/node-exporterservices and their named volumes were deleted, and theGRAFANA_APP_HOSTTraefik rule was removed. - Manage now connects to existing Grafana/Prometheus/Alertmanager instances
and never ships its own stack. The previous in-compose stack is preserved as
an optional, deploy-it-yourself example in
docker-compose.observability.yml(config undermonitoring/, documented indocs/observability-runbooks.md). - Removed the now-orphaned combined
monitoring/prometheus/prometheus.yml; the standalone stack usesmonitoring/prometheus/prometheus.standalone.yml. - Removed the Prometheus file-SD bridge (
PROMETHEUS_FILE_SD_DIR+ thewrite_prometheus_targetsfile writer). External Prometheus instances now consume node-exporter targets viahttp_sd_configsagainstGET /api/monitoring/prometheus-targets. The webhook receiver is log-only.
BREAKING
- Observability is configured entirely via the service registry; the backend
alertmanager_url/alertmanager_webhook_urland frontendVITE_GRAFANA_URL/VITE_PROMETHEUS_URLenvironment variables, plusPROMETHEUS_FILE_SD_DIR, were removed. Re-create your Alertmanager / Grafana / Prometheus instances on the Services page after upgrading. The only observability env var remaining isPROMETHEUS_ENABLED(toggles Manage's own/metricsendpoint).
Added — Service registry
- Runtime service registry persisted in the backend SQLite database. External services (Grafana, Prometheus, Jellyfin, Nextcloud, SSH task runner) are now configured in the app instead of via environment variables.
- Services page (
/services) to create, list, and delete service instances. - Service detail pages (
/services/:serviceType/:serviceId) to edit name/enabled state, rotate secrets, and view the widgets a service provides. - Service definitions live as Pydantic modules in
backend/.../integrations/, each declaring its config schema, secret fields, and widget kinds. - Multi-instance support: multiple Grafana/Jellyfin/etc. instances per type.
- SSH task runner service records run history in a new
service_task_runstable, shown on the runner's service page.
Changed
- Dashboard widgets are now service-bound (reference a service instance + widget kind) or built-in (backups, static text). The "Add widget" flow is pick-service → pick-widget-kind → configure.
- Deleting a service cascade-deletes widgets that reference it.
Security
- Service secrets (API keys, tokens, passphrases) are encrypted at rest with Fernet.
BREAKING
-
Saved Actions (server tasks) now target
ssh_tasksservice instances instead of monitoring machines. Thedefault_machine_idfield on saved tasks was replaced withdefault_service_id; the legacysaved_task_runstable was dropped and run history now lives inservice_task_runs. Re-create SSH task runner services on the Services page and re-link saved actions after upgrading. -
MANAGE_ENCRYPTION_KEYis now required to start the backend. Generate one with:python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())" -
The
GRAFANA_URLandPROMETHEUS_URLbackend environment variables were removed; Grafana/Prometheus URLs now live on service records configured in the UI. Re-create them on the Services page after upgrading. -
The legacy widget/addon-pages model (
/addons/:addonId,/api/widgets/types,/api/widgets/sources) was removed in favor of the service registry. -
Default dashboard widget seeding was removed; a fresh install starts with an empty dashboard. Add widgets from the dashboard's edit dialog after configuring services.
Notes / follow-ups
Machine-level Jellyfin/Jellyseerr app config still powers the Media/Users/Files pages. Migrating those onto the service registry is a separate follow-up change.Done (2026-06-23): Jellyfin is no longer a machine service, and the dead machine-levelmedia_root/path_prefixfields were removed. See the Jellyfin migration entry indocs/REQUIREMENTS.md.
Follow-up #2 — remove dead machine media_root/path_prefix + Jellyfin service
Completes the Jellyfin migration onto the service registry. Jellyfin is no
longer a machine services tag (DEFAULT_SERVICES is now ["monitoring", "files"]), and the dead machine-level media_root/path_prefix fields were
removed from the settings store, MonitoringMachineInput, frontend types, and
the Settings UI. Jellyfin is configured exclusively as a service-registry
instance. The global REMOTE_MEDIA_ROOT/REMOTE_PATH_PREFIX config properties
and path_utils.py remain (files/media-index still use them for Jellyfin→SSH
path resolution). Existing DB rows may still carry these keys in config_json;
they are inert and get dropped on the next machine save.
Follow-up #1 — remove dead machine Jellyfin/Jellyseerr fields
With Jellyfin/Jellyseerr now resolved from the service registry, the machine-level
Jellyfin/Jellyseerr fields are dead config. Removed from dependencies.py (dead
_jellyseerr_client_for; _resolve_machine simplified to SSH-only),
services/settings_store.py, routers/settings.py (MachineInput), frontend
types, the Settings.tsx form, and frontend test fixtures. Existing DB rows may
still carry these keys in config_json; they are inert and get dropped on the
next machine save. No data migration required.