044d386ac7
Every HTTP client passed an integer timeout to requests, applying the same value to BOTH connect and read phases. A slow Jellyfin /Items page or qBit /sync/maindata blew through the 10s read budget → ReadTimeoutError. Split into a (connect=5s, read=60s default) tuple via shared http_timeout() helper. The media index build worker uses a 180s read floor. Existing services with low timeout_seconds benefit from bumping to 60+.
173 lines
8.8 KiB
Markdown
173 lines
8.8 KiB
Markdown
# Changelog
|
|
|
|
All notable changes to Manage. Breaking changes are marked with **BREAKING**.
|
|
|
|
## [Unreleased]
|
|
|
|
### Fixed — HTTP read timeouts
|
|
|
|
- Service HTTP clients now use a `(connect, read)` timeout tuple (connect 5s,
|
|
read 60s default) instead of a single integer, resolving `ReadTimeoutError`
|
|
on slow Jellyfin index builds and qBittorrent stats. The media index build
|
|
worker uses a 180s read floor so slow `/Items` pages on large libraries
|
|
don't time out mid-build.
|
|
- The shared `http_timeout()` helper (`clients/http_timeout.py`) decouples
|
|
connect (fail-fast on dead hosts) from read (generous for slow responses).
|
|
- Integration `timeout_seconds` defaults were raised from 5/10s to 15/60s.
|
|
- Existing services with a low `timeout_seconds` may benefit from bumping it
|
|
to 60+ via the service editor.
|
|
|
|
### **BREAKING** — Prometheus queries now route through Grafana gateway
|
|
|
|
- The `prometheus` service config changed: `base_url` is replaced by
|
|
`grafana_url` + `datasource_uid`, and the `api_key` secret is replaced by
|
|
`grafana_api_key` (a Grafana service account token or API key with read
|
|
access to the Prometheus datasource). All metric widget queries (`chart`,
|
|
`gauge`, `mean`, `metric`) now issue `POST {grafana_url}/api/ds/query`
|
|
instead of direct Prometheus HTTP calls.
|
|
- **Migration:** Reconfigure existing `prometheus` services — replace
|
|
`base_url` with `grafana_url` (your Grafana instance URL), add the
|
|
`grafana_api_key` secret, and optionally set `datasource_uid` (defaults
|
|
to `"prometheus"`).
|
|
|
|
### Added — Direct Prometheus charting
|
|
|
|
- **Prometheus is now the direct source for in-app charts.** New widget kinds
|
|
on the `prometheus` service: `chart` (multi-series line chart via recharts,
|
|
backed by `/api/v1/query_range`), `gauge` (instant scalar with configurable
|
|
threshold bands), and `mean` (client-side average over a time window).
|
|
|
|
### **BREAKING** — Grafana service type removed
|
|
|
|
- The `grafana` service type, Grafana link widget, Grafana chart widget, and
|
|
`GET /api/monitoring/grafana-status` endpoint were **removed**. Manage now
|
|
queries Prometheus directly for all chart data.
|
|
- **Migration:** Delete any existing Grafana service instances and create
|
|
Prometheus service instances instead (pointing at your Prometheus URL). Any
|
|
configured `grafana/chart` widgets must be recreated as `prometheus/chart`
|
|
widgets. Grafana link widgets are gone — use Prometheus chart/metric widgets
|
|
instead.
|
|
|
|
### Added — Observability service registry
|
|
|
|
- **Alertmanager is now a service type.** Configure Alertmanager, Grafana, and
|
|
Prometheus instances in the UI on the Services page; all three are first-class
|
|
service-registry entries with dashboard widgets (`active_alerts`, Grafana link,
|
|
Prometheus metric).
|
|
- New monitoring endpoints resolve the configured service instance and probe its
|
|
health: `GET /api/monitoring/grafana-status`, `/prometheus-status`. The
|
|
`/alerts` and `/alertmanager-status` endpoints now take an optional
|
|
`service_id` and pick the first enabled alertmanager instance by default.
|
|
- The Observability page discovers Grafana/Prometheus/Alertmanager from the
|
|
registry and renders health cards; the dashboard `active_alerts` widget sums
|
|
firing alerts by severity.
|
|
|
|
### Changed — Observability is now external only
|
|
|
|
- **Removed** all observability services from `docker-compose.yml` and
|
|
`docker-compose.dev.yml`. They now deploy **only** the backend and frontend.
|
|
The `monitoring` network and the `prometheus`/`loki`/`alloy`/`grafana`/
|
|
`alertmanager`/`node-exporter` services and their named volumes were deleted,
|
|
and the `GRAFANA_APP_HOST` Traefik rule was removed.
|
|
- Manage now connects to **existing** Grafana/Prometheus/Alertmanager instances
|
|
and never ships its own stack. The previous in-compose stack is preserved as
|
|
an optional, deploy-it-yourself example in `docker-compose.observability.yml`
|
|
(config under `monitoring/`, documented in `docs/observability-runbooks.md`).
|
|
- Removed the now-orphaned combined `monitoring/prometheus/prometheus.yml`; the
|
|
standalone stack uses `monitoring/prometheus/prometheus.standalone.yml`.
|
|
- Removed the Prometheus file-SD bridge (`PROMETHEUS_FILE_SD_DIR` + the
|
|
`write_prometheus_targets` file writer). External Prometheus instances now
|
|
consume node-exporter targets via `http_sd_configs` against
|
|
`GET /api/monitoring/prometheus-targets`. The webhook receiver is log-only.
|
|
|
|
### **BREAKING**
|
|
|
|
- Observability is configured entirely via the service registry; the backend
|
|
`alertmanager_url`/`alertmanager_webhook_url` and frontend
|
|
`VITE_GRAFANA_URL`/`VITE_PROMETHEUS_URL` environment variables, plus
|
|
`PROMETHEUS_FILE_SD_DIR`, were **removed**. Re-create your Alertmanager /
|
|
Grafana / Prometheus instances on the Services page after upgrading. The only
|
|
observability env var remaining is `PROMETHEUS_ENABLED` (toggles Manage's own
|
|
`/metrics` endpoint).
|
|
|
|
### Added — Service registry
|
|
|
|
- Runtime **service registry** persisted in the backend SQLite database. External
|
|
services (Grafana, Prometheus, Jellyfin, Nextcloud, SSH task runner) are now
|
|
configured in the app instead of via environment variables.
|
|
- Services page (`/services`) to create, list, and delete service instances.
|
|
- Service detail pages (`/services/:serviceType/:serviceId`) to edit name/enabled
|
|
state, rotate secrets, and view the widgets a service provides.
|
|
- Service definitions live as Pydantic modules in `backend/.../integrations/`,
|
|
each declaring its config schema, secret fields, and widget kinds.
|
|
- Multi-instance support: multiple Grafana/Jellyfin/etc. instances per type.
|
|
- SSH task runner service records run history in a new `service_task_runs`
|
|
table, shown on the runner's service page.
|
|
|
|
### Changed
|
|
|
|
- Dashboard widgets are now **service-bound** (reference a service instance +
|
|
widget kind) or **built-in** (backups, static text). The "Add widget" flow is
|
|
pick-service → pick-widget-kind → configure.
|
|
- Deleting a service cascade-deletes widgets that reference it.
|
|
|
|
### Security
|
|
|
|
- Service secrets (API keys, tokens, passphrases) are **encrypted at rest** with
|
|
Fernet.
|
|
|
|
### **BREAKING**
|
|
|
|
- Saved Actions (server tasks) now target `ssh_tasks` service instances instead
|
|
of monitoring machines. The `default_machine_id` field on saved tasks was
|
|
replaced with `default_service_id`; the legacy `saved_task_runs` table was
|
|
dropped and run history now lives in `service_task_runs`. Re-create SSH task
|
|
runner services on the Services page and re-link saved actions after
|
|
upgrading.
|
|
- **`MANAGE_ENCRYPTION_KEY` is now required** to start the backend. Generate one
|
|
with:
|
|
|
|
```bash
|
|
python -c "from cryptography.fernet import Fernet; print(Fernet.generate_key().decode())"
|
|
```
|
|
|
|
- The `GRAFANA_URL` and `PROMETHEUS_URL` backend environment variables were
|
|
removed; Grafana/Prometheus URLs now live on service records configured in the
|
|
UI. Re-create them on the Services page after upgrading.
|
|
- The legacy widget/addon-pages model (`/addons/:addonId`,
|
|
`/api/widgets/types`, `/api/widgets/sources`) was removed in favor of the
|
|
service registry.
|
|
- Default dashboard widget seeding was removed; a fresh install starts with an
|
|
empty dashboard. Add widgets from the dashboard's edit dialog after
|
|
configuring services.
|
|
|
|
### Notes / follow-ups
|
|
|
|
- ~~Machine-level Jellyfin/Jellyseerr app config still powers the Media/Users/Files
|
|
pages. Migrating those onto the service registry is a separate follow-up change.~~
|
|
**Done (2026-06-23):** Jellyfin is no longer a machine service, and the dead
|
|
machine-level `media_root`/`path_prefix` fields were removed. See the
|
|
Jellyfin migration entry in `docs/REQUIREMENTS.md`.
|
|
|
|
## Follow-up #2 — remove dead machine `media_root`/`path_prefix` + Jellyfin service
|
|
|
|
Completes the Jellyfin migration onto the service registry. Jellyfin is no
|
|
longer a machine `services` tag (`DEFAULT_SERVICES` is now `["monitoring",
|
|
"files"]`), and the dead machine-level `media_root`/`path_prefix` fields were
|
|
removed from the settings store, `MonitoringMachineInput`, frontend types, and
|
|
the Settings UI. Jellyfin is configured exclusively as a service-registry
|
|
instance. The global `REMOTE_MEDIA_ROOT`/`REMOTE_PATH_PREFIX` config properties
|
|
and `path_utils.py` remain (files/media-index still use them for Jellyfin→SSH
|
|
path resolution). Existing DB rows may still carry these keys in `config_json`;
|
|
they are inert and get dropped on the next machine save.
|
|
|
|
## Follow-up #1 — remove dead machine Jellyfin/Jellyseerr fields
|
|
|
|
With Jellyfin/Jellyseerr now resolved from the service registry, the machine-level
|
|
Jellyfin/Jellyseerr fields are dead config. Removed from `dependencies.py` (dead
|
|
`_jellyseerr_client_for`; `_resolve_machine` simplified to SSH-only),
|
|
`services/settings_store.py`, `routers/settings.py` (`MachineInput`), frontend
|
|
types, the `Settings.tsx` form, and frontend test fixtures. Existing DB rows may
|
|
still carry these keys in `config_json`; they are inert and get dropped on the
|
|
next machine save. No data migration required.
|