chore(observability): remove remaining observability env vars, docs

Slice 5 (final) of observability-service-registry. Completes the move to
service-registry-only observability config: no observability service env
vars remain.

- config.py: removed alertmanager_url + alertmanager_webhook_url fields.
- docker-compose.yml / docker-compose.dev.yml: removed ALERTMANAGER_URL,
  ALERTMANAGER_WEBHOOK_URL (backend env), and VITE_GRAFANA_URL,
  VITE_PROMETHEUS_URL (frontend build args / dev env).
- frontend/Dockerfile: removed the VITE_GRAFANA_URL / VITE_PROMETHEUS_URL
  ARG, build-stage ENV, and dev-stage ENV lines.
- docs: REQUIREMENTS decision-log entry; CHANGELOG Added/Changed/BREAKING
  for the observability service registry; backend/README monitoring
  section (Observability page, services page config, http_sd_configs,
  new health endpoints, log-only webhook).

The only observability env var remaining is PROMETHEUS_ENABLED (Manage's
own /metrics toggle). Grep-gated: no live references to the removed
vars/fields in backend src, frontend src, compose, or Dockerfile.

ruff clean; 239 backend tests pass; frontend 0 lint errors, build clean,
72 tests pass.

.env.example is assistant-edit-blocked; user follow-up noted in the SDD
tasks: drop the removed vars there too.
This commit is contained in:
Developer
2026-06-24 08:49:53 +00:00
parent b200025daa
commit b1a66a1ab7
7 changed files with 35 additions and 25 deletions
+28 -5
View File
@@ -4,6 +4,20 @@ All notable changes to Manage. Breaking changes are marked with **BREAKING**.
## [Unreleased]
### Added — Observability service registry
- **Alertmanager is now a service type.** Configure Alertmanager, Grafana, and
Prometheus instances in the UI on the Services page; all three are first-class
service-registry entries with dashboard widgets (`active_alerts`, Grafana link,
Prometheus metric).
- New monitoring endpoints resolve the configured service instance and probe its
health: `GET /api/monitoring/grafana-status`, `/prometheus-status`. The
`/alerts` and `/alertmanager-status` endpoints now take an optional
`service_id` and pick the first enabled alertmanager instance by default.
- The Observability page discovers Grafana/Prometheus/Alertmanager from the
registry and renders health cards; the dashboard `active_alerts` widget sums
firing alerts by severity.
### Changed — Observability is now external only
- **Removed** all observability services from `docker-compose.yml` and
@@ -15,13 +29,22 @@ All notable changes to Manage. Breaking changes are marked with **BREAKING**.
and never ships its own stack. The previous in-compose stack is preserved as
an optional, deploy-it-yourself example in `docker-compose.observability.yml`
(config under `monitoring/`, documented in `docs/observability-runbooks.md`).
- The backend `alertmanager_url` default is now empty. The
`/api/monitoring/alerts` and `/alertmanager-status` endpoints return a
graceful "not configured" response when `ALERTMANAGER_URL` is unset.
- Removed the now-orphaned combined `monitoring/prometheus/prometheus.yml`; the
standalone stack uses `monitoring/prometheus/prometheus.standalone.yml`.
- `VITE_GRAFANA_URL`/`VITE_PROMETHEUS_URL` remain as optional frontend deep-link
overrides. `ALERTMANAGER_URL`/`ALERTMANAGER_WEBHOOK_URL` are optional.
- Removed the Prometheus file-SD bridge (`PROMETHEUS_FILE_SD_DIR` + the
`write_prometheus_targets` file writer). External Prometheus instances now
consume node-exporter targets via `http_sd_configs` against
`GET /api/monitoring/prometheus-targets`. The webhook receiver is log-only.
### **BREAKING**
- Observability is configured entirely via the service registry; the backend
`alertmanager_url`/`alertmanager_webhook_url` and frontend
`VITE_GRAFANA_URL`/`VITE_PROMETHEUS_URL` environment variables, plus
`PROMETHEUS_FILE_SD_DIR`, were **removed**. Re-create your Alertmanager /
Grafana / Prometheus instances on the Services page after upgrading. The only
observability env var remaining is `PROMETHEUS_ENABLED` (toggles Manage's own
`/metrics` endpoint).
### Added — Service registry