# Design — Grafana/Prometheus integration polish **Change:** `grafana-prometheus-polish` **Phase:** design **Date:** 2026-06-23 ## Current state - Backend integrations: `integrations/grafana.py` (config `base_url`, secret `api_key`, widget `link`) and `integrations/prometheus.py` (config `base_url`, secret `api_key`, widget `metric`). - Backend widget sources: `widgets/sources.py` has `GrafanaWidgetSource` and `PrometheusWidgetSource`. - Frontend `SERVICE_REGISTRY` (`frontend/src/integrations/registry.ts`) mirrors those widget kinds. - Observability page (`frontend/src/components/ObservabilityPage.tsx`) uses hard-coded `VITE_GRAFANA_URL` for Grafana links and shows only Alertmanager + Prometheus targets. - Monitoring router (`backend/src/media_library_viewer_api/routers/monitoring.py`) exposes `/alerts`, `/alertmanager-status`, `/alertmanager-webhook`, `/prometheus-targets`, `/machines`. ## Target state ### Backend 1. **New health endpoints** - `GET /api/monitoring/grafana-status` — probes `GET {base_url}/api/health` (no auth needed for health) and returns `{up, version, error?}`. - `GET /api/monitoring/prometheus-status` — probes `GET {base_url}/-/healthy` and returns `{up, version, error?}`. - If multiple instances of a type exist, probe the first enabled instance and include its `service_id` + `name` in the response. - If no instance is configured, return `up: false, error: "no_service_configured"`. 2. **New widget kinds** - Grafana: add `explore_link` kind with config `query` (optional), `datasource` (optional). Renders a deep-link to `/explore?orgId=1&left={...}` or similar. - Prometheus: add `range_chart` kind with config `promql`, `step` (e.g. `1m`), `range_seconds` (e.g. `3600`). Renders `{ series: [{ timestamp, value }] }` for the frontend to chart. The adapter performs a range query (`/api/v1/query_range`) and reshapes the result. 3. **Source adapters** - Extend `GrafanaWidgetSource.fetch` to handle `explore_link`. - Extend `PrometheusWidgetSource.fetch` to handle `range_chart`. ### Frontend 1. **Observability page** - Replace hard-coded `GRAFANA_BASE_URL` with discovery from `useServiceInstances("grafana")`. - Add `useGrafanaStatus()` and `usePrometheusStatus()` hooks using the new endpoints. - Add health cards for Grafana and Prometheus next to Alertmanager. - Grafana links (alerts panel, targets panel) use the discovered instance URL. 2. **Widget registry** - Add `explore_link` to `SERVICE_REGISTRY.grafana.widgets`. - Add `range_chart` to `SERVICE_REGISTRY.prometheus.widgets`. - Add presentational components `GrafanaExploreLinkWidget` and `PrometheusRangeChartWidget` (or extend existing components). 3. **Types** - Add status response types and widget config types. ### Docs - Update `docs/REQUIREMENTS.md` observability section to describe service-instance-driven Grafana/Prometheus status and new widget kinds. - `CHANGELOG.md` added/changed note. ## Slices 1. **Backend health endpoints + widget kind extensions.** 2. **Frontend Observability page + widget registry + new components.** 3. **Docs + changelog.** ## Risks - Grafana Explore URL encoding is version-dependent. Target Grafana 10+ and document the assumption. - Prometheus range queries can be expensive; cap `range_seconds` in the adapter (e.g. max 24h) and use a modest `step`. - Multiple Grafana/Prometheus instances: keep it simple — use the first enabled instance for status; widgets already bind to a specific service instance.