Files
manage/openspec/changes/archive/2026-07-08-prometheus-direct-charting/proposal.md
T
Developer 7efc06a629 chore(prometheus-direct-charting): archive verified+synced change
Move to openspec/changes/archive/2026-07-08-prometheus-direct-charting/
(git mv, history preserved). 9 artifacts: proposal/spec/design/tasks/
apply-progress/verify-report/sync-report/archive-report + delta spec.
Canonical openspec/specs/prometheus-charting/ remains.
2026-07-08 23:01:35 +00:00

12 KiB

SDD Proposal: Prometheus Direct Charting (drop Grafana middleman)

Change: prometheus-direct-charting Phase: proposal Date: 2026-07-08

1. Problem / Why Now

Manage already does in-app charting — but indirectly. Today the only chart path is:

GrafanaChartWidget (recharts) ← GrafanaWidgetSource._fetch_chart
   ← POST {grafana}/api/ds/query  (datasource.type = "prometheus", hardcoded)
   ← Grafana proxies to Prometheus

Two problems with this:

  1. Grafana is a pure middleman. The backend already hardcodes datasource.type: "prometheus" in every chart query. Grafana adds a hop, an API key, a different response shape (/api/ds/query frames), and a normalization layer — all to reach a Prometheus instance Manage could query directly via /api/v1/query_range. The existing PrometheusWidgetSource already proves Manage can talk to Prom straight; it just only does instant queries today.
  2. The project's documented rules are stale. openspec/config.yaml says "Do NOT re-implement charting in-app" and "No recharts/d3 is in use." Both are already false in code: recharts ^3.9.2 is declared and imported by GrafanaChartWidget. The config documents a reality the code left behind.

Meanwhile the operator wants two more metric visualizations Manage cannot currently render: a gauge and a single value / mean-over-time widget. Both belong naturally on the prometheus service, which today only exposes an instant-query numeric widget.

The clean answer is to stop routing charts through Grafana: query Prometheus directly, reuse the rendering infrastructure already written, add the two new modes, and remove the now-redundant Grafana surface.

2. Target Users and Situations

  • Primary users: Homelab operators who want metric visualizations on the Manage dashboard without bouncing to Grafana for a quick glance.
  • Workflow moments:
    • Glance at the dashboard: see a trend line, a gauge, or a mean value for a key PromQL query.
    • Build a custom at-a-glance panel from any PromQL expression Manage can already evaluate.
    • Decommission the Grafana hop for charts (one fewer external dependency in the chart path, one fewer API key to rotate).
  • Urgency: Medium. Charts already work via Grafana today; this is a simplification plus two new widget modes, not an outage fix.

3. Product Outcome

After this change, an authenticated user can:

  • Place a Prometheus Chart widget (line chart, multi-series) backed by a direct /api/v1/query_range call — same look as today's Grafana chart, no Grafana required.
  • Place a Prometheus Gauge widget rendering an instant PromQL scalar as a gauge.
  • Place a Prometheus Mean widget rendering a single value aggregated over a time window (e.g. avg_over_time(...), or a query_range aggregate).
  • Manage the dashboard without any Grafana service configured: the Grafana service type, link widget, chart widget, and status checks are removed.

4. Scope Boundaries and Non-Goals

In scope

  • Direct Prometheus range query path — backend hits /api/v1/query_range and returns the existing {series:[{label,points}]} shape so the frontend renderer is reused unchanged.
  • Shared series normalization — extract the metric-label → readable-label logic currently inside GrafanaWidgetSource._fetch_chart into a reusable helper used by both the chart path and (where relevant) the new modes.
  • Rebrand + rebindGrafanaChartWidgetPrometheusChartWidget, moved from the grafana service to the prometheus service as the chart widget kind.
  • New prometheus widget kinds: gauge and mean (semantics confirmed in the question round).
  • Grafana removal — delete integrations/grafana.py, GrafanaWidgetSource, GrafanaLinkWidget, get_grafana_status, the ObservabilityPage Grafana section, both registries' grafana entries, nav entries, and their tests.
  • config.yaml rewrite — replace the stale thin-dashboard / no-recharts wording with the reality: in-app charting via recharts is the sanctioned approach for Prometheus-backed series; Grafana is no longer referenced.
  • Migration note — CHANGELOG entry: existing Grafana service instances must be deleted and recreated as Prometheus services (true data migration is impossible; different URLs).

Non-goals (explicitly out of scope)

  • A general chart-widget framework with pluggable data-source middlewares. Chart rendering is reused; data sources are not abstracted into a swappable adapter registry. Two concrete sources (Prometheus range query, and later the service-storage harness) are wired directly where needed.
  • Grafana datasource proxy for non-Prometheus sources (Loki, InfluxDB, Postgres). If a real non-Prom need appears later, it is a separate change.
  • Embedding Grafana panels as images/iframes. Grafana is removed, not embedded.
  • Editing PromQL in a rich editor (autocomplete, metric explorer). Plain text input only, matching today's prometheus_metric widget.
  • qBittorrent widgets and the service-storage harness. Those are a separate change (service-storage-harness); only the qBit speed widget depends on this change's chart capability.
  • Touching prometheus_metric (instant numeric widget). It stays as-is; mean and chart are sibling kinds, not modes bolted onto it.
  • Re-indexing or migrating existing widget instance rows automatically. Existing grafana chart widgets are orphaned by the removal and must be recreated as prometheus chart widgets by the operator (documented in CHANGELOG).

5. High-Level Approach

5.1 Backend

  1. Prometheus range query — extend PrometheusWidgetSource (or add a sibling code path) to handle widget_kind == "chart":
    • GET {base_url}/api/v1/query_range?query=...&start=...&end=...&step=...
    • Parse Prom {value:[ts, val]} matrix into the existing {series:[{label, points}]} shape.
  2. Shared normalization — move the "metric labels → readable label" logic out of GrafanaWidgetSource._fetch_chart into widgets/series.py (or similar), so the Prom path and any future consumer reuse it.
  3. New kinds wiringgauge and mean resolve in PrometheusWidgetSource.fetch:
    • gauge: instant query (/api/v1/query), return {value, threshold?, ...} for a gauge renderer.
    • mean: range query aggregated to a single value (either PromQL avg_over_time via instant query, or client-side mean over a query_range window). Semantics decided in the question round.
  4. Integration updateintegrations/prometheus.py declares the three widget kinds (metric, chart, gauge, mean) with config schemas (promql, plus range params for chart/mean).
  5. Grafana removal — delete integrations/grafana.py, drop grafana from SERVICE_ADAPTERS, SERVICE_DEFINITIONS, remove _fetch_chart and the Grafana link logic. Remove get_grafana_status from routers/monitoring.py and the Grafana branch from ObservabilityPage.
  6. config.yaml rewrite — replace the stale charting rules with accurate wording.

5.2 Frontend

  1. Rebrand — rename GrafanaChartWidget.tsxPrometheusChartWidget.tsx; the recharts rendering (LineChart/Line/XAxis/YAxis/Tooltip/mergeSeries/CHART_COLORS) stays essentially unchanged.
  2. Registry — in integrations/registry.ts, move chart to the prometheus binding and add gauge + mean bindings; delete the entire grafana binding.
  3. New components:
    • PrometheusGaugeWidget.tsx — recharts <RadialBarChart> or a small SVG gauge; instant value.
    • PrometheusMeanWidget.tsx — single-value display (reuses MetricCard-style rendering) of the windowed mean.
  4. Nav + ObservabilityPage — remove Grafana nav entries and the Grafana status card.
  5. Typesfrontend/src/types/index.ts drops Grafana status types; no new endpoint types (data still flows through useWidgetData).

5.3 Type contracts

  • Backend: update integrations/prometheus.py widget-kind config models; remove Grafana models.
  • Frontend: remove GrafanaStatus type; widget payloads stay {series} / {value} shaped.

6. Success Criteria / Acceptance Criteria

  1. A user can configure a Prometheus service and place chart, gauge, mean, and metric widgets without any Grafana service present.
  2. The chart widget renders multi-series line charts from /api/v1/query_range with the same look as the prior Grafana-backed chart.
  3. The gauge widget renders an instant PromQL scalar as a gauge.
  4. The mean widget renders a single value aggregated over the configured window.
  5. No grafana references remain in backend/src or frontend/src (grep clean).
  6. openspec/config.yaml no longer claims "no recharts" or "do not chart in-app"; its wording matches the implementation.
  7. Existing pytest, npm run build, and npm run lint stay green; Grafana tests are removed, Prom chart/gauge/mean tests are added.
  8. CHANGELOG documents the migration (delete Grafana services, recreate as Prometheus).

7. Risks and Mitigations

Risk Mitigation
Response-shape regression. Prom /api/v1/query_range matrix differs from Grafana /api/ds/query frames. Extract normalization into a shared helper; add a backend test that feeds a sample Prom range response and asserts the {series} shape the frontend already consumes.
Orphaned configured widgets. Existing grafana/chart widget rows break at render. Acceptable + documented in CHANGELOG; the widget resolves to "unknown widget" gracefully (existing WidgetInstance error path). No silent data corruption.
Gauge rendering complexity. recharts gauges can be fiddly. Constrain gauge to a single scalar + optional thresholds; if recharts gauge proves heavy, fall back to a ~50-line SVG gauge (contained, no new dep).
Scope creep into a generic chart framework. Tempting to abstract data sources. Non-goal enforced: two direct wirings, no adapter registry.
Stale docs/tests lingering. "No grafana references" acceptance criterion (grep) catches leftovers.
Review budget (>400 lines). Slice into chained PRs (e.g. Slice 1: Prom range query + rebrand chart; Slice 2: gauge + mean; Slice 3: Grafana removal + config rewrite). Each slice leaves build/lint/test green.
DataGrid migration (config rule callout). Not applicable — no DataGrid work here. The key technical risk is the response-shape regression above.

8. Open Questions (for the proposal question round)

  • Q1 — mean semantics. Default: client-side mean over a query_range window (e.g. last 1h, step 30s → average all returned values). Alternative: require the user to write avg_over_time(...) in the PromQL and just run an instant query. Which UX do you want?
  • Q2 — gauge thresholds. Default: green/amber/red bands at user-configured thresholds (e.g. 70%/90%). Alternative: single color, no bands (simplest).
  • Q3 — Chart time window config. Default: expose from/to/step (or a simpler "window" preset like 1h/6h/24h). Confirm the granularity users can configure.
  • Q4 — Multi-series on mean/gauge. Default: single series only (one scalar). Confirm we do not try to render multi-series gauges.

9. Future Phases

  1. Service-storage harness integration — the qBittorrent speed widget (separate service-storage-harness change) reuses this change's chart rendering with an in-service data source wired directly.
  2. Rich PromQL editing — metric explorer / autocomplete.
  3. Non-Prom datasources — only if a concrete need (Loki logs in-app) appears.
  4. Threshold-based alerting from chart widgets — surface bands/lines from Alertmanager rules.