docs: add backup monitoring design spec

This commit is contained in:
2026-05-11 20:01:36 +02:00
parent 3718045d7b
commit 3fe116cf2e
@@ -0,0 +1,308 @@
# Backup Monitoring Design
**Date:** 2026-05-11
**Status:** Draft
## Overview
This document specifies a standalone backup monitoring module for the media library viewer. The system receives backup run reports from an external backup tool via HTTP, stores them in a structured yet extensible format, and provides both validation and automated alerting capabilities.
## Goals
- Receive backup statistics from an external tool via HTTP API
- Store backup job definitions and execution history
- Validate incoming data and generate alerts on anomalies
- Provide a dedicated frontend UI for backup monitoring
- Remain completely independent of the existing machine monitoring system
## Non-Goals
- Integrate with the existing monitoring machine infrastructure
- Run backup jobs or manage backup schedules
- Replace external backup tools
## Architecture
The backup monitoring system is a standalone module within the FastAPI backend with a dedicated React frontend page.
```
┌─────────────────┐ HTTP POST ┌──────────────────┐
│ Backup Tool │ ─────────────────> │ POST /api/ │
│ (external) │ Bearer token │ backups/report │
└─────────────────┘ └──────────────────┘
┌─────────────────┐ Query ┌──────────────────┐
│ React Frontend │ <───────────────── │ GET /api/ │
│ /backups │ OIDC/JWT │ backups/* │
└─────────────────┘ └──────────────────┘
┌──────────────────┐
│ SQLite Database │
│ backup_jobs │
│ backup_runs │
│ backup_alerts │
└──────────────────┘
```
## Data Model
### `BackupJob`
Represents a backup job definition. Auto-created when the first report arrives.
| Field | Type | Description |
|-------|------|-------------|
| `id` | UUID | Primary key |
| `name` | string | Unique job name (e.g., "photos-daily") |
| `source` | string | Source description (free text, e.g., "server-a:/data/photos") |
| `target` | string | Target description (free text, e.g., "server-b:/backups/photos") |
| `expected_schedule` | string | Cron expression or interval (e.g., "0 2 * * *") |
| `created_at` | datetime | When the job was first seen |
### `BackupRun`
Represents a single backup execution.
| Field | Type | Description |
|-------|------|-------------|
| `id` | UUID | Primary key |
| `job_id` | UUID | Foreign key to BackupJob |
| `started_at` | datetime | When the backup started |
| `ended_at` | datetime | When the backup ended (nullable) |
| `status` | enum | `success`, `failure`, `in_progress` |
| `bytes_transferred` | integer | Bytes transferred (nullable, >= 0) |
| `duration_ms` | integer | Duration in milliseconds (nullable) |
| `error_message` | string | Error details if status is failure (nullable) |
| `details_json` | JSON | Extensibility buffer for arbitrary fields |
| `created_at` | datetime | When the report was received |
### `BackupAlert`
Represents an alert condition detected for a job.
| Field | Type | Description |
|-------|------|-------------|
| `id` | UUID | Primary key |
| `job_id` | UUID | Foreign key to BackupJob |
| `run_id` | UUID | Related run, if applicable (nullable) |
| `alert_type` | enum | `missed_schedule`, `failed_status`, `anomaly_size`, `anomaly_duration` |
| `severity` | enum | `warning`, `critical` |
| `message` | string | Human-readable description |
| `acknowledged` | boolean | Whether a user has acknowledged this alert |
| `created_at` | datetime | When the alert was generated |
## API Specification
### Backup Tool Endpoints (Bearer Token Auth)
#### `POST /api/backups/report`
Submit a backup run report. Creates the job if it doesn't exist.
**Request Body:**
```json
{
"name": "photos-daily",
"source": "server-a:/data/photos",
"target": "server-b:/backups/photos",
"started_at": "2026-05-11T02:00:00Z",
"ended_at": "2026-05-11T02:15:30Z",
"status": "success",
"bytes_transferred": 10737418240,
"duration_ms": 930000,
"error_message": null,
"details": {
"compression_ratio": 0.75,
"files_count": 15420
}
}
```
**Validation Rules:**
- `name`: required, non-empty string
- `started_at`: required, ISO 8601 timestamp, must be in the past
- `status`: required, one of `success`, `failure`, `in_progress`
- If `status == "success"`: `ended_at` and `duration_ms` are required
- `bytes_transferred`: if provided, must be >= 0
- Unknown fields in `details` are stored in `details_json`
**Response:** `200 OK` with the created run object.
#### `POST /api/backups/report/start`
Mark a backup job as in_progress. Useful for long-running backups.
**Request Body:**
```json
{
"name": "photos-daily",
"source": "server-a:/data/photos",
"started_at": "2026-05-11T02:00:00Z"
}
```
**Response:** `200 OK` with the created run object (status: `in_progress`).
### Frontend Endpoints (OIDC/JWT Auth)
| Endpoint | Method | Description |
|----------|--------|-------------|
| `GET /api/backups/jobs` | GET | List all backup jobs with last run info |
| `GET /api/backups/jobs/{id}` | GET | Get job details with recent runs |
| `GET /api/backups/runs` | GET | List recent runs (filterable by job_id, status, date range) |
| `GET /api/backups/runs/{id}` | GET | Get run details |
| `GET /api/backups/alerts` | GET | List alerts (filterable by severity, acknowledged) |
| `POST /api/backups/alerts/{id}/acknowledge` | POST | Acknowledge an alert |
| `GET /api/dashboard/backups` | GET | Dashboard summary (job count, 24h success rate, active alerts) |
## Validation & Alerting
### Alert Generation
Alerts are generated in two ways:
1. **On report ingestion**: When a new run report is received
2. **By background poller**: Periodically checks for missed schedules
### Alert Types
| Alert Type | Trigger | Severity | Auto-Resolve |
|------------|---------|----------|--------------|
| `failed_status` | `status == "failure"` | `critical` | On next successful run |
| `missed_schedule` | No run within `expected_schedule * 1.5` | `warning` | On next run arrival |
| `anomaly_size` | `bytes_transferred == 0` or < 10% or > 300% of 7-day median | `warning` | On next normal run |
| `anomaly_duration` | `duration_ms` > 300% of 7-day median | `warning` | On next normal run |
### Alert State Machine
- Alerts are created once per condition (no duplicates for the same condition)
- `failed_status` alerts reference the specific failed run
- `missed_schedule` alerts reference the job (no run_id)
- When a new run resolves the condition, the alert is auto-resolved (add `resolved_at` timestamp)
- Users can manually acknowledge alerts to prevent repeated notifications
- Acknowledged alerts remain visible but are filtered out of active alert lists
## Frontend Design
### Navigation
Add a "Backups" item to the main navigation bar, linking to `/backups`.
### Dashboard Integration
A small "Backups" widget on the dashboard (`/`) showing:
- Total number of backup jobs
- 24-hour success rate
- Count of active (unacknowledged) alerts
- Time of last failed backup (if any)
### Backups Page (`/backups`)
Tabbed layout with three tabs:
#### Tab 1: Jobs
Sortable table with columns:
- Job name
- Source
- Target
- Expected schedule
- Last run status (color-coded chip)
- Last run time
- Next expected run time
Clicking a job row navigates to a job detail view showing historical runs.
#### Tab 2: Runs
Filterable table with columns:
- Job name
- Status (success/failure/in_progress)
- Duration
- Bytes transferred (human-readable, e.g., "10.7 GB")
- Timestamp
- Actions: view details
Filters: job name, status, date range.
#### Tab 3: Alerts
Table of active alerts with:
- Severity badge (warning/critical)
- Job name
- Alert type
- Message
- Created time
- Actions: acknowledge, view related run
### Charts
Reusing existing D3 chart infrastructure:
- Success/failure rate over time (line chart, stacked area)
- Bytes transferred per job over time (grouped bar or multi-line chart)
- Duration trends per job (line chart)
- Time range selection via brush (same pattern as monitoring charts)
## Authentication
### Backup Tool Authentication
- A single API key stored in application settings
- Backup tool sends `Authorization: Bearer <api_key>` header
- API key is auto-generated on first startup if not configured
- Key can be rotated via settings
### Frontend Authentication
- Uses existing OIDC/JWT authentication
- No additional authorization logic needed (all authenticated users can view backups)
## Error Handling
- Invalid report data returns `422 Unprocessable Entity` with detailed validation errors
- Duplicate reports for the same job within the same minute are idempotent (updates existing run)
- Malformed JSON in `details` field is accepted but stored as string in `details_json`
- Database errors are logged and return `500 Internal Server Error`
## Implementation Notes
### Database
- Use the existing SQLite database with WAL mode
- Create three new tables: `backup_jobs`, `backup_runs`, `backup_alerts`
- Add indexes on frequently queried columns: `backup_runs.job_id`, `backup_runs.status`, `backup_alerts.job_id`, `backup_alerts.acknowledged`
### Background Poller
- Extend the existing `MonitoringPoller` or create a new `BackupAlertPoller`
- Runs every 5 minutes (configurable)
- Checks for missed schedules and generates alerts
- Prunes old alerts (configurable retention, default 90 days)
### Extensibility
The `details_json` field allows the backup tool to include arbitrary metadata without schema changes. Future enhancements might include:
- Custom alert thresholds per job
- Alert notification channels (email, webhook)
- Backup verification (test restores)
- Cross-job dependency tracking
## Testing Strategy
- **Unit tests**: Validation logic, alert generation rules, median calculation
- **API tests**: Report submission, job creation, alert acknowledgment
- **Integration tests**: End-to-end flow from report submission to alert display
- **Frontend tests**: Component rendering, filtering, chart interactions
## Open Questions
1. Should we support multiple API keys (one per backup tool instance)?
2. Should alerts be sent via email or other notification channels?
3. Should we support backup job configuration via UI (pre-registering jobs)?
4. What retention policy should apply to backup runs (vs. alerts)?
## Related Documents
- [REQUIREMENTS.md](../../REQUIREMENTS.md) - Overall project requirements
- Existing monitoring system documentation (for patterns to follow)