11 KiB
Backup Monitoring Design
Date: 2026-05-11 Status: Draft
Overview
This document specifies a standalone backup monitoring module for the media library viewer. The system receives backup run reports from an external backup tool via HTTP, stores them in a structured yet extensible format, and provides both validation and automated alerting capabilities.
Goals
- Receive backup statistics from an external tool via HTTP API
- Store backup job definitions and execution history
- Validate incoming data and generate alerts on anomalies
- Provide a dedicated frontend UI for backup monitoring
- Remain completely independent of the existing machine monitoring system
Non-Goals
- Integrate with the existing monitoring machine infrastructure
- Run backup jobs or manage backup schedules
- Replace external backup tools
Architecture
The backup monitoring system is a standalone module within the FastAPI backend with a dedicated React frontend page.
┌─────────────────┐ HTTP POST ┌──────────────────┐
│ Backup Tool │ ─────────────────> │ POST /api/ │
│ (external) │ Bearer token │ backups/report │
└─────────────────┘ └──────────────────┘
│
▼
┌─────────────────┐ Query ┌──────────────────┐
│ React Frontend │ <───────────────── │ GET /api/ │
│ /backups │ OIDC/JWT │ backups/* │
└─────────────────┘ └──────────────────┘
│
▼
┌──────────────────┐
│ SQLite Database │
│ backup_jobs │
│ backup_runs │
│ backup_alerts │
└──────────────────┘
Data Model
BackupJob
Represents a backup job definition. Auto-created when the first report arrives.
| Field | Type | Description |
|---|---|---|
id |
UUID | Primary key |
name |
string | Unique job name (e.g., "photos-daily") |
source |
string | Source description (free text, e.g., "server-a:/data/photos") |
target |
string | Target description (free text, e.g., "server-b:/backups/photos") |
expected_schedule |
string | Cron expression or interval (e.g., "0 2 * * *") |
created_at |
datetime | When the job was first seen |
BackupRun
Represents a single backup execution.
| Field | Type | Description |
|---|---|---|
id |
UUID | Primary key |
job_id |
UUID | Foreign key to BackupJob |
started_at |
datetime | When the backup started |
ended_at |
datetime | When the backup ended (nullable) |
status |
enum | success, failure, in_progress |
bytes_transferred |
integer | Bytes transferred (nullable, >= 0) |
duration_ms |
integer | Duration in milliseconds (nullable) |
error_message |
string | Error details if status is failure (nullable) |
details_json |
JSON | Extensibility buffer for arbitrary fields |
created_at |
datetime | When the report was received |
BackupAlert
Represents an alert condition detected for a job.
| Field | Type | Description |
|---|---|---|
id |
UUID | Primary key |
job_id |
UUID | Foreign key to BackupJob |
run_id |
UUID | Related run, if applicable (nullable) |
alert_type |
enum | missed_schedule, failed_status, anomaly_size, anomaly_duration |
severity |
enum | warning, critical |
message |
string | Human-readable description |
acknowledged |
boolean | Whether a user has acknowledged this alert |
created_at |
datetime | When the alert was generated |
API Specification
Backup Tool Endpoints (Bearer Token Auth)
POST /api/backups/report
Submit a backup run report. Creates the job if it doesn't exist.
Request Body:
{
"name": "photos-daily",
"source": "server-a:/data/photos",
"target": "server-b:/backups/photos",
"started_at": "2026-05-11T02:00:00Z",
"ended_at": "2026-05-11T02:15:30Z",
"status": "success",
"bytes_transferred": 10737418240,
"duration_ms": 930000,
"error_message": null,
"details": {
"compression_ratio": 0.75,
"files_count": 15420
}
}
Validation Rules:
name: required, non-empty stringstarted_at: required, ISO 8601 timestamp, must be in the paststatus: required, one ofsuccess,failure,in_progress- If
status == "success":ended_atandduration_msare required bytes_transferred: if provided, must be >= 0- Unknown fields in
detailsare stored indetails_json
Response: 200 OK with the created run object.
POST /api/backups/report/start
Mark a backup job as in_progress. Useful for long-running backups.
Request Body:
{
"name": "photos-daily",
"source": "server-a:/data/photos",
"started_at": "2026-05-11T02:00:00Z"
}
Response: 200 OK with the created run object (status: in_progress).
Frontend Endpoints (OIDC/JWT Auth)
| Endpoint | Method | Description |
|---|---|---|
GET /api/backups/jobs |
GET | List all backup jobs with last run info |
GET /api/backups/jobs/{id} |
GET | Get job details with recent runs |
GET /api/backups/runs |
GET | List recent runs (filterable by job_id, status, date range) |
GET /api/backups/runs/{id} |
GET | Get run details |
GET /api/backups/alerts |
GET | List alerts (filterable by severity, acknowledged) |
POST /api/backups/alerts/{id}/acknowledge |
POST | Acknowledge an alert |
GET /api/dashboard/backups |
GET | Dashboard summary (job count, 24h success rate, active alerts) |
Validation & Alerting
Alert Generation
Alerts are generated in two ways:
- On report ingestion: When a new run report is received
- By background poller: Periodically checks for missed schedules
Alert Types
| Alert Type | Trigger | Severity | Auto-Resolve |
|---|---|---|---|
failed_status |
status == "failure" |
critical |
On next successful run |
missed_schedule |
No run within expected_schedule * 1.5 |
warning |
On next run arrival |
anomaly_size |
bytes_transferred == 0 or < 10% or > 300% of 7-day median |
warning |
On next normal run |
anomaly_duration |
duration_ms > 300% of 7-day median |
warning |
On next normal run |
Alert State Machine
- Alerts are created once per condition (no duplicates for the same condition)
failed_statusalerts reference the specific failed runmissed_schedulealerts reference the job (no run_id)- When a new run resolves the condition, the alert is auto-resolved (add
resolved_attimestamp) - Users can manually acknowledge alerts to prevent repeated notifications
- Acknowledged alerts remain visible but are filtered out of active alert lists
Frontend Design
Navigation
Add a "Backups" item to the main navigation bar, linking to /backups.
Dashboard Integration
A small "Backups" widget on the dashboard (/) showing:
- Total number of backup jobs
- 24-hour success rate
- Count of active (unacknowledged) alerts
- Time of last failed backup (if any)
Backups Page (/backups)
Tabbed layout with three tabs:
Tab 1: Jobs
Sortable table with columns:
- Job name
- Source
- Target
- Expected schedule
- Last run status (color-coded chip)
- Last run time
- Next expected run time
Clicking a job row navigates to a job detail view showing historical runs.
Tab 2: Runs
Filterable table with columns:
- Job name
- Status (success/failure/in_progress)
- Duration
- Bytes transferred (human-readable, e.g., "10.7 GB")
- Timestamp
- Actions: view details
Filters: job name, status, date range.
Tab 3: Alerts
Table of active alerts with:
- Severity badge (warning/critical)
- Job name
- Alert type
- Message
- Created time
- Actions: acknowledge, view related run
Charts
Reusing existing D3 chart infrastructure:
- Success/failure rate over time (line chart, stacked area)
- Bytes transferred per job over time (grouped bar or multi-line chart)
- Duration trends per job (line chart)
- Time range selection via brush (same pattern as monitoring charts)
Authentication
Backup Tool Authentication
- A single API key stored in application settings
- Backup tool sends
Authorization: Bearer <api_key>header - API key is auto-generated on first startup if not configured
- Key can be rotated via settings
Frontend Authentication
- Uses existing OIDC/JWT authentication
- No additional authorization logic needed (all authenticated users can view backups)
Error Handling
- Invalid report data returns
422 Unprocessable Entitywith detailed validation errors - Duplicate reports for the same job within the same minute are idempotent (updates existing run)
- Malformed JSON in
detailsfield is accepted but stored as string indetails_json - Database errors are logged and return
500 Internal Server Error
Implementation Notes
Database
- Use the existing SQLite database with WAL mode
- Create three new tables:
backup_jobs,backup_runs,backup_alerts - Add indexes on frequently queried columns:
backup_runs.job_id,backup_runs.status,backup_alerts.job_id,backup_alerts.acknowledged
Background Poller
- Extend the existing
MonitoringPolleror create a newBackupAlertPoller - Runs every 5 minutes (configurable)
- Checks for missed schedules and generates alerts
- Prunes old alerts (configurable retention, default 90 days)
Extensibility
The details_json field allows the backup tool to include arbitrary metadata without schema changes. Future enhancements might include:
- Custom alert thresholds per job
- Alert notification channels (email, webhook)
- Backup verification (test restores)
- Cross-job dependency tracking
Testing Strategy
- Unit tests: Validation logic, alert generation rules, median calculation
- API tests: Report submission, job creation, alert acknowledgment
- Integration tests: End-to-end flow from report submission to alert display
- Frontend tests: Component rendering, filtering, chart interactions
Open Questions
- Should we support multiple API keys (one per backup tool instance)?
- Should alerts be sent via email or other notification channels?
- Should we support backup job configuration via UI (pre-registering jobs)?
- What retention policy should apply to backup runs (vs. alerts)?
Related Documents
- REQUIREMENTS.md - Overall project requirements
- Existing monitoring system documentation (for patterns to follow)