Files
alex 4a7f24348c feat: container monitoring backend core (PR-1)
- Add instance_events and health_checks tables with Alembic migration
- InstanceEventBus: typed pub/sub singleton with wildcard support
- HealthMonitor: async background loop polling containers every 15s
- SSE endpoint GET /events/stream with auth and connection limits
- Lifecycle hooks in tool_instances.py (create/start/stop/restart/delete)
- Structured JSON logging with correlation IDs
- 15 new unit tests (EventBus, HealthMonitor, MonitoringModels)

Quality gates: pytest 15 new passed, ruff clean
2026-05-29 10:25:00 +02:00

231 lines
14 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# SDD Proposal: Container Monitoring & Notification System
## Status
**Phase:** proposal
**Date:** 2026-05-28
**Owner:** Gentle AI
**Based on:** Exploration `container-monitoring-notifications`
---
## 1. Problem Statement
Users start containers via docker compose, but when something goes wrong — a build error, a missing container, a crashed process, a failed tunnel — there is **zero visibility**. Failures are buried in server logs. The only hint is a generic 4004 error in the terminal or a stale status badge that only updates when the frontend happens to poll (every 30 seconds).
Current pain points:
- **Silent failures**: A container exits or a tunnel dies and the user doesn't know until they manually refresh.
- **No push notifications**: The frontend polls every 30s; status changes have up to 30s latency.
- **No lifecycle audit trail**: There's no record of when a container started, stopped, or crashed.
- **No health history**: The `probe_result` JSON blob is overwritten on every restart — no trend data.
- **Ephemeral logs**: Container stdout/stderr is only available via `docker logs` on-demand; nothing is persisted.
This gap was surfaced by the terminal feature: when containers fail to build or start, the terminal shows a 4004 error with no explanation, leaving users stuck.
---
## 2. Goals
1. **Real-time status push**: Users see container lifecycle events (start, stop, error, health change) within seconds, not 30s.
2. **Proactive health monitoring**: Background task continuously monitors running containers and tunnels, not just at startup.
3. **User notifications**: Toast / alert notifications when containers fail, crash, or become unhealthy.
4. **Event audit trail**: Append-only log of all instance lifecycle transitions.
5. **Health history**: Time-series snapshots of container + tunnel health for debugging trends.
6. **Structured logging**: JSON logs with correlation IDs and instance IDs for traceability.
---
## 3. Non-Goals
- **Auto-restart of crashed containers** (out of scope for MVP; may be added later).
- **Multi-replica API support** (in-memory event bus is sufficient for now; Redis/NATS deferred).
- **Log aggregation / shipping to external systems** (e.g., Loki, ELK — structured JSON logs only).
- **Email / SMS / Slack notifications** (in-app toast only for MVP).
- **Container resource metrics** (CPU, memory, disk — Docker stats not in scope).
- **Replacing docker CLI with Docker SDK** (keep existing shell-out pattern).
---
## 4. User Stories
### US-MON-001: Container Start Notification
> As a user, when I start a container, I want to see a "Container starting..." toast so I know the system is working, followed by a "Container running" toast when it's ready.
### US-MON-002: Build Failure Alert
> As a user, when a container fails to build or start, I want an immediate toast with the error message and exit code so I don't have to dig through server logs.
### US-MON-003: Tunnel Failure Detection
> As a user, when a Cloudflare tunnel dies while my container is running, I want a real-time notification so I can restart it.
### US-MON-004: Health Status History
> As a user, when my container is flapping between healthy and unhealthy, I want to see a history of health checks to diagnose the issue.
### US-MON-005: Lifecycle Audit
> As a platform operator, I want an audit log of who started/stopped/restarted which container and when, for troubleshooting and accountability.
---
## 5. Proposed Solution
### Architecture Overview
```
┌─────────────────┐ SSE ┌──────────────────┐
│ Frontend │◄─────────────│ FastAPI │
│ (toast + │ events │ SSE endpoint │
│ status badges)│ │ /events/stream │
└─────────────────┘ └────────┬─────────┘
┌───────────┴───────────┐
│ InstanceEventBus │
│ (in-memory) │
└───────────┬───────────┘
│ publish
┌─────────────────────┼─────────────────────┐
│ │ │
┌────────▼────────┐ ┌───────▼────────┐ ┌────────▼────────┐
│ Lifecycle hooks │ │ Health Monitor │ │ Instance CRUD │
│ (start/stop/ │ │ (asyncio loop) │ │ (create/delete)│
│ restart/delete)│ │ │ │ │
└─────────────────┘ └───────┬────────┘ └─────────────────┘
┌──────────▼──────────┐
│ Docker + Tunnel │
│ (poll every 15s) │
└─────────────────────┘
```
### Components
#### 5.1 InstanceEventBus (in-memory singleton)
- Pattern: Same singleton style as `TerminalManager`.
- Publishes typed events: `instance.created`, `instance.started`, `instance.stopped`, `instance.health_changed`, `instance.error`.
- Subscribers: SSE endpoint broadcasts to connected clients; health monitor subscribes for its own coordination.
#### 5.2 Background Health Monitor
- Pattern: `asyncio` loop, modeled after `TerminalManager._idle_check_loop` (every 60s → every 15s).
- For each running instance:
1. Call `docker inspect` for container status + exit code.
2. For web tools, curl the public URL for tunnel health.
3. Compare with last known state.
4. On change: update DB `status`, write `health_checks` row, publish event to bus.
- On container crash/OOM: publish `instance.error` with exit code and stderr snippet.
#### 5.3 SSE Endpoint
```
GET /events/stream
```
- FastAPI `StreamingResponse` with `text/event-stream`.
- Authenticated (same cookie/JWT as existing API).
- Sends JSON event payload per line.
- Frontend reconnects with exponential backoff on disconnect.
#### 5.4 Frontend Toast Layer
- New lightweight toast component (e.g., `sonner` or custom).
- Single SSE connection on app mount.
- Filters events by relevance (errors always shown; start/stop shown briefly).
- Also updates instance status badges in real-time (no more 30s polling lag).
#### 5.5 Database Additions
**New table: `instance_events`** — append-only audit log
```
id UUID PK
instance_id UUID FK → tool_instances.id ON DELETE CASCADE
event_type VARCHAR(50) -- created, started, stopped, restarted, deleted, health_changed, error
status VARCHAR(50) -- snapshot of instance status at time of event
message TEXT -- human-readable description / error message
created_by UUID FK → users.id
metadata JSONB -- exit_code, probe_output, tunnel_url, etc.
created_at TIMESTAMPTZ DEFAULT now()
```
**New table: `health_checks`** — periodic health snapshots
```
id UUID PK
instance_id UUID FK → tool_instances.id ON DELETE CASCADE
container_status VARCHAR(50) -- running, exited, dead, etc.
container_healthy BOOLEAN
tunnel_healthy BOOLEAN
exit_code INT
probe_status VARCHAR(50)
probe_output TEXT
checked_at TIMESTAMPTZ DEFAULT now()
```
#### 5.6 Structured Logging
- Switch API container logs to JSON format.
- Fields: `timestamp`, `level`, `logger`, `message`, `instance_id`, `event_type`, `correlation_id`.
- Container stdout/stderr remains in Docker; we do not duplicate it.
---
## 6. Key Decisions
| Decision | Choice | Rationale |
|----------|--------|-----------|
| **Event transport** | **SSE** (not WebSocket) | One-way server→client push is all we need. SSE is simpler, uses HTTP, works through proxies, and FastAPI supports it natively. WebSocket is overkill and only used for terminal bidirectional streams. |
| **Event bus** | **In-memory** (not Redis/NATS) | No new infrastructure. Single API process assumption holds today. TerminalManager already uses in-memory state. Defer distributed bus to when horizontal scaling is needed. |
| **Health monitoring** | **Background asyncio poll** (not Docker events API) | Docker CLI events API requires a persistent stream and is tricky with shell-outs. A simple poll loop every 15s is predictable, testable, and matches our existing `docker inspect` usage. |
| **Frontend polling** | **Eliminate for status** (keep for list refresh) | Instance list may still poll occasionally, but status changes and errors push via SSE. Reduces server load and gives instant UX. |
| **Notification scope** | **In-app toast only** | No external integrations for MVP. Keeps scope tight. Toast library (e.g., `sonner`) is a small dependency. |
| **Log persistence** | **Structured JSON to stdout only** | We do not build a log storage system. Docker already retains container logs. Our structured API logs can be shipped later if needed. |
| **Auto-restart** | **Out of scope** | Detect and notify, but do not automatically restart crashed containers. User must explicitly restart to avoid surprise side effects. |
---
## 7. Risks
| Risk | Severity | Likelihood | Mitigation |
|------|----------|------------|------------|
| **Docker CLI brittleness under load** | Medium | Medium | Keep poll interval conservative (15s). Reuse existing `docker.py` service; do not add new CLI patterns. Monitor `execute_compose_command` latency. |
| **SSE connection leaks** | Medium | Low | Use FastAPI background task cleanup. Close stream on client disconnect. Limit max connections per user (e.g., 5). |
| **Memory growth from event bus** | Low | Low | Event bus holds only subscriber references, not event history. Health monitor does not retain old check results. |
| **Tunnel PID fragility** | High | High | Existing risk, not introduced by this change. Health monitor will at least *detect* leaked/orphaned tunnels and surface them. |
| **Frontend SSE reconnect storms** | Medium | Low | Exponential backoff on reconnect. Jitter to prevent thundering herd. |
| **Database write amplification** | Medium | Medium | Health checks every 15s × N running instances. Write only on state change, not every poll. `health_checks` table may grow; add retention policy (e.g., 30 days) in follow-up. |
| **Scope creep into full observability** | High | Medium | Explicitly exclude metrics dashboards, log storage, alerting rules, and PagerDuty-style on-call. Stay focused on lifecycle events + toast. |
| **Multi-replica incompatibility** | Low | Low | Document that in-memory bus won't work across replicas. Add Redis/NATS only when scaling need is proven. |
---
## 8. Acceptance Criteria
- [ ] **AC-1:** `POST /instances/{id}/start` publishes `instance.started` event; frontend shows "Container starting..." toast.
- [ ] **AC-2:** If container fails during start (exit code ≠ 0), `instance.error` event is published within 5s; frontend shows error toast with message + exit code.
- [ ] **AC-3:** Background health monitor runs every 15s and detects container crashes, OOMs, and tunnel failures.
- [ ] **AC-4:** On health state change (e.g., `running → unhealthy`), `instance.health_changed` event pushes via SSE and updates status badge without page refresh.
- [ ] **AC-5:** `instance_events` table records every lifecycle transition with `event_type`, `status`, `message`, and `created_by`.
- [ ] **AC-6:** `health_checks` table records a row on every state change (not every poll) with `container_status`, `tunnel_healthy`, `exit_code`, `checked_at`.
- [ ] **AC-7:** Frontend establishes one SSE connection on app load and receives events for all user's instances.
- [ ] **AC-8:** API logs are emitted in JSON format with `instance_id`, `event_type`, and `correlation_id` fields.
- [ ] **AC-9:** No regression in existing terminal WebSocket, instance CRUD, or tunnel functionality.
- [ ] **AC-10:** Backend tests cover event bus publish/subscribe, health monitor state transitions, and SSE endpoint auth.
---
## Effort Estimate
| Phase | Files | Lines (est) | Complexity |
|-------|-------|-------------|------------|
| DB migrations + models (`instance_events`, `health_checks`) | 3 | 150 | Low |
| InstanceEventBus backend | 2 | 200 | Low |
| Health monitor background task | 2 | 300 | Medium |
| SSE endpoint + auth | 2 | 200 | Medium |
| Lifecycle hook instrumentation | 3 | 200 | Low |
| Frontend toast component + SSE client | 4 | 400 | Medium |
| Real-time status badge updates | 3 | 150 | Low |
| Structured logging refactor | 2 | 100 | Low |
| Tests | 4 | 400 | Medium |
| **Total** | **25** | **~2100** | **Medium** |
**Review workload forecast:** ~2100 lines exceeds the 400-line budget. Recommend **chained PRs**:
1. **Backend core**: Event bus, health monitor, DB migrations, SSE endpoint (~1000 lines)
2. **Frontend**: Toast component, SSE client, real-time badge updates (~700 lines)
3. **Integration + logging**: Structured JSON logs, lifecycle hooks, tests (~400 lines)
---
## Next Recommended Phase
**Design** — Detail the `InstanceEventBus` interface, health monitor state machine, SSE payload schema, and toast UX behavior. Then proceed to `tasks.md` for implementation breakdown.