# SDD Tasks: Container Monitoring & Notification System
## Review Workload Forecast
| Field | Value |
|-------|-------|
| Estimated changed lines | ~2,100 total (PR-1 ~1,000; PR-2 ~700; PR-3 ~400) |
| 400-line budget risk | High |
| Chained PRs recommended | Yes |
| Suggested split | PR 1 (Backend Core) → PR 2 (Frontend UI) → PR 3 (Integration + Polish) |
| Delivery strategy | auto-chain |
| Chain strategy | stacked-to-main |
```
Decision needed before apply: No
Chained PRs recommended: Yes
Chain strategy: stacked-to-main
400-line budget risk: High
```
> **Note:** PR-1 and PR-2 exceed the 400-line review budget. Within each PR, tasks are grouped into autonomous work units that can be reviewed independently. If review fanout is available, consider splitting PR-1 into (a) DB + EventBus + SSE and (b) HealthMonitor + Lifecycle Hooks + Logging. PR-2 can be split into (a) useEvents + ToastProvider and (b) Badge updates + Polling removal.
---
## PR-1: Backend Core
**Goal:** Establish the backend infrastructure for real-time container monitoring: database schema, in-memory event bus, background health monitor, SSE endpoint, structured logging, and lifecycle instrumentation.
**Estimated Lines:** ~1,000
**Review Risk:** High
---
### MON-PR1-001: Create Alembic migration for monitoring tables
**Description:**
Write a single Alembic revision that creates `instance_events` and `health_checks` with all columns, constraints, and indexes defined in the spec.
**Files to modify:**
- `apps/api/alembic/versions/2026_05_28_add_monitoring_tables.py` *(new)*
**Acceptance criteria:**
- [ ] Migration creates `instance_events` table with columns: `id`, `instance_id`, `event_type`, `status`, `message`, `created_by`, `metadata`, `created_at`.
- [ ] Migration creates `health_checks` table with columns: `id`, `instance_id`, `container_status`, `container_healthy`, `tunnel_healthy`, `exit_code`, `probe_status`, `probe_output`, `checked_at`.
- [ ] All 5 indexes from the spec are created.
- [ ] `upgrade()` and `downgrade()` are both implemented and pass `alembic upgrade head` / `alembic downgrade -1`.
- [ ] Migration depends on current `head` revision.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** None
---
### MON-PR1-002: Create SQLAlchemy models for InstanceEvent and HealthCheck
**Description:**
Add SQLAlchemy models matching the migration schema, following the existing `UUIDPrimaryKeyMixin` + `Base` pattern (no `TimestampMixin` on `InstanceEvent`; `created_at` uses `server_default`).
**Files to modify:**
- `apps/api/src/models/instance_event.py` *(new)*
- `apps/api/src/models/health_check.py` *(new)*
- `apps/api/src/models/__init__.py`
**Acceptance criteria:**
- [ ] `InstanceEvent` model matches spec schema with correct FKs (`ON DELETE CASCADE` / `SET NULL`).
- [ ] `HealthCheck` model matches spec schema with correct FK (`ON DELETE CASCADE`).
- [ ] Both models exported in `models/__init__.py`.
- [ ] `alembic revision --autogenerate` produces no drift against the hand-written migration.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** MON-PR1-001
---
### MON-PR1-003: Add correlation ID context variable and middleware
**Description:**
Implement an async context variable `CORRELATION_ID` and a FastAPI middleware that reads `X-Request-ID` or generates a new UUID on every request. This must be available before structured logging and event publishing.
**Files to modify:**
- `apps/api/src/services/correlation.py` *(new)*
- `apps/api/src/main.py`
**Acceptance criteria:**
- [ ] `CORRELATION_ID: contextvars.ContextVar[str]` exists with `get_correlation_id()` helper.
- [ ] `CorrelationIdMiddleware` sets the context var from `X-Request-ID` header or `uuid.uuid4()`.
- [ ] Middleware is registered in `main.py` before all routes.
- [ ] Calling `get_correlation_id()` inside a request handler returns the same ID for the full request lifecycle.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** None
---
### MON-PR1-004: Refactor API logging to structured JSON format
**Description:**
Replace the plain-text formatter in `logging_config.py` with a JSON formatter that includes `timestamp`, `level`, `logger`, `message`, `correlation_id`, `instance_id`, and `event_type`. Add a `logging.Filter` that reads from `CORRELATION_ID`.
**Files to modify:**
- `apps/api/src/logging_config.py`
**Acceptance criteria:**
- [ ] Log output is valid JSON lines with required fields.
- [ ] `correlation_id` is populated automatically from the context var.
- [ ] `instance_id` and `event_type` are included when passed as `extra=` to the logger.
- [ ] Request/response middleware logs remain functional but now emit JSON.
- [ ] Unhandled exception middleware logs tracebacks as JSON.
- [ ] `uvicorn.access` stays at `WARNING` to reduce noise.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** MON-PR1-003
---
### MON-PR1-005: Implement InstanceEventBus singleton with typed pub/sub
**Description:**
Create the in-memory event bus as a module-level singleton following the `TerminalManager` pattern. Support typed subscription, unsubscribe, and exception-isolated delivery.
**Files to modify:**
- `apps/api/src/services/event_bus.py` *(new)*
**Acceptance criteria:**
- [ ] `InstanceEventBus` is a singleton (`__new__` + lock).
- [ ] `subscribe(event_type, callback)` returns an unsubscribe callable.
- [ ] `publish(event_type, payload)` delivers to all subscribers in the same event loop iteration.
- [ ] If a subscriber raises, the exception is logged with `correlation_id` and delivery continues to remaining subscribers.
- [ ] `InstanceEventPayload` TypedDict matches the spec schema exactly.
- [ ] `unsubscribe_all(event_type)` exists for test teardown.
**Estimated effort:** Small (3–4 hours)
**Dependencies:** MON-PR1-003
---
### MON-PR1-006: Implement HealthMonitor background polling task
**Description:**
Build the background monitor that polls Docker + tunnel health every 15 seconds, compares against in-memory state, and only writes to DB / publishes events on actual state changes.
**Files to modify:**
- `apps/api/src/services/health_monitor.py` *(new)*
- `apps/api/src/services/docker.py` *(read-only; reuse `get_container_status`)*
**Acceptance criteria:**
- [ ] `HealthMonitor` accepts `event_bus: InstanceEventBus` and is a singleton-style service.
- [ ] `start()` is idempotent; creates an asyncio task for `_poll_loop()`.
- [ ] `stop()` cancels the task and clears `_last_known_state`.
- [ ] Poll interval is `15.0` seconds (configurable for tests).
- [ ] Queries all instances where `status NOT IN ("pending", "stopped", "error")`.
- [ ] Per instance: calls `get_container_status()`, then HTTP HEAD/GET to `public_url` if present.
- [ ] On state change: updates `tool_instances.status`, inserts `health_checks` row, publishes `instance.health_changed` or `instance.error`.
- [ ] On no change: skips all DB writes and event publication.
- [ ] Per-instance exceptions are caught, logged as structured JSON, and the loop continues.
- [ ] `_last_known_state` is a `dict[UUID, HealthSnapshot]` dataclass.
**Estimated effort:** Medium (5–7 hours)
**Dependencies:** MON-PR1-002, MON-PR1-005
---
### MON-PR1-007: Implement SSE streaming endpoint with auth and connection limits
**Description:**
Create `/events/stream` using FastAPI `StreamingResponse` with `text/event-stream`. Enforce authentication and a max of 5 concurrent connections per user.
**Files to modify:**
- `apps/api/src/api/events.py` *(new)*
- `apps/api/src/api/__init__.py`
**Acceptance criteria:**
- [ ] `GET /events/stream` returns `401` before stream start if auth is missing/invalid.
- [ ] Returns `429` if user already has 5 open SSE connections.
- [ ] Sends SSE `event:` and `data:` lines formatted per spec.
- [ ] Sends `:ping` comment every 30 seconds.
- [ ] Per-connection `asyncio.Queue(maxsize=100)` drops oldest events if client is slow.
- [ ] On disconnect (`asyncio.CancelledError` or client close), unsubscribes from `InstanceEventBus` and releases the connection slot.
- [ ] Router is exported from `api/__init__.py`.
**Estimated effort:** Medium (4–6 hours)
**Dependencies:** MON-PR1-005
---
### MON-PR1-008: Instrument lifecycle hooks in tool_instances.py
**Description:**
Add event publishing and audit-row writes at all lifecycle transition points in `tool_instances.py`. Create a thin `lifecycle_hooks.py` service to keep `tool_instances.py` readable.
**Files to modify:**
- `apps/api/src/services/lifecycle_hooks.py` *(new)*
- `apps/api/src/api/tool_instances.py`
**Acceptance criteria:**
- [ ] After DB commit on `POST /instances` → `instance.created` event + `instance_events` row.
- [ ] After DB commit on start begins → `instance.started` event + row.
- [ ] After probe success → `instance.health_changed` (`running`) event + row.
- [ ] After container exits during start → `instance.error` event + row.
- [ ] After DB commit on stop → `instance.stopped` event + row.
- [ ] After DB commit on restart → `instance.restarted` event + row.
- [ ] After DB commit on delete → `instance.deleted` event + row.
- [ ] `created_by` is set to `current_user.id` for user actions; `NULL` for system-detected transitions.
- [ ] `correlation_id` from the request context is propagated into the event payload.
**Estimated effort:** Medium (4–6 hours)
**Dependencies:** MON-PR1-002, MON-PR1-005, MON-PR1-003
---
### MON-PR1-009: Wire up HealthMonitor, EventBus, and events router in application startup
**Description:**
Register the new events router and start/stop the `HealthMonitor` within FastAPI lifespan events.
**Files to modify:**
- `apps/api/src/main.py`
**Acceptance criteria:**
- [ ] `events_router` is included in the main FastAPI app with appropriate prefix.
- [ ] `HealthMonitor` is instantiated with the global `InstanceEventBus` and started during app startup.
- [ ] `HealthMonitor.stop()` is called during app shutdown.
- [ ] No import cycles introduced.
- [ ] App boots and passes a smoke test (`GET /health` still works).
**Estimated effort:** Small (1–2 hours)
**Dependencies:** MON-PR1-006, MON-PR1-007
---
### MON-PR1-010: Backend unit tests — EventBus
**Description:**
Write pytest unit tests for `InstanceEventBus` covering pub/sub, exception isolation, and unsubscribe.
**Files to modify:**
- `tests/unit/test_event_bus.py` *(new)*
**Acceptance criteria:**
- [ ] `test_publish_delivers_to_all_subscribers`: 3 callbacks registered, all receive payload.
- [ ] `test_subscriber_exception_isolation`: callback A raises, B still receives event.
- [ ] `test_unsubscribe_removes_callback`: after unsubscribe, callback is not called.
- [ ] `test_publish_to_empty_subscriber_list`: no error raised.
- [ ] Tests use a fresh `InstanceEventBus` instance (reset singleton state in fixture).
**Estimated effort:** Small (2–3 hours)
**Dependencies:** MON-PR1-005
---
### MON-PR1-011: Backend unit tests — HealthMonitor
**Description:**
Write pytest unit tests for `HealthMonitor` state-transition logic using mocked Docker and HTTP responses.
**Files to modify:**
- `tests/unit/test_health_monitor.py` *(new)*
**Acceptance criteria:**
- [ ] `test_detects_container_crash`: mock `get_container_status` → `exited`, `exit_code=137`; asserts DB status becomes `error`, event published, `health_checks` row inserted.
- [ ] `test_detects_tunnel_failure`: mock tunnel HEAD → 502; asserts status → `unhealthy`, `tunnel_healthy=false` in DB.
- [ ] `test_detects_recovery`: mock running + tunnel 200 after unhealthy; asserts status → `running`, `health_checks` row inserted.
- [ ] `test_skips_writes_when_no_state_change`: two identical polls; asserts only one `health_checks` row.
- [ ] `test_docker_exception_resilience`: mock raises `CalledProcessError`; asserts no exception propagates, loop continues.
- [ ] Uses `db_session` and `event_bus` fixtures; mocks poll interval to `0.1s`.
**Estimated effort:** Medium (4–5 hours)
**Dependencies:** MON-PR1-006, MON-PR1-010
---
### MON-PR1-012: Backend integration tests — SSE endpoint
**Description:**
Write integration tests for the SSE endpoint covering auth, streaming, connection limits, and disconnect cleanup.
**Files to modify:**
- `tests/integration/test_sse_endpoint.py` *(new)*
**Acceptance criteria:**
- [ ] `test_sse_requires_auth`: `GET /events/stream` without cookie → `401`.
- [ ] `test_sse_streams_event`: authenticated client connects; backend publishes event; client receives valid SSE line within 1s.
- [ ] `test_sse_enforces_connection_limit`: open 6 connections; 6th returns `429`.
- [ ] `test_sse_disconnect_unsubscribes`: connect, close client, publish event; assert subscriber count is 0 and no error logged.
- [ ] Uses `authenticated_client` fixture.
**Estimated effort:** Medium (4–5 hours)
**Dependencies:** MON-PR1-007
---
## PR-2: Frontend UI
**Goal:** Build the frontend event consumption layer: SSE client hook, toast notification system, and real-time status badge updates.
**Estimated Lines:** ~700
**Review Risk:** High
---
### MON-PR2-001: Install sonner and create event TypeScript types
**Description:**
Add `sonner` to the frontend dependencies and create the `InstanceEventPayload` TypeScript interface that mirrors the backend spec.
**Files to modify:**
- `apps/web/package.json`
- `apps/web/src/types/events.ts` *(new)*
**Acceptance criteria:**
- [ ] `sonner` is added to `dependencies` (not `devDependencies`).
- [ ] `InstanceEventPayload` interface includes all required fields: `event`, `instance_id`, `status`, `message`, `metadata`, `timestamp`, `correlation_id`.
- [ ] `metadata` sub-type includes optional fields: `exit_code`, `tunnel_url`, `probe_output`, `error_type`, `previous_status`.
- [ ] `pnpm install` (or equivalent) succeeds and lockfile updated.
**Estimated effort:** Small (1–2 hours)
**Dependencies:** PR-1 merged (backend SSE endpoint must exist)
---
### MON-PR2-002: Implement useEvents() SSE hook with reconnect backoff
**Description:**
Create a React hook that opens an `EventSource` to `/events/stream`, handles reconnections with exponential backoff + jitter, and exposes parsed events.
**Files to modify:**
- `apps/web/src/hooks/use-events.ts` *(new)*
**Acceptance criteria:**
- [ ] Hook connects to `${API_BASE_URL}/events/stream` with credentials included.
- [ ] Parsed events are returned in a reactive list/array.
- [ ] `connected` boolean reflects `EventSource` ready state.
- [ ] On error/disconnect: waits `delay = min(30000, 1000 * 2^attempts) * (0.8 + Math.random() * 0.4)` before reconnect.
- [ ] On `401` response: stops reconnecting and redirects to login.
- [ ] On `429` response: adds extra 5s penalty before next retry.
- [ ] Hook cleans up `EventSource` on unmount.
- [ ] `reconnectCount` is exposed for debugging.
**Estimated effort:** Medium (4–5 hours)
**Dependencies:** MON-PR2-001
---
### MON-PR2-003: Implement toast rules and deduplication logic
**Description:**
Create a pure module that maps SSE event types to toast configurations and deduplicates rapid duplicate events.
**Files to modify:**
- `apps/web/src/components/toast-rules.ts` *(new)*
**Acceptance criteria:**
- [ ] `instance.started` → `info` toast, message `"Container starting..."`, duration 3s.
- [ ] `instance.health_changed` to `running` → `success` toast, message `"Container running"`, duration 3s.
- [ ] `instance.health_changed` to `unhealthy` → `warning` toast, message `"Container unhealthy"`, duration 5s.
- [ ] `instance.error` → `error` toast, uses event `message` + `metadata.exit_code` if present, duration 10s (or persistent if sonner supports it).
- [ ] Deduplication: same `(instance_id, event_type)` within 1s produces only one toast.
- [ ] Function is pure and testable without React rendering.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** MON-PR2-001
---
### MON-PR2-004: Implement ToastProvider component
**Description:**
Build a global toast provider that wraps `sonner`'s ``, consumes `useEvents()`, and renders toasts via the rules module.
**Files to modify:**
- `apps/web/src/components/toast-provider.tsx` *(new)*
- `apps/web/src/components/app-shell.tsx`
**Acceptance criteria:**
- [ ] `` mounts `` and calls `useEvents()`.
- [ ] Incoming events are passed through `toast-rules.ts` mapping.
- [ ] Mounted inside `AppShell` so it is active on every authenticated page.
- [ ] Deduplication state is managed internally (e.g., `Map` of last toast timestamp).
- [ ] Does not cause re-renders of the entire app on every SSE event (uses narrow subscription or memoization).
**Estimated effort:** Small (3–4 hours)
**Dependencies:** MON-PR2-002, MON-PR2-003
---
### MON-PR2-005: Replace health polling with real-time SSE updates in instance list
**Description:**
Remove the 30-second health polling loop from `instance-list.tsx` and `session-card.tsx`. Consume `useEvents()` to update status badges in real time. Retain a 60-second lightweight list refresh.
**Files to modify:**
- `apps/web/src/components/instance-list.tsx`
- `apps/web/src/components/session-card.tsx`
- `apps/web/src/api/sessions.ts`
**Acceptance criteria:**
- [ ] `setInterval` health polling (every 30s) is removed from `instance-list.tsx`.
- [ ] `session-card.tsx` badge colors map to statuses: `running` → green, `starting`/`probing` → blue, `unhealthy` → amber, `error` → red, `stopped` → gray.
- [ ] Badge text and color update within 1s of receiving the matching SSE event.
- [ ] `api/sessions.ts` still exports `checkInstanceHealth` for on-demand use (do not delete the function).
- [ ] A 60s list refresh poll remains for resilience (full list re-fetch, not per-instance health).
- [ ] Multiple instances update independently (no global refresh on single-instance event).
**Estimated effort:** Medium (4–5 hours)
**Dependencies:** MON-PR2-002
---
### MON-PR2-006: Frontend unit tests — useEvents hook
**Description:**
Write tests for the `useEvents` hook using mocked `EventSource` to verify reconnect logic and event parsing.
**Files to modify:**
- `apps/web/src/hooks/use-events.test.ts` *(new)*
**Acceptance criteria:**
- [ ] `test_reconnects_with_backoff`: simulate `EventSource` error; assert reconnect delay follows exponential pattern up to 30s cap.
- [ ] `test_parses_sse_event`: simulate incoming `message` event with JSON payload; assert hook state contains parsed event.
- [ ] `test_stops_on_401`: simulate 401; assert `EventSource` is closed and reconnect stops.
- [ ] `test_cleans_up_on_unmount`: unmount component; assert `EventSource.close()` called.
**Estimated effort:** Small (3–4 hours)
**Dependencies:** MON-PR2-002
---
### MON-PR2-007: Frontend unit tests — toast rules
**Description:**
Write tests for `toast-rules.ts` covering mapping correctness and deduplication.
**Files to modify:**
- `apps/web/src/components/toast-rules.test.ts` *(new)*
**Acceptance criteria:**
- [ ] `test_maps_error_event_to_error_toast`: asserts type, message includes exit code, duration.
- [ ] `test_maps_running_health_change_to_success_toast`: asserts type, message, duration.
- [ ] `test_deduplicates_within_one_second`: two identical events at t=0 and t=0.5 → one toast call.
- [ ] `test_allows_duplicate_after_one_second`: two identical events at t=0 and t=1.1 → two toast calls.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** MON-PR2-003
---
## PR-3: Integration + Polish
**Goal:** Validate the end-to-end event flow, add cross-stack integration tests, tune performance, update documentation, and ensure zero regression.
**Estimated Lines:** ~400
**Review Risk:** Medium
---
### MON-PR3-001: Integration tests — lifecycle event flow
**Description:**
Write backend integration tests that exercise real lifecycle endpoints and assert both DB audit rows and event bus publications.
**Files to modify:**
- `tests/integration/test_lifecycle_hooks.py` *(new)*
**Acceptance criteria:**
- [ ] `test_start_publishes_started_event`: call start endpoint; assert `instance_events` row with `event_type="started"` and event bus subscriber receives `instance.started`.
- [ ] `test_stop_publishes_stopped_event`: call stop endpoint; assert `event_type="stopped"` row and subscriber receives `instance.stopped`.
- [ ] `test_restart_publishes_restarted_event`: call restart endpoint; assert `event_type="restarted"`.
- [ ] `test_delete_publishes_deleted_event`: call delete endpoint; assert `event_type="deleted"`.
- [ ] `test_created_by_set_to_user_id`: user-initiated actions have `created_by` populated.
- [ ] Uses `authenticated_client`, `db_session`, and a test subscriber on `InstanceEventBus`.
**Estimated effort:** Medium (4–5 hours)
**Dependencies:** PR-1 merged, PR-2 merged
---
### MON-PR3-002: End-to-end tests — container start to toast
**Description:**
Write an E2E test (Playwright or Cypress) that starts a container and verifies the toast sequence in the browser.
**Files to modify:**
- `tests/e2e/container_monitoring.spec.ts` *(new)*
**Acceptance criteria:**
- [ ] User clicks Start on an instance.
- [ ] Toast "Container starting..." appears within 3s.
- [ ] After readiness probe passes, toast "Container running" appears within 10s.
- [ ] No manual page refresh is performed between steps.
- [ ] Test is tagged `@monitoring` for selective CI runs.
**Estimated effort:** Medium (4–6 hours)
**Dependencies:** PR-1 merged, PR-2 merged
---
### MON-PR3-003: End-to-end tests — container crash detection
**Description:**
Write an E2E test that kills a running container externally and verifies the error toast + badge update.
**Files to modify:**
- `tests/e2e/container_monitoring.spec.ts`
**Acceptance criteria:**
- [ ] Start a container and wait for "running" state.
- [ ] Kill the container via Docker CLI (or API call) from the test setup.
- [ ] Error toast appears within 5s.
- [ ] Status badge changes from green "running" to red "error" without page refresh.
- [ ] `instance_events` table contains `event_type="error"` with `exit_code`.
**Estimated effort:** Medium (4–6 hours)
**Dependencies:** MON-PR3-002
---
### MON-PR3-004: Performance tuning — connection limits and queue bounds
**Description:**
Verify and harden performance constraints: SSE queue cap, heartbeat ping, and connection-per-user limit.
**Files to modify:**
- `apps/api/src/api/events.py`
- `apps/web/src/hooks/use-events.ts`
**Acceptance criteria:**
- [ ] Per-connection `asyncio.Queue` is capped at 100 events; oldest dropped on overflow.
- [ ] SSE ping (`:ping`) is sent every 30s and confirmed with a test.
- [ ] Max 5 connections per user is enforced and load-tested (even 10 rapid tab opens).
- [ ] Frontend reconnect jitter prevents thundering herd (simulate 50 clients disconnect/reconnect).
- [ ] Document any latency findings; no regressions in existing terminal WS.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** PR-1 merged, PR-2 merged
---
### MON-PR3-005: Documentation updates
**Description:**
Add user-facing and developer-facing documentation for the monitoring system.
**Files to modify:**
- `docs/features/container-monitoring.md` *(new)*
- `docs/api/events.md` *(new)*
- `docs/architecture/event-bus.md` *(new)*
**Acceptance criteria:**
- [ ] `docs/features/container-monitoring.md` explains real-time status, toasts, and health history to users.
- [ ] `docs/api/events.md` documents `GET /events/stream` auth, headers, reconnection strategy, and event payload schema.
- [ ] `docs/architecture/event-bus.md` documents the in-memory bus design, health monitor loop, and state machine.
- [ ] README or nav index updated with links to new docs.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** PR-1 merged, PR-2 merged
---
### MON-PR3-006: Final cleanup and regression validation
**Description:**
Run the full test suite, fix any flakes, remove debug logging, and verify no existing functionality is broken.
**Files to modify:**
- Any files with temporary debug code or TODOs introduced in PR-1/PR-2.
**Acceptance criteria:**
- [ ] `pytest` passes (unit + integration) with no failures.
- [ ] Frontend build passes with no TypeScript errors.
- [ ] Existing terminal WebSocket functionality verified manually or via existing E2E tests.
- [ ] Existing instance CRUD (create, start, stop, restart, delete) works end-to-end.
- [ ] Tunnel creation and recreation still function.
- [ ] No `console.log` or debug `logger.debug` left from development.
- [ ] All TODO comments resolved or converted to tracked issues.
- [ ] CHANGELOG or release notes entry added if project maintains one.
**Estimated effort:** Small (2–3 hours)
**Dependencies:** MON-PR3-001, MON-PR3-002, MON-PR3-003, MON-PR3-004
---
## Dependency Graph (PR Level)
```
PR-1: Backend Core
│
├─► MON-PR1-001 ──► MON-PR1-002
│
├─► MON-PR1-003 ──► MON-PR1-004
│ └─► MON-PR1-008
│
├─► MON-PR1-005 ──► MON-PR1-006 ──► MON-PR1-009
│ │
│ └─► MON-PR1-007 ──► MON-PR1-012
│
├─► MON-PR1-010
│
└─► MON-PR1-011
PR-2: Frontend UI (depends on PR-1 merged)
│
├─► MON-PR2-001 ──► MON-PR2-002 ──► MON-PR2-004
│ │
│ └─► MON-PR2-005
│
├─► MON-PR2-003 ──► MON-PR2-004
│
├─► MON-PR2-006
│
└─► MON-PR2-007
PR-3: Integration + Polish (depends on PR-1 + PR-2 merged)
│
├─► MON-PR3-001
│
├─► MON-PR3-002 ──► MON-PR3-003
│
├─► MON-PR3-004
│
├─► MON-PR3-005
│
└─► MON-PR3-006
```
---
## Task Summary
| PR | Task ID | Description | Effort |
|----|---------|-------------|--------|
| 1 | MON-PR1-001 | Alembic migration for monitoring tables | S |
| 1 | MON-PR1-002 | SQLAlchemy models for InstanceEvent and HealthCheck | S |
| 1 | MON-PR1-003 | Correlation ID context variable and middleware | S |
| 1 | MON-PR1-004 | Structured JSON logging refactor | S |
| 1 | MON-PR1-005 | InstanceEventBus singleton | S |
| 1 | MON-PR1-006 | HealthMonitor background polling task | M |
| 1 | MON-PR1-007 | SSE streaming endpoint | M |
| 1 | MON-PR1-008 | Lifecycle hook instrumentation | M |
| 1 | MON-PR1-009 | Wire up startup/shutdown and router registration | S |
| 1 | MON-PR1-010 | Unit tests — EventBus | S |
| 1 | MON-PR1-011 | Unit tests — HealthMonitor | M |
| 1 | MON-PR1-012 | Integration tests — SSE endpoint | M |
| 2 | MON-PR2-001 | Install sonner + TypeScript event types | S |
| 2 | MON-PR2-002 | useEvents() SSE hook | M |
| 2 | MON-PR2-003 | Toast rules and deduplication | S |
| 2 | MON-PR2-004 | ToastProvider component | S |
| 2 | MON-PR2-005 | Real-time badge updates + polling removal | M |
| 2 | MON-PR2-006 | Unit tests — useEvents hook | S |
| 2 | MON-PR2-007 | Unit tests — toast rules | S |
| 3 | MON-PR3-001 | Integration tests — lifecycle event flow | M |
| 3 | MON-PR3-002 | E2E tests — container start to toast | M |
| 3 | MON-PR3-003 | E2E tests — container crash detection | M |
| 3 | MON-PR3-004 | Performance tuning (limits, queue, jitter) | S |
| 3 | MON-PR3-005 | Documentation updates | S |
| 3 | MON-PR3-006 | Final cleanup and regression validation | S |
**Total tasks:** 25
**Total estimated effort:** ~95 hours (backend ~55h, frontend ~25h, integration ~15h)