# SDD Design: Container Monitoring & Notification System
## Status
**Phase:** design
**Date:** 2026-05-28
**Owner:** Gentle AI
**Scope:** Cross-cutting (backend + frontend)
**Est. Lines:** ~2,100 (recommend 3 chained PRs)
---
## 1. Component Architecture
### 1.1 InstanceEventBus — In-Memory Singleton Pub/Sub
**Pattern:** Module-level singleton, modeled after `TerminalManager` (`apps/api/src/services/terminal_manager.py`).
**Responsibilities:**
- Maintain a registry of typed subscribers (`instance.created`, `instance.started`, `instance.stopped`, `instance.restarted`, `instance.deleted`, `instance.health_changed`, `instance.error`).
- Deliver events to all subscribers in the same asyncio event loop iteration.
- Catch subscriber exceptions, log them with `correlation_id`, and continue delivery.
- Provide no persistence or queuing; offline subscribers miss events.
**Class:**
```python
class InstanceEventBus:
_instance: "InstanceEventBus | None" = None
_lock: asyncio.Lock = asyncio.Lock()
def __new__(cls) -> "InstanceEventBus": ...
def subscribe(
self,
event_type: str,
callback: Callable[[InstanceEventPayload], Awaitable[None] | None],
) -> Callable[[], None]: ...
def unsubscribe(self, event_type: str, callback_id: str) -> None: ...
async def publish(self, event_type: str, payload: InstanceEventPayload) -> None: ...
```
**Payload type:**
```python
class InstanceEventPayload(TypedDict):
event: str
instance_id: str
status: str | None
message: str | None
metadata: dict[str, Any]
timestamp: str # ISO 8601 UTC
correlation_id: str # UUID
```
**Location:** `apps/api/src/services/event_bus.py`
---
### 1.2 HealthMonitor — Asyncio Background Task
**Pattern:** Singleton background task, modeled after `TerminalManager._idle_check_loop()`.
**Responsibilities:**
- Poll every 15 seconds for all instances whose `status` is NOT IN `("pending", "stopped", "error")`.
- For each candidate:
1. Call `docker inspect` via `get_container_status()` in `docker.py`.
2. For web tools with `public_url`, perform HTTP HEAD/GET to check tunnel health.
3. Compare against last known in-memory state (`_last_known_state: dict[UUID, HealthSnapshot]`).
- On state change:
1. Update `tool_instances.status` in DB.
2. Insert row into `health_checks`.
3. Publish appropriate event to `InstanceEventBus`.
- Catch all exceptions per-instance, log structured error, and continue to next instance.
**Class:**
```python
class HealthMonitor:
def __init__(self, event_bus: InstanceEventBus) -> None: ...
def start(self) -> None:
"""Idempotent start of the background polling task."""
def stop(self) -> None:
"""Cancel the background task and clear state."""
async def _poll_loop(self) -> None: ...
async def _check_instance(self, session: AsyncSession, instance: ToolInstance) -> None: ...
async def _publish_state_change(
self,
instance: ToolInstance,
previous: HealthSnapshot,
current: HealthSnapshot,
) -> None: ...
```
**Location:** `apps/api/src/services/health_monitor.py`
---
### 1.3 SSEManager — FastAPI StreamingResponse
**Pattern:** Stateless generator endpoint that bridges `InstanceEventBus` to HTTP `text/event-stream`.
**Responsibilities:**
- Authenticate via existing cookie/JWT (`get_current_user_id`).
- Return `401` before starting stream if auth fails.
- Subscribe a per-connection async callback to `InstanceEventBus`.
- Yield SSE `data:` lines formatted as JSON.
- Send SSE comment `:ping` every 30 seconds to keep proxies alive.
- On disconnect (`asyncio.CancelledError` / client close), unsubscribe and release.
- Enforce max 5 concurrent SSE connections per user.
**Endpoint:**
```python
@router.get("/events/stream")
async def events_stream(
request: Request,
user_id: uuid.UUID = Depends(get_current_user_id),
) -> StreamingResponse:
...
```
**Location:** `apps/api/src/api/events.py`
---
### 1.4 LifecycleHookService — Instrumentation Points
**Responsibilities:**
- Thin wrapper around existing lifecycle endpoints in `tool_instances.py`.
- At each lifecycle action (create, start, stop, restart, delete), publish the corresponding typed event **after** the DB transaction commits.
- Record an `instance_events` audit row for every transition.
- Pass `created_by` (current user ID) for user-initiated actions; `NULL` for system-detected transitions.
**Integration points (all in `apps/api/src/api/tool_instances.py`):**
| Endpoint | Event Published | Status | Audit Row |
|----------|----------------|--------|-----------|
| `POST /instances` | `instance.created` | `"pending"` | Yes |
| `POST /instances/{id}/start` | `instance.started` | `"starting"` | Yes |
| Probe success | `instance.health_changed` | `"running"` | Yes |
| Container exits during start | `instance.error` | `"error"` | Yes |
| `POST /instances/{id}/stop` | `instance.stopped` | `"stopped"` | Yes |
| `POST /instances/{id}/restart` | `instance.restarted` | `"starting"` | Yes |
| `DELETE /instances/{id}` | `instance.deleted` | `"deleted"` | Yes |
**Helper:** `LifecycleHookService` class or module-level async functions in `apps/api/src/services/lifecycle_hooks.py`.
---
### 1.5 ToastComponent — Frontend Event Consumer
**Responsibilities:**
- Single global `` component mounted in `AppShell`.
- Subscribes to SSE via `useEvents()` hook.
- Filters incoming events and maps to toast rules:
- `instance.error` → error toast, persistent (min 10s).
- `instance.started` → info toast, 3s.
- `instance.health_changed` → `running` = success 3s; `unhealthy` = warning 5s.
- Deduplicates toasts for same `(instance_id, event_type)` within 1s.
- Exposes a `toast.dismiss(id)` API.
**Technology choice:** `sonner` (lightweight, headless-compatible) or a custom 150-line toast stack. **Decision:** Use `sonner` to minimize custom UI code.
**Locations:**
- `apps/web/src/components/toast-provider.tsx` — wraps `Toaster` + `useEvents`.
- `apps/web/src/components/toast-rules.ts` — event-to-toast mapping logic.
---
## 2. File Structure
### New Files
| File | Purpose |
|------|---------|
| `apps/api/src/services/event_bus.py` | `InstanceEventBus` singleton + `InstanceEventPayload` type |
| `apps/api/src/services/health_monitor.py` | `HealthMonitor` background task + `HealthSnapshot` dataclass |
| `apps/api/src/services/lifecycle_hooks.py` | Helper functions to publish lifecycle events and write audit rows |
| `apps/api/src/services/correlation.py` | Async context var `CORRELATION_ID` + middleware injection |
| `apps/api/src/api/events.py` | SSE endpoint `/events/stream` + connection limiter |
| `apps/api/src/models/instance_event.py` | SQLAlchemy `InstanceEvent` model |
| `apps/api/src/models/health_check.py` | SQLAlchemy `HealthCheck` model |
| `apps/api/alembic/versions/2026_05_28_add_monitoring_tables.py` | Alembic revision creating `instance_events` + `health_checks` + indexes |
| `apps/web/src/hooks/use-events.ts` | `useEvents()` hook: SSE connect, reconnect backoff, event parsing |
| `apps/web/src/components/toast-provider.tsx` | Global toast provider consuming SSE events |
| `apps/web/src/components/toast-rules.ts` | Event-to-toast mapping and deduplication logic |
| `apps/web/src/types/events.ts` | TypeScript `InstanceEventPayload` interface |
| `tests/unit/test_event_bus.py` | EventBus pub/sub, exception isolation, unsubscribe |
| `tests/unit/test_health_monitor.py` | State transition logic, DB write gating |
| `tests/integration/test_sse_endpoint.py` | SSE auth, streaming, disconnect cleanup |
### Modified Files
| File | Purpose |
|------|---------|
| `apps/api/src/api/tool_instances.py` | Inject lifecycle hook calls at create/start/stop/restart/delete; pass `correlation_id` through async context |
| `apps/api/src/main.py` | Import `events_router`; register at startup; start `HealthMonitor`; add `CorrelationIdMiddleware` |
| `apps/api/src/logging_config.py` | Replace plain-text formatter with JSON formatter; include `correlation_id`, `instance_id`, `event_type` fields |
| `apps/api/src/models/__init__.py` | Export `InstanceEvent`, `HealthCheck` for Alembic autogenerate |
| `apps/web/src/components/instance-list.tsx` | Remove 30s health polling; consume `useEvents` for real-time badge updates; retain 60s list refresh |
| `apps/web/src/components/session-card.tsx` | Update badge colors based on SSE `status` events |
| `apps/web/src/components/app-shell.tsx` | Mount `` |
| `apps/web/src/api/sessions.ts` | Remove `checkInstanceHealth` polling call (keep function for on-demand use) |
| `apps/web/package.json` | Add `sonner` dependency |
| `tests/conftest.py` (or api equivalent) | Add `event_bus` fixture and `health_monitor` fixture for tests |
---
## 3. Interface Design
### 3.1 EventBus
```python
# apps/api/src/services/event_bus.py
class InstanceEventBus:
"""In-memory typed event bus. Singleton per process."""
def subscribe(
self,
event_type: str,
callback: Callable[[InstanceEventPayload], Awaitable[None] | None],
) -> Callable[[], None]:
"""Register a callback for an event type. Returns an unsubscribe function."""
async def publish(self, event_type: str, payload: InstanceEventPayload) -> None:
"""Deliver payload to all subscribers of event_type."""
def unsubscribe_all(self, event_type: str) -> None:
"""Remove all subscribers for an event type (used in tests)."""
```
**Usage in SSE endpoint:**
```python
async def event_generator(user_id: uuid.UUID):
queue: asyncio.Queue[InstanceEventPayload] = asyncio.Queue()
async def on_event(payload: InstanceEventPayload) -> None:
await queue.put(payload)
unsubscribe = event_bus.subscribe("*", on_event) # or per-type
try:
while True:
payload = await asyncio.wait_for(queue.get(), timeout=30.0)
yield f"event: {payload['event']}\ndata: {json.dumps(payload)}\n\n"
finally:
unsubscribe()
```
### 3.2 HealthMonitor
```python
# apps/api/src/services/health_monitor.py
class HealthMonitor:
POLL_INTERVAL_SECONDS: float = 15.0
MAX_STARTUP_WAIT_SECONDS: float = 30.0
def __init__(self, event_bus: InstanceEventBus) -> None: ...
def start(self) -> None:
"""Idempotent. Creates `asyncio.Task` for `_poll_loop`."""
def stop(self) -> None:
"""Cancel task and clear `_last_known_state`."""
async def force_check(self, instance_id: uuid.UUID) -> None:
"""Immediate check for a single instance (used in tests)."""
```
### 3.3 SSEManager
```python
# apps/api/src/api/events.py
@router.get("/events/stream")
async def events_stream(
request: Request,
user_id: uuid.UUID = Depends(get_current_user_id),
) -> StreamingResponse:
...
```
**Headers returned:**
- `Content-Type: text/event-stream`
- `Cache-Control: no-cache`
- `Connection: keep-alive`
- `X-Accel-Buffering: no` (disable nginx buffering)
**Rate limit:** Max 5 concurrent connections per `user_id`. Return `429` if exceeded.
### 3.4 Frontend: useEvents() Hook
```typescript
// apps/web/src/hooks/use-events.ts
export interface UseEventsReturn {
events: InstanceEventPayload[];
connected: boolean;
reconnectCount: number;
error: Error | null;
}
export function useEvents(): UseEventsReturn {
// Establishes SSE connection to `${BASE_URL}/events/stream`
// with exponential backoff reconnect.
}
```
**Reconnect strategy (client-side):**
- Initial delay: `1000ms`
- Multiplier: `2×`
- Cap: `30000ms`
- Jitter: `±20%` (`delay * (0.8 + Math.random() * 0.4)`)
- Max reconnect attempts: unlimited (persistent connection)
### 3.5 Correlation ID Propagation
```python
# apps/api/src/services/correlation.py
import contextvars
CORRELATION_ID: contextvars.ContextVar[str] = contextvars.ContextVar("correlation_id")
def get_correlation_id() -> str:
try:
return CORRELATION_ID.get()
except LookupError:
return str(uuid.uuid4())
```
**Middleware:** `CorrelationIdMiddleware` reads `X-Request-ID` header or generates new UUID, sets `CORRELATION_ID`, and includes it in all logs via a custom `logging.Filter`.
---
## 4. Data Flow Diagrams
### 4.1 Container Start Flow
```
User clicks Start
│
▼
POST /instances/{id}/start
│
├──► DB: tool_instances.status = "starting"
│
├──► LifecycleHookService.publish("instance.started", {status: "starting", ...})
│ │
│ ▼
│ InstanceEventBus
│ │
│ ├──► SSEManager ──► Frontend toast: "Container starting..."
│ │
│ └──► InstanceEvent DB write (audit)
│
├──► docker compose up -d
│
├──► wait_for_container_running()
│ │
│ ├──► Success ──► DB.status = "running"
│ │ LifecycleHookService.publish("instance.health_changed",
│ │ {status: "running", previous_status: "starting"})
│ │ │
│ │ ▼
│ │ Frontend toast: "Container running"
│ │
│ └──► Failure ──► DB.status = "error"
│ LifecycleHookService.publish("instance.error",
│ {status: "error", metadata: {exit_code, ...}})
│ │
│ ▼
│ Frontend toast: Error (persistent)
```
### 4.2 Health Monitor Flow
```
HealthMonitor._poll_loop() (every 15s)
│
├──► SELECT * FROM tool_instances WHERE status NOT IN ("pending","stopped","error")
│
├──► For each instance:
│ │
│ ├──► get_container_status(container_id) ──► {State.Status, ExitCode, Health.Status}
│ │
│ ├──► if public_url: HTTP HEAD public_url ──► tunnel_healthy?
│ │
│ ├──► Compare with _last_known_state[instance_id]
│ │
│ ├──► If changed:
│ │ │
│ │ ├──► DB: UPDATE tool_instances SET status = ?
│ │ │
│ │ ├──► DB: INSERT INTO health_checks (...)
│ │ │
│ │ └──► EventBus.publish("instance.health_changed" OR "instance.error")
│ │ │
│ │ ▼
│ │ Frontend badge + toast update
│ │
│ └──► If unchanged: skip DB writes
│
└──► Catch exception per-instance ──► structured JSON log ──► continue next instance
```
### 4.3 SSE Flow
```
Frontend mount
│
▼
EventSource.open("GET /events/stream")
│
├──► Server: auth cookie validation
│ │
│ ├──► Invalid ──► 401 (no stream)
│ │
│ └──► Valid ──► check connection count ≤ 5
│ │
│ ├──► Exceeded ──► 429
│ │
│ └──► OK ──► StreamingResponse
│ │
│ ├──► Subscribe callback to EventBus
│ │
│ ├──► yield "event: ...\ndata: {...}\n\n"
│ │
│ ├──► yield ":ping\n" (every 30s)
│ │
│ └──► Client disconnect
│ │
│ ├──► asyncio.CancelledError
│ └──► Unsubscribe callback
│
└──► Network interruption ──► Frontend closes EventSource
│
├──► wait exponential backoff + jitter
│
└──► reopen EventSource (repeat from top)
```
---
## 5. State Machine
### 5.1 Instance Status Transitions
```
+-----------+
| pending |
+-----+-----+
│ create()
v
+-----------+ build/compose failure +-------+
| starting +-------------------------------->│ error │
+-----+-----+ +---+---+
│ probe passes / monitor finds running │ restart()
v v
+-----------+ crash / OOM / exit ≠ 0 +-----------+
+--->| running +-------------------------------->│ error |
| +-----+-----+ +-----------+
| │ tunnel/probe fail
| v
| +-----------+ recover (tunnel OK) +-----------+
+----+ unhealthy +-------------------------------->│ running |
+-----+-----+ +-----------+
│ stop()
v
+-----------+
| stopped |
+-----------+
│ delete()
v
[gone]
```
### 5.2 Transition Triggers
| From | To | Trigger | DB Update | Event Published | Audit Row |
|------|----|---------|-----------|-----------------|-----------|
| `pending` | `starting` | User clicks Start | Yes | `instance.started` | Yes |
| `starting` | `running` | Readiness probe passes | Yes | `instance.health_changed` | Yes |
| `starting` | `error` | Container exits during start | Yes | `instance.error` | Yes |
| `running` | `unhealthy` | Monitor: tunnel down or probe fail | Yes | `instance.health_changed` | Yes |
| `running` | `error` | Monitor: container crashed / OOM | Yes | `instance.error` | Yes |
| `unhealthy` | `running` | Monitor: recovery detected | Yes | `instance.health_changed` | Yes |
| `running` | `stopped` | User clicks Stop | Yes | `instance.stopped` | Yes |
| `unhealthy` | `stopped` | User clicks Stop | Yes | `instance.stopped` | Yes |
| `error` | `starting` | User clicks Restart | Yes | `instance.restarted` | Yes |
| any | `deleted` | User clicks Delete | Yes (then row removed) | `instance.deleted` | Yes |
**Rule:** The monitor only evaluates instances with `status` in `{"starting", "running", "unhealthy"}`. It does NOT evaluate `pending`, `stopped`, or `error`.
---
## 6. Database Schema
### 6.1 Table: `instance_events`
```sql
CREATE TABLE instance_events (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
instance_id UUID NOT NULL REFERENCES tool_instances(id) ON DELETE CASCADE,
event_type VARCHAR(50) NOT NULL,
status VARCHAR(50),
message TEXT,
created_by UUID REFERENCES users(id) ON DELETE SET NULL,
metadata JSONB NOT NULL DEFAULT '{}',
created_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_instance_events_instance_id ON instance_events(instance_id);
CREATE INDEX idx_instance_events_created_at ON instance_events(created_at DESC);
CREATE INDEX idx_instance_events_event_type ON instance_events(event_type);
```
**SQLAlchemy model:**
```python
# apps/api/src/models/instance_event.py
class InstanceEvent(UUIDPrimaryKeyMixin, Base):
__tablename__ = "instance_events"
instance_id: Mapped[uuid.UUID] = mapped_column(
Uuid(as_uuid=True), ForeignKey("tool_instances.id", ondelete="CASCADE"), nullable=False
)
event_type: Mapped[str] = mapped_column(String(50), nullable=False)
status: Mapped[str | None] = mapped_column(String(50), nullable=True)
message: Mapped[str | None] = mapped_column(Text, nullable=True)
created_by: Mapped[uuid.UUID | None] = mapped_column(
Uuid(as_uuid=True), ForeignKey("users.id", ondelete="SET NULL"), nullable=True
)
metadata: Mapped[dict[str, Any]] = mapped_column(JSONB, nullable=False, default=dict)
created_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), server_default=func.now(), nullable=False
)
```
### 6.2 Table: `health_checks`
```sql
CREATE TABLE health_checks (
id UUID PRIMARY KEY DEFAULT gen_random_uuid(),
instance_id UUID NOT NULL REFERENCES tool_instances(id) ON DELETE CASCADE,
container_status VARCHAR(50),
container_healthy BOOLEAN,
tunnel_healthy BOOLEAN,
exit_code INT,
probe_status VARCHAR(50),
probe_output TEXT,
checked_at TIMESTAMPTZ NOT NULL DEFAULT now()
);
CREATE INDEX idx_health_checks_instance_id ON health_checks(instance_id);
CREATE INDEX idx_health_checks_checked_at ON health_checks(checked_at DESC);
```
**SQLAlchemy model:**
```python
# apps/api/src/models/health_check.py
class HealthCheck(UUIDPrimaryKeyMixin, Base):
__tablename__ = "health_checks"
instance_id: Mapped[uuid.UUID] = mapped_column(
Uuid(as_uuid=True), ForeignKey("tool_instances.id", ondelete="CASCADE"), nullable=False
)
container_status: Mapped[str | None] = mapped_column(String(50), nullable=True)
container_healthy: Mapped[bool | None] = mapped_column(Boolean, nullable=True)
tunnel_healthy: Mapped[bool | None] = mapped_column(Boolean, nullable=True)
exit_code: Mapped[int | None] = mapped_column(Integer, nullable=True)
probe_status: Mapped[str | None] = mapped_column(String(50), nullable=True)
probe_output: Mapped[str | None] = mapped_column(Text, nullable=True)
checked_at: Mapped[datetime] = mapped_column(
DateTime(timezone=True), server_default=func.now(), nullable=False
)
```
### 6.3 Migration
**File:** `apps/api/alembic/versions/2026_05_28_add_monitoring_tables.py`
**Dependency:** Depends on the latest existing revision (e.g., `2026_05_28_add_terminal_sessions_table.py` or whichever is `head` at apply time).
**Operations:**
1. `CREATE TABLE instance_events`
2. `CREATE TABLE health_checks`
3. Create all 5 indexes.
4. No data backfill.
**Rollback:** `op.drop_index(...)`, `op.drop_table("health_checks")`, `op.drop_table("instance_events")`.
---
## 7. Error Handling Strategy
### 7.1 Docker CLI Timeout / Failure
**Where:** `HealthMonitor._check_instance()` calling `get_container_status()` or HTTP tunnel probe.
**Behavior:**
- Wrap call in `try/except Exception`.
- Log structured JSON error with `instance_id`, `correlation_id`, `error_type`, `message`.
- **Do NOT** update `tool_instances.status`.
- **Do NOT** insert `health_checks` row.
- **Do NOT** publish event.
- Continue to next instance in the poll loop.
```python
try:
status = await get_container_status(instance.container_id)
except Exception as exc:
logger.error(
"Health check failed",
extra={
"instance_id": str(instance.id),
"correlation_id": get_correlation_id(),
"error": str(exc),
},
)
return
```
### 7.2 SSE Disconnect
**Where:** `events_stream()` generator, proxy/network failure, client close.
**Behavior:**
- Detect disconnect via `asyncio.CancelledError` or `Starlette` disconnect sentinel.
- Unsubscribe from `InstanceEventBus` in `finally` block.
- **Do NOT** log error for normal disconnects (log at `INFO` level only).
- Release connection slot in per-user counter.
### 7.3 SSE Reconnect Storm
**Where:** Frontend `useEvents()` hook.
**Behavior:**
- Exponential backoff with jitter (see §3.4).
- If server returns `429`, add extra 5s penalty before retry.
- If server returns `401`, stop reconnecting and redirect to login.
### 7.4 Event Bus Subscriber Crash
**Where:** `InstanceEventBus.publish()` iterating callbacks.
**Behavior:**
- Each callback wrapped in `try/except Exception`.
- Log error with full payload and `correlation_id`.
- Continue to next subscriber.
- Publisher (`publish()` call) is never blocked by a slow/failing subscriber.
```python
for callback in self._subscribers[event_type]:
try:
if asyncio.iscoroutinefunction(callback):
await callback(payload)
else:
callback(payload)
except Exception:
logger.exception("Event subscriber failed", extra={"correlation_id": payload["correlation_id"]})
```
### 7.5 Auth Failure on SSE
**Where:** `events_stream()` before `StreamingResponse`.
**Behavior:**
- `get_current_user_id` raises `HTTPException(401)`.
- FastAPI returns `401 Unauthorized` **before** creating the stream.
- No `InstanceEventBus` subscription is created.
- No connection slot is consumed.
---
## 8. Testing Strategy
### 8.1 Unit Tests
| Test | File | What |
|------|------|------|
| EventBus publish delivers to all subscribers | `tests/unit/test_event_bus.py` | Register 3 callbacks; publish; assert all called with correct payload |
| EventBus subscriber exception isolation | `tests/unit/test_event_bus.py` | Register callback that raises; publish; assert other callbacks still called |
| EventBus unsubscribe removes callback | `tests/unit/test_event_bus.py` | Unsubscribe; publish; assert callback not called |
| HealthMonitor detects crash | `tests/unit/test_health_monitor.py` | Mock `get_container_status` to return `"exited"`, `exit_code=137`; assert DB updated to `error`, event published |
| HealthMonitor detects tunnel failure | `tests/unit/test_health_monitor.py` | Mock tunnel HEAD to 502; assert status → `unhealthy`, `health_checks` row inserted |
| HealthMonitor skip on no change | `tests/unit/test_health_monitor.py` | Two identical polls; assert only one `health_checks` row |
| HealthMonitor Docker exception resilience | `tests/unit/test_health_monitor.py` | Mock `get_container_status` to raise; assert no exception propagates, loop continues |
**Fixtures needed:**
- `event_bus`: fresh `InstanceEventBus()` instance (reset singleton state).
- `health_monitor`: `HealthMonitor(event_bus)` with mocked `POLL_INTERVAL_SECONDS = 0.1`.
- `db_session`: async SQLAlchemy session with rollback after each test.
### 8.2 Integration Tests
| Test | File | What |
|------|------|------|
| SSE endpoint requires auth | `tests/integration/test_sse_endpoint.py` | `GET /events/stream` without cookie → `401` |
| SSE endpoint streams events | `tests/integration/test_sse_endpoint.py` | Authenticated client connects; backend publishes event; client receives SSE line within 1s |
| SSE endpoint enforces connection limit | `tests/integration/test_sse_endpoint.py` | Open 6 connections; 6th returns `429` |
| SSE disconnect unsubscribes | `tests/integration/test_sse_endpoint.py` | Connect; close client; publish event; assert no error, subscriber count = 0 |
| Lifecycle hook publishes on start | `tests/integration/test_lifecycle_hooks.py` | Call start endpoint; assert `instance_events` row exists and event bus receives `instance.started` |
### 8.3 E2E Tests
| Test | File | What |
|------|------|------|
| Start container → toast appears | `tests/e2e/container_monitoring.spec.ts` (or Playwright) | Click Start; assert "Container starting..." toast; wait for probe; assert "Container running" toast |
| Container crash → error toast | `tests/e2e/container_monitoring.spec.ts` | Start container; kill container externally; assert error toast within 5s |
| Real-time badge update | `tests/e2e/container_monitoring.spec.ts` | Start container; badge green; kill container; badge turns red without refresh |
### 8.4 Frontend Unit Tests
| Test | File | What |
|------|------|------|
| useEvents reconnect backoff | `apps/web/src/hooks/use-events.test.ts` | Simulate `EventSource` error; assert reconnect delay doubles up to cap |
| Toast deduplication | `apps/web/src/components/toast-rules.test.ts` | Two identical events within 1s; assert only one toast shown |
| Event-to-toast mapping | `apps/web/src/components/toast-rules.test.ts` | Map each event type to correct toast type, message, duration |
---
## 9. Performance Considerations
### 9.1 SSE Connection Pool
- **Limit:** 5 concurrent SSE connections per user ID.
- **Reasoning:** Prevents tab-spam from exhausting server memory. A typical user has 1–3 tabs open.
- **Implementation:** In-memory `dict[uuid.UUID, int]` in `events.py`. In-memory is acceptable because single-process API is assumed.
### 9.2 Health Monitor Batching
- **Current approach:** `docker inspect` is called once per instance per poll cycle.
- **Optimization (future):** Batch `docker ps --format json` to get all container statuses in a single CLI invocation, then match by `container_name`. **Not implemented in MVP** to keep changes minimal; document as follow-up.
- **DB writes:** Only on state change. The monitor compares against `_last_known_state` in memory before touching the DB.
### 9.3 Event Bus Memory Profile
- **No event history:** The bus holds only subscriber callable references (lightweight).
- **No queues:** SSE connections use per-connection `asyncio.Queue` capped at 100 items; if a client is slow, drop oldest events to prevent unbounded growth.
```python
queue: asyncio.Queue[InstanceEventPayload] = asyncio.Queue(maxsize=100)
```
### 9.4 Database Write Amplification
- **Health checks:** Written only on state change, not every 15-second poll.
- **Growth estimate:** 100 instances × 10 state changes/day × 365 days ≈ 365k rows/year. Acceptable for PostgreSQL.
- **Retention (follow-up):** Add a scheduled cleanup job or pg_partman for `health_checks` older than 30 days.
### 9.5 Frontend Polling Reduction
- **Before:** Health poll every 30s per running instance = 2 req/min/instance.
- **After:** One SSE connection per browser tab, zero polling for status. Fallback list refresh every 60s retained for resilience.
- **Server load reduction:** For 50 running instances across all users, eliminates ~100 health-check HTTP requests per minute.
### 9.6 JSON Logging Overhead
- JSON formatter adds ~20% CPU overhead vs plain text for high-volume logs. Mitigate by:
- Keeping `uvicorn.access` at `WARNING`.
- Not logging every SSE ping.
- Using `orjson` for JSON serialization if available (fallback to stdlib `json`).
---
## 10. Rollout Plan
| PR | Contents | Estimated Lines | Review Risk |
|----|----------|-----------------|-------------|
| **PR 1: Backend core** | DB migrations, models, `InstanceEventBus`, `HealthMonitor`, SSE endpoint, correlation ID middleware, JSON logging | ~1,000 | Medium |
| **PR 2: Frontend** | `useEvents` hook, `ToastProvider`, `sonner` integration, badge real-time updates, remove 30s health polling | ~700 | Medium |
| **PR 3: Integration + tests** | Lifecycle hook instrumentation in `tool_instances.py`, unit + integration tests, E2E tests | ~400 | Low |
**Dependency order:** PR 1 → PR 2 → PR 3. PR 2 can be developed in parallel but must merge after PR 1.
---
## 11. Open Questions / Decisions
| ID | Decision | Status |
|----|----------|--------|
| D1 | Use `sonner` for toasts (vs custom implementation) | **Decided:** `sonner` — reduces custom UI code by ~300 lines |
| D2 | In-memory event bus (vs Redis/NATS) | **Decided:** In-memory — matches `TerminalManager` pattern; defer distributed bus |
| D3 | SSE instead of WebSocket | **Decided:** SSE — one-way push, simpler auth, HTTP-compatible |
| D4 | Batch `docker ps` for health monitor | **Deferred:** Keep per-instance `docker inspect` for MVP; document optimization |
| D5 | `health_checks` retention policy | **Deferred:** 30-day retention to be added in follow-up |