- Add clone_mode and branch fields to tool_instances - Add ssh_key_id to git_repositories for per-repo SSH key assignment - Implement host-side git cloning with branch selection (default: main) - Mount SSH keys into containers for git operations in clone mode - Add dirty state check on clone-mode instance deletion with confirmation - Update SessionsPage with mount/clone selector, branch input, SSH key display - Add SSH key selector to repository creation form - Add dirty delete confirmation modal with changed files list - Update API schemas and endpoints for new fields - Sync delta specs to main specs (git-repo, tool-instances, repo-clone-mode) - Archive completed OpenSpec change: repo-clone-mode-with-ssh - Document git requirement for custom tool types Quality gates: Frontend typecheck and build passed OpenSpec: repo-clone-mode-with-ssh archived with all tasks complete
4.6 KiB
Context
The current instance management has critical gaps in health monitoring that lead to poor user experience:
-
Silent startup failures: When
docker compose upexecutes, the API immediately marks the instance as "running" without verifying the container actually reached a healthy state. Containers that crash on startup or fail to bind to their port appear "running" in the UI but serve 502 errors. -
Tunnel-only health checks: The existing health check at
GET /instances/{id}/healthonly performs an HTTP HEAD request to the tunnel URL. This cannot distinguish between:- Tunnel is broken (cloudflared process died) → should recreate tunnel
- Tool crashed inside container → should show container error
- Tool returns 502 because it's still starting → should wait for readiness probe
-
Unused readiness probes: The
readiness_probe.pyservice was built during the tool-workshop change but is never called during instance startup. Tool types can configure readiness probes (e.g.,curl -f http://localhost:8080/health) but these are ignored. -
Blind auto-recovery: The frontend shows a "Recreate Tunnel" button when the health check fails, but this recreates the tunnel even when the application itself is returning 502 errors, wasting time and confusing users.
Goals / Non-Goals
Goals:
- Verify containers actually start successfully before marking instances as "running"
- Distinguish container health from tunnel health in monitoring
- Integrate readiness probes into the instance startup flow
- Only recreate tunnels when the tunnel itself is broken, not when the tool returns errors
- Provide clear error messages when instances fail to start
Non-Goals:
- Persistent tunnels (keeping temporary cloudflared tunnels)
- Automatic restart of crashed containers (Docker already does this with restart policies)
- Health check WebSocket push (polling is sufficient)
- Changing the Docker compose architecture
Decisions
1. Startup verification via Docker API
- After
docker compose up, polldocker psfor 30 seconds to verify container state transitions to "running" - If container exits or stays in "restarting" loop, mark instance as "error" with exit code
- Rationale: Direct Docker API check is more reliable than HTTP checks during startup when ports may not be bound yet
2. Readiness probe as gate to "running" status
- Instance status flow:
pending→starting(container up) →running(probe passed) - If probe fails after timeout, status becomes
unhealthy(noterror- container is still up) - Rationale: Distinguishes "container won't start" from "container started but app isn't ready yet"
3. Container + Tunnel dual health checks
- Health endpoint returns both
container_status(from Docker API) andtunnel_status(HTTP check) - Frontend shows different badges: "container unhealthy" vs "tunnel error"
- Rationale: Users need to know if they should wait (app starting) or recreate tunnel
4. Smart tunnel failure detection
- Connection errors (ECONNREFUSED, ETIMEDOUT, DNS failure) → tunnel is broken → allow recreate
- HTTP 502/503/504 → application error → show "app error" badge, don't recreate
- HTTP 200-399 → healthy
- Rationale: 502 from the tool means the tunnel is working fine, the tool just isn't responding
5. Readiness probe configuration from ToolType
- Use existing
readiness_probeJSON field on ToolType model - Default probe for web tools:
curl -f http://localhost:{port} - Default probe for terminal tools: none (skip probe, mark running immediately)
- Rationale: Leverages existing infrastructure, provides sensible defaults
Risks / Trade-offs
[Risk] Startup polling adds latency → Mitigation: Poll every 2 seconds with 30 second max timeout. Most containers start in <5 seconds.
[Risk] Docker API calls from API container → Mitigation: API container already has Docker CLI access for managing instances. Using docker ps is consistent with existing patterns.
[Risk] False "unhealthy" from slow-starting tools → Mitigation: 30 second default timeout with configurable override per tool type. Frontend shows "starting..." status during probe.
[Risk] Probe commands may not exist in container → Mitigation: Probe failures log stderr. If probe command missing, container still starts but marked as running without probe validation.
Migration Plan
No database migration needed. This change:
- Adds new status values ("starting", "unhealthy") to existing
statusenum - Uses existing
readiness_probecolumn ontool_typestable - Changes health check API response format (adds fields, doesn't remove)
Open Questions
None.