Complete rewrite of the terminal pipeline for VS Code Server-level responsiveness. Key improvements: Backend: - Replace blocking select.select(0.1) with asyncio.add_reader() for event-driven PTY reading (eliminates ~110ms polling latency) - Add output batching (2ms window) to reduce WebSocket frame overhead - Add flow control: client acks processed bytes, server pauses PTY reads at 64KB threshold, resumes at 32KB - Add 5s ack timeout fallback to prevent stuck sessions Frontend: - Switch WebSocket to binary mode (binaryType = 'arraybuffer') - Eliminate Blob -> arrayBuffer async conversion overhead - Add flow control ack messages (every 4096 bytes or 100ms) - Add xterm-addon-webgl with graceful DOM fallback - Add performance tuning (scrollback=10000, fastScrollSensitivity) SDD artifacts: - openspec/explorations/terminal-responsiveness.md - openspec/proposals/terminal-responsiveness.md - openspec/specs/terminal-responsiveness.md - openspec/designs/terminal-responsiveness.md - openspec/tasks/terminal-responsiveness.md Quality gates: pytest (19 passed, 1 skipped), tsc --noEmit clean
6.5 KiB
SDD Exploration: Responsive Web Terminal
Status
Phase: explore
Date: 2026-06-02
Owner: el Gentleman (parent session)
Scope: Terminal I/O latency, rendering performance, connection stability
Goal
Achieve VS Code Server-level terminal responsiveness: near-local latency on keystrokes, smooth scrolling, no jank on output bursts, and instant resize reactions.
Current Architecture
Data Flow
Container shell → docker exec PTY → host PTY master fd → select.select(0.1s)
→ Python read loop (10ms sleep fallback) → WebSocket.send_bytes()
→ WebSocket (Blob mode) → frontend arrayBuffer decode → xterm.js.write()
Key Files
| File | Responsibility |
|---|---|
apps/web/src/components/terminal.tsx |
xterm.js, WebSocket client, FitAddon |
apps/api/src/api/terminal.py |
WebSocket endpoint, auth, read/write/heartbeat loops |
apps/api/src/services/terminal_session.py |
PTY creation, docker exec subprocess, I/O |
apps/api/src/services/terminal_manager.py |
Session lifecycle, persistence, idle cleanup |
Current Bottlenecks
1. Blocking Read with 100ms Timeout
# terminal_session.py:read_output()
readable, _, _ = select.select([self._master_fd], [], [], 0.1)
if readable:
data = os.read(self._master_fd, 4096)
Problem: select.select blocks up to 100ms when no data is available. With the read loop in terminal.py doing asyncio.sleep(0.01) between calls, worst-case latency from shell output to WebSocket is ~110ms.
VS Code approach: node-pty uses libuv's epoll/kqueue watchers — event-driven, no polling timeout.
2. WebSocket Blob → arrayBuffer Conversion
// terminal.tsx
ws.onmessage = (event) => {
if (event.data instanceof Blob) {
event.data.arrayBuffer().then((buffer) => {
const data = new Uint8Array(buffer);
termRef.current?.write(data);
});
}
}
Problem: Blob → arrayBuffer is async and adds microtask latency. Also forces GC pressure from transient Blobs.
VS Code approach: Uses ws.binaryType = "arraybuffer" — receives ArrayBuffer directly, zero-copy into Uint8Array.
3. No asyncio-Native PTY Reading
The PTY master fd is read with synchronous os.read() inside an async coroutine. This blocks the event loop thread for the duration of the read.
VS Code approach: node-pty's C++ binding hooks into libuv's event loop natively — true async I/O.
4. Docker Exec Subprocess Overhead
# terminal_session.py:start()
self.process = await asyncio.create_subprocess_exec(
"docker", "exec", "-it", "-e", "TERM=xterm",
self.container_id, "bash", "-c", shell_cmd,
stdin=self._slave_fd, stdout=self._slave_fd, stderr=self._slave_fd,
)
Problem: Spawns a new docker exec process on the host. Adds process startup latency and an extra process hop.
Alternative: Docker Engine API's attach endpoint with logs=0&stream=1&stdin=1&stdout=1&stderr=1 — streams directly to the API container via Unix socket. No host subprocess.
5. No Flow Control / Backpressure
If a command dumps output faster than the WebSocket can send (e.g., cat /dev/urandom | base64), data piles up in:
- The PTY kernel buffer (limited, ~4KB)
- Python's deque circular buffer (10KB)
- WebSocket's internal buffer (unbounded in some implementations)
- xterm.js parser queue
VS Code approach: Implements explicit flow control — pauses the PTY when the client buffer exceeds a threshold, resumes when drained.
6. xterm.js Renderer
Current: DOM renderer (default).
VS Code approach: Canvas renderer with WebGL addon for GPU-accelerated rendering.
Measurement Baseline
Before optimization, we need metrics:
| Metric | How to Measure | Target |
|---|---|---|
| Input latency | Time from keypress to character appearing | < 16ms (1 frame) |
| Output throughput | Bytes/sec for cat /dev/zero |
> 1 MB/s |
| Resize latency | Time from resize message to shell reacting | < 50ms |
| Reconnection time | Time from disconnect to full replay | < 200ms |
| Frame drops | Dropped frames during yes command |
0 |
Improvement Directions
Direction A: Low-Latency Read Loop (Quick Win)
Replace select.select + os.read with asyncio native approach:
- Use
loop.add_reader()to register a callback when fd is readable - Or use
asyncio.to_thread()with blockingos.readand immediate wake - Eliminate the 100ms timeout and 10ms sleep
Direction B: WebSocket Binary Mode (Quick Win)
Set ws.binaryType = "arraybuffer" on frontend, send binary frames directly.
Eliminates Blob → arrayBuffer conversion.
Direction C: Docker Engine API Attach (Medium)
Replace docker exec subprocess with direct container attach via Docker SDK or HTTP API:
from docker import DockerClient
client = DockerClient()
container = client.containers.get(container_id)
socket = container.attach_socket(params={...})
# socket is a raw TCP/Unix socket — read with asyncio
Pros: No subprocess overhead, direct stream to container
Cons: Requires Docker SDK or raw HTTP over Unix socket; needs docker group permissions
Direction D: Flow Control (Medium)
Add backpressure mechanism:
- Measure WebSocket send buffer depth on backend
- Pause reading from PTY when buffer exceeds threshold (e.g., 64KB)
- Resume when buffer drains below threshold
- Frontend: measure xterm.js parser queue depth, pause via control message
Direction E: WebGL Renderer (Quick Win)
Add xterm-addon-webgl:
import { WebglAddon } from 'xterm-addon-webgl';
term.loadAddon(new WebglAddon());
Pros: GPU-accelerated, much faster for large output bursts
Cons: Falls back to canvas/DOM if WebGL unavailable; slightly higher init time
Direction F: Output Batching (Quick Win)
Batch small writes before sending over WebSocket:
- Collect output for 1-2ms
- Send as single binary frame
- Reduces WebSocket frame overhead for high-frequency small writes (e.g., progress bars)
Recommended Next Steps
- Measure baseline with synthetic benchmarks
- Implement Directions A + B + F (low-risk, high-impact)
- Evaluate Direction C (Docker API attach) vs keeping docker exec
- Add Direction D (flow control) if throughput tests show issues
- Add Direction E (WebGL) as frontend enhancement
Risks
- Docker API attach may not support PTY mode as cleanly as
docker exec -it - WebGL addon may have compatibility issues on older GPUs
- Flow control adds complexity; premature optimization risk
- Changes to core I/O loop could introduce stability regressions