Files
Claude-Code-Monitor/docs/superpowers/specs/2026-08-03-lane-runtime-isolation-design.md
T
nntrivi2001 9d145865dd feat(lanes): per-lane database, Redis, and .env isolation (A1+A2)
Gives each lane its own slot-derived runtime (ports, detached process
lifecycle, profile-driven hooks) and its own database/Redis logical
index/.env file, so two lanes running the same repo's stack at once no
longer share state. Machine-level DB/Redis credentials live at
~/.ccam/secrets.env (mode 0600, never returned by any route); a hook's
output is redacted of that password (raw and URL-encoded forms) before
it reaches a log file or the lane_hook_output websocket broadcast.
Wired into provision/up/reset/remove; reset accepts --keep-db to skip
the drop/recreate/migrate/reseed block entirely.
2026-08-04 10:03:40 +07:00

10 KiB
Raw Blame History

A1 — Lane runtime isolation: slots, ports, profile hooks, detached lifecycle

Status: approved 2026-08-03. Roadmap it belongs to: docs/superpowers/plans/2026-08-03-shipyard-parity-lanes.md.

Problem

A lane isolates code and nothing else. server/lib/worktree.js gives each lane its own git worktree and branch, so two agents never overwrite each other's files — but they share every other resource: the same ports, the same database, the same .env, the same upload directory. Booting two lanes' stacks at once does not work.

Shipyard (AgentWorkflow/) solved this with per-lane ports, database, Redis logical index and .env. That is not containerization — every lane runs on the host as the same user. Its isolation is resource namespacing at the application layer.

Scope

Shipyard parity is six independent subsystems (see the roadmap). A — runtime isolation — goes first because it blocks the rest, and splits into three:

  • A1 (this spec) — a lane can boot its stack on isolated ports and directories, and the stack survives a dashboard restart. Profiles are hand-written.
  • A2 — data isolation: .env seeding, database create/drop/migrate/seed, Redis index.
  • A3 — stack detection + ccam lane profile init scaffolding, designed from the real profiles written during A1/A2.

The split point matters. Shipyard was genericized but its profiles/ is still empty (PROFILE=_template, SOURCE_REPO="", state/ empty) — a clean seam that never ran against a real app. Designing presets before a real profile exists repeats exactly that.

A1 done means: ccam lane up boots a lane's stack on ports no other lane uses, and ccam update (or any dashboard restart) leaves running stacks alone.

Decisions

Choice Rationale
Profile location <repo>/.ccam/profile/ CCAM is multi-repo (lanes keyed by cwd); one global PROFILE is wrong the moment two repos have lanes. A profile that lives in the repo is already present in a fresh worktree.
Engine Node, server/lib/ execFile with argument arrays (repo rule: never build a git/shell command as a string). Shares the SQLite DB and the websocket. Hooks stay shell scripts; the engine spawns them.
Resource model Declared per repo PORTS="api fe worker", LANE_DIRS="…". An app with no database declares none, so A2 generates no database hooks — no dead code.
Process ownership Fully detached pid files outside the worktree; a dashboard restart does not touch a running stack.
Port allocation Base + slot, step aside on collision, record the real number Predictable in the common case, does not hard-fail when a port is busy.
Adopted lanes May up/down; every write path refused up only runs what the user can already run themselves. Writes into their cwd (A2) are refused outright.
Runtime state DB owns allocation, disk owns handles, up/down computed on read A process dies to OOM, kill, or a reboot — caching a truth you do not control buys ghost state. Slot allocation is the opposite: fully controlled, needs to be race-free, so it belongs in the DB.

Invariants

  1. The runtime never writes stage, status, or notes — only slot and ports. In CCAM status=running means an agent is working, not a stack is up; conflating the two breaks classifyLiveness. Boot failures surface through GET /runtime and the hook log. Same reasoning as "the console never writes a lane's stage".
  2. CCAM still does not orchestrate. A1 adds primitives (up/down/hook) that a session calls. The dashboard chains nothing, evaluates no gate, retries nothing.
  3. No shell command built as a string. execFile + argv arrays everywhere. Hooks are the deliberate exception — they are scripts, spawned as bash <hook> <args…> with a fixed argv, never interpolated.
  4. profile.env is parsed, never sourced. Sourcing arbitrary shell from a repo into the dashboard process is a code-execution path. Hooks are executed on purpose; config is only read.
  5. Adopted lanes stay non-destroyable (assertManaged in worktree.js). A1 adds no bypass.

Design

Schema (server/db.js)

ALTER TABLE lanes ADD COLUMN slot INTEGER;
ALTER TABLE lanes ADD COLUMN ports TEXT NOT NULL DEFAULT '{}';
CREATE UNIQUE INDEX IF NOT EXISTS idx_lanes_slot ON lanes(slot) WHERE slot IS NOT NULL;

Added with the existing probe-each-column-independently migration pattern. ports is parsed by the same hydrate() branch that already handles stages/links, corrupt-blob fallback included.

slot and ports join PROVISIONING_FIELDS and stay out of PATCHABLE, for the reason kind is excluded: a client that can set slot can change the slot-derived DB_NAME A2 will use — which aims a drop at another database.

No column records "is up". Deliberate.

Modules

File Responsibility
server/lib/ports.js isListening(port) via a short-timeout net.connect; listenerPids(port) via lsof, falling back to ss, returning [] with a one-time warning when neither exists.
server/lib/lane-slots.js allocateSlot() — lowest free in [1, LANE_MAX_SLOTS] (default 9), called inside the existing withLaneLock, with the unique index as the backstop. releaseSlot() on remove only — not on reset, or a lane's ports and database move under a running session. resolvePorts(slot, profile) — prefers PORT_BASE_<name> + slot, then +100, +200… (the last digit stays the slot, so the number still reads as "lane 3"), up to 10 tries before EPORTBUSY naming the occupying pid.
server/lib/lane-profile.js resolveProfile(lane) — looks in lane.cwd/.ccam/profile/ first, lane.source_repo/.ccam/profile/ second, because the profile lives in the repo and a branch that edits its hooks must run its own. runHook(lane, name, args, {onLine}) — fixed hook-name allowlist, bash <hook> with an argv array, Shipyard's env contract verbatim so its profiles port unchanged (LANE = the slot, LANE_DIR, RUN_DIR, LOG_DIR, PROFILE_DIR, SOURCE_REPO, <NAME>_PORT per declared port), harness_spawn exported as a bash function writing the same $RUN_DIR/<name>.pid + $LOG_DIR/<name>.log, the GIT_* scrub from worktree.js:git(), and output streamed both to the log file and to a new lane_hook_output websocket message.
server/lib/lane-runtime.js upLane (defensive down → mkdir run/log/LANE_DIRSboothealth; no bootstrap, which belongs to provision/reset as in lane-up.sh), downLane (recursive pid-tree kill from the pid files — ported from lane-down.sh:kill_tree, which exists because detached uvicorn/celery prefork children were being missed — then a port-listener backstop; idempotent, sleep-free), runtimeFacts (computed: kill -0 per pid file, isListening per port, plus last-error.json).

State lives at LANES_ROOT/.state/lane<slot>/{run,logs}/, outside the worktree: lane reset runs git clean -fd, which would sweep pid files mid-flight.

Adopt-after-restart costs no code. Pid files are on disk, detached processes keep running, and runtimeFacts recomputes on read. That is the payoff of making up/down computed rather than stored.

Routes (server/routes/lanes.js)

POST /api/lanes/:id/up          { build? }   202 → background → broadcastLane
POST /api/lanes/:id/down        { mcp? }     200
GET  /api/lanes/:id/runtime                  200
POST /api/lanes/:id/hook/:name  { args? }    202
GET  /api/lanes/:id/logs/:svc   ?tail=N      200

All behind sameOriginGuard and the existing withLaneLock. GET /runtime follows the contract of GET /:id/git: a lane with no profile returns {available:false} with HTTP 200 — a normal state, not a fault. It probes ports and stats pid files, which is why it is its own endpoint rather than part of the polled GET /api/lanes.

:name and :svc come from allowlists, never from the request; log paths are realpath-confined to LOG_DIR; args is an array of strings passed straight to argv, never joined.

Errors

Condition Code Behaviour
No .ccam/profile/ ENOPROFILE 400, naming both paths searched
Unknown hook ENOHOOK 400
Hook exits non-zero lane_hook_output + last-error.json; reflected by /runtime
Every candidate port busy EPORTBUSY names the occupying pid
health fails EUNHEALTHY leaves the processes running — killing them destroys the evidence — and points at the log
Slot pool exhausted ESLOTS 409
Name outside an allowlist 400, nothing spawned

Verification

npm run test:server (required) and npm run test:client (review the snapshot diff before regenerating). A fixture repo with a three-line profile: boot.sh starts python3 -m http.server $FE_PORT through harness_spawn, health.sh curls it with retries.

Covered:

  • Real lifecycleup/runtime healthy with the port listening → simulate a dashboard restart (rebuild the module) → still healthy, not rebooted → down → port free, pid files gone.
  • Slots — nine allocations yield 1..9; releasing 4 makes the next allocation 4; concurrent allocations through withLaneLock never collide.
  • Port step-aside — occupy the preferred port, assert +100 is chosen and recorded in ports.
  • Profile parsing — missing keys fall back to defaults; no profile returns null without throwing; a value containing $(id) stays literal and is not executed.
  • GuardsPATCH /api/lanes/:id ignores slot/ports; a hook name outside the allowlist returns 400 and spawns nothing.
  • Env — with GIT_DIR set on the parent, the hook does not see it.

Manual: create a lane on ccam-lanes itself, write a three-line profile, ccam lane up, open the port, restart the server, confirm the stack survived.

Out of scope

.env seeding and databases (A2). Detection and scaffolding (A3). Per-feature state, proof gallery, cross-lane locks, the ship-feature skill, integrations (subsystems BF).

Container isolation is a separate design, not a later phase: hooks and MCP servers would have to run inside the container, so the cwd a hook reports becomes a container path and no longer matches the lane's host cwd — breaking resolveLaneByCwd, which is what binds a session to a lane at all.