#2308 · Idle provider bridge processes are never retired for non-Codex providers
Verdict: REPRODUCED · Root-cause confidence: high
Scope note. The issue was written against 8bd6acc9. Between the report and the base commit, #2325 (e42a4ef48, merged 2026-08-24) deleted the Codex thread-scoped process special case. On 494f66526 the Codex carve-out the issue describes no longer exists: every provider's bridge, Codex included, stays resident with zero threads until the daemon exits. The title's "non-Codex" qualifier is stale; the leak is now universal.
1. TL;DR
Every environment that has ever run an agent thread keeps one bridge-worker node process per provider alive for the life of the host daemon, even after the thread is stopped and the agent child has exited. A user with 37 environments sees ~1 GB of node processes that serve nothing. The runtime spawns one bridge per (environment, provider artifact) and has exactly one retirement path on thread release, retireSupersededBridgeProcessIfIdle, which only fires when the bridge's plugin artifact has been replaced by a newer one. A current-artifact bridge with zero threads matches no retirement path, and nothing else (no idle timer, no periodic sweep) ever revisits it. The behaviour is deliberate "keep the process warm for the next thread" design — three existing tests assert it — but the design has no upper bound on warmth, so the process count grows monotonically with environments touched. I reproduced it live on a dev instance (bridge pid 20879 alive with 0 children for minutes after each bb thread stop, across two spawn/stop cycles, 9½ minutes in total) and with a vitest that fails on the base commit for both claude-code and codex.
2. Claims vs findings
| Claim from the issue | Status | Evidence |
|---|---|---|
| When the last thread on a bridge is released, the bridge is shut down only if it is a thread-scoped Codex process. | Stale (was true at 8bd6acc9; not at base) | e42a4ef48 (#2325) removed isThreadScopedCodexProcess and the Codex branch of releaseIdleProviderProcess. At base the function is a one-liner that calls retireSupersededBridgeProcessIfIdle for every provider (packages/agent-runtime/src/runtime.ts:485-489). Codex bridges now leak the same way — my unit repro fails for codex too. |
| Claude Code and ACP bridges stay resident with zero threads until daemon exit. | Verified | Live: bridge 20879 (claude-code) alive with 0 children at +3 s and +95 s after the first bb thread stop, still alive 3 min 45 s later when a second thread reused it, and after the second stop a 10-second watch shows it with 0 children for 190 s (section 4 steps 5, 9, 10). No timer or sweep exists that could ever retire it (section 5). |
| Non-Codex bridges are keyed per environment, so the retained set grows with environment count. | Verified | Process key is providerId#bridge:<artifact-hash> (packages/agent-runtime/src/runtime.ts:410-418) inside a runtime entry that is keyed by environment: this.entries.set(args.environmentId, entry) (apps/host-daemon/src/runtime-manager.ts:1033) / get(environmentId) (apps/host-daemon/src/runtime-manager.ts:366-368); the comment at apps/host-daemon/src/runtime-manager.ts:358-360 states the intent ("Provider runtimes are environment-scoped and may host multiple threads"). Second thread in the same environment reused bridge 20879 instead of spawning (section 4 step 6). |
| One spawn/stop cycle in an environment with no loaded runtime permanently adds one bridge process. | Verified | Section 4: 0 bridges before (step 2), 1 stranded bridge after spawn+stop, still present across a second spawn/archive/stop cycle and 3 min after the second stop (step 10). |
| ~78 MB per stranded bridge; 12 stranded bridges = 960 MB on the reporter's machine. | Unverified (plausible) | I cannot see the reporter's machine. On my dev instance the bridge runs via tsx from source and sat at 190–256 MB RSS while idle, dropping to 96 MB after a GC about 90 s after the second stop (watch); a packaged host.mjs bridge will be smaller. The per-environment multiplication is what matters and is verified. |
retireSupersededBridgeProcessIfIdle returns early unless the artifact hash is superseded. | Verified | packages/agent-runtime/src/runtime-provider-process.ts:295-310: if (currentKey === undefined || currentKey === providerProcess.processKey) return; |
| All callers route through that guard (archive, staged rewind, failed construction, thread/stop user + reaper). | Verified | At base: L506 (abandonFailedSessionConstruction), L1319 (archive/unarchive), L1610 (staged rewind discard), L2325 and L2335 (stopThread, which the reaper calls at L2632). All five go to the same one-line function. |
bb thread archive is not equivalent: the agent process keeps running, so the bridge keeps its thread. | Verified | Section 4 step 7: after bb thread archive thr_jxijqce242 the claude child 29995 stayed alive under bridge 20879. Cause: the claude-code plugin declares supportsThreadArchive: false (plugins/provider-claude-code/server.ts:70), so the server never forwards archive to the runtime (apps/server/src/services/threads/thread-commands.ts:170-179). |
| The #1604 session reaper is enabled and working; this is the bridge the reaper leaves behind. | Partially verified | The reaper path does end in the same no-op: reapIdleProviderSessions → runtime.stopThread (packages/agent-runtime/src/runtime.ts:2593-2650) → releaseIdleProviderProcess, and the test at packages/agent-runtime/src/runtime.process-lifecycle.test.ts:1044 explicitly asserts the bridge survives a reap ("resumes it later on the same process"). But it is not enabled for non-Codex providers by default. findReapableIdleProviderSession (packages/agent-runtime/src/runtime.ts:1008-1018) only considers a claude-code/ACP session when providerSessionReapingEnabled is true; without the experiment it reaps Codex sessions only (runtimeConfig.providerId !== CODEX_PROVIDER_ID → return null). That flag is the providerSessionReaping experiment, default false (packages/domain/src/experiments.ts:34), read by the server at apps/server/src/internal/session.ts:43 and fetched by the daemon sweep via getRuntimePolicy().providerSessionReaping (apps/host-daemon/src/app.ts:700-704). With default settings a stopped-via-archive claude-code session is therefore never reaped; only an explicit bb thread stop (or daemon exit) releases it. This does not change the root cause (stop already strands the bridge), but the "reaper leaves the bridge behind" framing only applies with the experiment on. |
Proposed fix: drop the Codex guard so zero threads always shuts the bridge down; retireSupersededBridgeProcessIfIdle becomes redundant. | Refuted as written | There is no guard left to drop. Making releaseIdleProviderProcess unconditionally shut down a zero-thread process breaks 8 existing tests that assert warm reuse (section 6, log), including the reaper's resume-on-same-process contract. The correct shape is a bounded idle retirement (timer or sweep), which the issue itself offers as an alternative. |
| Respawn cost 0.66–1.22 s for the 2.5 MB host.js import. | Unverified | Not measured; irrelevant to a grace-timer fix. |
3. Environment
- bb
494f66526913557ab076e048218236f0a6610927(branch main;packages/bb-app0.39.0).git fetch origin main:origin/mainis the same commit — nothing newer to check. - macOS 26.5.2 (25F84) arm64, Node v22.23.1, pnpm from PATH.
- Providers: claude-code (Claude Code 2.1.241 at
~/.local/bin/claude), codex-cli 0.149.1; others installed but unused. - Dev instance (own worktree,
scripts/bb-dev-app current): App :14861, Server :22861, Host daemon :30861 (pid 17027), data dir~/.bb-dev/bb-machines-HOST.getbb.app-checkouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd. Stopped and deleted after the investigation. (The first version of this report used a sibling worktree on ports 19273/27273; every live artifact below was re-captured on this instance.) - Scratch repo
/tmp/bb-2308-qa, projectproj_iwzvzf2fev, environmentenv_uhhfyvdjc9, hosthost_shpa39kdsp. Threadsthr_7ptec8hm8d(first cycle) andthr_jxijqce242(second cycle).
4. Minimal reproduction
4a. Unit test (fails on base, no providers needed)
Drop runtime.idle-bridge-retirement.test.ts into packages/agent-runtime/src/ and run it from packages/agent-runtime. It uses the repo's scripted-echo bridge, so it exercises the real RuntimeProviderProcessManager and real child processes.
cd packages/agent-runtime pnpm exec vitest run src/runtime.idle-bridge-retirement.test.ts
// Repro for get-bb/bb#2308: the provider bridge process is never retired when
// its last thread is released. Place this file at
// packages/agent-runtime/src/runtime.idle-bridge-retirement.test.ts
// and run it from packages/agent-runtime with
// pnpm exec vitest run src/runtime.idle-bridge-retirement.test.ts
//
// On 494f66526 the "retires" cases FAIL (the bridge stays resident with zero
// threads); the "keeps" case passes. A fix that shuts the bridge down when
// threadIds reaches 0 makes all three pass.
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import type { ThreadEvent } from "@bb/domain";
import {
createScriptedEchoProcessLog,
createScriptedEchoRuntime,
fullRuntimeOptions,
waitForRuntimeState,
} from "./test/runtime-test-harness.js";
describe("#2308 idle bridge retirement", () => {
let tmpDir: string;
beforeEach(() => {
tmpDir = mkdtempSync(join(tmpdir(), "bb-2308-"));
});
afterEach(() => {
rmSync(tmpDir, { recursive: true, force: true });
});
function createRuntime(args: { pluginId: string }) {
const events: ThreadEvent[] = [];
const processLog = createScriptedEchoProcessLog();
const runtime = createScriptedEchoRuntime({
runtime: {
workspacePath: tmpDir,
env: processLog.env,
onEvent: (event) => events.push(event),
},
launch: { pluginId: args.pluginId, scripted: { identifyProcess: true } },
});
const spawned = () =>
processLog.read().filter((l) => l.startsWith("spawn:")).length;
const exited = () =>
processLog.read().filter((l) => l.startsWith("exit:")).length;
return { events, runtime, spawned, exited };
}
for (const [pluginId, providerId] of [
["provider-claude-code", "claude-code"],
["provider-codex", "codex"],
] as const) {
it(`retires the ${providerId} bridge when thread/stop releases its last thread`, async () => {
const { runtime, spawned, exited } = createRuntime({ pluginId });
try {
await runtime.startThread({
environmentId: "env-1",
threadId: "t1",
projectId: "p1",
providerId,
options: fullRuntimeOptions,
});
expect(spawned()).toBe(1);
expect(runtime.listRunningProviders()).toEqual([providerId]);
// The user stops the thread (`bb thread stop`), or the idle reaper
// does (reapIdleProviderSessions -> stopThread). Either way the bridge
// now owns zero threads.
await runtime.stopThread({ threadId: "t1" });
expect(runtime.hasThread("t1")).toBe(false);
// BUG: the bridge is still resident with nothing to serve.
await waitForRuntimeState({
label: "the idle bridge process exited",
predicate: () =>
exited() === 1 && runtime.listRunningProviders().length === 0,
timeoutMs: 3_000,
});
} finally {
await runtime.shutdown();
}
});
}
it("keeps the bridge while another thread still runs on it", async () => {
const { runtime, spawned, exited } = createRuntime({
pluginId: "provider-claude-code",
});
try {
for (const threadId of ["t1", "t2"]) {
await runtime.startThread({
environmentId: "env-1",
threadId,
projectId: "p1",
providerId: "claude-code",
options: fullRuntimeOptions,
});
}
expect(spawned()).toBe(1);
await runtime.stopThread({ threadId: "t1" });
expect(exited()).toBe(0);
expect(runtime.listRunningProviders()).toEqual(["claude-code"]);
} finally {
await runtime.shutdown();
}
});
});
Output on 494f66526 (full log). The assertion that fails is the waitForRuntimeState at line 73: the bridge never logs exit: and listRunningProviders() still reports the provider after its only thread was stopped. The third case ("keeps the bridge while another thread still runs") passes, so the failure is specifically "zero threads, still resident".
RUN v4.1.1 /Users/USER/.bb-machines/HOST.getbb.app/checkouts/bb/.claude/worktrees/wf_846839f8-f8a-21/packages/agent-runtime
❯ @bb/agent-runtime src/runtime.idle-bridge-retirement.test.ts (3 tests | 2 failed) 7107ms
× retires the claude-code bridge when thread/stop releases its last thread 3655ms
× retires the codex bridge when thread/stop releases its last thread 3243ms
⎯⎯⎯⎯⎯⎯⎯ Failed Tests 2 ⎯⎯⎯⎯⎯⎯⎯
FAIL @bb/agent-runtime src/runtime.idle-bridge-retirement.test.ts > #2308 idle bridge retirement > retires the claude-code bridge when thread/stop releases its last thread
FAIL @bb/agent-runtime src/runtime.idle-bridge-retirement.test.ts > #2308 idle bridge retirement > retires the codex bridge when thread/stop releases its last thread
Error: Timed out after 3000ms waiting for the idle bridge process exited
❯ waitForRuntimeConditionUnsafe src/test/runtime-wait-helpers.ts:187:9
185| const failureDetail = options.describeFailure?.();
186| const detail = failureDetail ? `\n${failureDetail}` : "";
187| throw new Error(
| ^
188| `Timed out after ${timeoutMs}ms waiting for ${label}${detail}`,
189| );
❯ waitForRuntimeState src/test/runtime-wait-helpers.ts:195:3
❯ src/runtime.idle-bridge-retirement.test.ts:73:9
⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯⎯[1/2]⎯
Test Files 1 failed (1)
Tests 2 failed | 1 passed (3)
Start at 09:50:17
Duration 9.15s (transform 1.35s, setup 0ms, import 1.64s, tests 7.11s, environment 0ms)
EXIT 1
4b. Live, on a dev instance (what the user sees)
Helper scripts (all under 2308/repro/): bbdev.sh runs the bb CLI against your isolated instance — it takes the checkout from BB_REPO (default: the current directory) and the ports from scripts/bb-dev-app env, so nothing in it needs editing; run it as BB_REPO=/path/to/bb bash bbdev.sh …. setup-project.sh creates the scratch repo and project the same way, and wait-idle.sh polls a thread until it is idle. bridges.sh lists the host daemon's children and their children, printing the executable name plus the last 100 characters of each command line so a process is identifiable from the snapshot itself: …/plugins/provider-claude-code/bridge-data is a provider bridge (bridge-worker-entry.ts), …/plugins/<id>/host-data is a plugin-host worker, claude … is the agent child, and the esbuild --service … --ping children are tsx's dev-only compile helpers (ignore them). Find the daemon pid with lsof -nP -iTCP:$BB_HOST_DAEMON_PORT -sTCP:LISTEN.
- Start the instance and create a project on a scratch repo (
setup-project.shdoes exactly this; the host id comes frombb machine list --json):scripts/bb-dev-app current # prints Server http://localhost:22861, Host daemon :30861 mkdir /tmp/bb-2308-qa && cd /tmp/bb-2308-qa && git init -q && echo hi > README.md && git add . && git commit -qm init curl -s -X POST http://localhost:22861/api/v1/projects -H 'content-type: application/json' \ -d '{"name":"qa","source":{"type":"local_path","path":"/tmp/bb-2308-qa","hostId":"host_shpa39kdsp"}}' HOST_ID=host_shpa39kdsp {"id":"proj_iwzvzf2fev","kind":"standard","name":"qa","gitRemoteUrl":null,"createdAt":1787590143021,"updatedAt":1787590143021,"sources":[{"id":"src_4xy4duan5e","projectId":"proj_iwzvzf2fev","type":"local_path","hostId":"host_shpa39kdsp","path":"/tmp/bb-2308-qa","isDefault":true,"createdAt":1787590143021,"updatedAt":1787590143021}]} - Baseline (
snapshot): the daemon (pid 17027, fromlsof -nP -iTCP:30861 -sTCP:LISTEN) has no bridge children — only tsx's esbuild helper, the parcel watcher and two plugin-host workers (…/plugins/keep-awake/host-data,…/plugins/provider-acp/host-data):daemon pid 17027 (09:49:03) PID PPID RSS_KB ETIME COMMAND (last 110 chars) 17033 17027 24592 00:24 ...ules/.pnpm/@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 85312 00:23 ...outs/bb/.claude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 20064 00:23 ...ules/.pnpm/@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17152 17027 109328 00:22 ...155abd/plugins/keep-awake/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-keep-awake-z3wYGm child 17165 17152 18880 00:22 ...ules/.pnpm/@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17196 17027 154496 00:21 ...bd/plugins/provider-acp/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-provider-acp-O1gWBL - Spawn one claude-code thread:
BB_REPO=$PWD bash bbdev.sh thread spawn --project proj_iwzvzf2fev --environment /tmp/bb-2308-qa \ --provider claude-code --prompt "Reply only with ok." --json { "id": "thr_7ptec8hm8d", "projectId": "proj_iwzvzf2fev", "environmentId": null, "providerId": "claude-code", "title": null, "titleFallback": "Reply only with ok.", "sectionId": null, … (20 more lines)Two…/plugins/provider-claude-code/bridge-dataprocesses appear (both arebridge-worker-entry.tsrunning the plugin'shost.mjs): 20603 is the daemon's provider-maintenance runtime (used for model listing; it only has an esbuild helper child) and 20879 is the environment runtime's bridge, with theclaudeagent 20880 under it (snapshot, 33 s after spawn):daemon pid 17027 (09:49:42) PID PPID RSS_KB ETIME EXE + last 100 chars of COMMAND 17033 17027 24608 01:03 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 85584 01:02 node ...laude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 20096 01:02 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17152 17027 98336 01:01 node ...gins/keep-awake/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-keep-awake-z3wYGm child 17165 17152 18880 01:01 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17196 17027 107552 01:00 node .../provider-acp/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-provider-acp-O1gWBL 20603 17027 195840 00:35 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data child 20604 20603 17456 00:35 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 20879 17027 255696 00:33 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data child 20880 20879 505392 00:32 claude ...nes-HOST.getbb.app-checkouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/thread-storage"]}}} - Wait for
status: "idle"(bash wait-idle.sh thr_7ptec8hm8d, orbash bbdev.sh thread show thr_7ptec8hm8d --json), then stop it:$ BB_REPO=$PWD bash bbdev.sh thread stop thr_7ptec8hm8d Thread thr_7ptec8hm8d stopped
- Actual: 3 s later the
claudeagent 20880 is gone but bridge 20879 is still there with no children. Expected: a bridge with nothing to serve goes away (immediately, or after a bounded grace period). (snapshot)daemon pid 17027 (09:50:03) PID PPID RSS_KB ETIME EXE + last 100 chars of COMMAND 17033 17027 24608 01:24 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 85680 01:23 node ...laude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 20112 01:23 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17152 17027 98432 01:22 node ...gins/keep-awake/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-keep-awake-z3wYGm child 17165 17152 18880 01:22 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17196 17027 107552 01:21 node .../provider-acp/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-provider-acp-O1gWBL 20603 17027 195936 00:55 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data child 20604 20603 17456 00:55 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 20879 17027 217520 00:53 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data95 s later (snapshot) the maintenance bridge 20603 has been retired byPROVIDER_MAINTENANCE_IDLE_TIMEOUT_MS(60 s) — proof that the daemon knows how to retire an idle bridge — while the environment bridge 20879 stays (age 2:31, 0 children, 217 MB RSS):daemon pid 17027 (09:51:41) PID PPID RSS_KB ETIME EXE + last 100 chars of COMMAND 17033 17027 24608 03:02 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 86240 03:01 node ...laude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 20624 03:01 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17152 17027 98944 03:00 node ...gins/keep-awake/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-keep-awake-z3wYGm child 17165 17152 19392 03:00 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17196 17027 107552 02:59 node .../provider-acp/host-data /var/folders/xx/xxxxxxxxxxxxxxxxxxxxxxxxxx/T/bb-host-provider-acp-O1gWBL 20879 17027 217520 02:31 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data - 3 min 45 s after the stop, spawn a second thread in the same environment (
spawn-2.json,thr_jxijqce242). No new environment bridge: 20879 (now 4:43 old) is reused and gets a freshclaudechild 29995. A new maintenance bridge 29912 appears (and is gone again by step 7, 60 s later, as before). This confirms both the per-environment keying and that "warm reuse" is what the resident process buys. By now the two plugin-host workers from step 2 have also been retired by their own 5-minute idle bound. (snapshot)daemon pid 17027 (09:53:53) PID PPID RSS_KB ETIME EXE + last 100 chars of COMMAND 17033 17027 24608 05:14 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 86960 05:13 node ...laude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 20736 05:13 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 20879 17027 190048 04:43 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data child 29995 20879 497648 00:04 claude ...nes-HOST.getbb.app-checkouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/thread-storage"]}}} 29912 17027 439312 00:07 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data - Archive that thread once idle (
BB_REPO=$PWD bash bbdev.sh thread archive thr_jxijqce242→Thread thr_jxijqce242 archived). The agent child 29995 is still alive: claude-code does not support provider-side archive, so the session is untouched and the bridge still owns the thread. Archiving therefore does not cause the stranding; it defers it to the idle reaper (30 min) if theproviderSessionReapingexperiment is enabled (it is off by default, see the claims table), and otherwise until an explicit stop or daemon exit. (snapshot)daemon pid 17027 (09:54:59) PID PPID RSS_KB ETIME EXE + last 100 chars of COMMAND 17033 17027 24608 06:20 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 87248 06:19 node ...laude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 20832 06:19 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 20879 17027 190480 05:49 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data child 29995 20879 504576 01:10 claude ...nes-HOST.getbb.app-checkouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/thread-storage"]}}} - Stop it (
BB_REPO=$PWD bash bbdev.sh thread stop thr_jxijqce242→Thread thr_jxijqce242 stopped). The agent exits; bridge 20879 remains, now 6:03 old with zero threads and nothing that will ever retire it short of daemon shutdown (snapshot):daemon pid 17027 (09:55:13) PID PPID RSS_KB ETIME EXE + last 100 chars of COMMAND 17033 17027 24608 06:34 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 87296 06:33 node ...laude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 20832 06:33 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 20879 17027 190496 06:03 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data - Watch it (
bash watch-bridge.sh 20879 190,log): one line every 10 s with the bridge's age, RSS and child count. It never exits and never gets a child; RSS drops from 190 MB to 96 MB at +90 s (a GC, not a retirement):09:55:25 +0s pid/ppid/etime/rss_kb:20879 17027 06:15 190496 children: 0 09:55:35 +10s pid/ppid/etime/rss_kb:20879 17027 06:25 190496 children: 0 09:55:45 +20s pid/ppid/etime/rss_kb:20879 17027 06:35 190496 children: 0 09:55:55 +30s pid/ppid/etime/rss_kb:20879 17027 06:45 190496 children: 0 09:56:06 +40s pid/ppid/etime/rss_kb:20879 17027 06:56 190496 children: 0 09:56:16 +50s pid/ppid/etime/rss_kb:20879 17027 07:06 190496 children: 0 09:56:26 +60s pid/ppid/etime/rss_kb:20879 17027 07:16 190496 children: 0 09:56:36 +70s pid/ppid/etime/rss_kb:20879 17027 07:26 190496 children: 0 09:56:46 +80s pid/ppid/etime/rss_kb:20879 17027 07:36 190496 children: 0 09:56:56 +90s pid/ppid/etime/rss_kb:20879 17027 07:46 96032 children: 0 09:57:06 +100s pid/ppid/etime/rss_kb:20879 17027 07:56 96032 children: 0 09:57:16 +110s pid/ppid/etime/rss_kb:20879 17027 08:06 96032 children: 0 09:57:26 +120s pid/ppid/etime/rss_kb:20879 17027 08:16 96032 children: 0 09:57:36 +130s pid/ppid/etime/rss_kb:20879 17027 08:26 96032 children: 0 09:57:47 +140s pid/ppid/etime/rss_kb:20879 17027 08:37 96032 children: 0 09:57:57 +150s pid/ppid/etime/rss_kb:20879 17027 08:47 96032 children: 0 09:58:07 +160s pid/ppid/etime/rss_kb:20879 17027 08:57 96032 children: 0 09:58:17 +170s pid/ppid/etime/rss_kb:20879 17027 09:07 96032 children: 0 09:58:27 +180s pid/ppid/etime/rss_kb:20879 17027 09:17 96032 children: 0 09:58:37 +190s pid/ppid/etime/rss_kb:20879 17027 09:27 96032 children: 0
- Final state, 3 min 28 s after the second stop (
snapshot;+95 sis the same): the daemon's only children are its dev-tooling processes and the stranded bridge. Nothing else will happen to 20879 until the daemon exits.daemon pid 17027 (09:58:38) PID PPID RSS_KB ETIME EXE + last 100 chars of COMMAND 17033 17027 13408 09:59 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 17096 17027 51104 09:58 node ...laude/worktrees/wf_846839f8-f8a-21/packages/host-watcher/src/parcel-subprocess/parcel-child-entry.ts child 17099 17096 12896 09:58 esbuild .../@esbuild+darwin-arm64@0.28.1/node_modules/@esbuild/darwin-arm64/bin/esbuild --service=0.28.1 --ping 20879 17027 96032 09:28 node ...ckouts-bb-.claude-worktrees-wf_846839f8-f8a-21-ed43ce155abd/plugins/provider-claude-code/bridge-data
No screenshots: the bug has no UI surface; the evidence is process-table snapshots, saved verbatim under 2308/repro/.
5. Root cause
Mechanism. The runtime keys one bridge process per provider artifact inside an environment-scoped runtime, and explicitly never scopes a process to a thread (packages/agent-runtime/src/runtime.ts:403-418):
/**
* One process per provider artifact: every thread of a provider in this
* environment runs on the same bridge process, and the bridge supervises
* whatever children it needs … The runtime never scopes a process to a
* thread.
*/
function resolveProviderProcessKey(args) {
return `${args.providerId}#bridge:${bridgeLaunchProcessKey(args.bridgeLaunch)}`;
}
When a thread is released (stop, reaper, archive, rewind discard, failed construction) the only thing that runs is (packages/agent-runtime/src/runtime.ts:478-489):
/**
* Releasing a thread is the moment a process can become retirable: a
* bridge process superseded by a plugin update was only being kept alive
* by the threads still running on it. A current process stays up for the
* provider's next thread; …
*/
async function releaseIdleProviderProcess(proc) {
await providerProcesses.retireSupersededBridgeProcessIfIdle(proc);
}
and that helper returns early for a current-artifact process (packages/agent-runtime/src/runtime-provider-process.ts:295-310):
async retireSupersededBridgeProcessIfIdle(providerProcess) {
if (providerProcess.identity.threadIds.size > 0) return;
const currentKey = this.currentProcessKeyByProviderId.get(providerProcess.providerId);
if (currentKey === undefined || currentKey === providerProcess.processKey) return; // ← always taken for a current bridge
await this.shutdownProvider({ … });
}
So "zero threads on a current bridge" is, by design, a no-op. What makes it a leak is that nothing ever revisits the process later:
- The daemon's periodic sweep (
apps/host-daemon/src/app.ts:65-66: every 5 min, sessions idle ≥ 30 min) only reaps sessions viastopThread; it lands in the same no-op. And for non-Codex providers it does not even get that far unless theproviderSessionReapingexperiment (default off,packages/domain/src/experiments.ts:34) is on:packages/agent-runtime/src/runtime.ts:1008-1018returnsnullfor every non-Codex session when the flag is false. RuntimeManager.evictIdleEnvironments()(apps/host-daemon/src/runtime-manager.ts:1259-1285) would shut idle runtimes (and their bridges) down, but it is only called from tests — grep finds no production caller.evictIdleRuntimeEntries()runs only when the base shell env changes (apps/host-daemon/src/runtime-manager.ts:640-650).- The two neighbouring pools that do have idle bounds show the intended pattern: the provider-maintenance runtime dies after 60 s idle (
apps/host-daemon/src/runtime-manager.ts:51,apps/host-daemon/src/runtime-manager.ts:888-912) and plugin-host workers after 5 min (apps/host-daemon/src/plugin-host-manager.ts:123; my daemon log shows "Host plugin worker stopped … host plugin worker became idle"). Environment bridges have no equivalent.
Why the symptom follows. A bridge is created on the first thread in an environment and released only by daemon shutdown, so the resident set is the set of environments that have ever hosted a thread for that provider since the daemon started — 29 bridges for 37 environments on the reporter's machine, 12 of them with zero threads because their sessions had been stopped or reaped.
History. The warm-process intent dates from c5b53caab (#1640), which added the test "keeps the current-hash bridge process when a thread is released" (packages/agent-runtime/src/runtime.process-lifecycle.test.ts:347-368). Before e42a4ef48 Codex was exempt because its process was thread-scoped (codex:thread:<id> keys) and was shut down at zero threads; #2325 collapsed Codex onto the shared one-bridge-per-artifact model and, with it, onto the same leak. Diff of runtime.ts in e42a4ef48.
Deeper issue. Process lifetime is tied to "is anything using it right now" with no notion of "has anything used it recently". Every other pooled resource in the daemon has an idle bound; this one doesn't. The fix is to add the bound, not to remove warm reuse.
6. Proposed fix (first principles)
What I tried first (and why it is wrong). The issue's patch — shut down whenever threadIds.size === 0 — applied to base (diff) makes my repro pass but fails 8 existing tests across four files (summary):
❯ @bb/agent-runtime src/runtime.acp-topology.test.ts (1 test | 1 failed) 1035ms
× releases the thread on the bridge when a construction times out on the runtime's side 1034ms
❯ @bb/agent-runtime src/runtime.codex-topology.test.ts (4 tests | 2 failed) 3836ms
× interrupt-stop settles the turn on the wire and then releases the thread's child 681ms
× releases the thread on the bridge when a construction times out on the runtime's side 647ms
❯ @bb/agent-runtime src/runtime.lifecycle.test.ts (28 tests | 2 failed) 11527ms
× forgets a thread whose thread/resume request was rejected 738ms
× keeps the provider running after thread stop 367ms
❯ @bb/agent-runtime:isolated src/runtime.process-lifecycle.test.ts (35 tests | 3 failed) 15008ms
× keeps the codex provider process when one session construction fails 298ms
× reaps a codex session after a terminal provider error before turn start 456ms
× reaps an idle codex session and resumes it later on the same process 443ms
⎯⎯⎯⎯⎯⎯⎯ Failed Tests 8 ⎯⎯⎯⎯⎯⎯⎯
FAIL @bb/agent-runtime src/runtime.acp-topology.test.ts > acp process topology > releases the thread on the bridge when a construction times out on the runtime's side
❯ src/runtime.acp-topology.test.ts:104:44
FAIL @bb/agent-runtime src/runtime.codex-topology.test.ts > codex process topology > interrupt-stop settles the turn on the wire and then releases the thread's child
❯ src/runtime.codex-topology.test.ts:314:44
FAIL @bb/agent-runtime src/runtime.codex-topology.test.ts > codex process topology > releases the thread on the bridge when a construction times out on the runtime's side
❯ src/runtime.codex-topology.test.ts:334:44
FAIL @bb/agent-runtime src/runtime.lifecycle.test.ts > createAgentRuntime lifecycle > thread setup and configuration > forgets a thread whose thread/resume request was rejected
❯ src/runtime.lifecycle.test.ts:240:51
FAIL @bb/agent-runtime src/runtime.lifecycle.test.ts > createAgentRuntime lifecycle > turn execution and thread commands > keeps the provider running after thread stop
❯ src/runtime.lifecycle.test.ts:914:46
FAIL @bb/agent-runtime:isolated src/runtime.process-lifecycle.test.ts > createAgentRuntime process lifecycle > keeps the codex provider process when one session construction fails
❯ src/runtime.process-lifecycle.test.ts:902:44
FAIL @bb/agent-runtime:isolated src/runtime.process-lifecycle.test.ts > createAgentRuntime process lifecycle > reaps a codex session after a terminal provider error before turn start
❯ src/runtime.process-lifecycle.test.ts:1035:46
FAIL @bb/agent-runtime:isolated src/runtime.process-lifecycle.test.ts > createAgentRuntime process lifecycle > reaps an idle codex session and resumes it later on the same process
❯ src/runtime.process-lifecycle.test.ts:1103:46
Test Files 4 failed | 4 passed (8)
Tests 8 failed | 101 passed (109)
EXIT 1
These are not stale assertions: "reaps an idle codex session and resumes it later on the same process" is the contract the #1604 reaper was built on (reap the agent child, keep the bridge warm), and "keeps the codex provider process when one session construction fails" / the ACP and Codex "construction times out" cases protect against a failed start tearing down a bridge that other work is about to use. forgets a thread whose thread/resume request was rejected breaks because the respawned bridge re-runs its scripted failure — a harness artifact, but it shows the respawn is observable. The naive change also re-introduces a spawn+initialize round trip between every stop and the next start in an environment.
Recommended change: bounded idle retirement of zero-thread bridges.
packages/agent-runtime/src/runtime-provider-process.ts: addidleSinceMs: number | nulltoRuntimeProviderProcess. Set it at spawn; set it tonowwheneveridentity.threadIdsbecomes empty (the release sites already callforgetThreadRuntimeStateForProviderState, so hook it there or inreleaseIdleProviderProcess); clear it when a thread is attached (threadIds.addin the identity registry /ensureProviderfor a thread).- Add
retireIdleProviderProcesses({ idleForMs, nowMs })to the manager: for each process withthreadIds.size === 0,idleSinceMs !== null,nowMs - idleSinceMs ≥ idleForMs, not inproviderStarting, and no pending requests, call the existingshutdownProvider(it already bumpsexpectedShutdownExpectationsso the exit is not reported as a crash). KeepretireSupersededBridgeProcessIfIdleas is — a superseded bridge should still go immediately. - Expose it on
AgentRuntimeand call it from the existing daemon sweep: inruntime.reapIdleProviderSessionsafter the session loop (or as a sibling call inRuntimeManager.reapIdleProviderSessions,apps/host-daemon/src/runtime-manager.ts:611-634). The sweep already runs every 5 min; use a shorter threshold for bridges than for sessions — 5 min idle matches the plugin-host worker bound, or reuse the 60 s maintenance bound. Make it independent of theproviderSessionReapingexperiment: a zero-thread bridge holds no session state, so retiring it loses nothing. - Tests: the repro in 4a adapted to call the sweep (fails before, passes after); one test that a bridge younger than the threshold, or with a thread attached, survives the sweep; one that a sweep during
ensureProviderstartup does not kill the starting process. All 109 existing agent-runtime lifecycle tests keep passing because none of them advance a clock past the threshold.
What could go wrong. (a) Race with a construction in flight: threadIds is empty between ensureProvider and the thread/identity notification — the providerStarting / pending-request checks plus the idle threshold cover it. (b) Retiring a bridge whose exit the runtime then reports through onProcessExit as unexpected — shutdownProvider already handles this for the superseded case. (c) Daemon/server wire shape: none changes, so no HOST_DAEMON_PROTOCOL_VERSION bump is needed; if the threshold is made a server policy field, bump it. (d) Alternative of wiring evictIdleEnvironments() into the sweep is simpler but tears down the whole runtime entry (turn state, staged rewinds, status watchers) and is broader than the bug; I would not start there.
7. PR review
No open pull request is linked to this issue.
8. Related issues
- #1604 Idle agent processes are never reclaimed for non-Codex providers (open) — the agent-child half; its reaper is what leaves these bridges behind.
- #1660 bb host process grew to 77GB RSS and froze the machine (open) — different mechanism (orphaned dev processes), same family of unbounded host-side process growth.
- #1363 Provider processes need one host-daemon lease owner tied to active turns (open) — a lease design would subsume this if it carried an idle bound.
- #2325 (merged) — removed the Codex exemption; this issue's title should drop "non-Codex".
9. Appendix
Artifacts
runtime.idle-bridge-retirement.test.ts— repro test;unit-test-base.txt— its output on base.candidate-fix-naive.diff,unit-test-with-fix-summary.txt,unit-test-with-fix.log— the issue's proposed patch and the 8 tests it breaks.ps-baseline.txt,ps-after-spawn.txt,ps-after-stop.txt,ps-after-stop-plus-95s.txt,ps-after-spawn-2.txt,ps-after-archive-2.txt,ps-after-stop-2.txt,ps-after-stop-2-plus-95s.txt,ps-after-stop-2-plus-190s.txt— process-table snapshots (bridges.shoutput);ps-bridge-watch.txt— 10-second watch of the stranded bridge (watch-bridge.shoutput).project-create.txt,spawn-1.json,spawn-2.json,stop-1.txt,archive-2.txt,stop-2.txt— CLI/API transcripts.- Scripts:
bbdev.sh,setup-project.sh,wait-idle.sh,bridges.sh,watch-bridge.sh,run-naive-fix.sh(applies the issue's patch, runs the eight test files, reverts). e42a4ef48-runtime.diff— how #2325 changed runtime.ts;issue-body.md— the issue as filed;dev-app-start.log(first run),revise-dev-app-start.log,revise-build.log(this run).
Helper scripts
#!/bin/bash
# Runs the bb CLI against YOUR isolated dev instance (the one started by
# `scripts/bb-dev-app current` in your bb checkout). Nothing is hardcoded:
# set BB_REPO to your checkout (default: the directory this script is run
# from) and the server/daemon ports come from `scripts/bb-dev-app env`.
# Usage: BB_REPO=/path/to/bb bash bbdev.sh <bb args...>
BB_REPO=${BB_REPO:-$PWD}
cd "$BB_REPO" || exit 1
eval "$(scripts/bb-dev-app env)" # exports BB_SERVER_URL, BB_HOST_DAEMON_PORT, BB_PROJECT_ID
exec node packages/scripts/dist/commands/run-cli.js "$@"
#!/bin/bash
# Lists the host daemon's direct children (provider bridges, plugin-host
# workers, the parcel watcher) with their own children (agent processes).
# The command column shows the executable's basename, then the LAST 100
# characters of the command line, so the script or data dir that identifies
# each process is visible:
# node ... /plugins/provider-<id>/bridge-data -> a provider bridge (bridge-worker-entry.ts)
# node ... /plugins/<id>/host-data -> a plugin-host worker (plugin-host-worker.ts)
# claude ... -> the claude-code agent child of a bridge
# esbuild --service ... --ping -> tsx's esbuild helper (dev-only, ignore)
# Usage: bash bridges.sh <host-daemon-pid>
# Find the pid with: lsof -nP -iTCP:$BB_HOST_DAEMON_PORT -sTCP:LISTEN
DAEMON=${1:?host daemon pid}
cmd() { ps -o command= -p "$1" | awk '{n=split($1,a,"/"); exe=a[n]; s=$0; if (length(s)>100) s="..." substr(s,length(s)-99); print exe " " s}'; }
echo "daemon pid $DAEMON ($(date +%H:%M:%S))"
printf '%-6s %-6s %-7s %-8s %s\n' PID PPID RSS_KB ETIME 'EXE + last 100 chars of COMMAND'
for p in $(pgrep -P "$DAEMON"); do
ps -o pid=,ppid=,rss=,etime= -p "$p" | tr -d '\n'; echo -n " "; cmd "$p"
for c in $(pgrep -P "$p"); do
echo -n " child "; ps -o pid=,ppid=,rss=,etime= -p "$c" | tr -d '\n'; echo -n " "; cmd "$c"
done
done
Commands run (in order)
gh issue view 2308 --repo get-bb/bb --json ... # issue + comments (none) pnpm install --frozen-lockfile --prefer-offline; pnpm exec turbo run build git fetch origin main; git rev-parse origin/main # == 494f66526 git log -S isThreadScopedCodexProcess --oneline -- packages/agent-runtime/src/runtime.ts # e42a4ef48, c5b53caab, 2a84ecfdc git show e42a4ef48 -- packages/agent-runtime/src/runtime.ts grep -rn evictIdleEnvironments apps packages # runtime-manager.ts + tests only cp repro/runtime.idle-bridge-retirement.test.ts packages/agent-runtime/src/ cd packages/agent-runtime && pnpm exec vitest run src/runtime.idle-bridge-retirement.test.ts # 2 failed / 1 passed on base scripts/bb-dev-app current # Server :22861, Host daemon :30861 BB_REPO=$PWD bash repro/setup-project.sh # /tmp/bb-2308-qa, proj_iwzvzf2fev lsof -nP -iTCP:30861 -sTCP:LISTEN # daemon pid 17027 bash repro/bridges.sh 17027 > ps-baseline.txt BB_REPO=$PWD bash repro/bbdev.sh thread spawn --project proj_iwzvzf2fev --environment /tmp/bb-2308-qa --provider claude-code --prompt "Reply only with ok." --json bash repro/bridges.sh 17027 (repeated after each step) BB_REPO=$PWD bash repro/wait-idle.sh thr_7ptec8hm8d; BB_REPO=$PWD bash repro/bbdev.sh thread stop thr_7ptec8hm8d sleep 92; bash repro/bridges.sh 17027 > ps-after-stop-plus-95s.txt BB_REPO=$PWD bash repro/bbdev.sh thread spawn ... (same project/environment) ... # thr_jxijqce242 BB_REPO=$PWD bash repro/wait-idle.sh thr_jxijqce242; BB_REPO=$PWD bash repro/bbdev.sh thread archive thr_jxijqce242 BB_REPO=$PWD bash repro/bbdev.sh thread stop thr_jxijqce242 bash repro/watch-bridge.sh 20879 190 > ps-bridge-watch.txt; bash repro/bridges.sh 17027 > ps-after-stop-2-plus-190s.txt BB_REPO=$PWD bash repro/run-naive-fix.sh > unit-test-with-fix.log # applies candidate-fix-naive.diff, 8 failed / 101 passed, reverts pnpm dev:stop; rm -rf <data dir> /tmp/bb-2308-qa
Daemon log excerpt (plugin-host workers do get an idle bound; from the first run's daemon log — in this run the same retirement is visible as workers 17152/17196 present at 09:51:41 and gone at 09:53:53)
{"level":30,"component":"host-daemon","pluginId":"keep-awake","reason":"host plugin worker became idle","uptimeMs":300472,"exitCode":0,"msg":"Host plugin worker stopped"}
{"level":30,"component":"host-daemon","pluginId":"provider-acp","reason":"host plugin worker became idle","uptimeMs":304965,"exitCode":0,"msg":"Host plugin worker stopped"}
10. Verification
What the independent verifier did. Followed both reproductions literally on base 494f66526 in a separate worktree and dev instance (server :20871, daemon :28871). The unit test gave the same 2 failed / 1 passed with the same waitForRuntimeState timeout for claude-code and codex; the issue's naive patch gave the same 8 failures / 101 passed; live, thread stop left a zero-child provider-claude-code/bridge-data process resident while the maintenance bridge was retired at 60 s, and a second spawn in the same environment reused it. Every code permalink and line number was checked against the base commit; no linked PR exists; origin/main was still the base commit.
Findings and what changed in this revision.
- Major — the "#1604 reaper is enabled and working" row was over-claimed. Confirmed in code:
findReapableIdleProviderSession(packages/agent-runtime/src/runtime.ts:1008-1018) returnsnullfor every non-Codex session unlessproviderSessionReapingEnabledis true, and that is theproviderSessionReapingexperiment, defaultfalse(packages/domain/src/experiments.ts:34), read atapps/server/src/internal/session.ts:43and fetched by the daemon sweep atapps/host-daemon/src/app.ts:700-704. The claims row is now "Partially verified" with that gating spelled out, step 7 and the section 5 sweep bullet say the reaper only applies with the experiment on, and section 6 keeps the recommendation that bridge retirement must not depend on the experiment. The root cause is unchanged: an explicit stop strands the bridge with or without the experiment. - Minor — snapshots did not identify bridges.
bridges.shnow prints the executable name plus the last 100 characters of each command line, and every live artifact was re-captured on a fresh instance (section 3) so…/plugins/provider-claude-code/bridge-dataversus…/plugins/<id>/host-datais readable in the snapshots themselves. A 10-secondwatch-bridge.shlog was added as continuous evidence. - Minor — hardcoded ports in bbdev.sh. The script now derives the checkout from
BB_REPOand the ports fromscripts/bb-dev-app env;setup-project.shandwait-idle.shwere added so the whole live repro runs without editing anything. - Minor — weak citation for per-environment keying. The claims row now points at
this.entries.set(args.environmentId, entry)(apps/host-daemon/src/runtime-manager.ts:1033) andget(environmentId)(apps/host-daemon/src/runtime-manager.ts:366-368) instead of only the comment.
Re-run results (this revision, worktree wf_846839f8-f8a-21, same base). Unit test on base: 2 failed / 1 passed (log). Naive patch: 8 failed / 101 passed, identical list (summary). Live: bridge 20879 survived two stop cycles and was last seen 9 min 28 s old with 0 children, 3 min 28 s after the final stop. One wrinkle worth recording: while re-running, a background capture of the +95 s snapshot initially appeared to be missing the bridge; re-reading the file on disk showed the bridge line was present (the capture had simply been read before its last line was flushed), and the later reuse of the same pid 20879 with a consistent start time confirms it never exited. Dev instance stopped, data dir and scratch repo deleted, ports 14861/22861/30861 confirmed free afterwards.