← reports

#2328 · Automatic recovery storms a dead Codex thread with repeated no-active-session errors

Bug Priority: High Effort: not set providers threads provider-codex open on GitHub 2026-08-24 · base 494f66526

Verdict: REPRODUCED (the bb-side wedge and the unbounded identical-rejection storm; the "automatic recovery" sender is the reporter's own plugin, not bb) · Root-cause confidence: high

1. TL;DR

A Codex thread gets into a state where every new message is rejected within ~25 ms with No active codex session for thread "<id>", producing one client/turn/rejected plus one system/error per attempt and flipping the thread to error, forever. The thread is not actually dead: its rollout is intact and a bb thread stop followed by one more message resumes it fine. What is broken is a split-brain between two daemon-side components: the host daemon's agent-runtime still has the thread registered (so it never asks the codex bridge to resume the session), while the codex bridge has already released the session (so it refuses every turn/start). Nothing in bb notices or repairs this divergence; the rejection is a generic -32000 error with no typed code or recovery hint, so the runtime cannot tell "no session" from any other failure.

I reproduced one concrete way the split-brain arises with real Codex on a dev instance: the runtime's generic JSON-RPC timeout for thread/stop is 30 s, but the codex bridge's interrupt path may legitimately take up to 60 s (child turn/interrupt) + 5 s (settlement) before it releases the session and answers. When the app-server is slow to answer the interrupt (here: paused for 43 s; in the wild: a hung or heavily loaded app-server mid-command, cf. the reporter's sibling issue #2327), the runtime gives up at 30 s and keeps the thread registered; the bridge finishes the interrupt later and releases the session. From then on the thread is wedged.

The "RECOVERY WAKE" messages themselves are not sent by bb. They come from the reporter's own plugin pixexid/bb-collab, whose thread.failed handler sends one agent-only wake every time the thread enters error, with no per-failure dedup or backoff. bb's contribution to the storm is that each rejected send emits run.failed → thread.failed again, so the plugin re-arms itself on bb's own rejection of its previous attempt. bb's first-party provider-retry plugin does not participate (it only retries rate-limit failures, bounded per reset window).

2. Claims vs findings

Claim from the issueStatusEvidence
An errored Codex thread receives repeated automatic agent-only RECOVERY WAKE new-turn requests.Partially refutedThe string RECOVERY WAKE does not exist in get-bb/bb. It is emitted by the reporter's plugin pixexid/bb-collab (server.ts, recoverErroredThread, triggered from bb.events.on("thread.failed")). bb's only first-party automatic retry (plugins/provider-retry) sends "Please continue." only for subscription rate-limit failures and is bounded by attemptedWindows. The repeated sends are external; the repeated rejections are bb's.
Every request is rejected with No active codex session, immediately followed by system/error.VerifiedLive: 9 wakes → 9 × (client/turn/rejected + system/error) with that exact message; thread status error after each (storm-events.json, screenshots below). Source: settleThreadCommandFailure appends both events for a failed turn.submit.
The recovery mechanism keeps retrying the same unrecoverable thread.Verified (mechanism is external, the re-arm is bb's)bb-collab dedups only while a send is in flight (recoveryInFlight). bb emits thread.failed on every applied transition into error (emitPluginThreadLifecycleOutcome), including a rejected start that never reached the provider, so the loop is closed by bb's own rejection.
Thread status error, environment ready, zero queued messages, last turn/completed preserved.VerifiedSame state in my repro: environment ready, last turn/completed {status: interrupted} at seq 65, no queued messages, status error.
The thread is "dead" / unrecoverable.RefutedThe provider session is restorable. bb thread stop on the errored thread (server release path → runtime forgets the thread) followed by one message resumed the rollout and ran a normal turn (heal-events.json, screenshot 5). The thread is wedged, not dead.
After threads.stop the storm continued; the next recovery was rejected with already has an active writer.Unverified (consistent with the mechanism)already has an active writer is a Codex app-server error (present in the codex 0.149.1 binary, absent from bb). After a stop the runtime forgets the thread, the next send resumes a fresh app-server from the rollout; whether Codex then rejects turn/start is Codex-side state (see #2327). In my repro the post-stop send succeeded.
Archiving the thread stopped the storm.Verified by codeensureThreadIsWritable rejects sends to archived threads before any dispatch, and bb-collab checks archivedAt before sending. Not a fix, just removes the target.
38 recovery requests, 10 within ~3 minutes.UnverifiedReporter's counts; I did not have their data. My run shows the cadence is bounded only by the sender: each rejection takes the daemon ~11–70 ms.

3. Environment

4. Minimal reproduction

4a. Unit-level (deterministic, no Codex needed)

A minimal provider bridge models the two codex-bridge behaviours involved: thread/stop releases the session immediately but answers late; turn/start with no session answers -32000 "No active codex session for thread …" (same code path shape as requireLiveSessionForTurn). The test drives the real @bb/agent-runtime through a stop whose reply arrives after the runtime's 30 s timer.

  1. Copy slow-stop-no-session-bridge.ts to packages/agent-runtime/src/test/bridges/ and runtime.issue-2328-no-session-wedge.test.ts to packages/agent-runtime/src/.
  2. Run from the package:
    cd packages/agent-runtime && pnpm exec vitest run src/runtime.issue-2328-no-session-wedge.test.ts
  3. Result on 494f66526 (vitest-output.txt):
    ✓ keeps the thread registered after a timed-out stop, then rejects every turn without ever re-resuming
    × EXPECTED (fails today): a timed-out stop leaves no stale registration behind
      AssertionError: expected true to be false   // runtime.hasThread("t2") is still true after the stop timed out
    The first test passes because it asserts the buggy behaviour: after stopThread rejects with JSON-RPC request timed out: thread/stop, hasThread stays true, five runTurn calls are each rejected with /No active codex session/, and the bridge's request record shows five turn/start and zero thread/resume. An explicit resumeThread then makes the same turn succeed. The second test asserts the desired invariant and fails on main.

4b. Live end-to-end with real Codex (what a user hits)

  1. Start a dev instance, create a project on a scratch repo, spawn a codex thread that runs a long command:
    bb thread spawn --project <proj> --provider codex --permission-mode full --title "2328 repro" \
      --prompt "Run the shell command 'sleep 300' and wait for it to finish. Then reply only with ok." --json
    Wait until the event log shows item/started commandExecution … sleep 300.
  2. Find the thread's codex app-server process under your host daemon (daemon → bridge → node …/codex app-server → vendor codex app-server binary). Make it slow to answer the interrupt, then stop the thread, wait so that the runtime's 30 s timer fires but the bridge's 60 s child timer does not, and let the app-server continue (live-stop-sequence.sh):
    kill -STOP <app-server pid>
    bb thread stop thr_6932zwarfj --json        # returns {"ok":true} after 30 s (see note)
    sleep 12                                     # total pause ≈ 43 s: > 30 s runtime timeout, < 60 s bridge child timeout
    kill -CONT <app-server pid>
    Expected: the stop interrupts the turn and the thread is idle and usable.
    Actual (log, dev log):
    t0 10:10:34 kill -STOP 79592 (codex app-server)
    t1 10:10:34 bb thread stop thr_6932zwarfj
    { "ok": true, "threadId": "thr_6932zwarfj" }
    t2 10:11:05 stop command returned; waiting 12s
    t3 10:11:17 kill -CONT 79592
    
    [host-daemon] online host RPC failed {"type":"thread.stop"}  err: "JSON-RPC request timed out: thread/stop"
    [host-daemon] Online host RPC {"commandType":"thread.stop","errorCode":"command_failed","handlerMs":30002.3,"ok":false}
    [server] Awaited thread stop command failed {"intent":"interrupt","threadId":"thr_6932zwarfj"}
    # after SIGCONT the interrupt lands: seq 61 system/thread/interrupted {reason: manual-stop}, seq 65 turn/completed {status: interrupted}; thread status: idle
  3. Send any message (this is what bb-collab's wake does: threads.send with mode: "auto"):
    bb thread tell thr_6932zwarfj --mode auto --json "RECOVERY WAKE 1 - reconcile state before resuming."
    Expected: the turn runs (the session is resumable from its rollout).
    Actual:
    66 client/turn/requested
    67 client/turn/rejected  {reason: command_failed, message: 'No active codex session for thread "thr_6932zwarfj"'}
    68 system/error          {code: thread_command_failed, message: 'Command turn.submit failed', detail: 'No active codex session for thread "thr_6932zwarfj"'}
    status after wake 1: error
    [host-daemon] Online host RPC {"commandType":"turn.submit","errorCode":"command_failed","handlerMs":32.9,"ok":false}
  4. Repeat the send 8 more times, each after the previous one settles (storm.sh; this is the shape of bb-collab's thread.failed → send loop). Actual (storm-output.txt, storm-events.json):
    Counter({'client/turn/requested': 9, 'client/turn/rejected': 9, 'system/error': 9})
    status after wake 1..9: error (×9)
    turn.submit handlerMs: 32.9, 29.1, 17.8, 25.4, 11.1, 17.4, 25.8, 70, 12.9   # no resume is ever attempted
  5. Heal check (proves the runtime registration was the stale side):
    bb thread stop thr_6932zwarfj --json     # thread is in status error → server takes the release path
    bb thread tell thr_6932zwarfj --mode auto --json "RECOVERY WAKE 1 - reconcile state before resuming."
    Actual: seq 94 thread/identity (a fresh thread/resume of the same rollout), seq 97 turn/started, the agent answers normally; daemon turn.submit … handlerMs 1919.9, ok:true.
Thread idle after the stop that timed out on the runtime side
Step 2: after the 43 s pause the interrupt landed ("Worked for 1m 3s · Stopped manually"); the thread looks idle and healthy. The runtime's stop RPC has already timed out at 30 s.
First wake rejected with No active codex session
Step 3: the first follow-up is rejected: "Command turn.submit failed — No active codex session for thread "thr_6932zwarfj"". The composer now says "Retry by sending a follow-up message", which will fail identically.
Nine wakes, nine identical rejections
Step 4: the storm. Every wake produces one identical rejection pair; nothing changes between attempts because no component ever tries to rebuild the session.
After bb thread stop the next wake runs
Step 5: after bb thread stop on the errored thread, the very next wake resumes the rollout and runs (the agent even finds the original sleep 300 still alive). Same thread, same rollout: it was never dead.

Note on step 2: the CLI and the POST /threads/:id/stop route report {"ok": true} even though the daemon stop timed out, because an interrupt-intent stop swallows its failure (runAwaitedThreadStopCommand: "An interrupt swallows a failure: its stopping status is durable"). The user has no signal that the runtime/bridge just diverged.

Repro files: 2328/repro/ (bridge, test, scripts, event dumps, full dev log, candidate-fix diff).

5. Root cause

5a. The wedge: runtime registration outlives the bridge session, and nothing reconciles them

On the host daemon, two components hold per-thread session state:

The daemon decides whether to resume a session purely from the runtime's view (command-handlers/thread.ts#L168-L175):

async function resumeThreadRuntimeIfMissing(args) {
  if (entry.runtime.hasThread(command.threadId)) {
    return;                       // ← never asks the bridge
  }
  …
  await entry.runtime.resumeThread({ … providerThreadId: resumeContext.providerThreadId … });
}

So whenever the registry says "registered" but the bridge has no entry, every turn.submit goes straight to turn/start and is rejected. The rejection is a generic BRIDGE_ERROR (-32000) with no recovery hint (handleTurnStart → rejectWithCodexError), and runTurn's catch only clears the pending-start marker and marks the session idle (runtime.ts#L2186-L2192); it never drops the registration. Compare steerTurn, which has a typed NO_ACTIVE_TURN to act on; there is no typed "no session" code in BRIDGE_JSON_RPC_ERRORS (errors.ts#L13-L26). The daemon's idle-session reaper cannot heal it either: it calls stopThread, which requires a live process and only runs after the idle window, and the server never sends a release for an error thread on its own.

Every rejected submit is settled by the server as a failed command: client/turn/rejected + system/error + run.failed (thread-lifecycle.ts#L771-L810), and run.failed into error emits the plugin event thread.failed (plugin-thread-events.ts#L43-L54). That is the re-arm signal bb-collab listens to.

5b. One proven way the divergence is created: stop timeout asymmetry

The runtime's generic JSON-RPC timeout is 30 s (runtime-json-rpc.ts#L324-L334, args.timeoutMs ?? 30_000) and stopThread does not override it (runtime.ts#L2302-L2340). On success it forgets the thread; on any failure it rethrows without forgetting. The codex bridge's interrupt path (handleThreadStop) sends turn/interrupt to the app-server with CHILD_REQUEST_TIMEOUT_MS = 60_000, then waits up to INTERRUPT_SETTLEMENT_TIMEOUT_MS = 5_000, and only then calls releaseSession(session) and answers. Any interrupt that takes between 30 s and 65 s therefore ends with: runtime timed out (registration kept), bridge released (session gone). The live repro above hits exactly this window (app-server paused 43 s). A real app-server can be that slow when it is wedged inside a long command execution, which is precisely the sibling report #2327 from the same reporter ("turn/completed can precede command completion…" followed by the same two errors).

Other divergence paths exist with the same symptom and are not covered by a timeout fix alone: any lost/late thread/stop reply (bridge stdout hiccup, daemon event-loop stall), and any future bridge-side release that the runtime does not observe. The underlying defect is that the runtime treats its registry as the truth and the protocol gives it no typed way to learn it is wrong.

5c. Why it looks like "automatic recovery storms the thread"

bb-collab's recoverErroredThread (server.ts) runs on every thread.failed whose thread has status === "error", guards only against a concurrent in-flight send, and calls bb.sdk.threads.send({mode: "auto", input: [{visibility: "agent-only", text: "RECOVERY WAKE — …"}]}). bb rejects in ~25 ms, emits thread.failed, the plugin sends again. The cadence the reporter saw ("sub-10-second") is the plugin's own threads.get/environments.status round trips; bb imposes no bound.

6. Proposed fix (first principles)

  1. Make "no session" a typed protocol outcome and act on it in the runtime (root cause). Add e.g. SESSION_NOT_FOUND: -32004 to BRIDGE_JSON_RPC_ERRORS (provider-bridge-protocol) and return it from the codex bridge's requireLiveSessionForTurn / handleTurnSteer (and the pi/ACP equivalents) instead of the generic BRIDGE_ERROR. In runtime.runTurn/steerTurn, on that code: forgetThreadRuntimeStateForProviderState(proc.identity, threadId) before rethrowing, so the daemon's next turn.submit takes the resumeThreadRuntimeIfMissing path and rebuilds the session from the rollout (a resume replaces whatever the bridge still holds, see constructThreadSession). This heals every divergence cause, not just the timeout. Risk: a bridge that wrongly reports no-session for a live turn would cause a resume that kills that turn; the code should only be emitted when the bridge truly holds nothing. This changes the bridge wire contract (new error code), so document it in docs/provider-bridge-protocol.md; it does not change server↔daemon payloads, so no HOST_DAEMON_PROTOCOL_VERSION bump is needed for this part.
  2. Do not keep a registration the runtime cannot vouch for after a stop it gave up on. In stopThread, distinguish a bridge rejection (JsonRpcResponseError: the bridge answered and still owns the session — keep state; two existing tests depend on this) from no answer (timeout / process exit): forget the thread so the next command resumes. I validated this variant: with candidate-fix-runtime-stop-timeout.diff the "EXPECTED" test passes, the buggy-behaviour test fails at its hasThread assertion as intended, and the four existing runtime suites pass (115/115). A naive "forget on any stop failure" breaks preserves active turn state when the stop request fails and keeps releasing later sessions after one release fails (log). Additionally give thread/stop a request timeout that covers the bridge's documented interrupt budget (≥ 65 s for codex) so the common case does not race at all.
  3. Do not re-arm external recovery on bb's own rejection (server policy). A turn.submit that never produced a provider turn (rejected before turn/input/accepted) is a command failure, not a provider failure; consider not emitting a fresh thread.failed when the thread was already in error with the same error detail, or include error/requestId context so plugins can dedup. Expose a recovery-disposition (or at least bb thread stop's release semantics) as the documented "start a fresh provider session" action; today the user-facing composer hint "Retry by sending a follow-up message" is wrong for this state.
  4. Surface stop failures. POST /threads/:id/stop returns {ok:true} after a 30 s daemon timeout; log at warn is not enough when the consequence is a wedged thread.

The plugin-side bound (single-flight per interruption identity, backoff, terminal disposition) the issue asks for is bb-collab's responsibility; items 1–2 remove the state that makes its retries pointless, item 3 stops bb from feeding the loop.

7. PR review

No open pull request is linked to this issue.

8. Related issues

9. Appendix

Daemon RPC outcomes for the live thread (from dev-instance-full.log)

thread.start   handlerMs 5422.6   ok:true
thread.stop    handlerMs 30000.9  errorCode command_failed   (attempt 1: pause 67 s > 60 s child timeout → bridge's interrupt errored, no release, next wake SUCCEEDED — control case)
thread.stop    handlerMs 30002.3  errorCode command_failed   (attempt 2: pause 43 s → bridge released after the runtime gave up)
turn.submit    handlerMs 32.9 / 29.1 / 17.8 / 25.4 / 11.1 / 17.4 / 25.8 / 70 / 12.9   errorCode command_failed  (9 wakes)
turn.submit    handlerMs 1919.9   ok:true   (after bb thread stop: resume + turn)

Control case screenshot (attempt 1, wake succeeded because the pause exceeded the bridge's own 60 s timeout): assets/2328-02-first-wake-rejected.png (file name predates the result; it shows the "ok" reply). assets/2328-01-idle-after-stop-timeout.png is the idle state after attempt 1.

Bridge-side error message in the daemon log

[host-daemon] online host RPC failed {"type":"turn.submit"}
  err: { "type": "JsonRpcResponseError", "message": "No active codex session for thread \"thr_6932zwarfj\"",
         at jsonRpcResponseError (packages/provider-bridge-protocol/src/bridge-kit/runtime-json-rpc.ts:178)
         at settleJsonRpcResponse (…/runtime-json-rpc.ts:296)
         at handleStdoutLine (packages/agent-runtime/src/runtime.ts:1463) }
[server] … "status": 502, "body": { "message": "No active codex session for thread \"thr_6932zwarfj\"", "code": "command_failed", "retryable": false }

Where the reporter's wake comes from (pixexid/bb-collab server.ts, excerpt)

bb.events.on("thread.failed", async (payload) => {
  …
  if (status === "error") { … await recoverErroredThread(id, payload.thread.projectId, identity.holder, identity.lane); }
});
const recoverErroredThread = async (threadId, projectId, holder, lane) => {
  if (recoveryInFlight.has(recoveryKey)) { bb.log.warn(`error-recovery wake suppressed: … reason=recovery-in-flight`); return null; }
  …
  await withRecoverySendTimeout(threadId, () => bb.sdk.threads.send({
    threadId, mode: "auto",
    input: [{ type: "text", visibility: "agent-only", text: `RECOVERY WAKE — reconcile state before resuming. …`, mentions: [] }],
  }));
  …
};

Commands run (abridged)

gh issue view 2328 --comments ; gh issue view 2327 --json title,body
git grep -n "RECOVERY WAKE" ; git grep -n "No active codex session" ; gh search code "RECOVERY WAKE" --repo pixexid/bb-collab
strings …/@openai/codex-darwin-arm64/vendor/aarch64-apple-darwin/bin/codex | grep "active writer"   # → " already has an active writer"
pnpm install --frozen-lockfile --prefer-offline ; pnpm exec turbo run build
cd packages/agent-runtime && pnpm exec vitest run src/runtime.issue-2328-no-session-wedge.test.ts
pnpm exec turbo run typecheck --filter=@bb/agent-runtime
scripts/bb-dev-app current ; curl -X POST $BB_SERVER_URL/api/v1/projects … ; bb thread spawn … ; live-stop-sequence.sh 79592 thr_6932zwarfj 12 ; storm.sh thr_6932zwarfj 8 ; heal.sh thr_6932zwarfj
git diff packages/agent-runtime/src/runtime.ts > candidate-fix-runtime-stop-timeout.diff ; git checkout -- packages/agent-runtime/src/runtime.ts
pnpm dev:stop

2026-09-30 verification: registration after delayed stop and session removal

Verdict: REPRODUCED for the bb-side registration mismatch and repeated rejection of bounded caller-driven attempts. Root-cause confidence: high within this synthetic runtime-package scope. This does not reproduce an autonomous external-plugin storm. The original report, its live-provider observations and screenshots remain historical evidence, preserved above; no new visual or production-runtime claim is made here.

Base: 0e7b518f135d43005dae201ef34ebb3001607eb4, fetched from trusted get-bb/bb origin/main. Linux 6.18.44 x86_64, Node 22.19.0, pnpm 9.15.0, Vitest 4.1.1. Both separate clean checkouts were pinned to that SHA before adding the same two synthetic fixture files. Workspace capacity was 1,776,158 free inodes and approximately 25 GiB free before setup; 1,686,333 inodes remained after verification. Prior issue evidence was retained.

Eligibility: open native Bug, Priority High, Effort Medium; existing labels providers, threads, provider-codex and confirmed-repro. All three comments and the paginated timeline were read. Linked PR #3460 is closed and merged; no linked open PR or open bb PR mentioning #2328 was found. The historical report had no new edits. Visible GitHub activity did not show overlapping work; internal SlopCop runtime state was not queried under the no-real-runtime constraint. Both repositories are public.

Two personal clean runs

The same agent personally ran the identical fixture in two clean same-SHA checkouts with separate frozen installations and fresh synthetic state for every test. Run A began at 18:16:43 UTC and Run B at 18:17:15 UTC on 2026-09-30. Each passed 4/4 tests; result summaries, request counts and retained events agree. Normal frozen installs and the repository's Turbo generate:test-bridges build succeeded. The selected Turbo tests ran their normal generation/native-module prerequisites and were forced executions, not cached results.

The actual createAgentRuntime library, provider-worker transport, registration map, JSON-RPC timeout, stop, turn and resume implementations ran unchanged inside the test process. A small synthetic bridge module was loaded by the repository's normally built test worker. No bb application, server, daemon instance, real Codex app-server or external plugin ran. All traffic used stdio; no network ports or provider credentials were needed. Each test used a new local directory and shut down its own runtime.

The delayed-stop test opens a real assembled active turn, waits until the synthetic child has received thread/stop, advances only the test process's timeout clock by 30,001 ms, and then releases the child's independently gated late reply. The child removes its synthetic session and emits the interrupted boundary before replying. This exercises the actual timeout rather than changing its constant or waiting 30 seconds on the wall clock. It does not measure real Codex interruption latency.

Expected versus actual

Expected: an uncertain stop followed by session removal must not permanently leave the runtime claiming a usable registration; a genuinely rejected stop should preserve a session that still works. Actual: a timed-out stop leaves hasThread true even after the synthetic bridge finishes the stop and deletes its session. Three bounded calls to the actual runTurn each receive No active codex session; no thread/resume is sent. Explicit resumeThread restores the same synthetic session identity and the next turn completes.

RESULT {"case":"delayed-stop","stopOutcome":"JSON-RPC request timed out: thread/stop","registeredAfterStop":true,"rejected":3,"resumesBeforeRepair":0,"resumesAfterRepair":1,"successfulAfterRepair":true,"turnRequests":5}
RESULT {"case":"session-removal","stopOutcome":"not-requested","registeredAfterStop":true,"rejected":3,"resumesBeforeRepair":0,"resumesAfterRepair":1,"successfulAfterRepair":true,"turnRequests":5}
RESULT {"case":"stop-rejected","stopOutcome":"rejected","registeredAfterStop":true,"rejected":0,"resumesBeforeRepair":0,"resumesAfterRepair":0,"successfulAfterRepair":true,"turnRequests":2}
RESULT {"case":"normal-stop","stopOutcome":"resolved","registeredAfterStop":false,"rejected":0,"resumesBeforeRepair":0,"resumesAfterRepair":1,"successfulAfterRepair":true,"turnRequests":2}

Each mismatch case has exactly five turn requests: the baseline, three deliberate rejected follow-ups, and the successful post-resume turn. A 150 ms real-time observation after the three failures finds no extra turn requests. This is a bounded observation of this runtime fixture, not proof that every higher-level retry mechanism is absent. Each case retains its baseline completion/interruption and successful final completion in the captured event history.

The external recovery plugin was neither fetched nor run. The current source search found no first-party RECOVERY WAKE sender. The reported live count, cadence, database rejection/error pairs, thread.failed re-arming, archive behavior and writer-lock errors are not dynamically reverified. The test caller supplies every repeated attempt; bb's demonstrated contribution is failing to repair its stale registration when those attempts encounter an untyped missing-session error.

Current root cause and code evidence

Small fix proposal, not implemented: give missing-session refusals a typed recovery disposition so the runtime can invalidate the stale registration and make a subsequent submission eligible for resume. Treat uncertain stop outcomes (timeout/process loss) separately from explicit bridge rejection, preserving the rejected-stop control above. Align stop timeout policy with the provider's interrupt budget, and guard against late replies or late cleanup affecting a newer session generation. External retry controllers separately need a bounded per-incident disposition; repairing the registration mismatch does not itself prove their loop is bounded.

Next test: after a proposed change, require registration invalidation or a single controlled resume for delayed-stop and session-removal cases, retain the rejected-stop control, and test late old-generation cleanup racing a new session. A separately scoped server test should count persisted rejection/failure events and prove bounded retry re-arming; it must not infer those results from this library fixture.

Repeatable commands and complete fixtures

Use Node 22.19.0 and pnpm 9.15.0. The explicit store path is the writable store used here; choose a writable equivalent elsewhere. Raw requests, events and results stay outside the reports repository in fresh issue-2328-state-* directories. Exploratory fixture runs timed out before reproducing behavior because of malformed synthetic request IDs/startup setup; those logs are retained locally and excluded from the two successful verification runs. The final request IDs and identity/reset sequence conform to trusted repository schemas.

git clone https://github.com/get-bb/bb.git run-a
cd run-a
git checkout --detach 0e7b518f135d43005dae201ef34ebb3001607eb4
node --version  # v22.19.0
pnpm --version  # 9.15.0
pnpm install --frozen-lockfile --store-dir /workspace/.pnpm-store
pnpm exec turbo run generate:test-bridges --filter=@bb/agent-runtime
# Save the two complete inline files below at their stated paths.
pnpm exec turbo run test --filter=@bb/agent-runtime --force -- src/issue-2328.test.ts
# Repeat in a second fresh clone named run-b at the same SHA.
# Install independently and copy only the two fixture files, never prior state.
packages/agent-runtime/src/issue-2328-fixture.mjs
import { appendFileSync, existsSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
const root=process.env.ISSUE2328_STATE;
const mode=process.env.ISSUE2328_MODE;
const sessions=new Set();
let turn=0;
const send=(value)=>process.stdout.write(JSON.stringify({jsonrpc:'2.0',...value})+'\n');
const reply=(id,result={})=>send({id,result});
const error=(id,message)=>send({id,error:{code:-32000,message}});
const deltas=(threadId,values)=>send({method:'thread/delta',params:{threadId,deltas:values}});
function handleLine(line) {
 const m=JSON.parse(line);const p=m.params??{};
 appendFileSync(join(root,'requests.jsonl'),JSON.stringify({method:m.method,threadId:p.threadId??null})+'\n');
 if(m.method==='initialize') return reply(m.id,{protocolVersion:2,capabilities:{sessionRestore:true,fork:'checkpoint',approvalEnforcedBy:'runtime',grammarVersions:[3,3],steerMode:'inject'}});
 if(m.method==='thread/start'||m.method==='thread/resume') {
  sessions.add(p.threadId);send({method:'thread/identity',params:{threadId:p.threadId,providerThreadId:'synthetic-session',sessionRestorable:true}});deltas(p.threadId,[{kind:'session.reset'}]);return reply(m.id,{providerThreadId:'synthetic-session',sessionRestorable:true});
 }
 if(m.method==='thread/stop') {
  if(mode==='stop-rejected')return error(m.id,'synthetic stop rejected');
  const finish=()=>{
   sessions.delete(p.threadId);
   if(p.activeTurnId)deltas(p.threadId,[{kind:'turn.boundary',providerTurnId:p.activeTurnId,status:'interrupted'}]);
   reply(m.id);writeFileSync(join(root,'stop-finished'),'done');
  };
  if(mode==='delayed-stop') {
   writeFileSync(join(root,'stop-received'),'ready');
   const timer=setInterval(()=>{if(existsSync(join(root,'release-stop'))){clearInterval(timer);finish();}},10);
  } else finish();
  return;
 }
 if(m.method==='turn/start') {
  if(existsSync(join(root,'remove-session')))sessions.delete(p.threadId);
  if(!sessions.has(p.threadId))return error(m.id,`No active codex session for thread "${p.threadId}"`);
  const providerTurnId='synthetic-turn-'+(++turn);
  deltas(p.threadId,[{kind:'turn.open',providerTurnId},{kind:'input.accepted',clientRequestId:p.clientRequestId}]);
  reply(m.id);
  if(!p.input.some(x=>x.type==='text'&&x.text==='hold'))deltas(p.threadId,[{kind:'turn.boundary',providerTurnId,status:'completed'}]);
  return;
 }
 reply(m.id);
}
export const experimental_providerBridge={experimental_apiVersion:1,handleLine,onClose:()=>process.exit(0),onSigterm:()=>process.exit(0)};
packages/agent-runtime/src/issue-2328.test.ts
import { existsSync, mkdtempSync, readFileSync, unlinkSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { expect, it, vi } from 'vitest';
import { z } from 'zod';
import type { ThreadEvent } from '@bb/domain';
import { createScriptedEchoRuntime, fullRuntimeOptions } from './test/runtime-test-harness.js';
import { promptTextInput } from './test/prompt-input.js';
const realSetTimeout=setTimeout;
const sleep=(ms:number)=>new Promise(resolve=>realSetTimeout(resolve,ms));
async function until(check:()=>boolean) {
 const limit=Date.now()+8000;
 while(!check()){if(Date.now()>limit)throw new Error('Synthetic fixture timeout');await sleep(10);}
}
const requestRow=z.object({method:z.string(),threadId:z.string().nullable()});
for(const mode of ['delayed-stop','session-removal','stop-rejected','normal-stop']) {
 it('2328 '+mode,async()=>{
  const root=mkdtempSync(fileURLToPath(new URL('./issue-2328-state-',import.meta.url)));
  const events:ThreadEvent[]=[];
  const runtime=createScriptedEchoRuntime({runtime:{workspacePath:root,bridgeBundleDir:fileURLToPath(new URL('../dist/test-bridges/',import.meta.url)),env:{ISSUE2328_STATE:root,ISSUE2328_MODE:mode},onEvent:event=>events.push(event)},launch:{modulePath:fileURLToPath(new URL('./issue-2328-fixture.mjs',import.meta.url))}});
  const args={environmentId:'synthetic-environment',projectId:'synthetic-project',threadId:'synthetic-thread',providerId:'fake',options:fullRuntimeOptions};
  const methods=()=>readFileSync(join(root,'requests.jsonl'),'utf8').trim().split('\n').map(line=>requestRow.parse(JSON.parse(line)).method);
  let request=0;
  const run=(text='synthetic input')=>runtime.runTurn({threadId:args.threadId,clientRequestId:'creq_abcdefghj'+'kmnpqrst'[request++],input:[promptTextInput({text})],options:fullRuntimeOptions});
  try {
   const start=await runtime.startThread(args);
   await run(mode==='delayed-stop'?'hold':'baseline');
   await until(()=>events.some(e=>e.type===(mode==='delayed-stop'?'turn/started':'turn/completed')));
   let stopOutcome='not-requested';
   if(mode==='delayed-stop') {
    vi.useFakeTimers({toFake:['setTimeout','clearTimeout']});
    const stopping=runtime.stopThread({threadId:args.threadId}).then(()=> 'resolved',e=>String(e.message));
    await until(()=>existsSync(join(root,'stop-received')));
    await vi.advanceTimersByTimeAsync(30001);
    stopOutcome=await stopping;expect(stopOutcome).toMatch(/JSON-RPC request timed out: thread\/stop/);
    vi.useRealTimers();writeFileSync(join(root,'release-stop'),'release');
    await until(()=>existsSync(join(root,'stop-finished')) && runtime.getActiveTurnId(args.threadId)===null);
   } else if(mode==='session-removal') writeFileSync(join(root,'remove-session'),'remove');
   else if(mode==='stop-rejected') {
    await expect(runtime.stopThread({threadId:args.threadId})).rejects.toThrow('synthetic stop rejected');stopOutcome='rejected';
   } else {await runtime.stopThread({threadId:args.threadId});stopOutcome='resolved';}
   const registeredAfterStop=runtime.hasThread(args.threadId);
   expect(registeredAfterStop).toBe(mode!=='normal-stop');
   let rejected=0;
   if(mode==='delayed-stop'||mode==='session-removal') {
    for(let i=0;i<3;i++) {await expect(run()).rejects.toThrow('No active codex session');rejected++;expect(runtime.hasThread(args.threadId)).toBe(true);}
    expect(methods().filter(m=>m==='thread/resume')).toHaveLength(0);
    const before=methods().filter(m=>m==='turn/start').length;await sleep(150);expect(methods().filter(m=>m==='turn/start')).toHaveLength(before);
   }
   const resumesBeforeRepair=methods().filter(m=>m==='thread/resume').length;
   if(mode==='session-removal')unlinkSync(join(root,'remove-session'));
   if(mode!=='stop-rejected')await runtime.resumeThread({...args,providerThreadId:start.providerThreadId});
   const completedBefore=events.filter(e=>e.type==='turn/completed').length;
   await run('after repair');await until(()=>events.filter(e=>e.type==='turn/completed').length===completedBefore+1);
   const result={case:mode,stopOutcome,registeredAfterStop,rejected,resumesBeforeRepair,resumesAfterRepair:methods().filter(m=>m==='thread/resume').length,successfulAfterRepair:true,turnRequests:methods().filter(m=>m==='turn/start').length};
   await runtime.shutdown();expect(runtime.hasThread(args.threadId)).toBe(false);
   writeFileSync(join(root,'evidence.json'),JSON.stringify({result,events,methods:methods()},null,2));
   process.stderr.write('RESULT '+JSON.stringify(result)+'\n');
  } finally {vi.useRealTimers();await runtime.shutdown();writeFileSync(join(root,'final-events.json'),JSON.stringify(events,null,2));}
 },15000);
}

Trust and remaining limits: issue bodies, comments, links and embedded code were treated only as untrusted evidence. Neither historical report scripts nor issue-supplied commands, external plugin code or linked PR branches were executed. The fixture was authored from trusted current contracts and test helpers. No production source or dependency was changed; there was no investigation-agent delegation, production fix, PR, manually started workflow or extra issue comment. This confirms the bb-side mismatch under the stated synthetic sequences; the reporter's original divergence path and external automatic storm remain outside this verification.