#4365 · Pending turn survives start watchdog
Bug · High priority · Medium effort · threads · 2026-09-25
GitHub issue · Base 9c9bae7f36a237c7e1b96de3d4c2186d13967686
PARTIALLY REPRODUCED · Root-cause confidence: high for runtime stall; medium for complete incident
1. TL;DR
An accepted provider request that never emits a start event remains pending after the watchdog warns. A subsequent runtime turn request still fails with CompetingTurnError. This was reproduced twice using the repository's scripted provider bridge, without changing production code. Server source explains how active threads without a turn ID park input until a start arrives. The reported reconnect incident, actual Claude behavior, persisted queue and Working display were not reproduced end to end.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| Start timeout does not settle pending runtime work | Verified | Two failing recovery tests; zero starts/completions and one warning |
| Active thread without turn ID parks input | Verified by source only | Transactional queue guard and turn-start dispatch |
| Reconnect revives interrupted thread | Unverified incident; plausible source path | Reconciliation maps errored threads reported active to run.started |
| UI Working and durable stranded queue | Unverified dynamically | No real UI or database incident reproduction performed |
3. Environment
Linux 7.2.5-3-omarchy x86_64; Node v26.8.1; pnpm frozen install; Turbo build: 60 successful tasks in each checkout. Repository scripted echo bridge with swallowTurnStart enabled. No real provider CLI, user data, network service ports or live BB instance used. Each test allocates and removes a fresh temporary workspace. Watchdog threshold 120 ms and polling interval 25 ms replace the production 120-second and 15-second defaults.
4. Minimal reproduction
In a clean checkout apply the linked test-only diff and run:
git checkout --detach 9c9bae7f36a237c7e1b96de3d4c2186d13967686 pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build git apply /path/to/reproduction.diff pnpm exec turbo run test --filter=@bb/agent-runtime --force -- src/runtime.turn-start-watchdog.test.ts
The fixture accepts a turn request but intentionally suppresses its start notification. After the warning, wait another 240 ms and attempt another runTurn. Expected under automatic recovery: retry is accepted. Actual:
timeout recovery {"warnings":1,"starts":0,"completions":0}
AssertionError: promise rejected "CompetingTurnError: Refusing to start a c…" instead of resolving
CompetingTurnError: Refusing to start a competing turn for thread "t1" while another turn is active or starting
Tests 1 failed | 1 passed (2)This assertion tests the proposed recovery behavior, not an existing guarantee that a timeout proves the provider stopped. It demonstrates the blocked state; it does not justify clearing the pending flag alone.
Test diff · Complete test · First log · Second log
Complete reproduction test
import { mkdtempSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { afterEach, beforeEach, describe, expect, it } from "vitest";
import type { ThreadEvent } from "@bb/domain";
import {
createScriptedEchoRuntime,
fullRuntimeOptions,
wait,
type LaunchBoundAgentRuntime,
} from "./test/runtime-test-harness.js";
import { promptTextInput } from "./test/prompt-input.js";
describe("turn-start watchdog", () => {
let tmpDir: string;
let runtime: LaunchBoundAgentRuntime | null = null;
beforeEach(() => {
tmpDir = mkdtempSync(join(tmpdir(), "bb-runtime-watchdog-"));
});
afterEach(async () => {
await runtime?.shutdown();
runtime = null;
rmSync(tmpDir, { recursive: true, force: true });
});
async function waitFor<T>(
resolve: () => T | undefined,
timeoutMs = 3_000,
): Promise<T> {
const deadline = Date.now() + timeoutMs;
for (;;) {
const value = resolve();
if (value !== undefined) {
return value;
}
if (Date.now() > deadline) {
throw new Error("timed out waiting");
}
await wait(15);
}
}
it("surfaces a visible error when an accepted turn never starts", async () => {
const events: ThreadEvent[] = [];
runtime = createScriptedEchoRuntime({
runtime: {
workspacePath: tmpDir,
onEvent: (event) => events.push(event),
turnStartWatchdog: { thresholdMs: 120, intervalMs: 25 },
},
launch: { scripted: { swallowTurnStart: true } },
});
await runtime.startThread({
environmentId: "env-1",
threadId: "t1",
projectId: "p1",
providerId: "fake",
options: fullRuntimeOptions,
});
await runtime.runTurn({
clientRequestId: "creq_222222222w",
threadId: "t1",
input: [promptTextInput({ text: "hello" })],
options: fullRuntimeOptions,
});
const watchdogEvent = await waitFor(() =>
events.find(
(event) =>
event.type === "system/error" &&
event.code === "provider_turn_start_timeout",
),
);
expect(watchdogEvent.threadId).toBe("t1");
await wait(240);
console.log("timeout recovery", JSON.stringify({
warnings: events.filter(event => event.type === "system/error" && event.code === "provider_turn_start_timeout").length,
starts: events.filter(event => event.type === "turn/started").length,
completions: events.filter(event => event.type === "turn/completed").length,
}));
await expect(runtime.runTurn({
clientRequestId: "creq_222222222y",
threadId: "t1",
input: [promptTextInput({ text: "retry" })],
options: fullRuntimeOptions,
})).resolves.toBeUndefined();
expect(events.some((event) => event.type === "turn/started")).toBe(false);
await wait(120);
expect(
events.filter(
(event) =>
event.type === "system/error" &&
event.code === "provider_turn_start_timeout",
),
).toHaveLength(1);
});
it("stays silent when the turn starts within the threshold", async () => {
const events: ThreadEvent[] = [];
runtime = createScriptedEchoRuntime({
runtime: {
workspacePath: tmpDir,
onEvent: (event) => events.push(event),
turnStartWatchdog: { thresholdMs: 150, intervalMs: 25 },
},
});
await runtime.startThread({
environmentId: "env-1",
threadId: "t1",
projectId: "p1",
providerId: "fake",
options: fullRuntimeOptions,
});
await runtime.runTurn({
clientRequestId: "creq_222222222x",
threadId: "t1",
input: [promptTextInput({ text: "hello" })],
options: fullRuntimeOptions,
});
await waitFor(() => events.find((event) => event.type === "turn/started"));
await wait(250);
expect(
events.some(
(event) =>
event.type === "system/error" &&
event.code === "provider_turn_start_timeout",
),
).toBe(false);
});
});
5. Root cause
packages/agent-runtime/src/runtime.ts:279 keeps pending starts in a map. The watchdog only marks watchdogFired and emits system/error; it does not settle or cancel the request. packages/agent-runtime/src/runtime.ts:788 rejects competing requests whenever that map still contains the thread. packages/agent-runtime/src/runtime.ts:840 clears pending starts on lifecycle events, which the stalled fixture never sends.
apps/server/src/services/threads/thread-turn-starting.ts:29 parks input when durable status is active and the active turn ID is null. apps/server/src/services/threads/queued-message-dispatch.ts:298 retries those waits on a start event. Idle recovery at line 511 visits idle threads, so an active thread does not qualify. These source findings support the reported queue mechanism but are not an integrated server reproduction.
apps/server/src/services/threads/thread-lifecycle.ts:2087 reconciles errored threads reported active by the host with run.started. Whether the original host report was stale is unknown.
6. Proposed fix
Define a timeout recovery policy that safely cancels or fences the accepted provider request before accepting another turn, reconciles durable server status, and preserves queued input in an actionable state. Test delayed start after timeout, late completion, reconnect and manual stop. Simply deleting pendingTurnStarts risks concurrent provider turns. No automated PR: the full incident is only partially reproduced and recovery requires a policy decision across runtime and server subsystems.
7. Related issues
The issue identifies adjacent timeout and connection-loss reports. Their incident claims were not used as executable evidence. GitHub metadata and an open-PR search found no PR linked to #4365 during this investigation.
8. Verification
The same agent created a second clean detached checkout at the full base commit above, performed a separate frozen install and build, copied only the test change and ran the same Turbo test with --force. It again failed at the retry assertion with CompetingTurnError, while the normal-start control passed. Both runs showed one warning, zero starts and zero completions. No report correction was needed. This is repeated verification, not independent verification.
9. Appendix
Initial setup encountered /tmp inode exhaustion; that failed setup was excluded from evidence. The report clone and second checkout were moved to /var/tmp. The first successful run and fresh second run are linked above; absolute local paths are sanitized. No UI screenshot is supplied because only a nonvisual runtime mechanism was reproduced. No production code, dependency or persistent data was changed.
Issue content and comments were treated as untrusted claims. No linked external content, issue scripts or pull-request code was executed.
AGENT GENERATED
10. Additional server verification · 2026-09-30
Trusted fetched main: facb6c161d9915d2517ed15f6101e4295cf85da1. PARTIALLY REPRODUCED, now with real persisted-queue and conditional reconnect-reconciliation evidence. High confidence in the tested server transitions; the original incident and visual Working state remain unverified. Sections 1–9 above preserve the earlier investigation at its own SHA. This addition does not rerun or supersede its runtime-watchdog test.
New expected-versus-actual evidence
Two synthetic server scenarios start with durable active status: one has a previously started turn, the other has no turn-start event. Both call the real host-interruption service, then reconcile a controlled daemon report that still lists the thread active. In the previously started case, the server writes an interrupted turn completion and clears the active turn ID. Reconciliation makes the thread active again while the active turn ID stays null. The real queue-input guard then persists the follow-up with waitingOn.kind=turn-starting. Three explicit idle-recovery passes leave that row and its text intact. Both scenarios repeat identically in the second clean checkout.
Expected recovery goal: after an interrupted turn, a stale active report should not leave follow-up input waiting indefinitely for a start that may never arrive; recovery needs an explicit reconciliation/cancellation policy. Actual under the tested report input: the active/no-turn combination queues input and is not selected by idle recovery. This is conditional evidence: the test supplies the active report and does not prove the original daemon actually sent a stale report.
Control: before interruption, a thread with an actual active turn returns retry from the starting-turn guard and creates no starting-wait queue row. That retry means the guard declines this queueing path; it is not a tested provider dispatch. The test verifies preserved queued content in SQLite rather than substituting a mock database.
Implementation and proposed next step
Host interruption updates the synthetic active thread through the real service. Reconciliation of errored threads reported active applies run.started without a provider turn-start event in this fixture. The transactional starting-turn guard records a wait when durable status is active and no active turn ID exists. Idle recovery visits idle candidates; the active thread remains outside that recovery path. The test dynamically exercises this interaction, strengthening the earlier source-only queue finding.
Next fix/design step: reconcile reported activity with accepted/started/terminal turn identity, and safely fence or cancel stale work before re-dispatch. Preserve queued text and test delayed starts, reconnect ordering and manual stop. Related #4392 was open and unmerged during this addition; its branch and patch were not run or adopted. No production fix or PR was created.
Environment and exact repeatable steps
Linux x86_64, Node 24.19.0, pinned pnpm 9.15.0, Vitest 4.1.1. Two clean detached checkouts, separate frozen installations/caches and fresh harness SQLite/temporary state for each scenario. Dependency downloads were shared. Forced Turbo test graphs executed upstream tasks. No live provider, agent, host connection or user runtime was used. Synthetic host/session records support the real server services; interruption and reconciliation are called directly rather than through an actual network disconnect.
Save the complete test below in both checkouts at apps/server/test/threads/issue4365.test.ts. The normalized commands are the executed install/test sequence.
WORK=$(mktemp -d) git clone https://github.com/get-bb/bb.git "$WORK/base" git -C "$WORK/base" worktree add --detach "$WORK/first" facb6c161d9915d2517ed15f6101e4295cf85da1 git -C "$WORK/base" worktree add --detach "$WORK/second" facb6c161d9915d2517ed15f6101e4295cf85da1 # Save the test below in both checkouts. cd "$WORK/first" npm_config_cache="$WORK/npm-first" npm_config_devdir="$WORK/node-gyp-first" XDG_CACHE_HOME="$WORK/cache-first" pnpm install --frozen-lockfile --store-dir "$WORK/dependency-store" pnpm exec turbo run test --filter=@bb/server --force -- issue4365.test.ts --silent=false # Same-agent second clean reproduction: cd "$WORK/second" npm_config_cache="$WORK/npm-second" npm_config_devdir="$WORK/node-gyp-second" XDG_CACHE_HOME="$WORK/cache-second" pnpm install --frozen-lockfile --store-dir "$WORK/dependency-store" pnpm exec turbo run test --filter=@bb/server --force -- issue4365.test.ts --silent=false
import {getThread,listEvents,listQueuedThreadMessages} from "@bb/db";
import {it,expect} from "vitest";
import {withTestHarness} from "../helpers/test-app.js";
import {seedThreadFixture,seedThreadRuntimeState,seedTurnStarted} from "../helpers/seed.js";
import {interruptActiveThreadsForHost,reconcileDaemonReportedThreads} from "../../src/services/threads/thread-lifecycle.js";
import {getActiveTurnId} from "../../src/services/threads/thread-events.js";
import {queueInputForStartingTurn} from "../../src/services/threads/thread-turn-starting.js";
import {runQueuedMessageDispatch} from "../../src/services/threads/queued-message-dispatch.js";
import type {QueuedDispatchMessage} from "../../src/services/threads/queue-waits.js";
const message:QueuedDispatchMessage={input:[{type:"text",text:"Synthetic queued follow-up"}],execution:{model:"gpt-5",reasoningLevel:"medium",permissionMode:"full",serviceTier:"default",source:"client/turn/requested"},senderThreadId:null,origin:null,originPluginId:null,requestedBy:null,payload:{kind:"inline"},systemNotice:null};
it.each([false,true])("persists a start wait after active reconnect; prior turn=%s",async(priorTurn)=>{
await withTestHarness(async h=>{
const {host,thread,environment}=seedThreadFixture(h,{thread:{status:"active"}});
seedThreadRuntimeState(h.deps,{threadId:thread.id,environmentId:environment.id,providerThreadId:"synthetic-provider"});
if(priorTurn)seedTurnStarted(h.deps,{threadId:thread.id,environmentId:environment.id,turnId:"synthetic-turn"});
const before=getActiveTurnId(h.deps,thread.id);
if(priorTurn){const control=queueInputForStartingTurn(h.deps,{claimed:null,input:message,threadId:thread.id});expect(control.kind).toBe("retry");expect(listQueuedThreadMessages(h.db,thread.id)).toHaveLength(0);}
interruptActiveThreadsForHost(h.deps,{hostId:host.id,reason:"host-daemon-restarted",cause:"host-connection-lost"});
const interruptedStatus=getThread(h.db,thread.id)?.status;expect(getActiveTurnId(h.deps,thread.id)).toBeNull();
await reconcileDaemonReportedThreads(h.deps,{hostId:host.id,activeThreadIds:[thread.id],undeliveredEventThreadIds:[],sameDaemonInstance:true});
const status=getThread(h.db,thread.id)?.status;expect(status).toBe("active");expect(getActiveTurnId(h.deps,thread.id)).toBeNull();
const queued=queueInputForStartingTurn(h.deps,{claimed:null,input:message,threadId:thread.id});expect(queued.kind).toBe("queued");
for(let n=0;n<3;n++)await runQueuedMessageDispatch(h.deps,{kind:"idle-recovery",now:Date.now()+600000+n*600000});
const rows=listQueuedThreadMessages(h.db,thread.id);expect(rows).toHaveLength(1);expect(JSON.parse(rows[0].waitingOn!)).toEqual({kind:"turn-starting"});expect(JSON.parse(rows[0].content)).toEqual(message.input);
const events=listEvents(h.db,{threadId:thread.id});const completions=events.filter(e=>e.type==="turn/completed").map(e=>JSON.parse(e.data).status);if(priorTurn)expect(completions).toContain("interrupted");
console.log("ISSUE4365",JSON.stringify({priorTurn,beforeActiveTurn:before!==null,interruptedStatus,reconciledStatus:status,activeTurnAfterReconcile:null,queued:queued.kind,waitingOn:JSON.parse(rows[0].waitingOn!),content:JSON.parse(rows[0].content),recoveryAttempts:3,queueCount:rows.length,completedStatuses:completions}));
});
});
Both-run actual results
The same agent personally ran the identical final test in the second clean checkout at the same SHA. Both runs exited 0 with two tests passing and matching structured outcomes. The assertions describe the undesirable persisted state, so a passing test establishes that state; it does not claim recovery succeeded. This is same-agent second clean reproduction, not independent verification.
first
[
{
"priorTurn": false,
"beforeActiveTurn": false,
"interruptedStatus": "error",
"reconciledStatus": "active",
"activeTurnAfterReconcile": null,
"queued": "queued",
"waitingOn": {
"kind": "turn-starting"
},
"content": [
{
"type": "text",
"text": "Synthetic queued follow-up"
}
],
"recoveryAttempts": 3,
"queueCount": 1,
"completedStatuses": []
},
{
"priorTurn": true,
"beforeActiveTurn": true,
"interruptedStatus": "error",
"reconciledStatus": "active",
"activeTurnAfterReconcile": null,
"queued": "queued",
"waitingOn": {
"kind": "turn-starting"
},
"content": [
{
"type": "text",
"text": "Synthetic queued follow-up"
}
],
"recoveryAttempts": 3,
"queueCount": 1,
"completedStatuses": [
"interrupted"
]
}
]
@bb/server:test: Test Files 1 passed (1) @bb/server:test: Tests 2 passed (2) @bb/server:test: Duration 7.46s (transform 4.64s, setup 1.67s, import 4.75s, tests 895ms, environment 0ms) Tasks: 9 successful, 9 total Cached: 0 cached, 9 total Time: 35.569s
second
[
{
"priorTurn": false,
"beforeActiveTurn": false,
"interruptedStatus": "error",
"reconciledStatus": "active",
"activeTurnAfterReconcile": null,
"queued": "queued",
"waitingOn": {
"kind": "turn-starting"
},
"content": [
{
"type": "text",
"text": "Synthetic queued follow-up"
}
],
"recoveryAttempts": 3,
"queueCount": 1,
"completedStatuses": []
},
{
"priorTurn": true,
"beforeActiveTurn": true,
"interruptedStatus": "error",
"reconciledStatus": "active",
"activeTurnAfterReconcile": null,
"queued": "queued",
"waitingOn": {
"kind": "turn-starting"
},
"content": [
{
"type": "text",
"text": "Synthetic queued follow-up"
}
],
"recoveryAttempts": 3,
"queueCount": 1,
"completedStatuses": [
"interrupted"
]
}
]
@bb/server:test: Test Files 1 passed (1) @bb/server:test: Tests 2 passed (2) @bb/server:test: Duration 7.23s (transform 4.62s, setup 1.63s, import 4.72s, tests 700ms, environment 0ms) Tasks: 9 successful, 9 total Cached: 0 cached, 9 total Time: 34.899s
Remaining limits and corrections
No claim is made about the actual original host report, multi-hour timing, pending-question interaction ordering, provider watchdog behavior on this new SHA, live websocket reconnect, delayed provider events, manual stop, or rendered Working/queue labels. No UI was exercised and no new screenshot is supplied. The recovery calls use future now values to invoke the selected recovery path; they are not elapsed-time measurements and do not simulate every periodic subsystem. End-to-end verification remains incomplete.
Initial installs exhausted inodes; only regenerable dependency directories from completed investigations were removed, preserving all evidence. The installs then succeeded. An initial fixture supplied a half-populated optional requester record; the real database reader rejected it. The corrected fixture uses the normal unset requester value and was rerun identically in both checkouts. These setup failures are excluded from bug evidence.
The old appendix’s four relative raw-artifact links are present in the current tracked reports tree and remain preserved. The earlier test and quoted output also remain inline. This addition includes its own complete test and actual results inline and adds no raw-artifact files.
Issue content, comments and PR content were untrusted claims only. No issue commands, scripts, patches, branches or external URLs were executed or fetched. All records and messages were synthetic; no real user data, credentials or agents were accessed. Raw local evidence is retained outside the public repository.