#4368 · Worker cleanup eligibility
Bug · Medium priority · Low effort · plugins · workflows
2026-09-25 · Trusted main 9c9bae7f36a237c7e1b96de3d4c2186d13967686 · Issue
ALREADY FIXED · Root-cause confidence: high
TL;DR
Cleanup previously held a list whose retirement decisions could become stale while another thread stopped. A newly attached worker could then be interrupted and archived. Main now rechecks the current database state immediately before stopping each candidate. The landed service regression passes in two clean trusted checkouts; this update supersedes the earlier report verdict.
Claims vs findings
| Claim | Finding |
|---|---|
| Cleanup can use a stale retirement decision | Historical mechanism supported by trusted repository history; repaired on current main. |
| Still present on main | No longer current: commit 58b133daed21e7ed2b4e15198395cc733c1e48ab landed the fix. |
| Specific desktop/provider timeline | Unverified; no live provider or macOS run. |
Environment
Linux x86_64; Node v26.8.1. Frozen dependencies and normal Turbo build in both checkouts. Existing SDK host harness with actual migrated SQLite storage. No app servers, provider calls, ports, or real runtime data used.
Minimal reproduction
In each clean checkout of the trusted commit above:
pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build pnpm exec turbo run test --filter=bb-plugin-workflows -- src/service-policy.test.ts
Expected and actual on current main: the worker stays running until completion and is archived afterward. Both runs:
Test Files 1 passed (1) Tests 45 passed (45)
The first attempt encountered temporary filesystem inode exhaustion, unrelated to product behavior. The successful first run used a private temporary directory and Turbo --env-mode=loose --force to pass TMPDIR. The second clean checkout passed with the standard command.
Complete trusted regression file
it("does not retire a worker that attaches while cleanup archives an earlier one", async () => {
const test = setup();
harnesses.push(test.harness);
let releaseSpawn = () => {};
const spawnGate = new Promise<void>((resolve) => {
releaseSpawn = resolve;
});
let releaseStop = () => {};
const stopGate = new Promise<void>((resolve) => {
releaseStop = resolve;
});
let metadata: NonNullable<
Parameters<typeof test.bb.sdk.threads.spawn>[0]["pluginMetadata"]
> = {};
test.harness.sdk.stub(
"threads.spawn",
async (args: Parameters<typeof test.bb.sdk.threads.spawn>[0]) => {
metadata = args.pluginMetadata!;
await spawnGate;
return { id: "zz-live-worker" } as never;
},
);
test.harness.sdk.stub("threads.list", async () =>
metadata.workflowWorker ? ([{ id: "zz-live-worker" }] as never) : [],
);
test.harness.sdk.stub("threads.getPluginMetadata", async () => metadata);
test.harness.sdk.stub(
"threads.stop",
async ({ threadId }: { threadId: string }) => {
if (threadId === "aa-retired-worker") await stopGate;
return { ok: true } as never;
},
);
const stoppedThreads = () =>
test.harness.sdk
.callsTo("threads.stop")
.map(([args]) => (args as { threadId: string }).threadId);
const run = await test.start(source(`return await agent("live work");`));
const controller = new AbortController();
const worker = test.service.runWorker(controller.signal);
try {
await eventually(() =>
expect(
test.db.prepare(`SELECT thread_id FROM workflow_workers`).get(),
).toEqual({ thread_id: "zz-live-worker" }),
);
expiredRunWithWorkers(test.db, "retired-run", ["aa-retired-worker"]);
await eventually(() =>
expect(stoppedThreads()).toContain("aa-retired-worker"),
);
releaseSpawn();
await eventually(() =>
expect(getCall(test.db, run.id, 0)).toMatchObject({
status: "running",
childThreadId: "zz-live-worker",
}),
);
releaseStop();
await eventually(() =>
expect(test.archived).toContain("aa-retired-worker"),
);
await new Promise((resolve) => setTimeout(resolve, 50));
expect(stoppedThreads()).not.toContain("zz-live-worker");
expect(test.archived).not.toContain("zz-live-worker");
expect(getCall(test.db, run.id, 0)?.status).toBe("running");
test.service.onThreadIdle("zz-live-worker", "done");
await eventually(() =>
expect(getRunRequired(test.db, run.id).status).toBe("succeeded"),
);
await eventually(() => expect(test.archived).toContain("zz-live-worker"));
} finally {
releaseSpawn();
releaseStop();
controller.abort();
await worker;
}
});
Root cause
Retirement query includes unattached workers for orphan cleanup. Awaiting an earlier stop lets attachment invalidate that snapshot. Cleanup now rechecks eligibility synchronously with isWorkerRetired before dispatching the stop, closing this gap.
Proposed fix
Already landed: preserve spawn protection and re-evaluate retirement per candidate. No additional change justified.
PR review
#4369 merged at 2026-09-25T22:11:25Z as 58b133daed21e7ed2b4e15198395cc733c1e48ab. Trusted main contains the synchronous guard and shared retirement predicate. Its service tests cover both the attaching worker and retained orphan cleanup. No PR branch was checked out or executed.
Verification
The same agent ran both checks. Checkout one was the clean thread worktree at the recorded SHA. Checkout two was a new detached temporary worktree at the identical SHA. Both builds and all 45 service-policy tests passed. This is service-level verification, not a live provider reproduction. No pre-fix failure was rerun in this update.
Related issues
No additional duplicate established. No open matching PR was found; the repair PR is merged.
Appendix
First test log · Second test log. Older reproduction artifacts remain available but were not executed for this update. Issue commands and links were treated as untrusted; only trusted origin/main code was executed.