Latest verification: 2026-09-30: ALREADY FIXED for workflow-service cleanup on current main. Historical report below is preserved; current provider process and RAM behavior remain unverified.

← reports

#2284 · Workflows plugin: finished unstructured worker threads are never stopped, leaking idle provider processes

Bug Priority: High Effort: Small perf threads workflows open on GitHub 2026-08-24 · base 494f66526

Verdict: REPRODUCED · Root-cause confidence: high

1. TL;DR

A bb workflow (the built-in workflows plugin) runs each agent() call in a hidden child thread. When the child is a plain, unstructured call (no outputSchema), the plugin records its final text as the call result, marks the call succeeded, wakes the orchestrator, and then does nothing with the child thread: handleThreadIdle in plugins/workflows/src/service.ts never calls stopChild(threadId) on that path, while the structured-result path (submitStructuredResult) does. An idle hidden thread keeps its provider session loaded on the host daemon, and for ACP providers that session is one OS process per thread (opencode acp here, 500–850 MB RSS each). Nothing else in the plugin stops those children later either: run settlement, cancellation and plugin shutdown only stop calls that are still queued/running. The daemon's 30-minute idle reaper does not help: under default settings (providerSessionReaping experiment off) it only ever releases Codex sessions. I reproduced this on a dev instance with the exact provider from the issue: a 6-call batch limited to 2 concurrent agents held 6 live opencode acp processes (3.4 GB) while only 1–2 calls were running, and bb thread stop on a finished worker released its process immediately without touching the running ones, exactly as the reporter described. The omission has existed since the plugin was introduced (#733). PR #2090 is not a fix for this issue.

2. Claims vs findings

Claim from the issueStatusEvidence
Unstructured workers are settled succeeded in handleThreadIdle and wake the orchestrator, but stopChild(threadId) is never called.VerifiedCode: service.ts#L1017-L1026. Unit test issue-2284-idle-worker-stop.test.ts fails on base: threads.stop is called 0 times for each of three unstructured workers (section 4, step 1). Host-daemon log of the live run shows 3 thread.start commands and no thread.stop.
The code excerpt in the issue (from dist/server.js) reflects the real source.VerifiedIdentical control flow in plugins/workflows/src/service.ts at base (structured branch at L1086-L1091 calls await stopChild(threadId); unstructured else at L1017-L1023 does not).
BB keeps provider connections alive while threads remain open, so completed worker threads retain their subprocess indefinitely.Verified (with one nuance)Live: finished workers' opencode acp processes stayed alive for as long as the batch kept running (runs 2 and 3). No plugin or server code path ever stops a succeeded call's child; the daemon reaper skips non-Codex providers by default (runtime.ts#L1007-L1017, experiments.ts#L34). Nuance: in my runs the leaked processes did die a few seconds after the run ended, because the origin thread's completion-notification turn made the daemon swap the (by then fully idle) environment runtime for a different skill catalog (section 5, "what finally killed them"). That can only happen once no worker is active, so it cannot relieve a batch in progress, and it is not a plugin action.
With concurrency 4 and 40 completed tasks, 48 opencode acp processes (~560 MB RSS each, ~26 GB) were alive.Verified at smaller scaleRun 2 (6 calls, maxConcurrentAgents=2): at 09:25:19 six processes, 3.1–3.4 GB total, with 5 succeeded + 1 running; per-process RSS 460–850 MB (run2-census-timeline.txt). Linear in completed calls: 48 × ~560 MB ≈ 27 GB is consistent.
bb thread stop <thread-id> on the idle workers released the processes without disturbing the active workers.VerifiedRun 3: stopping finished worker thr_7zx3kkwyr6 mid-batch removed only its process; the other finished worker and the two running ones stayed (run3-stop-one-child.out). Server side, a stop of an idle thread is a runtime release, not an interruption (thread-lifecycle.ts#L1393-L1445).
PR #2090 "adds host-daemon idle session reaping with a 30-minute timeout".Partially accurateThe 30-min/5-min reaper already exists at base (app.ts#L65-L66); #2090 only removes the experiment gate so non-Codex restorable sessions are eligible. The reporter's conclusion stands: a 30-minute sweep cannot help a batch that exhausts RAM in minutes.
Proposed fix: call await stopChild(threadId) in handleThreadIdle when settling an unstructured task.Verified sufficient for the reported path; one more gapWith that change the repro test passes and the plugin suite (226 tests) stays green (section 6). The structured fallback path (worker returned valid JSON as text instead of calling bb_workflow_result, L959-L969) has the same omission and needs the same call; the issue does not mention it.

3. Environment

4. Minimal reproduction

Step 1 — unit level (no provider needed): the plugin never stops an unstructured worker

  1. Copy issue-2284-idle-worker-stop.test.ts to plugins/workflows/src/ and run it from the repo root:
    pnpm -C plugins/workflows exec vitest run src/issue-2284-idle-worker-stop.test.ts
    It builds the real createWorkflowService on an in-memory plugin DB (createFakePluginHost from @get-bb/plugin-sdk/testing), starts a run with three sequential agent() calls, feeds each spawned child an idle event with output "result N", and counts threads.stop calls per child.
    expected: each finished worker gets exactly one threads.stop
    actual (base 494f66526):
     ❯ src/issue-2284-idle-worker-stop.test.ts (3 tests | 2 failed)
       × stops an unstructured worker as soon as its call is settled
         AssertionError: expected +0 to be 1   // stopsFor(harness, "child-1") stays 0
       × stops a structured worker settled via the freeform JSON fallback
         AssertionError: expected +0 to be 1
       ✓ control: a worker that submits bb_workflow_result IS stopped (existing behavior)
    Full output: repro-test-base.log. The control test shows the structured tool path already stops its worker, pinning the asymmetry.

Step 2 — live: one leaked opencode acp process per finished call

  1. Start a dev instance from a checkout at 494f66526 and note its URLs/data dir (scripts/bb-dev-app current, then scripts/bb-dev-app env). Make sure opencode is on PATH so the acp-opencode provider is listed (pnpm bb:dev provider list). The workflows plugin ships disabled; switch it on and lower the per-run concurrency so the batch is larger than the concurrency:
    BB_SERVER_URL=http://localhost:21252 bash 2284/repro/plugin-ctl.sh on workflows
    curl -s -X PUT $BB_SERVER_URL/api/v1/plugins/workflows/settings -H 'content-type: application/json' \
      -d '{"values":{"maxConcurrentAgents":"2"}}'
  2. Create a scratch git repo and a project on it (host id from pnpm bb:dev machine list --json), then spawn one cheap origin thread in that project (a workflow needs an origin thread with an environment):
    mkdir -p /tmp/bb-2284-qa && git -C /tmp/bb-2284-qa init -q && touch /tmp/bb-2284-qa/README.md \
      && git -C /tmp/bb-2284-qa add -A && git -C /tmp/bb-2284-qa -c user.name=qa -c user.email=qa@example.com commit -qm init
    curl -s -X POST $BB_SERVER_URL/api/v1/projects -H 'content-type: application/json' \
      -d '{"name":"qa","source":{"type":"local_path","path":"/tmp/bb-2284-qa","hostId":"<host id>"}}'
    pnpm bb:dev thread spawn --project <project id> --provider acp-opencode --model opencode/big-pickle \
      --permission-mode full --title "origin 2284" --prompt "Reply only with ok." --json      # note the thread id
    pnpm bb:dev thread wait <origin thread id>
  3. Start the batch workflow (batch-workflow.js: six unstructured agent("Reply only with the word ok.") calls on acp-opencode inside parallel()) through the plugin CLI endpoint — the same path as bb workflows run --script from inside a thread — and start a 3-second process census next to it:
    export BB_SERVER_URL=http://localhost:21252 ORIGIN_THREAD=<origin thread id> PROJECT_ID=<project id>
    bash 2284/repro/run-workflow.sh 2284/repro/batch-workflow.js
    # → {"exitCode":0,"stdout":"{\"runId\":\"wfr_066a2bfe-05bf-4907-a86e-756005690469\",\"name\":\"issue-2284-batch\",\"status\":\"queued\"}\n"}
    bash 2284/repro/census-loop.sh wfr_066a2bfe-05bf-4907-a86e-756005690469 run2-census-timeline.txt 3 150
    census.sh lists every live opencode acp process with RSS and the BB_THREAD_ID the daemon put in its environment (macOS ps -E).
  4. Compare the census with the run's call counters at the same instant.
    expected: live opencode processes == calls still RUNNING (≤ 2)
    actual (2284/repro/run2-census-timeline.txt):
    === 09:25:19 run: running calls={"total":6,"queued":0,"running":1,"succeeded":5,"failed":0,"cancelled":0}
    83691   499536    00:57      thr_bmn6hk7iwf     opencode acp
    83699   491104    00:57      thr_5ydc7srqsd     opencode acp
    84676   602032    00:46      thr_2wu7iq3cqq     opencode acp
    84847   645072    00:45      thr_avb9qm9vun     opencode acp
    86074   463408    00:34      thr_ggpgqttx3h     opencode acp
    86198   468224    00:34      thr_qvjnz4b6ei     opencode acp
    total: 6 process(es), 3095 MB RSS
    Five calls are finished (their rows in the plugin DB are succeeded with finished_at set — run2-workflow-calls.txt) and their threads are idle/hidden (run2-threads.txt), yet all six agent processes are resident. The host-daemon log for the whole session contains only thread.start commands, never a thread.stop.
  5. (The reporter's workaround.) Run 3 repeats the batch and, as soon as two calls have succeeded while two are running, stops one finished worker with the CLI (stop-one-child.sh):
    09:28:06 run status: succeeded=2 running=2; stopping finished worker thr_7zx3kkwyr6
    --- census BEFORE stop:
    4639    632320    00:17      thr_7zx3kkwyr6     opencode acp      ← finished (succeeded)
    4764    627024    00:15      thr_utph6c3kdc     opencode acp      ← finished (succeeded)
    6752    336784    00:01      thr_qzdenzmj3p     opencode acp      ← running
    7065    47824     00:00      thr_32ugrpy25v     opencode acp      ← running
    Thread thr_7zx3kkwyr6 stopped
    --- census AFTER stop (09:28:09):
    4764    626768    00:18      thr_utph6c3kdc     opencode acp      ← still leaked
    6752    775104    00:04      thr_qzdenzmj3p     opencode acp
    7065    564736    00:03      thr_32ugrpy25v     opencode acp
    bb thread stop on an idle thread is a pure runtime release, so the stopped worker's process is gone and the running workers are untouched; the other finished worker stays leaked because nothing asks for it to be released.
The origin thread in the dev app after runs 1 and 2, showing the two workflow completion notifications
The origin thread (claude-code) in the dev app after runs 1 and 2. The bug is not visual; this shows where the workflow completion notifications land. The prose under each notification is the origin thread's own agent replying to the notification (it ran its own ad-hoc checks; its remarks about a "bridge-worker pool" are its guesses, not evidence used here).

Repro files: 2284/repro/ (workflow scripts, helper shell scripts, the vitest file, census timelines for all three runs, plugin-DB and thread-table dumps, proposed-fix diff).

5. Root cause

The omission. handleThreadIdle is the only place an unstructured call is settled. Its else branch records the worker's final text and wakes the waiter, and returns without releasing the child thread (service.ts#L914-L1026):

    } else {
      settleCall(db, {
        id: call.id,
        status: "succeeded",
        result: output ?? "",
        error: null,
      });
    }
    wakeCall(call);
  }                                   // ← no stopChild(threadId)

Every terminal path for a structured call does release the worker: submitStructuredResult (L1086-L1095) and the repair-exhausted / correction-failed branches of handleThreadIdle (L984-L988, L1012-L1016). Two paths do not: the unstructured else above, and the structured fallback that accepts valid JSON from the worker's plain text (L959-L970). git log -S stopChild shows the asymmetry has been there since the plugin landed in #733 (57a027724); the only later change (#826, ef8c867a6) removed a stall watchdog that used to stop stalled workers.

Why nothing cleans up later. Run-level cleanup only targets calls that are still open: settleRunAndStopOutstanding stops the children of settleRun's returned rows, which are queued/running only (service.ts#L1206-L1220, data.ts#L343-L372); stop(runId) uses activeChildThreadsForRun, same filter (data.ts#L773-L783); shutdown/recovery uses recoverInterruptedRuns, also running only. A child whose call already reads succeeded is invisible to all of them.

Why an idle hidden thread costs a whole process. The server keeps a thread's runtime loaded after its turn; only an explicit stop releases it (stopThreadForCurrentState → releaseIdleThreadRuntime → daemon thread.stop {intent:"release"}, thread-lifecycle.ts#L1393-L1473). In the ACP bridge each bb thread owns one spawned agent process (sessionsByBbThreadId, bridge.ts#L202-L203; release path releaseSession → connection.kill(), L2088-L2104). Claude Code and Pi are also one process per thread. The daemon's periodic reaper (every 5 min, sessions idle ≥ 30 min, app.ts#L65-L66) considers non-Codex sessions only when the providerSessionReaping experiment is on (runtime.ts#L1007-L1017), and it defaults to off (experiments.ts#L34) — that is issue #1604. So for the reporter's acp-opencode workers the only thing that ever releases a finished worker is a human running bb thread stop.

What finally killed the processes in my runs (not a fix, but worth knowing). In all three runs every leaked process disappeared 1–3 s after the run reached succeeded. The host-daemon log shows the cause each time (run2-devlog-at-end.txt): the plugin's completion notification starts a turn on the origin thread, which carries a different injected-skill catalog than the workers (workers get skills: [] from bb.agents.configure). The daemon's RuntimeManager then replaces the environment's runtime to switch catalogs, which is allowed only when the environment has no active runtime work, and shuts the old runtime down — taking the ACP bridge and all its idle agent subprocesses with it (runtime-manager.ts#L765-L784, "Removing unused injected skill staging directory" at 09:18:20 / 09:25:20 / 09:28:33). That swap cannot happen while any worker is still running, so it never relieves a batch in progress, and it is incidental: a workflow whose origin is not notified, or whose completion coincides with other activity in the environment, keeps the processes. It also explains why the reporter saw the leak "resolve" only by stopping threads manually.

Deeper issue. The plugin SDK has no "this worker is done, release it" primitive, and the server keeps sessions of hidden plugin-owned threads loaded with the same policy as user-visible threads. The workflows plugin therefore has to remember to stop every child itself, and it forgot the most common path. A server-side policy (release the runtime of a hidden, plugin-originated thread as soon as it goes idle, unless its owner keeps it) would remove the class of bug; that is a larger change than this issue needs.

6. Proposed fix (first principles)

Stop the worker on the two settle paths that forget it. Minimal patch against base (proposed-fix.diff):

--- a/plugins/workflows/src/service.ts
+++ b/plugins/workflows/src/service.ts
@@ -966,6 +966,9 @@ export function createWorkflowService(
             error: null,
           });
           wakeCall(call);
+          // The worker is finished; release its provider session so the host
+          // does not keep one idle agent process per completed call (#2284).
+          await stopChild(threadId);
           return;
         }
       }
@@ -1021,6 +1024,11 @@ export function createWorkflowService(
         result: output ?? "",
         error: null,
       });
+      wakeCall(call);
+      // Unstructured calls settle here and nowhere else: stop the finished
+      // worker so its provider process is released (#2284).
+      await stopChild(threadId);
+      return;
     }
     wakeCall(call);
   }

Verification: with the patch applied, the repro test passes (3/3, repro-test-with-fix.log) and the whole workflows plugin suite passes (14 files, 226 tests, workflows-suite-with-fix.log). The patch was reverted afterwards; the test file is left in the worktree for the fix PR to adopt.

What to watch: wakeCall must stay before stopChild so the orchestrator continues while the release round-trips to the daemon (the handler task is tracked in handlerTasks, so shutdown still awaits it). stopChild tolerates a thread that is already gone (isMissingThread) and dedupes concurrent stops, so calling it from the idle handler cannot double-stop the structured-tool path (submitStructuredResult already stopped that worker; its later idle event hits the call.resultJson !== null branch, which does not call it again). Stopping an idle thread is a release, not an interruption, so the worker's timeline and final output remain readable for bb workflows history. One behavioral change: a finished worker can no longer be "talked to" cheaply (its session is unloaded, resuming restarts the agent), which is the intended trade-off. Keep the workers' sessionRestorable semantics in mind if someone later adds a "retry last call in the same thread" feature. Also consider a belt-and-braces stopChildren of all children when a run settles (not only queued/running), so a future settle path cannot reintroduce the leak.

7. PR review

#2090 · Graduate idle provider session reaping (ScaleLeanChris, branch fix/provider-session-reaping-default, head d73645726)

What it changes. Removes the providerSessionReaping experiment (domain key, DB default, web/mobile/desktop settings UI, docs, fixtures), removes the server's /internal/runtime-policy endpoint and the daemon's getRuntimePolicy, makes the daemon's 30-minute reaper consider every provider session that reports sessionRestorable (Codex keeps its legacy thread-scoped fallback), keeps the open-work guards (backgroundWorkState.hasOpenThreadWork, adapter hasOpenThreadWork), and bumps HOST_DAEMON_PROTOCOL_VERSION 146 → 147. 31 files, +34/−241. It claims Fixes #1604; the reporter of #2284 commented on it but it is not linked as a fix for #2284.

Does it address this issue's root cause? No. It does not touch plugins/workflows. At best it turns "never released" into "released 30–35 minutes after the worker went idle", and only for ACP agents whose bridge reports sessionRestorable (ACP sets it from the agent's session/load support, bridge.ts#L2006; whether opencode acp advertises it was not verified here). A 40-task batch at concurrency 4 finishes its first 36 idle workers in far less than 30 minutes, so RAM is exhausted before the first sweep could act. The reporter's assessment in the issue is correct.

Findings.

WhereSeverityFinding
packages/host-daemon-contract/src/protocol.ts:118HighBumps the protocol to 147, but main is already at 164 (base 494f66526). The branch is 82 commits behind (merge base c942421a4) and git merge-tree reports content conflicts in runtime.ts, runtime.process-lifecycle.test.ts, protocol.ts and contract.test.ts (GitHub: mergeable: CONFLICTING). After a rebase the bump must become 165 with a fresh changelog comment, and the runtime diff must be re-derived against the post-#2325 provider-plugin runtime, not just conflict-resolved.
packages/agent-runtime/src/runtime.ts (reaper gating)MediumBehavior change without an off switch: every restorable Claude Code / Pi / ACP session is now unloaded after 30 idle minutes for every installation, and the PR deletes the only knob. #1604's own report suggested flipping the default rather than deleting the setting. Silent for users who rely on warm sessions; at minimum the release note / docs should say so, and a config knob (not an "experiment") to disable or tune the window would be prudent.
packages/domain/src/experiments.ts + experimentsSchema (strict key enum)Medium-lowRemoving the key makes PUT /api/v1/settings/experiments reject any client that still sends providerSessionReaping (z.record over experimentKeySchema). The web app is served by the same server so it is fine; an older installed mobile app build (separately released) sending the key will get 400 on every experiments save until updated. Reading persisted rows is safe (getExperiments filters by known keys).
apps/host-daemon/src/app.ts / server-client.tsOKRemoving /internal/runtime-policy together with the protocol bump is correct: an older daemon is forced to update instead of looping on the missing endpoint.
TestsOK on its own baseChecked out in my worktree and ran pnpm exec turbo run test --filter=@bb/agent-runtime --filter=@bb/host-daemon --filter=@bb/host-daemon-contract: 7 tasks successful (pr-2090-tests.log). The added lifecycle regression ("reaps a restorable non-Codex session") is meaningful. Nothing in the PR exercises a hidden workflow worker or the burst scenario of #2284.
LayeringOKPolicy stays where it was (daemon reaper, server no longer distributes a toggle). No casts or unknown smuggling introduced.

Verdict: REQUEST CHANGES — as a #1604 change it needs a rebase onto current main, a protocol bump to 165, and a decision about an opt-out; as a response to #2284 it should not be considered at all. #2284 needs the one-line plugin fix in section 6.

8. Related issues

9. Appendix

Commands run (in order)

gh issue view 2284 --repo get-bb/bb --json ...            # issue text, no comments
pnpm install --frozen-lockfile --prefer-offline; pnpm exec turbo run build
pnpm -C plugins/workflows exec vitest run src/issue-2284-idle-worker-stop.test.ts   # fails on base
git fetch origin main; git log 494f66526..origin/main -- plugins/workflows           # empty
gh pr view 2090 --json ...; gh pr diff 2090 > 2284/pr-2090.diff
scripts/bb-dev-app current; scripts/bb-dev-app env
pnpm bb:dev provider list --json; pnpm bb:dev machine list --json; pnpm bb:dev project list --json
curl $BB_SERVER_URL/api/v1/system/config                   # experiments.providerSessionReaping=false
bash 2284/repro/plugin-ctl.sh on workflows
bash 2284/repro/run-workflow.sh 2284/repro/leak-workflow.js         # run 1: wfr_91201162-…
bash 2284/repro/wait-for-run.sh wfr_91201162-…; bash 2284/repro/census.sh (mid-run and after)
sqlite3 <data>/bb.db "select id,status,provider_id,visibility,title from threads"
sqlite3 <data>/plugins/workflows/data.db "select call_index,status,child_thread_id,… from workflow_calls where run_id=…"
curl -X PUT $BB_SERVER_URL/api/v1/plugins/workflows/settings -d '{"values":{"maxConcurrentAgents":"2"}}'
bash 2284/repro/run-workflow.sh 2284/repro/batch-workflow.js        # run 2: wfr_066a2bfe-…
bash 2284/repro/census-loop.sh wfr_066a2bfe-… run2-census-timeline.txt 3 150
pnpm bb:dev thread spawn --provider acp-opencode --model opencode/big-pickle … # cheap origin for run 3
bash 2284/repro/run-workflow.sh 2284/repro/batch-workflow.js        # run 3: wfr_66ae9e12-…
bash 2284/repro/stop-one-child.sh wfr_66ae9e12-…                      # bb thread stop on a finished worker
doobie --headless < 2284/repro/screenshot-origin.js                 # assets/2284-origin-thread.png
pnpm dev:stop
(apply proposed-fix.diff) pnpm -C plugins/workflows exec vitest run    # 226 passed; then git checkout -- service.ts
git fetch origin pull/2090/head:pr-2090; git merge-tree --write-tree 494f66526 pr-2090   # 4 CONFLICTs
git checkout pr-2090; pnpm install; pnpm exec turbo run test --filter=@bb/agent-runtime --filter=@bb/host-daemon --filter=@bb/host-daemon-contract
git checkout worktree-wf_846839f8-f8a-2; pnpm install; rm -rf <data dir> /tmp/bb-2284-qa

Run 1 (sequential, 3 calls) — census while call 3 was running

PID     RSS_KB    ELAPSED    BB_THREAD_ID       COMMAND
65323   847552    00:48      thr_s44tzchsc5     opencode acp      ← call 0 succeeded 09:17:58
65876   656528    00:21      thr_zciq2rj2ua     opencode acp      ← call 1 succeeded 09:18:09
66895   588464    00:10      thr_ixm4x5q5zy     opencode acp      ← call 2 running
total: 3 process(es), 2043 MB RSS

Run status afterwards: run-final-status.json ("result":["ok","ok","ok"]). Host-daemon RPC command types for the session (from dev.log): 4 host.read_file, 3 thread.start, 1 provider.list_models — no thread.stop.

Run 2 — dev.log at the moment the processes vanished

[09:25:20] DEBUG: [host-daemon] Online host RPC {"commandType":"host.read_file","errorCode":"ENOENT",…}
[09:25:20] DEBUG: [host-daemon] Removing unused injected skill staging directory {"catalogHash":"5b68f779…","stagingRootPath":"…/runtime/global-skills"}
[09:25:20] DEBUG: [host-daemon] Using cached host artifact {"cacheDir":"…/plugin-host-artifacts/provider-claude-code",…}

The same three lines appeared at 09:18:20 (run 1) and at the end of run 3 (origin on opencode): the origin's notification turn starts, ensureCompatibleEntry sees a different skill catalog on an environment with no active work, replaceEntryForSkillCatalog shuts the old runtime down (and its ACP agent subprocesses with it), and the new runtime launches the origin's provider bridge.

PR 2090 merge-tree against base

CONFLICT (content): Merge conflict in packages/agent-runtime/src/runtime.process-lifecycle.test.ts
CONFLICT (content): Merge conflict in packages/agent-runtime/src/runtime.ts
CONFLICT (content): Merge conflict in packages/host-daemon-contract/src/protocol.ts
CONFLICT (content): Merge conflict in packages/host-daemon-contract/test/contract.test.ts

Repro test (inline)

// Repro for get-bb/bb#2284: a workflow worker that finishes an UNSTRUCTURED
// agent() call (no outputSchema) is settled as `succeeded` in
// handleThreadIdle, but the plugin never calls `threads.stop` for it, so the
// hidden worker thread keeps its provider process alive until something else
// (run cancellation, plugin shutdown, or the daemon's idle reaper) stops it.
//
// Structured calls that come back through `bb_workflow_result`
// (submitStructuredResult) DO call stopChild; this test pins the gap.
import { createFakePluginHost } from "@get-bb/plugin-sdk/testing";
import { afterEach, describe, expect, it } from "vitest";
import { getCall, getRunRequired, migrations } from "./data.js";
import { createWorkflowService } from "./service.js";
import { DEFAULT_WORKFLOW_SETTINGS } from "./settings.js";

async function eventually(assertion: () => void | Promise<void>, timeoutMs = 4_000): Promise<void> {
  const deadline = Date.now() + timeoutMs;
  while (true) {
    try { await assertion(); return; }
    catch (error) { if (Date.now() >= deadline) throw error; await new Promise((r) => setTimeout(r, 10)); }
  }
}

function source(body: string, name: string): string {
  return `export const meta = { name: ${JSON.stringify(name)}, description: "issue 2284" };
  ${body}`;
}

function setup() {
  let childCount = 0;
  const { bb, harness } = createFakePluginHost({
    pluginId: "workflows",
    sdk: {
      threads: {
        get: async ({ threadId }) => ({ id: threadId, environmentId: "environment-1", providerId: "acp-opencode",
          status: threadId === "origin" ? "idle" : "active" }) as never,
        output: async () => ({ output: null }),
        defaultExecutionOptions: async () => ({ model: "gpt-test", reasoningLevel: "medium", permissionMode: "full",
          serviceTier: "default", source: "default" }),
        spawn: async () => { childCount += 1; return { id: `child-${childCount}` } as never; },
        send: async () => ({ ok: true }),
        stop: async () => ({ ok: true }),
      },
      providers: {
        list: async () => [{ id: "acp-opencode", displayName: "OpenCode", logoUrl: null, available: true,
          capabilities: { supportsThreadArchive: true, supportsThreadRename: true, supportsServiceTier: true,
            supportsNativeUserQuestion: false, supportsFork: true, permissionModes: ["full"] }, composerActions: [] }],
        models: async () => ({ providers: [], models: [{ id: "gpt-test", model: "gpt-test", displayName: "GPT Test",
          description: "test", supportedReasoningEfforts: [{ reasoningEffort: "medium", description: "test" }],
          defaultReasoningEffort: "medium", isDefault: true }], selectedOnlyModels: [], modelLoadError: null }),
      },
      environments: { get: async () => ({ id: "environment-1", projectId: "project-test", hostId: "host-1", path: "/workspace" }) as never },
    },
  });
  const db = bb.storage.database();
  bb.storage.migrate(db, migrations);
  const service = createWorkflowService(bb, db, DEFAULT_WORKFLOW_SETTINGS);
  return { harness, db, service, childCount: () => childCount };
}

function stopsFor(harness: ReturnType<typeof setup>["harness"], threadId: string): number {
  return harness.sdk.callsTo("threads.stop")
    .filter(([input]) => (input as { threadId: string }).threadId === threadId).length;
}

describe("issue #2284: finished workers release their provider session", () => {
  const harnesses: Array<ReturnType<typeof setup>["harness"]> = [];
  afterEach(async () => { await Promise.all(harnesses.map((h) => h.dispose())); harnesses.length = 0; });

  it("stops an unstructured worker as soon as its call is settled", async () => {
    const test = setup(); harnesses.push(test.harness);
    const run = await test.service.start({ projectId: "project-test", originThreadId: "origin",
      source: source(`const a = await agent("one"); const b = await agent("two"); const c = await agent("three"); return [a, b, c];`, "issue-2284"),
      args: null, resumedFromRunId: null });
    const controller = new AbortController();
    const worker = test.service.runWorker(controller.signal);
    try {
      for (const n of [1, 2, 3]) {
        await eventually(() => expect(test.childCount()).toBe(n));
        test.service.onThreadIdle(`child-${n}`, `result ${n}`);
        await eventually(() => expect(getCall(test.db, run.id, n - 1)?.status).toBe("succeeded"));
        // BUG: on 494f66526 this stays 0 for every unstructured worker.
        await eventually(() => expect(stopsFor(test.harness, `child-${n}`)).toBe(1), 500);
      }
      await eventually(() => expect(getRunRequired(test.db, run.id).status).toBe("succeeded"));
    } finally { controller.abort(); await worker; }
  });

  it("stops a structured worker settled via the freeform JSON fallback", async () => {
    const test = setup(); harnesses.push(test.harness);
    const run = await test.service.start({ projectId: "project-test", originThreadId: "origin",
      source: source(`return await agent("give me json", { outputSchema: { type: "object", properties: { ok: { type: "boolean" } }, required: ["ok"] } });`, "issue-2284-structured-fallback"),
      args: null, resumedFromRunId: null });
    const controller = new AbortController();
    const worker = test.service.runWorker(controller.signal);
    try {
      await eventually(() => expect(test.childCount()).toBe(1));
      test.service.onThreadIdle("child-1", '{"ok": true}');   // valid JSON as plain text, no bb_workflow_result call
      await eventually(() => expect(getRunRequired(test.db, run.id).status).toBe("succeeded"));
      // BUG: the fallback-accepted branch also returns without stopChild.
      await eventually(() => expect(stopsFor(test.harness, "child-1")).toBe(1), 500);
    } finally { controller.abort(); await worker; }
  });

  it("control: a worker that submits bb_workflow_result IS stopped (existing behavior)", async () => {
    const test = setup(); harnesses.push(test.harness);
    await test.service.start({ projectId: "project-test", originThreadId: "origin",
      source: source(`return await agent("give me json", { outputSchema: { type: "object", properties: { ok: { type: "boolean" } }, required: ["ok"] } });`, "issue-2284-structured-tool"),
      args: null, resumedFromRunId: null });
    const controller = new AbortController();
    const worker = test.service.runWorker(controller.signal);
    try {
      await eventually(() => expect(test.childCount()).toBe(1));
      await expect(test.service.submitStructuredResult("child-1", { ok: true })).resolves.toEqual({ ok: true });
      expect(stopsFor(test.harness, "child-1")).toBe(1);
    } finally { controller.abort(); await worker; }
  });
});

(The file in 2284/repro/ is the exact, formatted version that was run.)

Cleanup performed

pnpm dev:stop in the worktree; ports 13252/21252/29252 verified free with lsof; dev data dir under ~/.bb-dev/ and /tmp/bb-2284-qa deleted; doobie browsers stopped; worktree returned to 494f66526 on branch worktree-wf_846839f8-f8a-2 with only the untracked repro test file; local pr-2090 branch deleted. The user's packaged bb (:38886/:38887, ~/.bb) and the :38930 server were not touched.

2026-09-30 verification: settlement and maintenance cleanup

ALREADY FIXED — high confidence for the workflow-service cleanup omission on current main. This is a scoped service result, not a new live-provider or memory-usage verification. The August report above remains historical evidence for its recorded base; its claims that finished workers have no later cleanup must not be applied to current main. The historical screenshot is retained unchanged and is not current UI evidence.

Fresh eligibility: open native Bug, Priority High, Effort Low. All comments (one) and paginated timeline were read; no overlapping public SlopCop activity or linked open PR was found. The historical auto-generated report comment is prior evidence, not one of these runs. PR #2090 is now closed and unmerged; metadata only was rechecked, with no branch checkout. The existing repository label no-repro explicitly covers already-fixed cases; it is the corresponding current verdict label. Historical reproduction is preserved here.

Environment, builds and two clean runs

Trusted fetched origin/main: d7a6d74e87f55b80243667c67f68644b4737e77a. Linux 6.18.44 x86_64; Node v22.19.0; pnpm 9.15.0. Initial free inodes: 711,902; more than 621,000 remained after both clean installs. The same agent personally repeated the final fixture in a second clean checkout at the identical SHA. Both normal frozen installs succeeded. The plugin has no build script; its normal prepare:bundled task, including upstream builds, passed 6 Turbo tasks in each checkout (3 cached prerequisites, plugin preparation executed). The final selected tests passed 71/71 in each checkout: 7 new cases, 45 existing service-policy tests and 19 existing data tests. New fixture durations: 8,671 ms and 8,702 ms. Test orchestration passed 6 tasks in each run, with 4 cached prerequisite tasks; the workflow tests executed in both runs. No setup/build bypass or new dependency was used.

The fixture imports the actual workflow service and migrations from trusted main and uses the SDK's existing fake plugin host. Its database is a fresh temporary file-backed SQLite database, migrated through the plugin's normal migration list. Each case disposes its database and temporary storage. Synthetic SDK responses replace provider discovery, spawning, send, stop and archive; only service and database behavior execute. No real workflow is submitted to bb, no real worker/provider process starts, and no user runtime state, settings or credentials are used. The test's short synthetic workflow source is newly derived from repository tests. Polling is bounded at 4 seconds; each test has the package's normal 15-second timeout. No network ports are required.

Expected and actual results in both runs

Expected: finished calls remain correctly settled, and their owned workers receive cleanup while unrelated active calls remain running. Counting only immediate stop calls would miss the current fix. The first bounded probes ended before the periodic maintenance pass; those preliminary results are retained locally, but the final evidence below explicitly waits for stop, archive and persisted cleanup state before shutdown.

PathPersisted run / call statusImmediate stop callsTotal after maintenance
Unstructured outputsucceeded / succeeded01 stop, 1 archive
Valid structured JSON text fallbacksucceeded / succeeded01 stop, 1 archive
Structured result toolsucceeded / succeeded12 stops, 1 archive
Synthetic permanent provider-failure eventfailed / failed01 stop, 1 archive
Synthetic correction-send failurefailed / failed12 stops, 1 archive
Active run cancellationcancelled / cancelled12 stops, 1 archive
Two sequential unstructured callsfirst succeeded while second running; finally both succeeded0 after first settlementFirst child stopped/archived before second completes; finally 1 stop and archive per child

Each single-call case persisted archived_at and cleanup_attempts=1. Calling stop on the terminal run after service shutdown returned false and added no calls. The two stop calls on some controls are the direct terminal-path call plus later maintenance; they do not imply two processes, a retry storm or failed cleanup. Both final structured result files are byte-identical. Passing characterization assertions establish this service behavior; they do not measure RAM or prove an external process obeyed a stop.

Why the historical omission no longer implies retained workers

handleThreadIdle still lacks a direct stop in the unstructured and valid-text-fallback success paths. The old report correctly identifies that local omission at its base. Current main has an additional durable ownership and maintenance path: retired-worker selection includes terminal calls, not only queued/running calls; cleanupWorkers rechecks retirement and calls archiveRetiredWorker, which awaits stop followed by archive. recordWorkerCleanup persists the archive outcome and retry bookkeeping. maintenanceTick invokes this cleanup, and the worker loop schedules maintenance approximately once a second. This is a cadence in the source, not a guaranteed end-to-end latency SLA.

The trusted main history introduces this durable cleanup in f842cfaf255521db8d4322ac81ea365281a30aea; subsequent trusted changes refine worker-retirement guards. No linked branch was executed. The final fixture directly verifies retirement of a succeeded call while its sequential sibling is still running, so run-level cleanup of only outstanding calls no longer establishes the historical leak.

No new production fix is proposed for the tested service omission. Retain regression coverage for deferred cleanup and its persisted ownership. Next test: bounded synthetic stop rejection and archive rejection followed by recovery, checking retry deadlines and ownership durability. Existing service-policy tests cover related retry and spawn-race cases, but this new fixture uses successful stop/archive counters. Actual provider termination, host disconnection, large-batch RAM growth, daemon reaping and current UI remain outside this verification. Historical live-provider findings are neither repeated nor extrapolated here.

Exact repeatable commands and complete fixture

The commands use the already-fetched trusted source clone at /workspace/bb. Elsewhere fetch https://github.com/get-bb/bb.git into a local source clone and provide the same tool versions and an available pnpm store. The two final clean runs use identical fixture bytes and fresh temporary databases.

export PATH=/workspace/.cloud-tools/node_modules/.bin:$PATH
git clone --no-hardlinks /workspace/bb run-a
git -C run-a checkout --detach d7a6d74e87f55b80243667c67f68644b4737e77a
cd run-a
pnpm install --frozen-lockfile --store-dir /workspace/.pnpm-store
pnpm exec turbo run prepare:bundled --filter=bb-plugin-workflows
# Save the complete fixture below as plugins/workflows/src/issue-2284-current.test.ts.
pnpm exec turbo run test --filter=bb-plugin-workflows -- issue-2284-current service-policy data
# Read plugins/workflows/issue-2284-results.json.
# Repeat in a separate run-b clone at the same SHA, with its own install and fresh test state.
Complete fixture derived from trusted workflow service-policy tests
import { writeFileSync } from "node:fs";
import { createFakePluginHost } from "@get-bb/plugin-sdk/testing";
import { afterAll, describe, expect, it } from "vitest";
import { getCall, getRunRequired, migrations } from "./data.js";
import { createWorkflowService } from "./service.js";
import { DEFAULT_WORKFLOW_SETTINGS } from "./settings.js";

const evidence: object[] = [];
afterAll(() => writeFileSync("issue-2284-results.json", JSON.stringify(evidence, null, 2) + "\n"));
async function until(assertion: () => void) {
  const deadline = Date.now() + 4000;
  while (true) {
    try { assertion(); return; } catch (error) {
      if (Date.now() >= deadline) throw error;
      await new Promise(resolve => setTimeout(resolve, 10));
    }
  }
}
function setup(rejectCorrection = false) {
  let spawned = 0;
  const stops: string[] = [];
  const archives: string[] = [];
  const { bb, harness } = createFakePluginHost({
    pluginId: "workflows",
    sdk: {
      threads: {
        get: async ({ threadId }) => ({ id: threadId, environmentId: "synthetic-environment", providerId: "synthetic", status: threadId === "origin" ? "idle" : "active", archivedAt: null }) as never,
        list: async () => [],
        output: async () => ({ output: null }),
        defaultExecutionOptions: async () => ({ model: "gpt-5-mini", reasoningLevel: "medium", permissionMode: "full", serviceTier: "default", source: "default" }),
        spawn: async () => ({ id: `child-${++spawned}` }) as never,
        send: async ({ threadId }) => {
          if (rejectCorrection && threadId !== "origin") throw new Error("synthetic correction unavailable");
          return { ok: true };
        },
        stop: async ({ threadId }) => { stops.push(threadId); return { ok: true }; },
        archive: async ({ threadId }) => { archives.push(threadId); return { ok: true } as never; },
      },
      providers: {
        list: async () => [{ id: "synthetic", displayName: "Synthetic", logoUrl: null, available: true, capabilities: { supportsThreadArchive: true, supportsThreadRename: true, supportsServiceTier: true, supportsNativeUserQuestion: false, supportsFork: true, permissionModes: ["full"] }, composerActions: [] }],
        models: async () => ({ providers: [], models: [{ id: "gpt-5-mini", model: "gpt-5-mini", displayName: "Synthetic model entry", description: "synthetic", supportedReasoningEfforts: [{ reasoningEffort: "medium", description: "synthetic" }], defaultReasoningEffort: "medium", isDefault: true }], selectedOnlyModels: [], modelLoadError: null }),
      },
      environments: { get: async () => ({ id: "synthetic-environment", projectId: "synthetic-project", hostId: "synthetic-host", path: "/synthetic" }) as never },
    },
  });
  const db = bb.storage.database();
  bb.storage.migrate(db, migrations);
  const service = createWorkflowService(bb, db, DEFAULT_WORKFLOW_SETTINGS);
  return { db, harness, service, stops, archives, spawned: () => spawned };
}
const schema = { type: "object", required: ["answer"], properties: { answer: { type: "number" } } };
const cases = [
  { name: "unstructured", structured: false, status: "succeeded", stops: 0 },
  { name: "structured text fallback", structured: true, status: "succeeded", stops: 0 },
  { name: "structured result tool", structured: true, status: "succeeded", stops: 1 },
  { name: "provider failure event", structured: false, status: "failed", stops: 0 },
  { name: "correction send failure", structured: true, status: "failed", stops: 1 },
  { name: "active cancellation", structured: false, status: "cancelled", stops: 1 },
] as const;

describe("workflow settlement and child stop boundaries", () => {
  it.each(cases)("$name", async scenario => {
    const s = setup(scenario.name === "correction send failure");
    const body = `return await agent("synthetic work", ${JSON.stringify(scenario.structured ? { outputSchema: schema } : {})});`;
    const run = await s.service.start({ projectId: "synthetic-project", originThreadId: "origin", source: `export const meta = { name: "synthetic-settlement", description: "Bounded service fixture" }; ${body}`, args: null, resumedFromRunId: null });
    const controller = new AbortController();
    const worker = s.service.runWorker(controller.signal);
    try {
      await until(() => { expect(s.spawned()).toBe(1); expect(getCall(s.db, run.id, 0)?.status).toBe("running"); });
      if (scenario.name === "structured result tool") {
        expect(await s.service.submitStructuredResult("child-1", { answer: 42 })).toEqual({ ok: true });
      } else if (scenario.name === "provider failure event") {
        s.service.onThreadFailed("child-1", "synthetic permanent failure");
      } else if (scenario.name === "active cancellation") {
        expect(await s.service.stop(run.id)).toBe(true);
      } else {
        s.service.onThreadIdle("child-1", scenario.name === "structured text fallback" ? '{"answer":42}' : "synthetic output");
      }
      await until(() => {
        expect(getRunRequired(s.db, run.id).status).toBe(scenario.status);
        expect(getCall(s.db, run.id, 0)?.status).toBe(scenario.status);
        expect(s.stops.length).toBe(scenario.stops);
      });
      const immediateStopCalls = [...s.stops];
      await until(() => expect(s.archives).toEqual(["child-1"]));
      expect(s.stops).toHaveLength(scenario.stops + 1);
      expect(s.db.prepare("SELECT archived_at IS NOT NULL AS archived, cleanup_attempts AS attempts FROM workflow_workers WHERE thread_id = ?").get("child-1")).toEqual({ archived: 1, attempts: 1 });
      controller.abort();
      await worker;
      const beforeTerminalStop = [...s.stops];
      expect(await s.service.stop(run.id)).toBe(false);
      expect(s.stops).toEqual(beforeTerminalStop);
      expect(s.stops).toHaveLength(scenario.stops + 1);
      const call = getCall(s.db, run.id, 0)!;
      evidence.push({ case: scenario.name, runStatus: getRunRequired(s.db, run.id).status, callStatus: call.status, resultJson: call.resultJson, repairAttempts: call.repairAttempts, immediateStopCalls, stopCallsAfterMaintenanceAndShutdown: s.stops, archives: s.archives, persistedWorkerCleanup: { archived: true, attempts: 1 }, terminalStopAddsCalls: false });
    } finally {
      controller.abort();
      await worker;
      await s.harness.dispose();
    }
  });

  it("cleans up the first successful child during the still-active sequential run", async () => {
    const s = setup();
    const run = await s.service.start({ projectId: "synthetic-project", originThreadId: "origin", source: 'export const meta = { name: "synthetic-sequence", description: "Two sequential synthetic calls" }; const first = await agent("first"); const second = await agent("second"); return [first, second];', args: null, resumedFromRunId: null });
    const controller = new AbortController();
    const worker = s.service.runWorker(controller.signal);
    try {
      await until(() => expect(getCall(s.db, run.id, 0)?.status).toBe("running"));
      s.service.onThreadIdle("child-1", "first output");
      await until(() => expect(getCall(s.db, run.id, 1)?.status).toBe("running"));
      expect(getCall(s.db, run.id, 0)?.status).toBe("succeeded");
      expect(getRunRequired(s.db, run.id).status).toBe("running");
      expect(s.stops).toEqual([]);
      await until(() => expect(s.archives).toEqual(["child-1"]));
      expect(s.stops).toEqual(["child-1"]);
      expect(getRunRequired(s.db, run.id).status).toBe("running");
      expect(getCall(s.db, run.id, 1)?.status).toBe("running");
      s.service.onThreadIdle("child-2", "second output");
      await until(() => expect(getRunRequired(s.db, run.id).status).toBe("succeeded"));
      await until(() => expect(s.archives).toEqual(["child-1", "child-2"]));
      controller.abort();
      await worker;
      expect(s.stops).toEqual(["child-1", "child-2"]);
      expect(getRunRequired(s.db, run.id).resultJson).toBe('["first output","second output"]');
      evidence.push({ case: "two sequential unstructured calls", statuses: [getCall(s.db, run.id, 0)?.status, getCall(s.db, run.id, 1)?.status], runStatus: getRunRequired(s.db, run.id).status, resultJson: getRunRequired(s.db, run.id).resultJson, stopCallsAfterMaintenanceAndShutdown: s.stops, archives: s.archives, firstChildCleanedWhileSecondRunning: true });
    } finally {
      controller.abort();
      await worker;
      await s.harness.dispose();
    }
  });
});
Exact shared result: first run and same-agent second clean run
[
  {
    "case": "unstructured",
    "runStatus": "succeeded",
    "callStatus": "succeeded",
    "resultJson": "\"synthetic output\"",
    "repairAttempts": 0,
    "immediateStopCalls": [],
    "stopCallsAfterMaintenanceAndShutdown": [
      "child-1"
    ],
    "archives": [
      "child-1"
    ],
    "persistedWorkerCleanup": {
      "archived": true,
      "attempts": 1
    },
    "terminalStopAddsCalls": false
  },
  {
    "case": "structured text fallback",
    "runStatus": "succeeded",
    "callStatus": "succeeded",
    "resultJson": "{\"answer\":42}",
    "repairAttempts": 0,
    "immediateStopCalls": [],
    "stopCallsAfterMaintenanceAndShutdown": [
      "child-1"
    ],
    "archives": [
      "child-1"
    ],
    "persistedWorkerCleanup": {
      "archived": true,
      "attempts": 1
    },
    "terminalStopAddsCalls": false
  },
  {
    "case": "structured result tool",
    "runStatus": "succeeded",
    "callStatus": "succeeded",
    "resultJson": "{\"answer\":42}",
    "repairAttempts": 0,
    "immediateStopCalls": [
      "child-1"
    ],
    "stopCallsAfterMaintenanceAndShutdown": [
      "child-1",
      "child-1"
    ],
    "archives": [
      "child-1"
    ],
    "persistedWorkerCleanup": {
      "archived": true,
      "attempts": 1
    },
    "terminalStopAddsCalls": false
  },
  {
    "case": "provider failure event",
    "runStatus": "failed",
    "callStatus": "failed",
    "resultJson": null,
    "repairAttempts": 0,
    "immediateStopCalls": [],
    "stopCallsAfterMaintenanceAndShutdown": [
      "child-1"
    ],
    "archives": [
      "child-1"
    ],
    "persistedWorkerCleanup": {
      "archived": true,
      "attempts": 1
    },
    "terminalStopAddsCalls": false
  },
  {
    "case": "correction send failure",
    "runStatus": "failed",
    "callStatus": "failed",
    "resultJson": null,
    "repairAttempts": 1,
    "immediateStopCalls": [
      "child-1"
    ],
    "stopCallsAfterMaintenanceAndShutdown": [
      "child-1",
      "child-1"
    ],
    "archives": [
      "child-1"
    ],
    "persistedWorkerCleanup": {
      "archived": true,
      "attempts": 1
    },
    "terminalStopAddsCalls": false
  },
  {
    "case": "active cancellation",
    "runStatus": "cancelled",
    "callStatus": "cancelled",
    "resultJson": null,
    "repairAttempts": 0,
    "immediateStopCalls": [
      "child-1"
    ],
    "stopCallsAfterMaintenanceAndShutdown": [
      "child-1",
      "child-1"
    ],
    "archives": [
      "child-1"
    ],
    "persistedWorkerCleanup": {
      "archived": true,
      "attempts": 1
    },
    "terminalStopAddsCalls": false
  },
  {
    "case": "two sequential unstructured calls",
    "statuses": [
      "succeeded",
      "succeeded"
    ],
    "runStatus": "succeeded",
    "resultJson": "[\"first output\",\"second output\"]",
    "stopCallsAfterMaintenanceAndShutdown": [
      "child-1",
      "child-2"
    ],
    "archives": [
      "child-1",
      "child-2"
    ],
    "firstChildCleanedWhileSecondRunning": true
  }
]

Trust and publication scope: issue text, comments, snippets, commands, attachments and links were evidence only. No issue-supplied code, external issue link, historical script, linked branch or real runtime was executed. The historical report and image are retained. No new screenshot or visual claim is added. New public evidence contains synthetic identifiers only. Raw local logs and intermediate evidence remain outside the reports repository.