← reports

#2289 · Pi provider: stale model catalog after models.json change; plugin reload never retires superseded bridge workers

Bug Priority: High Effort: n/a (project fields not readable with this token) perf providers provider-pi host open on GitHub 2026-08-24 · base 494f66526

Verdict: REPRODUCED (headline bug; one secondary claim refuted) · Root-cause confidence: high

1. TL;DR

Add a model to pi's models.json while bb is running and bb provider models pi (and the app's model picker, which reads the same execution-options endpoint) keeps listing the old catalog. bb plugin reload provider-pi says "running" and does not help. The bridge worker process that serves model/list is not what holds the stale data: it holds a long-lived pi --mode rpc --no-session child, and pi answers get_available_models from the snapshot it built at boot (ModelRuntime.getAvailableSnapshot()); pi only re-reads models.json inside refresh(), which bb never triggers and which the pi RPC protocol does not expose. The bb bridge memoizes that child per cwd for 5 minutes of idle time, re-armed on every request, and the daemon keeps the bridge worker itself alive across plugin reloads because its process identity is the artifact digest plus declaration facts, both unchanged by a reload. On top of that the server memoizes every successful provider.list_models answer for 10 minutes (keyed by plugin registration revision, so a reload does bust that layer, but nothing else does). The "repeated reloads leak duplicate workers" part of the issue does not reproduce at the base commit: three reloads left exactly one provider-pi worker under the daemon; the extra workers the reporter saw are most plausibly the per-environment thread runtime's worker (by design, one per runtime) and, on 0.39.0, a maintenance runtime that had no idle shutdown yet.

2. Claims vs findings

Claim from the issueStatusEvidence
Adding a model to models.json is invisible to bb provider models pi while standalone pi sees it immediately.VerifiedStep 2 below: bb lists only bb2289/probe-one after probe-two was added; a fresh pi --list-models shows both. Unit test catalog.models-json-staleness.test.ts fails on base with exactly this diff.
"The Pi provider's bridge worker loads ~/.pi/agent/models.json once at process boot."Partially correctThe bridge worker never reads models.json. Its memoized pi --mode rpc --no-session catalog child does, once, at pi boot (ModelRuntime.create → ModelConfig.load), and get_available_models returns getAvailableSnapshot() without re-reading. Verified directly against pi 0.84.2 (PATH) and the pinned 0.84.0 with pi-rpc-stale-catalog.sh: two get_available_models calls around an edit both return ['probe-one'].
bb plugin reload provider-pi reports success but the host daemon keeps serving model/list from the old worker.VerifiedStep 3: reload printed provider-pi@0.1.0 running; the same worker pid 10089 (spawned 09:06:20) survived and the list stayed stale. Daemon log shows the same artifact digest 1337f13b… before and after every reload, so the process key is unchanged.
Repeated reloads leak duplicate bb-provider-bridge-worker … provider-pi processes.Refuted at base commitStep 5: three consecutive reloads → exactly one provider-pi worker under my daemon (pid 20831) with one pi child. The worker count only grows when a thread starts (its environment runtime gets its own worker, pid 35020, by design: one process per provider per runtime). Other provider-pi workers visible in ps belonged to other dev daemons on this machine (ppid 99045, 5958), not mine.
The process-key hash covers "bundled:<id>" for bundled Pi plus the capabilities hash.Refuted (immaterial)At base the launch source digest is the sha256 of the built host artifact (plugin-runtime.ts, artifact.digest), not a bundled: literal; the key is pi#bridge:<digest16>.<fingerprint(capabilities, providerOptions)>. The conclusion stands either way: the key cannot change when models.json changes.
retireStaleBridgeProcesses / retireSupersededBridgeProcessIfIdle only retire on a differing key, and plugin reload does not reach bridge lifecycle.VerifiedCode: runtime-provider-process.ts L270-311 compare processKey only; plugin-service.ts reload() L1871 re-runs loadOne and touches only the plugin host worker (plugin-host-manager.ts L426-447 replaces it by generation), never the provider bridge worker. Live: worker pid unchanged across reload.
SIGTERM to the tracked worker → next model/list spawns a fresh worker that re-reads models.json.Verified with a caveatKilling only the pi child (step 4) did not refresh the CLI listing because the server's 10-minute providerModelList memo answered without calling the daemon (no pi child was even respawned). Kill + bb plugin reload provider-pi (memo key bumped) → probe-two appeared. Step 8 also shows a worker recycled by the daemon's 60 s maintenance-runtime idle shutdown picks up the new file.
"Any other surviving superseded worker keeps running idle until manually killed."Not reproduced at baseAt base a maintenance-runtime worker with no threads is shut down after 60 s idle (PROVIDER_MAINTENANCE_IDLE_TIMEOUT_MS, added in d74974183 / #1879 after desktop-v0.39.0). On 0.39.0 there was no such timeout, which is consistent with what the reporter saw on that version. Sibling issue #2308 tracks the general "idle bridge processes never retired" problem.

3. Environment

4. Minimal reproduction

Repro files: 2289/repro/ — bbdev.sh (CLI against the isolated instance), add-model.py, pi-rpc-stale-catalog.sh (pi-only probe), reload-n-times.sh, warm-vs-idle.sh, catalog.models-json-staleness.test.ts (+ vitest-out.txt).

4a. Fastest: pi alone does not re-read models.json inside one RPC process

  1. Run pi-rpc-stale-catalog.sh. It writes a models.json with one custom model into an isolated agent dir, starts pi --mode rpc --no-session, asks get_available_models, edits the file to add a second model while the child is alive, asks again.
    $ AGENT_DIR=/tmp/bb-2289-pi-agent bash pi-rpc-stale-catalog.sh
    pi version: 0.84.2
    edited models.json: added probe-two
    response id=1 bb2289 models=['probe-one']
    response id=2 bb2289 models=['probe-one']        <-- expected ['probe-one', 'probe-two']
    Same with the pinned binary: PATH=plugins/provider-pi/node_modules/.bin:$PATH → pi version: 0.84.0, identical output (log).

4b. Unit-level: the bridge's memoized catalog (fails on base)

Save as plugins/provider-pi/src/bridge/catalog.models-json-staleness.test.ts, then cd plugins/provider-pi && pnpm exec vitest run src/bridge/catalog.models-json-staleness.test.ts. Uses the real pinned pi, an isolated agent dir, no network. (Note: the dev server watches builtin plugin sources, so write this file before starting a dev instance or expect a "1 file changed · rebuilt host" reload.)

/**
 * Repro for get-bb/bb#2289: the memoized `pi --mode rpc --no-session`
 * catalog child answers `get_available_models` from the snapshot pi built
 * at boot, so a model added to `<agent dir>/models.json` after the first
 * `model/list` never shows up until that child is gone.
 */
import { mkdtempSync, mkdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { fileURLToPath } from "node:url";
import { afterAll, beforeAll, expect, it } from "vitest";
import { BB_PI_EXTENSION_SOURCE } from "./bb-pi-extension.js";
import { closeAllPiCatalogs, getPiCatalog } from "./catalog.js";
import { PI_BRIDGE_COMMAND_ENV } from "./rpc-child.js";

const pinnedPi = fileURLToPath(new URL("../../node_modules/.bin/pi", import.meta.url));

let scratch: string;
let agentDir: string;
let extensionPath: string;
const savedEnv: Record<string, string | undefined> = {};

function modelsJson(ids: string[]): string {
  return JSON.stringify({ providers: { bb2289: {
    baseUrl: "http://127.0.0.1:9/v1", api: "openai-completions", apiKey: "not-a-real-key",
    models: ids.map((id) => ({ id, name: id })),
  } } }, null, 2);
}

beforeAll(() => {
  scratch = mkdtempSync(join(tmpdir(), "bb-2289-catalog-"));
  agentDir = join(scratch, "agent");
  mkdirSync(agentDir);
  writeFileSync(join(agentDir, "models.json"), modelsJson(["probe-one"]));
  extensionPath = join(scratch, "bb-pi-extension.mjs");
  writeFileSync(extensionPath, BB_PI_EXTENSION_SOURCE, "utf8");
  for (const key of [PI_BRIDGE_COMMAND_ENV, "PI_CODING_AGENT_DIR", "PI_OFFLINE"]) savedEnv[key] = process.env[key];
  process.env[PI_BRIDGE_COMMAND_ENV] = pinnedPi;
  process.env.PI_CODING_AGENT_DIR = agentDir;
  process.env.PI_OFFLINE = "1";
});

afterAll(() => {
  closeAllPiCatalogs();
  for (const [key, value] of Object.entries(savedEnv)) {
    if (value === undefined) delete process.env[key]; else process.env[key] = value;
  }
  rmSync(scratch, { recursive: true, force: true });
});

it("model/list reflects a model added to models.json after the catalog child booted", async () => {
  const catalog = await getPiCatalog(scratch, extensionPath);
  const before = (await catalog.listModels()).models.map((m) => m.model);
  expect(before).toEqual(["bb2289/probe-one"]);

  writeFileSync(join(agentDir, "models.json"), modelsJson(["probe-one", "probe-two"]));
  expect(JSON.parse(readFileSync(join(agentDir, "models.json"), "utf8")).providers.bb2289.models).toHaveLength(2);

  // Same memoized catalog (same cwd): pi is asked again, answers from its boot snapshot.
  const after = (await catalog.listModels()).models.map((m) => m.model);
  expect(after).toEqual(["bb2289/probe-one", "bb2289/probe-two"]);   // <-- fails on base
}, 60_000);
FAIL  src/bridge/catalog.models-json-staleness.test.ts > model/list reflects a model added to models.json after the catalog child booted
AssertionError: expected [ 'bb2289/probe-one' ] to deeply equal [ 'bb2289/probe-one', …(1) ]
- Expected
+ Received
  [
    "bb2289/probe-one",
-   "bb2289/probe-two",
  ]
 ❯ src/bridge/catalog.models-json-staleness.test.ts:81:17
 Test Files  1 failed (1)   Tests  1 failed (1)   Duration 1.13s

4c. End to end on a running bb (CLI)

  1. Prepare an isolated pi agent dir and start a dev instance with it:
    mkdir -p /tmp/bb-2289-pi-agent-live
    cat > /tmp/bb-2289-pi-agent-live/models.json <<'EOF'
    { "providers": { "bb2289": { "baseUrl": "http://127.0.0.1:9/v1", "api": "openai-completions",
      "apiKey": "not-a-real-key", "models": [ { "id": "probe-one", "name": "probe-one" } ] } } }
    EOF
    PI_CODING_AGENT_DIR=/tmp/bb-2289-pi-agent-live scripts/bb-dev-app current
    eval "$(scripts/bb-dev-app env)"
  2. First listing (spawns the bridge worker and its pi catalog child):
    $ pnpm bb:dev provider models pi
    Models for pi:
    
    Model             Name       Default
    ----------------  ---------  -------
    bb2289/probe-one  probe-one  *
    ps: worker bridge-worker-entry.ts …/provider-pi/1337f13b…/host.mjs provider-pi pid 10089 (ppid = host daemon 9024), child pi pid 10238.
  3. Add a model, compare standalone pi with bb:
    $ python3 add-model.py /tmp/bb-2289-pi-agent-live/models.json probe-two
    $ PI_CODING_AGENT_DIR=/tmp/bb-2289-pi-agent-live PI_OFFLINE=1 pi --list-models | grep bb2289
    bb2289    probe-one  128K     16.4K    no        no
    bb2289    probe-two  128K     16.4K    no        no
    $ pnpm bb:dev provider models pi
    Model             Name       Default
    ----------------  ---------  -------
    bb2289/probe-one  probe-one  *                      <-- expected probe-two too
  4. Reload the plugin, list again:
    $ pnpm bb:dev plugin reload provider-pi
    provider-pi@0.1.0  running
      source: builtin:provider-pi
    $ pnpm bb:dev provider models pi
    Model             Name       Default
    ----------------  ---------  -------
    bb2289/probe-one  probe-one  *                      <-- still stale; worker pid 10089 unchanged
  5. Kill only the pi catalog child, list again (still stale: the server memo answers, the daemon is not called, no pi child is respawned); then reload (bumps the memo key) and list:
    $ kill -TERM 10238
    $ pnpm bb:dev provider models pi          # stale, served from the 10-minute server memo
    bb2289/probe-one  probe-one  *
    $ pnpm bb:dev plugin reload provider-pi
    $ pnpm bb:dev provider models pi
    bb2289/probe-one  probe-one  *
    bb2289/probe-two  probe-two                          <-- fresh pi child re-read models.json
  6. Three reloads in a row, then count workers (log): exactly one provider-pi worker under the dev daemon (pid 20831, ppid 9024) and one pi child. No duplicates.
  7. Warm vs idle worker (warm-vs-idle.sh, log). Cycle A ran after the maintenance runtime had idled > 60 s (its worker was already gone); cycle B ran 2 s later against the warm worker:
    === 09:16:04 cycle probe-four: add probe-four, bb plugin reload provider-pi, bb provider models pi
    bb2289/probe-one … bb2289/probe-four                  <-- fresh worker 57438 (started 09:16:05) sees it
    === 09:16:07 cycle probe-five: add probe-five, bb plugin reload provider-pi, bb provider models pi
    bb2289/probe-one … bb2289/probe-four                  <-- probe-five missing; same worker 57438

Not a visual bug; no screenshots. The app's model picker reads GET /api/v1/system/execution-options, the same memoized path the CLI uses (sdk.providers.models), so it shows the same stale list.

5. Root cause

Four independent caches sit between models.json and the picker. Only the last one is invalidated by a plugin reload.

5.1 pi: get_available_models is a boot-time snapshot (the actual data staleness)

In pi 0.84.x RPC mode the command returns session.modelRuntime.getAvailableSnapshot() (@earendil-works/pi-coding-agent/dist/modes/rpc/rpc-mode.js L380-383). The snapshot is rebuilt only by ModelRuntime.refresh(), which is where this.config = await ModelConfig.load(this.modelsPath) happens (dist/core/model-runtime.js L499-500); refresh() runs at boot and after credential/provider registration changes (L551, L587, L594). There is no RPC command that triggers it and no file watcher. A fresh pi process (what the reporter calls "standalone pi resolves it immediately") loads the file at boot, so it always looks fresh.

5.2 pi bridge: one memoized catalog child per cwd

plugins/provider-pi/src/bridge/catalog.ts L93-131: catalogsByCwd memoizes a pi --mode rpc --no-session --extension … child; spawnChild() reuses it while !child.exited, and fetchRaw() sends get_available_models to it on every model/list, provider/health and bare-id model resolution (bridge.ts L586-606, L622-624, L678-679). The child is only closed after catalogIdleMs() = 5 minutes without a request (L166-186), and every request re-arms that timer. The header comment even notes "No catalog network-refresh control exists over RPC: pi refreshes at its own startup". Nothing in the bridge looks at models.json.

const fetchRaw = async (): Promise<PiRpcModel[]> => {
  const data = (await spawnChild().requestOk({
    type: "get_available_models",
  })) as { models?: unknown[] } | undefined;
  // Idle counts from the answer: a slow boot must not evict the child.
  touch();
  …

5.3 host daemon: the bridge worker's identity cannot change on reload

runtime.ts L410-418 keys the process as `${providerId}#bridge:${bridgeLaunchProcessKey(bridgeLaunch)}`, and bridge-launch-process-key.ts L50-57 builds that from the artifact digest prefix plus a fingerprint of capabilities and providerOptions. runtime-provider-process.ts L270-311 retires a thread-less process only when its key differs from the newly ensured one. A builtin plugin reload re-runs loadOne (plugin-service.ts L1871-1888) and produces the same artifact digest (daemon log: 1337f13b… at 09:06:20, 09:07:11, 09:08:47, 09:10:30, 09:12:34), so listModels (runtime.ts L2458-2478) keeps hitting the same worker and the same catalog child. The reload does restart the plugin's host worker (plugin-host-manager.ts L426-447, new generation), which is why it looks like "something restarted".

What does bound the worker's life at base: model/list runs in the daemon's provider-maintenance runtime (app.ts L768-773), which is torn down, bridge workers included, after 60 s without a maintenance request (runtime-manager.ts L51, L888-912). Observed: worker 45894 spawned 09:12:34 was gone by 09:14:04. This timeout was added in d74974183 (#1879), after desktop-v0.39.0, so the reporter's 0.39.0 build kept the worker forever; at base the stale window is "as long as something keeps asking within 60 s" (health probes for the settings page and the execution-options provider discovery both go through the same runtime and the same catalog.rawModels()).

5.4 server: 10-minute memo of the model list

lifecycle-dedupers.ts L11-17 (PROVIDER_MODEL_LIST_MEMO_TTL_MS = 10 * 60_000) and execution-options.ts L620-640: the key is [hostId, daemonSessionId, registrationRevision, command]. This is why killing the pi child alone changed nothing for the CLI (step 4: no provider.list_models RPC reached the daemon) and why the reporter's "picker updated after killing the worker" needs either a reload, a daemon reconnect, or 10 minutes. It also means that even after the fix below, an edit is visible at most 10 minutes later unless the user reloads the plugin.

Why the symptom follows

Edit → pi child still holds the boot snapshot (5.1) → bridge keeps asking that child (5.2) → daemon keeps that bridge worker for the same key across reloads, and for as long as requests arrive within 60 s (5.3) → server re-serves the last answer for 10 minutes unless the registration revision changes (5.4). A plugin reload clears 5.4 only.

Deeper issue

The bridge's resolvePiModel (bridge.ts L652-692) refuses <provider>/<id> when a warm catalog for that cwd knows the provider but not the id. With a stale catalog this can reject a model that pi itself would run. I did not hit it live because a thread's environment runtime uses a separate bridge process whose catalog was cold (the probe-three thread spawned, booted a fresh pi session and failed only with the expected "Connection error" against the fake endpoint), but a model/list with a workspace cwd through the same process would arm it.

6. Proposed fix (first principles)

Fix it in the pi bridge (the plugin owns provider translation; the daemon should not learn about pi's config files, and putting models.json into the daemon's process key, as the issue suggests first, would make a host-wide bridge identity depend on a per-user file and still leave the server memo stale).

  1. Invalidate the catalog child when models.json changes. In catalog.ts, resolve pi's agent dir the way pi does (PI_CODING_AGENT_DIR else ~/.pi/agent) and record a content hash (or mtimeMs + size) of models.json when the child is spawned. In fetchRaw()/probe(), compare before sending; on a difference child.kill() and let spawnChild() respawn (it already handles child.exited). The file is tiny, hashing it per request is cheap, and the respawn cost (~1.4 s, what the first probe already costs) is paid once per edit. Missing file ↔ present file counts as a change. This fixes 5.1 and 5.2 together and also keeps resolvePiModel honest.
  2. Let the server see it. The 10-minute memo (5.4) still hides the change. Either shorten the TTL for providers whose catalog is user-editable, or add a cheap "catalog fingerprint" to the memo key; the pragmatic option is to keep the TTL and document that bb plugin reload provider-pi (which already bumps the key) is the refresh path, now that the bridge respawns the child behind it.
  3. Regression test: the vitest above (real pinned pi, isolated agent dir) passes with (1); add a second case that the child pid changed exactly once and that no extra pi process survives (closeAllPiCatalogs).

What could go wrong: editors that write via rename can produce a transient missing file (hash "absent" → respawn → hash present → respawn again; harmless but wasteful, so compare against the last successful spawn's hash). Models added through pi extensions/registerProvider are not in models.json and stay invisible to the fingerprint; that is a pi-side concern (an RPC refresh command upstream would be the cleanest long-term answer and the bridge could call it instead of respawning). Do not include the fingerprint in the daemon process key: that would spawn a second bridge worker per edit while threads pin the old one.

7. PR review

No open PRs are linked to this issue.

8. Related issues

9. Appendix

Daemon log: artifact digest across reloads

09:06:20 Downloading 1337f13bb6e9   (first model/list spawns worker 10089)
09:07:11 Using cached 1337f13bb6e9   (after reload #1)
09:08:47 Using cached 1337f13bb6e9   (after reload #2)
09:10:30 Using cached 1337f13bb6e9   (thread.start, environment runtime worker 35020)
09:12:34 Using cached 1337f13bb6e9   (model/list after the maintenance runtime idled; new worker 45894)
[09:06:21] Online host RPC {"commandType":"provider.list_models","handlerMs":1610.7,"ok":true}
[09:07:02] plugin provider-pi@0.1.0 loaded            (reload #1; list at 09:07:11 hit the same worker, <1 s so not logged)
[09:08:39] plugin provider-pi@0.1.0 loaded            (reload #2)
[09:08:49] Online host RPC {"commandType":"provider.list_models","handlerMs":1438.5,"ok":true}   (cold pi child → probe-two visible)
[09:09:18] [09:09:19] [09:09:20] plugin provider-pi@0.1.0 loaded   (reload ×3; worker count stayed 1)

Process snapshots

# after step 1 (my daemon is pid 9024)
10089  9024  …/bridge-worker-entry.ts …/plugin-host-artifacts/provider-pi/1337f13b…/host.mjs provider-pi …
10238 10089  pi                                   (catalog child: pi --mode rpc --no-session --extension …)
# after 3 reloads (step 5)
20831  9024  provider-pi worker     20994 20831 pi
# after thread.start (step 6): a second worker for the thread's environment runtime, by design
35020  9024  provider-pi worker     35029 35020 pi (the thread's session child)

Thread started on a model only present in the edited file

thread spawn --provider pi --model bb2289/probe-three (thr_9267nerxqu) was accepted, its fresh pi session child read the file, and the turn failed with provider/error: "Connection error." against the fake endpoint — i.e. the per-thread runtime is not affected by the maintenance catalog's staleness unless a catalog is warm in that same bridge process.

Commands run (abridged)

gh issue view 2289 --repo get-bb/bb --json …
pnpm install --frozen-lockfile --prefer-offline ; pnpm exec turbo run build
git fetch origin main ; git log 494f66526..origin/main --oneline      # empty
AGENT_DIR=/tmp/bb-2289-pi-agent bash pi-rpc-stale-catalog.sh            # pi 0.84.2
PATH=plugins/provider-pi/node_modules/.bin:$PATH AGENT_DIR=/tmp/bb-2289-pi-agent-0840 bash pi-rpc-stale-catalog.sh   # pi 0.84.0
PI_CODING_AGENT_DIR=/tmp/bb-2289-pi-agent-live scripts/bb-dev-app current
curl -s -X POST $BB_SERVER_URL/api/v1/projects -d '{"name":"qa","source":{"type":"local_path","path":"/tmp/bb-2289-qa","hostId":"host_6b6az84tvc"}}'
bbdev.sh provider models pi ; add-model.py … probe-two ; bbdev.sh provider models pi ; bbdev.sh plugin reload provider-pi ; …
kill -TERM <pi child> ; bbdev.sh provider models pi ; bbdev.sh plugin reload provider-pi ; bbdev.sh provider models pi
reload-n-times.sh 3 ; warm-vs-idle.sh
cd plugins/provider-pi && pnpm exec vitest run src/bridge/catalog.models-json-staleness.test.ts
pnpm dev:stop ; rm -rf ~/.bb-dev/…-0c45c8842b56 /tmp/bb-2289-* ; lsof -nP -iTCP -sTCP:LISTEN | grep -E ':(14908|22908|30908)'   # free

Raw logs: step1, step2, step3 reload, step3 list, step4 kill, step4 kill+reload, step5, step6, step8, pi-only probe (0.84.2), pi-only probe (0.84.0), dev-app start, vitest output.

2026-09-30 verification: catalog invalidation and child lifecycle

Verdict: REPRODUCED for stale model results at the actual bb catalog boundary when the child serves a boot-time model snapshot. Root-cause confidence: high within that scope. Duplicate bridge-worker leakage remains not reproduced; this verification does not assert a leak. The original 2026-08-24 report and all its historical evidence above are retained; one historical permalink end line was corrected from 58 to 57 to match that file. Its live Pi, server reload and daemon observations are historical, not new measurements.

Base: 0e7b518f135d43005dae201ef34ebb3001607eb4, fetched from trusted get-bb/bb origin/main. Two separate clean checkouts were pinned to that SHA. Linux 6.18.44 x86_64; Node 22.19.0; pnpm 9.15.0; Vitest 4.1.1. Workspace capacity before setup was 1,866,140 free inodes; after both installations and checks it was 1,776,218. The earlier #2122 evidence was retained.

Eligibility and scope: #2289 was still an open native Bug, Priority High, Effort Medium, with confirmed-repro. The issue, its one comment and complete paginated timeline were read as untrusted evidence. No linked open PR or open bb PR mentioning #2289 was found. The report had no new activity since its historical publication. No overlapping GitHub workflow activity or visible SlopCop activity was found; private SlopCop runtime state was not queried under the no-real-runtime constraint. Both source and reports repositories are public.

Two personal clean runs

The same agent personally repeated the identical faithful fixture in a second clean checkout with a separate frozen install and fresh synthetic directories/processes. Run A started at 18:04:34 UTC and Run B at 18:05:01 UTC on 2026-09-30. Each passed 5/5 tests, with the identical output below. Every scenario wrote its own synthetic models.json, lifecycle log and result under a fresh issue-2289-state-* directory. No TCP port or real provider was used; transport was local stdio plus the extension pipes.

The actual getPiCatalog, PiRpcChild, cache map, idle timers, retry and close code ran unchanged. Only the Pi child was replaced by a small protocol fixture. Its snapshot mode reads a synthetic model file once; its live-answer mode rereads on every model request. The installed, lockfile-pinned Pi 0.84.0 source was inspected, not executed: its RPC get_available_models calls getAvailableSnapshot(), while refresh() loads the model configuration again. This supports the fixture's snapshot premise, but this run does not establish real-provider or UI behavior end to end.

Normal pnpm install --frozen-lockfile succeeded in both checkouts. The SDK runtime was built with Turbo; the selected Turbo tests also ran their normal build/generation and native-module prerequisites. Both final tests were forced executions, not cached results. No production source, dependency manifest or lockfile was modified, and no dependency was added.

Expected versus actual

Expected: a supported configuration-refresh path should make a newly added model visible without abandoning catalog children; repeated same-directory requests should not spawn duplicates. Actual: changing the model file alone leaves a warm snapshot child serving only synthetic/one. Twelve concurrent requests reuse one catalog and one child. An explicit catalog close, an exited child retried by the catalog, or idle eviction makes the subsequent generation return both synthetic/one and synthetic/two. The names are synthetic fixture identifiers, not real model IDs.

RESULT {"case":"warm-and-close","before":["synthetic/one"],"after":["synthetic/one"],"fresh":["synthetic/one","synthetic/two"],"concurrentReads":12,"spawns":2,"exits":2}
RESULT {"case":"live-answer-control","after":["synthetic/one","synthetic/two"],"spawns":1,"exits":1}
RESULT {"case":"child-exit-retry","after":["synthetic/one","synthetic/two"],"spawns":2,"exits":2}
RESULT {"case":"idle-eviction","after":["synthetic/one","synthetic/two"],"idleMs":150,"spawns":2,"exits":2}
RESULT {"case":"cwd-and-cleanup","cycles":3,"peakChildren":2,"spawns":6,"exits":6,"remaining":0}

An exploratory fixture initially used a blocking file reader for the extension pipe, which delayed process exit and interfered with the shortened idle timer. That harness failure is preserved in local logs but is not evidence of a product leak. The final fixture uses the repository's nonblocking socket transport pattern (plugins/provider-pi/src/bridge/bb-pi-extension.ts: 226-L233); both final clean runs passed with that identical implementation.

Root cause supported on this SHA

Current server caveat: the historical report's registration-revision memo description is not the current implementation. Today createLifecycleDedupers installs ProviderModelCatalogStore (apps/server/src/lifecycle-dedupers.ts: 19-L30). Its fingerprint uses provider identity, launch source/options and environment passthrough (apps/server/src/services/providers/provider-model-catalog-store.ts: 160-L175); the picker can serve a catalog while refreshing in the background after ten minutes, and validation has a separate age/missing-model rule (apps/server/src/services/providers/provider-model-catalog-store.ts: 218-L248). This is source inspection only. No claim is made that plugin reload currently invalidates the full server-to-child chain.

Small fix proposal, not implemented: fingerprint the resolved Pi model configuration at the catalog layer and retire/await the catalog-only child before creating a replacement when that fingerprint changes, or negotiate an explicit model refresh with the provider. Share replacement work across concurrent callers. Keep active conversation children separate. Also ensure the server's catalog refresh/invalidation path reaches this refreshed provider catalog; a local child fix alone does not prove immediate picker freshness. Preserve the current successful idle and explicit-close behavior.

Next test: after such a change, invert the warm-file-edit assertion to require the added model while preserving single replacement and full cleanup; then add removal, atomic replacement, malformed-file handling and concurrent invalidation cases. A separately authorized server/daemon integration test should verify the supported refresh UI/CLI path and distinguish maintenance workers from conversation workers before making any duplicate-worker claim.

Exact commands and complete fixtures

Use Node 22.19.0 and pnpm 9.15.0. The store path below is the writable store used here; use a writable equivalent elsewhere. All raw evidence remains outside the reports repository. The fixture opens only its synthetic file and protocol pipes and ignores the supplied extension file rather than loading provider code.

git clone https://github.com/get-bb/bb.git run-a
cd run-a
git checkout --detach 0e7b518f135d43005dae201ef34ebb3001607eb4
node --version  # v22.19.0
pnpm --version  # 9.15.0
pnpm install --frozen-lockfile --store-dir /workspace/.pnpm-store
pnpm exec turbo run build --filter=@get-bb/plugin-sdk
# Save the two complete inline files below at their stated paths.
pnpm exec turbo run test --filter=bb-plugin-provider-pi --force -- src/bridge/issue-2289.test.ts
# Repeat in a second fresh clone named run-b at the same SHA.
# Install separately and copy only the two fixture files, never prior state.
plugins/provider-pi/src/bridge/issue-2289-fixture.mjs
import { appendFileSync, existsSync, readFileSync, unlinkSync, writeSync } from 'node:fs';
import { createInterface } from 'node:readline';
import { Socket } from 'node:net';
import { join } from 'node:path';
const root = process.argv[2];
const mode = process.argv[3];
const log = (event) => appendFileSync(join(root, 'lifecycle.jsonl'), JSON.stringify({event, pid:process.pid}) + '\n');
const readModels = () => JSON.parse(readFileSync(join(root, 'models.json'), 'utf8')).providers.synthetic.models.map(m => ({...m,provider:'synthetic',input:['text'],reasoning:false}));
const snapshot = readModels();
log('spawn');
process.on('exit', () => log('exit'));
process.on('SIGTERM', () => process.exit(0));
const channel = createInterface({input:new Socket({fd:4,readable:true,writable:false})});
channel.on('line', line => {
  const request = JSON.parse(line);
  writeSync(3, JSON.stringify({kind:'reply',id:request.id,result:{scopedModelIds:[]}})+'\n');
});
createInterface({input:process.stdin}).on('line', line => {
  const request = JSON.parse(line);
  log(request.type);
  if (request.type === 'get_available_models' && existsSync(join(root,'exit-once'))) {
    unlinkSync(join(root,'exit-once'));
    process.exit(0);
  }
  const data = request.type === 'get_available_models' ? {models:mode === 'live' ? readModels() : snapshot} : {snapshotIds:snapshot.map(m => m.id)};
  process.stdout.write(JSON.stringify({type:'response',id:request.id,command:request.type,success:true,data})+'\n');
});
plugins/provider-pi/src/bridge/issue-2289.test.ts
import { existsSync, mkdirSync, mkdtempSync, readFileSync, writeFileSync } from 'node:fs';
import { join } from 'node:path';
import { fileURLToPath } from 'node:url';
import { afterEach, expect, it, vi } from 'vitest';
import { z } from 'zod';
import { closeAllPiCatalogs, getPiCatalog, peekPiCatalog } from './catalog.js';
import { PI_BRIDGE_ARGS_ENV, PI_BRIDGE_COMMAND_ENV } from './rpc-child.js';

const rowSchema = z.object({event:z.string(),pid:z.number().int().positive()});
const roots: string[] = [];
const delay = (ms:number) => new Promise(resolve => setTimeout(resolve,ms));
async function until(check:()=>boolean) {
  const deadline = Date.now()+8000;
  while (!check()) { if (Date.now()>deadline) throw new Error('Synthetic lifecycle timeout'); await delay(10); }
}
function setup(mode='snapshot',idle='5000') {
  const root=mkdtempSync(fileURLToPath(new URL('./issue-2289-state-',import.meta.url)));
  roots.push(root);
  const cwd=join(root,'workspace');mkdirSync(cwd);
  const extension=join(root,'unused-extension.mjs');writeFileSync(extension,'export {};\n');
  vi.stubEnv(PI_BRIDGE_COMMAND_ENV,process.execPath);
  vi.stubEnv(PI_BRIDGE_ARGS_ENV,JSON.stringify([fileURLToPath(new URL('./issue-2289-fixture.mjs',import.meta.url)),root,mode]));
  vi.stubEnv('BB_PI_CATALOG_IDLE_MS',idle);
  const rows=()=>existsSync(join(root,'lifecycle.jsonl')) ? readFileSync(join(root,'lifecycle.jsonl'),'utf8').trim().split('\n').filter(Boolean).map(line=>rowSchema.parse(JSON.parse(line))) : [];
  const count=(event:string)=>rows().filter(row=>row.event===event).length;
  const alive=()=>count('spawn')-count('exit');
  const write=(ids:string[])=>writeFileSync(join(root,'models.json'),JSON.stringify({providers:{synthetic:{models:ids.map(id=>({id,name:id}))}}}));
  write(['one']);
  return {root,cwd,extension,rows,count,alive,write};
}
function record(root:string,result:object) {
  writeFileSync(join(root,'result.json'),JSON.stringify(result,null,2));
  process.stderr.write('RESULT '+JSON.stringify(result)+'\n');
}
afterEach(async()=>{
  await closeAllPiCatalogs();
  vi.unstubAllEnvs();
});
it('same-cwd cache retains a stale child until explicit close',async()=>{
  const f=setup();
  const catalog=await getPiCatalog(f.cwd,f.extension);
  const before=(await catalog.listModels()).models.map(m=>m.id);
  f.write(['one','two']);
  const copies=await Promise.all(Array.from({length:12},()=>getPiCatalog(f.cwd,f.extension)));
  expect(copies.every(c=>c===catalog)).toBe(true);
  const after=await Promise.all(copies.map(async c=>(await c.listModels()).models.map(m=>m.id)));
  expect(before).toEqual(['synthetic/one']);
  expect(after.every(ids=>JSON.stringify(ids)===JSON.stringify(before))).toBe(true);
  expect((await catalog.probe()).snapshotIds).toEqual(['one']);
  expect(f.count('spawn')).toBe(1);expect(f.alive()).toBe(1);
  await closeAllPiCatalogs();expect(f.alive()).toBe(0);
  const replacement=await getPiCatalog(f.cwd,f.extension);
  const fresh=(await replacement.listModels()).models.map(m=>m.id);
  expect(fresh).toEqual(['synthetic/one','synthetic/two']);
  await closeAllPiCatalogs();expect(f.alive()).toBe(0);
  record(f.root,{case:'warm-and-close',before,after:after[0],fresh,concurrentReads:12,spawns:f.count('spawn'),exits:f.count('exit')});
},15000);
it('live-answer control proves model answers themselves are not memoized',async()=>{
  const f=setup('live');const catalog=await getPiCatalog(f.cwd,f.extension);
  await catalog.listModels();f.write(['one','two']);
  const after=(await catalog.listModels()).models.map(m=>m.id);
  expect(after).toEqual(['synthetic/one','synthetic/two']);expect(f.count('spawn')).toBe(1);
  await closeAllPiCatalogs();expect(f.alive()).toBe(0);
  record(f.root,{case:'live-answer-control',after,spawns:f.count('spawn'),exits:f.count('exit')});
},15000);
it('an exited catalog child is replaced once and the failed request retried',async()=>{
  const f=setup();const catalog=await getPiCatalog(f.cwd,f.extension);await catalog.listModels();
  f.write(['one','two']);writeFileSync(join(f.root,'exit-once'),'synthetic trigger');
  const after=(await catalog.listModels()).models.map(m=>m.id);
  expect(after).toEqual(['synthetic/one','synthetic/two']);expect(f.count('spawn')).toBe(2);expect(f.alive()).toBe(1);
  expect((await catalog.probe()).snapshotIds).toEqual(['one','two']);
  await closeAllPiCatalogs();expect(f.alive()).toBe(0);
  record(f.root,{case:'child-exit-retry',after,spawns:f.count('spawn'),exits:f.count('exit')});
},15000);
it('idle cache eviction kills the child and a fresh lookup reads edits',async()=>{
  const f=setup('snapshot','150');const catalog=await getPiCatalog(f.cwd,f.extension);await catalog.listModels();
  f.write(['one','two']);
  await until(()=>peekPiCatalog(f.cwd)===null && f.alive()===0);
  const fresh=await getPiCatalog(f.cwd,f.extension);expect(fresh).not.toBe(catalog);
  const after=(await fresh.listModels()).models.map(m=>m.id);
  expect(after).toEqual(['synthetic/one','synthetic/two']);
  await closeAllPiCatalogs();expect(f.alive()).toBe(0);
  record(f.root,{case:'idle-eviction',after,idleMs:150,spawns:f.count('spawn'),exits:f.count('exit')});
},15000);
it('separate cwd catalogs are intentional and three close/recreate cycles leave no children',async()=>{
  const f=setup();const second=join(f.root,'workspace-two');mkdirSync(second);
  for(let cycle=0;cycle<3;cycle++) {
    const [a,b]=await Promise.all([getPiCatalog(f.cwd,f.extension),getPiCatalog(second,f.extension)]);
    expect(a).not.toBe(b);await Promise.all([a.listModels(),b.listModels()]);expect(f.alive()).toBe(2);
    await closeAllPiCatalogs();expect(f.alive()).toBe(0);
  }
  expect(f.count('spawn')).toBe(6);expect(f.count('exit')).toBe(6);
  record(f.root,{case:'cwd-and-cleanup',cycles:3,peakChildren:2,spawns:f.count('spawn'),exits:f.count('exit'),remaining:0});
},15000);

Remaining limits and trust: no real Pi process, bb application, daemon, server, provider account, production state or real user model file was used. Plugin reload, picker rendering, server-cache behavior and daemon worker counts were not dynamically reverified. No new visual claims or screenshots are added. The fixture was authored from trusted current repository contracts and test helpers; issue-supplied commands, tests, patches and linked branches were not executed, and external issue links were not fetched. There were no investigation agents, production fixes, PRs or manually started workflows. Historical duplicate-worker leakage remains unconfirmed.