#4154 · ACP probe failure stops later model discovery
Verdict: REPRODUCED. Root-cause confidence: high.
1. TL;DR
A model picker can lose reasoning choices for healthy models when an earlier model rejects its discovery request. The ACP bridge catches errors around the entire sequential probe loop, so the first rejection ends discovery. Missing entries receive the generic medium effort and the bridge returns a successful catalog. A synthetic agent reproduced this exact behavior twice on unchanged production code; a control with no rejection passed.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| One rejected probe prevents later probes | Verified | Both request logs contain only probe-a and probe-b; probe-c is absent. |
| Later healthy models receive only medium | Verified | The returned catalog gives probe-c a medium fallback; all three have low/medium/high in the passing control. |
| Failure is silently accepted | Verified at bridge boundary | The result is successful JSON-RPC, and the catch has no logging. See root-cause links. |
| Partial catalog can be persisted as successful | Supported by source | The server persists successful model-list responses. Database persistence was not exercised in this reproduction. |
| A particular real account returns HTTP 429 | Unverified | No real account or live provider process was used. |
3. Environment
Public get-bb/bb main at the full commit above; macOS Darwin arm64; Node 22.22.3; Corepack pnpm 9.15.0; bb-app package 0.43.4. Two separate clean clones, named base and verify, were checked at the same commit before adding only reproduction files. Frozen installs and full Turbo builds were run. The host pnpm launcher was broken, so a temporary Corepack shim supplied pnpm without changing dependencies or tracked configuration.
No BB server, network listener, database, or user runtime data was used. ACP ran over child-process stdio. Each test used a fresh temporary workspace, removed by the existing test cleanup. No real agent version applies to the synthetic fixture.
4. Minimal reproduction
- Clone and install trusted main at the recorded commit.
- Apply the authored regression insertion and add the synthetic ACP fixture.
- Run the focused test. It has a successful control and a rejecting-middle-model variant.
git clone https://github.com/get-bb/bb.git bb-repro cd bb-repro git checkout 94d77da09568d7f0f7cdb23b84686e743acfb0ad pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build curl -fsS https://get-bb.github.io/reports/issues/4154/repro/regression.patch -o /tmp/4154-regression.patch git apply /tmp/4154-regression.patch curl -fsS https://get-bb.github.io/reports/issues/4154/repro/issue-4154-agent.mjs -o packages/provider-bridge-acp/src/bridge/issue-4154-agent.mjs pnpm exec turbo run test --filter=@bb/provider-bridge-acp -- --testNamePattern='issue 4154'
Expected: the rejecting variant preserves low/medium/high for probe-c and requests all three models. Actual: probe-c receives medium only, and requests stop at probe-b.
probed: ["probe-a", "probe-b"] probe-a: low, medium, high probe-b: medium probe-c: medium Tests 1 failed | 1 passed | 348 skipped (350)
The test intentionally exits 1 on affected main. Download regression.patch, fixture, and commands.
Regression test insertion
it.each([false, true])("issue 4154 preserves later model efforts (reject middle: %s)", async (rejectMiddle) => {
const log = join(workspaceDir, "probes.txt");
const modelListId = sendModelList({
agent: { command: process.execPath, args: [resolve(dirname(FAKE_AGENT_PATH), "issue-4154-agent.mjs")] },
envVars: { REPRO_REQUEST_LOG: log, REPRO_REJECT_MIDDLE: rejectMiddle ? "1" : "0" },
});
const response = await waitForResponse(modelListId);
const probed = readFileSync(log, "utf8").trim().split("\n");
console.error(JSON.stringify({ rejectMiddle, probed, response }));
expect(response.result).toMatchObject({
models: [
{ id: "probe-a", supportedReasoningEfforts: [
{ reasoningEffort: "low" }, { reasoningEffort: "medium" }, { reasoningEffort: "high" },
] },
{ id: "probe-b", supportedReasoningEfforts: rejectMiddle
? [{ reasoningEffort: "medium" }]
: [{ reasoningEffort: "low" }, { reasoningEffort: "medium" }, { reasoningEffort: "high" }] },
{ id: "probe-c", supportedReasoningEfforts: [
{ reasoningEffort: "low" }, { reasoningEffort: "medium" }, { reasoningEffort: "high" },
] },
],
});
expect(probed).toEqual(["probe-a", "probe-b", "probe-c"]);
});
Synthetic ACP fixture
import { createInterface } from "node:readline";
import { appendFileSync } from "node:fs";
const models = ["probe-a", "probe-b", "probe-c"];
const configOptions = [
{ id: "model", category: "model", type: "select", currentValue: models[0],
options: models.map(value => ({ value, name: value })) },
{ id: "effort", category: "thought_level", type: "select", currentValue: "low",
options: ["low", "medium", "high"].map(value => ({ value })) },
];
createInterface({ input: process.stdin }).on("line", line => {
const request = JSON.parse(line);
let result;
if (request.method === "initialize") {
result = { protocolVersion: 1, agentCapabilities: {}, authMethods: [] };
} else if (request.method === "session/new") {
result = { sessionId: "repro-session", configOptions };
} else if (request.method === "session/set_config_option") {
appendFileSync(process.env.REPRO_REQUEST_LOG, request.params.value + "\n");
if (process.env.REPRO_REJECT_MIDDLE === "1" && request.params.value === models[1]) {
process.stdout.write(JSON.stringify({ jsonrpc: "2.0", id: request.id,
error: { code: -32603, message: "synthetic model rejection" } }) + "\n");
return;
}
result = { configOptions };
} else {
return;
}
process.stdout.write(JSON.stringify({ jsonrpc: "2.0", id: request.id, result }) + "\n");
});
5. Root cause
bridge.ts:976–1001 wraps the entire asynchronous loop in one catch. A rejection escapes the loop and returns the partial map without logging. The successful first model remains; neither the rejected model nor any later model is added.
for (const model of modelsToProbe) {
const configState = await args.connection.request(...);
supportByModel.set(model.value, ...);
}
...
} catch {
return supportByModel.size > 0 ? supportByModel : null;
}
model-catalog.ts:218–228 maps missing entries to the medium-only fallback. bridge.ts:904–920 caches and returns the catalog successfully. The server treats a resolved RPC as success (catalog store:467–474 and persists that successful catalog, subject to its normal host eligibility checks. The underlying issue is an error boundary spanning independent probes, not the rate limit itself.
6. Proposed fix
Catch rejection per model, log its model ID and error, and continue probing the remaining models. Keep the failed model on the existing fallback. Preserve the overall deadline and stop issuing requests after timeout or connection shutdown. Regression coverage should include first and middle failures, all failures, and deadline termination.
7. PR review
Open PR #4149
Static review only: its current diff places error handling inside the loop, logs the failed model, and checks a timeout flag before continuing after timeout. This addresses the reproduced root cause. Its three changed files cover the bridge and test fixture. The PR description mentions additional catalog-description work absent from the current diff; that prose is not treated as evidence. No PR code or tests were checked out or run, and this is not a merge-readiness verdict. No duplicate fix PR was opened because this open PR already covers the same root cause.
8. Related issues
Repository search also found #2503 concerning effort discovery and timeouts. This reproduction finishes well before the discovery deadline and triggers through an explicit rejection; it does not establish the cause or current status of that other issue.
9. Verification
The same agent repeated the reproduction in a second clean temporary clone at 94d77da09568d7f0f7cdb23b84686e743acfb0ad. Before inserting the same authored test, git status was clean. A separate frozen install and build were run, and the focused Turbo command was executed again with a new temporary test workspace. The control passed and the rejection variant failed with the same missing low/high choices and the same two-request log. No production files were changed. No report correction was required. This is a second clean run by the same agent, not an independent review.
10. Appendix
First run log · Second run log. Logs retain test output with temporary checkout paths normalized. Full builds are preparation checks, not proof of end-to-end UI behavior. The full existing test suite was not run because no production fix was made.
Issue content and linked PR metadata/diffs were treated as untrusted claims. No issue-provided code or command was executed. The regression and fixture were authored from trusted repository contracts. Real provider 429s, UI rendering, and stored account history are outside this report's verified scope.