#3899 · ACP reasoning catalog fallback
PARTIALLY REPRODUCED · Root-cause confidence: medium
1. TL;DR
BB exposes a single managed medium level when reasoning metadata is missing. A direct test of the trusted catalog code reproduces this result in two clean checkouts. The same test confirms that low, medium, high and xhigh are retained when advertised. This establishes the fallback mechanism, but does not establish why the reporter's provider supplies or receives incomplete metadata. No live provider, native TUI, CLI rejection or UI interaction was reproduced.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| Catalog can expose only managed medium | Verified conditionally | Missing thought-level metadata and missing per-model probe results both produce medium. |
| Provider-specific models lose native variants | Unverified | No provider wire capture or native TUI run; synthetic model IDs used. |
| Four advertised levels are lost in mapping | Not observed in control | The mapper retains all four recognized levels. |
| Picker does nothing and CLI rejects high | Unverified | No live application started. |
3. Environment
Trusted get-bb/bb origin/main at 8d35c2776bc7fd33f41a64e4c6dd7ebd2f021887; Darwin arm64; Node v22.22.3. Frozen dependency installation succeeded through Corepack with a temporary pnpm wrapper, after the default pnpm executable failed with MODULE_NOT_FOUND. Repository build: 58 successful tasks. No dependencies added. No real provider, ports, credentials or runtime data used.
4. Minimal reproduction
- Create a clean checkout at the base commit.
- Run
corepack pnpm install --frozen-lockfile --prefer-offlineandcorepack pnpm exec turbo run buildwith a working pnpm executable on PATH. - Copy the test artifact to
packages/provider-bridge-acp/src/bridge/issue-3899.test.ts. - Run
corepack pnpm exec turbo run test --filter=@bb/provider-bridge-acp -- --testNamePattern='records the missing-control'.
Expected and actual: all three catalog assertions pass. This is a characterization test of a conditional fallback, not a failing regression demonstrating the provider-specific bug.
missing control: medium unprobed model: medium advertised control: low,medium,high,xhigh Test Files 1 passed | 18 skipped (19) Tests 1 passed | 319 skipped (320)
import { expect, it } from "vitest";
import { buildAcpNativeReasoningSupport, buildModelCatalogFromConfigOptions, findAcpThoughtLevelConfigOption } from "./model-catalog.js";
it("records the missing-control fallback and preserves advertised levels", () => {
const missing = buildAcpNativeReasoningSupport(findAcpThoughtLevelConfigOption([]));
expect(missing.supportedReasoningEfforts.map(e => e.reasoningEffort)).toEqual(["medium"]);
const advertised = buildAcpNativeReasoningSupport({
id: "effort", category: "thought_level", type: "select", currentValue: "medium",
options: ["low", "medium", "high", "xhigh"].map(value => ({ value })),
});
expect(advertised.supportedReasoningEfforts.map(e => e.reasoningEffort)).toEqual(["low", "medium", "high", "xhigh"]);
const models = buildModelCatalogFromConfigOptions({
id: "model", category: "model", type: "select", currentValue: "synthetic/a",
options: [{value: "synthetic/a"}, {value: "synthetic/b"}],
}, new Map([["synthetic/b", advertised]]));
expect(models.map(m => m.supportedReasoningEfforts.map(e => e.reasoningEffort))).toEqual([["medium"], ["low", "medium", "high", "xhigh"]]);
process.stdout.write("missing control: medium; unprobed model: medium; advertised control: low,medium,high,xhigh\n");
});
5. Root cause
buildAcpNativeReasoningSupport maps recognized thought-level options and uses the fixed medium fallback only when the control is absent. An explicitly present empty control yields no supported levels. Catalog construction also applies medium when the per-model support map lacks an entry. Thus the observed single level alone cannot distinguish missing upstream metadata from incomplete probing.
Discovery selects each model with session/set_config_option and reads the returned thought-level options. It returns a partial map after errors or the five-second discovery budget. Models not reached retain the fallback. This timeout path is a source finding, not reproduced timing evidence. The OpenCode dialect only customizes command-event normalization; it does not independently discover native variants.
Reasoning selection needs an advertised or configured control to send a value back to ACP. Merely adding catalog levels could expose choices that never affect execution. Confidence is high in the tested fallback mechanism, medium in its relevance to this issue, and low in selecting the upstream-omission versus incomplete-probe explanation.
6. Proposed next test and fix scope
Capture isolated session/new and model-selection responses for an affected model, along with probe completion and timing. If the response advertises levels but the catalog lacks them, address that proven discovery failure and add a failing regression. If the agent omits a usable control, establish its supported selection mechanism before expanding the catalog. Do not hard-code levels based on model family. No PR: the provider-specific failure is unverified, so the simple-fix reproduction and failing-regression requirements are not met.
7. Verification
The same agent repeated the test in a second freshly created detached worktree at the exact base commit. It used its own frozen dependency installation and no application state. The targeted Turbo run executed with zero cached tasks in both checkouts and passed once in each. The first checkout additionally ran the existing catalog suite plus the artifact: 32 tests passed. No corrections were necessary; the report deliberately limits the verdict to partial reproduction.
First run · Second run · Catalog suite · Build
8. Related issues and PRs
No open linked PR was found in closing references, timeline cross-references, or an open-PR search for this issue number. Other issue claims were not used as evidence.
9. Appendix
Only trusted origin/main source and the locally authored test were executed. Issue content was treated as untrusted evidence; supplied commands and external links were not executed or fetched. Artifact logs replace machine-specific paths. No screenshots are included because the visual symptom was not exercised. The repository and issue were classified via GitHub properties; existing classification values were absent.