#4367 · Reasoning discovery deadline produces an actionable fallback

Bug · Priority: Medium · Effort: Medium · providers · provider-acp · 2026-09-25

GitHub issue · Base 9c9bae7f36a237c7e1b96de3d4c2186d13967686

Verdict: REPRODUCED · Root-cause confidence: high

1. TL;DR

A slow ACP model catalog exhausts the bridge’s discovery deadline before every model’s reasoning choices are read. Unvisited models are returned with a single Medium choice, even when the agent supports several levels. That fallback is also actionable: starting a session with the catalog’s advertised medium value sets the agent’s real effort to medium. This was reproduced through the real bridge and the repository’s synthetic ACP process twice; the reporter’s authenticated provider inventory and browser appearance were not tested.

2. Claims vs findings

ClaimFindingEvidence
Slow catalogs lose per-model levelsVerified60 models at 120 ms per switch exceed the 5 s budget; the final model has only Medium.
The fallback can configure real effortVerifiedA new session for the final model echoes selected-effort:medium after starting with medium.
Exact model counts and latency in the reported provider installationUnverifiedNo external provider or account was used. Coverage is time-dependent, not a fixed 80-model cap.
Proposed external metadata extension solves the issueUnverifiedNo external branch, patch, or URL was fetched or run.

3. Environment

Trusted origin/main at the commit above; Linux 7.2.5-3-omarchy x86_64; Node v26.8.1. Both clean detached checkouts completed frozen installs and Turbo builds (60 successful build tasks each). No app server, browser, authenticated provider, or user runtime data was used. Tests create fresh temporary workspaces and local bridge transport endpoints and clean them through the existing harness.

4. Minimal reproduction

Use a clean checkout of get-bb/bb at the recorded commit. Save the inline reproduction patch below as repro.patch outside the repository, then run:

git checkout 9c9bae7f36a237c7e1b96de3d4c2186d13967686
pnpm install --frozen-lockfile --prefer-offline
pnpm exec turbo run build
git apply /path/to/repro.patch
pnpm exec turbo run test --force --filter=@bb/provider-bridge-acp -- src/bridge/bridge.test.ts -t 'issue 4367' --testTimeout=15000

The patch adds a configurable delay to the repository’s fake agent and extends its existing large-catalog bridge test. No production code changes. The original fast catalog test did not exercise exhaustion of the real discovery budget.

Expected: fake/gen-59 advertises Low, Medium, High. Actual in both runs:

supportedReasoningEfforts: [{"reasoningEffort":"medium","description":"Reasoning effort is managed by the connected ACP agent."}]
defaultReasoningEffort: "medium"
session-effort selected-effort:medium
AssertionError: expected [ { reasoningEffort: 'medium', …(1) } ] to deeply equal [ …(3) ]
Tests  1 failed | 100 skipped (101)

The session check explicitly supplies the observed catalog default, medium; it does not exercise the composer. The agent fixture supports low/medium/high for that model, and its initial choice is low. The assertion about the session passes before the catalog assertion fails.

Test excerpt (apply the patch for the executable test with its existing harness):

  it("issue 4367 preserves reasoning across slow model catalogs", async () => {
    const modelListId = sendModelList({
      envVars: {
        FAKE_ACP_MODEL_CONFIG: "1",
        FAKE_ACP_THOUGHT_LEVEL_CONFIG: "1",
        FAKE_ACP_MODEL_COUNT: "60",
        FAKE_ACP_MODEL_DELAY_MS: "120",
      },
    });

    const result = (await waitForResponse(modelListId)).result as {
      models: {
        id: string;
        supportedReasoningEfforts: { reasoningEffort: string }[];
      }[];
    };
    expect(result.models).toHaveLength(60);
    const lastGenerated = result.models.find(
      (model) => model.id === "fake/gen-59",
    );
    console.log("catalog-tail", JSON.stringify(lastGenerated));
    const { providerThreadId } = await startThread({
      envVars: {
        FAKE_ACP_MODEL_CONFIG: "1",
        FAKE_ACP_THOUGHT_LEVEL_CONFIG: "1",
        FAKE_ACP_MODEL_COUNT: "60",
      },
      model: "fake/gen-59",
      reasoningLevel: "medium",
    });
    sendTurnRequest("turn/start", providerThreadId, {
      input: [{ type: "text", text: "echo-selected-effort", mentions: [] }],
    });
    await waitForTurnCompleted();
    expect(agentMessageTexts()).toContain("selected-effort:medium");
    console.log("session-effort", agentMessageTexts().join(";"));
    expect(lastGenerated?.supportedReasoningEfforts).toEqual([
      { reasoningEffort: "low", description: "low" },
      { reasoningEffort: "medium", description: "medium" },
      { reasoningEffort: "high", description: "high" },
    ]);
  });

5. Root cause

Sequential discovery processes priority models then catalog order. Its five-second timer kills the connection and returns only support accumulated so far. It does not put the current model first unless it is already a configured priority.

Catalog construction replaces each missing map entry with the generic medium effort and medium default. Missing knowledge therefore becomes a selectable concrete value. Session reasoning selection maps that supplied medium to the live thought-level option and sends session/set_config_option.

const reasoning = reasoningByModel?.get(option.value) ?? {
  supportedReasoningEfforts: ACP_NATIVE_REASONING_EFFORTS,
  defaultReasoningEffort: "medium" as ReasoningLevel,
};

6. Proposed fix and simple-fix assessment

Keep unknown reasoning support distinct from a user-selected medium, and obtain capabilities for the selected model before offering choices. Decide how to refresh incomplete catalogs within a bounded discovery latency. Bulk per-model metadata could help only with an agreed agent contract. Increasing the timeout alone moves the cutoff and may delay discovery substantially.

No PR: a complete correction needs a product/architecture decision about unknown capabilities and discovery latency, or a new metadata contract. That fails the automation’s simple-fix conditions. No production fix was made or pushed. Git diff --check passed for the two-file reproduction (20 added/deleted lines).

7. Related issues

#4154 concerns early termination after a failed probe; this report independently demonstrates termination by elapsed time. The issue timeline links #2062, concerning model display distinctions. No open PR referencing #4367 was found in issue timeline metadata or the open-PR search.

8. Verification

The same agent repeated the reproduction in a second clean detached checkout named verify at the same full commit. It performed a separate frozen install and build, applied only the reproduction patch, and ran the command above with --force to bypass cached test results. Both runs returned the same Medium-only final model, applied medium to the live test-agent session, and failed the same catalog assertion. This is a repeat check by the same agent, not an independent review.

Initial attempts hit Vitest’s default five-second test timeout; the final command explicitly allows 15 seconds so the product’s five-second cutoff can finish and the assertion can run. Those harness timeouts were not counted as bug evidence. A first checkout attempt also hit temporary-filesystem inode exhaustion; both final runs used clean checkouts on a filesystem with capacity.

9. Appendix

Complete reproduction patch, authored from trusted repository code:

diff --git i/packages/provider-bridge-acp/src/bridge/bridge.test.ts w/packages/provider-bridge-acp/src/bridge/bridge.test.ts
index 1745d27de..1b9156657 100644
--- i/packages/provider-bridge-acp/src/bridge/bridge.test.ts
+++ w/packages/provider-bridge-acp/src/bridge/bridge.test.ts
@@ -996,12 +996,13 @@ describe("acp bridge", () => {
     });
   });
 
-  it("probes per-model reasoning across large catalogs instead of falling back", async () => {
+  it("issue 4367 preserves reasoning across slow model catalogs", async () => {
     const modelListId = sendModelList({
       envVars: {
         FAKE_ACP_MODEL_CONFIG: "1",
         FAKE_ACP_THOUGHT_LEVEL_CONFIG: "1",
         FAKE_ACP_MODEL_COUNT: "60",
+        FAKE_ACP_MODEL_DELAY_MS: "120",
       },
     });
 
@@ -1015,6 +1016,22 @@ describe("acp bridge", () => {
     const lastGenerated = result.models.find(
       (model) => model.id === "fake/gen-59",
     );
+    console.log("catalog-tail", JSON.stringify(lastGenerated));
+    const { providerThreadId } = await startThread({
+      envVars: {
+        FAKE_ACP_MODEL_CONFIG: "1",
+        FAKE_ACP_THOUGHT_LEVEL_CONFIG: "1",
+        FAKE_ACP_MODEL_COUNT: "60",
+      },
+      model: "fake/gen-59",
+      reasoningLevel: "medium",
+    });
+    sendTurnRequest("turn/start", providerThreadId, {
+      input: [{ type: "text", text: "echo-selected-effort", mentions: [] }],
+    });
+    await waitForTurnCompleted();
+    expect(agentMessageTexts()).toContain("selected-effort:medium");
+    console.log("session-effort", agentMessageTexts().join(";"));
     expect(lastGenerated?.supportedReasoningEfforts).toEqual([
       { reasoningEffort: "low", description: "low" },
       { reasoningEffort: "medium", description: "medium" },
diff --git i/packages/provider-bridge-acp/src/bridge/fake-acp-agent.mjs w/packages/provider-bridge-acp/src/bridge/fake-acp-agent.mjs
index 4704b8018..3f3043d20 100755
--- i/packages/provider-bridge-acp/src/bridge/fake-acp-agent.mjs
+++ w/packages/provider-bridge-acp/src/bridge/fake-acp-agent.mjs
@@ -774,6 +774,7 @@ async function handleMessage(message) {
           });
           return;
         }
+        await new Promise((resolve) => setTimeout(resolve, Number(process.env.FAKE_ACP_MODEL_DELAY_MS ?? "0")));
         selectedModel = value;
         send({ jsonrpc: "2.0", id: message.id, result: configState() });
         return;

Both final runs produced the expected-versus-actual output shown above. Full local logs remain in the investigation directory; repository policy requires inline evidence rather than published raw logs or test artifacts. No screenshot is supplied because this is protocol-level evidence, not a browser reproduction. Issue content was treated as untrusted claims; no issue-provided commands or external prototype code were executed.

> AGENT GENERATED