#4073 · Provider model discovery and session synchronization

BugMedium priorityMedium effortprovidersprovider-claude-codeGitHub issue

2026-09-22 · trusted origin/main cb0bdbc3d843f200c46bda324cdf25f64b512290

Verdict: PARTIALLY REPRODUCED · Root-cause confidence: medium

1. TL;DR

The report describes a newer provider model missing from BB's picker and a session-local model change leaving the displayed selection stale. On current main, a new model supplied by discovery is correctly included; the reported discovery failure could not be reproduced from the available evidence. A focused test does prove that authoritative initialization model information and local-command output are discarded by the Claude translator. That removes these events as a way to synchronize BB's selection, but does not establish the actual event sequence emitted by the reporter's CLI. Both clean checkouts gave the same partial result; no live UI reproduction or screenshot was obtained.

2. Claims vs findings

ClaimStatusEvidence
A newly available model is absent from the pickerUnverifiedThe synthetic newly discovered model is included in the current-main catalog. The reporter's discovery response and executable resolution were not captured.
A session-local switch does not update BB's selectionPartially verifiedBoth runs show init model information and local-command output produce no outgoing deltas. The real slash-command sequence and rendered picker were not exercised.
The provider actually used the requested modelUnverifiedA conversational self-identification is not authoritative runtime evidence. No authenticated session was run.

3. Environment

Darwin arm64; Node 22.22.3; repository pnpm 9.15.0 through Corepack; locked Claude Agent SDK 0.3.245. Trusted target repository visibility is public. Source commit: cb0bdbc3d843f200c46bda324cdf25f64b512290. Checkouts: /tmp/bb-issue-4073-repro and /tmp/bb-issue-4073-verify. Frozen installation and the full Turbo build succeeded in each checkout (59 build tasks). The system pnpm launcher was broken; a temporary Corepack shim directory supplied the repository-pinned pnpm without modifying dependencies.

No dev app, provider login, user runtime data, ports, or application data directories were used. No screenshot is available: this report verifies the data path only, not the visual symptom. The updated CLI version and release environment described by the reporter were not reproduced.

4. Minimal reproduction

Clone the trusted repository and use the pinned commit. Save the test artifact at the path below. Ensure the repository's pnpm is available to child processes; on this host corepack enable --install-directory /tmp/issue-4073-tools and a temporary PATH prefix repaired that environment issue.

git clone https://github.com/get-bb/bb.git bb-4073
cd bb-4073
git checkout --detach cb0bdbc3d843f200c46bda324cdf25f64b512290
corepack pnpm install --frozen-lockfile --prefer-offline
corepack pnpm exec turbo run build
# Save the linked test as plugins/provider-claude-code/src/issue-4073.repro.test.ts
corepack pnpm exec turbo run test --force --filter=bb-plugin-provider-claude-code -- --run src/issue-4073.repro.test.ts
corepack pnpm exec turbo run test --filter=bb-plugin-provider-claude-code -- --run src/model-list.test.ts src/delta-translation.test.ts

The first three assertions characterize actual behavior. The final diagnostic assertion deliberately requires the reported initialization model to survive translation, without proposing a new event type. Expected: outgoing data retains the model identity. Actual output in both runs:

AssertionError: expected '[]' to contain 'claude-future-6'
Expected: "claude-future-6"
Received: "[]"
Tests  1 failed | 3 passed (4)

claude-future-6 is a synthetic test identifier taken from the repository's existing discovery-test pattern; it makes no claim about model availability. The local-command message follows the locked SDK type and is constructed locally, not copied from the issue.

import { describe, expect, it } from "vitest";
import type { SDKLocalCommandOutputMessage } from "@anthropic-ai/claude-agent-sdk";
import { buildClaudeCodeModels } from "./model-list.js";
import { createClaudeDeltaTranslator } from "./delta-translation.js";
import { loadFixture } from "./delta-test-harness.js";

const discoveredModel = "claude-future-6";

describe("issue 4073 data-path reproduction", () => {
  it("offers an uncurated model when discovery supplies it", () => {
    const catalog = buildClaudeCodeModels([{
      value: "future",
      resolvedModel: discoveredModel,
      displayName: "Future 6",
      description: "Newly discovered model",
    }]);
    expect(catalog.models.some((model) => model.model === discoveredModel)).toBe(true);
  });

  it("does not invent an undiscovered model in the curated fallback", () => {
    expect(buildClaudeCodeModels([]).models.some((model) => model.model === discoveredModel)).toBe(false);
  });

  it("drops local-command output before it can update the selection", () => {
    const translator = createClaudeDeltaTranslator({ sandboxEnabled: false });
    const message = {
      type: "system",
      subtype: "local_command_output",
      content: "Model selection changed within the provider session",
      uuid: "00000000-0000-4000-8000-000000000001",
      session_id: "repro-session",
    } satisfies SDKLocalCommandOutputMessage;
    expect(translator.translate({
      jsonrpc: "2.0",
      method: "sdk/message",
      params: { threadId: "repro-thread", message },
    }, { threadId: "repro-thread" })).toEqual([]);
  });

  it("preserves an authoritative initialization model in outgoing deltas", () => {
    const translator = createClaudeDeltaTranslator({ sandboxEnabled: false });
    const deltas = translator.translate({
      jsonrpc: "2.0",
      method: "sdk/message",
      params: {
        threadId: "repro-thread",
        message: { ...loadFixture("system-init.json"), model: discoveredModel },
      },
    }, { threadId: "repro-thread" });
    expect(JSON.stringify(deltas)).toContain(discoveredModel);
  });
});

5. Root cause and limits

Catalog construction appends newly discovered models to the curated list. The discovery probe obtains that list from SDK initialization. Therefore a stale curated catalog alone does not prove why a successfully discovered newer model would be absent. The curated fallback is fixed in source and cannot invent an undiscovered model. The two-minute bridge memo can also delay refresh, but no cache-specific failure was reproduced.

System-message translation has no model-selection handling for initialization or local-command output and falls through to an empty array. Visibility metadata explicitly classifies both as noise, so the envelope's unhandled-event fallback also emits nothing. Existing initialization tests already expect no events. This behavior is directly reproduced, with high confidence.

case "init":
...
case "local_command_output":
...
  return { kind: `sdk/system:${event.subtype}`, coverage: "noise" };

Live settings application sends BB-side model changes into the provider via setModel. The picker binding reads BB's supplied active/selected model. Discarding provider events explains why these particular signals cannot update that display; attributing the reporter's complete sequence to this mechanism remains an inference.

6. Proposed next test and fix boundary

Capture discovery and raw SDK messages from an isolated authenticated session using the affected executable, then compare BB's selected model before and after a provider-side change. Determine whether the runtime exposes a structured effective-model update and whether BB should display or persist that session override. Do not parse natural-language command output or simply add a catalog entry as a substitute for synchronization.

No PR was opened. The full reported failure was only partially reproduced, and choosing persistence semantics or adding a provider-to-BB selection contract exceeds the authorized simple-fix conditions. No production files changed, no source branch was pushed, and no linked open PR was found through issue timeline metadata or PR search.

7. Related issues

#3571 concerns catalog context variants; #1724 concerns duplicated provider metadata. Both were open when checked; neither is established as a duplicate of this runtime synchronization report.

8. Verification

The same agent repeated the reproduction in a second clean detached checkout at the identical trusted commit. Before adding the test artifact, git status --short was empty. A fresh frozen install and full build succeeded. The same test command used --force to avoid a cached result and again produced three passes and the one expected failure. No ports or data directories were needed. This is a repeat by the same agent, not independent verification. No result correction was necessary; the verdict remains partial.

The relevant existing suites passed: 57 tests in model-list.test.ts and delta-translation.test.ts. No fix was applied.

9. Appendix

Preparation used git fetch origin main, git rev-parse origin/main, and two detached git worktree add operations at the recorded SHA. Repository evidence was inspected with targeted text reads and searches. GitHub reads checked repository visibility, issue classification, labels, comments, related issue state, and linked PR metadata. The issue text was treated solely as untrusted claims; none of its commands or links were executed. No processes or dev data required cleanup.

> AGENT GENERATED