#4082 · Model discovery and catalog freshness

Bug · Priority: Medium · Effort: Medium · providers · provider-claude-code · partial-repro
2026-09-22 · base 6ecd23570b70c12653c0bb457c7c2d50eecdf227
GitHub issue

Verdict: PARTIALLY REPRODUCED · Root-cause confidence: medium

1. TL;DR

A model absent from both the curated catalog and the CLI discovery response cannot appear in the returned list. The committed catalog lacks the newly reported model, and its dependency range remains unchanged. However, current main explicitly discovers an installed Claude executable and supplies its path to the SDK probe; it does not invariably use the bundled CLI. A synthetic model returned by that probe reaches the model list and becomes the default when discovery marks it as such. The reporter's account-specific response and visible picker failure were not reproduced.

2. Claims vs findings

ClaimFindingEvidence
The curated catalog needs a code update for a new fallback entry.VerifiedStatic catalog data; an empty discovery response adds no unknown entries.
Updating an installed CLI cannot affect discovery because the bundled CLI is always used.Refuted on trusted mainExecutable resolver checks explicit configuration, PATH, then known installation paths. The probe forwards the result.
Unknown models can be appended.VerifiedBoth clean runs retain a synthetic unknown model returned at the SDK boundary.
The SDK dependency range is behind the reported release.Partly verifiedManifest range is ^0.3.245. Release chronology, claimed newer version behavior and account availability were not verified.
The reported new model is missing from a real user's picker after their CLI update.UnverifiedNo authenticated CLI session, reporter environment, or UI capture used.

3. Environment

macOS Darwin arm64; Node v22.22.3; Vitest 4.1.1; provider package 0.1.0. Two fresh detached worktrees, named base and verify, at the full commit above. No server, ports, live provider calls, user credentials, or persistent runtime data were used. SDK query is mocked; the production resolver, probe wrapper and model-list builder run unchanged. The executable fixture is never launched.

The host pnpm launcher was broken. A temporary PATH shim delegated pnpm to Corepack. Frozen installs then succeeded in both checkouts. The full base build passed (59 tasks); targeted Turbo verification built required dependencies in the second checkout.

4. Minimal reproduction

Download the test artifact. Repeat these commands in two separate clones. The test checks the actual bridge call options and returned list, without making an upstream service request.

git clone https://github.com/get-bb/bb.git bb-repro
cd bb-repro
git checkout --detach 6ecd23570b70c12653c0bb457c7c2d50eecdf227
corepack pnpm install --frozen-lockfile --prefer-offline
corepack pnpm exec turbo run build
# Copy the linked test into plugins/provider-claude-code/src/issue-4082.test.ts
corepack pnpm exec turbo run test --force --filter=bb-plugin-provider-claude-code -- issue-4082.test.ts

Expected and actual in each checkout:

Test Files  1 passed (1)
Tests  2 passed (2)

The positive discovery check passes on unchanged production code. The second check confirms the bounded absence case, not a regression: the builder cannot synthesize unknown models from empty discovery.

import { mkdtempSync, writeFileSync, chmodSync, rmSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { expect, it, vi } from "vitest";
import { buildClaudeCodeModels } from "./model-list.js";
import { listClaudeCodeBridgeModels } from "./bridge/model-list.js";

const queryMock = vi.hoisted(() => vi.fn());
vi.mock("@anthropic-ai/claude-agent-sdk", () => ({ query: queryMock }));

it("includes a previously unknown model returned by the installed CLI probe", async () => {
  const dir = mkdtempSync(join(tmpdir(), "provider-discovery-"));
  const executable = join(dir, "claude");
  writeFileSync(executable, "#!/bin/sh\nexit 0\n");
  chmodSync(executable, 0o755);
  const close = vi.fn();
  queryMock.mockReturnValue({
    initializationResult: async () => ({ models: [{
      value: "default",
      resolvedModel: "claude-future-6",
      displayName: "Future 6",
      description: "Synthetic discovery fixture",
    }] }),
    close,
  });
  try {
    const result = await listClaudeCodeBridgeModels({ PATH: dir, HOME: dir });
    expect(queryMock).toHaveBeenCalledWith(expect.objectContaining({
      options: expect.objectContaining({ pathToClaudeCodeExecutable: executable }),
    }));
    expect(result.models).toContainEqual(expect.objectContaining({
      model: "claude-future-6", isDefault: true,
    }));
    expect(close).toHaveBeenCalledOnce();
  } finally {
    rmSync(dir, { recursive: true, force: true });
    vi.clearAllMocks();
  }
});

it("cannot invent an unknown model absent from discovery and the curated catalog", () => {
  expect(buildClaudeCodeModels([]).models.some(model => model.model === "claude-future-6")).toBe(false);
});

5. Root cause

The verified mechanism is a finite curated catalog merged with discovery results. If a model is absent from both inputs, it is absent from the output. This explains the fallback limitation, but does not establish why a particular real CLI failed to report an available model.

const pathToClaudeCodeExecutable = resolveClaudeCodeExecutable({ env });
...(pathToClaudeCodeExecutable ? { pathToClaudeCodeExecutable } : {})
const initialization = await session.initializationResult();
return buildClaudeCodeModels(initialization.models);

6. Proposed next test and fix gate

In an isolated authenticated environment, record the executable selected by the host, that executable's version, and its sanitized initialization model response. Compare direct CLI availability with that response. If the model is returned, trace downstream cache and picker behavior; if it is absent, investigate CLI selection and upstream discovery first. A catalog refresh requires verified model capabilities; changing the default is a product decision. Dependency or release automation changes fall outside this rule's simple-fix scope.

No production fix or PR: no failing regression demonstrates the claimed discovery defect on trusted main. No linked open PR was found in cross-reference metadata or the issue-number PR search.

7. Verification

The same agent ran the exact artifact in a second clean detached checkout at 6ecd23570b70c12653c0bb457c7c2d50eecdf227. The verification run and the forced base run each executed two tests successfully with zero cache hits. Separate temporary executable directories were created and deleted by each run. This is a repeated check by the same agent, not independent verification. The findings narrow the issue's explanation: installed CLI selection works at the tested boundary. No claims of real upstream availability or UI reproduction are made.

8. Related issues

Repository search also found issue #4073 covering a similar missing-model symptom. It was used only as triage context, not reproduction evidence. Existing catalog-choice requests show that model policy and discovery behavior are distinct concerns.

9. Appendix

Source links were checked against the recorded base. A final origin/main fetch advanced to 78804e79d with no changes in the provider subsystem since the recorded base. Issue text was treated as untrusted claims; no suggested patch, fork code, script, or external issue-supplied URL was executed or fetched. No visual artifact is included because this is a provider-boundary investigation, not a demonstrated visual reproduction.