← reports

#4861 · Configured Codex model is ignored during default selection

Bug Medium priority Low effort providers provider-codex · GitHub issue · October 4, 2026 · base 09ff18a73

Verdict: REPRODUCED · Root-cause confidence: high · Scope: the default-selection defect at the real BB bridge boundary, with a hermetic Codex subprocess.

1. TL;DR

When BB needs a model default, its Codex bridge forwards the catalog's default flag without consulting the configured model. A catalog and an effective configuration can therefore disagree, and the catalog wins. The regression invokes the real bridge with a repository-owned fake app-server that exposes this mismatch; it fails on unchanged production code in two clean checkouts. The server subsequently chooses the model bearing that flag for a project without remembered execution defaults. This investigation does not reproduce the reporter's account-specific first-turn rejection or inspect a real user's configuration.

2. Claims vs findings

ClaimStatusEvidence
BB uses catalog default flags without reading effective Codex configuration.VerifiedThe real bridge test preserves the first model's default flag despite a different configured model. Both clean runs fail identically.
A new project with no remembered default uses the catalog default.Verified in sourceThe server's thread-create default resolver selects the first model with isDefault; no live BB thread is needed for the bridge reproduction.
A stock or custom native Codex catalog can disagree with effective configuration.Unverified against native CodexThe mismatch is controlled input to the hermetic subprocess, not an observation of the reporter's native process.
An unusable catalog entry produces an account or quota rejection.UnverifiedNo account credentials, model execution, or billing endpoints are accessed.
Explicit model selection is unaffected by this failure.Supported by source scopeThe defective path is default resolution. No end-to-end explicit model turn was executed.

3. Environment

4. Minimal reproduction

The regression patch below changes test/support files only. It does not alter production code. Save its contents as regression.patch outside the checkout, then:

git clone https://github.com/get-bb/bb.git bb-4861
cd bb-4861
git checkout --detach 09ff18a73bd0a8037f43c6ed86c28eb02228089c
pnpm install --frozen-lockfile --prefer-offline
pnpm exec turbo run build
git apply ../regression.patch
pnpm exec turbo run test --filter=bb-plugin-provider-codex -- src/bridge/bridge.model-list.test.ts

The configured-model case supplies two entries. The first is the catalog default, while the effective configuration names the second by its model slug rather than its unrelated identifier. The test sends model/list through the real JSON-RPC bridge and expects the second entry alone to be default.

Expected default flags: catalog-model=false, configured-model=true
Actual default flags:   catalog-model=true,  configured-model=false

FAIL src/bridge/bridge.model-list.test.ts > resolves catalog defaults with 'configured model'
Test Files  1 failed (1)
Tests       1 failed | 7 passed (8)

Test-only regression patch

diff --git a/plugins/provider-codex/src/bridge/bridge.model-list.test.ts b/plugins/provider-codex/src/bridge/bridge.model-list.test.ts
index eaec92e8f..9c77c1b47 100644
--- a/plugins/provider-codex/src/bridge/bridge.model-list.test.ts
+++ b/plugins/provider-codex/src/bridge/bridge.model-list.test.ts
@@ -37,6 +37,70 @@ it("reuses one initialized app-server across model catalog requests", async () =
   expect(second.result).toEqual(first.result);
 });
 
+it.each([
+  {
+    name: "configured model",
+    configRead: { config: { model: "configured-model" } },
+    configReadError: false,
+    expectedDefaults: [false, true],
+  },
+  {
+    name: "unset model",
+    configRead: { config: { model: null } },
+    configReadError: false,
+    expectedDefaults: [true, false],
+  },
+  {
+    name: "model outside the catalog",
+    configRead: { config: { model: "unlisted-model" } },
+    configReadError: false,
+    expectedDefaults: [true, false],
+  },
+  {
+    name: "unavailable configuration API",
+    configRead: { config: { model: "configured-model" } },
+    configReadError: true,
+    expectedDefaults: [true, false],
+  },
+  {
+    name: "malformed configuration response",
+    configRead: { config: { model: 42 } },
+    configReadError: false,
+    expectedDefaults: [true, false],
+  },
+])("resolves catalog defaults with $name", async (scenario) => {
+  const workDir = await mkdtemp(join(tmpdir(), "bb-codex-model-default-"));
+  temporaryDirectories.push(workDir);
+  const scriptPath = join(workDir, "script.json");
+  const catalog = [
+    { id: "catalog-id", model: "catalog-model", isDefault: true },
+    { id: "configured-id", model: "configured-model", isDefault: false },
+  ];
+  await writeFile(
+    scriptPath,
+    JSON.stringify({
+      modelList: { data: catalog },
+      configRead: scenario.configRead,
+      configReadError: scenario.configReadError,
+      turns: [],
+    }),
+  );
+  stubFakeCodexAppServer(scriptPath);
+
+  harness.sendRequest(1, "model/list", {});
+  const response = await harness.waitForResponse(1);
+
+  expect(response.error).toBeUndefined();
+  expect(response.result).toMatchObject({
+    models: catalog.map((model, index) => ({
+      id: model.id,
+      model: model.model,
+      isDefault: scenario.expectedDefaults[index],
+    })),
+    selectedOnlyModels: [],
+  });
+});
+
 it("replaces the cached app-server after a model catalog failure", async () => {
   const workDir = await mkdtemp(join(tmpdir(), "bb-codex-model-list-"));
   temporaryDirectories.push(workDir);
diff --git a/plugins/provider-codex/src/bridge/fake-codex-app-server.mjs b/plugins/provider-codex/src/bridge/fake-codex-app-server.mjs
index af2d2d83b..012cba0ab 100644
--- a/plugins/provider-codex/src/bridge/fake-codex-app-server.mjs
+++ b/plugins/provider-codex/src/bridge/fake-codex-app-server.mjs
@@ -428,6 +428,10 @@ async function handleRequest(message) {
         respond(id, { data: [] });
         return;
       }
+      if (script?.modelList) {
+        respond(id, script.modelList);
+        return;
+      }
       respond(id, {
         data: [
           {
@@ -444,6 +448,13 @@ async function handleRequest(message) {
         ],
       });
       return;
+    case "config/read":
+      if (script?.configReadError) {
+        respondError(id, -32601, "Configuration read unavailable");
+      } else {
+        respond(id, script?.configRead ?? { config: { model: null } });
+      }
+      return;
     case "skills/extraRoots/set":
       respond(id, {});
       return;

Complete regression test, also retained locally with the raw evidence:

import { mkdtemp, rm, writeFile } from "node:fs/promises";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { afterEach, beforeEach, expect, it, vi } from "vitest";
import { BRIDGE_JSON_RPC_ERRORS } from "@get-bb/plugin-sdk/provider-bridge";
import { experimental_createBridgeJsonRpcTestHarness as createBridgeJsonRpcTestHarness } from "@get-bb/plugin-sdk/provider-bridge/testing";
import { experimental_killAllChildrenForTests, handleLine } from "./bridge.js";
import { stubFakeCodexAppServer } from "./fake-codex-app-server-harness.js";

let harness: ReturnType<typeof createBridgeJsonRpcTestHarness>;
const temporaryDirectories: string[] = [];

beforeEach(() => {
  stubFakeCodexAppServer();
  harness = createBridgeJsonRpcTestHarness(handleLine);
});

afterEach(async () => {
  experimental_killAllChildrenForTests();
  harness.restore();
  vi.unstubAllEnvs();
  await Promise.all(
    temporaryDirectories
      .splice(0)
      .map((path) => rm(path, { recursive: true, force: true })),
  );
});

it("reuses one initialized app-server across model catalog requests", async () => {
  harness.sendRequest(1, "model/list", {});
  const first = await harness.waitForResponse(1);
  harness.sendRequest(2, "model/list", {});
  const second = await harness.waitForResponse(2);

  expect(first.error).toBeUndefined();
  expect(second.error).toBeUndefined();
  expect(second.result).toEqual(first.result);
});

it.each([
  {
    name: "configured model",
    configRead: { config: { model: "configured-model" } },
    configReadError: false,
    expectedDefaults: [false, true],
  },
  {
    name: "unset model",
    configRead: { config: { model: null } },
    configReadError: false,
    expectedDefaults: [true, false],
  },
  {
    name: "model outside the catalog",
    configRead: { config: { model: "unlisted-model" } },
    configReadError: false,
    expectedDefaults: [true, false],
  },
  {
    name: "unavailable configuration API",
    configRead: { config: { model: "configured-model" } },
    configReadError: true,
    expectedDefaults: [true, false],
  },
  {
    name: "malformed configuration response",
    configRead: { config: { model: 42 } },
    configReadError: false,
    expectedDefaults: [true, false],
  },
])("resolves catalog defaults with $name", async (scenario) => {
  const workDir = await mkdtemp(join(tmpdir(), "bb-codex-model-default-"));
  temporaryDirectories.push(workDir);
  const scriptPath = join(workDir, "script.json");
  const catalog = [
    { id: "catalog-id", model: "catalog-model", isDefault: true },
    { id: "configured-id", model: "configured-model", isDefault: false },
  ];
  await writeFile(
    scriptPath,
    JSON.stringify({
      modelList: { data: catalog },
      configRead: scenario.configRead,
      configReadError: scenario.configReadError,
      turns: [],
    }),
  );
  stubFakeCodexAppServer(scriptPath);

  harness.sendRequest(1, "model/list", {});
  const response = await harness.waitForResponse(1);

  expect(response.error).toBeUndefined();
  expect(response.result).toMatchObject({
    models: catalog.map((model, index) => ({
      id: model.id,
      model: model.model,
      isDefault: scenario.expectedDefaults[index],
    })),
    selectedOnlyModels: [],
  });
});

it("replaces the cached app-server after a model catalog failure", async () => {
  const workDir = await mkdtemp(join(tmpdir(), "bb-codex-model-list-"));
  temporaryDirectories.push(workDir);
  const scriptPath = join(workDir, "script.json");
  await writeFile(
    scriptPath,
    JSON.stringify({
      modelListFailOnceMarkerPath: join(workDir, "failed-once"),
      turns: [],
    }),
  );
  stubFakeCodexAppServer(scriptPath);

  harness.sendRequest(1, "model/list", {});
  const failed = await harness.waitForResponse(1);
  harness.sendRequest(2, "model/list", {});
  const recovered = await harness.waitForResponse(2);

  expect(failed.error?.message).toContain(
    "Codex model/list returned no supported models.",
  );
  expect(recovered.error).toBeUndefined();
  expect(recovered.result).toMatchObject({
    models: [
      {
        displayName: "Fake model",
      },
    ],
  });
});

it("rejects a model catalog request with the missing-executable code when codex cannot be spawned", async () => {
  vi.stubEnv(
    "BB_CODEX_BRIDGE_APP_SERVER_COMMAND",
    join(tmpdir(), "bb-codex-does-not-exist"),
  );
  vi.stubEnv("BB_CODEX_BRIDGE_APP_SERVER_ARGS", "[]");

  harness.sendRequest(1, "model/list", {});
  const response = await harness.waitForResponse(1);

  expect(response.error?.code).toBe(BRIDGE_JSON_RPC_ERRORS.MISSING_EXECUTABLE);
  expect(response.error?.message).toContain("could not find the Codex CLI");
});

5. Root cause

Three trusted code locations explain the result:

  1. models.ts, lines 143–164 parses identities and copies the upstream flag without any configured-model input: isDefault: raw.isDefault === true.
  2. bridge.ts, lines 1459–1472 requests only model/list and returns models: parseModelsResponse(result). No configuration is consulted in this path.
  3. thread-create.ts, lines 105–125 chooses catalog.models.find((model) => model.isDefault) ?? catalog.models[0] when resolving defaults.

The parser is faithfully translating catalog metadata, but the bridge treats that metadata as the effective execution preference. The bridge is the local owner of Codex app-server translation, so the repair belongs there; no server/daemon or public bridge wire fields need to change.

6. Proposed fix and validation

Read config/read over the existing initialized model-list connection, validate only config.model at that boundary, and change default flags only if the configured model matches a listed model slug. Retain the catalog's existing defaults when configuration is absent, unavailable, invalid, or points outside the catalog. This protects older servers and keeps a valid catalog usable despite a configuration read error.

The local implementation changes three files in the existing provider-codex subsystem: 93 additions and one deletion, 94 total changed text lines. It adds no dependency, generated file, migration, stored-data change, public protocol, or authentication/permission change.

Before fix, both clean runs: 1 failed | 7 passed (8)
After fix, focused suite:    8 passed (8)
Relevant existing suite:    32 files passed; 340 tests passed
pnpm exec turbo run test lint typecheck --filter=bb-plugin-provider-codex
git diff --check

The parameterized regression additionally verifies unset configuration, an unlisted model, a configuration RPC error, and malformed configuration. Existing model-catalog recovery, missing executable, and connection reuse tests also pass.

7. Verification

The same agent repeated the reproduction in a second clean detached worktree at the full recorded SHA. That checkout received its own frozen install and successful full Turbo build, followed by the saved test-only patch and the exact focused command in section 4. The second run exits 1 and produces the same two reversed default flags and the same single failing configured-model case. No production edits were present in either pre-fix run. This is a second clean run, not an independent review. No report correction was necessary.

The second checkout uses no ports or runtime data directory; test fixture paths are fresh temporary directories. Code permalinks above were checked against the trusted base files and line numbers. This nonvisual bug requires no screenshots.

8. Related work

No linked open pull request was present in the issue's cross-reference metadata or the open-PR search when investigation and implementation began. No contributor branch, linked patch, script, or external issue URL was fetched or executed. The fix was developed from trusted repository source and the observed regression.

9. Appendix

Raw logs and test-support files are retained locally, not committed to this public report repository. The inline test, test-only patch, and observed results make the report self-contained.

Other setup/check commands: git fetch origin main; git rev-parse origin/main; git ls-remote for the target repository main; git worktree add --detach for each checkout; uname -sr; node --version; pnpm --version; git diff --numstat origin/main. GitHub classification and open-PR metadata were read through the GitHub API. Both full Turbo builds completed successfully.

Trust boundary: Issue content was treated only as untrusted claims. No embedded instructions, contributor branches, or account data were used. All reproduction code is trusted main plus the agent's test-only changes. Account-specific rejection and native Codex configuration/catalog behavior remain outside the verified scope.

> AGENT GENERATED