#4318 · Codex fallback selection

Bug · Priority Medium · Effort Low · providers · provider-codex · no-repro

Issue · 2026-09-25 · Base 0baa605b32a00619c1d7e3f32be6553ebcf8244a

ALREADY FIXED · Root-cause confidence: high

TL;DR

The reported failure concerns automatic titles after a transient Codex inference error. Earlier main selected a fallback that the reporter says its ChatGPT account rejected. Trusted main now selects a different model pair, following merged PR #4327. Two clean-checkout runs confirm the rejected model is never selected in the tested retry paths. These are code-level tests with simulated host responses, not live account entitlement checks.

Claims vs findings

ClaimFindingEvidence
Current main retries the rejected modelRefuted on recorded baseThe service model list changed in merged commit 3b319202f250b015c278c1894a398a13c0118b9b.
Transient failures trigger fallbackVerifiedFocused tests exercise rate_limited, service_unavailable and invalid_response.
Failure can leave a thread without a generated titleVerified in codeThe title generator returns null on failed inference.
ChatGPT rejects the previous fallback, with the reported frequencyUnverified liveNo credentials, real accounts, deployed logs or database were accessed.

Environment

Linux x86_64, Node v26.8.1, pnpm frozen install, Vitest 4.1.1. Two fresh detached worktrees at the recorded base; no server, ports, runtime data directory or provider process. Both full Turbo builds passed: 60/60 tasks. An initial temporary checkout failed because the system temporary filesystem lacked inodes; fresh checkouts on the home filesystem succeeded.

Minimal reproduction

  1. Check out the recorded trusted main commit in a clean worktree.
  2. Save the focused test shown below as plugins/provider-codex/src/issue-4318.test.ts.
  3. Run:
    pnpm install --frozen-lockfile --prefer-offline
    pnpm exec turbo run build
    pnpm exec turbo run test --filter=bb-plugin-provider-codex --force -- --run src/ai-service.test.ts src/issue-4318.test.ts

Expected on fixed main: a retry selects a different model, avoids the rejected fallback and returns the simulated title. Actual in both runs:

Test Files  2 passed (2)
     Tests  8 passed (8)

Three new cases and five existing tests passed. The new test models rejection as a hypothesis; it does not establish which models a real account accepts.

import { createFakePluginHost } from "@get-bb/plugin-sdk/testing";
import { expect, it } from "vitest";
import { registerCodexAiService } from "./ai-service.js";
import { codexAiCompleteInputSchema } from "./ai/host-contract.js";

it.each(["rate_limited", "service_unavailable", "invalid_response"] as const)(
  "does not select the rejected fallback after %s",
  async (code) => {
    const models: string[] = [];
    const { bb, harness } = createFakePluginHost({
      sdk: { system: { config: async () => ({ primaryHostId: "test-host" }) } },
      experimental_callHostRpc: (call) => {
        const { model } = codexAiCompleteInputSchema.parse(call.input);
        models.push(model);
        if (models.length === 1) return { ok: false, code, message: "Retryable failure" };
        if (model === "gpt-5.4-mini") return { ok: false, code: "request_failed", message: "Unsupported fallback" };
        return { ok: true, text: "Generated title" };
      },
    });
    registerCodexAiService(bb);
    const complete = harness.registrations.aiServiceRegistrations[0]?.complete;
    if (!complete) throw new Error("Missing completion service");
    await expect(complete("Generate a short thread title", {
      signal: new AbortController().signal,
    })).resolves.toBe("Generated title");
    expect(models).toHaveLength(2);
    expect(models).not.toContain("gpt-5.4-mini");
    expect(models[0]).not.toBe(models[1]);
  },
);

Root cause

The hardcoded model list formerly included the rejected fallback. The retry loop advances for selected transient failures and throws if the final result fails. The title generator returns null for a failed task. This explains the reported symptom conditional on the reported backend rejection.

The ChatGPT request path and API-key path use different endpoints. This report does not infer universal model availability from that fact. Current main's model list removes the specific rejected fallback.

Proposed fix

Use a build containing merged commit 3b319202f250b015c278c1894a398a13c0118b9b. No additional patch is justified for the specific stale model selection on current main. For runtime assurance, the next experiment is an authorized live ChatGPT-account fallback request. Both candidate models can still fail for quota or availability reasons; successful title generation is not guaranteed.

PR review

#4327 is merged. Its two-file diff replaces the model pair and updates the existing retry regression test. Static review found that it removes the exact fallback selection implicated here. Eight relevant tests pass on trusted main in each checkout. No PR branch was checked out or executed; live model access remains outside this verification.

Verification

The same agent repeated the test in a second freshly created detached worktree at the same base. Frozen install and full build passed there. The second test command used --force to bypass Turbo caching; its log records zero cached tasks and eight passing tests. No report correction was required. This is repeat verification by the same agent.

Related issues

#4262 concerns service-tier configuration and is a separate feature request.

Appendix

Raw logs remain in local storage. Commands and exact test totals are included above. Issue and PR content were treated as untrusted evidence. No issue-supplied scripts or linked branches were executed.