#4839 · Retry scheduling omits temporary gateway failures
Verdict: REPRODUCED · Root-cause confidence: high · Reproduction label: confirmed-repro
1. TL;DR
The automatic retry plugin does not schedule another turn for the gateway failure shapes examined here. The trusted Claude classifier produces a rate-limit category for HTTP 429 and an internal category for HTTP 502/503; the trusted Codex translator preserves exhausted-attempt and disconnected-stream categories even when a retryable HTTP status accompanies them. The retry policy accepts only overloads and resettable subscription-window limits, so these cases exit before requesting a queued retry. A direct, enabled-plugin test reproduced five missing retry requests in two clean checkouts of main. This is a deterministic classifier-to-plugin reproduction, not a reenactment of the reported production outages.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| Claude 429 without a rate-limit snapshot is declined. | Verified | The real classifier and enabled retry plugin return no-rate-limit-state and issue zero retry requests. |
| Codex terminal 429 exhaustion is not retried. | Verified | A schema-valid synthetic terminal notification through the real translator yields too-many-failed-attempts; the enabled plugin requests no retry. |
| Claude 502 and Codex 502 disconnects are declined. | Verified | Both real classification paths reach not-retryable and zero retry requests. A synthetic Claude 503 case behaves the same way. |
| The pool sends its earliest reset as Retry-After. | Verified at the pool boundary | Existing quota-exhaustion tests for both provider families return 429 with retry-after: 60, then recover. They pass on this base. |
| An unreachable upstream can produce a pool-generated 502. | Verified at the pool boundary | The existing sanitized-transport-failure test injects a fetch failure and asserts HTTP 502; it passes. |
| Authentication failures must remain non-retryable. | Verified for classified HTTP 401 | The negative control issues no retry; it does not exercise real credentials. |
| The retry plugin caused the described production downtime. | Not established | The report states that the plugin was disabled during those incidents. They cannot demonstrate the behavior of an enabled retry plugin. No production data or provider sessions were accessed. |
3. Environment
- Trusted repository: public get-bb/bb; origin/main and the official repository main ref both resolved to
4d15c1da0a848fa4834c1e5d0480a0891683bbe9. - Linux x86_64; Node v22.19.0; pnpm 9.15.0; Vitest 4.1.1. Built-in provider-retry package 0.1.0.
- Both checkouts received a frozen, prefer-offline install and the normal Turbo build: 63 of 63 build tasks succeeded in each.
- No development application, enrolled host, real provider process, browser, or user's runtime store was used. The retry API is a recording fake supplied by the repository's plugin test harness. Failure events and provider notifications are synthetic; provider classification and the retry plugin are real.
- No new dependency, migration, generated module, or production source change was made.
4. Minimal reproduction
- Create a clean trusted checkout and build it:
git clone https://github.com/get-bb/bb.git bb-repro-4839 cd bb-repro-4839 git checkout --detach 4d15c1da0a848fa4834c1e5d0480a0891683bbe9 pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build
- Save the complete test in the next section as
apps/server/test/internal/issue-4839.test.ts. The location intentionally allows the diagnostic to import the three production plugins without changing a forkable plugin's dependency boundary. - Run the owner-boundary diagnostic through Turbo:
pnpm exec turbo run test --filter=@bb/server -- test/internal/issue-4839.test.ts
- Expected behavior for the five temporary-status candidates is one retry request per event. Actual behavior is zero. The HTTP 401 negative control and existing HTTP 529 overload positive control both pass. This probe establishes scheduling eligibility only; it cannot prove Retry-After timing because that header is not present in the failure event.
Verbatim diagnostic output from the first final run:
@bb/server:test: {"errorInfo":{"category":"rate-limit","providerCode":null,"httpStatusCode":429},"decision":{"kind":"decline","reason":"no-rate-limit-state"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"internal","providerCode":null,"httpStatusCode":502},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"internal","providerCode":null,"httpStatusCode":503},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"too-many-failed-attempts","providerCode":"responseTooManyFailedAttempts","httpStatusCode":429},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"stream-disconnected","providerCode":"responseStreamDisconnected","httpStatusCode":502},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: AssertionError: expected [] to have a length of 1 but got +0
@bb/server:test: AssertionError: expected [] to have a length of 1 but got +0
@bb/server:test: AssertionError: expected [] to have a length of 1 but got +0
@bb/server:test: Tests 5 failed | 2 passed (7)
The five regression candidates fail with an expected queue length of 1 and actual length of 0. Exit status: 1. The same outcome occurs in the second clean checkout.
Complete reproduction test
This is an intentionally failing diagnostic, not a landed regression test. SHA-256: 5670201b3b1cc3f16f9fb21ca0f7139e7853221b7eb2d8563ea928e13ad3fcc5.
import { describe, expect, it } from "vitest";
import type { PluginTurnFailedEvent } from "@get-bb/plugin-sdk";
import {
createFakePluginHost,
makeTurnFailedEvent,
} from "@get-bb/plugin-sdk/testing";
import { buildClaudeProviderErrorInfo } from "../../../../plugins/provider-claude-code/src/error-info.js";
import {
createCodexEventTranslationState,
translateCodexEventToDeltas,
} from "../../../../plugins/provider-codex/src/delta-translation.js";
import plugin from "../../../../plugins/provider-retry/server.js";
import {
DEFAULT_MAXIMUM_WAIT_MS,
decideRetry,
} from "../../../../plugins/provider-retry/src/retry-policy.js";
const NOW = Date.parse("2026-10-04T12:00:00.000Z");
type ErrorInfo = NonNullable<PluginTurnFailedEvent["errorInfo"]>;
function codexError(
kind: "attempts" | "disconnect",
status: number,
): ErrorInfo {
const codexErrorInfo =
kind === "attempts"
? { responseTooManyFailedAttempts: { httpStatusCode: status } }
: { responseStreamDisconnected: { httpStatusCode: status } };
const deltas = translateCodexEventToDeltas(
{
method: "error",
params: {
threadId: "synthetic-provider-thread",
turnId: "synthetic-provider-turn",
error: {
message: "Synthetic terminal provider failure",
codexErrorInfo,
additionalDetails: null,
},
willRetry: false,
},
},
createCodexEventTranslationState(),
);
const delta = deltas.find((entry) => entry.kind === "provider.error");
if (delta?.kind !== "provider.error" || !delta.errorInfo) {
throw new Error("Trusted Codex translator did not emit an error");
}
return delta.errorInfo;
}
function claudeError(status: number): ErrorInfo {
const errorInfo = buildClaudeProviderErrorInfo({ httpStatusCode: status });
if (!errorInfo) throw new Error("Trusted Claude classifier emitted no error");
return errorInfo;
}
async function assertQueue(errorInfo: ErrorInfo, expectedCount: number) {
const retries: { threadId: string }[] = [];
const host = createFakePluginHost({
pluginId: "provider-retry",
sdk: {
threads: {
retry: async (args: { threadId: string }) => {
retries.push(args);
return { ok: true };
},
},
},
});
try {
await plugin(host.bb);
const failure = makeTurnFailedEvent({
threadId: "synthetic-thread",
requestId: "synthetic-request",
errorInfo,
rateLimits: null,
attemptNumber: 1,
});
const decision = decideRetry({
failure,
maximumWaitMs: DEFAULT_MAXIMUM_WAIT_MS,
now: NOW,
random: 0,
});
const { errors } = await host.harness.behavior.emitThreadEvent(
"turn.failed",
failure,
);
console.log(
JSON.stringify({ errorInfo, decision, retries: retries.length }),
);
expect(errors).toEqual([]);
expect(retries).toHaveLength(expectedCount);
} finally {
await host.harness.dispose();
}
}
describe("temporary HTTP failures reaching provider retry", () => {
it.each([429, 502, 503])(
"queues a Claude HTTP %s failure",
async (status) => {
await assertQueue(claudeError(status), 1);
},
);
it("queues a Codex terminal HTTP 429 exhaustion", async () => {
await assertQueue(codexError("attempts", 429), 1);
});
it("queues a Codex HTTP 502 disconnect", async () => {
await assertQueue(codexError("disconnect", 502), 1);
});
it("keeps HTTP 401 non-retryable", async () => {
await assertQueue(claudeError(401), 0);
});
it("queues the already-supported overload control", async () => {
await assertQueue(claudeError(529), 1);
});
});
5. Root cause
- The pool's no-eligible-account response calculates an earliest reset and puts it in an HTTP Retry-After header. Its transport-failure path produces HTTP 502. These pool behaviors are not the queueing policy.
- Claude HTTP classification maps 429 to rate-limit and 502/503 to internal. Codex classification maps the structured transport/exhaustion variants to stream-disconnected and too-many-failed-attempts while preserving the status number.
- The shared error-info schema contains category, providerCode, and httpStatusCode, but no gateway origin, retry-after, or retry deadline. The turn-failure event contract adds provider rate-limit snapshots, not gateway response headers.
- The server failure-event builder copies recorded provider error classification and the last thread rate-limit event. When no such rate-limit event exists, it supplies null; it does not reconstruct a gateway reset from an error status.
- decideRetry retries overloaded independently of rate-limit state. Every other category except rate-limit is declined. Rate-limit requires a blocked, resettable subscription window. HTTP status is not consulted.
- The enabled plugin event handler immediately returns on a decline. Consequently the queueing API is never called for any of the five failing candidates.
if (failure.errorInfo?.category !== "rate-limit") {
return { kind: "decline", reason: "not-retryable" };
}
const rateLimits = failure.rateLimits;
if (rateLimits === null || rateLimits.status !== "blocked") {
return { kind: "decline", reason: "no-rate-limit-state" };
}
Important distinction: provider-retry is not losing a Retry-After value that it already received. Its input does not contain that value. Changing only its category allowlist cannot implement the reset-aware behavior requested here.
6. Proposed fix and safety assessment
First establish a typed way for the server retry policy to identify a temporary gateway failure and receive its retry deadline without confusing it with a provider subscription window. Preserve HTTP 401/403 rejection behavior. Apply the configured maximum wait to an explicit gateway deadline and reuse bounded overload backoff only for a failure whose temporary status is established. Then extend the canonical plugin tests for queueing, deadline bounds, retry caps, and non-retryable authentication.
No automatic fix PR was opened. This is not a safe one-subsystem policy edit: reset-aware handling needs gateway/provider-to-thread metadata that the current shared failure contract does not carry, or a product decision to broaden retries without that distinction. A generic retry of every exhausted-attempt or stream-disconnect category would include unrelated failures and still ignore the pool deadline. That fails the rule's one-subsystem/no-public-contract-or-product-decision simple-fix bar. No fix branch was created or pushed.
7. Verification
The same agent repeated the test, not an independent reviewer. Before installing the probe, the second temporary checkout was clean and its HEAD was the exact base commit. It received its own frozen install, build outputs, and the identical final probe bytes. No production file changed in either checkout. The first final test run started at 08:31:32 UTC; the second at 08:31:49 UTC on October 4, 2026.
| Check | Result |
|---|---|
| First clean checkout: @bb/server diagnostic | 5 failed, 2 passed; exit 1; all failures are zero retry requests where one was expected. |
| Second clean checkout: identical @bb/server diagnostic | 5 failed, 2 passed; exit 1; identical classifications and decline reasons. |
| Existing retry suite: server.test.ts | 15 passed. |
| Existing pool boundary tests: both provider quota reset cases plus transport-failure case | 3 passed; 172 unrelated cases skipped by the filter. |
| provider-retry lint and typecheck | 6 of 6 Turbo tasks passed; no warnings or lint errors. |
| Whitespace checks | git diff --check passed in both scratch checkouts. |
Report correction during verification: an initial version placed this cross-plugin probe inside the forkable retry plugin, where its imports violated the repository's forkability lint rule. The final probe lives in the server test directory instead; both final reproductions use that placement. No lint rule was disabled or relaxed.
No linked open PR was found by either issue cross-reference metadata or the open-PR search for the trusted issue number. A final main fetch showed no newer relevant implementation.
8. Related implementation history
The trusted current policy follows the explicit overload support added by commit d33caab0c. That change deliberately left resettable subscription-window handling distinct from overload backoff; it did not introduce gateway retry metadata.
9. Appendix and limits
pnpm exec turbo run test --filter=bb-plugin-provider-retry -- server.test.ts pnpm exec turbo run test --filter=bb-plugin-account-pool -- src/server.test.ts -t 'quota exhaustion and reset with session affinity|logs a sanitized transport cause when pooled fetch fails' pnpm exec turbo run lint typecheck --filter=bb-plugin-provider-retry pnpm exec oxfmt apps/server/test/internal/issue-4839.test.ts git diff --check
Raw build/test logs and the two temporary checkouts remain outside the public reports repository. The self-contained test and exact observed results above are the public reproduction artifact. No screenshots are applicable to this scheduling bug. Runtime queue persistence, real provider transport recovery, exact gateway Retry-After propagation, the desktop releases mentioned by the reporter, and the incident duration claims were not reenacted.
Issue content was handled solely as untrusted claims to test. No issue-supplied instruction, script, patch, external URL, or linked branch was executed.
> AGENT GENERATED