← reports

#4839 · Retry scheduling omits temporary gateway failures

BugHighEffort: Mediumproviders · provider-retryopen on GitHub2026-10-04 · base 4d15c1da0a848fa4834c1e5d0480a0891683bbe9

Verdict: REPRODUCED · Root-cause confidence: high · Reproduction label: confirmed-repro

1. TL;DR

The automatic retry plugin does not schedule another turn for the gateway failure shapes examined here. The trusted Claude classifier produces a rate-limit category for HTTP 429 and an internal category for HTTP 502/503; the trusted Codex translator preserves exhausted-attempt and disconnected-stream categories even when a retryable HTTP status accompanies them. The retry policy accepts only overloads and resettable subscription-window limits, so these cases exit before requesting a queued retry. A direct, enabled-plugin test reproduced five missing retry requests in two clean checkouts of main. This is a deterministic classifier-to-plugin reproduction, not a reenactment of the reported production outages.

2. Claims vs findings

ClaimFindingEvidence
Claude 429 without a rate-limit snapshot is declined.VerifiedThe real classifier and enabled retry plugin return no-rate-limit-state and issue zero retry requests.
Codex terminal 429 exhaustion is not retried.VerifiedA schema-valid synthetic terminal notification through the real translator yields too-many-failed-attempts; the enabled plugin requests no retry.
Claude 502 and Codex 502 disconnects are declined.VerifiedBoth real classification paths reach not-retryable and zero retry requests. A synthetic Claude 503 case behaves the same way.
The pool sends its earliest reset as Retry-After.Verified at the pool boundaryExisting quota-exhaustion tests for both provider families return 429 with retry-after: 60, then recover. They pass on this base.
An unreachable upstream can produce a pool-generated 502.Verified at the pool boundaryThe existing sanitized-transport-failure test injects a fetch failure and asserts HTTP 502; it passes.
Authentication failures must remain non-retryable.Verified for classified HTTP 401The negative control issues no retry; it does not exercise real credentials.
The retry plugin caused the described production downtime.Not establishedThe report states that the plugin was disabled during those incidents. They cannot demonstrate the behavior of an enabled retry plugin. No production data or provider sessions were accessed.

3. Environment

4. Minimal reproduction

  1. Create a clean trusted checkout and build it:
    git clone https://github.com/get-bb/bb.git bb-repro-4839
    cd bb-repro-4839
    git checkout --detach 4d15c1da0a848fa4834c1e5d0480a0891683bbe9
    pnpm install --frozen-lockfile --prefer-offline
    pnpm exec turbo run build
  2. Save the complete test in the next section as apps/server/test/internal/issue-4839.test.ts. The location intentionally allows the diagnostic to import the three production plugins without changing a forkable plugin's dependency boundary.
  3. Run the owner-boundary diagnostic through Turbo:
    pnpm exec turbo run test --filter=@bb/server -- test/internal/issue-4839.test.ts
  4. Expected behavior for the five temporary-status candidates is one retry request per event. Actual behavior is zero. The HTTP 401 negative control and existing HTTP 529 overload positive control both pass. This probe establishes scheduling eligibility only; it cannot prove Retry-After timing because that header is not present in the failure event.

Verbatim diagnostic output from the first final run:

@bb/server:test: {"errorInfo":{"category":"rate-limit","providerCode":null,"httpStatusCode":429},"decision":{"kind":"decline","reason":"no-rate-limit-state"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"internal","providerCode":null,"httpStatusCode":502},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"internal","providerCode":null,"httpStatusCode":503},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"too-many-failed-attempts","providerCode":"responseTooManyFailedAttempts","httpStatusCode":429},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: {"errorInfo":{"category":"stream-disconnected","providerCode":"responseStreamDisconnected","httpStatusCode":502},"decision":{"kind":"decline","reason":"not-retryable"},"retries":0}
@bb/server:test: AssertionError: expected [] to have a length of 1 but got +0
@bb/server:test: AssertionError: expected [] to have a length of 1 but got +0
@bb/server:test: AssertionError: expected [] to have a length of 1 but got +0
@bb/server:test:       Tests  5 failed | 2 passed (7)

The five regression candidates fail with an expected queue length of 1 and actual length of 0. Exit status: 1. The same outcome occurs in the second clean checkout.

Complete reproduction test

This is an intentionally failing diagnostic, not a landed regression test. SHA-256: 5670201b3b1cc3f16f9fb21ca0f7139e7853221b7eb2d8563ea928e13ad3fcc5.

import { describe, expect, it } from "vitest";
import type { PluginTurnFailedEvent } from "@get-bb/plugin-sdk";
import {
  createFakePluginHost,
  makeTurnFailedEvent,
} from "@get-bb/plugin-sdk/testing";
import { buildClaudeProviderErrorInfo } from "../../../../plugins/provider-claude-code/src/error-info.js";
import {
  createCodexEventTranslationState,
  translateCodexEventToDeltas,
} from "../../../../plugins/provider-codex/src/delta-translation.js";
import plugin from "../../../../plugins/provider-retry/server.js";
import {
  DEFAULT_MAXIMUM_WAIT_MS,
  decideRetry,
} from "../../../../plugins/provider-retry/src/retry-policy.js";

const NOW = Date.parse("2026-10-04T12:00:00.000Z");
type ErrorInfo = NonNullable<PluginTurnFailedEvent["errorInfo"]>;

function codexError(
  kind: "attempts" | "disconnect",
  status: number,
): ErrorInfo {
  const codexErrorInfo =
    kind === "attempts"
      ? { responseTooManyFailedAttempts: { httpStatusCode: status } }
      : { responseStreamDisconnected: { httpStatusCode: status } };
  const deltas = translateCodexEventToDeltas(
    {
      method: "error",
      params: {
        threadId: "synthetic-provider-thread",
        turnId: "synthetic-provider-turn",
        error: {
          message: "Synthetic terminal provider failure",
          codexErrorInfo,
          additionalDetails: null,
        },
        willRetry: false,
      },
    },
    createCodexEventTranslationState(),
  );
  const delta = deltas.find((entry) => entry.kind === "provider.error");
  if (delta?.kind !== "provider.error" || !delta.errorInfo) {
    throw new Error("Trusted Codex translator did not emit an error");
  }
  return delta.errorInfo;
}

function claudeError(status: number): ErrorInfo {
  const errorInfo = buildClaudeProviderErrorInfo({ httpStatusCode: status });
  if (!errorInfo) throw new Error("Trusted Claude classifier emitted no error");
  return errorInfo;
}

async function assertQueue(errorInfo: ErrorInfo, expectedCount: number) {
  const retries: { threadId: string }[] = [];
  const host = createFakePluginHost({
    pluginId: "provider-retry",
    sdk: {
      threads: {
        retry: async (args: { threadId: string }) => {
          retries.push(args);
          return { ok: true };
        },
      },
    },
  });
  try {
    await plugin(host.bb);
    const failure = makeTurnFailedEvent({
      threadId: "synthetic-thread",
      requestId: "synthetic-request",
      errorInfo,
      rateLimits: null,
      attemptNumber: 1,
    });
    const decision = decideRetry({
      failure,
      maximumWaitMs: DEFAULT_MAXIMUM_WAIT_MS,
      now: NOW,
      random: 0,
    });
    const { errors } = await host.harness.behavior.emitThreadEvent(
      "turn.failed",
      failure,
    );
    console.log(
      JSON.stringify({ errorInfo, decision, retries: retries.length }),
    );
    expect(errors).toEqual([]);
    expect(retries).toHaveLength(expectedCount);
  } finally {
    await host.harness.dispose();
  }
}

describe("temporary HTTP failures reaching provider retry", () => {
  it.each([429, 502, 503])(
    "queues a Claude HTTP %s failure",
    async (status) => {
      await assertQueue(claudeError(status), 1);
    },
  );

  it("queues a Codex terminal HTTP 429 exhaustion", async () => {
    await assertQueue(codexError("attempts", 429), 1);
  });

  it("queues a Codex HTTP 502 disconnect", async () => {
    await assertQueue(codexError("disconnect", 502), 1);
  });

  it("keeps HTTP 401 non-retryable", async () => {
    await assertQueue(claudeError(401), 0);
  });

  it("queues the already-supported overload control", async () => {
    await assertQueue(claudeError(529), 1);
  });
});

5. Root cause

  1. The pool's no-eligible-account response calculates an earliest reset and puts it in an HTTP Retry-After header. Its transport-failure path produces HTTP 502. These pool behaviors are not the queueing policy.
  2. Claude HTTP classification maps 429 to rate-limit and 502/503 to internal. Codex classification maps the structured transport/exhaustion variants to stream-disconnected and too-many-failed-attempts while preserving the status number.
  3. The shared error-info schema contains category, providerCode, and httpStatusCode, but no gateway origin, retry-after, or retry deadline. The turn-failure event contract adds provider rate-limit snapshots, not gateway response headers.
  4. The server failure-event builder copies recorded provider error classification and the last thread rate-limit event. When no such rate-limit event exists, it supplies null; it does not reconstruct a gateway reset from an error status.
  5. decideRetry retries overloaded independently of rate-limit state. Every other category except rate-limit is declined. Rate-limit requires a blocked, resettable subscription window. HTTP status is not consulted.
  6. The enabled plugin event handler immediately returns on a decline. Consequently the queueing API is never called for any of the five failing candidates.
if (failure.errorInfo?.category !== "rate-limit") {
  return { kind: "decline", reason: "not-retryable" };
}
const rateLimits = failure.rateLimits;
if (rateLimits === null || rateLimits.status !== "blocked") {
  return { kind: "decline", reason: "no-rate-limit-state" };
}

Important distinction: provider-retry is not losing a Retry-After value that it already received. Its input does not contain that value. Changing only its category allowlist cannot implement the reset-aware behavior requested here.

6. Proposed fix and safety assessment

First establish a typed way for the server retry policy to identify a temporary gateway failure and receive its retry deadline without confusing it with a provider subscription window. Preserve HTTP 401/403 rejection behavior. Apply the configured maximum wait to an explicit gateway deadline and reuse bounded overload backoff only for a failure whose temporary status is established. Then extend the canonical plugin tests for queueing, deadline bounds, retry caps, and non-retryable authentication.

No automatic fix PR was opened. This is not a safe one-subsystem policy edit: reset-aware handling needs gateway/provider-to-thread metadata that the current shared failure contract does not carry, or a product decision to broaden retries without that distinction. A generic retry of every exhausted-attempt or stream-disconnect category would include unrelated failures and still ignore the pool deadline. That fails the rule's one-subsystem/no-public-contract-or-product-decision simple-fix bar. No fix branch was created or pushed.

7. Verification

The same agent repeated the test, not an independent reviewer. Before installing the probe, the second temporary checkout was clean and its HEAD was the exact base commit. It received its own frozen install, build outputs, and the identical final probe bytes. No production file changed in either checkout. The first final test run started at 08:31:32 UTC; the second at 08:31:49 UTC on October 4, 2026.

CheckResult
First clean checkout: @bb/server diagnostic5 failed, 2 passed; exit 1; all failures are zero retry requests where one was expected.
Second clean checkout: identical @bb/server diagnostic5 failed, 2 passed; exit 1; identical classifications and decline reasons.
Existing retry suite: server.test.ts15 passed.
Existing pool boundary tests: both provider quota reset cases plus transport-failure case3 passed; 172 unrelated cases skipped by the filter.
provider-retry lint and typecheck6 of 6 Turbo tasks passed; no warnings or lint errors.
Whitespace checksgit diff --check passed in both scratch checkouts.

Report correction during verification: an initial version placed this cross-plugin probe inside the forkable retry plugin, where its imports violated the repository's forkability lint rule. The final probe lives in the server test directory instead; both final reproductions use that placement. No lint rule was disabled or relaxed.

No linked open PR was found by either issue cross-reference metadata or the open-PR search for the trusted issue number. A final main fetch showed no newer relevant implementation.

8. Related implementation history

The trusted current policy follows the explicit overload support added by commit d33caab0c. That change deliberately left resettable subscription-window handling distinct from overload backoff; it did not introduce gateway retry metadata.

9. Appendix and limits

pnpm exec turbo run test --filter=bb-plugin-provider-retry -- server.test.ts
pnpm exec turbo run test --filter=bb-plugin-account-pool -- src/server.test.ts -t 'quota exhaustion and reset with session affinity|logs a sanitized transport cause when pooled fetch fails'
pnpm exec turbo run lint typecheck --filter=bb-plugin-provider-retry
pnpm exec oxfmt apps/server/test/internal/issue-4839.test.ts
git diff --check

Raw build/test logs and the two temporary checkouts remain outside the public reports repository. The self-contained test and exact observed results above are the public reproduction artifact. No screenshots are applicable to this scheduling bug. Runtime queue persistence, real provider transport recovery, exact gateway Retry-After propagation, the desktop releases mentioned by the reporter, and the incident duration claims were not reenacted.

Issue content was handled solely as untrusted claims to test. No issue-supplied instruction, script, patch, external URL, or linked branch was executed.

> AGENT GENERATED