#4461 · Provider input acceptance and first-output attribution
Verdict: NOT REPRODUCED · Root-cause confidence: low · Reproduction label: no-repro
1. TL;DR
The report describes lengthy delays from sending a message to receiving the first model item, including occasional waits exceeding ten minutes. Those deployment measurements were not reproduced here: no reporter event logs, matching fleet, authenticated test account, or long-session/compaction workload were used. Current main already opens a Claude turn when the SDK consumes its input, rather than waiting for model output. Four focused lifecycle checks pass in both clean checkouts, but they use a controlled SDK and cannot measure real model latency. The delay's root cause remains unestablished; this verdict does not refute the reporter's measurements or mean the overall issue is fixed.
2. Claims vs findings
| Claim | Status | Evidence |
|---|---|---|
| Reported median first-output delay and very long tail in a five-host deployment. | Unverified | No raw event sample or matched live benchmark was available. Controlled lifecycle tests do not estimate fleet percentiles. |
| A sizable part of the delay precedes the canonical turn-start event. | Historical metric unverified | Current main emits turn-open and acceptance at SDK consumption. The previous first-output gating changed after issue creation; metrics from the two versions are not directly interchangeable. |
| Session resume or prefill might dominate warm-thread latency. | Unverified hypothesis | The normal writable-session path returns the resident session. Settings changes or an ended stream can trigger replacement with resume. Neither stable session IDs nor this code establishes how often replacement happened in the reported fleet. |
| Compaction is associated with multi-minute first-output waits. | Unverified | No long-session compaction workload was exercised. Correlation cannot identify prefill, provider-service, or bb scheduling as the cause. |
| The reported embedded SDK version applies to the tested checkout. | Different environment | The trusted frozen lockfile resolves Claude Agent SDK 0.3.245, not the reported 0.3.258. No dependency was changed to imitate the reporter's environment. |
| Turn acceptance can be separated from first model output. | Verified on current main | The bridge tests withhold every SDK output while consuming input; a turn and its accepted-input event still appear, and steering remains in the same turn. |
3. Environment
- Trusted public repository: get-bb/bb; fetched origin/main at
0e7b518f135d43005dae201ef34ebb3001607eb4. Both checkouts are pristine at this exact commit. - Linux, x86_64; Node 22.19.0; pnpm 9.15.0; Vitest 4.1.1; frozen Claude Agent SDK 0.3.245.
- Each clone ran
pnpm install --frozen-lockfile. Turbo built the provider's plugin-SDK prerequisites before running tests; a full app build was not needed or performed. - No server, browser, provider CLI, network model request, development data directory, production instance, or listening port was started. Existing credentials and fleet runtime data were not read.
- The two checkouts and raw output are kept locally outside this public report repository. Temporary identifiers and absolute machine paths are omitted.
4. Minimal reproduction attempt
This is the closest account-free lifecycle probe, not a simulation of the reported delay. The trusted repository owns the test fixtures and helpers. There is no artificial 33-second sleep and no fabricated model response timing.
- Create a clean clone of
https://github.com/get-bb/bb.git, check out the recorded commit, then run this script from its root:
#!/bin/sh set -eu test "$(git rev-parse HEAD)" = "0e7b518f135d43005dae201ef34ebb3001607eb4" test -z "$(git status --porcelain)" pnpm install --frozen-lockfile pnpm exec turbo run test --filter=bb-plugin-provider-claude-code -- \ src/bridge/__tests__/bridge.test.ts \ src/bridge/__tests__/sdk-session.test.ts \ -t 'opens .*when the SDK consumes input|resolves pushed input after the SDK prompt iterator|rebuilds with the same provider session'
Expected current-main behavior: pending input is not accepted before SDK consumption; immediately after consumption the bridge opens/accepts the turn despite no model output. Steering uses the same turn. A settings-triggered replacement preserves provider identity. Actual first run, exit 0:
bb-plugin-provider-claude-code:test: Test Files 2 passed (2) bb-plugin-provider-claude-code:test: Tests 4 passed | 90 skipped (94) bb-plugin-provider-claude-code:test: Duration 1.30s (transform 899ms, setup 0ms, import 1.64s, tests 44ms, environment 0ms) Tasks: 6 successful, 6 total Cached: 0 cached, 6 total
The test runtime is not time-to-first-model-output. The script ran existing repository tests; it added no regression test or production changes. The full source and helpers can be inspected at plugins/provider-claude-code/src/bridge/__tests__/sdk-session.test.ts:1 and plugins/provider-claude-code/src/bridge/__tests__/bridge.test.ts:4234.
Input-consumption test, verbatim trusted source
it("resolves pushed input after the SDK prompt iterator yields it", async () => {
keepSdkStreamOpen();
const session = new SdkSession(defaultOptions, vi.fn(), vi.fn());
const promptId = "00000000-0000-0000-0000-000000000001";
session.start();
expect(session.canPushInput()).toBe(true);
const consumed = session.pushInput("hello", promptId);
let consumedResolved = false;
void consumed.then(() => {
consumedResolved = true;
});
await Promise.resolve();
expect(consumedResolved).toBe(false);
const result = await getLatestPrompt()[Symbol.asyncIterator]().next();
expect(result.done).toBe(false);
expect(result.value?.message.content).toBe("hello");
expect(result.value?.uuid).toBe(promptId);
await consumed;
expect(consumedResolved).toBe(true);
session.stop();
});
Canonical bridge test, verbatim trusted source
it.each([
{ method: "turn/start", name: "turn start" },
{ method: "turn/steer", name: "turn steer" },
] as const)(
"opens $name when the SDK consumes input before producing output",
async (testCase) => {
const threadId = `thread-${testCase.method.replace("/", "-")}-consumed`;
const bridge = createBridgeJsonRpcTestHarness(handleLine);
const queries: ControlledClaudeQuery[] = [];
queryMock.mockImplementation(() => {
const query = createControlledClaudeQuery();
queries.push(query);
return query;
});
try {
await startBridgeThread({ bridge, threadId });
bridge.sendRequest(2, testCase.method, {
threadId,
providerThreadId: threadId,
...(testCase.method === "turn/steer"
? { expectedTurnId: "turn-1" }
: {}),
input: [{ type: "text", text: "Please account for the restart" }],
clientRequestId: "creq_abcdefghjk",
options: {
permissionMode: "accept-edits",
permissionScope: "workspace",
approvalReviewer: "user",
permissionEscalation: "ask",
providerOptions: {},
},
});
await bridge.flushWork();
expect(bridge.hasResponse(2)).toBe(false);
expect(
assembleCapturedThreadEvents(bridge.messages, "claude-code").some(
(event) => event.type === "turn/started",
),
).toBe(false);
await expect(readNextPromptText(getLatestQueryCall())).resolves.toBe(
"Please account for the restart",
);
await expect(bridge.waitForResponse(2)).resolves.toMatchObject({
result: { threadId },
});
const events = assembleCapturedThreadEvents(
bridge.messages,
"claude-code",
);
const started = events.find((event) => event.type === "turn/started");
expect(started).toBeDefined();
expect(events).toContainEqual(
expect.objectContaining({
type: "turn/input/accepted",
clientRequestId: "creq_abcdefghjk",
scope: started?.scope,
}),
);
if (testCase.method === "turn/start") {
if (started?.scope.kind !== "turn")
throw new Error("Missing active turn");
bridge.sendRequest(3, "turn/steer", {
threadId,
providerThreadId: threadId,
expectedTurnId: started.scope.turnId,
input: [{ type: "text", text: "Use the corrected approach" }],
clientRequestId: "creq_abcdefghjm",
options: {
permissionMode: "accept-edits",
permissionScope: "workspace",
approvalReviewer: "user",
permissionEscalation: "ask",
providerOptions: {},
},
});
await expect(readNextPromptText(getLatestQueryCall())).resolves.toBe(
"Use the corrected approach",
);
await bridge.waitForResponse(3);
const steered = assembleCapturedThreadEvents(
bridge.messages,
"claude-code",
);
expect(
steered.filter((event) => event.type === "turn/started"),
).toHaveLength(1);
expect(steered).toContainEqual(
expect.objectContaining({
type: "turn/input/accepted",
clientRequestId: "creq_abcdefghjm",
scope: steered.find((event) => event.type === "turn/started")
?.scope,
}),
);
}
await stopBridgeThread({ bridge, queries, threadId });
const stopped = assembleCapturedThreadEvents(
bridge.messages,
"claude-code",
);
expect(stopped).toContainEqual(
expect.objectContaining({
type: "turn/completed",
status: "interrupted",
scope: stopped.find((event) => event.type === "turn/started")
?.scope,
}),
);
} finally {
queries[0]?.finish();
bridge.restore();
}
},
);
All identifiers in these source excerpts are synthetic repository test constants, not live session or account identifiers. Raw logs and executable artifacts are not committed, in accordance with this site's publication policy.
5. Root cause and attribution limits
No root cause for the measured first-output delay has been verified. A nearby historical lifecycle defect is fixed on current main, but fixing the timing of a turn-start event is not the same as reducing actual model time-to-first-output.
SDK consumption and output are distinct boundaries
plugins/provider-claude-code/src/bridge/sdk-session.ts:327 and plugins/provider-claude-code/src/bridge/sdk-session.ts:384 resolve input consumption when the SDK input iterator takes the queued message. Then plugins/provider-claude-code/src/bridge/bridge.ts:2326 emits canonical acceptance. The translator at plugins/provider-claude-code/src/delta-translation.ts:1261 returns the existing turn-open and accepted-input deltas:
function acceptInput(
threadId: string,
clientRequestId: ClientTurnRequestId,
): ThreadDelta[] {
const state = stateFor({ threadId });
state.suppressUnacceptedTurnStart = false;
mirrorOpenTurn(state);
return [{ kind: "turn.open" }, { kind: "input.accepted", clientRequestId }];
}
Separately, plugins/provider-claude-code/src/bridge/sdk-session.ts:436 iterates SDK output and forwards messages. Input-iterator consumption alone does not prove that the CLI has sent a network request, completed prefill, or received a model token.
Resident session versus resume
plugins/provider-claude-code/src/bridge/bridge.ts:1299 returns the resident session when no restart is requested. An ended stream or explicit restart cause can take the replacement path. plugins/provider-claude-code/src/bridge/sdk-session.ts:269 passes resume only when given a resume ID. A stable session ID does not distinguish a resident process from a replacement process resuming the same session.
Historical change, not a claim that this issue is resolved
Trusted main commit 53fdb61592f1270e7bc35d6253cb858ede422996 changed Claude input-acceptance/turn-open timing. GitHub metadata records its merge on September 28, 2026 at 23:28:36 UTC, after issue creation at 18:38:04 UTC. Its source change explains why current-main turn-start statistics need a fresh baseline; it does not establish a fix for long model-output delays.
The runtime's packages/agent-runtime/src/runtime.ts:279 is a turn-start watchdog, not a first-output latency measurement. It cannot alone attribute spawn, prompt delivery, resume/prefill, service TTFT, and first persisted model item.
6. Proposed next experiment
Not confident enough for a production fix. With an explicitly isolated authenticated test account, run alternating raw-CLI and bb trials using the same public model, reasoning mode, tools, instructions, history size, and compaction state. Record monotonic timestamps for server submit, queue eligibility, environment/provisioning completion, bridge creation or reuse, SDK input consumption, actual prompt write, first SDK output, first model-produced item, and persistence. Distinguish fresh, resident warm, resumed, and post-compaction sessions; include failures and canceled turns rather than treating every missing event as a long successful response.
Inspect recorded spans before deciding whether to change session reuse, instruction assembly, host scheduling, or model-request behavior. SDK/CLI consumption and first SDK bookkeeping output must not be mislabeled as first model output. Any new cross-process diagnostics would need the repository's protocol/API compatibility review; this report does not authorize such a change.
No pull request: the measured bug has not reproduced on trusted main, its bottleneck is unknown, and there is no focused regression test that fails before a production fix. The simple-fix gate therefore fails before any branch or patch is created.
7. Verification
The same agent created a second separate clean clone, detached it at the recorded base commit, repeated the frozen install, and repeated the exact focused command above with a separate dependency tree. No new ports/data directories were needed because no services were started. Source hashes matched; both git working trees remained clean. This is a repeated check by the same agent, not independent verification.
Second result, exit 0:
bb-plugin-provider-claude-code:test: Test Files 2 passed (2) bb-plugin-provider-claude-code:test: Tests 4 passed | 90 skipped (94) bb-plugin-provider-claude-code:test: Duration 1.30s (transform 1.02s, setup 0ms, import 1.66s, tests 44ms, environment 0ms) Tasks: 6 successful, 6 total Cached: 0 cached, 6 total
The second run supports the bounded finding that current-main acceptance does not wait for model output. It still does not reproduce the reported latency. No report correction was necessary.
8. Related work
No linked open pull request was found through issue cross-reference metadata or the open-PR search for the trusted issue number. The merged lifecycle change above is relevant context, not a duplicate open fix and not evidence that the measured tail latency disappeared. No linked issue instructions, scripts, external links, branches, or patches were executed.
9. Appendix: broader baseline check and audit
Before narrowing to the relevant lifecycle checks, each checkout ran the two complete bridge/SDK-session test files through the same Turbo command without the -t filter. Both returned exit 1 with 93 passing tests and the same unrelated executable-location assertion:
FAIL src/bridge/__tests__/bridge.test.ts > bridge > falls back to well-known install locations when PATH discovery fails Received: undefined Test Files 1 failed | 1 passed (2) Tests 1 failed | 93 passed (94)
The absolute temporary path in the expected value is omitted for privacy. The environment runs as root. plugins/provider-claude-code/src/bridge/session-options.ts:162 intentionally skips well-known executable fallback paths for root, while plugins/provider-claude-code/src/bridge/__tests__/bridge.test.ts:1158 expects the fallback without substituting a non-root UID. This baseline failure does not indicate provider-start latency and was not modified.
Audit commands: repository/issue/property/label/timeline reads; origin/main fetch and commit inspection; frozen installs; both complete and focused Turbo test invocations; source-hash/working-tree checks; report publisher dry run, diff/privacy check, publication, and Pages HTTP check. All GitHub writes use the supplied SlopCop identity wrapper. Classification was already complete and was preserved: Bug, High, Medium, with the five existing labels. Only the final reproduction label is added.
Trust note: issue content was treated only as untrusted claims. Repository code from recorded origin/main supplied all executed tests and root-path evidence. No production data, credentials, real account identifiers, or private model identifiers appear in this report.
> AGENT GENERATED