#3713 · Reasoning stream identity mismatch

Bug · Priority: Medium · Effort: Medium · providers · threads · provider-claude-code

GitHub issue · 2026-09-15 · base d5f04df266cbac8983a0e9d3518fc6740e048471

PARTIALLY REPRODUCED · Root-cause confidence: high for provider identity mismatch

1. TL;DR

The provider can complete a different reasoning item from the one it opened. A repository-derived test proves this for nonzero stream indices in two clean checkouts. The zero-index control passes. Source inspection also supports a separate timeline range-matching concern, but neither the HTTP 500 nor the reported live-session frequency was reproduced here.

2. Claims vs findings

ClaimFindingEvidence
Thinking stream completion can use the wrong item IDVerifiedIndices 2 and 7 fail; index 0 passes in both runs.
Every multi-thinking live turn encounters the problemUnverifiedNo live Claude session; test covers the accepted event shape only.
Orphan reasoning can be extended to turn completionSupported by sourceFinalization uses supplied event sequence and time.
Timeline details return HTTP 500, including late background completionUnverified at runtimeStart-range rewriting, exact matching, and error throw remain in source.

3. Environment

Trusted public get-bb/bb origin/main at the commit above; macOS arm64; Node v22.22.3; Corepack pnpm 9.15.0; Vitest 4.1.1. Two separate source checkouts with separate node_modules installations. No provider CLI, server, database, ports, or user runtime data were used.

The normal frozen install failed because its native dependency hook referenced a missing node-gyp entry point. Frozen installation with lifecycle scripts disabled succeeded. The normal Turbo build was attempted and failed because the local pnpm launcher referenced a missing module. Focused source-mode tests succeeded in executing; these limitations prevent any full-build claim.

4. Minimal reproduction

  1. Clone the trusted repository and detach at d5f04df266cbac8983a0e9d3518fc6740e048471.
  2. Install the locked dependencies with corepack pnpm install --frozen-lockfile --prefer-offline. On this host, use --ignore-scripts after the recorded native hook failure.
  3. Save this test as plugins/provider-claude-code/src/issue-3713.test.ts.
  4. From that package directory run corepack pnpm exec vitest run --config vitest.config.ts src/issue-3713.test.ts.

This direct Vitest invocation is a deliberate investigation bypass after the normal Turbo build failed. Production source is unchanged. The test derives from trusted extraction and harness code, using a content-block start event to exercise the same stream channel without copying the issue's test.

import { expect, it } from "vitest";
import { createClaudeDeltaHarness } from "./delta-test-harness.js";

it.each([0, 2, 7])("settles the reasoning stream at index %i", (index) => {
  const harness = createClaudeDeltaHarness();
  const events = harness.translate({
    type: "stream_event",
    session_id: "repro-session",
    event: {
      type: "content_block_start",
      index,
      content_block: { type: "thinking", thinking: "Inspect the build." },
    },
  });
  events.push(...harness.translate({
    type: "assistant",
    session_id: "repro-session",
    message: {
      role: "assistant",
      content: [{ type: "thinking", thinking: "Inspect the build." }],
    },
  }));
  const opened = events.flatMap(event => event.type === "item/started" ? [event.item.id] : []);
  const closed = events.flatMap(event => event.type === "item/completed" ? [event.item.id] : []);
  expect(opened).toHaveLength(1);
  expect(closed).toEqual(opened);
});

Expected: all three cases close the opened ID. Actual, both runs:

AssertionError: expected [ 'cl-test-i2' ] to deeply equal [ 'cl-test-i1' ]
Tests  2 failed | 1 passed (3)

First run · Second run

5. Root cause

plugins/provider-claude-code/src/sdk-extraction.ts:156 reads a stream event's explicit index for both starts and deltas. Assistant thinking extraction instead enumerates the message content array. A single thinking entry has array index zero, regardless of the index used by its stream.

plugins/provider-claude-code/src/delta-translation.ts:846 uses the assistant array index in the close channel; plugins/provider-claude-code/src/delta-translation.ts:911 uses the stream index in the open/delta channel. Those are different assembler keys. The test proves that the unmatched close creates another item ID, leaving the original item without its matching completion.

Separate source evidence: packages/thread-view/src/reasoning-lifecycle-projection.ts:203 takes the finalizing event's sequence and time; apps/server/src/services/threads/timeline.ts:1871 lowers the requested start to the turn start; packages/thread-view/src/build-thread-timeline.ts:1235 requires equal start and end; apps/server/src/services/threads/timeline.ts:1919 throws when no match is found. This supports the reported failure mechanism but is not an HTTP reproduction.

6. Proposed fix

Preserve stream identity when assistant blocks complete, scoped by parent tool and stream lifetime. Add cases for repeated indices, concurrent parent tools, missing deltas, and text differences before choosing a correlation policy. Separately test timeline details using requested segment bounds and turn-start context, including late task completion.

No production fix or pull request was created: the reported repair spans the provider and timeline subsystems, outside this automation's single-subsystem limit. The full build also remains unverified on this host.

7. Verification

The same agent repeated the test in a second clean detached worktree at the exact base commit, with a separate frozen dependency installation. Before adding the test, the source was clean. The same command produced the same two failures and passing control. This verifies only the provider mismatch; the report deliberately retains a partial verdict for the untested HTTP and live-provider claims. No independent reviewer or workflow was used.

8. Related issues and pull requests

GitHub cross-reference metadata and an open-PR search for issue 3713 returned no linked open pull requests at review time. No PR code or external issue links were fetched. Related-issue distinctions stated in the issue were not independently verified.

9. Appendix

Normal install failure · Turbo build failure. Paths are sanitized. The zero-index control establishes that this is an identity mismatch rather than a test harness startup failure.

Untrusted-data handling: issue instructions, suggested scripts, patches, branches, and links were not executed or followed. Only trusted upstream source and the test authored from that source were executed. No live processes needed cleanup.