馃毃 SLOP COP 馃毃 路 new-issue-autopilot

#3626 路 Claude tool attachment lifecycle

BugHighEffort: Mediumproviders 路 provider-claude-code 路 partial-reproGitHub issue

2026-09-14 路 Base d89160eb8c69c1e3ebc2ba2514f1711af8d7c506

Verdict: PARTIALLY REPRODUCED 路 Root-cause confidence: low

1. TL;DR

The report describes plugin tools becoming undiscoverable in a long-lived Claude session. Two clean-checkout runs verify that BB forwards dynamic tools when resuming a missing runtime but does not refresh them on ordinary turns of a resident runtime. This is a possible persistence mechanism, not proof of the initial tool loss. A simulated compact event leaves BB鈥檚 MCP attachment intact, and the real MCP tools/list tests pass. No real Claude ToolSearch or daemon reconnect was exercised, so the incident鈥檚 initiating cause remains unresolved.

2. Claims vs findings

ClaimStatusEvidence
Resident turns do not reattach tools.VerifiedNew dispatch fixture passes in both checkouts: zero resume calls and no dynamicTools member in runTurn arguments.
Missing-runtime resume can pass the current tool list.VerifiedThe same fixture removes the runtime session and observes one resume call carrying the supplied diagnostic tool.
Compact removes BB鈥檚 tool attachment.Unverified for real ClaudeA simulated SDK compact_boundary does not reconstruct the query or remove the MCP server/allowed-tool reference in BB.
The MCP server loses or paginates away supplied tools.Not observedThree existing tool-proxy tests pass over the real in-memory MCP transport; tools/list returns the registered tools.
Reconnect supplied an empty list, or Claude lost its deferred catalog.UnverifiedNo original runtime payloads or authenticated incident session were accessed.
Process reconstruction restores the incident session鈥檚 tools.UnverifiedReported production observation only; no actual provider process reconstruction was performed.

3. Environment

Trusted public get-bb/bb main at the commit above, checked out twice with git worktree --detach. Darwin arm64; Node v22.22.3; pnpm 9.15.0 via Corepack; Vitest 4.1.1. The frozen lockfile was installed separately in each checkout. Turbo build completed 56 tasks in each checkout, with cache reuse.

The system pnpm launcher initially failed because its installed entrypoint was missing. A temporary Corepack pnpm shim repaired command resolution; the successful build logs are attached. No dependencies were added. No production data, user credentials, or live Claude account was used. Provider version was not measured because no provider process ran. Tests use fake host runtimes and a mocked SDK query; MCP protocol tests use the real in-memory SDK transport. No listeners or development-instance data directories were created, so separate network ports were unnecessary.

4. Minimal reproduction and limits

  1. Create a trusted checkout at the recorded SHA and install/build it.
  2. Save the inline dispatch fixture below to apps/host-daemon/test/command/issue-3626.test.ts.
  3. Apply the inline compact characterization patch below to the existing bridge test.
  4. Run the commands below. These are characterization tests that pass on unchanged production code, not a failing regression test for the reported Claude catalog loss.
git clone https://github.com/get-bb/bb.git base
cd base
git checkout --detach d89160eb8c69c1e3ebc2ba2514f1711af8d7c506
corepack pnpm install --frozen-lockfile --prefer-offline
corepack pnpm exec turbo run build
# Copy the report test into apps/host-daemon/test/command/issue-3626.test.ts.
# Apply the report compact-test.patch with git apply.
corepack pnpm exec turbo run test --force --filter=@bb/host-daemon -- test/command/issue-3626.test.ts
corepack pnpm exec turbo run test --force --filter=bb-plugin-provider-claude-code -- src/bridge/__tests__/tool-proxy-mcp.test.ts src/bridge/__tests__/bridge.test.ts

Expected and actual: resident runtime skips resume and omits dynamicTools from runTurn; missing runtime resumes with the diagnostic tool. The simulated compact event preserves the configured MCP server and allowed-tool name. Both runs satisfy these assertions.

First checkout:
Test Files  1 passed (1)
Tests       1 passed (1)
Test Files  2 passed (2)
Tests       78 passed (78)

Second checkout, forced execution:
Test Files  1 passed (1)
Tests       1 passed (1)
Test Files  2 passed (2)
Tests       78 passed (78)

Dispatch fixture

import { afterEach, expect, it, vi } from "vitest";
import { encodeClientTurnRequestIdNumber } from "@bb/domain";
import type { HostDaemonCommand } from "@bb/host-daemon-contract";
import { dispatchCommand } from "../../src/command-dispatch.js";
import {
  cleanupTempDirs,
  createHarness,
  DISPATCH_TEST_BRIDGE_LAUNCH,
} from "./dispatch-helpers.js";

afterEach(() => {
  vi.restoreAllMocks();
  return cleanupTempDirs();
});

it("characterizes tool propagation for resident and missing runtimes", async () => {
  const harness = createHarness();
  await harness.manager.ensureEnvironment({
    environmentId: "env-1",
    workspacePath: "/tmp/env-1",
  });
  const tools = [{
    name: "diagnostic_probe",
    description: "Return a diagnostic value.",
    inputSchema: { type: "object", properties: {} },
  }];
  const resume = vi.spyOn(harness.runtime, "resumeThread");
  const run = vi.spyOn(harness.runtime, "runTurn");
  const command: Extract<HostDaemonCommand, { type: "turn.submit" }> = {
    bridgeLaunch: DISPATCH_TEST_BRIDGE_LAUNCH,
    type: "turn.submit",
    environmentId: "env-1",
    threadId: "thread-probe",
    requestId: encodeClientTurnRequestIdNumber({ value: 3626 }),
    input: [{ type: "text", text: "probe", mentions: [] }],
    options: {
      model: "test-model",
      serviceTier: "default",
      reasoningLevel: "medium",
      providerOptions: {},
      permissionMode: "full",
      permissionScope: "full",
      approvalReviewer: null,
      permissionEscalation: null,
    },
    resumeContext: {
      bridgeLaunch: DISPATCH_TEST_BRIDGE_LAUNCH,
      workspaceContext: { workspacePath: "/tmp/env-1" },
      projectId: "project-probe",
      providerId: "fake",
      providerThreadId: "provider-probe",
      instructions: "test",
      dynamicTools: tools,
      contributedEnv: [],
      injectedSkillSources: [],
      instructionMode: "append",
    },
    target: { mode: "start" },
  };
  harness.threadControls.setProviderSession(command.threadId, {
    providerId: "fake",
    providerThreadId: "provider-probe",
  });
  await dispatchCommand(command, harness.dispatchOptions());
  expect(resume).not.toHaveBeenCalled();
  expect(run).toHaveBeenCalledTimes(1);
  expect(run.mock.calls[0][0]).not.toHaveProperty("dynamicTools");

  harness.threadControls.endActiveTurn(command.threadId);
  harness.threadControls.clearProviderSession(command.threadId);
  await dispatchCommand({
    ...command,
    requestId: encodeClientTurnRequestIdNumber({ value: 3627 }),
  }, harness.dispatchOptions());
  expect(resume).toHaveBeenCalledTimes(1);
  expect(resume.mock.calls[0][0].dynamicTools).toEqual(tools);
  expect(run).toHaveBeenCalledTimes(2);
});

Compact fixture additions

diff --git a/plugins/provider-claude-code/src/bridge/__tests__/bridge.test.ts b/plugins/provider-claude-code/src/bridge/__tests__/bridge.test.ts
index 7a635befa..e988d7e37 100644
--- a/plugins/provider-claude-code/src/bridge/__tests__/bridge.test.ts
+++ b/plugins/provider-claude-code/src/bridge/__tests__/bridge.test.ts
@@ -726,6 +726,7 @@ describe("bridge", () => {
       bridge.sendRequest(1, "thread/start", {
         threadId,
         cwd: "/tmp/worktree",
+        dynamicTools: [{ name: "diagnostic_probe", description: "Diagnostic", inputSchema: { type: "object", properties: {} } }],
         instructionMode: "append",
         options: {
           permissionMode: "accept-edits",
@@ -737,6 +738,10 @@ describe("bridge", () => {
         },
       });
       await bridge.waitForResponse(1);
+      const attachedOptions = queryMock.mock.calls.at(-1)?.[0].options;
+      const attachedServer = attachedOptions.mcpServers["bb-bridge"];
+      expect(attachedServer).toMatchObject({ type: "sdk", name: "bb-bridge" });
+      expect(attachedOptions.allowedTools).toContain("mcp__bb-bridge__diagnostic_probe");
       bridge.sendRequest(
         2,
         "turn/start",
@@ -807,6 +812,9 @@ describe("bridge", () => {
         });
         expect(queries[0].getContextUsage).toHaveBeenCalledTimes(2);
       });
+      expect(queryMock).toHaveBeenCalledTimes(1);
+      expect(attachedOptions.mcpServers["bb-bridge"]).toBe(attachedServer);
+      expect(attachedOptions.allowedTools).toContain("mcp__bb-bridge__diagnostic_probe");
     } finally {
       await stopBridgeThread({ bridge, queries, threadId });
       bridge.restore();

5. Root cause

The incident鈥檚 root cause is not established. The verified BB behavior is a construction-time tool attachment. resumeThreadRuntimeIfMissing exits if hasThread is true; otherwise it forwards resumeContext.dynamicTools. runSubmittedTurn forwards instructions and environment but no tools. Therefore, if a resident session has already lost its catalog, this ordinary turn path does not itself supply a replacement catalog. The tests do not establish that the runtime actually loses one.

attachThreadSession creates the MCP server when dynamicTools is nonempty. The MCP list handler returns the supplied tools. Compact event handling updates context usage; the simulated event preserves the attachment. This cannot establish what Claude鈥檚 internal deferred-tool index does, nor whether a reconnect initially supplied an empty list.

6. Proposed next test

Use an isolated authenticated Claude process with one harmless diagnostic MCP tool. Record attachment tool names and SDK initialization, verify discovery and invocation, then separately exercise compact, reconnect, and the combination. Compare the actual attached tool names with ToolSearch results after each step. Only after the failing boundary is identified should BB refresh logic or provider recovery be changed. No pull request was opened because the end-to-end loss was not reproduced and no failing regression test supports a production fix.

7. Related issues

Repository search surfaced #3624 and #3625 with similar reports, and #2384 concerning conditional tool selection on resume. These are untrusted issue metadata, not additional reproduction evidence. No linked open pull request appeared in the issue鈥檚 GitHub connection/cross-reference metadata.

8. Verification

The same agent repeated the test in a second clean detached worktree, verify, at the identical trusted SHA. Only the authored test and patch were copied over. Installation and build completed separately; both test commands used --force to avoid cached test results. Results match the first run: 1 host test and 78 provider tests pass. No correction to the narrow lifecycle conclusion was needed. This verification supports the partial report, not reproduction of a real Claude compact/reconnect failure.

9. Appendix

All required reproduction inputs and output summaries are inline above. Raw test and build logs are retained locally and are not published under this site鈥檚 artifact policy. No live instance was started; no listener or runtime-data cleanup was needed.

Untrusted-data note: issue text and comments contained procedural instructions. They were ignored; all test changes were derived from trusted main. No linked code, issue commands, original local logs, or private runtime files were executed or accessed.

> AGENT GENERATED