← reports

#4557 · ACP cancellation failure and session recovery

Bug High Effort: Medium providers provider-acp · GitHub issue · September 30, 2026

Trusted base: 0e7b518f135d43005dae201ef34ebb3001607eb4

Verdict: NOT REPRODUCED · Root-cause confidence: low · Reproduction label: no-repro

1. TL;DR

The reported incident is an active-turn interruption followed by a Python exception and an apparently unusable Hermes conversation. BB does send ACP cancellation before delivering an active-turn follow-up. An investigation-only probe using BB's existing subprocess test adapter shows that a cancellation RPC error fails the current turn, but BB can then send a new prompt and complete a later turn on the same session. Four existing steer tests also pass. Both results were repeated by the same investigator in a second clean checkout of the identical trusted commit. Hermes is not installed in this environment, so neither its Python exception nor its retained internal queue was reproduced; this verdict does not invalidate the field report or establish that an upstream fix is installed.

2. Claims vs findings

ClaimStatusEvidence
An active-turn follow-up triggers ACP cancellation.VerifiedThe trusted steer handler queues input and invokes requestSteerCancel; the existing hung-prompt steer test completes with the replacement input in both runs.
Cancellation can fail an active turn through an adapter RPC error.Verified for a synthetic adapter error onlyThe probe returns a deliberate JSON-RPC cancellation error. The bridge emits a provider-error delta and the turn settles as failed. This is not a reproduction of the reported Python exception.
Hermes raises an exception when a final response is null.UnverifiedNo Hermes executable, installation, Python traceback, or upstream source was available in the trusted target repository. A comment describing upstream commits is a claim, not independently verified evidence here.
After the failure, a later prompt is queued internally while BB appears idle.Unverified for Hermes; not observed with the test adapterThe probe's next turn succeeds and the adapter's prompt log contains both the original prompt and the later follow-up. No live server queue, reporter runtime, or macOS process was accessed.
The affected installed adapter already contains an upstream fix.UnverifiedNo affected-host version or post-update reproduction was available.

3. Environment

4. Repeatable investigation probe

This probes BB's cancellation/error/recovery boundary, not the unavailable Hermes implementation. Production source stays unchanged. The only modifications are one investigation test and an opt-in error response in the existing test adapter. The adapter clears its own active prompt before replying with the deliberate error; therefore this probe cannot establish recovery from an adapter that incorrectly keeps its internal running state.

  1. Create a clean checkout at the recorded commit:
    git clone https://github.com/get-bb/bb.git bb-4557-base
    cd bb-4557-base
    git checkout --detach 0e7b518f135d43005dae201ef34ebb3001607eb4
    pnpm install --frozen-lockfile --prefer-offline
    pnpm exec turbo run build
  2. In packages/provider-bridge-acp/src/bridge/fake-acp-agent.mjs, insert the following immediately after activePromptId = null; in the session/cancel handler and before the normal cancelled result:
    if (process.env.FAKE_ACP_CANCEL_ERROR === "1") {
      send({
        jsonrpc: "2.0",
        id,
        error: { code: -32603, message: "Test adapter cancellation failure" },
      });
      return;
    }
  3. Add the following test inside the existing describe("acp bridge", ...) in packages/provider-bridge-acp/src/bridge/bridge.test.ts. It uses that file's existing helpers:
    it("accepts a later turn after an adapter rejects cancellation during steer", async () => {
      const promptLog = join(workspaceDir, "cancel-error-prompt-log.jsonl");
      const { providerThreadId } = await startThread({
        envVars: {
          FAKE_ACP_CANCEL_ERROR: "1",
          FAKE_ACP_PROMPT_LOG: promptLog,
        },
      });
      const turnId = sendTurnRequest("turn/start", providerThreadId, {
        input: [{ type: "text", text: "hang", mentions: [] }],
      });
      expect((await waitForResponse(turnId)).error).toBeUndefined();
    
      const steerId = sendTurnRequest("turn/steer", providerThreadId, {
        expectedTurnId: "turn-1",
        input: [{ type: "text", text: "redirect", mentions: [] }],
      });
      expect((await waitForResponse(steerId)).error).toBeUndefined();
      expect(await waitForTurnCompleted()).toMatchObject({ status: "failed" });
      expect(emittedDeltaKinds()).toContain("provider.error");
      expect(loggedPrompts(promptLog)).toEqual(["hang"]);
    
      const nextTurnId = sendTurnRequest("turn/start", providerThreadId, {
        input: [{ type: "text", text: "follow-up", mentions: [] }],
      });
      expect((await waitForResponse(nextTurnId)).error).toBeUndefined();
      await waitFor(
        () => threadEventsOfType("turn/completed").length === 2 ? true : undefined,
        "second turn completion",
      );
      expect(threadEventsOfType("turn/completed").at(-1)).toMatchObject({
        status: "completed",
      });
      expect(agentMessageTexts()).toContain("echo:follow-up");
      expect(loggedPrompts(promptLog)).toEqual(["hang", "follow-up"]);
    });
  4. Run the probe plus the existing steer tests:
    pnpm exec turbo run test --filter=@bb/provider-bridge-acp -- \
      src/bridge/bridge.test.ts \
      -t 'cancels a hung prompt|accepts a later turn after an adapter rejects cancellation|keeps partial output|delivers stacked steers|cancels a stacked steer'

Probe expectation: the injected cancellation error fails the original turn, BB does not send the queued redirect, and a newly submitted follow-up completes. Actual: all assertions pass; no Python exception or frozen later turn is produced.

First corrected run:
Test Files  1 passed (1)
Tests  5 passed | 98 skipped (103)
Tasks: 2 successful, 2 total

Second clean corrected run:
Test Files  1 passed (1)
Tests  5 passed | 98 skipped (103)
Tasks: 2 successful, 2 total

The four existing cases cover a hung prompt, partial output before interruption, stacked steers, and a replacement prompt that also hangs. They are positive checks for normal ACP cancellation, not evidence about Hermes Python state. Raw test files and logs are retained locally rather than committed to this public report repository. The complete additional probe is inline here.

5. Root cause and limits

Incident root cause remains unconfirmed. The following bridge mechanism is verified from trusted code and the subprocess probe:

  1. Steer handler: queues the follow-up, requests cancellation, and acknowledges the steer RPC.
  2. Cancellation: sends session/cancel while the original prompt RPC is pending.
  3. RPC error handling: an adapter error rejects the pending request as AcpAgentResponseError.
  4. Turn failure cleanup: clears the pending-prompt flag, drops queued input, resets cancellation, emits the error, and clears the bridge's active prompt kind. It does not automatically discard the whole provider session.
  5. Error translation: the provider error settles the active turn. Clearing BB's active turn does not prove the external process cleared its own internal running state.
  6. Hermes launch contract: BB launches the external hermes acp command; the Python adapter is not part of this repository.
session.queuedInputs.push(...);
requestSteerCancel(session);

session.connection.notify("session/cancel", {
  sessionId: session.providerThreadId,
});

catch (error) {
  session.promptRequestPending = false;
  dropTurnInput(pending, "ACP turn failed before the prompt was sent");
  dropQueuedTurnInputs(session, "ACP turn failed before the steer was sent");
  session.cancelRequested = false;
  if (!session.stopping && !session.connection.exited) {
    emitSessionError(session, error instanceof Error ? error.message : String(error));
  }
  session.activePromptKind = null;
  return;
}

A hypothetical adapter cleanup failure would explain why BB can settle a failed turn while an external adapter still believes it is busy. That is an inference, not a verified throw site. No external source URL from the issue or its comment was fetched, and no upstream fix claim is promoted to an ALREADY FIXED verdict.

6. Proposed next experiment

Capture a sanitized Python traceback and the exact installed Hermes adapter revision in an isolated macOS setup, then reproduce cancellation while a tool is pending and submit a fresh prompt afterward. This distinguishes an adapter failing to clear its running state from a BB session-lifecycle defect. If the adapter owns the failure, verify exception-safe cleanup and final-response normalization in that implementation; if BB owns it, design a focused failing bridge/runtime regression before choosing a repair. Automatically destroying every session on any RPC error would change recovery behavior without enough evidence here.

No safe fix PR: the reported bug was not reproduced on trusted BB main with the real Hermes process, and no verified BB root cause or pre-fix failing regression exists. The 48 added investigation lines are test/fixture code only; no production patch was made or pushed.

7. Verification

The same investigator created a second clean local clone at the exact recorded commit with no shared working-tree modifications. Its origin was set to the trusted target repository. The only copied changes were the two investigation test files; their final contents were compared byte-for-byte between checkouts. Production files were unchanged in both. The second checkout received its own frozen install, full Turbo build, and fixture-created temporary resources.

pnpm install --frozen-lockfile --prefer-offline
pnpm exec turbo run build
pnpm exec turbo run test --force --filter=@bb/provider-bridge-acp -- \
  src/bridge/bridge.test.ts \
  -t 'cancels a hung prompt|accepts a later turn after an adapter rejects cancellation|keeps partial output|delivers stacked steers|cancels a stacked steer'
git diff --check

Result: both builds succeeded; the corrected five-case probe passed separately in both checkouts. The second test run used --force to prevent a cached test result. This is a repeat by the same investigator, not independent verification.

Probe correction: the first draft incorrectly expected the cancellation detail nested inside turn/completed.error. Both initial runs failed that test assertion, while showing the correct failed-turn status. Trusted translation code emits the detail as a separate provider-error delta. The test was corrected to assert failed-turn status and the provider-error delta; production code was not changed. The corrected test reaches and passes the later-turn assertions. No product fix is inferred from that test-authoring error.

8. Related issues and PR metadata

The issue's cross-reference timeline and an open-PR search for its number returned no linked open PR. No linked branch was fetched or run. Other ACP issues were read only for classification context; no shared root cause was established. Existing Type Bug, Priority High, Effort Medium, and labels providers/provider-acp were already present and preserved.

9. Appendix: evidence and safety

Commands used for investigation: trusted repository clone/fetch and detached checkout; GitHub read APIs for repository visibility, issue properties/comments, labels and PR cross-references; command -v checks for Hermes; Node/pnpm/OS version checks; frozen installs; Turbo builds and focused tests; trusted-source reads; byte comparisons and git diff --check. Raw paths, usernames, machine identifiers, reporter conversation content and private session identifiers are intentionally absent from this public report.

The issue and comments were treated as untrusted claims. Suggested scripts, external links and upstream code were not executed or fetched. No real BB instance, credentials, live user data, dependency changes, workflow, or delegated agent was used. No production process or provider installation was altered.