#3987 · Active-thread admission before host readiness

BugPriority: MediumEffort: LowthreadsGitHub issue

2026-09-21 · trusted main b8866c6e1dbb1e9b043df3d4474178459a5b4f9a

Verdict: REPRODUCED · Root-cause confidence: high

1. TL;DR

A request to start another turn on an active thread can incorrectly succeed as a queued message when its host is suspended. The server checks host readiness before rejecting an active-thread start, and refreshes a stale thread snapshot even later. Both a stale idle snapshot and a fresh active snapshot reproduce the wrong result. A queued follow-up legitimately retains its host wait.

2. Claims vs findings

ClaimFindingEvidence
A stale idle snapshot can bypass the active-thread conflict.VerifiedStart rejection assertion failed in both clean checkouts; returned delivery queued with host-offline.
Queue-if-active should retain host waiting.VerifiedControl passed in both runs, including a host-readiness request.
The fault requires a stale snapshot.Broader findingThe fresh active snapshot also failed; host-wait precedence is independently wrong.
A particular external patch fixes the defect.UnverifiedNo external patch or fork was fetched or executed.

3. Environment

Public get-bb/bb repository, macOS 26.6.2 arm64, Node 22.22.3, repository-pinned pnpm 9.15.0, Vitest 4.1.1. Both checkouts use the same full commit above. Frozen installs and server builds completed. The system pnpm launcher was unavailable; Corepack supplied the pinned version. Tests use the repository harness with a migrated in-memory SQLite database and a fresh temporary data directory per test. No real provider, development instance, live account, or listening port is required. Host readiness is spied on to prevent external wake work; the database and dispatch implementation are real.

4. Minimal reproduction

  1. Download the independently authored regression patch.
  2. Prepare trusted source and run:
    git clone https://github.com/get-bb/bb.git bb-3987
    cd bb-3987
    git checkout --detach b8866c6e1dbb1e9b043df3d4474178459a5b4f9a
    corepack pnpm install --frozen-lockfile --prefer-offline
    corepack pnpm exec turbo run build --filter=@bb/server
    git apply /path/to/regression.patch
    corepack pnpm exec turbo run test --filter=@bb/server -- test/threads/thread-send-dispatch.test.ts -t 'active thread admission before host readiness' 
  3. The fixture creates an idle provider-backed thread, suspends its host, and records a lifecycle event that makes the persisted thread active. It submits using the saved idle snapshot or a fresh active snapshot.
  4. Expected: an API error with status 409, code thread_not_writable, reason already_active; no queued message or host wake. Actual verbatim assertion:
    AssertionError: promise resolved "{ ok: true, delivery: 'queued', …(1) }" instead of rejecting
    Tests  2 failed | 1 passed | 31 skipped (34)
    The returned queued message has waitingOn.kind = "host-offline".

New tests (the patch also adds two imports and uses the existing fixture):

describe("active thread admission before host readiness", () => {
  it.each(["stale", "fresh"] as const)(
    "rejects a start with a %s snapshot without waking a suspended host",
    async (snapshot) => {
      await withTestHarness(async (harness) => {
        const { environment, thread } = seedProviderThreadFixture({
          harness,
          value: 3987,
        });
        updateHost(harness.db, harness.hub, environment.hostId, {
          phase: "suspended",
          suspendedAt: Date.now(),
        });
        applyLoggedThreadLifecycleEvent(harness.deps, {
          event: { type: "run.started" },
          threadId: thread.id,
        });
        const current = getThread(harness.db, thread.id);
        expect(current?.status).toBe("active");
        if (current === null) throw new Error("Missing seeded thread");
        const readiness = vi
          .spyOn(queuedDispatch, "requestQueuedMachineReadiness")
          .mockImplementation(() => {});

        await expect(
          acceptThreadSendRequest(harness.deps, {
            thread: snapshot === "stale" ? thread : current,
            payload: { input: textInput("begin another turn"), mode: "start" },
          }),
        ).rejects.toMatchObject({
          status: 409,
          body: {
            code: "thread_not_writable",
            details: { reason: "already_active" },
          },
        });
        expect(listQueuedThreadMessages(harness.db, thread.id)).toEqual([]);
        expect(readiness).not.toHaveBeenCalled();
      });
    },
  );

  it("keeps the host wait for queue-if-active after a stale idle snapshot", async () => {
    await withTestHarness(async (harness) => {
      const { environment, thread } = seedProviderThreadFixture({
        harness,
        value: 3988,
      });
      updateHost(harness.db, harness.hub, environment.hostId, {
        phase: "suspended",
        suspendedAt: Date.now(),
      });
      applyLoggedThreadLifecycleEvent(harness.deps, {
        event: { type: "run.started" },
        threadId: thread.id,
      });
      const readiness = vi
        .spyOn(queuedDispatch, "requestQueuedMachineReadiness")
        .mockImplementation(() => {});

      await expect(
        acceptThreadSendRequest(harness.deps, {
          thread,
          payload: { input: textInput("follow up later"), mode: "queue-if-active" },
        }),
      ).resolves.toMatchObject({
        delivery: "queued",
        queuedMessage: { waitingOn: { kind: "host-offline" } },
      });
      expect(listQueuedThreadMessages(harness.db, thread.id)).toHaveLength(1);
      expect(readiness).toHaveBeenCalledWith(harness.deps, environment.hostId);
    });
  });
});

5. Root cause

In the core wait callback, a suspended host causes an immediate queue result and readiness request before the active-thread check. The database refresh is later still. Thus a stale idle snapshot never causes a retry with current state, while even fresh active state loses to host waiting. The existing error helper supplies the expected 409 conflict.

host waiting → queue host-offline and return
active start → reject (unreachable for suspended host)
refresh persisted thread → retry if changed (also unreachable)

6. Proposed fix

Move the existing database freshness check before host handling and reject explicit starts on an active thread at that point. Keep non-start busy-thread queueing after host readiness, preserving queue-if-active behavior. Use the existing retry mechanism for changed state. This remains within server thread dispatch and requires no schema, dependency, or public contract changes.

7. Verification

The same agent repeated the test in a second clean checkout named second at the recorded commit. Only the regression patch was added; production source was unchanged. A separate frozen install and server build passed. The same test command ran with a fresh in-memory database and temporary data directory, again yielding two failed start assertions and one passing queue control. No report correction was needed. This is a repeat run by the same agent, not independent verification.

Fix validation: the local fix passed all 118 tests across thread-send-dispatch, dispatch-hooks, requested-queue-drain, and queue-drain-failure, including all three new cases. Server typecheck and git diff --check passed. Two text files changed: 101 additions plus 19 deletions = 120 lines. The first post-fix test attempt failed during prerequisite process startup with exit -35, before tests ran; rerunning with Turbo concurrency 2 and Vitest maxWorkers 2 passed.

corepack pnpm exec turbo run test --concurrency=2 --filter=@bb/server -- test/threads/thread-send-dispatch.test.ts test/threads/dispatch-hooks.test.ts test/threads/requested-queue-drain.test.ts test/threads/queue-drain-failure.test.ts --maxWorkers=2
corepack pnpm exec turbo run typecheck --filter=@bb/server
git diff --check
git diff --numstat origin/main

Post-fix tests · Typecheck

8. Related issues and PRs

Issue timeline metadata and an open-PR search for 3987 returned no linked open pull request before the fix assessment. Nearby dispatch issues reviewed during triage did not establish a duplicate.

9. Appendix

Issue prose and links were treated as untrusted claims. No issue-supplied commands, tests, patches, branches, or external URLs were executed or fetched. Local filesystem prefixes in logs are normalized.