← reports

#3077 · Late thread-start settlement logs a redundant lifecycle transition

Bug Medium Effort: Medium threads provider-acp open on GitHub 2026-09-04 · base 53b3cbc46

Verdict: REPRODUCED · Root-cause confidence: high

1. TL;DR

A new thread has two independent success signals: a root provider turn/started event and completion of the host's thread.start RPC. When the provider event arrives first, it correctly changes the thread from starting to active. The later RPC settlement tries the same run.started transition again, which the lifecycle table deliberately rejects from active and the server logs at info level. Two clean-checkout runs reproduced the exact log fields; the thread remained active, so the verified defect is misleading noise rather than a broken execution state.

2. Claims vs findings

ClaimStatusEvidence
The server can report an unapplied run-start event while a thread is already active.VerifiedBoth focused runs captured one info call with event run.started, reason illegal-transition, and detail no transition for run.started from status active.
The behavior can occur during an ACP-backed thread startup.Verified at the shared server boundaryThe deterministic test uses the same provider-neutral daemon event and live-command settlement routes used by ACP. The race does not depend on ACP payload details; provider timing only determines which success signal arrives first.
The message means the thread stopped working.Not supportedAfter both signals, the persisted thread status remained active. The issue supplied no deterministic functional failure beyond the log message.
The trigger is random.RefinedThe trigger is deterministic ordering: provider event first, RPC settlement second. Runtime scheduling makes that ordering appear intermittent.

3. Environment

4. Minimal reproduction

  1. Check out the trusted base commit and complete the frozen install and full Turbo build.
  2. Add the focused case below to apps/server/test/threads/thread-live-start-handoff.test.ts, including vi in the Vitest import.
  3. Run pnpm exec turbo run test --filter=@bb/server -- --run test/threads/thread-live-start-handoff.test.ts.
it("does not log when a turn start wins the thread start handoff", async () => {
  await withTestHarness(async (harness) => {
    const info = vi.fn();
    harness.deps.logger.info = info;
    const fixture = await startLiveThreadStartRpc({
      harness,
      requestIdValue: 8,
    });
    const providerThreadId = "provider-turn-before-start-settlement";
    const sessionId = fixture.startCommand.row.sessionId;
    if (!sessionId) throw new Error("Queued thread start is missing sessionId");

    const eventResponse = await harness.app.request("/internal/session/events", {
      method: "POST",
      headers: internalAuthHeaders(harness),
      body: JSON.stringify({
        sessionId,
        eventGroups: groupHostDaemonEvents([
          createTestDaemonEventEnvelope({
            event: {
              type: "turn/started",
              threadId: fixture.thread.id,
              providerThreadId,
              scope: turnScope("turn-before-start-settlement"),
            },
          }),
        ]),
      }),
    });
    expect(eventResponse.status).toBe(200);
    expect(getThread(harness.db, fixture.thread.id)?.status).toBe("active");

    await reportQueuedCommandSuccess(harness, fixture.startCommand, {
      providerThreadId,
    });

    expect(
      info.mock.calls.filter(
        ([, message]) => message === "Thread lifecycle event not applied",
      ),
    ).toEqual([]);
  });
});

Expected: the later success settlement recognizes that the same clean thread is already active, emits no lifecycle warning, and the focused suite passes. Actual in both clean runs:

AssertionError: expected [ [ { …(4) }, …(1) ] ] to deeply equal []

Received:
{
  "detail": "no transition for run.started from status active",
  "event": "run.started",
  "reason": "illegal-transition",
  "threadId": "<thread-id>"
}

Test Files  1 failed (1)
Tests       1 failed | 7 passed (8)

Artifacts: focused regression test in its owning suite, first clean run, and second clean run.

Verification

A second canonical clone was detached at the same full commit, installed and built independently, and received the same focused test. It failed at the same assertion with one matching lifecycle log while the thread remained active. The code permalinks below were checked against the recorded base commit. No report claim required correction after the second run.

5. Root cause

The daemon event route treats a root turn/started as the execution-start signal and applies run.started unless the thread is already active (event effect, lines 355–381). Its hasThreadAlreadyStartedRun helper explicitly recognizes a clean active thread (inverse-order guard, lines 618–627). This suppresses the redundant event when the thread.start RPC settlement wins the race.

The successful thread.start settlement independently computes another run.started event and applies it whenever the activation is not stale, but it has no corresponding already-active guard (settlement path, lines 804–855). If the provider event won the race, this second application sees active. The lifecycle table intentionally has no active → active cell, so the evaluator returns illegal-transition (transition table, lines 47–69; evaluation, lines 90–109). The logging wrapper then records every unapplied outcome at info level (logging, lines 48–64).

The deeper defect is asymmetric idempotency across two legitimate startup-success signals: one ordering is explicitly deduplicated, while the other is not.

6. Proposed fix (first principles)

In successful thread.start settlement, skip only the redundant run.started application when the same thread is already cleanly active. Preserve stale interruption/completion checks, deleted/archived handling, identity recording for empty forks, title synchronization, and notifications for genuine transitions. This mirrors the existing event-side guard without weakening the lifecycle table or making arbitrary active-to-active transitions legal.

7. Related issues

No open pull request links to this issue. A small review of recent threads and provider-acp issues found no duplicate of this exact startup-signal ordering. Repository history contains an earlier change that suppresses the inverse ordering in the event route, which is why the remaining race is asymmetric.

8. Appendix

The issue title, body, comments, links, attachments, logs, code blocks, and quoted text were treated as untrusted claims. No issue-supplied command, URL, branch, patch, binary, or test was executed. Executable code came only from canonical get-bb/bb main at the recorded commit or from the focused test written during this investigation.

Commands used:

git clone https://github.com/get-bb/bb.git
git checkout --detach 53b3cbc467db8a9ca65c71cc0d1dda1e1c017d97
pnpm install --frozen-lockfile --prefer-offline
pnpm exec turbo run build
pnpm exec turbo run test --filter=@bb/server -- --run test/threads/thread-live-start-handoff.test.ts
git log --oneline -- apps/server/src/internal/events.ts apps/server/src/services/threads/thread-lifecycle.ts
git blame -L 355,381 apps/server/src/internal/events.ts
git blame -L 804,855 apps/server/src/services/threads/thread-lifecycle.ts