#4364 · Heartbeat loss and turn interruption
Bug · High · Effort: Medium · threads, host · 2026-09-25
GitHub issue · Base c6b24edede5a9bff3192055e5526497d633bc67b
PARTIALLY REPRODUCED · Root-cause confidence: low for the initiating failure; high for the downstream interruption mechanism.
1. TL;DR
Withholding server heartbeat acknowledgements makes the daemon reconnect. Keeping the host disconnected then makes the server interrupt active work and emit the reported connection-lost message. Both mechanisms were exercised with existing trusted-main tests in two clean checkouts. The tests inject the missing acknowledgements and disconnect; they do not reproduce why the real deployment lost acknowledgements, a seven-hour Claude session, or the reported delayed question completion. No safe automatic repair follows from this evidence.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| Missing acknowledgements cause a reconnect | Verified under injection | Host connection suite: 17 tests pass in each checkout. |
| Prolonged host absence emits interrupted completion, error, then interruption event | Verified under injection | Server reconciliation suite: 12 tests pass in each checkout. |
| A brief reconnect necessarily loses a turn | Not supported as a general claim | Same-instance reconnect with active thread preserves the turn in the server test. |
| Real heartbeat failure and seven-hour question lifecycle | Unverified | No incident host, network capture, matching macOS instance, or real provider session used. |
3. Environment
Trusted origin/main at the full commit above; Linux, Node v26.8.1, pnpm 9.15.0. Both checkouts used frozen installs and normal Turbo builds. No live app, provider, production data, or network listener was started; harnesses use temporary data and in-memory databases. The first temporary checkout attempt exhausted /tmp inodes; two successful checkouts were then created on a filesystem with capacity.
4. Minimal reproduction
- Clone get-bb/bb and check out the recorded commit in a fresh directory.
- Run the artifact below from its root.
- Repeat in a second fresh checkout at the same commit.
set -eu pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build pnpm exec turbo run test --force --filter=@bb/host-daemon -- src/server-connection.test.ts pnpm exec turbo run test --force --filter=@bb/server -- test/internal/background-task-reconciliation.test.ts
Command artifact. Expected and actual for both runs:
Host: Test Files 1 passed (1); Tests 17 passed (17) Server: Test Files 1 passed (1); Tests 12 passed (12)
The host test “reconnects when server heartbeat acknowledgements stop” acknowledges at 25 seconds, verifies no reconnect after 30 more seconds, then advances five seconds and asserts code 1013 and the timeout reason. The server test “records a lost host connection after the live event window elapses without a reconnect” invokes the real disconnect handler, verifies activity survives the short grace, then advances the active-work grace and asserts the three timeline events and error state. These are passing characterization tests, not a failing regression for a proven root defect.
Exact trusted test sources: host suite and server suite. The following is the key server test, unchanged:
it("records a lost host connection after the live event window elapses without a reconnect", async () => {
await withTestHarness(async (harness) => {
const { session, thread } = seedActiveTurnThread(harness);
vi.useFakeTimers();
handleDaemonSocketClosed(harness.deps, { sessionId: session.id });
await vi.advanceTimersByTimeAsync(DAEMON_DISCONNECT_GRACE_MS + 1);
expect(getThread(harness.deps.db, thread.id)?.status).toBe("active");
expect(
listEvents(harness.deps.db, { threadId: thread.id })
.filter((row) => row.type !== "turn/started")
.map((row) => row.type),
).toEqual([]);
await vi.advanceTimersByTimeAsync(
DAEMON_ACTIVE_WORK_DISCONNECT_GRACE_MS - DAEMON_DISCONNECT_GRACE_MS,
);
expect(getThread(harness.deps.db, thread.id)?.status).toBe("error");
expect(
listEvents(harness.deps.db, { threadId: thread.id })
.filter((row) => row.type !== "turn/started")
.map((row) => ({
data: JSON.parse(row.data),
type: row.type,
})),
).toEqual([
expect.objectContaining({
type: "turn/completed",
}),
expect.objectContaining({
data: expect.objectContaining({
code: "thread_command_failed",
message:
"Thread interrupted because the connection to the host was lost",
detail: "Please retry the thread to continue.",
}),
type: "system/error",
}),
expect.objectContaining({
data: {
reason: "host-daemon-restarted",
cause: "host-connection-lost",
},
type: "system/thread/interrupted",
}),
]);
});
});
5. Root cause
apps/host-daemon/src/server-connection.ts:828 compares the last acknowledged timestamp against the lease, logs the timeout, clears the heartbeat timer and reconnects. apps/server/src/ws/daemon-protocol.ts:282 sends an acknowledgement upon receiving a heartbeat. Neither fact identifies whether the incident lost delivery, stalled processing, or encountered another runtime condition.
apps/server/src/constants.ts:1 sets a 5-second pending-interaction grace and a 30-second active-work grace. apps/server/src/internal/session-owner-side-effects.ts:148 schedules both on disconnect. The short grace interrupts pending interactions; apps/server/src/internal/session-owner-side-effects.ts:267 interrupts active threads after the long grace when no daemon is registered, using host-daemon-restarted with cause host-connection-lost. apps/server/src/services/threads/thread-lifecycle.ts:413 maps the cause to the connection-lost error and retry detail. That explains the visible symptom but is not a diagnosis of the initiating transport failure.
6. Proposed next experiment
Capture synchronized daemon heartbeat-send/ack-receive timing and server receive/ack-send timing alongside event-loop delays and connection transitions in an isolated deployment. Exercise a pending question across disconnect and same-instance reconnect, checking both pending-interaction and tool-item terminal states. Do not automatically retry a possibly mutating turn. Extending a timeout without identifying the cause would hide the symptom. No PR: the initiating bug is not reproduced and no failing regression justifies a contained repair.
7. Verification
The same agent repeated frozen install, build, and both suites in a second clean checkout at the exact commit. Turbo --force prevented test cache replay. Both suites passed again: 29 tests total. No report correction was needed; the verdict remains partial. This is a repeat verification by the same agent, not an independent review.
8. Related issues
#3407 also discusses heartbeat timeouts, but no shared initiating cause was established. No open PR referencing #4364 was found in the issue timeline or open-PR search at investigation time.
9. Appendix
All issue text and comments were treated as untrusted claims. No linked attachment, script, or PR code was executed. Source evidence comes exclusively from trusted main. No source fix or new test was added.