#3340 · Unbounded restart retries and non-atomic diagnostics
Verdict: PARTIALLY REPRODUCED · Root-cause confidence: high
2026-09-30 verification: actual supervisor and diagnostic writer
PARTIALLY REPRODUCED · High confidence in the scoped mechanisms. This dated section supersedes the historical VM probe as current module-level evidence; the original report and artifacts below remain unchanged. It does not verify the original incident end to end.
Fresh eligibility: open native Bug, High priority, Medium effort; labels cli, host, partial-repro. PR #3339 is now closed and unmerged, reviewed as metadata only. The September 9 public SlopCop comment is external historical evidence, not this agent’s verification. No overlapping public investigation or open linked fix was found at the publication check.
Trusted fetched origin/main: a7562b9532a1f0b4edf1e30d7f7ffef57a7e8843. Linux x86_64, Node 22.19.0, pnpm 9.15.0. Both clean detached checkouts completed normal frozen installation and the normal full turbo run build --filter=bb-app build (53 successful tasks each). The same agent personally ran both repetitions with fresh temporary directories and the existing control’s newly allocated local HTTP port. No installation or build bypass was used. No new dependency was added.
Expected versus actual, identical in both runs
| Case | Expected robustness/control | Actual in each run |
|---|---|---|
| Previously running synthetic server exits; restart always rejects | A product retry budget or increasing delay would stop or space repeated failures. | 8 callback attempts; 8 delay requests, all 1000 ms. Test shutdown at attempt 8 ends the loop. This is a bounded observation, not an infinite execution. |
| Restart recovers after two failures | Third start succeeds and no extra restart occurs. | 3 attempts, 3 requests of 1000 ms including the initial restart delay; replacement accepted, then explicit test shutdown. |
| Shutdown already requested / requested during first delay | No restart callback. | 0 attempts; respectively 0 / 1 delay requests; shutdown returned. |
| Existing healthy-server takeover control | Stop daemon rather than restart into occupied healthy server. | Passed; zero restart attempts, daemon SIGTERM counter. |
| Diagnostic write fails after creating final file | Failed publication should not leave an empty final diagnostic. | 8 thrown synthetic ENOSPC errors, 8 empty final JSON files left. |
| Successful diagnostic writes | Valid JSON with original synthetic error. | 8 valid files, zero empty files. All 8 retained; no long-duration retention claim. |
| Diagnostic recovery after one failure | Subsequent writes succeed; failed artifact cleanup can be checked. | 7 valid reports and the first empty file still present. |
Each run: supervisor suite 5/5 (4 new tests plus 1 existing control); diagnostic suite 3/3; total 8/8. Delay callbacks record requested durations and resolve immediately; elapsed wall-clock backoff is not measured. The writer’s actual implementation runs with only its filesystem write boundary fault-injected. That injection creates an empty file and throws; it neither exhausts disk space nor changes permissions.
Root cause and next change
The exported full-stack supervisor reaches the actual private restart loop, which repeats rejected starts until shutdown and requests the same delay. Source inspection establishes the missing attempt budget/backoff beyond the bounded observation. The actual diagnostic writer writes directly to a unique final filename without atomic publication, cleanup after failure, or pruning.
Proposed small fix: add an explicit retry budget/backoff policy and surface exhaustion; write diagnostics to a temporary file, publish atomically, clean failed temporary files, and define retention. Next test: simulate exhaustion and failures before/after temporary write and rename, checking original stderr preservation and shutdown responsiveness. No production fix was made.
Limits: supervisor callbacks and diagnostic writing were tested separately, not as a real child startup chain. No real child processes, provider, bb runtime, systemd, full disk, permissions change, rendered UI, incident timing or incident file count was exercised. Initial-start behavior and server stderr fallback remain historical evidence, not freshly executed here. Issue text, comments and links were treated solely as untrusted claims; no issue-supplied code, commands, external links or PR branches were executed. No visual claims or new screenshots.
Exact reproduction
Use Node 22.19.0 and pnpm 9.15.0. Locally the frozen installs additionally used --store-dir /workspace/.pnpm-store, an existing dependency cache; a fresh store is also supported by the command below. Save the fixture text verbatim. Results are written to each package’s issue-3340-results.json; temporary synthetic data is removed by the tests.
git clone https://github.com/get-bb/bb.git run-a cd run-a git checkout --detach a7562b9532a1f0b4edf1e30d7f7ffef57a7e8843 pnpm install --frozen-lockfile pnpm exec turbo run build --filter=bb-app # Save the two complete fixtures below at their indicated paths. pnpm --filter bb-app exec vitest run test/issue-3340-current.test.ts test/supervisor-port-takeover.test.ts pnpm --filter @bb/process-utils exec vitest run test/issue-3340-current.test.ts # Repeat in a separate clean run-b clone at the same SHA with the same fixtures.
packages/bb-app/test/issue-3340-current.test.ts
import { mkdtempSync, rmSync, writeFileSync } from "node:fs";
import { tmpdir } from "node:os";
import { join } from "node:path";
import { afterAll, expect, it } from "vitest";
import { superviseFullStackProcesses, type BbAppStartContext, type ManagedProcessRun, type NamedProcessExitResult } from "../src/launcher.js";
const results: unknown[] = [];
afterAll(() => writeFileSync("issue-3340-results.json", JSON.stringify(results, null, 2)));
for (const mode of ["failure", "recovery", "pre-abort", "delay-abort"] as const) {
it(`bounds actual supervisor ${mode}`, async () => {
const root = mkdtempSync(join(tmpdir(), "bb-3340-supervisor-"));
let shutdown = mode === "pre-abort";
let attempts = 0;
const delays: number[] = [];
const makeRun = (processName: "server" | "daemon") => {
let finish!: (v: NamedProcessExitResult) => void;
const run: ManagedProcessRun = { exit: new Promise(resolve => { finish = resolve; }), terminate: async signal => finish({ processName, result: {code:null, signal} }) };
return {run, finish: () => finish({processName, result:{code:1, signal:null}})};
};
const server = makeRun("server");
const daemon = makeRun("daemon");
const replacement = makeRun("server");
const processes = {serverRun: server.run as ManagedProcessRun | null, daemonRun: daemon.run as ManagedProcessRun | null};
const context = Object.fromEntries(["appDistDir","configFile","daemonBundleDir","daemonEntry","daemonLockDir","daemonLockFile","dataDir","dbPath","envFile","logDir","packageRoot","serverEntry"].map(key => [key, join(root,key)])) as unknown as BbAppStartContext;
Object.assign(context, {appVersion:"0.0.0-test",serverPort:0,daemonPort:0,serverUrl:"http://127.0.0.1:0"});
const supervision = superviseFullStackProcesses({context,processes,
isShutdownRequested: () => shutdown || attempts >= 8,
delayMilliseconds: async ({ms}) => {delays.push(ms); if(mode === "delay-abort") shutdown = true;},
isHealthyServerAnswering: async () => false,
readServerMovedFile: async () => null,
onServerMoved: async () => {throw new Error("Unexpected move");},
startDaemon: async () => {throw new Error("Unexpected daemon restart");},
startServer: async () => {
attempts++;
if(mode === "recovery" && attempts === 3) {
processes.serverRun = replacement.run;
setTimeout(() => {shutdown = true; replacement.finish();}, 0);
return replacement.run;
}
throw new Error("Synthetic bounded start failure");
}
});
server.finish();
try {
const outcome = await supervision;
expect(outcome).toBe("shutdown");
expect(attempts).toBe(mode === "failure" ? 8 : mode === "recovery" ? 3 : 0);
expect(delays).toEqual(Array(mode === "failure" ? 8 : mode === "recovery" ? 3 : mode === "delay-abort" ? 1 : 0).fill(1000));
results.push({mode,attempts,delays,outcome});
} finally {
shutdown = true; server.finish(); daemon.finish(); replacement.finish();
await supervision;
rmSync(root,{recursive:true,force:true});
}
}, 5000);
}
packages/process-utils/test/issue-3340-current.test.ts
import { mkdtempSync, readdirSync, readFileSync, rmSync, writeFileSync } from "node:fs";
import { join } from "node:path";
import { tmpdir } from "node:os";
import { afterAll, expect, it, vi } from "vitest";
const fault = vi.hoisted(() => ({enabled:false}));
vi.mock("node:fs", async importOriginal => {
const actual = await importOriginal<typeof import("node:fs")>();
return {...actual, writeFileSync: (...args: Parameters<typeof actual.writeFileSync>) => {
if(fault.enabled && typeof args[0] === "string" && args[0].includes("process-synthetic-server-")) {
actual.writeFileSync(args[0], "");
throw Object.assign(new Error("Synthetic write failure"), {code:"ENOSPC"});
}
return actual.writeFileSync(...args);
}};
});
import { writeSafeProcessDiagnosticReport } from "../src/index.js";
const results: unknown[] = [];
afterAll(() => writeFileSync("issue-3340-results.json", JSON.stringify(results,null,2)));
for(const mode of ["failure","success","recovery"] as const) {
it(`actual diagnostic writer ${mode}`, () => {
const root = mkdtempSync(join(tmpdir(), "bb-3340-diagnostic-"));
let failures = 0;
try {
for(let i=0;i<8;i++) {
fault.enabled = mode === "failure" || (mode === "recovery" && i === 0);
try {
writeSafeProcessDiagnosticReport({logsDir:root,processName:"synthetic-server",kind:"startupFailure",error:new Error("Synthetic start failure"),now:()=>new Date("2026-09-30T00:00:00Z"),createReportId:()=>`attempt-${i}`});
} catch(error) {expect((error as NodeJS.ErrnoException).code).toBe("ENOSPC"); failures++;}
}
const files = readdirSync(root);
const contents = files.map(file => readFileSync(join(root,file),"utf8"));
const empty = contents.filter(content => content.length === 0).length;
const valid = contents.filter(Boolean).map(content => JSON.parse(content));
expect(files.length).toBe(8);
expect(failures).toBe(mode === "failure" ? 8 : mode === "recovery" ? 1 : 0);
expect(empty).toBe(failures);
expect(valid.every(report => report.kind === "startupFailure" && report.error.message === "Synthetic start failure")).toBe(true);
results.push({mode,attempts:8,failures,files:files.length,empty,valid:valid.length});
} finally {fault.enabled = false; rmSync(root,{recursive:true,force:true});}
});
}
Sanitized result objects (identical in both clean runs)
[
{
"mode": "failure",
"attempts": 8,
"delays": [
1000,
1000,
1000,
1000,
1000,
1000,
1000,
1000
],
"outcome": "shutdown"
},
{
"mode": "recovery",
"attempts": 3,
"delays": [
1000,
1000,
1000
],
"outcome": "shutdown"
},
{
"mode": "pre-abort",
"attempts": 0,
"delays": [],
"outcome": "shutdown"
},
{
"mode": "delay-abort",
"attempts": 0,
"delays": [
1000
],
"outcome": "shutdown"
}
]
[
{
"mode": "failure",
"attempts": 8,
"failures": 8,
"files": 8,
"empty": 8,
"valid": 0
},
{
"mode": "success",
"attempts": 8,
"failures": 0,
"files": 8,
"empty": 0,
"valid": 8
},
{
"mode": "recovery",
"attempts": 8,
"failures": 1,
"files": 8,
"empty": 1,
"valid": 7
}
]1. TL;DR
A managed child that exits after startup can enter a restart loop whose failures have no attempt budget and always request a one-second delay. Diagnostic reports are written directly to their final JSON filenames with no retention or cleanup after a failed write. Two clean-checkout source-function probes each reached our 32-attempt safety stop, retained all 32 successful reports, and left 32 empty JSON files under an injected storage-write failure. A real disk-full incident and Linux systemd behavior were not replayed. The server already preserves its original startup error on stderr when diagnostic writing throws.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| Failed managed restarts continue without backoff or a limit | Verified at source-function level | 32 attempts; every requested delay 1000 ms; only the probe's shutdown flag terminates the loop. |
| Storage write failures accumulate empty JSON reports | Verified with fault injection | Open/truncate succeeds, then injected ENOSPC throws before content is written: 32 of 32 files remain empty. |
| Diagnostic history is unbounded | Verified | 32 successful writes leave 32 reports; source has no pruning. |
| Original startup evidence must survive diagnostic failure | Already supported by server handler | Injected diagnostic error still produces original stderr text and exit code 1. |
| Any initial startup failure loops inside the launcher | Not supported on this base | Initial startServer failure is caught and exits with code 1. A Restart=always service can restart that supervisor externally. The internal unlimited loop is reached after a previously running child exits. |
| Incident counts, real ENOSPC, and systemd terminal-state handling | Unverified | No Linux service manager or full filesystem was exercised. |
3. Environment
Darwin arm64; Node v22.22.3. Two detached, clean worktrees at the exact commit above, named first and second inside a newly created temporary work directory. No provider, database, listening port, imported store, live core, or user runtime data was used. Each probe creates and removes its own new temporary filesystem directory. The probe uses only Node built-ins and unchanged source functions, with TypeScript annotations stripped by Node.
Full install/build status is recorded in the appendix. These probes do not depend on a successful monorepo build, and no existing-suite pass is claimed.
4. Minimal reproduction
- Fetch trusted main and create two detached worktrees at the recorded commit.
- Save probe.mjs outside both checkouts.
- Run the probe in each checkout as below. It exits 1 when the four safety expectations fail. No disk needs to be filled.
git fetch origin main git worktree add --detach /tmp/bb-3340-first e13605d71e261d863e907801987f7df9d6a6243b git worktree add --detach /tmp/bb-3340-second e13605d71e261d863e907801987f7df9d6a6243b node --disable-warning=ExperimentalWarning probe.mjs /tmp/bb-3340-first node --disable-warning=ExperimentalWarning probe.mjs /tmp/bb-3340-second
Expected: a finite retry budget before the probe safety stop, an increasing delay, removal of older reports, and no final JSON after a failed write. The 32-attempt boundary is a test safety limit, not a proposed product policy. Actual output from both runs:
{"attempts":32,"delays":[1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000,1000],"safetyStop":true,"result":null}
FAIL repeated failures stop before the 32-attempt safety limit: Expected values to be strictly equal:
true !== false
FAIL retry delay increases: The expression evaluated to a falsy value:
assert.ok(new Set(delays).size > 1)
{"successfulWrites":32,"retained":32}
FAIL retention removes older reports: The expression evaluated to a falsy value:
assert.ok(retained < 32)
{"stderr":["Error: controlled startup failure\n"],"exitCode":1}
PASS startup evidence survives diagnostic write failure
{"failedWrites":32,"published":32,"empty":32}
FAIL failed writes leave no published JSON files: Expected values to be strictly equal:
32 !== 0
Failed safety assertions: 4
The probe slices the exact named functions out of trusted source and executes them in a VM. Retry starts, delays, and cosmetic logging are controlled. Diagnostic writes use a real temporary directory; the ENOSPC case deliberately creates a real empty file and then throws at the write boundary. This isolates the cleanup defect without claiming to reproduce kernel-level ENOSPC. Database access is neither used nor mocked.
Complete reproduction source
import assert from 'node:assert/strict';
import * as fs from 'node:fs';
import { tmpdir } from 'node:os';
import { join, resolve } from 'node:path';
import { randomUUID } from 'node:crypto';
import { stripTypeScriptTypes } from 'node:module';
import { runInNewContext } from 'node:vm';
const root = resolve(process.argv[2] ?? '.');
const scratch = fs.mkdtempSync(join(tmpdir(), 'bb-3340-probe-'));
const launcher = fs.readFileSync(join(root, 'packages/bb-app/src/launcher.ts'), 'utf8');
const utilities = fs.readFileSync(join(root, 'packages/process-utils/src/index.ts'), 'utf8');
const restartSource = launcher.slice(launcher.indexOf('async function restartManagedProcess('), launcher.indexOf('export async function terminateManagedFullStackProcesses('));
assert.ok(restartSource.length > 100);
const retryConstant = launcher.match(/const MANAGED_PROCESS_RESTART_RETRY_DELAY_MS = [\d_]+;/)?.[0];
assert.ok(retryConstant);
const restart = runInNewContext(stripTypeScriptTypes(`${retryConstant}\n${restartSource}\nrestartManagedProcess`), {
beginStep() {}, endStep() {}, logManagedProcessStartupFailureContext() {},
formatManagedProcessLabel: x => x, formatManagedProcessName: x => x,
green: x => x, red: x => x,
});
const diagnosticSource = utilities.slice(utilities.indexOf('function createCurrentDiagnosticDate('), utilities.indexOf('export function installSafeProcessDiagnostics(')).replace('export function ', 'function ');
assert.ok(diagnosticSource.length > 100);
const constants = utilities.match(/const MAX_DIAGNOSTIC_[A-Z_]+ = \d+;/g).join('\n');
function diagnosticWriter(writeFileSync) {
return runInNewContext(stripTypeScriptTypes(`${constants}\n${diagnosticSource}\nwriteSafeProcessDiagnosticReport`), {
mkdirSync: fs.mkdirSync, writeFileSync, join, randomUUID, process, Error, AggregateError,
});
}
let failures = 0;
function check(name, test) {
try { test(); console.log(`PASS ${name}`); }
catch (error) { failures++; console.log(`FAIL ${name}: ${error.message}`); }
}
try {
let attempts = 0;
let safetyStop = false;
const delays = [];
const result = await restart({
context: {}, processName: 'server', isShutdownRequested: () => safetyStop,
start: async () => { attempts++; throw new Error('controlled startup failure'); },
delayMilliseconds: async ({ms}) => { delays.push(ms); if (attempts === 32) safetyStop = true; },
});
console.log(JSON.stringify({attempts, delays, safetyStop, result}));
check('repeated failures stop before the 32-attempt safety limit', () => assert.equal(safetyStop, false));
check('retry delay increases', () => assert.ok(new Set(delays).size > 1));
const normalDir = join(scratch, 'normal');
const writer = diagnosticWriter(fs.writeFileSync);
for (let i = 0; i < 32; i++) writer({ logsDir: normalDir, processName: 'server', kind: 'startupFailure', error: new Error('controlled startup failure') });
const retained = fs.readdirSync(normalDir).length;
console.log(JSON.stringify({successfulWrites: 32, retained}));
check('retention removes older reports', () => assert.ok(retained < 32));
const server = fs.readFileSync(join(root, 'apps/server/src/index.ts'), 'utf8');
const startupSource = server.slice(server.indexOf('function reportStartupFailure('), server.indexOf('async function main('));
const stderr = [];
const fakeProcess = {stderr: {write: value => stderr.push(value)}, exitCode: undefined};
const reportFailure = runInNewContext(stripTypeScriptTypes(`${startupSource}\nreportStartupFailure`), {
writeSafeProcessDiagnosticReport() { throw new Error('controlled diagnostic failure'); },
diagnosticsLogsDir: scratch, process: fakeProcess, Error,
});
const startupError = new Error('controlled startup failure');
startupError.stack = 'Error: controlled startup failure';
reportFailure(startupError);
console.log(JSON.stringify({stderr, exitCode: fakeProcess.exitCode}));
check('startup evidence survives diagnostic write failure', () => {
assert.deepEqual(stderr, ['Error: controlled startup failure\n']);
assert.equal(fakeProcess.exitCode, 1);
});
const failingDir = join(scratch, 'write-failure');
const failedWriter = diagnosticWriter((path) => {
fs.closeSync(fs.openSync(path, 'w'));
throw Object.assign(new Error('controlled storage write failure'), {code: 'ENOSPC'});
});
for (let i = 0; i < 32; i++) {
assert.throws(() => failedWriter({ logsDir: failingDir, processName: 'server', kind: 'startupFailure', error: new Error('controlled startup failure') }), {code: 'ENOSPC'});
}
const published = fs.readdirSync(failingDir);
const empty = published.filter(name => fs.statSync(join(failingDir, name)).size === 0).length;
console.log(JSON.stringify({failedWrites: 32, published: published.length, empty}));
check('failed writes leave no published JSON files', () => assert.equal(published.filter(name => name.endsWith('.json')).length, 0));
} finally {
fs.rmSync(scratch, {recursive: true, force: true});
}
console.log(`Failed safety assertions: ${failures}`);
process.exitCode = failures ? 1 : 0;
5. Root cause
restartManagedProcess loops only on the shutdown flag. Its catch block neither counts failures nor distinguishes unrecoverable failures; it requests the constant 1000 ms delay each time. The supervisor invokes this loop after a server exit. A finite probe cannot prove infinite execution by itself; the unbounded loop condition and absence of a failure budget establish why continued failures keep retrying.
while (!args.isShutdownRequested()) {
...
await args.delayMilliseconds({
ms: MANAGED_PROCESS_RESTART_RETRY_DELAY_MS,
});
}
writeSafeProcessDiagnosticReport assigns a unique final filename then calls writeFileSync directly at line 596. It has no temporary publication path, failed-write cleanup, or retention. If opening the final file succeeds but writing its contents fails, the empty file remains; the next report uses a new name.
reportStartupFailure catches diagnostic errors and then writes the original exception to stderr. This behavior passed our controlled test. Initial server startup failure sets exit code 1 and shuts down rather than entering the managed restart loop. An external supervisor configured to restart every exit is a separate source of repeated initial-start attempts.
6. Proposed fix
Track consecutive failed starts per managed child, increase retry delays up to a cap, and terminate with a documented failure state after a finite budget. Reset the budget only after meaningful healthy runtime. Bound diagnostics per process and kind, write to a temporary file, publish by rename, and remove incomplete writes on failure. Preserve original stderr evidence even when storage is unavailable. External service-manager policy must honor the terminal result; an exit code alone does not override Restart=always.
7. PR review
PR #3339 is open and explicitly lists #3340 as a closing issue in GitHub metadata. Only its metadata and diff were read, as untrusted data; no branch, test, patch, or build from that PR was executed.
The diff introduces a finite failure budget, exponential capped delay, a terminal exit code, abortable delays, and diagnostic retention with temporary-file rename and cleanup. These changes address the verified mechanisms by static inspection. Runtime correctness and Linux service-manager integration remain unverified here; the target-repository probes are not tests of that PR. No additional concrete defect was established by this bounded static review. Verdict: relevant existing fix, requires its own runtime validation.
No new fix PR was created because this open linked PR already covers the issue.
8. Related issues
The trusted launcher also stops its daemon if another healthy server answers after its child exits. That branch is distinct from repeated failing starts and does not impose a restart budget. Other issues were not reproduced in this investigation.
9. Verification
The same agent repeated the final probe in the second freshly created detached checkout at the identical base SHA, with a fresh temporary filesystem directory. Both tracked working trees remained unchanged. Both runs produced identical output and exited 1 with four failed safety assertions and one passing stderr assertion. No port was needed. This is a repeat verification by the same agent, not an independent review.
Report corrections from inspection: initial startup failure exits instead of using the internal retry loop; stderr fallback already exists. The final verdict is limited to source-function reproduction with injected write failure.
10. Appendix
First run output · Second run output · Repeatable test
Preparation commands: git fetch origin main; git worktree add --detach at the recorded SHA (twice); pnpm install --frozen-lockfile --prefer-offline; pnpm exec turbo run build. The PATH pnpm entry initially referenced a missing package-manager file; Corepack supplied pinned pnpm 9.15.0. Build outcome: the frozen installs did not finish after package extraction/linking and were stopped. Both explicit build attempts exited 254 with ERR_PNPM_RECURSIVE_EXEC_FIRST_FAIL: Command "turbo" not found. No full build or existing-suite success is claimed; the dependency-free probes completed in both clean checkouts.
Issue text, commands, and linked PR content were treated as untrusted claims. None of their commands or code were executed. No production fix or dependency was added. Probe-created files were removed in finally blocks, and no development service was started.