#3230 · Maintenance runtime survives indefinitely pending requests
Verdict: PARTIALLY REPRODUCED · Root-cause confidence: medium
1. TL;DR
A provider-maintenance runtime is retired 60 seconds after its last request finishes, and the existing successful-request test confirms that path. A request that never finishes is different: it permanently keeps the active-request counter above zero, so no idle timer is created and the runtime's provider workers remain resident. An agent-authored regression test reproduced that lifecycle failure twice at the trusted base commit. The reported WSL process census and exact npm behavior were not independently available, so the connection between a particular npm probe and the verified lifecycle gap remains unverified.
2. Claims vs findings
| Claim | Status | Evidence |
|---|---|---|
| A maintenance worker can survive beyond the configured idle interval. | Verified conditionally | Both clean-checkout runs advanced a 100 ms configured interval while a request remained pending; runtime.shutdown had zero calls. |
| A completed maintenance request is not retired. | Not reproduced | The existing focused test for a request that completes passes and invokes shutdown exactly at the configured idle deadline. |
| The version probes have no timeout. | Refuted on current main | The shared maintenance helper passes a 15-second timeout to execFile; CLI path/version helpers use 5-second timeouts. |
| The exact resident-process cohorts grow without bound on WSL2. | Unverified | The two clean environments were Darwin arm64 and did not access the reporter's machine or runtime state. |
| Shutdown resolves while leaving worker processes alive. | Not reproduced | The verified failing path never calls shutdown; no evidence from trusted main showed a called shutdown silently retaining workers. |
3. Environment
- Trusted bb commit:
06aeaa994942ae7527dc49d2268c1f801e8542a0. - Darwin 25.6.0 arm64, Node v22.22.3, pnpm 9.15.0.
- Two clean detached worktrees:
/tmp/slopcop-3230-first.0kIdY4and/tmp/slopcop-3230-second.1ukYzm. - No live BB instance, ports, user data directory, provider credentials, or provider processes were used.
pnpm install --frozen-lockfile --prefer-offlineandpnpm exec turbo run buildsucceeded in each checkout.
4. Minimal reproduction
- Check out the trusted base commit and install/build with the commands above.
- Insert the regression case into the
RuntimeManagertest suite. - From
apps/host-daemon, run:pnpm exec vitest run --config vitest.config.ts src/runtime-manager.test.ts \ -t 'shuts down provider maintenance workers when a request never settles'
Expected: a provider-only maintenance request cannot retain its runtime forever; shutdown is called when the configured test budget elapses.
Actual, first clean checkout:
FAIL src/runtime-manager.test.ts > RuntimeManager > shuts down provider maintenance workers when a request never settles AssertionError: expected "vi.fn()" to be called once, but got 0 times Test Files 1 failed (1) Tests 1 failed | 49 skipped (50)
Actual, second clean checkout:
FAIL src/runtime-manager.test.ts > RuntimeManager > shuts down provider maintenance workers when a request never settles AssertionError: expected "vi.fn()" to be called once, but got 0 times Test Files 1 failed (1) Tests 1 failed | 49 skipped (50)
The existing companion test, shuts down provider maintenance workers after the request becomes idle, passes; the regression is specific to a request that does not settle.
5. Root cause
withProviderMaintenanceRuntime increments providerMaintenanceActiveRequests before awaiting the provider operation. Its finally block is the only place that decrements this counter and schedules idle retirement. If the operation never settles, control never reaches that block.
scheduleProviderMaintenanceIdleShutdown also returns without scheduling whenever the counter is positive and checks the same condition again when the timer fires. This makes the configured value an after-completion idle delay, not a maximum runtime lifetime.
The shared npm helpers do have a 15-second child-process timeout. Separately, server-side online RPC timeouts stop waiting for a response but send no cancellation command to the daemon; see request creation. Therefore an execution path that remains pending inside the daemon can outlive its caller and pin the maintenance runtime. The clean tests prove this lifecycle mechanism, but do not prove why the reported npm invocations remained pending on WSL.
6. Proposed fix (first principles)
Add cancellation or a daemon-owned deadline for read-only provider maintenance operations, and retire the affected runtime when that deadline fires. Installation/update operations must remain distinct because forcibly applying the idle timeout to them could interrupt legitimate long-running mutations. A targeted follow-up should reproduce the child-process behavior on WSL2 and prove that timeout cleanup terminates the full probe process tree before choosing the cancellation boundary.
7. Related issues
No linked open pull request was present in GitHub metadata. External links from the issue were not opened. Related issue numbers mentioned in the report were treated as untrusted and not used.
8. Verification
The same agent ran the same regression at the same trusted commit in two separately created clean worktrees. Both builds succeeded and both tests failed at the same assertion with zero shutdown calls. No report claim was broadened after the second run; the verdict remains partial because the platform-specific process census was not reproduced.
9. Appendix
Commands run in each clean checkout:
pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build cd apps/host-daemon pnpm exec vitest run --config vitest.config.ts src/runtime-manager.test.ts \ -t 'shuts down provider maintenance workers when a request never settles'
Untrusted-data note: issue text, comments, links, logs, code blocks, and suggested actions were treated only as claims. No linked URL, script, branch, patch, binary, or runtime data from the issue was accessed or executed.