← reports

#3230 · Maintenance runtime survives indefinitely pending requests

BugPriority: MediumEffort: Lowhostprovidersperf open on GitHub 2026-09-07 · base 06aeaa994942

Verdict: PARTIALLY REPRODUCED · Root-cause confidence: medium

1. TL;DR

A provider-maintenance runtime is retired 60 seconds after its last request finishes, and the existing successful-request test confirms that path. A request that never finishes is different: it permanently keeps the active-request counter above zero, so no idle timer is created and the runtime's provider workers remain resident. An agent-authored regression test reproduced that lifecycle failure twice at the trusted base commit. The reported WSL process census and exact npm behavior were not independently available, so the connection between a particular npm probe and the verified lifecycle gap remains unverified.

2. Claims vs findings

ClaimStatusEvidence
A maintenance worker can survive beyond the configured idle interval.Verified conditionallyBoth clean-checkout runs advanced a 100 ms configured interval while a request remained pending; runtime.shutdown had zero calls.
A completed maintenance request is not retired.Not reproducedThe existing focused test for a request that completes passes and invokes shutdown exactly at the configured idle deadline.
The version probes have no timeout.Refuted on current mainThe shared maintenance helper passes a 15-second timeout to execFile; CLI path/version helpers use 5-second timeouts.
The exact resident-process cohorts grow without bound on WSL2.UnverifiedThe two clean environments were Darwin arm64 and did not access the reporter's machine or runtime state.
Shutdown resolves while leaving worker processes alive.Not reproducedThe verified failing path never calls shutdown; no evidence from trusted main showed a called shutdown silently retaining workers.

3. Environment

4. Minimal reproduction

  1. Check out the trusted base commit and install/build with the commands above.
  2. Insert the regression case into the RuntimeManager test suite.
  3. From apps/host-daemon, run:
    pnpm exec vitest run --config vitest.config.ts src/runtime-manager.test.ts \
      -t 'shuts down provider maintenance workers when a request never settles'

Expected: a provider-only maintenance request cannot retain its runtime forever; shutdown is called when the configured test budget elapses.

Actual, first clean checkout:

FAIL  src/runtime-manager.test.ts > RuntimeManager > shuts down provider maintenance workers when a request never settles
AssertionError: expected "vi.fn()" to be called once, but got 0 times
Test Files  1 failed (1)
Tests  1 failed | 49 skipped (50)

Actual, second clean checkout:

FAIL  src/runtime-manager.test.ts > RuntimeManager > shuts down provider maintenance workers when a request never settles
AssertionError: expected "vi.fn()" to be called once, but got 0 times
Test Files  1 failed (1)
Tests  1 failed | 49 skipped (50)

The existing companion test, shuts down provider maintenance workers after the request becomes idle, passes; the regression is specific to a request that does not settle.

5. Root cause

withProviderMaintenanceRuntime increments providerMaintenanceActiveRequests before awaiting the provider operation. Its finally block is the only place that decrements this counter and schedules idle retirement. If the operation never settles, control never reaches that block.

scheduleProviderMaintenanceIdleShutdown also returns without scheduling whenever the counter is positive and checks the same condition again when the timer fires. This makes the configured value an after-completion idle delay, not a maximum runtime lifetime.

The shared npm helpers do have a 15-second child-process timeout. Separately, server-side online RPC timeouts stop waiting for a response but send no cancellation command to the daemon; see request creation. Therefore an execution path that remains pending inside the daemon can outlive its caller and pin the maintenance runtime. The clean tests prove this lifecycle mechanism, but do not prove why the reported npm invocations remained pending on WSL.

6. Proposed fix (first principles)

Add cancellation or a daemon-owned deadline for read-only provider maintenance operations, and retire the affected runtime when that deadline fires. Installation/update operations must remain distinct because forcibly applying the idle timeout to them could interrupt legitimate long-running mutations. A targeted follow-up should reproduce the child-process behavior on WSL2 and prove that timeout cleanup terminates the full probe process tree before choosing the cancellation boundary.

7. Related issues

No linked open pull request was present in GitHub metadata. External links from the issue were not opened. Related issue numbers mentioned in the report were treated as untrusted and not used.

8. Verification

The same agent ran the same regression at the same trusted commit in two separately created clean worktrees. Both builds succeeded and both tests failed at the same assertion with zero shutdown calls. No report claim was broadened after the second run; the verdict remains partial because the platform-specific process census was not reproduced.

9. Appendix

Commands run in each clean checkout:

pnpm install --frozen-lockfile --prefer-offline
pnpm exec turbo run build
cd apps/host-daemon
pnpm exec vitest run --config vitest.config.ts src/runtime-manager.test.ts \
  -t 'shuts down provider maintenance workers when a request never settles'

Untrusted-data note: issue text, comments, links, logs, code blocks, and suggested actions were treated only as claims. No linked URL, script, branch, patch, binary, or runtime data from the issue was accessed or executed.