Reports

#4186 · Ready environment adoption triggers stale-selection cancellation

Bug · Priority: Medium · Effort: Medium · workspaces · 2026-09-23 · Issue

REPRODUCED · Root-cause confidence: high

1. TL;DR

A newly preparing thread can adopt a ready checkout and then destroy its environment record. Adoption preserves the original provider selection, but preparation compares it with the new request and cancels when they differ. Cancellation bypasses the no-retirement policy before the new thread is attached. A focused server regression reproduces this on main in two separate checkouts using a real migrated SQLite database.

2. Claims vs findings

ClaimFindingEvidence
An older branch selection causes destruction on adoptionVerifiedBoth baseline runs return destroyed where ready is expected.
No-retirement policy prevents ordinary expiryVerified, but cancellation overrides itTest uses retireGraceMs: null.
Files remain intactConsistent with source; not exercised on a real checkoutCheckout remove returns removed without filesystem operations.
Composer/CLI timing, WSL and historical incidentsUnverifiedNo live user install accessed; reproduction isolates the server lifecycle.

3. Environment

Trusted origin/main: 853e1e9a1ca68c3978c88d16deb05ee4154cadd8. macOS arm64, Node 22.22.3, pnpm 9.15.0 via an isolated Corepack shim, Vitest 4.1.1. Frozen installation and all 60 build tasks passed in each checkout. The initial system pnpm executable was broken; the shim resolved that tooling issue. The in-process provider returns the existing path and a successful no-op removal, matching the checkout behavior relevant here. No real agent, external server, network ports or user runtime data were used. Each harness creates and removes a fresh temporary data directory and uses migrated SQLite.

4. Minimal reproduction

Use a fresh trusted checkout and download the regression patch. It adds one test to the repository's existing provider-orchestration suite, reusing its setup helpers.

git checkout 853e1e9a1ca68c3978c88d16deb05ee4154cadd8
pnpm install --frozen-lockfile --prefer-offline
pnpm exec turbo run build
git apply /absolute/path/to/regression.patch
pnpm exec turbo run test --filter=@bb/server -- test/services/environments/provider-orchestration.test.ts --testNamePattern='does not cancel an adopted checkout' 

The test seeds a ready environment with an older branch input, prepares another thread on that path, waits for adoption, and performs the next preparation pass and lifecycle sweep.

Expected: "ready"
Received: "destroyed"
Test Files  1 failed (1)
Tests  1 failed | 57 skipped (58)

Test body (imports and setup are supplied by the patched repository file):

  it("does not cancel an adopted checkout with a previous selection", async () =>
    withTestHarness(async (harness) => {
      const remove = vi.fn<PluginEnvironmentProviderDeclaration["remove"]>(
        async () => ({ status: "removed" }),
      );
      const fixture = setup(harness, {
        create: async () => ({
          status: "created",
          path: "/tmp/adoption-regression",
          ownsPath: false,
        }),
        remove,
        policy: { retireGraceMs: null },
      });
      const existing = seedEnvironment(harness.deps, {
        projectId: fixture.context.project.id,
        hostId: fixture.host.id,
        path: "/tmp/adoption-regression",
        status: "ready",
        providerOwnsPath: false,
        environmentProviderId: fixture.record.provider.id,
        environmentProviderPluginId: "test",
        environmentProviderInstanceKey: "original-key",
      });
      harness.db.update(environments).set({
        environmentProviderSelection: {
          machine: fixture.context.machine,
          inputs: { branch: { kind: "new", baseBranch: "main" } },
        },
      }).where(eq(environments.id, existing.id)).run();
      fixture.ask();
      await fixture.settled();
      expect(fixture.row().id).toBe(existing.id);
      const next = fixture.ask();
      await sweepProviderEnvironment(harness.deps, existing.id);
      expect(getEnvironment(harness.db, existing.id)?.status).toBe("ready");
      expect(next.action).toBe("ready");
      expect(remove).not.toHaveBeenCalled();
    }));

5. Root cause

  1. bindEnvironmentPath transfers preparation ownership to the new thread while retaining the existing provider selection and lifecycle identity.
  2. prepareProviderEnvironment compares that retained selection with current inputs even though the row is ready. The mismatch enters cancellation.
  3. cancelProviderEnvironmentCreation records running teardown and immediate retirement.
  4. The lifecycle sweep observes cancellation with no attached live threads and removes the record despite null retirement grace.
  5. Checkout removal reports success without deleting files.
environmentProviderSelection: existing.environmentProviderId === null
  ? current.environmentProviderSelection
  : existing.environmentProviderSelection

The retained selection is lifecycle metadata, not reliable evidence that this thread's completed adoption request changed. Existing cleanup-preservation coverage requires retaining it.

6. Proposed fix and validation

Limit input-selection cancellation to environments that are not ready. Continue checking provider identity and host identity, and preserve original cleanup metadata. The local fix changes four production lines and removes two; the regression adds 40 lines. Total: 46 changed text lines in two files in the server environment subsystem. No stored-data, schema, protocol or dependency change.

The regression and relevant existing suites passed: 83 tests across provider orchestration, shared workspace creation, and provisioning recovery. Server typecheck passed. The before-fix failure was recorded before any production edit. Fixed test output · Typecheck output.

7. Verification

The same agent repeated the reproduction in a second clean detached worktree at the full base commit, under a separate temporary work directory. After its own frozen install and full build, only the regression patch was added; production code remained unchanged. The exact focused command above failed with the same ready-versus-destroyed assertion (3.12 seconds). The first focused run took 3.33 seconds. No report correction was needed. This is a repeated check by the same agent, not independent verification.

8. Related issues and PRs

GitHub issue timeline metadata and an open-PR search for issue 4186 returned no linked open pull request at investigation time. No other issue's claims were used as evidence for this result.

9. Appendix

First baseline log · Second baseline log. Logs retain test output while removing local home and checkout paths. Additional checks: pnpm exec turbo run typecheck --filter=@bb/server, git diff --check, and git diff --numstat origin/main. No live development server was started; harness cleanup ran after each test.

Issue content was treated as untrusted evidence. No supplied issue commands, patches or linked branches were executed.