Reports

#4160 · Pruning loses storage cleanup ownership

BugPriority: MediumEffort: Highthreads, workspacesIssue2026-09-23

REPRODUCED · Root-cause confidence: high

1. TL;DR

Deleting a thread after its destroyed environment is pruned removes the database record without requesting storage cleanup from its host. The environment foreign key is cleared during pruning, and deletion treats this missing reference as proof that storage is already gone. A route-level regression test reproduces the lost cleanup request and premature hard deletion on main. A control with the same destroyed environment still present correctly retains a tombstone and sends a cleanup request. The same agent repeated the test in a second clean checkout with identical results.

2. Claims vs findings

ClaimFindingEvidence
Environment pruning clears a surviving thread's environment reference.VerifiedReal SQLite pruning deletes one environment; the test asserts a null thread reference.
Deleting that thread skips host cleanup and loses the row.VerifiedHTTP 200, row absent, zero cleanup commands in both runs. Control: row present, one command.
Host storage deletion removes the thread directory.Verified by source inspectionHost handler removes the contained per-thread path recursively. A live host was not launched.
Thousands of historical directories and several GB are affected on the reporter's machine.UnverifiedNo access to the reporter's runtime data; scale and installed release were not tested.
The null-reference path emits no missing-environment warning.Verified by source inspectionThe early return precedes the warning branch; logger output was not separately asserted.

3. Environment

Trusted public get-bb/bb main at 94d77da09568d7f0f7cdb23b84686e743acfb0ad, workspace package version 0.43.4; Darwin arm64, Node v22.22.3, pnpm 9.15.0, Vitest 4.1.1. Both checkouts completed the frozen install and Turbo build (60 successful tasks each). The broken global pnpm launcher was bypassed with npm exec --yes --package=pnpm@9.15.0 -- pnpm; no repository dependencies changed.

The harness uses a migrated in-memory SQLite database, a fresh temporary server data directory, and captured host RPCs. Requests run through Hono in process; no listening ports or live provider sessions are needed. Sentinel directories use a new mkdtemp root on each case and are removed in a finally block. The sentinel demonstrates the absence of cleanup, not a live daemon end-to-end filesystem test.

4. Minimal reproduction

  1. Check out the recorded trusted commit and build it.
  2. Save the regression test embedded below at the path below.
  3. Run the focused test. Expected failure is the pruned case's tombstone assertion; the unpruned control passes.
git clone https://github.com/get-bb/bb.git bb-4160
cd bb-4160
git checkout --detach 94d77da09568d7f0f7cdb23b84686e743acfb0ad
pnpm install --frozen-lockfile --prefer-offline
pnpm exec turbo run build
# Save the linked test as apps/server/test/threads/issue-4160.test.ts
pnpm exec turbo run test --filter=@bb/server -- test/threads/issue-4160.test.ts

The fixture archives an idle thread, marks its environment destroyed and removed with an expired timestamp, then invokes the real prune function before the real DELETE route. This isolates the post-teardown state without spawning a worktree or waiting seven days.

Expected: both cases retain deletion work and dispatch one storage command until the host acknowledges cleanup. Actual, both clean runs:

{"prune":false,"httpStatus":200,"rowPresent":true,"deleteCommands":1,"sentinel":"storage awaits host cleanup"}
{"prune":true,"httpStatus":200,"rowPresent":false,"deleteCommands":0,"sentinel":"storage awaits host cleanup"}

AssertionError: retain the tombstone until a host confirms cleanup: expected null not to be null
Tests  1 failed | 1 passed (2)

Raw logs and the standalone test are retained in the local investigation backup; the complete test and decisive output are embedded here so this report is self-contained.

Complete regression test
import { mkdtemp, mkdir, writeFile, readFile, rm } from "node:fs/promises";
import { tmpdir } from "node:os";
import { join } from "node:path";
import {
  environments,
  threads,
  getThread,
  pruneDestroyedEnvironments,
  DEFAULT_DESTROYED_ENVIRONMENT_EVENT_DETACH_BATCH_SIZE,
  DEFAULT_DESTROYED_ENVIRONMENT_PRUNE_BATCH_SIZE,
  DESTROYED_ENVIRONMENT_TTL_MS,
} from "@bb/db";
import { eq } from "drizzle-orm";
import { expect, it } from "vitest";
import { seedHostSession, seedProjectWithSource, seedEnvironment, seedThread } from "../helpers/seed.js";
import { withTestHarness } from "../helpers/test-app.js";
import { listQueuedThreadCommands } from "../helpers/commands.js";

it.each([false, true])("retains deletion work until host cleanup, prune=%s", async (prune) => {
  const storageRoot = await mkdtemp(join(tmpdir(), "bb-4160-storage-"));
  try {
    await withTestHarness(async (harness) => {
      const { host } = seedHostSession(harness.deps);
      const { project } = seedProjectWithSource(harness.deps, { hostId: host.id });
      const environment = seedEnvironment(harness.deps, {
        hostId: host.id,
        projectId: project.id,
        environmentProviderId: "personal-workspace",
        isGitRepo: false,
      });
      const thread = seedThread(harness.deps, {
        projectId: project.id,
        environmentId: environment.id,
        status: "idle",
      });
      harness.db.update(threads).set({ archivedAt: Date.now() - 100_000 })
        .where(eq(threads.id, thread.id)).run();
      harness.db.update(environments).set({
        status: "destroyed",
        teardownStatus: "removed",
        updatedAt: Date.now() - DESTROYED_ENVIRONMENT_TTL_MS - 100_000,
      }).where(eq(environments.id, environment.id)).run();
      const storage = join(storageRoot, thread.id);
      await mkdir(storage);
      await writeFile(join(storage, "sentinel.txt"), "storage awaits host cleanup");
      if (prune) {
        expect(pruneDestroyedEnvironments(harness.db, harness.hub, {
          updatedBefore: Date.now() - DESTROYED_ENVIRONMENT_TTL_MS,
          eventBatchSize: DEFAULT_DESTROYED_ENVIRONMENT_EVENT_DETACH_BATCH_SIZE,
          limit: DEFAULT_DESTROYED_ENVIRONMENT_PRUNE_BATCH_SIZE,
        }).deleted).toBe(1);
        expect(getThread(harness.db, thread.id)?.environmentId).toBeNull();
      }
      const response = await harness.app.request("/api/v1/threads/" + thread.id, {
        method: "DELETE",
        headers: { "content-type": "application/json" },
        body: JSON.stringify({ childThreadsConfirmed: false }),
      });
      expect(response.status).toBe(200);
      const remaining = getThread(harness.db, thread.id);
      const commands = listQueuedThreadCommands(harness, "thread.storage.delete", thread.id);
      const sentinel = await readFile(join(storage, "sentinel.txt"), "utf8");
      process.stdout.write(JSON.stringify({
        prune,
        httpStatus: response.status,
        rowPresent: remaining !== null,
        deleteCommands: commands.length,
        sentinel,
      }) + "\n");
      expect(sentinel).toBe("storage awaits host cleanup");
      expect(remaining, "retain the tombstone until a host confirms cleanup").not.toBeNull();
      expect(remaining?.storageDeletedAt).toBeNull();
      expect(commands).toHaveLength(1);
    });
  } finally {
    await rm(storageRoot, { recursive: true, force: true });
  }
});

5. Root cause

Pruning selects old destroyed environments whose provider teardown is complete; it does not preserve them for surviving thread storage. Deleting the environment clears the thread reference through the database foreign key:

packages/db/src/data/sweeps.ts:739–750

  const staleEnvironmentIds = db
    .select({ id: environments.id })
    .from(environments)
    .where(
      and(
        eq(environments.status, "destroyed"),
        sql`(${environments.environmentProviderId} is null or ${environments.teardownStatus} = 'removed')`,
        lt(environments.updatedAt, args.updatedBefore),
      ),
    )
    .orderBy(asc(environments.updatedAt), asc(environments.id))
    .limit(args.limit)

packages/db/src/data/sweeps.ts:793–797

        const deleteResult = tx
          .delete(environments)
          .where(eq(environments.id, environmentId))
          .run();
        return { deleted: deleteResult.changes, detachedEvents: 0 };

packages/db/src/schema.ts:584–586

    environmentId: text("environment_id").references(() => environments.id, {
      onDelete: "set null",
    }),

The DELETE route marks the thread deleted and calls storage cleanup using that now-null environment reference:

apps/server/src/routes/threads/base.ts:464–490

  del(routes.delete, async (context, payload) => {
    const thread = requirePublicThread(deps.db, context.req.param("id"));
    requireChildThreadsConfirmation({
      action: "delete",
      confirmed: payload.childThreadsConfirmed,
      deps,
      thread,
    });
    const dependents = listLifecycleThreadTree(deps.db, thread.id);
    markThreadDeleted(deps.db, deps.hub, { threadId: thread.id });
    for (const dependent of dependents) {
      const deleted = getThread(deps.db, dependent.id);
      if (!deleted) continue;
      emitPluginThreadDeleted(deleted);
      cancelAbandonedProviderCreations(deps, deleted.id);
      deps.terminalSessions.closeDeletedThreadTerminals({
        threadId: deleted.id,
      });
      requestThreadStorageDeletion(
        deps,
        deleted,
        deleted.environmentId === null
          ? null
          : getEnvironment(deps.db, deleted.environmentId),
      );
    }
    return context.json({ ok: true });

The null-reference branch marks storage deleted and finalizes immediately. It cannot tell a thread that never had storage from one whose environment was pruned:

apps/server/src/services/threads/thread-lifecycle.ts:1083–1115

export function requestThreadStorageDeletion(
  deps: CommandResultSideEffectsDeps,
  thread: Pick<Thread, "environmentId" | "id">,
  environment: { hostId: string; id: string } | null,
): void {
  deps.pendingInteractions.interruptPendingInteractionsForThreadIds({
    threadIds: [thread.id],
    reason: "thread-deleted",
  });
  abortPluginToolCallsForThreads([thread.id], "thread-deleted");
  if (thread.environmentId === null) {
    markThreadStorageDeleted(deps.db, { threadId: thread.id });
    finalizeStoppedThread(deps, { threadId: thread.id });
    return;
  }
  if (environment === null) {
    deps.logger.warn(
      { environmentId: thread.environmentId, threadId: thread.id },
      "Thread storage deletion environment is unavailable",
    );
    return;
  }
  if (!inFlightThreadRpcGuard.claim(thread.id, "thread.storage.delete")) {
    return;
  }
  void runLiveHostCommand(deps, {
    command: buildThreadStorageDeleteCommand({
      environmentId: environment.id,
      threadId: thread.id,
    }),
    hostId: environment.hostId,
    timeoutMs: LIVE_DAEMON_COMMAND_TIMEOUT_MS,
  })

Finalization only waits for a storage acknowledgement when the environment reference is non-null, permitting hard deletion in this case:

apps/server/src/services/threads/thread-lifecycle.ts:1962–1971

    clearThreadProvisionSchedule(finalizedThread.id);
    if (
      finalizedThread.environmentId !== null &&
      finalizedThread.storageDeletedAt === null
    )
      return;
    if (providerEnvironmentHasPendingWork(deps.db, finalizedThread.id)) return;
    deleteThread(deps.db, deps.hub, finalizedThread.id);
    if (finalizedThread.environmentId !== null)
      refreshProviderRetirement(deps, finalizedThread.environmentId);

The host-side operation itself needs the thread ID and its host-local storage root:

apps/host-daemon/src/command-handlers/thread.ts:54–64

export async function deleteThreadStorage(
  command: CommandOf<"thread.storage.delete">,
  options: CommandDispatchOptions,
): Promise<void> {
  const storagePath = requireContainedPath(
    options.threadStorageRootPath,
    path.join(options.threadStorageRootPath, command.threadId),
    "Thread storage path escapes the storage root",
  );
  await fs.rm(storagePath, { recursive: true, force: true });
}

Once the thread row is gone, the normal deleted-thread lifecycle sweep cannot retry it. The sweep selects existing deleted rows.

6. Proposed fix and simple-fix assessment

Persist storage ownership independently of the environment lifetime, retain a deletion tombstone until the owning host confirms cleanup, and retry offline hosts. The already-orphaned population requires a separate recovery policy because the authoritative relationship has been lost. An alternative is retaining environments until dependent storage is reclaimed, but that changes retention behavior and cannot restore ownership already discarded.

No PR opened. A complete correction needs a stored-data/ownership or retention design decision, outside this automation's simple-fix constraints. No production code was changed or pushed. No open linked PR appeared in issue timeline metadata or the open-PR search. Reporting directory sizes is a separate feature and was not implemented.

7. Verification

The same agent created two detached clean worktrees at the recorded commit, named first and second under a unique temporary root. Each received only the new test, its own frozen dependency install, and its own build. The focused Turbo command above executed freshly in each checkout; both logged a test cache miss and distinct runner process IDs. Both produced one passing control and one failing pruned case, with the exact observations above. Each harness created fresh database and temporary filesystem state; no ports were bound.

An initial test draft omitted the required DELETE JSON body and received HTTP 400 in both checkouts. That harness error was corrected using the repository's request contract before collecting the final evidence. No root-cause claim relies on those initial failures.

8. Related issues

#1924 addressed archiving after environment pruning. Its existing archive test covers a different operation; allowing archive does not restore storage cleanup ownership. This report does not assess the separate pruning-performance issue.

9. Appendix

The adjacent archive and thread-stop retry suites passed all 9 tests across 2 files. Both builds completed 60 tasks. All source links were checked against the recorded commit. Production files remained unchanged. Issue content was treated as untrusted claims; no external issue link, proposed patch, or issue-provided command was executed.

Preparation used git fetch origin main and two git worktree add --detach calls at the recorded SHA. The origin's old repository name resolves through GitHub to public get-bb/bb. Verification commands were the frozen install, Turbo build, focused test above, and pnpm exec turbo run test --filter=@bb/server -- test/threads/archive-pruned-environment.test.ts test/threads/thread-stop-retry.test.ts. Repository reads used git, rg, and GitHub metadata queries. No user runtime or secrets were accessed.

AGENT GENERATED