#4865 · Skill directory publication rejects modeled access-denied renames
Trusted main: 09ff18a73bd0a8037f43c6ed86c28eb02228089c.
PARTIALLY REPRODUCED · Root-cause confidence: medium
1. TL;DR
A temporary directory containing a required skill is fully written before the host publishes it with a directory rename. Both the downloaded-tree cache and the runtime skill catalog make just one publication attempt and propagate an EPERM result instead of retrying or recognizing a valid concurrent destination. Four fault-injection cases reproduce that rejection on trusted main; a subsequent submission succeeds with intact skill bytes. The same results occur in a second clean checkout. This is a Linux test of the application’s response to a modeled Windows filesystem result, not a Windows or Defender reproduction, so the verdict is partial and no fix PR is opened.
2. Claims vs findings
| Claim | Finding | Evidence and limit |
|---|---|---|
| A publication access-denied result aborts required-skill installation. | Verified at the staging boundary | Both real publication paths reject the injected EPERM after exactly one attempt. |
| Resubmitting after the denial clears succeeds. | Verified in the model | The next call publishes the expected bytes in all four cases. |
| The temporary completed directory is removed on failure. | Verified in the model | No .tmp- directories remain in the affected publication root after the rejected call. |
| A destination collision returning access denied is not recognized. | Verified in the model | Copying a complete candidate to the final destination before returning EPERM still rejects the first staging call. |
| Windows produces the reported code while Defender scans fresh files. | Unverified | No Windows kernel, Defender, desktop installation, or scanner handle is available in this environment. The test injects the error, not the OS condition. |
| Windows maps existing-directory collisions to this error. | Unverified natively | The model only verifies how the application responds if it receives that error with a complete destination already present. |
| Related PR #4856 also omits access-denied handling. | Verified statically only | The read-only diff recognizes EEXIST/ENOTEMPTY and does not change catalog publication. Its code was not checked out or executed. |
3. Environment
- Public repository:
get-bb/bb; base commit recorded above. - Linux; Node
v22.19.0; pnpm9.15.0; Vitest4.1.1. - Frozen dependency install and full Turbo build pass in both checkouts: 63 build tasks successful in each run.
- No provider, real turn, development server, user runtime directory, browser, or network-facing plugin is involved.
- Every case creates a distinct temporary
bb-4865-probe-*data directory and removes it in afinallyblock.
4. Minimal reproduction
This reproduces the application-side failure using a narrowly scoped filesystem spy. It does not reproduce a Windows handle lock. The actual hashing, validation, file writes, publication caller, cleanup, and resubmission all execute trusted-main code.
- Prepare trusted main in a fresh checkout:
git clone https://github.com/get-bb/bb.git bb-4865-repro cd bb-4865-repro git checkout --detach 09ff18a73bd0a8037f43c6ed86c28eb02228089c pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build
- Save the linked reproduction test as
apps/host-daemon/src/issue-4865-repro.test.ts. - Run it through the repository’s normal test orchestration:
pnpm exec turbo run test --filter=@bb/host-daemon -- src/issue-4865-repro.test.ts
The test intercepts only publication renames under skill-store or global-skills. A transient case returns one access-denied error and allows later native renames. A collision case creates an intact destination using real filesystem copying before returning that same error, representing a concurrent complete installer. This is an error-result model, not an actual concurrent Windows process.
Expected: if that filesystem result is supported, the first staging operation should retry a transient denial or accept a validated complete destination, preserving the requested skill bytes. Permanent denials must still fail.
Actual (verbatim structured observations, identical in both runs):
{"root":"skill-store","failure":"transient","first":"rejected: EPERM","firstRenameAttempts":1,"remainingTemps":0,"resubmission":"published intact bytes"}
{"root":"global-skills","failure":"transient","first":"rejected: EPERM","firstRenameAttempts":1,"remainingTemps":0,"resubmission":"published intact bytes"}
{"root":"skill-store","failure":"collision","first":"rejected: EPERM","firstRenameAttempts":1,"remainingTemps":0,"resubmission":"published intact bytes"}
{"root":"global-skills","failure":"collision","first":"rejected: EPERM","firstRenameAttempts":1,"remainingTemps":0,"resubmission":"published intact bytes"}
AssertionError: expected 'rejected: EPERM' to be 'fulfilled'
Test Files 1 failed (1)
Tests 4 failed (4)
The failures are intentional regression assertions. Before asserting that the first call should fulfill, every case successfully stages again and checks the actual published text against an independently specified fixture.
Complete test
import { createHash } from "node:crypto";
import fs from "node:fs/promises";
import { tmpdir } from "node:os";
import path from "node:path";
import { expect, it, vi } from "vitest";
import type {
HostDaemonInjectedSkillSource,
HostDaemonSkillTree,
} from "@bb/host-daemon-contract";
import { stageInjectedSkillSources } from "./injected-skills.js";
const skillText =
"---\nname: rename-probe\ndescription: Rename probe.\n---\n\nprobe bytes\n";
const skillBytes = Buffer.from(skillText);
const treeHash = createHash("sha256")
.update("bb-skill-tree-v1\0file\0SKILL.md\0")
.update("644\0")
.update(String(skillBytes.length))
.update("\0")
.update(skillBytes)
.digest("hex");
const tree: HostDaemonSkillTree = {
treeHash,
entries: [
{
path: "SKILL.md",
mode: 0o644,
contentBase64: skillBytes.toString("base64"),
},
],
};
it.each([
["skill-store", "transient"],
["global-skills", "transient"],
["skill-store", "collision"],
["global-skills", "collision"],
] as const)(
"publishes %s despite a modeled Windows %s rename denial",
async (root, failure) => {
const dataDir = await fs.mkdtemp(path.join(tmpdir(), "bb-4865-probe-"));
const nativeRename = fs.rename;
const denial = Object.assign(
new Error("Modeled directory publication denial"),
{ code: "EPERM" },
);
let denials = 0;
let renameAttempts = 0;
const rename = vi
.spyOn(fs, "rename")
.mockImplementation(async (from, to) => {
if (path.basename(path.dirname(String(from))) === root) {
renameAttempts += 1;
if (denials === 0) {
denials += 1;
if (failure === "collision") {
await fs.cp(from, to, { recursive: true });
}
throw denial;
}
}
await nativeRename(from, to);
});
try {
const sourceRootPath = path.join(dataDir, "source");
await fs.mkdir(sourceRootPath);
await fs.writeFile(path.join(sourceRootPath, "SKILL.md"), skillBytes);
const source: HostDaemonInjectedSkillSource =
root === "skill-store"
? {
kind: "tree",
sourceType: "data-dir",
name: "rename-probe",
description: "Rename probe.",
treeHash,
entryPath: "SKILL.md",
}
: {
kind: "workspace-path",
sourceType: "project",
name: "rename-probe",
description: "Rename probe.",
sourceRootPath,
skillFilePath: path.join(sourceRootPath, "SKILL.md"),
};
const stage = () =>
stageInjectedSkillSources({
dataDir,
injectedSkillSources: [source],
fetchSkillTree: async () => tree,
logger: { warn: () => undefined },
});
const first = await stage().then(
() => "fulfilled",
(error: unknown) =>
error === denial ? "rejected: EPERM" : "rejected: unexpected error",
);
const firstRenameAttempts = renameAttempts;
const remainingTemps = (
await fs.readdir(path.join(dataDir, "runtime", root))
).filter((name) => name.startsWith(".tmp-"));
const second = await stage();
const publishedText = await fs.readFile(
path.join(second.skillRoots[0]!.path, "rename-probe", "SKILL.md"),
"utf8",
);
console.info(
JSON.stringify({
root,
failure,
first,
firstRenameAttempts,
remainingTemps: remainingTemps.length,
resubmission:
publishedText === skillText
? "published intact bytes"
: "wrong bytes",
}),
);
expect(publishedText).toBe(skillText);
expect(first).toBe("fulfilled");
} finally {
rename.mockRestore();
await fs.rm(dataDir, { recursive: true, force: true });
}
},
);
5. Root cause
Catalog publication and tree-cache publication each perform one rename after writing the complete candidate directory. Their catches only recognize two collision codes:
await fs.rename(tempRootPath, stageRootPath);
...
isFsErrorWithCode(error, "EEXIST") ||
isFsErrorWithCode(error, "ENOTEMPTY")
...
await fs.rm(tempRootPath, { recursive: true, force: true });
throw error;
The tree-cache branch uses the same error-code gate. Therefore the modeled access-denied result always takes the cleanup-and-rethrow branch, regardless of whether the destination is absent or an intact concurrent winner exists. The rejection is confirmed dynamically; whether Windows supplies that result for the reported scanner or collision scenarios remains unverified.
Required tree staging logs a pull failure and rethrows the error. Runtime skill configuration returns that staging operation rather than making the skill optional. These source links explain how the error can prevent runtime setup; this investigation does not execute a real turn.submit or prove the desktop error presentation.
Confidence is high in this application-side mechanism and medium in the overall reported root cause because the native trigger is not observed. A second caveat is that cleanup itself can fail when a real scanner still holds a file; the model leaves removal available and does not establish native cleanup behavior.
6. Proposed fix and simple-fix decision
First establish the native Windows conditions using an independently authored test with a separately held file handle and an actual complete-destination race. If confirmed, add a bounded, Windows-specific publication retry for transient sharing/access-denied results to both writers. Before accepting a concurrent destination, validate its expected completion artifacts; do not treat any existing path as success. Preserve non-Windows and non-transient failures, bound permanent-lock latency, and make temporary-directory cleanup robust to the same lock.
Required future checks include lock release during backoff, a permanent lock reaching a deadline, intact and incomplete collision destinations, unrelated filesystem errors, and catalog as well as cache publication. The four supplied cases provide repeatable application-side failing evidence, but cannot validate Windows sharing behavior or a backoff budget.
No fix PR: the native Windows/Defender failure was not reproduced in this Linux-only environment; shipping a Windows retry policy without that validation does not meet the conservative reproduced-bug gate. No production change, dependency, branch, commit, or fix push is made. No currently open PR links to #4865 in its issue timeline or the open-PR search.
7. Related PR review — #4856
#4856 is an open, related fix for incomplete skill-cache collisions, not a fix linked to #4865. Only its metadata and diff were read as untrusted data.
- It validates cache completion, repairs incomplete destinations, and retries after recognized collisions.
- Static finding: in
apps/host-daemon/src/injected-skills.ts, the new inner rename catch still excludes access-denied errors. Conditional impact: if Windows returns such an error, the incomplete-cache repair branch is not reached. - It does not modify the catalog writer’s publication operation.
- Verdict: does not statically cover the modeled access-denied failure. No runtime verdict on that PR is claimed; none of its tests or code were run.
8. Verification
The same agent repeated the test in a second clean, detached worktree at the exact recorded base commit. Its tracked files were clean before installation; a new frozen install and full Turbo build succeeded. Only the independently authored reproduction test was then added. Each test created its own fresh temporary data directory; no ports or persistent runtime data were used.
pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build pnpm exec turbo run test --filter=@bb/host-daemon -- src/issue-4865-repro.test.ts Build: 63 successful tasks. Reproduction: exit 1; four intended failures. All four structured observations match the first run.
The baseline staging suite also passes in the first checkout: pnpm exec turbo run test --filter=@bb/host-daemon -- src/injected-skills.test.ts reports 15 tests passed. No report correction was necessary after the second run. This is repeat verification by the same agent, not independent verification. The final verdict is intentionally partial.
9. Related issues
#4855 concerns incomplete tree-cache destinations. That differs from the modeled temporary denial and complete-destination collision investigated here; #4856 addresses the former. These related identifiers were inspected only through GitHub metadata and a read-only diff.
10. Appendix and trust boundary
Saved evidence: first run, second clean run, baseline staging tests, and build results. Paths in logs are redacted; results are otherwise unchanged. This is nonvisual, so no screenshot is appropriate.
Issue content, comments, and related PR contents were treated as untrusted claims. No issue-provided script, patch, command, binary, branch, or linked external URL was run or fetched. The reproduction test was authored from trusted repository evidence. No production BB instance or credential was accessed.
Investigation commands: read issue properties/comments/timeline and repository field/label definitions with GitHub API; inspect similar issue classifications; fetch trusted main; inspect the staging implementation and its callers; read related PR metadata/diff without checkout; create a detached verification worktree; clone the report repository; frozen-install and full-build both checkouts; format the authored test; run the focused reproduction in both checkouts; run baseline staging tests; confirm no tracked production diff. GitHub classification and verdict writes use the SlopCop wrapper. Publication uses the report skill publisher without its issue-comment or issue-label flags.
> AGENT GENERATED