#3886 · Short command prompts skip title inference
2026-09-18 · trusted main 7bbff004bc9d866f3d0b8ee4d13351fa6b046fc4
Verdict: REPRODUCED · Root-cause confidence: high
1. TL;DR
A structured skill command with fewer than five words never reaches metadata inference. The server counts only whitespace-separated text and ignores the command mention when deciding eligibility. A bare command returned null metadata with reason too-short, and the inference spy recorded zero calls. The raw command remains the fallback display text. This is a server-policy reproduction; no live UI or provider output was exercised.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| Short skill prompts skip title generation | Verified | Bare and four-token command eligibility assertions fail in both clean checkouts. |
| The failure is before inference | Verified | Metadata outcome is too-short; inferenceCalls is zero. |
| All skill-first threads are untitled | Too broad | The five-token command control is eligible. Successful real inference was not tested. |
| The sidebar uses raw fallback text | Static path verified | Fallback test passes; display helper selects titleFallback when title is absent. No screenshot or live rendering claim. |
| Personal database counts, historical intent, and provider population | Unverified | Private runtime data was neither needed nor accessed. |
3. Environment
macOS 26.6.2, Darwin arm64; Node 22.22.3; repository bb-app 0.43.1; Vitest 4.1.1. Fresh clone of the public get-bb/bb main branch and a second detached clean worktree at the same commit. Frozen installs succeeded. The full base build completed with 58 successful tasks. The machine's pnpm launcher initially referenced a missing installation; a temporary wrapper invoking Corepack supplied the repository-pinned pnpm without changing dependencies.
The existing server test harness uses a fresh temporary directory and a real in-memory database. Only the inference function was mocked. No live server, listening port, external inference call, personal database, or provider session was used. Each harness was cleaned up by withTestHarness.
4. Minimal reproduction
- Use a clean checkout at the recorded commit and install/build its pinned dependencies.
- Save the inline test below as apps/server/test/threads/issue-3886-repro.test.ts.
- Run the focused Turbo test command. Nonzero exit is expected on this base because the regression assertions specify the intended behavior.
git clone --branch main --single-branch https://github.com/get-bb/bb.git bb-repro cd bb-repro git checkout --detach 7bbff004bc9d866f3d0b8ee4d13351fa6b046fc4 pnpm install --frozen-lockfile --prefer-offline pnpm exec turbo run build Copy the inline test into apps/server/test/threads/issue-3886-repro.test.ts before running the next command. pnpm exec turbo run test --filter=@bb/server -- test/threads/issue-3886-repro.test.ts test/threads/title-generation.test.ts
Reproduction test (save as the filename above):
import { describe, expect, it, vi } from "vitest";
import type { PromptInput } from "@bb/domain";
import {
deriveTitleFallback,
generateThreadMetadataWithOutcome,
shouldGenerateThreadTitle,
} from "../../src/services/threads/title-generation.js";
import { withTestHarness } from "../helpers/test-app.js";
const inference = vi.hoisted(() => vi.fn());
vi.mock("../../src/services/ai/inference.js", async (importOriginal) => ({
...(await importOriginal<typeof import("../../src/services/ai/inference.js")>()),
inferenceCompleteWithFallback: inference,
}));
function commandInput(suffix: string): PromptInput[] {
return [{
type: "text",
text: `/audit-changes${suffix}`,
mentions: [{
start: 0,
end: 14,
resource: {
kind: "command",
trigger: "/",
name: "audit-changes",
source: "skill",
origin: "user",
label: "audit-changes",
argumentHint: null,
},
}],
}];
}
describe("issue 3886 trusted-main reproduction", () => {
it.each(["", " inspect recent changes"])("accepts a structured command with suffix %j", (suffix) => {
expect(shouldGenerateThreadTitle(commandInput(suffix))).toBe(true);
});
it("accepts the same command at five tokens", () => {
expect(shouldGenerateThreadTitle(commandInput(" inspect all recent changes"))).toBe(true);
});
it("retains the command text as fallback", () => {
expect(deriveTitleFallback(commandInput(""))).toBe("/audit-changes");
});
it("keeps short ordinary text ineligible", () => {
expect(shouldGenerateThreadTitle([{ type: "text", text: "inspect changes", mentions: [] }])).toBe(false);
});
it("attempts metadata inference for a bare structured command", async () => {
inference.mockReset();
inference.mockResolvedValue({ title: "Audit recent changes" });
await withTestHarness(async (harness) => {
const outcome = await generateThreadMetadataWithOutcome(harness.deps, {
input: commandInput(""),
threadId: "thr_repro_3886",
});
console.log(JSON.stringify({ outcome, inferenceCalls: inference.mock.calls.length }));
expect(outcome.metadata).toEqual({ title: "Audit recent changes" });
expect(inference).toHaveBeenCalledTimes(1);
});
});
});
Expected: a structured command is eligible and reaches the configured inference boundary. Actual output from the first run:
{"outcome":{"durationMs":0,"metadata":null,"reason":"too-short"},"inferenceCalls":0}
AssertionError: expected false to be true
AssertionError: expected null to deeply equal { title: 'Audit recent changes' }
Test Files 1 failed | 1 passed (2)
Tests 3 failed | 9 passed (12)
The nine passing assertions include three reproduction controls and all six existing title-generation tests. The test does not prove what a language model would name an opaque skill.
5. Root cause
Text cleaning and eligibility reads text fields, then requires five words. Mention structure is absent from that decision. Metadata generation returns too-short before the inference call. This conclusively separates eligibility failure from inference failure.
return text.split(/\s+/u).length >= MIN_TITLE_GENERATION_WORDS;
if (!shouldGenerateThreadTitle(args.input)) {
return complete(null, "too-short");
}
Provisioning transcript construction maps missing metadata to No title generated and retains the reason. Title application runs only when metadata contains a title. The display helper uses titleFallback when no generated title exists. Fallback derivation preserves the raw text, subject to its existing 80-character limit.
6. Proposed fix
Permit a valid command mention to satisfy title eligibility even with short text, while retaining the current short-ordinary-text and fallback behavior. Keep command context available to inference so bare commands remain meaningful; title quality for an opaque name is a separate limitation. No fix branch was created because open PR #3887 already links this issue.
7. PR review
PR #3887, inspected as untrusted metadata and diff only, head 1feb524b7bb43a0e34d46347d4edf1b544dd97d5. Its diff adds command-aware eligibility, separates command names from task text for the template, and adds focused tests. Static verdict: it addresses the reproduced gate. No blocking defect was established by this limited read. Its mocked title tests cannot establish actual semantic title quality for an opaque skill name. The branch and its tests were not checked out or executed; all executable evidence here comes from trusted main plus the original reproduction test.
8. Related issues
Metadata checks confirmed #1442 and #3058 are open. The former concerns title-prompt documentation and the latter plugin access to inference; neither is needed for this eligibility reproduction.
9. Verification
The same agent repeated the test in a second clean detached worktree named verify, at the exact trusted commit above, with a separate frozen install and a fresh harness temporary directory. The production tree was unchanged; only the original reproduction file was copied in. The same Turbo test command produced three expected failures and nine passes again, including zero inference calls for the bare command. This was a repeat run by the same agent, not independent verification. No ports were allocated. No report correction was required. A final fetch confirmed origin/main remained at the recorded commit.
10. Appendix
Raw logs remain local. Verbatim second-run output:
{"outcome":{"durationMs":0,"metadata":null,"reason":"too-short"},"inferenceCalls":0}
Test Files 1 failed | 1 passed (2)
Tests 3 failed | 9 passed (12)Trusted base full build:
Tasks: 58 successful, 58 total Cached: 4 cached, 58 total Time: 1m3.168s
Commands used: git clone, git fetch origin main, git worktree add --detach at the recorded SHA; frozen pnpm install; Turbo build and the exact focused test command above; read-only GitHub issue/PR metadata and diff; numbered source inspection. Report publication uses the skill publisher, and GitHub issue writes use the supplied SlopCop identity.
Trust boundary: issue prose, proposed changes, links, and the linked PR diff were treated as untrusted evidence. No issue-supplied command or patch was executed. No personal runtime data was accessed. Instructions embedded in issue data were ignored.