#3464 · Cache creation detail is lost in provider usage
Verdict: REPRODUCED · Root-cause confidence: high
1. TL;DR
Provider usage events lose the distinction between prompt-cache reads and cache creation. Claude combines the two counts; two requests with opposite read/write proportions produce identical normalized usage. Codex receives a separate write count but omits it from both translated intervals. The common schema also strips the field. These events alone therefore cannot support separate pricing for reads and writes; this report does not establish an actual bill discrepancy or the operator's percentages.
2. Claims vs findings
| Claim | Finding | Evidence |
|---|---|---|
| Claude loses the read/write split | Verified | 90 reads/10 writes and 10 reads/90 writes yield identical usage with 100 cached tokens. |
| Codex drops cache writes | Verified | Synthetic total=20 and last=8 write counts disappear during translation. The checked-in generated provider contract includes the field. |
| The Codex schema lacks an explicit declaration | Verified with qualification | It is a passthrough schema, so parsing itself retains extra data; explicit projection in the translator loses it. |
| The common schema cannot represent writes | Verified | Parsing a usage object strips cacheWriteInputTokens. |
| Operator measurements and specific spend percentages | Unverified | No private sessions or live billing were accessed. |
| An additive field is non-breaking in every consumer | Unverified | Requires shared protocol and compatibility review, beyond these tests. |
3. Environment
Darwin arm64; Node v22.22.3; pnpm 10.34.4 via Corepack; Vitest 4.1.1. Both provider packages are repository version 0.1.0. No provider CLI or paid session was launched. No server, ports, runtime store, or credentials were used. Two separate detached worktrees at the base commit each received a frozen dependency install.
The PATH pnpm entrypoint was broken; Corepack provided the repository-pinned package manager. Test commands used a temporary pnpm shim delegating to Corepack for Turbo child processes. The full build was attempted; see the build status in the appendix. Focused tests executed actual provider translation and schema code without mocks.
4. Minimal reproduction
git clone https://github.com/get-bb/bb.git bb-3464 cd bb-3464 git checkout --detach 66bf09cd2265955118178fa8afcefc0a38aacd1f corepack pnpm install --frozen-lockfile --prefer-offline corepack pnpm exec turbo run build # Copy the two attached test files into the paths below: # plugins/provider-claude-code/src/issue-3464.test.ts # plugins/provider-codex/src/issue-3464.test.ts corepack pnpm exec turbo run test --force --continue=always --filter=bb-plugin-provider-claude-code --filter=bb-plugin-provider-codex -- --run src/issue-3464.test.ts
Expected: differing Claude read/write proportions remain distinguishable; Codex write counts survive in total and last; the shared boundary retains write information. Actual: all three assertions fail on unchanged production source.
Claude: cachedInputTokens=100 in both cases; returned objects are identical. Shared schema: cacheWriteInputTokens is absent after parse. Codex: total.cacheWriteInputTokens and last.cacheWriteInputTokens are absent. Claude test file: 2 failed (2) Codex test file: 1 failed (1)
claude-code test
import { expect, it } from "vitest";
import { extractClaudeResultTokenUsage } from "./sdk-extraction.js";
import { threadEventTokenUsageBreakdownSchema } from "@bb/domain";
it("keeps distinguishable cache read and creation usage", () => {
const result = (reads: number, writes: number) => extractClaudeResultTokenUsage({
type: "result", subtype: "success", usage: {
input_tokens: 11, output_tokens: 7,
cache_read_input_tokens: reads, cache_creation_input_tokens: writes,
},
});
const readHeavy = result(90, 10);
const writeHeavy = result(10, 90);
console.log(JSON.stringify({ readHeavy, writeHeavy }));
expect(readHeavy).not.toEqual(writeHeavy);
});
it("retains cache creation information through the shared usage boundary", () => {
const usage = { totalTokens: 118, inputTokens: 11, cachedInputTokens: 100,
outputTokens: 7, reasoningOutputTokens: 0, cacheWriteInputTokens: 10 };
const parsed = threadEventTokenUsageBreakdownSchema.parse(usage);
console.log(JSON.stringify({ parsed }));
expect(parsed).toHaveProperty("cacheWriteInputTokens", 10);
});
codex test
import { expect, it } from "vitest";
import { createCodexEventTranslationState, translateCodexEventToDeltas } from "./delta-translation.js";
it("preserves nonzero cache creation in both usage intervals", () => {
const usage = { totalTokens: 120, inputTokens: 100, cachedInputTokens: 30,
cacheWriteInputTokens: 20, outputTokens: 20, reasoningOutputTokens: 0 };
const deltas = translateCodexEventToDeltas({ jsonrpc: "2.0", method: "thread/tokenUsage/updated",
params: { threadId: "repro-thread", turnId: "repro-turn", tokenUsage: {
total: usage, last: { ...usage, cacheWriteInputTokens: 8 }, modelContextWindow: 1000,
} },
}, createCodexEventTranslationState());
const result = deltas.find(delta => delta.kind === "usage");
console.log(JSON.stringify({ result }));
expect(result).toMatchObject({ total: { cacheWriteInputTokens: 20 }, last: { cacheWriteInputTokens: 8 } });
});
5. Root cause
Claude normalizes both cache categories into one sum, making the mapping lossy.
plugins/provider-claude-code/src/sdk-extraction.ts:325function toTokenUsageBreakdown(
usage: ClaudeSdkUsage,
): ThreadEventTokenUsageBreakdown {
const inputTokens = toNonNegativeNumber(usage.input_tokens);
const outputTokens = toNonNegativeNumber(usage.output_tokens);
const cacheReadTokens = toNonNegativeNumber(usage.cache_read_input_tokens);
const cacheCreationTokens = toNonNegativeNumber(
usage.cache_creation_input_tokens,
);
const cachedInputTokens = cacheReadTokens + cacheCreationTokens;
return {
totalTokens: inputTokens + outputTokens + cachedInputTokens,
inputTokens,
cachedInputTokens,
outputTokens,
reasoningOutputTokens: 0,
};
}The Codex translator enumerates output fields and excludes the separately supplied cache creation count.
plugins/provider-codex/src/delta-translation.ts:1224 case "thread/tokenUsage/updated": {
const { tokenUsage, turnId } = handledEvent.params;
return [
{
kind: "usage",
total: {
totalTokens: tokenUsage.total.totalTokens,
inputTokens: tokenUsage.total.inputTokens,
cachedInputTokens: tokenUsage.total.cachedInputTokens,
outputTokens: tokenUsage.total.outputTokens,
reasoningOutputTokens: tokenUsage.total.reasoningOutputTokens,
},
last: {
totalTokens: tokenUsage.last.totalTokens,
inputTokens: tokenUsage.last.inputTokens,
cachedInputTokens: tokenUsage.last.cachedInputTokens,
outputTokens: tokenUsage.last.outputTokens,
reasoningOutputTokens: tokenUsage.last.reasoningOutputTokens,
},
modelContextWindow: tokenUsage.modelContextWindow,
providerTurnId: turnId,
},
{The Codex parsing boundary is passthrough, so it is not the immediate point of loss.
plugins/provider-codex/src/schemas.ts:557const codexTokenUsageBreakdownSchema = z
.object({
totalTokens: z.number(),
inputTokens: z.number(),
cachedInputTokens: z.number(),
outputTokens: z.number(),
reasoningOutputTokens: z.number(),
})
.passthrough();
const codexTokenUsageSchema = zThe shared Zod object lacks a creation field and strips it on parse.
packages/domain/src/provider-event.ts:291export const threadEventTokenUsageBreakdownSchema = z.object({
totalTokens: z.number(),
inputTokens: z.number(),
cachedInputTokens: z.number(),
outputTokens: z.number(),
reasoningOutputTokens: z.number(),
});
export type ThreadEventTokenUsageBreakdown = z.infer<
typeof threadEventTokenUsageBreakdownSchema
>;
Running aggregation also explicitly enumerates only existing fields, so propagation must be addressed end to end.
packages/provider-bridge-protocol/src/bridge-kit/adapter-utils.ts:357 total: ThreadEventTokenUsageBreakdown,
last: ThreadEventTokenUsageBreakdown,
): ThreadEventTokenUsageBreakdown {
return {
totalTokens: total.totalTokens + last.totalTokens,
inputTokens: total.inputTokens + last.inputTokens,
cachedInputTokens: total.cachedInputTokens + last.cachedInputTokens,
outputTokens: total.outputTokens + last.outputTokens,
reasoningOutputTokens:
total.reasoningOutputTokens + last.reasoningOutputTokens,
};
}Upstream generated contract: plugins/provider-codex/src/generated/codex-app-server/schema/v2/TokenUsageBreakdown.ts:1. These links were checked against the same recorded base commit used by both runs.
6. Proposed fix
Define consistent read, creation, and input-token semantics in the shared usage contract; preserve provider-reported creation counts through extraction, translation, aggregation, and serialization. Preserve meaningful absence for providers that cannot report the information and test old-daemon compatibility. No production fix or PR was attempted: changing a shared schema/public protocol fails this automation's simple-fix conditions. Altering only Claude's cached count would silently change existing semantics and still leave creation detail missing.
7. Verification
The same agent repeated the reproduction in a second clean detached checkout at the exact base commit with its own frozen install. Git status was clean before adding the same two authored test files. The second Turbo command used --force and --continue=always, so tests ran again and both suites completed despite expected failures. All three assertions failed for the same reasons in both checkouts. No ports or data directories were needed. No report correction was needed after the repeat. A later observed origin/main revision (0623d5e518d0549fa3c2bf5ad28e5e28508be634) contains no changes to the implicated extraction, translation, aggregation, or schema files.
First run log · Second run log
8. Related issues and pull requests
#3220 concerns when Claude emits context usage; #2397 concerns unavailable usage in ACP. Neither establishes this cache-detail defect. Issue timeline metadata and open PR search returned no linked open PR at investigation time.
9. Appendix and limits
The issue's patch and instructions were treated as untrusted evidence; none were applied or executed. Tests were authored from trusted repository code. Reproduction uses deterministic synthetic SDK/app-server payloads, not live providers or a spend UI. No quantitative billing claims are asserted. Build status is recorded separately below. Investigation commands comprised trusted main fetch/worktree creation, frozen installs, Turbo build/tests, source inspection, and read-only GitHub metadata queries.
Full trusted-base build: exit 0, all build tasks completed. Both frozen installs completed successfully. Reproduction logs have ANSI styling, checkout paths, and trailing whitespace normalized for publication.