#3464 · Cache creation detail is lost in provider usage

Bug · Priority: Medium · Effort: High · providers · provider-claude-code · provider-codex
2026-09-11 · Base: 66bf09cd2265955118178fa8afcefc0a38aacd1f · Issue

Verdict: REPRODUCED · Root-cause confidence: high

1. TL;DR

Provider usage events lose the distinction between prompt-cache reads and cache creation. Claude combines the two counts; two requests with opposite read/write proportions produce identical normalized usage. Codex receives a separate write count but omits it from both translated intervals. The common schema also strips the field. These events alone therefore cannot support separate pricing for reads and writes; this report does not establish an actual bill discrepancy or the operator's percentages.

2. Claims vs findings

ClaimFindingEvidence
Claude loses the read/write splitVerified90 reads/10 writes and 10 reads/90 writes yield identical usage with 100 cached tokens.
Codex drops cache writesVerifiedSynthetic total=20 and last=8 write counts disappear during translation. The checked-in generated provider contract includes the field.
The Codex schema lacks an explicit declarationVerified with qualificationIt is a passthrough schema, so parsing itself retains extra data; explicit projection in the translator loses it.
The common schema cannot represent writesVerifiedParsing a usage object strips cacheWriteInputTokens.
Operator measurements and specific spend percentagesUnverifiedNo private sessions or live billing were accessed.
An additive field is non-breaking in every consumerUnverifiedRequires shared protocol and compatibility review, beyond these tests.

3. Environment

Darwin arm64; Node v22.22.3; pnpm 10.34.4 via Corepack; Vitest 4.1.1. Both provider packages are repository version 0.1.0. No provider CLI or paid session was launched. No server, ports, runtime store, or credentials were used. Two separate detached worktrees at the base commit each received a frozen dependency install.

The PATH pnpm entrypoint was broken; Corepack provided the repository-pinned package manager. Test commands used a temporary pnpm shim delegating to Corepack for Turbo child processes. The full build was attempted; see the build status in the appendix. Focused tests executed actual provider translation and schema code without mocks.

4. Minimal reproduction

git clone https://github.com/get-bb/bb.git bb-3464
cd bb-3464
git checkout --detach 66bf09cd2265955118178fa8afcefc0a38aacd1f
corepack pnpm install --frozen-lockfile --prefer-offline
corepack pnpm exec turbo run build
# Copy the two attached test files into the paths below:
# plugins/provider-claude-code/src/issue-3464.test.ts
# plugins/provider-codex/src/issue-3464.test.ts
corepack pnpm exec turbo run test --force --continue=always --filter=bb-plugin-provider-claude-code --filter=bb-plugin-provider-codex -- --run src/issue-3464.test.ts

Expected: differing Claude read/write proportions remain distinguishable; Codex write counts survive in total and last; the shared boundary retains write information. Actual: all three assertions fail on unchanged production source.

Claude: cachedInputTokens=100 in both cases; returned objects are identical.
Shared schema: cacheWriteInputTokens is absent after parse.
Codex: total.cacheWriteInputTokens and last.cacheWriteInputTokens are absent.
Claude test file: 2 failed (2)
Codex test file: 1 failed (1)

claude-code test

Download test

import { expect, it } from "vitest";
import { extractClaudeResultTokenUsage } from "./sdk-extraction.js";
import { threadEventTokenUsageBreakdownSchema } from "@bb/domain";

it("keeps distinguishable cache read and creation usage", () => {
  const result = (reads: number, writes: number) => extractClaudeResultTokenUsage({
    type: "result", subtype: "success", usage: {
      input_tokens: 11, output_tokens: 7,
      cache_read_input_tokens: reads, cache_creation_input_tokens: writes,
    },
  });
  const readHeavy = result(90, 10);
  const writeHeavy = result(10, 90);
  console.log(JSON.stringify({ readHeavy, writeHeavy }));
  expect(readHeavy).not.toEqual(writeHeavy);
});

it("retains cache creation information through the shared usage boundary", () => {
  const usage = { totalTokens: 118, inputTokens: 11, cachedInputTokens: 100,
    outputTokens: 7, reasoningOutputTokens: 0, cacheWriteInputTokens: 10 };
  const parsed = threadEventTokenUsageBreakdownSchema.parse(usage);
  console.log(JSON.stringify({ parsed }));
  expect(parsed).toHaveProperty("cacheWriteInputTokens", 10);
});

codex test

Download test

import { expect, it } from "vitest";
import { createCodexEventTranslationState, translateCodexEventToDeltas } from "./delta-translation.js";

it("preserves nonzero cache creation in both usage intervals", () => {
  const usage = { totalTokens: 120, inputTokens: 100, cachedInputTokens: 30,
    cacheWriteInputTokens: 20, outputTokens: 20, reasoningOutputTokens: 0 };
  const deltas = translateCodexEventToDeltas({ jsonrpc: "2.0", method: "thread/tokenUsage/updated",
    params: { threadId: "repro-thread", turnId: "repro-turn", tokenUsage: {
      total: usage, last: { ...usage, cacheWriteInputTokens: 8 }, modelContextWindow: 1000,
    } },
  }, createCodexEventTranslationState());
  const result = deltas.find(delta => delta.kind === "usage");
  console.log(JSON.stringify({ result }));
  expect(result).toMatchObject({ total: { cacheWriteInputTokens: 20 }, last: { cacheWriteInputTokens: 8 } });
});

5. Root cause

Claude normalizes both cache categories into one sum, making the mapping lossy.

plugins/provider-claude-code/src/sdk-extraction.ts:325
function toTokenUsageBreakdown(
  usage: ClaudeSdkUsage,
): ThreadEventTokenUsageBreakdown {
  const inputTokens = toNonNegativeNumber(usage.input_tokens);
  const outputTokens = toNonNegativeNumber(usage.output_tokens);
  const cacheReadTokens = toNonNegativeNumber(usage.cache_read_input_tokens);
  const cacheCreationTokens = toNonNegativeNumber(
    usage.cache_creation_input_tokens,
  );
  const cachedInputTokens = cacheReadTokens + cacheCreationTokens;

  return {
    totalTokens: inputTokens + outputTokens + cachedInputTokens,
    inputTokens,
    cachedInputTokens,
    outputTokens,
    reasoningOutputTokens: 0,
  };
}

The Codex translator enumerates output fields and excludes the separately supplied cache creation count.

plugins/provider-codex/src/delta-translation.ts:1224
    case "thread/tokenUsage/updated": {
      const { tokenUsage, turnId } = handledEvent.params;
      return [
        {
          kind: "usage",
          total: {
            totalTokens: tokenUsage.total.totalTokens,
            inputTokens: tokenUsage.total.inputTokens,
            cachedInputTokens: tokenUsage.total.cachedInputTokens,
            outputTokens: tokenUsage.total.outputTokens,
            reasoningOutputTokens: tokenUsage.total.reasoningOutputTokens,
          },
          last: {
            totalTokens: tokenUsage.last.totalTokens,
            inputTokens: tokenUsage.last.inputTokens,
            cachedInputTokens: tokenUsage.last.cachedInputTokens,
            outputTokens: tokenUsage.last.outputTokens,
            reasoningOutputTokens: tokenUsage.last.reasoningOutputTokens,
          },
          modelContextWindow: tokenUsage.modelContextWindow,
          providerTurnId: turnId,
        },
        {

The Codex parsing boundary is passthrough, so it is not the immediate point of loss.

plugins/provider-codex/src/schemas.ts:557
const codexTokenUsageBreakdownSchema = z
  .object({
    totalTokens: z.number(),
    inputTokens: z.number(),
    cachedInputTokens: z.number(),
    outputTokens: z.number(),
    reasoningOutputTokens: z.number(),
  })
  .passthrough();

const codexTokenUsageSchema = z

The shared Zod object lacks a creation field and strips it on parse.

packages/domain/src/provider-event.ts:291
export const threadEventTokenUsageBreakdownSchema = z.object({
  totalTokens: z.number(),
  inputTokens: z.number(),
  cachedInputTokens: z.number(),
  outputTokens: z.number(),
  reasoningOutputTokens: z.number(),
});
export type ThreadEventTokenUsageBreakdown = z.infer<
  typeof threadEventTokenUsageBreakdownSchema
>;

Running aggregation also explicitly enumerates only existing fields, so propagation must be addressed end to end.

packages/provider-bridge-protocol/src/bridge-kit/adapter-utils.ts:357
  total: ThreadEventTokenUsageBreakdown,
  last: ThreadEventTokenUsageBreakdown,
): ThreadEventTokenUsageBreakdown {
  return {
    totalTokens: total.totalTokens + last.totalTokens,
    inputTokens: total.inputTokens + last.inputTokens,
    cachedInputTokens: total.cachedInputTokens + last.cachedInputTokens,
    outputTokens: total.outputTokens + last.outputTokens,
    reasoningOutputTokens:
      total.reasoningOutputTokens + last.reasoningOutputTokens,
  };
}

Upstream generated contract: plugins/provider-codex/src/generated/codex-app-server/schema/v2/TokenUsageBreakdown.ts:1. These links were checked against the same recorded base commit used by both runs.

6. Proposed fix

Define consistent read, creation, and input-token semantics in the shared usage contract; preserve provider-reported creation counts through extraction, translation, aggregation, and serialization. Preserve meaningful absence for providers that cannot report the information and test old-daemon compatibility. No production fix or PR was attempted: changing a shared schema/public protocol fails this automation's simple-fix conditions. Altering only Claude's cached count would silently change existing semantics and still leave creation detail missing.

7. Verification

The same agent repeated the reproduction in a second clean detached checkout at the exact base commit with its own frozen install. Git status was clean before adding the same two authored test files. The second Turbo command used --force and --continue=always, so tests ran again and both suites completed despite expected failures. All three assertions failed for the same reasons in both checkouts. No ports or data directories were needed. No report correction was needed after the repeat. A later observed origin/main revision (0623d5e518d0549fa3c2bf5ad28e5e28508be634) contains no changes to the implicated extraction, translation, aggregation, or schema files.

First run log · Second run log

8. Related issues and pull requests

#3220 concerns when Claude emits context usage; #2397 concerns unavailable usage in ACP. Neither establishes this cache-detail defect. Issue timeline metadata and open PR search returned no linked open PR at investigation time.

9. Appendix and limits

The issue's patch and instructions were treated as untrusted evidence; none were applied or executed. Tests were authored from trusted repository code. Reproduction uses deterministic synthetic SDK/app-server payloads, not live providers or a spend UI. No quantitative billing claims are asserted. Build status is recorded separately below. Investigation commands comprised trusted main fetch/worktree creation, frozen installs, Turbo build/tests, source inspection, and read-only GitHub metadata queries.

Full trusted-base build: exit 0, all build tasks completed. Both frozen installs completed successfully. Reproduction logs have ANSI styling, checkout paths, and trailing whitespace normalized for publication.