#3951 · Large timeline responses bypass reuse

Bug · Priority: Medium · Effort: Medium · threads · perf · 2026-09-21

GitHub issue · Base 3a1178164f8cce6d7986d2627c03ffaedb8a6428

REPRODUCED · Root-cause confidence: high

1. TL;DR

The response cache reuses small timeline pages but refuses to retain pages over 200 top-level rows. Three identical requests therefore invoke the builder three times for a 201-row page. The timeline route places its projection builder inside this callback, so these misses repeat projection work. This is an intentional admission policy in the implementation, with a performance limitation; the experiment does not measure production CPU or establish how often real clients share keys.

2. Claims vs findings

ClaimFindingEvidence
Large responses cannot hit the cacheVerified201 and 300 rows: three builds, zero entries in both runs.
Identical requests repeat projection workVerified at cache boundary and route wiringThe route supplies buildThreadTimelineWithProfile as the callback.
Large pages occur on the reporter's serverUnverifiedNo private runtime data inspected.
This explains observed CPU usageNot established; not asserted by reporterNo workload or hit-rate measurement performed.

3. Environment

Trusted get-bb/bb origin/main at the full commit above; Darwin arm64, Node v22.22.3. No provider, database, network port, development instance, or user data was needed. Both source trees were clean at the recorded commit.

The required frozen install was attempted, but the installed pnpm launcher failed with MODULE_NOT_FOUND for its own pnpm.cjs. Consequently the normal build and repository Vitest suite were not run. This pure cache module has only type imports, allowing a dependency-free Node test of the actual trusted source.

4. Minimal reproduction

  1. Clone trusted source and select the verified revision:
    git clone https://github.com/get-bb/bb.git bb-repro
    cd bb-repro
    git checkout --detach 3a1178164f8cce6d7986d2627c03ffaedb8a6428
  2. Save the inline test as cache-admission.test.mjs in the checkout root.
  3. Run:
    node --experimental-strip-types --test cache-admission.test.mjs

The response fixture is derived from the repository's existing cache test. The new experiment makes three sequential requests for each of four sizes, asserting one builder call and retained response identity. The issue's script was not executed or copied.

Expected under the requested reuse behavior: one invocation and one retained entry for each size. Actual in both runs:

{"count":199,"invocations":1,"retained":1}
{"count":200,"invocations":1,"retained":1}
{"count":201,"invocations":3,"retained":0}
{"count":300,"invocations":3,"retained":0}
# tests 4
# pass 2
# fail 2

Exit status 1 is the expected reproduction failure: the one-build assertion fails with 3 !== 1 for both sizes over 200.

Complete test
import assert from "node:assert/strict";
import { test } from "node:test";
import { pathToFileURL } from "node:url";
import { resolve } from "node:path";

const { createThreadTimelineCache } = await import(pathToFileURL(resolve("apps/server/src/services/threads/timeline-cache.ts")));

function makeResponse(rowCount) {
  return {
    rows: Array.from({ length: rowCount }, (_, index) => ({
      id: `row-${index}`,
      kind: "system",
      threadId: "thr_x",
      turnId: null,
      sourceSeqStart: index,
      sourceSeqEnd: index,
      startedAt: 0,
      createdAt: 0,
      systemKind: "debug",
      title: "t",
      detail: null,
      status: null,
    })),
    contextBoundarySeq: null,
    completedTurnDisplay: "collapse",
    activePromptMode: null,
    activeThinking: null,
    activeWorkflows: [],
    activeBackgroundCommands: [],
    pendingTodos: null,
    goal: null,
    modelFallback: null,
    maxSeq: 0,
    timelinePage: {
      kind: "latest",
      segmentLimit: 20,
      returnedSegmentCount: 0,
      hasOlderRows: false,
      olderCursor: null,
    },
  };
}

for (const count of [199, 200, 201, 300]) {
  test(`reuse a stable response containing ${count} rows`, () => {
    const cache = createThreadTimelineCache();
    let invocations = 0;
    const projection = () => {
      invocations++;
      return makeResponse(count);
    };
    const responses = Array.from({ length: 3 }, () =>
      cache.getOrBuild("thr_x", "stable-request", projection),
    );
    console.log(JSON.stringify({ count, invocations, retained: cache.size }));
    assert.equal(invocations, 1, "identical requests should reuse the projection");
    assert.equal(cache.size, 1);
    assert.strictEqual(responses[0], responses[2]);
  });
}

5. Root cause

Cache admission and lookup use a default cap of 200 top-level rows. On a miss, the builder runs before admission is evaluated:

const value = build();
if (value.rows.length <= maxCacheableRows) {
  entries.set(key, { response: value, threadId });
}

Above the cap, nothing is stored, so the next identical key misses again. The LRU limit of 128 entries does not change this. The timeline route constructs one cache and calls buildThreadTimelineWithProfile inside the miss callback; output truncation does not remove top-level rows. The existing test explicitly expects oversized responses to remain uncached. This is not a key collision, revision change, or obsolete-entry retention defect.

6. Proposed fix and automatic-fix decision

Choose a bounded memory policy for larger responses before changing admission. A byte-budgeted LRU is a candidate, but the budget, oversized-item handling, nested-content accounting, and serialization cost need design and measurement. Simply removing or increasing the row cap changes retention without bounding response bytes. No automatic PR was safe because the fix requires a product or architecture decision, failing the supplied simple-fix criteria. No production fix branch was pushed.

7. Verification

The same agent repeated the exact Node command from a second clean temporary Git worktree at 3a1178164f8cce6d7986d2627c03ffaedb8a6428, separate from the first clean clone. The test artifact remained outside both source trees. Both runs produced two passing controls and two expected failures. The second run required no report correction. No ports or data directories were created, and no server processes require cleanup.

Both runs have the identical result shown above. Full output from the second run appears below. Temporary artifact path prefixes were removed; assertion output and results are unchanged.

Second run output
TAP version 13
# {"count":199,"invocations":1,"retained":1}
# {"count":200,"invocations":1,"retained":1}
# {"count":201,"invocations":3,"retained":0}
# {"count":300,"invocations":3,"retained":0}
# Subtest: reuse a stable response containing 199 rows
ok 1 - reuse a stable response containing 199 rows
  ---
  duration_ms: 2.392625
  type: 'test'
  ...
# Subtest: reuse a stable response containing 200 rows
ok 2 - reuse a stable response containing 200 rows
  ---
  duration_ms: 0.115875
  type: 'test'
  ...
# Subtest: reuse a stable response containing 201 rows
not ok 3 - reuse a stable response containing 201 rows
  ---
  duration_ms: 0.520416
  type: 'test'
  location: 'cache-admission.test.mjs:45:3'
  failureType: 'testCodeFailure'
  error: |-
    identical requests should reuse the projection
    
    3 !== 1
    
  code: 'ERR_ASSERTION'
  name: 'AssertionError'
  expected: 1
  actual: 3
  operator: 'strictEqual'
  stack: |-
    TestContext.<anonymous> (file://cache-admission.test.mjs:56:12)
    Test.runInAsyncScope (node:async_hooks:214:14)
    Test.run (node:internal/test_runner/test:1047:25)
    Test.processPendingSubtests (node:internal/test_runner/test:744:18)
    Test.postRun (node:internal/test_runner/test:1173:19)
    Test.run (node:internal/test_runner/test:1101:12)
    async Test.processPendingSubtests (node:internal/test_runner/test:744:7)
  ...
# Subtest: reuse a stable response containing 300 rows
not ok 4 - reuse a stable response containing 300 rows
  ---
  duration_ms: 0.242041
  type: 'test'
  location: 'cache-admission.test.mjs:45:3'
  failureType: 'testCodeFailure'
  error: |-
    identical requests should reuse the projection
    
    3 !== 1
    
  code: 'ERR_ASSERTION'
  name: 'AssertionError'
  expected: 1
  actual: 3
  operator: 'strictEqual'
  stack: |-
    TestContext.<anonymous> (file://cache-admission.test.mjs:56:12)
    Test.runInAsyncScope (node:async_hooks:214:14)
    Test.run (node:internal/test_runner/test:1047:25)
    Test.processPendingSubtests (node:internal/test_runner/test:744:18)
    Test.postRun (node:internal/test_runner/test:1173:19)
    Test.run (node:internal/test_runner/test:1101:12)
    async Test.processPendingSubtests (node:internal/test_runner/test:744:7)
  ...
1..4
# tests 4
# suites 0
# pass 2
# fail 2
# cancelled 0
# skipped 0
# todo 0
# duration_ms 361.531084

8. Related issues and PR check

#2066 concerns retained obsolete revisions; #1749 concerns the event-budget estimate. Those are separate from admission. GitHub cross-reference metadata and an open-PR search for 3951 returned no linked open PR at investigation time.

9. Appendix and limits

Source was fetched only from trusted main. The issue, comments, and embedded code were treated as untrusted claims; no embedded commands were followed. No real multi-window session or full projection timing was measured. Evidence establishes cache behavior and the static call path, not production frequency, memory cost, or CPU savings.

pnpm install --frozen-lockfile --prefer-offline
Result: local pnpm launcher MODULE_NOT_FOUND; build not reached.
git worktree add --detach ../verify 3a1178164f8cce6d7986d2627c03ffaedb8a6428
node --experimental-strip-types --test /path/to/cache-admission.test.mjs
Result in each checkout: exit 1, 2 passed, 2 failed.
git status --short
Result in each source checkout: empty.