Skip to content

fix(agent-sessions): net usage without copying lookup tables per reporter - #1197

Merged
JeremyFunk merged 1 commit into
mainfrom
fix/agent-sessions-netting-memory
Oct 1, 2026
Merged

JeremyFunk merged 1 commit into
mainfrom
fix/agent-sessions-netting-memory

Conversation

@JeremyFunk

@JeremyFunk JeremyFunk commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

The Agent Sessions list (aiSessionsPage) and the sidebar distributions (aiSessionsDistributions) fail with MEMORY_LIMIT_EXCEEDED (code 241) at the list profile's 1.5 GB for orgs with long agent runs. The facets read, which does no netting, succeeds over the same window.

Cause: the usage netting did its lookups inside lambdas that name a column.

  • per trace, for every trace in the window (before the page LIMIT): arrayMap(r -> … tokenLinks[tokenLinks[…[r.2]]] …, usageReporters)
  • per session: indexOf(childClaims…, r.1) and has(reporterIds, …)

ClickHouse copies a captured column once per array element, so memory is reporters x table size. ClickHouse names the first one in the error: while executing 'FUNCTION arrayElement(tokenLinks, tupleElement(r, 2))'.

Fix

  • lookupExpr: the same keyed lookup as a sort-merge over arrays. Table entries and needles are sorted together by key, and arrayFill carries the table's value down each run. No lambda names a column, so memory is linear.
  • The 4-hop ancestor climb is four chained lookups over one per-trace link table (t<SpanId> / c<SpanId> keys for tokens / cost).
  • Child claims and the "is a reporter" mark come from one sumMap, read back in one lookup.
  • Each lookup is a single expression (bind), because same-level aliases expand wherever they are named: the first attempt with aliases hit Query tree is too big and, once under the limit, ran 4x slower.
  • Session reporters use groupArrayArray(2000) instead of arraySlice(arrayFlatten(groupArray(…))), so the cap bounds the aggregate state too.

Netting rules are unchanged.

Verification

Check Old SQL New SQL
Production week, every session netted (949 sessions) hash 13809315504893199302 same hash
Same read, db.duration_ms (2-thread interactive profile, first run of each) 944 ms 1061 ms
Randomized span trees, local ClickHouse (wrappers reporting usage, shared response ids, multi-trace sessions; netting changes 3,986 raw calls to 2,270) — row-for-row equal
40 traces of 1,500–2,000 spans, local fails at 200 MB, 1.5 GB and 12 GB caps passes at 200 MB; 110 MB peak, 829 ms
New e2e case: one trace of 1,500 model calls under max_memory_usage = 100 MB fails (code 241) passes
  • ai-trace-index-materialization.clickhouse.e2e.test.ts: 10/10
  • catalog.clickhouse.e2e.test.ts: 405/405
  • packages/query-engine-integrations src/ai: 309/309

Not covered

  • deepestFailureCount still names failedSpans inside its lambda (has(tupleElement(failedSpans, 2), f.1)). Same class of cost for a trace with many failed spans; not what failed here, left for a follow-up.
  • The trace-level aggregation still holds every trace of the window before the page is cut. That part grows linearly with index rows.

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

Summary by CodeRabbit

  • Bug Fixes
    • AI session usage totals now account for token and cost reports across nested model calls, helping prevent usage from being counted more than once.
    • Session usage reporting has been updated to handle traces with many model calls, including traces with 1,500 calls.

…rter

The list and distributions reads looked up each reporter's ancestors and
charged claims inside a lambda that named the trace's link maps and the
session's claims. ClickHouse copies a captured column once per element, so
memory grew with reporters x table size, per trace for every trace in the
window, and a window of long agent runs exceeded the list profile's 1.5 GB.

The lookups are now a sort-merge over arrays (table entries and needles
sorted by key, arrayFill carrying the value down each run), written as one
expression per lookup so the query tree stays linear. The session's reporters
are capped by groupArrayArray instead of flatten-then-slice.

Same rows: identical output on a production week and on randomized span
trees against the previous SQL. A trace of 1,500 model calls now nets inside
100 MB in the ClickHouse e2e; the previous SQL fails that case.
@maple-review-bot

maple-review-bot Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Maple review

🟢 Confidence 4/5 · likely safe to merge
Clean, contained SQL rewrite; the one loose end is the __sql_baseline__/integrations.sql snapshot, which still records the old SQL and needs its regeneration run.
quality 100/100 · no findings · tests covered · risk medium

Rebuilds the Agent Sessions usage netting so ancestors and child claims come from sort-merge array lookups instead of lambdas that name a table per reporter, which is what blew the list read's 1.5 GB ceiling. The netting rules and the SQL-text, e2e and new low-memory assertions all still hold.

  • lookupExpr replaces per-reporter map lookups with one sort-merge over arrays
  • traceUsageColumns returns usageSpans, usageLinks and usageReporters in one call
  • nettedReportersExpr reads one sumMap back at four needles as lambda arguments
  • Session reporters capped by groupArrayArray(2000) rather than a slice after flatten
What was checked
  • Netting arithmetic maps 1:1 to the old element pairs (tokens/cost/five buckets, ai-span-columns.ts:250)
  • Climb semantics kept: 4 hops, '' fallback, per-measure t/c prefixes stripped by substring(t, 2)
  • Repo-wide grep: no references left to the removed usageLinksExpr, childClaimsExpr, reporterIds

7720f3e · Updated on every push. Reply "won't fix" to dismiss a finding, or mention @maple-review-bot to ask about one.

@coderabbitai

coderabbitai Bot commented Oct 1, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Advanced

Run ID: c40954ee-81a7-4ce6-a3a7-ddc9c7e4995d

📥 Commits

Reviewing files that changed from the base of the PR and between 5a6085d and 7720f3e.

📒 Files selected for processing (5)
  • packages/backend/src/services/warehouse/ai-trace-index-materialization.clickhouse.e2e.test.ts
  • packages/query-engine-integrations/src/ai/ai-sessions.test.ts
  • packages/query-engine-integrations/src/ai/ai-sessions.ts
  • packages/query-engine-integrations/src/ai/ai-span-columns.test.ts
  • packages/query-engine-integrations/src/ai/ai-span-columns.ts

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 2 remain after this review.


📝 Walkthrough

Walkthrough

AI session usage aggregation now resolves token and cost ancestry per trace, gathers reporters across traces, and nets claims at the session level. Tests cover the updated SQL and a trace with 1,500 model-call spans.

Changes

AI usage aggregation

Layer / File(s) Summary
Build trace reporter ancestry
packages/query-engine-integrations/src/ai/ai-span-columns.ts, packages/query-engine-integrations/src/ai/ai-span-columns.test.ts, packages/query-engine-integrations/src/ai/ai-sessions.ts
Trace usage columns combine reporter data with token and cost ancestor IDs. The trace-level tests check the link construction and ancestry lookups.
Collect and net session reporters
packages/query-engine-integrations/src/ai/ai-span-columns.ts, packages/query-engine-integrations/src/ai/ai-span-columns.test.ts, packages/query-engine-integrations/src/ai/ai-sessions.ts, packages/query-engine-integrations/src/ai/ai-sessions.test.ts, packages/backend/src/services/warehouse/ai-trace-index-materialization.clickhouse.e2e.test.ts
Session queries gather reporters across traces and net claims with a span-keyed claim table. Tests check the updated SQL and query results for a trace with 1,500 model-call spans.

Priority: ⬆️ High

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Bug fix

Suggested reviewers: makisuo

Merge Risk: ⚪ Minimal · up to 7720f

This change rewrites how Agent Sessions nets token and cost usage, to avoid memory-limit failures on long agent runs. No concrete correctness or stability defect was established at the current head, and the one open question about link truncation was shown to match the previous behavior.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: reducing memory use by avoiding lookup-table copying during agent-session usage netting.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 5 functions across 5 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Autopilot is currently an internal CodeRabbit preview.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@JeremyFunk
JeremyFunk merged commit 674a12b into main Oct 1, 2026
35 of 37 checks passed
@JeremyFunk
JeremyFunk deleted the fix/agent-sessions-netting-memory branch October 1, 2026 17:00
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant