A context optimization runtime for AI agents.
Spend context on what matters.
Beam is not a token compressor. Beam is the runtime that decides which information an agent should be allowed to see.
Most "compression" tools answer "how much can we remove?". Beam answers "what is the smallest context that still lets the agent finish the task?"
Those are different questions, and the second one is the one that matters:
- A 90% reduction that causes a retry, a wrong fix, or a missed error is a failure, not a win.
- A 40% reduction that preserves first-pass correctness is a major win.
Every competitor's own documentation concedes this. Caveman reports a user who measured a net token increase ([#145]) and an unreproducible A/B showing 4.3M tokens with it on versus 1M off ([#550]). Headroom documents that plain source code compresses 0% by design, and concedes its headline table was "not from a published, repeatable benchmark suite". CAVEWOMAN (arXiv 2606.24083) shows that input compression can raise net cost, because models compensate with longer responses even as accuracy collapses.
Beam takes that seriously enough to make it architectural.
Three ideas carry the whole system.
Fidelity is the controlling parameter: full, safe, compact, minimal.
It answers "what are we allowed to change", not "how small can we get". The
tier ladder refuses to escalate past the ceiling the active fidelity sets.
Generative summarization is never granted by a fidelity value — it is gated behind explicit opt-in — because it is the only rung that can invent statements absent from the source.
Context waste is not uniformly distributed. Before asking how to shrink something, Beam asks which parts would have changed the agent's next action.
| P0 | errors, failed assertions, security findings — never dropped |
| P1 | files, symbols, test names to act on — never dropped |
| P2 | dependencies, package layout |
| P3 | disambiguating context |
| P4 | already-delivered / repeated success output |
| P5 | progress bars, banners, ANSI |
P0/P1 protection is an invariant, not a setting. Budget pressure is resolved by escalating tiers or reporting failure — never by silently discarding an error message.
Every token count carries a TokenBasis. Estimates are never presented as
counts, and a reduction built on an estimate is not quotable as a cost claim.
A 90% token reduction is not a 90% cost reduction, and Beam will not let you
print it as one.
Rungs, in escalating risk order. Beam tries each one and keeps a stage only if it actually made the payload smaller.
| Tier | Name | Example |
|---|---|---|
| 0 | passthrough |
do nothing — a real decision |
| 1 | lossless_representation |
strip ANSI, normalize newlines |
| 2 | deterministic_reduction |
collapse repeated lines, aggregate passing tests |
| 3 | extractive_selection |
keep failures and changed hunks |
| 4 | virtualization |
keep raw bytes outside context; expose search/expand |
| 5 | learned_compression |
LLMLingua-style (opt-in) |
| 6 | generative_summarization |
never by fidelity alone |
cargo build --releasegit status | beam reduce # compact by default
cat build.log | beam reduce --fidelity safe
cargo test 2>&1 | beam measure # report, don't rewriteOn a 500-pass/1-fail cargo test run, Beam keeps every failure and drops
the successes:
$ cargo test 2>&1 | beam reduce
500 passing
running 501 tests
test auth::suite::case_broken ... FAILED
failures:
---- auth::suite::case_broken stdout ----
thread 'auth::suite::case_broken' panicked at src/auth.rs:214:9:
assertion `left == right` failed
left: 500
right: 200
failures:
auth::suite::case_broken
test result: FAILED. 500 passed; 1 failed; 0 ignored; 0 measured17,230 bytes in, 348 out — with the panic location and the exact expected-vs-got values intact.
Any reduced payload can be kept addressable rather than resident:
$ cargo test 2>&1 | beam reduce --recover
500 passing
running 501 tests
test auth::suite::case_broken ... FAILED
...
[beam: 348 of 17230 bytes shown (97% reduced); full output: beam://4563f4be...]The agent now knows 16,882 bytes exist and how to reach them. It can go straight to what it needs:
beam search beam://4563f4be... FAILED # which tests failed
beam range beam://4563f4be... 502 515 # fetch just those lines
beam expand beam://4563f4be... # the byte-exact original$ printf '\033[31mFAILED\033[0m tests/test_auth.py::test_refresh\nretrying\nretrying\nretrying\nretrying\n' | beam measure
tier=lossless_representation 22 -> 19 tokens (13.6% reduction, basis estimated:mixed); 85 -> 76 bytes; 0 artifact(s) storedThe basis estimated:mixed is not decoration. It is Beam refusing to pretend
an estimate is a measurement.
Note that four repeated lines save almost nothing, and Beam correctly reports that. A transform only earns its place when it actually pays:
$ printf '\033[31mFAILED\033[0m tests/test_auth.py::test_refresh\n' > /tmp/e
$ for i in $(seq 1 40); do echo 'retrying connection to redis:6379' >> /tmp/e; done
$ beam measure < /tmp/e
tier=deterministic_reduction 365 -> 37 tokens (89.9% reduction, basis estimated:mixed); 1459 -> 148 bytes; 0 artifact(s) stored| Crate | Role |
|---|---|
beam-core |
Domain types: fidelity, priority, tiers, budget, measurement |
beam-tokens |
Token counting, honest about exact vs estimated |
beam-optimize |
The reducers and the pipeline that walks the ladder |
beam-store |
Content-addressed artifact store, recovery gate |
beam-adapters |
Per-tool structural parsers (test runners) |
beam-cli |
The beam binary |
beam-bench |
Correctness-first benchmark harness |
beam-core has no I/O and no compression logic. It is the vocabulary every
other crate agrees on, which is what keeps Beam independently installable and
free of compile-time coupling to any sibling product.
Implemented and tested (177 tests, clippy pedantic clean):
- the fidelity, priority, tier, budget, and measurement models;
- Tier 1 — ANSI and control-character removal;
- Tier 2 — contiguous repeat collapse, first/last retention;
- Tier 3 — test-runner structural selection (cargo, pytest, Jest, Go);
- the ladder-walking pipeline, including the "never grow the output" guard;
- a content-addressed artifact store with byte-exact recovery, TTL policy, size caps, and a no-store mode for payloads that may contain secrets;
- the recovery gate, which stores the original before any lossy transform and degrades to pass-through when recovery is unavailable;
store,show,expand,search, andrangefor progressive disclosure;- Tier 4 — an inline marker telling the agent exactly what it is not seeing, with a handle it can follow.
Not yet implemented: Tiers 5-6 (learned and generative compression), the
remaining adapters, the context planner, and the agent integrations. See
docs/PLAN.md for phases 4-7.
docs/PLAN.md— the master plan: phases 0-7 and exit criteriadocs/research/— verbatim practitioner evidence with citationsdocs/BENCHMARKS.md— the correctness-first benchmarkdocs/RESEARCH.md— what the prior art actually does, what was stolen, and what was deliberately avoideddocs/INTEGRATIONS.md— which agent surfaces can rewrite tool output today, and which cannot
Apache-2.0. See LICENSE and
docs/LICENSING.md for the reasoning, including why
Apache-2.0 over the MIT used by sibling GrayCodeAI repositories.
Read AGENTS.md first.