Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Beam

A context optimization runtime for AI agents.

Spend context on what matters.

Beam is not a token compressor. Beam is the runtime that decides which information an agent should be allowed to see.

The thesis

Most "compression" tools answer "how much can we remove?". Beam answers "what is the smallest context that still lets the agent finish the task?"

Those are different questions, and the second one is the one that matters:

  • A 90% reduction that causes a retry, a wrong fix, or a missed error is a failure, not a win.
  • A 40% reduction that preserves first-pass correctness is a major win.

Every competitor's own documentation concedes this. Caveman reports a user who measured a net token increase ([#145]) and an unreproducible A/B showing 4.3M tokens with it on versus 1M off ([#550]). Headroom documents that plain source code compresses 0% by design, and concedes its headline table was "not from a published, repeatable benchmark suite". CAVEWOMAN (arXiv 2606.24083) shows that input compression can raise net cost, because models compensate with longer responses even as accuracy collapses.

Beam takes that seriously enough to make it architectural.

Design

Three ideas carry the whole system.

1. Fidelity, not aggressiveness

Fidelity is the controlling parameter: full, safe, compact, minimal. It answers "what are we allowed to change", not "how small can we get". The tier ladder refuses to escalate past the ceiling the active fidelity sets.

Generative summarization is never granted by a fidelity value — it is gated behind explicit opt-in — because it is the only rung that can invent statements absent from the source.

2. Information priority (P0–P5)

Context waste is not uniformly distributed. Before asking how to shrink something, Beam asks which parts would have changed the agent's next action.

P0 errors, failed assertions, security findings — never dropped
P1 files, symbols, test names to act on — never dropped
P2 dependencies, package layout
P3 disambiguating context
P4 already-delivered / repeated success output
P5 progress bars, banners, ANSI

P0/P1 protection is an invariant, not a setting. Budget pressure is resolved by escalating tiers or reporting failure — never by silently discarding an error message.

3. Honest measurement

Every token count carries a TokenBasis. Estimates are never presented as counts, and a reduction built on an estimate is not quotable as a cost claim. A 90% token reduction is not a 90% cost reduction, and Beam will not let you print it as one.

The optimization ladder

Rungs, in escalating risk order. Beam tries each one and keeps a stage only if it actually made the payload smaller.

Tier Name Example
0 passthrough do nothing — a real decision
1 lossless_representation strip ANSI, normalize newlines
2 deterministic_reduction collapse repeated lines, aggregate passing tests
3 extractive_selection keep failures and changed hunks
4 virtualization keep raw bytes outside context; expose search/expand
5 learned_compression LLMLingua-style (opt-in)
6 generative_summarization never by fidelity alone

Install and use

cargo build --release
git status | beam reduce                    # compact by default
cat build.log | beam reduce --fidelity safe
cargo test 2>&1 | beam measure             # report, don't rewrite

On a 500-pass/1-fail cargo test run, Beam keeps every failure and drops the successes:

$ cargo test 2>&1 | beam reduce
500 passing
running 501 tests
test auth::suite::case_broken ... FAILED
failures:
---- auth::suite::case_broken stdout ----
thread 'auth::suite::case_broken' panicked at src/auth.rs:214:9:
assertion `left == right` failed
  left: 500
 right: 200
failures:
    auth::suite::case_broken
test result: FAILED. 500 passed; 1 failed; 0 ignored; 0 measured

17,230 bytes in, 348 out — with the panic location and the exact expected-vs-got values intact.

Virtualized recovery

Any reduced payload can be kept addressable rather than resident:

$ cargo test 2>&1 | beam reduce --recover
500 passing
running 501 tests
test auth::suite::case_broken ... FAILED
...
[beam: 348 of 17230 bytes shown (97% reduced); full output: beam://4563f4be...]

The agent now knows 16,882 bytes exist and how to reach them. It can go straight to what it needs:

beam search beam://4563f4be... FAILED   # which tests failed
beam range  beam://4563f4be... 502 515  # fetch just those lines
beam expand beam://4563f4be...          # the byte-exact original
$ printf '\033[31mFAILED\033[0m tests/test_auth.py::test_refresh\nretrying\nretrying\nretrying\nretrying\n' | beam measure
tier=lossless_representation 22 -> 19 tokens (13.6% reduction, basis estimated:mixed); 85 -> 76 bytes; 0 artifact(s) stored

The basis estimated:mixed is not decoration. It is Beam refusing to pretend an estimate is a measurement.

Note that four repeated lines save almost nothing, and Beam correctly reports that. A transform only earns its place when it actually pays:

$ printf '\033[31mFAILED\033[0m tests/test_auth.py::test_refresh\n' > /tmp/e
$ for i in $(seq 1 40); do echo 'retrying connection to redis:6379' >> /tmp/e; done
$ beam measure < /tmp/e
tier=deterministic_reduction 365 -> 37 tokens (89.9% reduction, basis estimated:mixed); 1459 -> 148 bytes; 0 artifact(s) stored

Workspace layout

Crate Role
beam-core Domain types: fidelity, priority, tiers, budget, measurement
beam-tokens Token counting, honest about exact vs estimated
beam-optimize The reducers and the pipeline that walks the ladder
beam-store Content-addressed artifact store, recovery gate
beam-adapters Per-tool structural parsers (test runners)
beam-cli The beam binary
beam-bench Correctness-first benchmark harness

beam-core has no I/O and no compression logic. It is the vocabulary every other crate agrees on, which is what keeps Beam independently installable and free of compile-time coupling to any sibling product.

Status

Implemented and tested (177 tests, clippy pedantic clean):

  • the fidelity, priority, tier, budget, and measurement models;
  • Tier 1 — ANSI and control-character removal;
  • Tier 2 — contiguous repeat collapse, first/last retention;
  • Tier 3 — test-runner structural selection (cargo, pytest, Jest, Go);
  • the ladder-walking pipeline, including the "never grow the output" guard;
  • a content-addressed artifact store with byte-exact recovery, TTL policy, size caps, and a no-store mode for payloads that may contain secrets;
  • the recovery gate, which stores the original before any lossy transform and degrades to pass-through when recovery is unavailable;
  • store, show, expand, search, and range for progressive disclosure;
  • Tier 4 — an inline marker telling the agent exactly what it is not seeing, with a handle it can follow.

Not yet implemented: Tiers 5-6 (learned and generative compression), the remaining adapters, the context planner, and the agent integrations. See docs/PLAN.md for phases 4-7.

Design notes and research

License

Apache-2.0. See LICENSE and docs/LICENSING.md for the reasoning, including why Apache-2.0 over the MIT used by sibling GrayCodeAI repositories.

Contributing

Read AGENTS.md first.

About

A context optimization runtime for AI agents. Decides which information an agent is allowed to see.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages