Skip to content

Latest commit

 

History

History

Folders and files

NameName
Last commit message
Last commit date

parent directory

..
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Evidence

Measurements behind the claims in the README. Each file records the method as well as the result, so a reader can decide whether the number means what it says. A new measurement gets a new file.

These files are dated, but "never edited after the fact" is not true of the history and this page said it was until 2026-09-10. Twenty-three of the forty-two dated files carry more than one commit. Most of those edits are appended retractions, which is the intended shape. Several are not: rehydration-boundary-2026-09-05.md lost an eight-line section at ad7e884, and routing-baseline-2026-09-06.md lost a published 85.10% at a6b259b. Read a file as its reading on its date plus whatever was appended, and git log -p docs/evidence/ as the only complete record.

Document What it measures Headline
Calibration corpus, 2026-09-03 How well the replay engine reproduces the provider's own cache reads 398 of 402 turns across 11 sessions, all from this repository's own development on one machine
Calibration corpus, 2026-09-05 How well the engine reproduces the provider's own numbers, across 1363 transcripts from many unrelated projects rather than this repository's own development work. It called those transcripts "sessions"; see the 2026-09-06 correction Supersedes the 2026-09-03 corpus, which covered 11 self-referential sessions
Calibration corpus, 2026-09-06 The same engine, re-read, and a correction: the previous file counted transcript files and called them sessions 1450 transcripts from 78 sessions, 97.46%. The 1363 published as "sessions" was a file count; one session supplied 1020 of them
Calibration corpus, 2026-09-10 A later reading of the same engine on a grown corpus, not a correction of the 2026-09-06 figures. Records one defect it did not fix: the per-model Sessions column still carries lane counts. Addendum 2026-09-11: the headline never said how much of it was exact 1751 transcripts from 116 sessions, 97.79%. The rate moved 0.33 points across both a larger corpus and 67 engine commits, and this reading cannot separate those causes. The 97.79% counts exact reproductions together with turns where the provider served MORE prefix than predicted; a 2026-09-11 re-reading of the same corpus root splits a 97.89% match rate into 94.10% exact, 3.78% read more than predicted, 2.12% broken, and 258 transcripts clear the 95% calibration gate only because the second kind counts as a match
Launch KPIs, pre-registered 2026-09-12 What the 2026-09-14 launch is expected to produce, and the thresholds that decide what each outcome means, written before the day Baseline is 53 fetches from 2 distinct IPs and zero external users. Predicts a 25% front page, a median 38 installs if it lands, and a 15% to 25% chance of the one thing that matters: a contributed corpus that is not the maintainer's. Everything except the baseline is modelled from invented population shares and is labelled so. Results get appended, including the rows that were wrong
What a corpus contribution contains, 2026-09-10 The exact bytes cost --contribute produces, and the surface replay mcp exposes, both read from the types that build them 13 scalars, 394-432 bytes — smaller than one transcript line, and the package cannot transmit at all (an import test bans net). Amended 2026-09-12: 19 fields and under 600 bytes, after three build-identity fields landed for #284, and the "no text" wording is restated there rather than dropped. The disclosure is not the field list: totalUsd cross-matches a shared cost card, the pull-request transport binds the identity the payload omits, and a stable tag over repeated submissions is a spend time series
Keepalive is the wrong lever, 2026-09-10 Whether holding a prompt cache open with pings would pay on this corpus, tested against arXiv:2607.19214's own break-even formula, and whether arXiv:2601.06007's "exclude tool results" finding reproduces here $77.24 recoverable by keepalive against $508.83 that is not: 87% of cache-creation spend happens on gaps under five minutes, where the cache had not expired and something rewrote the prefix while it was warm. The published 46-minute band is 4.3x wider (196 min) at the 0.025 read multiple Replay's own table already carries for two models. blame puts tool traffic at ranks 1, 3 and 4 independently, agreeing with the second paper
Main-lane pricing, 2026-09-10 Why replay cost and the new replay burn cost column disagree by 2.8x on one corpus, and which requests fall between them 21,854 of 60,401 requests (36.2%) are dropped, concentrated in 8 of 1,812 files — the fan-out sessions, where the money is. cost prices one lane per transcript; burn prices all of them. Zero overlap between a session's in-file lanes and its subagents/ files, so they are not duplicates. Measured, not fixed: the fix moves every published dollar figure
The fan-out premium, 2026-09-06 What parallel subagent lanes cost, against a baseline where siblings share the cache write. Corrected the same day: most of the number is arithmetic 1.68x / 2.56x / 3.34x, sitting at 88-99% of a ceiling fixed by the group size and the provider's own multipliers. No corpus can produce a premium below 1, so the rise with width is the shape of the estimator. Only the dispersion ratio (~0.9, flat) is empirical
Compaction and the index, 2026-09-06 What context compaction discards, and what indexing transcripts saves 39 compactions keep a median 2.55%, discarding 30.5M tokens over 73.5 minutes of wall clock; the index takes replay cost from 6.474s to 0.046s
Wire families, 2026-09-06 Which request shapes a terminal GenAI client actually sends, captured off the wire from a live session Three families, not two. Partly retracted 2026-09-08: the "no local transcript" finding was wrong, ~/.grok/sessions holds 3.8 GB across 6,787 files, and a $406.07 figure derived from a claim about that store is withdrawn. The Grok CLI speaks OpenAI Responses (/responses), which Replay does not parse. It returns the four x-ratelimit-* headers this project had never captured, and remaining equalled limit on every call, so they did not measure consumption. An earlier version of the file called them a better instrument, on header names before any value was read
Lane isolation, 2026-09-06 How much of a fan-out session's re-billing comes from a changed prefix, and what the answer was while the instrument compared each sub-agent against a different one 4.2%, published the same morning as 98.8%. Both the 98.8% and the 11.0% intermediate are retracted in the file rather than removed; 31 of the 34 events had never happened
Ollama cache ceiling, 2026-09-08 Whether "100% of the prompt served from cache" is a claim any Ollama request can make, read off 40 MB of one machine's own server logs It cannot, and the refusal is deliberate. When llama.cpp finds the whole prompt resident it backs off exactly one token so it has something to evaluate, and logs the reason. The ceiling is (n-1)/n and the observed maximum equals it to six decimal places. Worse for reporting: n_past appears only on those requests, so a cached share computed from it measures a population defined by having been fully cached. replay burn printed 100% for a 99.7318% share and now prints two counts and no average
Surface census, 2026-09-08 What each agent actually writes to one disk, checked by opening the directories rather than inferring from the wire Everything Replay claims to read, it reads completely: Claude Code 1,681 of 1,681, Codex 150 across both session roots, Ollama's narrower glob correct. Two surfaces documented as absent are present: Grok 3.8 GB over 6,787 files, Cursor 118 agent transcripts. Neither carries usage, so the conclusion held and both stated reasons were false. A $406.07 figure derived from a summary about Grok is withdrawn
QM budget, 2026-09-08 Whether the spend budget in yc-software/qm — which throttles agents — computes a number that means what it is used for No, by a factor of eight. claude-harness.ts:570 collapses Anthropic's four usage fields into one inputTokens and budget.ts bills it at a flat $5/MTok, so cache reads cost the same as fresh input. Tested on 52,511 real requests: 8.1× aggregate, 16.3× on Sonnet. Ignoring the cache is 101% of the error; the flat rate is −1.1%. Separately, the five harnesses fill inputTokens from five different quantities, two of them local tokenizer estimates, all summed into one org budget
TUI example screens, 2026-09-08 Which TUI screens show the reader their own machine, checked by rendering all ten against a real corpus Five of ten describe nobody: context, advise, guards, model, safe. Not a defect in the analysis — the same corpus through the command line gives Bash 43.2%, 209k, x106 — but those screens have no case in the dispatch and fall through to a canned illustration. The installer opens this surface automatically, so it is the first thing a new user sees
Subscription allowance, 2026-09-09 What a cache break costs a reader who is not billed per token, which is most readers RETRACTED IN FULL 2026-09-09. Both terms are unsound: the utilization reading exists only as a source comment (zero ledger records carry any rate-limit header, ever), and the 3,778,706 multiplicand is retracted at source in lane-isolation — "Both figures are retracted". Kept as a specification for a measurement nobody has taken, and the figure is an 8.94x extrapolation of a trial that resolved five steps at 0.01 granularity — quote 0.36-0.54, not 0.45. Re-billed tokens do move anthropic-ratelimit-unified-*, so "rate-limit budget spent on nothing" is literal rather than a figure of speech. Whether cache reads carry the 0.10x discount against the allowance as they do against the bill is NOT MEASURED, and is the parameter everything here is sensitive to
Ollama cache observable, 2026-09-09 Whether a prompt-cache hit is observable at all on a surface that reports no cached tokens It is, in the duration. Neither Ollama API reports a cache: prompt_tokens and prompt_eval_count are identical warm and cold. prompt_eval_duration is not — 4211.9 +/-23.8 us/token cold against 15.4 +/-1.5 warm, 273x apart with no overlap between the extremes, and a prefix change sends it back to cold while the model stays loaded, which rules out warm-up. It tracks how much was re-evaluated, not just whether, so cached tokens are recoverable to within 8pp as count - duration/cold_rate. Concurrency tested and survives: 269x separation under 4-way parallel and under a mixed warm/cold batch, because prompt_eval_duration reports per-request work, not elapsed time — cold wall clock spanned 11.6-46.6s while the reported duration stayed within 0.3%. An estimator built on latency would have been destroyed by queueing
Break causes, 2026-09-06 Which causes actually re-bill tokens, across the whole corpus rather than a sample 735 breaks, 31.26M tokens. Client re-render 50.8%, TTL expiry 33.9%. The same measurement on the 40 largest sessions said TTL 75.2% - sorting by size selected for the cause
Routing baseline, 2026-09-06 What model replay route takes as its baseline, how far actual routing departs from it, and which half of "blocked on n=1" is real Baseline claude-opus-5, 85.67% of 20,497 fitted turns; deviation 14.33%. The transcript corpus is 1,606 transcripts from 115 sessions, not n=1; the probe series is n=1 reading per model and is the part actually blocked. Sigma's +/-85% error bar is the estimator's own floor, shown by measuring opus-5 against itself, so no size of corpus will shrink it. Fixed 2026-09-09 (#96): quadrature assumed the two sides were independent, but on the identity pair they are the same fit at rho = 1; it now reads +/-0%. The corpus still cannot compare two models, which was always the binding limit
Quota titration, 2026-09-06 Whether a cache break burns a flat-seat subscription quota the way it burns a metered bill Null, and stated as null. 3.09M tokens across matched cold-write and warm-read arms moved the 5h counter by zero steps; the instrument is too coarse to answer at this budget
Proxy added latency, 2026-09-03 What replay serve adds to a request, and why the first attempt could not measure it ~1.7ms p50 (corrected 2026-09-05; the 48µs originally published was noise)
Rehydration boundary under attack, 2026-09-05 Whether a poisoned agent can write a real credential outside the project Eight vectors refused; the harness was checked by weakening the boundary, which showed filepath.Clean alone defeats dot-dot and only EvalSymlinks stops the symlink escape
Spike: OpenAI-compatible path, 2026-09-05 What the proxy does with a request it was not built for It forwarded it correctly and applied nothing: a one-token spend cap refused none of three requests. Led to the warnings, then to the path being built
Spike: Cursor, 2026-09-05 Whether Cursor is a transcript problem or a provider problem A provider problem. Cursor stores 29,665 message rows and zero cache fields, so the transcript path could never produce cache forensics
Spike 4: the real provider, 2026-09-05 Whether the proxy works against the real provider at all Ten turns, all 200, 1,816,417 prompt tokens measured, zero credentials and zero message content in the ledger
Adversarial security review, 2026-09-04 An external reviewer reading the code and running the proxy end to end Six findings resolved, one open in part (the vault key sits beside the ciphertext), and what was verified to hold
codex-cache-breaks-2026-09-07 Codex cache breaks and their causes 80 breaks re-read 10,635,679 tokens cold; five client-side causes ruled out, the sixth is not logged
codex-quota-2026-09-07 A live quota signal, and its unit The counter moves, unlike every other surface checked; cached reads weigh far less, by an amount this corpus cannot pin down
cache-accounting-shapes-2026-09-07 What each provider's usage fields mean arithmetically The input field names the whole prompt on three surfaces and a remainder on the fourth
Cost after the parser fixes, 2026-09-10 Whether merging the four transcript-parser PRs moved any figure replay cost prints, run pre- and post-merge over the same corpus in the same minute No figure moved, and that is the expected answer. Cost prices the provider's usage object, so the image-block fix (#138) changes attribution and not spend — it shows up as unaccounted falling 133k×133 to 132k×132 in replay context. The Codex double-count fix (#137) applies to a wire family absent from this corpus, where route is first-party on all 116 sessions. The $3,382.13 → $3,411.92 move against an earlier baseline is one more session, not code
Avoidable concentration, 2026-09-10 Where the corpus-wide re-billed figure actually sits, across 116 sessions 25 fan-out sessions hold 98.8% of all re-billed spend; 91 single-lane sessions hold 1.2% ($2.03 total), and the top five sessions carry 91%. Two true sentences point opposite ways: 96 of 116 sessions have zero re-billed spend, and those 96 are 2.2% of spend. Corroborates the per-repository unit and leaves the "one-time fix" objection open, because cause attribution needs the proxy and this is estimated. The measuring session is in its own sample
Does the ±10% band decide the re-render headline?, 2026-09-11 Whether the 50.8% re-render share is an artifact of rerenderTolerance = 0.10, swept from 0% to infinity over the whole corpus No, and 50.8% is also not today's number. The band decides 8 of 597 classifications and 1.52 points; the whole range across every possible tolerance is 8.1 points. 589 of the 597 sit at exact integer equality, and all 589 compare two provider-reported numbers, because UnseenPrefix is measured from the first request's cache_read whenever the lane started warm — one break in the corpus is both band-decided and fit-estimated, worth 4,045 tokens. TTL, system and model rows are byte-identical at every tolerance. benefit-gap-analysis.md pre-registered the refutation — "0.10 to 0.15 moves more than 10% of breaks" — and 0.10 to 0.15 moves zero. Separately, and not softened: the same measurement today reads 42.4%, a third run reads 77.8%, and this reading cannot separate corpus growth from 234 engine commits
Mutation score, 2026-09-13 What fraction of the mutants this tree admits the test suite catches, over a generated population rather than the frozen catalogue of historical defects 275 of 375 viable mutants killed, 73.3% [68.6%, 77.6%], on a uniform random sample of 400 drawn with seed 20260913 from a population of 8,150 mutants across 206 production files at commit c0ed888. The first killed-over-generated figure here: the catalogue reports a numerator over 76 hand-chosen defects and the two denominators are not commensurable. Split by operator the suite tests what a branch decides and not where it sits, negate-conditional 85% against conditional-boundary 49%. Worst large packages cmd/replay 65% (29% of the whole population), internal/tui 67%, internal/analysis 68%. Nine equivalent mutants are named and removed from the denominator; 100 of the 109 survivors were not examined, so the tree-wide equivalent rate is unknown rather than small. A sample, not a census: 4.91% of the population
upgrade verifies the hash, not the signature Whether replay upgrade checks the Sigstore signature on checksums.txt, which the release publishes It does not. install.sh checks it where cosign exists and upgrade never does, so the two install routes make different promises and nothing said so
Break causes, third reading The launch reading of the cause table, taken because the launch draft still carried a retracted one 833 breaks, 44.2M re-billed tokens over 1,950 transcripts. Third reading in nine days; the ranking flipped and kept moving (re-render 50.8 to 42.4 to 39.7). The SHAPE held all three times: TTL expiries rare and 6.8x larger than re-renders
Two builds, one corpus Whether the headline figures move with the binary, measured directly rather than inferred across days The same corpus reads 4.99% on v0.5.4 and 2.75% on the shipping build. The current build reads 27,067 more requests and prices six more models. The published 5% must be re-read or carry its build
v0.5.4 reproduces byte for byte Whether a published tag rebuilds to the same bytes, which RELEASE-CRITERIA called unverified All four platforms reproduce. The toolchain must be read from the binary, and the build must run in a CLONE: a linked worktree stamps (devel) and differs, which is why an earlier attempt wrongly concluded it does not reproduce
Seeded blame, 2026-09-11 Whether Replay names the event that actually broke the prefix, against ten breaks whose cause was seeded and sealed before the run 10 of 10 correct. The first file here to test the front-page claim rather than a property of the instrument: seeds the EVENT and not the symptom, since asserting "cache_read is 0 so it says prefix changed" restates the classifier's own rule. Found one limit and confirmed the response to it is right — systemChanged compares byte counts, so an edit preserving its length (2026-09-10 to 2026-09-11) moves the prefix and no count, and the tool widens to the cause covering both halves rather than naming one at random. A reader is told less than the truth, never something false. Two mutations, each landing on exactly its own case, show the trial can fail. Says nothing about whether naming the event helps anyone, which remains unmeasured
Embedding collision, 2026-09-15 Whether coding-agent prompts that provably ask different things fall inside the cosine-similarity thresholds semantic caches ship with Median cosine 0.9999 between two prompts differing by one integer. At 0.95, 86.9% of 3,187 constructed distinct pairs collide. Pre-registered before the data was read, including the rule that a near-zero result meant stopping. Representation collision only: no cache decided anything and no answer was checked. Pairs are constructed rather than observed, one embedder, n=1 operator
Contribution path, 2026-09-15 Whether a user following the README can get a corpus submission into the pool, exercised end to end rather than read from code go install ...@latest produces commit: unknown and is refused 400 by production; make build produces a real commit and is accepted 201 locally. The endpoint is correct and the first broken invariant is at the producer. Separately: rulesVersion was syntactically valid and semantically false, naming a document that priced almost none of the corpus, detectable only as $0.97 on $13,518. Neither defect fixed. Production pool unchanged at 1
Reconciliation protocol, pre-registered 2026-09-15 How much of a provider's own metered figure Replay's reconstruction from local artifacts can account for, and what the remainder is Establishes nothing. It is a protocol, registered before any provider artifact was retrieved and before any comparison was run, frozen at commit bada9918. Section 0 governs: neither side is ground truth, so the experiment measures agreement or discrepancy between two representations of billing and never correctness. Window, rules document, tolerance, six discrepancy classes and stop conditions are all fixed in advance. The likeliest registered outcome is a STOP: a flat-fee subscription allowance is not comparable to a list-price reconstruction, and "these two records cannot be reconciled at this granularity" will be published as a result rather than treated as a failure

Important

The calibration corpus is 1450 transcripts from 78 sessions on one machine, not the independent sessions the roadmap gate asks for. It is published because an under-powered measurement stated plainly is worth more than a claim with nothing behind it.

Contributing your own corpus takes one command and shares no paths, project names or content. See Contributing a calibration corpus.


Documentation index · Repository README