Skip to content

WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress - #1014

Draft
MauroToscano wants to merge 1348 commits into
mainfrom
noepoch/whir
Draft

MauroToscano wants to merge 1348 commits into
mainfrom
noepoch/whir

Conversation

@MauroToscano

@MauroToscano MauroToscano commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

One WHIR (multilinear) proof per block, with no epochs (the prove-and-retire / VADCOP shape). This is a second prover next to #1010's epoch-based one. #1010 stays the reference and this branch does not touch it. The branch starts from #1010's head f3d359998; compare against f3d359998 to see only this work.

Status: draft, work in progress. One binary (FAST 416, #1010's head 88b0d31 merged in, ABBA × 2): the block tree 23.20 s whole vs #1010's epoch tree 31.80 s, −8.60 s. Since then: 20.89 s whole at d9d0ac5 (FAST 424: base 16.76 + recursion 4.13), with the batched argue on by default, ops dropped after commit (M1), the lean walk and paged executor memory (E1/E4), and one prepared stack per group. e06dce5 added narrow trace storage (M4: same bytes, time-neutral, max RSS 1× 36.2 → 29.5 GiB, 4.13× now proves). 3cfffe2 added the KECCAK and ECSM split (§6: the same tables and proof bytes on every block measured so far, FAST 830). 01da99f uploads phase A's next group beside the current commit (prover-only, same bytes, base −0.60 s at 1×, FAST 832). 59e9890 lays out the streamed chunks on three threads by default (E3; prover-only, same bytes; whole block −0.48 s to 19.75 s and max RSS −1.51 GiB at 1×, FAST 838). 82f9046 re-hashes phase B's kept-top paths in parallel (prover-only, same bytes; phase B −0.96 s, whole block −0.82 s to 18.90 s at 1×, FAST 843 + 844 pooled). d169edd lets the tree's leaves execute while their artifacts are built (prover-only, same bytes; level 0 −0.24 s, whole block −0.24 s to 18.54 s at 1×, FAST 841 + 842 pooled). 70eee3e lays out the rest in waves and builds KECCAK_RND as its tables (W, lane i-m4b; same bytes; max RSS 88.9 → 79.5 GiB at 4.13×, base −0.14 s at 1×). fac261f lets the tree's first finished leaf execute during phase B (R2-ii; prover-only, same bytes; whole block −0.24 s to 18.39 s at 1×, measured before the merge with W, FAST 845 + 846 pooled). 365e3ab packs the finish's tables as they are built and commits them narrow (b2, lane i-m4b; same bytes; max RSS 78.8 → 52.1 GiB at 4.13×, 28.5 → 22.5 at 1×; base −0.36 s at 1×). 351a773 adds a device readout to the tree test (test-only), and the measurement posture moves to the card memory pool's code default, which retains freed memory: a baseline shift, not a prover change (whole block 17.91 → 17.27 s at 1×, FAST 848 + 849 pooled; every number above was taken under the old posture and stays as measured). 9a18ab5 makes the prover refuse a block over the group maximum as its groups close, with the verifier's own error (lane i-m4b; prover-only, same bytes, FAST 858). 54aae28 builds the recursion tree's node programs beside the tree, each level proving as its own nodes arrive (A; prover-only, same bytes; median whole block 242.69 → 233.80 s, −8.89 s, BIG 611; 2.66× recursion −1.22 s; 1× unchanged), and closes a latent card-permit re-entry that parked BIG 569. d107a91 proves each node of the tree as soon as its own children are proved (C; prover-only, same bytes; median recursion −1.73 s, BIG 612; 2.66× recursion −0.49 s, BIG 614 + 615; 1× unchanged). 0d61081 spills the held tables to disk when memory is short (lane i-m4b; prover-only, same bytes; #1013's store and auto policy, on by default; nothing spills at 1× or 4.13× by default, and the 1× time is unchanged within noise; BIG 616, ULTRA 006). acfd54a counts phase 4's BITWISE lookups in slices (#1013's c369843, ported by lane i-m4b; prover-only, same bytes; median whole-run VmRSS 108 → 89 GiB, phase A −9.5 s, BIG 619; 1× base −0.30 s, ULTRA 010). c55dec4 lets the tree's top execute each child's share as that child is proved, while a child is still unproved (B′; prover-only, same bytes; median recursion −0.47 s, BIG 613 / 620; 1× unchanged, where the top executes whole, ULTRA 009). 806c211 lets each node of the tree execute from its program as soon as it is emitted, with only its prove waiting for its artifacts (b′; prover-only, same bytes; 1× recursion −0.27 s, ULTRA 012; 2.66× −0.41 s, BIG 624; median unchanged). Head cbfa7d7 grinds the block's WHIR proofs at 18 bits instead of 20 and opens 114 queries instead of 112 (§7, a format change approved by Mauro 10-05; proven bits unchanged at 130.393; 1× whole block −0.27 s, ULTRA 019; median whole block −3.42 s, BIG 628). Every format change below but §6 (the port of #1013's S0a, gated FAST 830) is approved by Mauro (10-02; §7 on 10-05); the cryptography team reviews them at the end (CRYPTO-REVIEW.md rows W-0…W-6, W-8, W-9).

What it is

  • prover::block_whir::prove_block_whir / verify_block_whir. The whole block is one MultiProof:
    • every table is cut into instances of at most 2^21 rows, and KECCAK_RND into 2^16-row instances;
    • memory is the monolithic PAGE argument, with no local-to-global bookend and no cross-epoch proof.
  • The tables are packed into groups of at most 3 stacked polynomials (2^27 each).
    • Phase A commits each group and retires it: the codewords go, and only the top of each tree stays on the host (the bottom 4 levels are dropped).
    • The roots block: every root goes into the transcript, then (z, α, β) are drawn once.
    • Phase B proves each group on its own fork S_post ‖ g: upload again, argue (LogUp-GKR), re-encode with the NTT alone, open. The dropped tree levels are rebuilt from the queried cosets and checked against the kept nodes.
    • The bus is checked once, over every table of the block.
  • Streamed phase A. The executor runs in 2^20-cycle windows feeding the shared WindowedTraceBuilder (noepoch/windowed-builder, also used by STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013). Every full chunk of CPU, MEMW_R, MEMW_A, MEMW, LOAD, LT, SHIFT and STORE is laid out and committed as soon as it exists.
  • The builder is deterministic: on block 25368371 the whole-run build and two windowed builds under host contention give identical per-table digests (FAST 411; 129 tables row for row, the six HashMap-ordered tables as row multisets).

⚠ Format changes (all approved by Mauro, 10-02 (§7 on 10-05), but §6, the port of #1013's S0a; for the cryptography team's end review; ledger rows W-0…W-6, W-8, W-9 in CRYPTO-REVIEW.md)

The block is one WHIR proof (BlockWhirProof, one MultiProof) over all of the block's tables.

  • Phase A commits the tables in groups as the trace is built. Each group is at most 3 stacked polynomials of 2^27, and each group is committed and retired as it closes.
  • The challenges. The transcript then absorbs the statement, the groups' roots and the derived prepared roots, and draws z, α, β.
  • Phase B proves each group on its own fork of the transcript (S_post ‖ g).

Every parameter below is a verifier-side constant. None is read from the proof.

constant value meaning
BLOCK_GROUP_POLYS 3 a group of two or more tables stacks into at most 3 polynomials of 2^27
BLOCK_MAX_GROUPS 64 at most 64 groups
BLOCK_MAX_TABLE_VARS 27 no table stated over 2^27 rows
BLOCK_MAX_KECCAK_RND 2^12 at most 4096 KECCAK_RND tables
BLOCK_KECCAK_RND_MAX_VARS 16 each at most 2^16 rows
BLOCK_MAX_ECDAS 2^12 at most 4096 ECDAS tables
BLOCK_ECDAS_MAX_VARS 17 each at most 2^17 rows
BLOCK_MAX_KECCAK 2^12 at most 4096 KECCAK tables
BLOCK_KECCAK_MAX_VARS 18 each at most 2^18 rows
BLOCK_MAX_ECSM 2^12 at most 4096 ECSM tables
BLOCK_ECSM_MAX_VARS 17 each at most 2^17 rows
ArgueFormat::BATCHED bins of 2^26 input cells change 5 below
PRODUCTION_WHIR_GRIND_BITS 18 (Q 114) each WHIR chain's query grind; change 7 below
LEAF_PERMS_CAP, BLOCK_FAN_IN 279,000 permutations; 3 the recursion tree's leaf cap and fan-in

1. The group partition is in the statement

What. The streamed prover packs tables into groups in the order their chunks complete. It writes the partition (per group, its tables' indices) into the statement. The transcript absorbs it before any root.

Verifier. The verifier checks:

  • that the partition is exact (every table once);
  • that each group of two or more tables fits within 3 stacked polynomials under the verifier's own stack cap;
  • that there are at most 64 groups.

It then rebuilds every group's stack layout itself.

Security. Proven bits are unchanged.

  • Each WHIR chain is 130.393 bits, and the proof minimum under the campaign's min-over-phases accounting is 128.946.
  • The chain heights are capped by BLOCK_MAX_TABLE_VARS (change 3), so no partition can make a chain taller than 2^27.

Proof bytes vary from run to run. Arrival order depends on thread timing, and the hash-ordered tables' row order (LT, EQ, BYTEWISE, BRANCH, MUL, DVRM) follows a hash state that is random per process. Mauro, 10-02: "the proof not having the same bytes it's fine, they never had the same bytes anyways".

  • Byte-identity tests run with LAMBDA_VM_FIXED_TRACE_HASH=1 (fixed hash keys) and LAMBDA_VM_DETERMINISTIC_GRIND=1. Under both, two processes prove the same bytes (FAST 423).
  • Fixed keys let a crafted program steer many operations into one bucket, so they are a test setting, not production's.

2. Prepared openings, one stack per group (approved by Mauro, 10-02: "it's a standard technique, go ahead")

In short. This is #1010's existing prepared DECODE opening, carried into the block format. It is extended to the dense genesis pages and stacked per group.

  • Costs: +0.14 s of base and −0.57 s of recursion (4.71 → 4.14 s).
  • The verifier recomputes every derived root from the ELF.

What. The recursion guest cannot evaluate DECODE's five 2^20 preprocessed columns, nor the dense genesis pages (millions of rows), on its own. So each group's prepared tables (DECODE and the dense genesis pages chosen by genesis_stack_plan) have their leading preprocessed columns stacked into one commitment.

  • Both sides derive this commitment from the ELF and the partition. The verifier recomputes every derived root from the ELF: verify_block_whir, and verify_block_tree's plan. The prover's own roots only shortcut the prover's side.
  • Its roots are absorbed after the groups' roots, before z.
  • It is opened on the group's fork after the group's own opening, each table's block at that table's point.

Negatives, on a guest with two dense genesis pages. Each is refused by the host verifier and by the recursion leaf, and admitted when the openings are skipped (the mutation):

  • two groups' openings swapped;
  • a block opened at another table's point.

Wrongly committed stacks are refused at the roots block: two pages swapped, a page holding another page's columns, another program's DECODE.

Cost. One stack per group costs +0.14 s of base (FAST 419; one commitment per table cost +0.49 s) and saves 0.57 s of recursion, by taking 90 k permutations out of the leaves. BlockFormat::prepared = false turns the openings off, for measurement only.

3. ECDAS split; table heights capped

What.

  • ECDAS is split into tables of at most 2^17 rows, as KECCAK_RND is at 2^16 (split_ecdas).
  • A scalar multiplication may straddle two tables. Its steps chain only through the Ecdas bus, keyed by the call's timestamp and the step's (round, op).
  • The frame (the host verifier and the tree plan) refuses:
    • a KECCAK_RND table over 2^16 rows;
    • an ECDAS table over 2^17 rows;
    • any table over 2^27 rows (REV-JUDGE G2). A 2^30 table would otherwise make a 2^30 chain at 127.39 bits.

Why. At a median block, one ECDAS table (2^19) argues on a 12 GiB tree. At p90 (2^20) it stacks into five polynomials on a 24 GiB tree.

Gates. FAST 534 and 535: the negatives, 3 mutations caught.

4. Recursion leaf: the leaf cap (G3) and the inverse share (W1)

  • G3. The tree plan refuses a group over LEAF_PERMS_CAP. With no fixed leaf count, it adds leaves while the heaviest leaf is over the cap. The block's own partition is unchanged.
  • W1, on by default. Each WHIR leaf publishes its bus share Σ p/q as p · ediv(1, q), which has no satisfying assignment for q = 0 whatever p is. The former ediv(p, q) left the share free at p = q = 0 (≈ 2^-160 under GKR soundness).
    • The leaf gains one XALU op per table, so every WHIR tree id changes. The base proof is untouched.
    • Gate: FAST 536. LFM_WHIR_SHARE_INVERSE=0 keeps the former form.

5. Batched argue, on by default since d9d0ac5 (BlockFormat::argue = ArgueFormat::BATCHED; Mauro, 10-02: "batch the constraints")

What. Each group's tables are argued together on the group's fork:

  • one lockstep LogUp-GKR ladder per bin, the tables packed first-fit-decreasing under 2^26 input cells;
  • one front-loaded constraint sumcheck per group;
  • no claim reduction for unshifted tables (every VM table).

Verifier side. The verifier derives the bins from the statement's shapes and its own cap; the variant is never read from the proof. The proof carries one BatchedArgue per group in BlockWhirProof.argues, and proof.tables is empty under this format. The recursion leaves verify the same format (lane i-batch2, N-4).

Security at the block's measured inputs (D-BATCH §3.2's formulas; re-checked at FAST 424's inputs: |T| ≤ 60 tables a group, ≤ 51 trees a bin, ≤ 2^28.87 input cells, bus messages ≤ 204 elements, N_C ≤ 413, D_max 4, n ≤ 21; L = 300):

term bits
LogUp fractional 147.22
GKR step batching 185.34
GKR sumcheck round 190.42
GKR line 186.33
zerocheck point 187.61
constraint β (N_C − 1 = 412) 183.31
claim batching 184.52
constraint rounds 190.00

Every term is above the WHIR phase, so the proof minimum is unchanged: 128.946 under the campaign accounting (130.393 for the WHIR fold under the calculator of record).

Measured.

  • FAST 348: phase-B argue −0.685 s, whole block −0.385 s (≈ −0.65 s attributable).
  • FAST 424 at the flip (d9d0ac5, A B B A, one binary): base 17.56 → 16.76 s (−0.80), whole block 21.85 → 20.89 s (−0.96), recursion −0.16 s. Every proof verified; determinism within each format 2/2.
  • Peaks: BIG 402/403 found none raised at 1.0 / 1.31 / 1.80×. BIG 395 at 1.80× with ops dropped: per-table 62.39 against batched 63.12 GiB (+0.72).

The knob. ArgueFormat::PerTable (one argument per table, the format before) stays measurable, with BLOCK_WHIR_ARGUE=per-table in the real-block tests. At the per-table format, the MultiProof, the prepared openings and the partition are digest-equal to the head before the batched merge (BIG 393).

6. KECCAK and ECSM split; their heights capped (the port of #1013's S0a; landed at 3cfffe2; approved by the lead under Mauro's any-block target, Mauro to confirm)

What.

  • KECCAK is split into tables of at most 2^18 rows and ECSM into tables of at most 2^17 (split_keccak, split_ecsm), as ECDAS is at 2^17.
  • A row of either is one whole call (a permutation, a scalar multiplication), so a cut falls between calls. A call reaches its rounds (KECCAK_RND), its steps and scalar bits (ECDAS) and its memory only through buses keyed by its timestamp.
  • The frame takes up to 2^12 tables of each and refuses a KECCAK table over 2^18 rows or an ECSM table over 2^17. Every other verifier keeps both to one table.

Why. A keccak-heavy block at the gas limit makes about 2^21 permutation calls. One KECCAK table that tall stacks into eight polynomials of 2^27, against a group budget of three. At its cap each table fits one polynomial.

Bytes. Every block measured so far has KECCAK ≤ 2^17 rows and ECSM ≤ 2^11, so it builds the same tables and the same proof (FAST 830: digest 07d1bd43… unchanged on block 25368371).

Gates. FAST 830: KECCAK and ECSM forced into 4 tables each on block 25368371 prove and verify, base and tree; the count and height negatives; mutation C caught.

7. WHIR grinding 20 → 18 bits, queries 112 → 114 (landed at cbfa7d7; approved by Mauro, 10-05)

What.

  • Each WHIR chain grinds 18 bits instead of 20 before each round's queries, and opens 114 queries instead of 112 to buy the two bits back. At every block height (15 to 27 variables) Q = num_queries(2, rounds, 128, 18).

  • One site, multilinear_prove::chain_config_under. Every side of the block reaches it through BlockFormat::chain_config:

    • the prover's group chains and prepared openings;
    • verify_block_whir;
    • the plan the recursion leaves' in-guest verifier is emitted from.

    The statement absorbs Q and the grind bits.

  • Scope: the block's WHIR proofs only. The recursion tree's own leaf, node and top proofs are STARK proofs (FRI) and keep their own grind.

  • It ports WHIR recursion on GPU (RPX) with ZisK-style proof formats: block 25368371 in 32.62 s #1010's grind 18 (landed there at 2687e48). Opt-out: LAMBDA_VM_ZF_WHIR_GRIND_BITS=20, which reproduces the previous proofs byte for byte.

Security. The calculator of record: Johnson bound, η = 1/300, the fold proof of work unground.

  • The binding WHIR phase is the unground fold at 27 variables, 130.393 bits at both settings. Neither Q nor the query grind reaches it.
  • The query phases move from 130.926 to 130.907 bits (15 to 27 variables).
  • The argue terms and the tree's FRI phases do not move. The proof minimum stays 128.946 under the campaign accounting.
  • The grind bits are a verifier constant:
    • a block ground at 18 bits is refused by a verifier at 20, and the reverse;
    • a leaf emitted at 20 refuses an 18-bit block;
    • at chain level, the host and the machine refuse fewer grind bits at the same Q.

Measured behind the knob, on c55dec4, one binary per box:

  • 1× (ULTRA 019, 8 + 8 runs after one warm-up):
    • phase B −0.344 s (t −12.9), from the grind falling 0.577 → 0.214 s;
    • whole block −0.27 s (18.71 → 18.44 s);
    • recursion −0.01 s, i.e. unchanged.
  • Median 25475471 at i-m4b's cap-128 exploration head (BIG 628, A B B A):
    • phase B −3.48 s;
    • whole block −3.42 s (226.51 → 223.09 s);
    • recursion −0.12 s;
    • both arms verified.
  • The leaves' cost. The leaves verify two more queries per chain: +1,041 permutations per three-polynomial group. On these two blocks no leaf table crossed a power of two, and the leaf counts did not change (3 at 1×, 23 at the median).

Bytes. The proof bytes change. Under the fixed trace hash and deterministic grind, the 1× top proof's digest is b9b4826d… at the default and 0da6fea9… under =20. Both trees are accepted (ULTRA 015, at this head).

Known limit. The leaf cap (279,000 permutations) sits above 2^18 LFM_HASH rows, and 18 bits moves every leaf about 1.7 % closer to that doubling.

  • On the median, the heaviest leaf's LFM_HASH is 256,549 of 262,144 rows: 2.1 % headroom.
  • A heavier block can cross it. That doubles the table, which costs that leaf's prove; it does not affect correctness.

Gates.

  • Laptop: zf_format::; the chain refusals and the legacy-bytes KAT; the pins the flip moved; the host block refusal.
  • ULTRA 019: the leaf refusal (box tier); the opt-out byte-identical to c55dec4.
  • ULTRA 015: the landing sha's identity, above.

One-binary A/B: the block tree vs #1010's epoch tree (FAST 416, noepoch/whir @ 454dad6)

Eight arms E B B E E B B E on the same binary (md5 checked after every arm):

whole block base last stage (E: root; B: top) whole − last stage harness verifies
E, #1010's epoch tree (n = 4) 31.80 s 22.48 s 1.05 s 30.75 s 0.92 s, in its timed path
B, the block tree (n = 4) 23.20 s 18.50 s 1.35 s 21.86 s 2.97 s, off the clock
B − E −8.60 s −3.98 s −8.89 s

Base, block 25368371 on FAST (each run against #1010's epoch base on the same binary)

run change block base
FAST 360 W1: non-streamed block proof 29.9 s
FAST 364 first streamed version 25.96 s
FAST 365 chunk generation off the walk thread 24.45 s
FAST 366 windowed executor 23.55 s
FAST 367 walker/accumulator split 22.68 s
FAST 368 walked windows kept whole, concatenated at finish 21.98 s
FAST 369 BITWISE counted per window, KECCAK_RND split by rows 20.73 s (epoch base 25.71 s, −4.98 s)
  • Phase B: 9.96 s. Host peak: 43 GiB.
  • Every block proof the windowed builder produced verified: 14 of 14 (FAST 363–369). Non-windowed: 6 of 6 (FAST 360–362).

Recursion (W3): proved and verified on block 25368371 (FAST 413)

  • Leaves are cut along groups: a group's opening covers all of its tables, so a group is the smallest unit a leaf can verify alone. The tree plan (lfm::whir_block::WhirBlockPlan) runs the host verifier's own statement checks, derives every shape from the AIR and every prepared root from the program, and never reads a proof. It prices each group in permutations and partitions the groups over the leaves (heaviest first, onto the least-loaded leaf).
  • A leaf:
  • Nodes are the STARK block's (block_node, shared byte for byte with STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013). They check the id, the state and the output equal across children, add the shares, and the top asserts zero.
  • Final check: verify_block_tree(elf, statement, top) derives the top program and checks the top proof against it. Every preset is pinned inside: the base's options, the block format, the tree's options, the leaf rule and the fan-in.
  • Tests:
    • every new check has a negative, and each negative a mutation showing that check is the one refusing;
    • the leaf's challenges equal the host's;
    • a tree over another partition is refused at the final check.
block 25368371 FAST 413: first build, serial FAST 414: pipelined (mean of 2 arms)
base (streamed, 9 groups, 27 chains, 12 prepared tables) 21.00 s 20.98 s
plan + leaf emission 0.23 s + 2.38 s, after the base inside the base (done at 12.2 s of 21.0)
leaves (3) 7.04 s, one at a time 3.36 s, three at a time
nodes 2 levels: 3.41 + 1.91 s one top over 3 children: 1.36 s
recursion 12.36 s 4.75 s
whole block 33.37 s 25.72 s
tree verify (derives every program from the ELF and the statement) accepted, 4.97 s accepted, 2.95 s

Every proof was verified by the harness off the clock. The base is unchanged by the emission running beside phase B (20.98 s vs 21.18 s in the arms without the change).

Next:

  • one prepared stack per group (group 8 carries 12 prepared chains);
  • the base levers: build_traces over segmented lists, and KECCAK_RND streamed.

Base levers after the A/B (FAST 417–422, each S/P on one binary)

lever base status
builder: windows concatenated in parallel −0.52 s (417), −0.21 s (419, replication in band) on by default
KECCAK_RND chunks streamed −0.14 s, no effect (this block's keccak work is at its end) off by default
LT's MEMW-derived ops streamed per window (FAST 420) +1.27 s, regression: the layout thread binds (9 more LT chunks, layout busy +2.47 s); LT heights deterministic across runs off by default
streamed chunks laid out on 3 threads, prepared columns after the groups (FAST 421, 422; BIG 390) −0.54 s (421), −0.22 s (422 replication); but the extra host memory scales with the block: +2.2 / +4.2 / +8.7 GiB max RSS at 1.0 / 1.3 / 1.8× (BIG 390) reverted at 11e2de5; on again since 59e9890, bounded (E3)
the rest of the run packed as it is laid out, AIR order (FAST 422) +0.21 s, regression: groups 6–8 still close together off

Memory (D-MEMORY, with i-mem / i-mem2)

stage effect status
M1: the builder drops each streamed chunk's ops as it leaves (drop_streamed_ops) FAST 501: 1× base −0.74 s, peak 47.4 → 36.6 GiB; the 1.20× block 25490321 now proves (49.2 GiB, OOM before). BIG 391 lean gate on the exact head: same statement shape kept vs dropped, both verify, 1.20× at 49.05 GiB on (f4aafea)
E1 lean walk + E4 paged executor memory (D-EXEC, lane i-exec) base −0.45 s (FAST 602); median windows phase 42.4 → 25.2 s (harness, FAST 599) on (5b4e6b4)
E1 v2: smaller per-step records (CPU op 96 B, MEMW_A ops as 48 B rows; D-EXEC, lane i-exec) byte-identical: proof digest equal to d9d0ac5's under the fixed trace hash + deterministic grind (FAST 426) on (773f134)
E3: layout workers bounded to K + 1 chunks unpacked (layout_ahead, D-EXEC) BIG 392: max RSS vs workers 0 +0.23 / −0.06 / +0.16 GiB at 1.0 / 1.3 / 1.8× (flat). FAST 425: no effect on time (whole 20.91 s both), since phase A no longer waits on layout on (3 workers since 59e9890)
one prepared stack per group prepared cost +0.49 → +0.14 s; recursion −0.57 s in
W: the rest laid out in ≤ 2 GiB waves; KECCAK_RND built as its 2^16-row tables at the finish (lane i-m4b) BIG 562: max RSS 88.88 → 79.54 GiB at 4.13×, 29.84 → 27.10 at 1×; FAST 852 + 853: base −0.14 s on (70eee3e)
b2: the finish packs each table as it builds it; phase A commits those tables narrow, never transposing them (lane i-m4b) BIG 563: max RSS 78.84 → 52.10 GiB at 4.13× (89.15 with W and b2 off), 28.53 → 22.45 at 1×; FAST 855 + 856: base −0.36 s on (365e3ab)

Narrow trace storage (M4, lane i-m4)

Prover-only; the proof bytes are the same. The proof digest equals the head's before M4 (07d1bd43…) under the fixed trace hash + deterministic grind (FAST 820).

  • What changes: between phase A's commit and phase B, each table of at least 2^16 cells is kept on the host packed at the width its values need, 1, 2, 4 or 8 bytes a column (BlockOptions::narrow, production Narrowing::CARD).
    • The pack runs on the card from the group's resident store, on a thread of its own, after the group's commit.
    • Phase B uploads the packed table and widens it on the card. A host reader widens on the host.
    • A wrong width map is refused by the kept-tree check (test plus mutation).
  • Time-neutral: whole block −0.04 s at 1×.

Max RSS, wide → narrow (BIG 560; the traces themselves shrink about 4×, e.g. 24.89 → 6.20 GiB at 1×):

block wide narrow
25368371 (1.0×) 36.23 GiB 29.47 GiB
25453112 (1.8×) 63.52 GiB 46.10 GiB
25482821 (2.66×) 80.80 GiB 52.83 GiB
25410821 (4.13×) — 86.93 GiB, proves and verifies

Before M4, #1014 ran out of memory at 3.03× (i-m4); the narrow line puts the 120.7 GiB edge near 5.5–6×.

Phase A uploads the next group beside the commit (3′, lane i-noepoch-w2)

Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 832; the commits, their order and their bytes are unchanged).

  • What changes: phase A used to upload, commit and retire each group in turn, so the card waited on every group's upload. Now the current group's commit runs on a thread of its own while phase A's thread takes the next group and uploads its columns (BlockOptions::upload_ahead, on by default; BLOCK_WHIR_UPLOAD_AHEAD=0 is the control).
    • The upload starts only after the commit has asked the card for its room, so it takes what the ledger has left; a store the ledger refuses goes up after the commit, as before (no store or commit room was refused in any run).
  • Time (FAST 832, A B B A on one binary): base −0.60 s, phase A −0.66 s, whole block −0.61 s; 46 % of the upload seconds hidden. Groups 4–5 cannot hide theirs: the inline layout thread closes them after the previous commit ends, which is the next term (E3, next section: on since 59e9890).
  • Memory: one more group's columns held at the peak — the previous group's wide columns and its packed bytes wait for its pack's install while the current group commits (FAST 834, BLOCK MEM terms). At most one group (≤ ≈ 3.8 GiB), constant per group, not growing with the block; max RSS +1.8 to +2.3 GiB at 1× (FAST 832, 833, 834).

Streamed chunks laid out on three threads (E3 on by default, lane i-noepoch-w2)

Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 838 and 835; the groups, their order and the packing are unchanged).

  • What changes: layout_workers 0 → 3 in BlockOptions::production (59e9890). The streamed chunks are laid out on three threads, with at most K + 1 = 3 of them unpacked at a time (layout_ahead, E3), instead of on the builder's one layout thread. BLOCK_WHIR_LAYOUT_WORKERS=0 (the inline layout, 01da99f's) is the control.
  • Why it pays now: with phase A uploading ahead (3′), the inline layout closed groups 4–5 only after the previous commit had ended, so the card waited on them; three workers close them in time.
  • Time (FAST 838, A B B A on two binaries, A = 01da99f, B = 59e9890): base −0.54 s (t −14.7), phase A −0.58 s, whole block −0.48 s (t −12.4; 20.23 → 19.75 s); phase B +0.06 s (one job, below what a single job resolves). FAST 835 (one binary, workers 0 vs 3) read whole −1.00 s.
  • Memory: max RSS −1.51 GiB at 1× (31.99 → 30.48 GiB, FAST 838; 835 read −0.68 GiB). The bound is what the earlier revert lacked: unbounded, the workers' host memory grew with the block (BIG 390); bounded, it stayed flat at 1.0 / 1.3 / 1.8× (BIG 392, measured before M4 and 3′, not re-run on this head).
  • 59e9890 also caps a readout stamp (an upload's "paid" seconds at the upload's own length); no proving change.

Phase B's kept-top paths re-hashed in parallel (lever 1, lane i-noepoch-w2)

Prover-only; the proof bytes are the same (digest 07d1bd43… at the default, FAST 840 at 733557a — 82f9046 adds only the knob; the paths, their order and the refusal are unchanged).

  • What changes: phase B serves each revived commitment's first-round paths from the kept tree top. The leaves under each queried block are gathered from the card, re-hashed on the host, and the block's subtree root is checked against the kept node. That re-hash ran one block at a time on the prover's thread while the card waited: 27 gaps of ≈ 37 ms, 1.01 s at 1× (nsys timeline, FAST 839). The blocks are now re-hashed in parallel and kept in block order; the root check still refuses before any path is returned. BLOCK_WHIR_REHASH_SERIAL=1 restores the serial re-hash (a measurement knob).
  • Time (FAST 843 + 844 pooled, 6 + 6 + 6 runs: one binary with the knob, serial S vs parallel P, plus 59e9890's binary as the reference R): re-hash 1.01 → 0.07 s; phase B −0.96 s (P − R; P − S −0.93 s, t −30.5); whole block −0.82 s (P − R, t −5.5; 19.72 → 18.90 s); phase A unchanged (this build vs R +0.09 s, t +1.0); max RSS unchanged.
  • FAST 840 (two binaries, 2 + 2 runs) read phase A +0.71 s, from one run that stalled in the layout; the code runs only in phase B, and the one-binary check shows no phase-A effect.
  • Also in: BLOCK OPEN SPLIT (with LAMBDA_VM_BASE_SPLIT=1), phase B's openings by host stage and the kept-top gather and re-hash seconds.

The tree's leaves execute while their artifacts are built (R2-i, lane i-noepoch-w2)

Prover-only; the proofs are the same (the top proof's digest equal with and without it, 0da6fea9…, under the fixed trace hash + deterministic grind, FAST 841).

  • What changes: WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's tree is proved by the W3 harness (prove_tree_pipelined, the way a prover would run it; there is no production tree driver yet). Its level 0 built all three leaves' artifacts first (three serial card holds, 0.24 s at 1×) and only then let the leaves execute and fill their traces (0.46 s until the first was ready) while the card sat idle (nsys timeline, FAST 839). Execution and fill do not read the artifacts; only the prove does. lfm_prove is now cut into its host half (lfm_execute_and_fill) and its card half (LfmFilled::prove, which asserts the artifacts' hasher is the one the traces were filled for), in the same order; the harness builds the artifacts on a thread of their own and each leaf executes and fills beside them, waiting for its artifacts only to prove. W3_EXEC_BESIDE_ARTIFACTS=0 is the control (the order before).
  • Time (FAST 841 + 842 pooled, 8 + 8 runs on one binary): level 0 −0.24 s (t −17.8), recursion −0.23 s, whole block −0.24 s (t −5.2; 18.78 → 18.54 s; per job −0.29 / −0.20); base and max RSS unchanged.
  • The first leaf now takes the card 0.48 s after the tree starts, not 0.76 s (one run of each arm, FAST 841's card-hold trace).

A second whole-block field. The W3 readout prints, on the clock between the base and the tree, the block report, a second block frame for the LT heights and the group tables: 0.13–0.14 s at 1× that a prover would not spend. "Whole block" keeps its meaning (every number above includes them); from d169edd on, W3 RECURSION also prints "whole excl. harness readouts" beside it (at d169edd, FAST 841 + 842 pooled: whole 18.54 s, whole excl. harness readouts 18.40 s).

The rest laid out in waves; KECCAK_RND built as its tables (W, lane i-m4b)

Prover-only; the proof bytes are the same (digest 07d1bd43… under the fixed trace hash + deterministic grind: FAST 854 at 0bcc1fc, and FAST 855's W arm on this head's code).

  • What changes:
    • The tables the finish builds (the "rest") were turned into columns all at once, and a table's column copy exists before its rows are freed. So for a few seconds most of the rest was held twice, and WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's memory log put the 1× and 4.13× peaks there (BIG 561). The rest is now laid out in AIR order, in waves of at most 2 GiB of rows, each wave in parallel (BlockOptions::rest_layout_bytes). The tables, their order and the groups are the same (test).
    • The finish built KECCAK_RND as one table of 1,480 columns, which split_keccak_rnd then copied into its 2^16-row tables. At 4.13× that copy was a second peak of the same height. The finish now builds the 2^16-row tables directly (BlockOptions::finish_keccak_rnd_chunks), the same tables as the split's (test), and hands none out during the windows, which would change the groups.
    • BLOCK_WHIR_REST_LAYOUT=all and BLOCK_WHIR_KR_FINISH_CHUNKS=0 restore the old behaviour.
  • Memory (BIG 562, one binary at 0bcc1fc, which 70eee3e merges with WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's later commits; A = both off, W = this; three layout workers; jemalloc never purges, the record posture):
block A W
25368371 (1.0×) 29.84 GiB 27.10 GiB
25482821 (2.66×) — 49.02 GiB
25410821 (4.13×) 88.88 GiB 79.54 GiB
  • The live peak (jemalloc active) at 4.13× fell 72.9 → 63.7 GiB and now sits at the finish's end.
  • Not as pre-registered, on magnitude: the bands were ≤ 76 GiB and a ≥ 10 GiB drop at 4.13×, and ≤ 27 GiB at 1×. The resident-set gain came from the KECCAK_RND copy alone. The waves bound the live column copies, but the resident set still grows ≈ 16 GiB during the layout, because the freed row buffers stay with the allocator while the column copies take new pages. That growth and the finish's 8-byte tables are the next cut's target (finish packing, being measured).
  • On the max-RSS line, the 120.7 GiB edge moves from ≈ 5.9× to ≈ 6.4–6.7× (a linear estimate).
  • Time (FAST 852 + 853 pooled, one binary per job, three layout workers): base −0.14 s against both off, phase A −0.12 s.
    • The time comes from the finish ending without the KECCAK_RND copy.
    • The waves add ≈ 0.3 s to the rest's layout, which phase A does not wait on.
    • Packing the rest as it is laid out (pack_rest_as_laid_out) was re-measured with the waves: no gain (+0.01 s), so it stays off.
  • Also in:
    • the block memory log, LAMBDA_VM_BLOCK_MEMLOG=1 (off by default). It prints the host memory term by term every half second and at each phase mark, with the line at jemalloc's active peak and each arena's bytes (those two in the lib tests, which install jemalloc).
    • the readouts BLOCK REST TABLES (each rest table's layout end and group) and each group's commit start (from@).

The tree's first finished leaf executes during phase B (R2-ii, lane i-noepoch-w2)

Prover-only; the proofs are the same (base digest 07d1bd43… and the top proof's digest equal with and without it, 0da6fea9… — FAST 841's — under the fixed trace hash + deterministic grind: FAST 845, and FAST 847 again on fac261f, the merge with W).

  • What changes: phase B now hands each group's share of the proof (its argue, its opening, its prepared opening) to an observer as that group's opening ends (block_prove_on_forks_observed; the old entry point passes a no-op, and the proof is the same with any observer). A leaf's arena is built from its groups' words alone (group_arena_words + leaf_arena; block_leaf_arena is built from them). The W3 harness turns the groups into words as they arrive; the first leaf whose groups are all opened while another group is still to come (leaf 0 = groups 1, 5, 6 at 1×, done after group 6) executes and fills its traces on a thread of its own, about 1.5 s before the base ends, and the tree picks it up. Off the clock its arena is checked against the finished proof's. W3_LEAF_DURING_PHASE_B=0 is the control.
  • Time (FAST 845 + 846 pooled, 8 + 8 runs on one binary, at c5d9cec before the merge with W): whole block −0.24 s (18.63 → 18.39 s; per job −0.21 / −0.28); phase B +0.02 s; max RSS unchanged. The card's wait before the first leaf's prove fell from 0.23 s to 0.003 s; level 0 gained 0.15 s of it, since the first prove now runs beside the other two leaves' execution and fill.
  • Memory: the host's high-water is set in phase A at every size (1×: 25.5 GiB active at its peak, 7.2 GiB active by group 6's end with 31.0 GiB resident; 4.13×: 63.7 GiB at the finish's end, i-m4b), so one leaf's traces during phase B's tail fit in what phase A already holds.

The finish's tables packed as they are built (b2, lane i-m4b)

Prover-only; the proof bytes are the same (digest 07d1bd43… under the fixed trace hash + deterministic grind at both arms: FAST 855 at 0cebe16, and FAST 857 again on 365e3ab, the merge with R2-ii).

  • What changes:
    • After W, the finish still built its tables 8 bytes per cell, and phase A transposed each one into columns before packing it on the card. At 4.13× that layout took 24.6 GiB of wide tables and grew the resident set by ≈ 16 GiB (BIG 562).
    • The finish now packs each table into the narrow layout as soon as it is built (BlockOptions::pack_finished). KECCAK_RND is packed in waves of four 2^16-row tables. Phase A takes those tables as narrow columns and uploads them as they are, so the card packs only the groups' other tables. Nothing packed is transposed.
    • KECCAK, ECSM and ECDAS stay wide (BLOCK REST PACKED counts the rest: 187 of 191 tables at 4.13×).
    • A packed table's preprocessed columns are checked word for word against the AIR's before it is committed; a wrong one is refused (test, with a mutation).
    • BLOCK_WHIR_PACK_FINISHED=0 restores W.
  • Memory (BIG 563, one binary at 0cebe16; O = W and b2 off, A = W, B = W + b2; the memory log on; jemalloc never purges):
block O A (W) B (W + b2)
25368371 (1.0×) — 28.53 GiB 22.45 GiB
25482821 (2.66×) — — 37.64 GiB
25410821 (4.13×) 89.15 GiB 78.84 GiB 52.10 GiB
  • At 4.13× the live peak (jemalloc active) fell 63.7 → 49.7 GiB. The resident set now sits within 2 GiB of it (retention +2.0 GiB, from +14.7), and the rest's layout adds +1.15 GiB to it instead of +16.1. Every pre-registered row is in but one: the 1× live peak, at 21.0 GiB against a band of ≤ 19, because at 1× it falls in the middle of the finish, where the commit pipeline's in-flight copies bind.
  • On the max-RSS line (52.10 GiB at 4.13×, 12.82 + 9.47 GiB per 1×), the 120.7 GiB edge moves from ≈ 6.4–6.7× to ≈ 11.1–11.4× (a linear estimate for the base alone; the median block is 9.78×).
  • Next in line at the 4.13× peak: the committed groups held narrow for phase B (18.5 GiB, ≈ 5.9 GiB per 1×), then the finish's tables still being built (16.5 GiB).
  • Time (FAST 855 + 856 pooled, 8 + 8 runs on one binary): base −0.36 s (t −6.2; per job −0.40 / −0.32), all of it in phase A: the rest is laid out in 0.13 s instead of 0.57.

Measurement posture: the card's pool now retains freed memory (baseline shift)

Not a prover change; the proofs are the same (the top proof's digest 0da6fea9… under both postures, fixed trace hash + deterministic grind, FAST 848; 351a773 adds only a test readout).

  • What the env set, and why: every WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014 box run (the canonical env since FAST 416) set LAMBDA_VM_MEMPOOL_RELEASE_MB=0. The card's stream-ordered memory pool then hands its freed blocks back to the driver at each synchronize, so a VRAM sampler's total − free reads the live working set. The code default is to retain every freed block (DEFAULT_MEMPOOL_RELEASE_THRESHOLD_BYTES = u64::MAX in math-cuda), and that is what a prover runs.
  • What release-0 cost: each synchronize paid for the release, and the next large allocation mapped the memory again. In FAST 839's trace, the time spent inside cuStreamSynchronize and cuMemAllocAsync while the card was idle was ≈ 0.25 s in phase A, 0.42 s in phase B (0.26 s of it in the encode: one 18–33 ms gap per group) and 0.34 s in the tree.
  • Measured (FAST 848 + 849, one binary at 351a773, release-0 vs the default, 8 + 8 runs): whole block 17.91 → 17.27 s, −0.64 s (t −6.0). Phase B −0.44 s (t −28, against the 0.42 s above); phase A −0.07 s (t −0.7; phase A's runs scatter); recursion −0.12 s (level 0 −0.09 s); base −0.51 s; max RSS −0.16 GiB.
  • Gates, every run: the tree verified; the log's posture line matched its arm; no host fallback, device commit error, decline or ledger refusal; the card's peaks 24.15 GiB (the ledger's reservation) and 22.8 GiB (the pool's live high-water), read from the ledger and the pool, not from total − free (the new W3 DEVICE line).
  • From 351a773 on, WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014's numbers are taken at the code default. Every earlier number in this description was taken under release-0 and stays as measured. VRAM figures come from the ledger or the pool's high-water, never from total − free.

The prover refuses a partition over the group maximum (lane i-m4b)

Prover-only; the proof bytes are the same (digest 07d1bd43…, FAST 858). The prover now refuses a block over BlockFormat::max_groups as its groups close, with the verifier's own InvalidTableCounts error, instead of proving a block the verifier refuses. The median block (90 groups against the cap of 64) spent its whole ≈ 190 s base on such a proof (BIG 564). A test covers both prover paths, with a mutation for each check.

The tree's nodes built beside it, each level proving as its nodes arrive (A, lane i-noepoch-w2)

Prover-only (the W3 harness); the proofs are the same (the top proof's digest 0da6fea9… — FAST 841's — with and without it, under the fixed trace hash + deterministic grind: FAST 889, and FAST 891 again after the permit fix below).

  • What changes:
    • The W3 tree (prove_tree_pipelined) built every node program and its artifacts on one thread, level after level, and level 0 joined that thread before any node proved, so level 1 waited for the whole tree's builds. At the median block, 32 s of serial node builds after the leaves' artifacts held level 0 open 10.7 s past its last leaf proof, and the card sat idle 11.4 s before level 1 (BIG 565 / 568's card trace). At 2.66× the wait was 1.5–2.1 s (FAST 888).
    • Now a builder runs beside the whole tree and puts each node's program and artifacts in a slot of its own as soon as they exist. A level's programs are emitted together on a pool of W3_EMIT_THREADS threads (default 4, the builder's own: on the global pool a prover's join can steal an emission and leave the card idle, STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013's FAST 454), and the builder's thread builds each one's artifacts as it arrives. Each node proves once its children are proved and its slot is filled. A builder that stops fails every empty slot. W3_NODE_PIPE=0 is the control.
    • This is STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013's slot pipeline (i-tree's per-node emission, −7.32 s at the median there) in the W3 harness, with three differences: W3's Published<T> serves as the slot; there is no per-group early emission (W3 builds every leaf's artifacts up front, so a level's programs are emitted together); and there is no host-only thread marking (WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014 has none).
    • The level builder is generic (build_levels, an emit on the pool and a finish on the builder's thread). A toy tree at the median's shape puts the serial builder's programs in their slots whatever order a pool finishes them in, and finishes every node off the rayon workers (test; finishing on a worker, or publishing in arrival order, fails it). A failed build fails every slot above it without a hang (test).
  • Time:
    • Median (BIG 611, block 25475471 at i-m4b's cap-128 exploration head with these commits, P O O P on one binary): whole block 242.69 → 233.80 s, −8.89 s, every pipeline run below every control run (the estimate from BIG 565 / 568 was −10.6 s). Recursion 52.85 → 43.89 s; level 1's start gap 12.52 → 1.73 s (card idle 11.4 → 1.5 s). Peak VRAM 24274 → 24242 MiB and VmRSS 112.44 → 111.95 GiB: no growth (the gates).
    • 2.66× (FAST 889 + 890 pooled, one binary, 4 + 4 runs): recursion −1.22 s (t −5.2; whole block −1.21 s). Level 1's start gap (the last leaf proof's end to level 1's first prove on the card) fell from 1.97 to 0.69 s. After the permit fix: −1.19 s (t −6.5, FAST 891).
    • 1× (8 + 8 runs): recursion +0.007 s (t +0.3). At 1× the tree has one node, built in the same 1.8 s either way, so there is nothing to take.
  • What is left: a level still starts only after its children's proofs and its node's own execute + fill (0.45 s at 1×, 0.63 s at 2.66×; at the median the card sits idle 0.9–1.1 s at each later level's start). That start term is the next one; it is not addressed here.
  • Part of this landing: the card permit's latent re-entry, closed (lane i-m4b's diagnosis and repro). W3 built its leaves' artifacts as rayon jobs that each held the card around a build that uses rayon. A holder that waits inside rayon runs queued jobs on its own thread, so a sibling build could take the card a second time on that thread: the permit's assert fired and BIG 569's run parked. Which run trips it was the scheduler's choice (568 ran clean on the same code). Now:
    • the armed permit refuses a hold on a rayon worker before it takes the card, so any hold inside a rayon job fails at once instead of by luck (i-m4b's repro test now asserts the refusal; a plain thread holding the card while rayon builds under it is tested too);
    • W3 builds the leaves' artifacts one after another on their own thread (each held the card throughout anyway);
    • the node pipeline builds artifacts only on the builder's plain thread;
    • the harness disarms the permit before its off-the-clock verify, whose WhirBlockPlan::programs builds as rayon jobs.
  • Also in: W3_NODE_TIMES=1 (off by default). It prints, per program, when it was built, taken and proved; its prove's split (execute, fill, the prove net of the card wait, the wait); and each node level's start against the level below. The stamps are taken either way and cost a handful of Instant reads.
  • Toward a production tree driver: WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014 has none, and prove_tree_pipelined is its de-facto driver. With the slots it now has a driver's shape: leaves from phase B, nodes from slots, levels as their nodes and children arrive. Moving it out of the test harness is a landing item for STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013 and WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014 both.

Each node proved as soon as its own children are (C, lane i-noepoch-w2)

Prover-only (the W3 harness); the proofs are the same (the 1× top proof's digest with and without it, under the fixed trace hash + deterministic grind, both trees accepted: 0da6fea9…, BIG 614 — the digest FAST 841 / 889 / 891 printed).

  • What changes:
    • After (A), a level still started only once its whole level below was proved, and then ran its first node's execute + fill with the card idle. At the median the card sat idle 1.4–1.6 s before level 1 and 0.9–1.1 s before each later level (BIG 611).
    • Now one pool of siblings workers proves every program of the tree, leaves and nodes, taking them in topological order (the leaves, then each node level) and publishing each proof into a slot of its own. A node waits only for its own children's slots, so it executes, fills and proves while the rest of the level below still proves. W3_DATAFLOW=0 is the control: each node waits for its whole level below (tested never to start early).
    • Workers that take programs in this order reach a node only after every program before it has been taken, so the earliest unfinished program never waits on one not yet taken: no deadlock at any number of workers (test; taking the nodes first deadlocks it). The siblings bound, and so the memory in flight, is unchanged. A failed or panicking prove fails every program above it without a hang (test).
    • A level's W3 LEVEL wall now runs from the level below's last proof to its own.
  • Time (one binary, arms alternated on one box):
    • Median (BIG 612, block 25475471, F O O F): recursion 43.78 → 42.05 s, −1.73 s, every dataflow run below every control run. The card's idle at level 1's start fell 1.4 → 0.0 s, and at level 2's 0.93 → 0.0 s, in both runs. Whole block −2.60 s (234.15 → 231.55 s), but −0.87 s of that is the base, which the knob cannot reach (it is read after the base): recursion is the number. VRAM flat (24306 → 24274 MiB); the tree phase's VmRSS max 120.05 GiB (one control run) / 112.07 GiB.
    • 2.66× (BIG 614 + 615 pooled, 6 + 6 runs): recursion 12.51 → 12.02 s, −0.49 s (t −4.5). Whole block −0.13 s (t −0.8): the base read +0.37 s (t +2.1), all of it in phase A's host front (executor and trace build), before the knob is read. Phase B was equal within 0.02 s, so nothing moved out of it. At 1× the base moved +0.04 s.
    • 1× (8 + 8 runs): recursion −0.02 s (t −0.5; per-run sd 0.08 s). At 1× the tree has one node, so there is nothing to take.
  • Mechanism, as registered (2.66× level-1 start, card idle with no hold active, ≤ 0.30 s): 0.31 s, OUT by 0.007 s (control 1.11 s).
    • Four of the six dataflow runs read 0.00. In the other two (0.88 / 0.96 s), the first level-1 node ready to prove had its program by 4.0 s, but its artifacts (a 0.14 s build on the card) waited 2.0 s behind level 0's proves and were ready only after level 0 ended. A node's slot holds its program and its artifacts together, so its execute + fill waited too, and ran with the card idle. In the four other runs the same build got the card between two leaf proves.
    • So the residual is the slot waiting for the artifacts before the execute, not the dataflow. At the median, level 1's idle read 0.00 in both runs. (Corrected: the first version of this section blamed the builder's emission; the build stamps include the card wait.)
  • A readout note: under dataflow the levels overlap, so a level's start gap (from the level below's last proof to the first prove hold starting after it) can count the next level's own proves. At the median, level 2's gap read 0.65 s with the card idle 0.00 s (BIG 612). The level starts are therefore read as card idle with no hold active.
  • What is left: the top node's execute + fill after its last child (1.0–1.1 s of card idle at the median, BIG 612). That is (B)'s target, next.
  • Toward a production tree driver: with (C), the driver's prove loop is one topological pool over the slots, with the levels taken from the plan's shape.

The held tables spill to disk, by default only when memory is short (lane i-m4b)

Prover-only; the proof bytes are the same: one shape and one proof digest (07d1bd43…) across the base and all four policies, under the fixed trace hash + deterministic grind (BIG 616). The 1× top proof's digest (0da6fea9…) and the tree verify at the landed merge (ULTRA 006).

  • What changes:
    • In phase A, after each group is installed, the narrow (packed) tables of every committed group past the first two can move to STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013's spill store. The store (crypto/stark/src/spill.rs) is imported byte-identical; it digests each slot and writes it with O_DIRECT on a writer thread, off the committer's path. A table moves only when the writer's queue has room at that moment, so the committer never waits.
    • Phase B reads the tables back ahead of their upload, within a window of two groups' bytes, and checks every slot's digest. A table whose columns are still out when its group is read is refused (SpillFailed). After a group's opening its packed columns are let go.
    • The policy is STARK no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1013's knob and rule: LAMBDA_VM_BLOCK_SPILL = auto (the default) | off | always | <GiB> (a resident budget). auto spills a table when the host's bytes plus a reserve plus the table would pass the target.
    • BLOCK SPILL prints the policy, the slots and bytes, the writer's split, the read-back, and the largest host reading auto decided on.
    • Tests:
      • the same bytes spilled and held;
      • a byte flipped in a spilled slot is refused, and so is a slot that does not come back;
      • a budget the block fits in spills nothing, and a zero budget spills every table past the first two;
      • unit tests of the knob, the rule, the cgroup v2/v1 readers and the working-set measure, each with a mutation that fails it.
  • Measured:
    • 4.13×, always (BIG 566, single runs): 22.74 GiB spilled, max RSS 52.53 → 40.12 GiB, whole block +1.15 s, drivers waited 0 s.
    • 4.13×, auto (BIG 567): by default nothing spills (target 110.7 GiB). Under a 35 GiB target it spills 21.21 GiB, max RSS 41.04 GiB, +2.91 s.
    • The median, on the cap-128 exploration head (BIG 568; the cap is not part of this landing):
      • auto spilled 26.98 GiB from group 44 of 90, and the tree verified.
      • VmRSS fell only 110.55 → 105.32 GiB. Under the never-purge allocator posture, freed pages stay resident, and the whole-run peak forms in phase A's finish, which the spill does not reach.
    • The landing gate (BIG 616 at 3d99c64, base WHIR no-epoch (prove-and-retire) prover: block 25368371 — work in progress #1014 54aae28, two jobs pooled, arms B, D = auto, O = off, S = always, T = auto at a 47.5 GiB target):
      • the card tests;
      • one shape and one proof digest across all five arms;
      • every auto run spilled 0 in both jobs, at 110.7 GiB and under the 47.5 GiB target (FAST's limit, kept as a real test on BIG's 128 GiB; the host read 25.3–26.8 GiB across those runs);
      • every run's environment the same size;
      • D − B −0.063 s (t −0.6), O − B −0.048 s (t −0.4), both inside the box's ±0.23 s band.
    • Earlier 1× gates on FAST (880 + 881, 882): always and auto at 0 spilled within noise.
    • ULTRA 006 at the landed merge 0d61081: 1× top digest 0da6fea9…, tree accepted, auto spilled 0 (host at most 27.65 GiB).
  • What is left: the median's peak is phase A's own (the finish's transient over the tables still held when it starts), not phase B's. That is the next memory term.

Phase 4 counts the BITWISE lookups in slices (#1013's c369843, lane i-m4b)

Prover-only; the bytes are the same:

  • the 187 whole-run trace digests (four programs at the small and default caps) are identical with the slices on and off;
  • the 1× proof digest is 07d1bd43… at both, under the fixed trace hash + deterministic grind (ULTRA 010).

What changes:

Measured: the median, BIG 619. One binary on the cap-128 exploration head, N P P N, auto spill on. N runs the whole-source path, P the slices.

N P (slices)
phase 4's transient (unnamed peak) 31.9 / 30.9 GiB 11.0 / 11.0
phase 4's length 13.3 / 15.4 s 6.6 / 6.5
active peak 93.1 / 89.2 GiB 83.5 / 83.1
VmRSS − active peak (retention) +13.2 / +18.1 GiB +5.1 / +4.3
base VmRSS, phase A / B 104.1 / 106.3 GiB (N1) 86.1 / 88.6, 85.0 / 87.4
whole-run VmRSS 107.6 / 108.9 GiB 89.1 / 89.9
the tree phase's cgroup peak 115.8 / 117.1 GiB 97.3 / 98.1
phase A 135.4 / 132.3 s 124.1 / 124.6
whole block 238.3 / 235.0 s 227.2 / 227.4
spilled (auto) 25.9 / 25.6 GiB 25.9 / 25.0
  • Phase A −9.5 s (whole block −9.3 s) comes from the finish, not the spill: every run spilled ≈ 25–26 GiB.
    • Phase 4 lost its two long poles (LT, KECCAK), so the finish builds the run 8.5–11 s sooner.
    • Phase A's wait for the rest's groups falls 33.8 / 32.0 → 20.8 / 22.6 s.
  • 1× (ULTRA 010, 8 + 8, one warm-up discarded): base −0.30 s (t −2.0), recursion −0.01 s.
  • The pre-registered bands for phase 4's transient (≤ 9 GiB), the active peak (≤ 83) and phase A (±3 s) were missed, all on the good side. Landed on the lead's judgment; there is a row in FAILED-LEVERS.
  • Proves 25416971 (12.4×) at cap 128, exploration (BIG 621, acfd54a + the cap, defaults): 40.7 G cells, 109 groups, base and tree verified; whole-run VmRSS 98.0 GiB, cgroup 106.2; 40.3 GiB spilled; whole block 278.1 s. On the line through the median and this block (≈ 1.18 GiB per G cells), cap 128 binds first, at ≈ 14.4× (≈ 47.8 G cells); memory binds at ≈ 15.5–16×.
  • What is left: the median's active peak is now the finish's phase 5. It holds the rest's tables until the finish ends (unnamed +24.0 GiB), beside the held narrow tables (33.8 GiB), which auto spills only once memory is short.

The top executes each child's share as that child is proved, while a child is still unproved (B and B′, lane i-noepoch-w2)

Prover-only (the W3 harness and the LFM executor); the proofs are the same. The 1× top digest is 0da6fea9… with the streamed top on and off, under the fixed trace hash + deterministic grind, and both trees are accepted (BIG 617; ULTRA 008 and 009; ULTRA 011 at c55dec4 itself).

  • What changes:
    • The top node executed its whole program after its last child was proved, although most of that program verifies one child at a time.
    • executor::StreamedExecution runs an execution as its arena groups land: each instruction runs in the wave in which the last group it reads lands (each instruction's group set comes from one forward pass), in the level schedule's order, into the slot the schedule gives it. So the witness is execute's, word for word.
    • Tests: the streamed execution is byte-identical in every landing order; a missing group, a group landed twice, or a wrong arena count or length is refused; finish checks that every instruction ran once. Dropping the group propagation (ReadBeforeWrite) or re-running instructions across waves (DoubleWrite) fails them.
    • The W3 top lands each child's arenas as that child is proved (W3_STREAM_TOP, on by default), so only the last child's share and the fill remain after the last child.
    • The readiness gate (B′): the top streams only if a child is still unproved when its program and artifacts arrive; otherwise it executes whole, on the control's path. W3 TOP EXECUTE prints which. A test covers a program ready before the last child (streamed), after it (whole), and a failed child (whole, which then reports the failure). Inverting the predicate fails the test.
  • Why the gate: the first version (B) always streamed. At 1× and 2.66× the top's slot arrives after its last child: its artifacts are a short card hold that queues behind the children's proves. So there was nothing to stream behind, and the streamed path costs more than one whole execution. 1× recursion rose by 0.22 s (t +5.8, BIG 617 + 618), and the top's start idle rose too (1.01 → 1.20 s). (B) was not landed. With the gate, the top executes whole at those sizes.
  • Time:
    • Median (block 25475471, BIG 613, S O O S O S on one binary): recursion −0.47 s (t −3.2); the card idle at the top's start fell 1.07 → 0.69 s. There the top's program is built 6–8 s before its first child, so it streams. BIG 620 (one gated run at the median) streamed, with recursion 41.91 s, inside 613's streamed range.
    • 1× (ULTRA 009, 8 + 8 after a warm-up): recursion +0.019 s (t +0.8), within the registered ±0.05 s; every gated run executed whole. At 2.66× the top's slot also arrives after its last child, so it executes whole there too.
  • What is left: the top's fill and the streamed path's own overhead (≈ 0.7 s of card idle before the median's top). Splitting a node's slot into its program and its artifacts, so that a node executes before its artifacts are built, is the next lever.

A node executes from its program; only its prove waits for its artifacts (b′, lane i-noepoch-w2)

Prover-only (the W3 harness); the proofs are the same: the 1× top digest is 0da6fea9… with W3_EXEC_EARLY on and off, under the fixed trace hash + deterministic grind, and both trees are accepted (ULTRA 012; ULTRA 013 at 806c211 itself).

  • What changes:
    • A node's slot held its program and its artifacts together, and its artifacts are a card hold that queues behind the proves below it. So a node could not execute until its artifacts were built, although executing needs only its program. At 1× and 2.66× the top's slot arrived 0.12–0.15 s after its last child, though its program had been emitted 0.03–2.1 s before it. At 2.66× a level-1 node's artifacts waited up to 2.1 s for the card while its children were already proved (BIG 614 / 617 / 618's card holds).
    • Now the builder publishes each node's program as soon as it is emitted: on the pool thread that emitted it, before the builder's thread builds its artifacts. The node executes and fills from the program once its children are proved, and only its prove waits for its own artifacts (node_flow). The traces are filled under the block hasher, which the node artifacts are built for. The top's streaming choice (B′) is taken when the top can execute. W3_EXEC_EARLY=0 is the control.
    • Tests:
      • A toy at the median's shape runs the real builder, scheduler and node flow at 1 to 6 workers, without a hang. It checks that every node gets its program within 40 ms of its emission or its take, executes before its artifacts where it can, and proves only with its own artifacts, after they are built.
      • Mutations fail it: waiting on another node's artifact slot; waiting for the artifacts before the execute; publishing a program after a sibling's finish.
      • The readiness predicate's test, the deadlock test (1–6 workers) and the failure-propagation test were re-run.
    • Part of this landing: proving filled traces against artifacts built for another hasher is refused (LfmProveError::HasherMismatch, naming both) instead of failing a release assert, and the streamed execution's finish returns an error instead of asserting its public words' capacity. Each has a test and a mutation.
  • The first build published each program from the builder's thread, between finishes that wait for the card. In one of BIG 622's six runs, a level-1 node got its program 2.3 s after it was emitted, so level 1's idle read 0.183 s against the registered ≤ 0.15. Publishing on the pool fixed it (BIG 624).
  • Time (one binary per box, arms alternated):
    • 1× (ULTRA 012, 8 + 8 after a warm-up): recursion −0.274 s (t −15.1; 4.125 → 3.851 s). The card idle at the top's start fell 0.536 → 0.239 s: the top's program now arrives before its last child, so B′ streams it.
    • 2.66× (BIG 624, 6 + 6 after a warm-up): recursion −0.410 s (t −2.8; 12.008 → 11.598 s). Level 1's start idle fell 0.197 → 0.007 s and the top's 0.795 → 0.525 s.
    • Median (BIG 625, N O O N): unchanged. In every run, every node is taken at least 1.27 s after its artifacts exist: the three workers are busy with the 23 leaves. So the split has nothing to act on there. 625 read −0.55 s, but the control arm itself moved +0.52 s from BIG 623 with no change on its path; over both jobs, −0.31 s at t ≈ −1.3. It passed its registered row by chance.
  • What is left: the top's fill and the streamed path's overhead (≈ 0.5–0.7 s of card idle at the top's start), and at the median the worker-bound leaf level.

The fix was made on fix2/gpu-tables-a23, from A2+A3's 8c925a4, so that
A2+A3's re-run measures A2+A3 alone. Here the test takes A4's Knobs
helper; its assertions are the same.

# Conflicts:
#	crypto/stark/src/multilinear_table.rs
L1N{j} (arity k) in place of "whir wide j", on its census panel, IDENTITY,
TIMING and host-peak lines and in its prove and harvest labels, so a tree
log's readers file the proof as level 1's node j (zf_summary.py's
L{level}N{j} rule). Labels reach prints and error messages only: no program
or proof byte moves, and a wrap's lines are unchanged.
Under LAMBDA_VM_LFM_WIDE=on the level-0 stage yields one node per fan_in
epochs, not one wrap per epoch, and the interior's report covers levels
2..=child_level: the two counts the fixture asserted assumed wraps (job 215's
FW arm stopped on the first). The wrap path's asserts are unchanged.
The level-0 lead-in's slots gain a span: slot k waits for epochs
0..(k+1)*span and its builder is handed that prefix; span 1 is today's wrap
lead-in, call for call. Under LAMBDA_VM_LFM_WIDE=on the production tree
starts the wide lead-in (start_whir_wide_lead_in, span = fan-in) where it
started none: helpers harvest each group's epochs as the base proves them and
emit the wide node with wide_prologue_emit, the emission the pool shares, so
the pool takes a prologue instead of building it. Job 215 measured level 1
opening with 2.7 s of host prologue and the card idle; this moves that work
into the base's tail. The wide level prints its start on the prove-split clock
so the gap to the first card hold is read from the log. Card-free unit test
for the spanning slots.
…A_VM_ARGUE_LEAN_TAIL)

A5. On FAST (job 213) the host's work between two device GKR layers is
264 us a layer, and 52 % of it is the tail's arithmetic: the nine rounds
the host finishes over a cube of 512 after the device hands the layer
back. The generic round extends every factor to t = 2, 3 by
multiplication and sums the whole relation at three points.

Under the knob (default off) a device layer's tail runs lean. Its weight
is eq(point, .) folded on the device's challenges, so at each round it
factors as l(t) * w(x), with l(t) = eq(rho_k, t) and w the sum of its two
halves. The round polynomial is l(t) * h(t) with h of degree two:
- the tail sums h(1) and h(2), extending the factors to t = 2 by
  addition;
- h(0) comes from the claim the round carries, carried through the
  device's rounds exactly as the verifier does;
- h(3) comes from the zero third difference;
- w is carried by additions and one scalar;
- only the four halves are folded, with the generic fold.
That is about 14 extension multiplications a pair where the generic
round makes 30.

Every sent value, challenge and bound factor is the generic round's
field element, so the transcript and the canonical bytes are unchanged.
Raw limbs of the round values may differ, and only a device layer takes
this path, so no host-only byte gate reaches it. The tail declines,
before absorbing anything, when the point does not match or some
1 - rho_k has no inverse.

Tests:
- host: the lean tail against the generic rounds (values, challenges,
  transcript state, bound factors) at 2^1..2^11, 0..3 device rounds in,
  over Goldilocks and its cubic extension. Five mutations (the h(3)
  difference, the claim chain, w's halves, the scalar, l(1)) each fail
  it;
- host: LAMBDA_VM_ARGUE_XCHECK sums every round the generic way too and
  refuses a faulted tail;
- device (cuda-only): argue_lean_tail proves a 2^15 tree to the same
  proof with every card layer lean, keeps host trees generic, and shows
  a wrong lean round refused by GKR's layer check and by the cross-check;
- device: the stark tall tables prove the same bytes with the tail lean
  (alone and with every argue knob), and the faulted tail is refused.

Under the base split an `ARGUE TAIL` line counts lean tails per record.
Brings fix2/gpu-tables-a23 up to 34c1760, the sha job 216 measured on
FAST: WHIR ABBA EFFECTIVE, whole run B - A -5.20 s (band [-5.5, -3.3]
HIT), base 41.65 -> 36.65 s, sum of ARGUE 19.7 -> 14.7 s, 1,364 tables
built on the card per B arm, every arm proved, verified and on the
record's identities, and the untimed cross-check arm clean. Device
checks 7/7, the corrected negative control's mutation step included.

The only conflict is gpu.rs's knob helpers: the A1 landing added
env_not_off and not_off where A2+A3 added its knob block, and both stay,
unchanged. LAMBDA_VM_ARGUE_DEVICE_TABLES is still off by default here;
the next commit turns it on.

# Conflicts:
#	crypto/multilinear/src/gpu.rs
LAMBDA_VM_ARGUE_DEVICE_TABLES is now on unless set to 0, which is the
opt-out and the old path exactly. Its WHIR A/B on the block (FAST, job
216, at 34c1760) read EFFECTIVE:
- the whole run 5.20 s faster (band [-5.5, -3.3]);
- base 41.65 -> 36.65 s and the argue 19.7 -> 14.7 s;
- 1,364 tables built on the card a run;
- every arm proved and verified on the record's identities, and the
  untimed cross-check arm, which compares every card table with the
  host's, clean.

The proof does not change. The card builds the zerocheck's eq weights
and the reduce's shift tables and batched columns as the same field
elements the host built, and the sumchecks run over them as before, so
the canonical bytes, the transcript and every challenge are the same.
The fixture identity test, the cross-check arm and the verified arms
showed it.

No pin moves. Without a device a weight given as its point is made
into the table eq_mle made before, in the same batch order, and the
reduce keeps its host path, so a host-only proof is the same to the raw
limb. The WHIR byte gate builds without cuda. No STARK path reaches the
argue.

A unit test pins that the knob reads its variable the default-on way
(unset on, 0 off). A mutation restoring the old reading fails it.
…ed A/B

Brings fix2/gpu-tables up to 8d2cb35 on top of land/gfs-a23:
- A4, a session's end read back in one copy (LAMBDA_VM_ARGUE_LEAN_READS);
- the GKR layer timers under LAMBDA_VM_BASE_SPLIT;
- A5, a device GKR layer's host tail finished lean
  (LAMBDA_VM_ARGUE_LEAN_TAIL).
Both knobs stay off by default here. The combined A/B measures them
together on top of what is landed (A1 and A2+A3 on), and they land
together if it is EFFECTIVE. No conflicts.
…vel 1)

The recursion's LFM proofs proved by the stacked-WHIR prover: W-LFM proofs
and their in-guest W-legs, the preprocessed-count binding, policy B, the
wide level-1 node with its base-tail lead-in, and the knobs
LAMBDA_VM_LFM_PROVER / LAMBDA_VM_LFM_WHIR_PREP / LAMBDA_VM_LFM_WIDE, still
default-off here. W4's ABBA at 4150afa: 56.65 s against 60.70 s.

The one conflict, crypto/stark/src/multilinear_table.rs's test module, is
the union of both sides' blocks: A1 and A2+A3's tests on the card, then the
count-trap and policy-B tests, inserted at the same base line.
On #1010's line a WHIR chain grinds before its queries only, with one spent
nonce a round (P2), and a W-LFM proof takes its chain config from the same
chain_config. The production-config anchor now asserts GrindBits::query_only(20),
and the W-leg's sizing pins drop each round's folding and OOD grinds and their
two nonce words (wrap 0, policy B: 38,440 -> 38,374 perms, 384,825 -> 383,715
ops, 73,961 -> 73,937 hints); the previous values are kept in the comments.
The F1 exactness tests pass unchanged, and every pin stays within the design's
0.5 % band.
LAMBDA_VM_LFM_PROVER defaults to whir, LAMBDA_VM_LFM_WHIR_PREP to prepared
(policy B), and LAMBDA_VM_LFM_WIDE, unset, follows the tree's prover: on
under whir, off under stark, whose tree has wraps. The opt-outs are
LAMBDA_VM_LFM_PROVER=stark (the per-table STARK recursion as it was),
LAMBDA_VM_LFM_WHIR_PREP=both and LAMBDA_VM_LFM_WIDE=off. D-WHIR W4, the
ABBA at 4150afa: 56.65 s against 60.70 s for the block's whole run.

The §2.4 count-trap tests build under policy A explicitly: only there is a
proof without its prepared opening well-formed, and under the new default
the prover refuses to leave out a prefix nothing settles.
…ts (S0)

A device round's per-thread slot file is the lowered program's live
set, and the slot budget divides by it: the widest batch, KECCAK_RND's,
holds 2,763 values and caps its early rounds at 8,096 threads. That is
exactly the head trace's 253x32 shape, 1.31 s in 14 launches.

The census (an ignored printing test) builds each VM table's zerocheck
program and the batch its session runs (the constraint rule beside the
bus's two), lowers them, and reports steps, slots and the thread ceiling
three ways:
- today;
- with roots, and the bus's interaction terms, folded into their sums
  as soon as they are computed;
- with every factor and constant read emitted again at each use instead
  of held.
At a random point it asserts that every variant is the same polynomial.

On the three batches that bind, folding early barely moves the live set
(KECCAK_RND 2,763 -> 2,344). Reloading the reads cuts it 6-13x
(KECCAK_RND 211, ECSM 269, ECDAS 161) at 50-60 % more steps: the held
values are shared column reads and constants, not roots.

IrShape::program_lean (the zerocheck folded early) and its host parity
test are included; nothing calls it outside the census yet.
…ult off)

A big batch holds so many values a thread that its first device rounds run a
few thousand threads: KECCAK_RND's 2,763 left 8,096 at the 512 MiB slot
budget, 94 ms a launch on the head's trace. The values are the order the batch
was written in: every shared read held from its first use to its last, every
root and every interaction's side held until the sum at the end.

`Program::on_demand` re-emits the same steps in Sethi-Ullman demand order and
emits every read again at each use. Each step is the same operation on the
same operands, so every value is the same; a sum's terms are added in as they
are made. On the four big VM batches the live set drops 2,763 -> 207
(KECCAK_RND), 1,782 -> 48 (ECSM), 1,291 -> 159 (ECDAS), 859 -> 14 (KECCAK),
at about 60 % more steps.

Under LAMBDA_VM_ARGUE_LEAN_PROGRAM (off by default; off is today's path) a
zerocheck whose lowered program holds more than 341 values a thread (under
64 k threads a round) runs the program on demand, with the session's slot file
sized for every interpolation node from the first round (`session_spread`).
Small batches, the byte gate's EQ fixture among them, keep today's program.

Under LAMBDA_VM_ARGUE_XCHECK a shadow session walks today's program over the
same factors each round and the prove is refused at the first disagreement:
the block's identity gate, since block proof bytes are not reproducible.

The base split gains an ARGUE ZEROCHECK line: big sessions, lean and checked
counts, and the big batches' early (half >= 2^13) and late device rounds, the
other batches' rounds and the host tail, timed.

Tests: host parity on every VM batch and the gate's four big batches; device
parity round by round on the four (box); the stark identity knob off and on,
alone and with every argue knob; a fault that the verifier (BatchMismatch)
and the cross-check (DeviceFailed) both refuse.
…the pure-WHIR landing

Merges fix2/lean-program @ 965e13d into land/pure-whir @ 6a6e266. S1a read
EFFECTIVE on FAST (job 223): the whole run 1.35 s faster, the big batches' early
device rounds 1,527 -> 442 ms, the argue 1.43 s faster, every arm proved and
verified with equal identities.

The merge brings S1a's ancestry from gfs/a45 with it, every knob default off:
A4 (LAMBDA_VM_ARGUE_LEAN_READS), A5 (LAMBDA_VM_ARGUE_LEAN_TAIL), the device GKR
layer timers under LAMBDA_VM_BASE_SPLIT, and the S0 census test.

One conflict, in stark's multilinear_table tests: both sides added tests at the
same place (the preprocessed-prefix tests here, the argue-knob tests there).
Both are kept.
LAMBDA_VM_ARGUE_LEAN_PROGRAM is now on unless set to 0, which is the opt-out
and the old path exactly. Its WHIR A/B on the block (FAST, job 223, at
965e13d) read EFFECTIVE:
- the whole run 1.35 s faster (band [-1.6, -0.5]);
- the big batches' early device rounds 1,527 -> 442 ms, the argue 1.43 s
  faster, the other batches' rounds +8 ms, and the late rounds held (no S1b);
- every arm proved and verified on the record's identities and census, and
  the untimed cross-check arm, which walks today's program beside every big
  session round by round, clean.

The proof does not change. The program on demand is the same steps on the
same operands in another order, so every round's values are the same field
elements, and so are the transcript and every challenge. The stark identity
test, the device parity test on the four big VM batches and the cross-check
arm showed it. The round sums of a big batch are added over another launch
shape, so their raw representatives, and a device proof's rkyv bytes, may
differ, as between any two launch shapes.

No pin moves. Only a batch whose lowered program holds more than 341 values
a thread takes the path, and only on a device: the WHIR byte gate builds
without cuda, and its EQ fixture holds 26. The W-LFM pins count the W-leg's
permutations, operations and hints, which no argue path changes. No STARK
path reaches the argue.

A unit test pins that the knob reads its variable the default-on way (unset
on, 0 off). A mutation restoring the old reading fails it.
Under pure WHIR the recursion proves its LFM programs with the base's argue,
so the lean program's gate reaches their zerocheck batches too. A census of
the W-LFM chips (the WHIR recursion chip set at the recursion's hasher, as
whir_lfm_airs builds it) finds one big batch: LFM_HASH holds 365 values a
thread today, a 61,286-thread ceiling, and 34 on demand, 657,930, at 41 %
more steps. LFM_BITDEC holds 233 and stays under the gate.

- every_w_lfm_batch_on_demand_is_the_same_program: host parity on every
  W-LFM batch, and LFM_HASH as the only big one.
- the device parity test now finds its batches by the gate, over the VM's
  and the W-LFM's AIRs, and pins the five.
- the_w_lfm_zerocheck_programs_and_their_live_sets: the printing census,
  ignored.
Both RPX implementations (prover lfm::rpo::Rpo256::mds, used by the block
path's RpxStarkHash, and crypto::hash::rpx::mds) built each output lane with
a core::array::from_fn closure. The closure's generic from_fn wrapper is
placed in a codegen unit of rustc's choosing and is inlined into mds only
when that unit happens to be mds's own. When it is not, every lane is an
out-of-line call that recomputes (j - i) mod 12 with a 64-bit multiply per
term: about a fifth more instructions per permutation.

That is the two-speed host verify on the STARK tree. A Linux x86-64 cross
build of the prover test crate at 8934b59, ef6d4be and 3fd644e shows
mds inlined (2499 B, no calls) only at ef6d4be, the one FAST build, and
twelve closure calls at the other two, the SLOW builds. Crypto's mds makes
the twelve calls in its current partitioning too.

Loops over a precomputed circulant compile the same way in every build. The
permutation's values are unchanged: the RPO and RPX known-answer vectors, the
two-implementation agreement test and a new test against the circulant
definition all pass.
Under `parallel` the grind is `find_any`. It returns any valid nonce, 0
included, and the nonce it returns is absorbed, so every later state varies
from run to run. Two tests assumed otherwise and were flaky there:
- a_ground_proof_verifies asserted every spent nonce was non-zero, but nonce
  0 passes an 8-bit PoW with probability 2^-8.
- a_query_only_chain_checks_its_query_nonce_in_every_round accepted only Ok
  or GrindingRejected for a flipped query nonce. A flip that still passes
  the PoW is absorbed, moves the transcript and fails elsewhere.

The tests now read the verifier's grind checks through a logging transcript
that records each check's state and nonce:
- A ground proof's checks are exactly its spent slots, in order, and each
  nonce passes the PoW at its own state.
- A forged nonce is the first value above the honest one that fails the PoW
  at that state, so the check itself must refuse it with GrindingRejected.

The three a_forged_*_nonce_is_rejected tests forged nonce + 1 and carried
the same latent flake; they now take the same forgery.

Tests only.
Two P2-W tests read a nonce that grinding picks with find_any under the
parallel feature: one asserted it non-zero, the other expected a flipped
query nonce that still passes the proof of work to verify. Both now assert
what the proof of work guarantees. Tests only.
Every device entry point calls device::backend() before it touches the card.
A process-wide, monotone count of those calls (backend_entries()) lets a test
bracket a phase and show it never reached the device. The first user is the
W-LFM prove's host prep, which LFM_CARD_AFTER_PREP=1 moves outside the card
permit.

This is one relaxed atomic add per backend() call, with no other change.
…REP, default off)

prove_traces_whir_opening takes the card permit as its first statement, so
the card is held idle, and closed to every other proof, while the prove builds
its tables on the host. At job 222 that prep was 2.05 s of the pure-WHIR
recursion's 9.89 s of W-LFM holds, 1.68 s of it at level 1.

The prep moves unchanged into prep_whir_tables: the plan's layouts, each
table's main columns moved into its layout, and the prefix check. It takes no
device handle.

LFM_CARD_AFTER_PREP=1 takes the permit after that function instead of before.
Unset, empty or 0 keeps today's order; any other value panics. The setting is
read once and named once on stdout. The split line's wall stays prep + the
held part in both settings, so the wait for the card is in neither.

Tests:
- Laptop: the knob parse and the test override.
- Box, #[ignore]: the fixture's wraps and a wide node, proved with the setting
  off and on, serially and three at a time with the permit armed, under the
  deterministic grind. Every proof must equal the default's bytes.
- Box, cuda: device::backend_entries() must not move across prep_whir_tables
  for any job. The paired control requires that it moves across the prove that
  follows.
The W-LFM prove now takes the card permit after prep_whir_tables, so its
host prep runs while another proof holds the card instead of holding the
card idle. LFM_CARD_AFTER_PREP=0 restores the permit before the prep.
Only the lock moves: card_after_prep_tests proves the same programs under
both settings, serially and three at a time, to the same bytes.

Measured on block 25368371 (FAST, A B B A): -0.65 s at fan-in 3 (job 234)
and -0.90 s together with fan-in 4 (job 236).
An unset LFM_CENSUS_FAN_IN now selects WHIR_WIDE_FAN_IN (4) when level 1
is wide, which it is by default under the W-LFM prover. Fifteen epochs
then make four wide level-1 nodes (4/4/4/3), and the block-artifact root
takes them directly beside the global wrap, so the interior level is gone.

The WHIR trees without a wide level 1 (LAMBDA_VM_LFM_PROVER=stark and
LAMBDA_VM_LFM_WIDE=off) keep WHIR_FAN_IN (3), and the STARK tree keeps
FAN_IN (2), so their shapes and program ids do not move.
LFM_CENSUS_FAN_IN=3 restores the previous pure-WHIR tree.

Measured on block 25368371 (FAST, A B B A): -0.70 s alone (job 235) and
-0.90 s with the card permit after the host prep (job 236).

The tree-shape and root tests gain fan-in-4 rows. The production root
(15 epochs at fan-in 4) joins the honest control, the fixed-size schema
and both L2G tamper tests.
On a wide tree, level 1 verifies epochs rather than proofs. With two to
fan-in epochs there is one wide node and no proofs below it, so root
option A, which takes level top-1's output, would hand the root that one
node against a fold shape expecting one digest per epoch; the drivers'
child-count assert then refuses the block. RootOption::for_tree maps A to
B there, so the root sits above the single node. One epoch needs no
mapping: its one node folds a single root to itself.

Tests: for 1..=40 epochs at fan-in 2..=4 the children a wide tree hands
the root equal what the option's fold shape refolds to, and A as named
fails exactly at 2..=fan-in epochs; the root over every small wide tree
(1..=6 epochs at fan-in 4) executes at the artifact's fixed width; the
moved-L2G tamper test covers one epoch and one node of three.
Both WHIR tree drivers now run the root option through
RootOption::for_tree, so a block of at most fan-in epochs composes under
the default wide tree. Such blocks failed the root's child-count assert:
up to 3 epochs at fan-in 3, up to 4 at fan-in 4. The 15-epoch block's
tree is unchanged.

The fixture tree becomes a body over a block (guest, private input,
epoch size, arity and the epoch count the run must reach); the
three-epoch fixture runs as before. the_whir_fixture_tree_proves_every_
small_block (box tier) proves 1..=6 epochs of a new guest,
test_private_input_spin, whose private input sets its spin count and so
its cycle count, each to a verified root at the default arity.
the_spin_guest_lands_every_small_epoch_count executes the guest on the
host and checks each input reaches its epoch count.
The WHIR production driver's arity bound becomes 2..=5 so that fan-in 5
can be measured: LFM_CENSUS_FAN_IN=5 puts 15 epochs into three wide
level-1 nodes of five, and the root takes them beside the global wrap.
The default stays 4. The two STARK drivers keep 2..=4: nothing above four
has been costed on them, and their arity-3 node did not fit the card.

The tree-shape test gains the fan-in-5 row (15 epochs: 2 levels, 4
nodes); the root's fold-shape tests run at fan-in 5 as well, and the
fan-in-5 root (15 epochs, 3 nodes) joins the root gates' shapes.
WHIR_WIDE_FAN_IN becomes 5: fifteen epochs make three wide level-1
nodes of five, and the root takes them beside the global wrap.
LFM_CENSUS_FAN_IN=4 or 3 opts out; the stark opt-out and
LAMBDA_VM_LFM_WIDE=off keep fan-in 3, and the STARK tree keeps 2.

Measured on block 25368371 (FAST, A B B A, job 238) against fan-in 4:
-1.55 s (level 1 -0.76, root -0.70). Level 1's census falls from
1,052.8 M to 852.6 M cells and the root's from 270.9 M to 149.6 M (its
hash table stays under 2^18). The W-LFM argue's reservation peaked at
24.3 GiB of the 25.7 GiB budget: 1.4 GiB of margin, so a block with
heavier epochs should be measured before relying on five.

The small-block fixture test runs at the new default: one epoch, one
node of two to five epochs, and two nodes at six. The root tests gain
the fan-in-5 root in the tamper arms and the small-tree execution test.
The WHIR_WIDE_FAN_IN doc comment, and b682091's commit message, read the
RESERVED HW figures, which the tree log prints in MiB, as GiB by dividing
by 1000. The W-LFM argue's reservation peaked at 24,342 MiB (23.8 GiB) of
a 25,688 MiB (25.1 GiB) budget at fan-in 5: 1,346 MiB of margin, not
1.4 GiB. Fan-in 4 peaked at 21,754 MiB. The ratio, and so the claim,
stands. Comment only.
When the card refuses a GKR fraction tree's carry and handing it back
saves enough to be worth it, input_layer_tree_impl asks for one promise
for the whole tree. If the budget refuses that too on the consume path,
logup::resident_tree gets None and the per-table argue builds the
table's factors and whole tree on the host: slower, never wrong, and
until now uncounted. The refusal is now counted where it happens, in
multilinear::gpu::gkr_tree_refusals and in math-cuda's argue-surface
device_fallbacks, and logged as it happens. A prefetch refused the same
way is not counted: it only means no prefetch, and the consume site asks
again.

The WHIR production tree prints "gkr tree refusals N" beside "device
fallbacks N", which now includes it. DEVICE_FALLBACKS' doc names its six
sites by function instead of line numbers that had moved, and says why
the other argue-side reservations in multilinear::gpu are not fallbacks:
the carry is speculative, a prefetch is optional, and reserve_room's two
callers keep the work on the card.

A test-only hand-back threshold (set_hand_back_threshold_for_tests) lets
a table smaller than the widest precompiles reach the whole-tree
promise. multilinear's call-site census pins the one counted site;
math-cuda's keeps its five.
A box test (cuda, ignored, run alone) on an ADD table of 2^12 rows,
whose factors go to the card. At the normal budget the card builds the
tree and nothing is counted. With every carry handed back and the rest
of the budget held by the test, a prefetch is declined uncounted, and
the consume path's refusal reads one GKR tree refusal and one argue-side
device fallback. The host's tree, the path the refusal hands the table
to, has the card's output fraction.
…re proved (W3_DATAFLOW)

The tree proved level by level: a node whose children were all proved waited
for the whole level below, and then for its own execute + fill with the card
idle. At the median after the node pipeline the card still sat idle 1.4-1.6 s
before level 1 and 0.9-1.1 s before each later level (BIG 611).

Now one pool of `siblings` workers proves every program of the tree, taking
them in topological order (the leaves, then each node level), each publishing
its proof into a slot; a node waits only for its own children's slots, so it
executes, fills and proves while the rest of the level below still proves.
Workers that take programs in this order reach a node only after every
program before it has been taken, so the earliest unfinished program never
waits on an untaken one: no deadlock at any number of workers (tested; taking
the nodes first deadlocks the test). The siblings bound, and so the memory in
flight, is unchanged. A failed or panicking prove fails every program above
it without a hang (tested). W3_DATAFLOW=0 is the control: each node waits for
its whole level below (tested to never start early). A level's W3 LEVEL wall
is now from the level below's last proof to its own.
…cf253d2)

The host measure #1014 copies from #1013 took max(VmHWM, the cgroup's
charge), and the charge counts page cache. FAST charged 40.58 GB at 17:28Z
with 15.18 GB of it inactive file pages, which would have made auto spill
the last groups of a 1x block that fits.

This re-copies #1013's measure from its landed head 035aef5 (commit
cf253d2): HostReading reads VmHWM, the charge and memory.stat's inactive
file pages (v2 inactive_file, v1 total_inactive_file, by their exact key),
and the host's bytes are max(VmHWM, charge - inactive_file). CgroupValue
lets cgroup_memory read a whole-file number or one memory.stat key. The
BLOCK SPILL line ends with the largest reading auto decided on, as #1013's
does. With #1013's two tests; the text differs only in citations and the
test's temporary directory. The rest of #1013's train (lazy queue budgets,
the allocator purge) is not copied: #1014 has no queue budgets.
#1013 added one test-hook field (ProveOverrides::tight_claims, behind
cfg(test)/test-utils) with its tight resident claims (1cc3a70, 00871b9).
#1014 keeps the store byte-identical to #1013's: git diff 035aef5 --
crypto/stark/src/{narrow,spill}.rs is empty. Nothing in #1014 reads the
field.
…group lands

StreamedExecution runs an LFM program as its arenas arrive in groups: one
forward pass gives every instruction the set of groups it reads through any
chain of memory (a Hint reads its arena's group; every operand address is
below its destination), and each instruction runs in the wave in which the
last of its groups lands, in the level schedule's order restricted to that
wave, writing the record slot the schedule gives it. What reads no arena runs
before any group lands. finish() refuses unless every group landed and every
instruction ran exactly once, then publishes the slots as execute() does.

The witness is execute()'s, word for word: the identity gate now also runs a
node-shaped program (three independent sponges and a step over all three) in
all six landing orders, and every identity case in two, against the serial
reference; on the node-shaped case at least two thirds of the program runs
before the last group lands. A wrong group set fails stop rather than
miscomputes: dropping the propagation through memory is ReadBeforeWrite, and
re-running an instruction in a later wave is DoubleWrite (both tested as
mutations). Malformed landings (twice, wrong arena count or length, finishing
early) are refused. The machine now holds its arenas as slices (an unlanded
one is empty, so reading it is ArenaOutOfBounds), and the fill half of
lfm_execute_and_fill is lfm_fill_executed, for an execution run elsewhere.
… child is proved (W3_STREAM_TOP)

After the node pipeline and dataflow, the card still sits idle ≈ 1.05 s at the
top node's start at the median (BIG 611 / 612): the top executes and fills
only after its last child is proved. The top now streams its execution
(StreamedExecution): its children are its arena groups, in order and the same
number of arenas each; it lands each child's arenas as that child's proof is
published (in arrival order, each once; a failed child fails the top with its
error), so only the last child's share, the fill and the prove remain after
the last child. The prove is the unstreamed one's over the same witness. The
execute the split reports is the part after the last child landed.

prove_dataflow now hands a program its children's slots unwaited, so a node
waits for them as it needs them; the others still wait for all of their
children first. W3_STREAM_TOP=0 is the control.
The spill code (stark, multilinear, block_whir) and the streamed top (executor, proof, the W3 tree) touch disjoint code; the one shared file, whir_block_tests.rs, merges without conflict (the spill side adds BlockSpillPolicy::Off to two test fixtures).
Phase 4 round-robins its collectors into at most eight 80 MiB histograms, but
LT, KECCAK and the other per-op sources were one collector each: at the median
block LT (2.6 G cells of its ops kept to the finish) and KECCAK were single
long poles, and each built one list of every lookup it sends (LT 8 lookups per
op, KECCAK about 5,000 per permutation) before counting it. Every source that
is a sum over its ops is now cut into slices of whole ops (LT, SHIFT, BRANCH,
BYTEWISE, EQ, STORE 2^20; MEMW_A 2^22; KECCAK 2^11; ECDAS 2^16), and MUL and
DVRM, which deduplicate per instance, into one slice per instance. The
histogram is a commutative sum, so BITWISE does not move: the 187 whole-run
trace digests (four programs at small and default caps) are equal before and
after.

(cherry picked from commit c369843)
…lectors (the A/B's control)

The cherry-picked c369843 cuts phase 4's per-op BITWISE sources into
slices. #1014 measures it as a memory lever: at the median, phase 4's
whole-source lists were the peak's largest term that does not spill
(unnamed 7.5 -> 31.8 GiB inside p4, BIG 568). So the path before stays
behind LAMBDA_VM_P4_SLICED=0 for a one-binary A/B; unset, the slices run.

The slice lengths now come from one function, p4_slice_len (MUL and DVRM:
one instance, since they deduplicate per instance; the per-op sums: a
bound on a slice's list of lookups). A test counts MUL, DVRM and LT whole
and slice by slice at those lengths and needs them equal; its MUL and DVRM
ops repeat inside instances, so a cut one row off does change the counts
(checked in the test), and a p4_slice_len that misaligns them fails it.
At 1x and 2.66x the top node's program and artifacts arrive after its last
child is proved (its artifacts are a short card hold that queues behind the
children's proves), so a streamed top had nothing to stream behind and ran
its forward pass and one wave per child after the last child: +0.19 s from
its slot to its prove at 1x, and 1x recursion +0.22 s (BIG 617 / 618). At
the median the slot fills 6-8 s before the first child and streaming paid
(-0.47 s, BIG 613).

Now the top streams only if a child is still unproved when its program and
artifacts arrive, and executes whole otherwise (the control's path, so the
proof is the same either way). W3 TOP EXECUTE prints which it did. A test
covers a program ready before the last child (streamed), after it (whole),
and a failed child (whole, which reports the failure).
…2/stream

The p4 port touches only prover/src/tables/trace_builder.rs; the streamed top and its readiness gate touch the executor, proof.rs and the W3 harness. No file in common, no conflict.
…for its artifacts (W3_EXEC_EARLY)

A node's slot held its program and its artifacts together, and the
artifacts are a card hold that queues behind the proves below: at 1x and
2.66x the top's slot came 0.12-0.15 s after its last child although its
program was emitted 0.03-2.1 s before it, and at 2.66x a level-1 node's
artifacts waited 2.0 s for the card while its children were already proved
(BIG 614 / 615 / 617 / 618's CARD HOLD lines). The node's execute and fill
need only its program.

Now the builder publishes each node's program as it is emitted, before its
finish builds the artifacts, and the node executes and fills from it as soon
as its children are proved; only the prove waits for the artifacts, in the
node's own slot (node_flow). The traces are filled under the block hasher,
which the node artifacts are built for; the prove asserts it. The top's
streaming choice is taken when it can execute. W3_EXEC_EARLY=0 is the
control (the execute also waits for the artifacts). The proof is the same:
the execution, the fill and the artifacts do not change, only when each runs.

A toy at the median's shape over build_levels, prove_dataflow and node_flow
shows, at 1 to 6 workers and without a hang, that every node proves only
with its own artifacts and after they are built, and that nodes execute
before them. The toy builder checks each program is published before its
finish starts.
…ts emission ends

The builder's thread published a node's program only when it took the
emission off the channel, between finishes, and a finish holds the card for
the node's artifacts: a program emitted while a sibling's artifacts waited
for the card was published only after them. At 2.66x a level-1 node emitted
at 4.04 s got its program at 6.30 s, behind its sibling's artifacts that
waited 2.1 s for the card, so it executed after level 0 and the card sat
idle 1.08 s (BIG 622, run 9).

Now each emission publishes its program on the pool thread that emitted it,
before it is sent to the builder's thread to finish. The toy now emits a
level in reverse and holds the card 120 ms in the first finish, and checks
every node gets its program within 40 ms of its emission or its take.
…s is refused, not asserted

LfmFilled::prove checked that the artifacts were built for the hasher the
traces were filled with by a release assert_eq. Since the W3 tree fills its
leaves and nodes before their artifacts exist, the pair comes from two
places, and a mismatch would panic the prover. It is now
LfmProveError::HasherMismatch, naming both hashers, returned before anything
reaches the card. The streamed execution's finish likewise returns an
internal error instead of asserting the public words' capacity before its
set_len. The proofs are unchanged.
…as a knob

Port of #1010's knob (fc0ff59) to the no-epoch WHIR block. The WHIR chains
ground 20 bits, a constant. The knob makes the bits a format parameter read at
the one config site, chain_config_under, which every side of the block reaches
through BlockFormat::chain_config: the prover's group chains and prepared
openings, verify_block_whir, and the plan the block leaves are emitted from.
All of them take the bits from the process format, never from a proof, and the
block statement absorbs Q and the grind bits. The query count already
subtracts the query grind, so 18 bits raises Q from 112 to 114 at every
production height (15..=27 variables); every proven phase keeps its minimum
(binding WHIR phase: the unground fold at 27 variables, 130.393; query phases
130.926 -> 130.907). The block tree's own leaf, node and top proofs are STARK
proofs under aggregation_wrap_options and do not read this knob.

Unset is 20: today's configs and proofs, byte for byte. The banner gains
whir_grind_bits=. A chain ground at fewer bits is refused by a stricter
verifier on both the host and the machine, and a proof opened at a Q the
verifier does not expect is refused.

(cherry picked from commit fc0ff59)
…nd 18 grind bits

Prints the emitted chain verifier's real rows per chip, production format,
n 21..=27, under LAMBDA_VM_ZF_WHIR_GRIND_BITS 20 (Q 112) and 18 (Q 114): the
per-chain deltas that size the recursion's table heights under the 18-bit arm.
Asserts nothing; ignored.

(cherry picked from commit 9d51421)
… host and in its leaves

- a_block_ground_at_18_bits_is_refused_by_a_verifier_at_20: a block proved at
  18 bits (Q 114) verifies under 18 and is refused by verify_block_whir at 20;
  a block proved at 20 is refused at 18.
- a_block_leaf_at_20_bits_refuses_a_block_ground_at_18 (box tier): over a
  block ground at 18, the plan at 18 emits a leaf that executes and the plan
  at 20 emits one that refuses it.
- grind_bits_group_cost_census (ignored instrument): the plan's charge for a
  group's stacked opening per stack height and polynomial count at 20 and 18
  bits, which predicts the W3 plan's costs and leaf count under the arm.
- The real-block tree driver prints each leaf's chips, real / padded rows
  (W3 LEAF CENSUS), so a run shows which leaf table, if any, doubled.
…used, not asserted

Two release asserts on the prover's path become typed errors:
lfm_prove_with_hasher's assert_eq between the artifacts' hasher and the one
asked for is now LfmProveError::HasherMismatch, returned before anything
runs; and the unstreamed execute's capacity assert before its public words'
set_len is now LfmExecError::Internal, through public_rows_fit, which the
streamed execution's finish uses too. Each has a test. The proofs are
unchanged.
…114; min proven bits unchanged at 130.393)

The default flip of LAMBDA_VM_ZF_WHIR_GRIND_BITS: the port of #1010's
5364aa2 and its pin follow-up 2687e48. PRODUCTION_WHIR_GRIND_BITS is 18;
LEGACY_WHIR_GRIND_BITS = 20 keeps the legacy format and is the opt-out
(LAMBDA_VM_ZF_WHIR_GRIND_BITS=20). It reaches #1014's WHIR proofs: the
block's group chains and prepared openings, verify_block_whir, and the W3
leaves' in-guest verifier. The W3 tree's own STARK LFM proofs do not read it.
The binding WHIR phase is unchanged (the unground fold at 27 variables,
130.393 bits); the query phases move 130.926 -> 130.907.

Measured behind the knob on c55dec4: ULTRA 019 1x phase B -0.344 s
(t -12.9), whole -0.270 s, recursion -0.010 s; BIG 628 median whole
-3.42 s, both arms verified. The opt-out reproduces the previous 1x proof
byte for byte (top digest 0da6fea9 under fixed hash keys and deterministic
grind); the new default's is b9b4826d.

The W-leg sizing pins follow the process's grind bits, as 2687e48 did on
#1010.
Grind 18 touches the WHIR format, the multilinear prover and their tests; the follow-up touches proof.rs and executor.rs (no file in common, no conflict).
…e tests

HostTable / host_table_forked, TableLegs / build_table_legs /
fri_layer_openings and the child harvest (RealChild, now HarvestedChild,
with its arenas) are what the W3 tree's driver needs to fill each node's
arenas from its children. They lived in epoch_tests, epoch_verify_tests
and per_table_aggregator_tests; they now live in lfm::harvest (the same
module as #1013's), and a proof that disagrees with its AIR is an Err
instead of an assert. The suites keep every name and its panicking form
through re-exports and shims. Same logic, same order.
the_whir_block_tree_on_a_real_block's body, with prove_tree_pipelined and
its machinery (the publish slots, build_levels / build_nodes, the dataflow,
node_flow, the streamed top), was the only way to prove a whole WHIR block
as a tree. It is now lfm::whir_block_tree::prove_whir_block_tree, with
every knob read once by WhirTreeConfig::from_env (same names, same
defaults: W3_*, BLOCK_WHIR_*), every line sent through a WhirTreeSink (the
harness prints to stdout as before, byte for byte, plus a first BLOCK
POSTURE line), and every panic on the run's path an Err. The harness test
is a wrapper that keeps its off-clock checks (the early leaf's arena, the
device line, the top digest, every tree proof and the base verified, the
block verifier). in_index_order moves to lfm::tree_run; the BLOCK_WHIR_*
env readers in block_whir are no longer test-only. The toy and dataflow
tests stay in whir_block_tests and import what moved.
…verifier

WhirBlockTreeProof is what a consumer of a whole W3 block receives: the
block's statement and the top node's proof, behind the magic, the layout
version and a pipeline tag shared with #1013's STARK block file (so the
two trees' files are never read as each other's). verify_whir_block_tree_proof
is whir_block::verify_block_tree over the claimed statement, unchanged:
the plan and the top program are still derived from the trusted ELF under
the block presets, and no format parameter is read from the file.
The shipped binary could not prove a whole WHIR block as a tree, and it
had no CUDA build. prove-block runs lfm::whir_block_tree::prove_whir_block_tree
(the driver the W3 harness runs) and writes a WhirBlockTreeProof;
verify-block checks it with verify_whir_block_tree_proof. For the two
block commands the binary sets the production posture (table parallelism,
the row cap, retention, the RPX WHIR hash, the tree cache cap, the
executor schedule, the grind search) for every knob the environment
leaves unset, before any thread exists; the WHIR hash is a format knob,
so verify-block sets it too. W3_LEAVES, W3_FAN_IN and BLOCK_WHIR_ARGUE
make a tree the block verifier does not derive and are refused. No purge:
#1014's harness has none. New tests: the arguments, the posture plan, and
the never-purge conf this binary compiles in.
…tamps

dataflow_proves_the_level_by_level_programs_in_any_completion_order asked
whether any node started before its level below ended by comparing
millisecond stamps, and at a laptop load of 11 a 2-worker run once found no
such node (i-cli, 1 of 5 runs). Now the last leaf, once taken, holds its
worker until some node starts or 1 s passes, and the test reads whether that
signal came: with two or more workers and dataflow a node takes the free
worker and starts while the leaf still proves; under the barrier, or with
one worker, none can. The proved texts are checked as before. Making the
barrier inert, or always on, fails it.
The same writer as #1013's G-pack: a generator writes its trace straight
into packed columns at guessed widths, the writer keeps the OR of every
word per column, narrows the columns given more bytes than they needed in
place, and names the widths the trace needs when a word did not fit. Its
bytes are NarrowMain::pack's for the same words.

Here a trace holds its packed columns as multilinear's NarrowColumns (the
same layout), so NarrowMain::into_parts is public and
TraceTable::try_from_narrow_main takes the columns, giving them back where
a trace cannot be held packed.
…ctly

Each generator's fill runs through generate_main! (tables::gpack), as on
#1013: TraceForm::Wide is today's 64-bit table; TraceForm::Narrow writes the
packed columns directly at the widths the table's earlier traces needed,
narrows or rewrites on a miss, and builds a kind's first trace wide then
packs it to learn them. The bytes are NarrowColumns::pack_row_major's of the
same words whatever the hint.

On the block (BlockOptions::gpack, LAMBDA_VM_BLOCK_GPACK=0 the old path):
- each streamed chunk is written packed by its job and laid out narrow from
  its packed columns (table_of_narrow), with no 64-bit table and no
  transposition, when the groups are held narrow;
- with pack_finished, the finish writes each table it packs packed
  (WindowedTraceBuilder::generate_packed), KECCAK_RND's chunks included;
- BLOCK GPACK reports how each trace was built.

Tests: every path of the writer gives the wide table packed; every table of
a windowed build in the block's configuration, on four programs, has the
packed bytes and words of the same build without G-pack; one column a width
too wide or shifted fails that comparison; the block proves with the same
partition and tables (the same bytes under the deterministic grind) with
G-pack on or off; a box-only test compares every table of a real block.
stark gained a libc dependency (ee252c1, the packed-trace store
imported from #1013) without the recursion guest workspace's lockfile,
so every make compile-recursion-elfs rewrote
bench_vs/lambda/recursion/Cargo.lock and left a tracked file modified.
This is the lockfile cargo writes; no dependency version moves.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant