Skip to content

Release audit 2026-08: green CI census, claim corrections, adversarial pre-talk pass - #112

Merged
dtumad merged 11 commits into
dtumad/v1.0-releasefrom
dtumad/release-audit-2026-08
Aug 12, 2026
Merged

Release audit 2026-08: green CI census, claim corrections, adversarial pre-talk pass#112
dtumad merged 11 commits into
dtumad/v1.0-releasefrom
dtumad/release-audit-2026-08

Conversation

@dtumad

@dtumad dtumad commented Aug 8, 2026

Copy link
Copy Markdown
Contributor

First unit of the 2026-08 release-readiness audit campaign (stacked on #110's head branch, like #111 was).

What's here

CI: the red audit job is fixed by construction. The axiom census is split by library so each CI job probes exactly the oleans its build target produces:

  • scripts/axiom_probe.lean — 455 main-library probes (import SP1Clean only); gated in the audit job via run_audit.sh --main-only.
  • scripts/axiom_probe_test.lean — the 21 SP1CleanTest conformance anchors; gated in the test job via --test-only right after lake test builds those oleans.
  • Split verified content-perfect against the committed census: 476 = 455 + 21, no entry added, removed, or with a changed axiom set. The main scope now hard-fails on any compiler-trusted constant (the quarantine tripwire); the test scope discloses its 21/21 native_decide carriers.

Four verified documentation overclaims corrected at the claim (or closed by a new artifact):

  1. Sail config disclosure: "six values disclosed as rfl lemmas" was false — three lemmas existed, and two of the six generated sites are ValidateConfig-internal lets no lemma can address. Added the fourth disclosure lemma (sys_pmp_usable_count_eq_zero) and corrected the prose.
  2. "All quoted signatures are mechanically checked": the checker verifies paths + declaration names, not signature text. Claim now says exactly that.
  3. Field bound: the capstone/grounding layer carries Fact (2^25 < p) — now stated as the three-threshold ladder, sourced to SP1's own memory.rs "field larger than 2·2^24" requirement.
  4. native_decide/bv_decide axiom names: ~12 sites cited Lean.ofReduceBool/trustCompiler, which no longer appear in a v4.32.2 census — updated to the generated ._native.*.ax_* constants a reviewer will actually find.

Plus staleness: Verifier.lean's ArkLib-toolchain comment, the hand-written-maxRecDepth count (two, both named), the ChipAssumptions keep-list LoadX0 double-count.

Verification

lake build SP1Clean (3645 jobs) + lake test (3691) + scripts/run_audit.sh (both scopes) all green at this tree; check_pins.sh / check_report_citations.sh / check_root_index.sh pass.

Further campaign units (adversarial per-chip spec audit, file-by-file sweep, real-row non-vacuity battery, OpcodeTable extraction cross-check, EF-findings closure map in report §11) will land on this branch.

🤖 Generated with Claude Code

dtumad and others added 4 commits August 8, 2026 13:25
…t it builds

The audit job builds only the default SP1Clean target, but the generated probe
imported SP1CleanTest modules whose oleans only `lake test` produces — so the
census failed in CI (unknown module prefix 'SP1CleanTest') while passing locally
after a test run. Align each gate with the oleans its job actually builds:

- gen_axiom_probe.py emits scripts/axiom_probe.lean (main library, import
  SP1Clean only, 455 probes) and scripts/axiom_probe_test.lean (the 21
  SP1CleanTest conformance anchors, explicit module imports).
- run_audit.sh grows --main-only / --test-only (default runs both scopes) and
  diffs per-scope snapshots: docs/snapshots/axiom-census.txt +
  docs/snapshots/axiom-census-test.txt. The main scope now hard-fails on ANY
  compiler-trusted constant (the quarantine tripwire); the test scope discloses
  its native_decide carriers (21/21) without forbidding a future stronger proof.
- CI: the audit job runs --main-only; the test job runs --test-only right after
  lake test.
- check_pins.sh sums both probe files, so the documented 476-declaration figure
  is unchanged and still gate-checked.

Split verified content-perfect against the committed census: 476 = 455 + 21,
no entry added, removed, or with a changed axiom set. Snapshots restamped at
this tree via run_audit.sh --update (AUDIT PASS).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Adversarial claims-vs-tree review of the release surface found four claims that
had drifted from what the tree proves; each is fixed at the claim or closed by a
new disclosure artifact:

- Sail config disclosure (report §3.2/§10 T4, sail-model-provenance.md): the
  'six values disclosed as rfl lemmas' claim was false — only three lemmas
  existed, and two of the six generated sites are let-bindings inside
  ValidateConfig that no lemma can address. Added the fourth top-level
  disclosure lemma (SailMem.sys_pmp_usable_count_eq_zero, deliberately not
  @[simp] — no proof consumes it) and corrected the prose to 'four rfl lemmas +
  two ValidateConfig let-sites (config validation only, visible in generated
  source)'.
- Citation checking (report header + §13): check_report_citations.sh verifies
  cited paths and declaration NAMES, not quoted signature text. The claim now
  says exactly that.
- Field bound (report §2): the grounding layer and supported_core_native_sound
  carry Fact (2^25 < p), not just the 2^17/2^24 thresholds the report named.
  Now stated as a three-threshold ladder, with 2^25 sourced to SP1's own
  memory.rs requirement ('field larger than 2*2^24', per the
  TimeExtraction.lean module docstring). KoalaBear satisfies all three.
- native_decide/bv_decide axiom naming (~12 sites: report, AGENTS.md,
  run_audit.sh + check_no_native_decide.sh echoes, four agent docs, two
  SP1CleanTest docstrings): at v4.32.2 the census surfaces compiler trust as
  generated ._native.(native_decide|bv_decide).ax_* constants, not the named
  Lean.ofReduceBool/Lean.trustCompiler axioms these docs cited — a reviewer
  grepping the documented names found nothing.

Staleness: Verifier.lean's 'ArkLib on Lean 4.31' comment is now
toolchain-relative; perf-findings' 'three hand-written maxRecDepth sites' is
now two with both named (verified against option_escapes_allowlist.txt: 11
total sites, 2 hand-written); ChipAssumptions' keep-list docstring no longer
double-counts LoadX0.

lake build SP1Clean (3645 jobs) + lake test (3691) + run_audit.sh both scopes:
all green at this tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…antification session

An adversarial statement-vs-prose session independently re-derived the
quantification of every capstone theorem (verdict: the theorems and the
reserved-name policy are honest; the 10-cell public-value prefix and the
no-ensemble-completeness discipline propagate correctly everywhere) and then
held every reader-facing sentence to that baseline. Two sentences in the §11
competitive comparison had dropped premises outright; seven more were precise
only if the reader already knew the theorems:

MISFRAMES (report §11): 'everything global is a theorem downstream of channel
balance' contradicted ProviderBindings' own docstrings — the per-location
init/finalize uniqueness facts are premises balance cannot force; the
consistency-row paraphrase omitted both disclosed premise conjuncts, making
the stated implication false. Both now name the §8.1 premises.

PRECISION: the 'M1/M2 to be derived from system tables' closure story now
distinguishes provider-content facts (executionCase-derivable) from the
permanent app-level contracts M1 also carries (SailConfigured,
SailCodeMemoryCompatible, program well-formedness); row-exactness is now
attributed to the intermediate grounding theorem rather than the exported
relation that discards it; the COMMIT-digest sentence no longer reads as a
discharged theorem (it is the publicCommitOperand obligation); the report's
headline sentence now scopes to 'each of the 25 instruction chips… composed
over a native 36-table ensemble' instead of inviting the composed
exact-upstream reading; §8.1 enumerates all 11 InitialBoundaryFacts fields;
'single genuinely open obligation' no longer implies the bundle's other
fields are closed; plus D3/D4/D5 nuances and the stale 'three-line fork'
phrase.

check_report_citations.sh + check_pins.sh green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ble helper notes, stale docstring

From three completed adversarial chip-family sessions (Branch/Jal/Jalr, the five
loads, Lt/Bitwise — all three verdicts: no soundness gap; the EF-precedent
classes JALR-masking / load-widths / SLTI-vacuity each close on named,
axiom-clean lemma chains):

- report §12 gains item 7: trap/exception executions are unrepresentable (each
  chip forces its width's alignment, matching SP1's executor), and jump/branch
  targets live in the 3-limb 48-bit PC space — a completeness restriction, not
  a soundness gap. Both facts were previously only implicit in the
  RETIRE_SUCCESS-shaped statements.
- LoadDouble/LoadX0 bridges: the spec_ld helpers' is_unsigned := true parameter
  is decoder-unreachable ((8, true) forbidden by valid_load_encdec) and inert;
  comments now say so instead of reading as a wrong-spec model.
- LoadByte bridge: the AdvanceReady bundle docstring now points at its
  discharge site (ChipContracts/MemoryChips closure lemmas incl. the 8-case
  lane-selection theorem) so the premises no longer read as assumptions.
- BitwiseChip/Formal.lean: the Spec docstring claimed the flag sum is not tied
  to is_real — contradicted by the circuit (Defs.lean asserts
  is_real − (is_xor + is_or + is_and) = 0); now cites the constraint and the
  selectorActive_of_mainConstraints closure.

lake build SP1Clean green (3645 jobs) at this tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@dtumad
dtumad force-pushed the dtumad/release-audit-2026-08 branch from 6d5c3d6 to 66f0458 Compare August 8, 2026 18:26
dtumad and others added 7 commits August 8, 2026 14:04
…ion fixes

The flagship new artifact: SP1CleanTest/NonVacuityReal.lean — for every one of
the 25 instruction chips, a named theorem exhibits a concrete is_real = 1 row
with non-degenerate operands satisfying the chip's complete flattened
constraint system (every subcircuit assertZero, at SP1's KoalaBear field, with
witness values from the chip's own main witness closures; Lt/Bitwise/UType
exercise both variants/outcomes; a guard theorem checks every assertion system
is nonempty). 34 named anchors, census-visible (test scope 21 -> 55; total
510), native_decide-discharged in the sanctioned quarantine. This closes the
'soundness true only vacuously on real rows' failure mode at the chip level —
the existing NonVacuity.lean witnessed 17/20 chips only at padding rows.
Report §9 rewritten accordingly.

Wave-1 adversarial-audit session fixes (12/12 sessions complete, 0 blockers):
- Addw completeness excludes ADDIW (imm_c = 0 witness pin): report §11/§12.4
  now name Bitwise/Lt/Addw + UType's rd≠x0 scope instead of Bitwise/Lt only;
  the ChipAssumptions docstring no longer reads as if Addw were structurally
  register-register.
- New report §5.3: the per-chip audit surface is the Spec PLUS its named
  closure lemmas (memory-chip lane selection, selector-activation, imm-row
  bindings) — the recurring pattern three sessions surfaced.
- Timestamps/boundary session: GroundingAdapter's compare_low polarity was
  inverted in the M2 docstring (the prev_high premise serves the compare=0
  high-limb branch); ProviderBindings' uniqueness docstring now describes the
  actual upstream mechanism (MemoryGlobal control chain + LtOperationUnsigned
  + PublicValues endpoint anchors) and records the locOf-cell vs exact-address
  granularity gap the eventual discharge must bridge; roadmap §2's closure
  attributions corrected (Program discharge routes through C1 + an unbuilt
  decode correspondence — the exact table has empty asserts; memory-image
  value truth routes through the vk cumulative-sum binding, not executionCase
  tables). New roadmap backlog entry: one-instruction end-to-end instance for
  joint premise satisfiability.
- decode/Coverage session: AGENTS.md now files the opcode enum at
  Model/Opcode.lean (53/53 opcodes + discriminants hand-verified against
  opcode.rs at the pin — the citable interim anchor for the hand-mirror).

lake build SP1Clean (3645) + lake test (3692) + run_audit.sh both scopes:
== AUDIT PASS == at this tree (main census unchanged at 455; test census
restamped with the battery).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…acuous on jumps

The 2026-08 adversarial audit's one substantive finding (user-approved fix):
two pulled byte-row shapes had no pusher among the provider tables, so byte-
channel balance — a premise of supported_core_native_sound — was unsatisfiable
for any witness containing a real Jal/Jalr/Branch/AluX0 row. Soundness held,
but only vacuously, on every jump-, branch-, or x0-ALU-containing execution:

- AluX0 pulls the LTU row ⟨4, 1, opcode, 29⟩ (its opcode < 29 range check) —
  SP1's own byte_table() has LTU, but only 5 of the 6 byte opcodes had
  provider circuits here;
- Jal/Jalr/Branch pull width-14 Range rows (the upstream-#102 pc/4
  divisibility checks) — RangeChip instantiated only widths 8/13/16.

Fix (faithfulness-improving — both tables exist upstream):
- ByteChip.Ltu: the sixth byte-opcode provider (MSB-template: range-checked
  operand pair, boolean ltu witness, the b − c + ltu·256 range trick forcing
  ltu = 1 ↔ b < c; push discharged via the new byteRowSpec_ltu iff-lemma);
- RangeChip.circuit14, exactly like circuit13;
- sp1ProviderTables 11 → 13, ensemble 36 → 38; every structural consumer
  repaired (FinishedChannels' 13-pair shape lemmas, ProviderBindings' shifted
  provider indices 35/36/37, TypedProgram/TypedMemoryBalance suffix
  decompositions, alignment lemmas) and every doc count updated (AGENTS.md,
  architecture, roadmap, verification-report D1/§6.1/§8). The 34-table
  exact-upstream cluster counts are untouched (different object).

Both new circuits depend on exactly [propext, Classical.choice, Quot.sound].
lake build SP1Clean (0/0/0) + lake test + run_audit.sh (AUDIT PASS, zero
census drift — no probed declaration's axiom set changed) + citation/pin
gates all green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… FamilyEvidence.total theorems

Both user-approved from the adversarial audit's findings:

- The Jalr Spec's '(cols.lsb = 0 ∨ cols.lsb = 1)' conjunct was dead weight —
  no bridge or advance consumer selected it, and the mask conjunct
  (toBitVec64 nextPcWord = ~~~1 &&& toBitVec64 value) alone pins the committed
  next_pc (the audit's worked malicious-lsb analysis). The boolean gate stays
  in the circuit; the Spec just stops re-exporting it. Soundness/completeness/
  bridge consumers repaired positionally.

- DivRem gains Unsigned64/Signed64/Unsigned32/Signed32 Evidence.total: for
  every (rs1, rs2) the family's evidence is inhabited at the canonical
  Family.result outputs — making the July audit's prose totality claim a
  standalone kernel theorem (all four axiom-clean, proved by divisor-zero/
  overflow/normal case split over the existing Math.lean closing lemmas).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…name fix

The COMMIT digest issue is now modeled at its true strength (user-authored
during the session pause; audit-reviewed):

- Contracts/CoreAIR.lean: CommitRowsSetFlags — the AIR-forced row→flag
  implications, deliberately one-way (a set rolling flag never implies row
  existence).
- Contracts/PublicValues.lean: SP1PublicValues.CommitTransitionValid — the
  four intra-shard transition laws (flag monotone, non-execution stability,
  nonzero-digest stability, committed-stability), explicitly permitting an
  unbacked 0→1 flag transition; plus the proved persistence theorem
  committedDigest_eq_last_of_flag (pure: once the flag is set, ledger
  continuity + transition laws freeze the digest to the terminal value).
- Execution.lean: FinalCommitRowsMatch / finalCommitRowsMatch_of_layout
  (inductive over the shard ledger) / finalCommitRowsMatch_of_execution /
  completeCommitDigestMatches_of_coveredExecution; the target relation gains
  commitRowsSetFlags/commitTransition; commitCovered_of_outputSafeVerifyingKey
  renamed to commitCovered_of_commitCoveringVerifyingKey.
- Soundness/CoreAIR.lean: three new CoreAIRRefinementObligations fields
  (publicCommitSetsFlag, deferredCommitSetsFlag, commitTransition) + assembly.
- gen_axiom_probe.py: dotted-name capture fix — 'theorem Foo.total' was
  collapsing to the (already-probed) structure name at the first '.', so
  dotted theorems were silently unprobed; \w+ → [\w.]+ on the DivRem globs.
- Census restamped: 465 main + 55 test = 520 (counts updated in README +
  axiom-ledger). committedDigest_eq_last_of_flag is axiom-pure; the
  execution-level layout theorems inherit only the disclosed Sail surface.

lake build SP1Clean (3645) + lake test + run_audit.sh both scopes: AUDIT PASS.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ections, zero code changes

The full-repo file-by-file sweep is complete: every hand-written Lean file
(~460 files across Math/Model/Native/FormalModel/Proofs/Faithful/Soundness/
SP1CleanTest, plus a 10-file Extracted header spot-audit) received a reviewer
pass. ZERO soundness escalations repo-wide — every finding is prose-level.
This commit applies the accumulated comment/docstring worklist (103 items,
comment-only; no statements, proofs, or deletions):

- Status-drift on closed proofs (the highest-value class): DivRem's
  'deferred/future' seam language (evidenceSoundness + completeness are
  proved); Mul + ShiftLeft/ShiftRight 'completeness seam open' ×7; the stale
  256M-heartbeat present-tense claims; TypedMemoryBalance's 'seams
  deliberately left open' (both closed, one in the same file).
- Phantom/retired names: Specs.Operation ×7, SP1WitnessDecode,
  check_heartbeats.sh, advance_of_regWrite→advance_write_core, Chain.lean→
  Walk.lean, addRow_engineFacts/addChip_rowAligned/TouchOK.pull_lt_push,
  utype/air.rs, PROOF_PATTERNS.md, five Faithful bridge spellings, the
  nonexistent Faithful/ITypeReader + Operations/AddOperation/Extracted paths.
- Pre-refactor design prose: the four bus-consistency 'Honest scope' headers
  (links now derived from balance, not threaded); reader-family witnessing
  descriptions (zero-witness input-takers); pre-#398 pull design; pre-W11
  polarity phrasing; the pre-cutover two-conjunct GatedExecution description;
  RowEffectDefs' misleading 'scheduled for deletion' banner (live-path
  interface); 'for Clean's VmTables' rationale ×13 (migration deferred).
- Wrong facts: MicroTime's pre/post-effect RAM conflation; AddChip/UType kind
  docstrings misattributing their advance routes; 'AddressOperation's 65
  columns' ×4 (it is 4 cells); Bitwise 'three' bridges (six); Coverage's
  DivRem-uncovered claim; table-33→35; one-row demo; AGENTS.md's -i-token
  justification, Trace/Witness description, and 22-of-25 Native note.
- Counts: native_decide occurrences 51→86 (now naming NonVacuityReal);
  Chips.lean/Readers.lean header enumerations; digest-model doc nits
  (AIR-forced wording, obligations-field list, stated-not-discharged).

lake build SP1Clean (0/0/0) + lake test (3692) + check_report_citations +
check_pins all green. Full per-item checklist in the campaign ledger.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… closed), extraction consolidation

Report §11.1 — the findings of the May-2026 public review of the predecessor
effort, mapped finding-by-finding to the mechanism here that addresses each
(JALR mask conjunct + lift lemmas; per-width load closure lemmas; circuit-
forced selectors + the real-row battery; UType's full stack; zero axioms/
sorryAx; pin+hash-gated extraction; the adversarial-review campaigns). Every
citation machine-checked. The 62-vs-51 opcode history is answered by the
decide-checked 50/3 coverage partition.

Docs full pass (every non-report doc): docs/README role table completed (+5
rows); overview gains the 38/13 ensemble figure, 521-census split, battery
mention, and the completeness-scope disclosure; release-audit gains the
split-census figures and a 10th trusted-components row for the opcode
hand-mirror (interim, now closed — see below); roadmap C1 gloss;
chip-standardization/bus-model/consolidation staleness; agents docs (stale
paths, retired registration language in porting-recipe, historical
maxHeartbeats framing in mul-learnings, rename-queue 40→42, allowlist header
9→8 maxRecDepth sites).

Trust-gap F8 closed: update_extracted.py gains a fail-closed opcode-table
writer (text-parse of opcode.rs AT THE SEMANTIC PIN via git show — regenerable
today, independent of overlay patch state) emitting
Extracted/OpcodeTable.lean (53 rows); the new
FormalModel/OpcodeTable.lean proves opcodeTable_matchesExtracted by plain
decide — census entry: depends on NO axioms. Census 520 → 521 (466 main + 55
test), restamped; counts updated across docs.

Extraction consolidation: generator template bugs fixed ('Reuses the no
helper module(s)' in both renderers; chip-oracle headers now name the
MemoryAccess carrier; transitive-import pruning drops 20+ superfluous reader
imports; the '#guard' comment corrected); dead legacy render path removed;
EXTRACT_AIR_ONLY regeneration verified byte-idempotent with only the intended
header/import diffs. The sibling ../sp1 branch
dtumad/lean-constraint-extraction carries the audited overlay as a proper
5-commit series (every file byte-verified against the overlay+patches;
build + smoke-tested) ready for the upstream PR that retires manual patching —
repo provenance deliberately unchanged until that merges
(docs/agents/extraction.md documents the path).

lake build SP1Clean (3647 jobs, 0/0/0) + run_audit.sh --main-only --update
(AUDIT PASS; the one census delta is the new axiom-free entry) + all fast
gates green. CI covers the untouched test scope.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
* audit semantic foundations and verifier boundary

* restamp independent audit census
@dtumad
dtumad merged commit eb6f44b into dtumad/v1.0-release Aug 12, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant