You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Reproducibility bundle for the paper 'One Prediction Set, Two Reported Results: Provenance Linkage and a Reproducible Benchmark-Audit Sequence for AI-Text Detection'. Seven-step audit sequence, executable checks, sanitized per-item predictions with SHA-256 manifest.
Faithfulness audit of miniF2F-v2 (v2s): 23 Lean 4 statements that do not say what their informal problems say, including 8 where the v2 correction introduced the defect.
Pre-registered census + provenance audit of the Turkish MMLU benchmark ecosystem (91-repo name-collision census, gating/license findings, 29-node provenance DAG) - paper + full evidence chain.
Faithfulness audit of CLEVER: seven Lean 4 specifications that do not say what their docstrings say, five of them machine-certified vacuous, with a script to reproduce the certificates.
An audit of the TRAIL benchmark's scorer. Both headline metrics divide by the gold count, so a predictor that emits every span and category scores 0.973 without reading anything.