Kind
Question
What qualitative errors distinguish MarinFold's predicted contact maps from the correct maps on all 97 eval-val proteins, and which errors plausibly explain its limited accuracy?
Hypothesis
Errors may reflect missing or misplaced long-range contact blocks, diffuse ensemble votes, or local register shifts. Their relative contribution should be measured rather than inferred from aggregate precision alone.
Background
Build on #245's frozen eval-val universe, #277's saved rollout-and-resample predictions, and #339's diversity/accuracy investigation.
Approach
Reuse saved predictions from an explicitly identified checkpoint. Reproduce canonical per-protein metrics against the frozen resolved-residue universe, then inspect representative low-, medium-, and high-accuracy maps. Quantify contact-range bias, local near misses, confidence and support for missed true contacts. Build a self-contained HTML report with predicted and true maps side by side for all 97 proteins, sortable metrics, error overlays and explicit coordinate/threshold conventions. Use eval-val only; no model execution or eval-test read is needed.
Success criteria
All 97 eval-val proteins present with verified provenance and canonical scoring. A browsable HTML report, per-protein diagnostic CSV, selected annotated cases, reproducible analysis, and conclusions that distinguish observed patterns from causal hypotheses.
Kind
Question
What qualitative errors distinguish MarinFold's predicted contact maps from the correct maps on all 97 eval-val proteins, and which errors plausibly explain its limited accuracy?
Hypothesis
Errors may reflect missing or misplaced long-range contact blocks, diffuse ensemble votes, or local register shifts. Their relative contribution should be measured rather than inferred from aggregate precision alone.
Background
Build on #245's frozen eval-val universe, #277's saved rollout-and-resample predictions, and #339's diversity/accuracy investigation.
Approach
Reuse saved predictions from an explicitly identified checkpoint. Reproduce canonical per-protein metrics against the frozen resolved-residue universe, then inspect representative low-, medium-, and high-accuracy maps. Quantify contact-range bias, local near misses, confidence and support for missed true contacts. Build a self-contained HTML report with predicted and true maps side by side for all 97 proteins, sortable metrics, error overlays and explicit coordinate/threshold conventions. Use eval-val only; no model execution or eval-test read is needed.
Success criteria
All 97 eval-val proteins present with verified provenance and canonical scoring. A browsable HTML report, per-protein diagnostic CSV, selected annotated cases, reproducible analysis, and conclusions that distinguish observed patterns from causal hypotheses.