Skip to content

exp: analyze MarinFold vs ESMFold2 performance across PDB-deduped monomers #324

Description

@znichols

Kind

  • Kind: evals

Question

Given matched ESMFold2 and MarinFold contact scores on the same 10k PDB-deduped experimental monomers, what biology/structure/provenance features explain the MarinFold-vs-ESMFold2 gap? Which protein subgroups does MarinFold perform relatively well or poorly on?

Hypothesis

ESMFold2 will remain the strongest aggregate predictor, but MarinFold will close the gap on identifiable protein slices where language-model priors are useful or single-sequence structure prediction is less dominant: synthetic/designed proteins, viral proteins, rare experimental modalities such as Solution NMR, and specific GO/CATH structural-function groups. Conversely, highly represented natural structural classes and some complex cellular machinery should be ESMFold2-favorable.

Background

This experiment builds a per-protein feature/delta analysis comparing ESMFold2 n_samples=100 against MarinFold 100-rollout checkpoints on the same PDB-deduped experimental monomer manifest. The central object is a 10k-row table with matched contact metrics and compact annotations for taxonomy, function, structural class, provenance, experimental method, and homolog-depth proxy.

Key comparisons:

  • ESMFold2 n_samples=100
  • contacts-v2 / delta-stream, 100 rollouts
  • exp277 full-epoch contacts-v1 baseline, 100 rollouts
  • exp117 old-recipe contacts-v1 baseline, 100 rollouts

Approach

Protein sample and ground truth contacts

  • Start from the experimental-PDB contacts-v1 deduped monomer corpus.
  • Sample a fixed 10,000-protein manifest with length filter 50 <= L <= 1024 and seed 32410.
  • Decode/materialize the corpus's emitted contacts-v1 PDB contacts as the ground-truth contact set.
  • Ground truth is therefore experimental PDB pyconfind contacts, not ESMFold2 predictions.
  • Contact convention: contact_degree >= 0.001 and abs(seq_i - seq_j) >= 6.

ESMFold2 predictions and scoring

  • Run ESMFold2 on CoreWeave H100 over the fixed 10k manifest.
  • For n_samples=100, persist predicted structures/provenance/timings to CoreWeave S3.
  • Convert predicted structures back into pyconfind contact-degree rankings.
  • Score those rankings against the same experimental-PDB GT contacts using the same R-precision/AUC machinery used for MarinFold.

Durable ESMFold2 scoring prefixes:

s3://marin-us-east-02a/protein-structure/MarinFold/exp324_esmfold2_pdb_sample/pdb_deduped_10k_n100_scored/score_shards/
s3://marin-us-east-02a/protein-structure/MarinFold/exp324_esmfold2_pdb_sample/pdb_deduped_10k_n100_scored_rescue_lesshalf_20260924/score_shards/

MarinFold rollout scoring

  • Build the same 10k target set from the fixed manifest and GT contacts.
  • Run 100-rollout contacts-v1/contact-v2 scoring for each checkpoint.
  • Store sparse vote/contact scores and aggregate per-protein metrics.
  • Compare only on the same proteins / same GT / same metric definitions as ESMFold2.

Durable MarinFold scoring prefix:

s3://marin-us-east-02a/protein-structure/MarinFold/exp324_esmfold2_pdb_sample/marinfold_10k_rollout_scores/

Feature annotation sources

The compiled analysis table joins the 10k manifest and performance metrics with compact feature annotations from:

  • RCSB/PDB metadata: PDB ID, chain, experimental method, resolution, titles/keywords, source organism.
  • Experimental-PDB pyconfind contacts: GT contact counts and contact-degree summaries.
  • NCBI taxonomy: lineage levels such as Eukaryota, Bacteria, Viruses, Riboviria, etc.
  • UniProt / UniProt-GOA: accession mapping and GO annotations.
  • GO-slim: compact functional/process/location tag set derived from GO annotations.
  • CATH: domain-level structural class/topology tags.
  • UniRef100/90/50 cluster sizes: cheap homolog-depth proxy; used instead of full ColabFold/MMseqs MSA depth after the latter proved too slow/rate-limited.
  • Synthetic/provenance heuristics: synthetic taxonomy plus PDB title/keyword evidence for engineered/de novo/designed constructs.

Success criteria

  • Produce a fixed 10k manifest and reproducible scoring pipeline for ESMFold2 n_samples=100 and MarinFold 100-rollout checkpoints.
  • Compile per-protein R-precision/AUC metrics with compact feature annotations for taxonomy, GO-slim, CATH, experimental method, synthetic status, and UniRef depth proxy.
  • Identify interpretable feature groups where MarinFold closes the ESMFold2 gap, and distinguish direction (MarinFold-favorable vs ESMFold2-favorable) from feature importance.
  • Keep generated statistics reproducible from tracked scripts and durable remote result prefixes rather than checking large generated outputs into git.

Results so far

Headline all-range R-precision on the shared 10k sample:

predictor all R
ESMFold2 n_samples=100 0.7224
exp277 full-epoch contacts-v1, 100 rollouts 0.5338
contacts-v2 / delta-stream, 100 rollouts 0.4909
exp117 old-recipe contacts-v1, 100 rollouts 0.3953

Back-of-the-envelope ANOVA/OLS feature importance ranks features by partial eta-squared for the paired gap MarinFold R - ESMFold2 R, adjusted for length, emitted-contact count, resolution, and UniRef50 depth. Direction is the fitted effect sign: positive is MarinFold-favorable, negative is ESMFold2-favorable.

Top contacts-v2 / delta-stream feature-importance hits:

feature family feature/tag direction partial eta² readout
taxonomy lineage_level_2 = synthetic_construct +0.202 0.0537 contacts-v2-favorable
taxonomy lineage_level_1 = synthetic +0.188 0.0529 contacts-v2-favorable
GO-slim organelle -0.090 0.0365 ESMFold2-favorable
CATH class alpha_beta +0.066 0.0213 contacts-v2-favorable
GO-slim ribosome -0.151 0.0160 ESMFold2-favorable
method SOLUTION NMR +0.113 0.0149 contacts-v2-favorable
CATH topology 3.40.50 +0.079 0.0148 contacts-v2-favorable
GO-slim structural molecule activity -0.116 0.0141 ESMFold2-favorable

Top exp277 feature-importance hits:

feature family feature/tag direction partial eta² readout
taxonomy lineage_level_2 = synthetic_construct +0.202 0.0303 exp277-favorable
taxonomy lineage_level_1 = synthetic +0.170 0.0276 exp277-favorable
GO-slim ribosome -0.162 0.0229 ESMFold2-favorable
GO-slim organelle -0.062 0.0217 ESMFold2-favorable
method SOLUTION NMR +0.135 0.0210 exp277-favorable
GO-slim structural molecule activity -0.118 0.0180 ESMFold2-favorable
CATH class alpha_beta +0.053 0.0170 exp277-favorable
CATH topology 3.40.50 +0.059 0.0100 exp277-favorable

Useful gap-closing slices:

category n ESMFold2 R contacts-v2 R exp277 R contacts-v2 win % exp277 win %
all proteins 9,943 0.722 0.491 0.534 9.1 11.1
synthetic 92 0.727 0.595 0.637 26.1 32.6
Solution NMR 622 0.558 0.430 0.485 29.1 38.9
Riboviria 195 0.389 0.212 0.242 23.1 26.7
Duplodnaviria 175 0.553 0.419 0.448 20.0 24.0
CATH topology 1.20.5 70 0.484 0.399 0.409 18.6 24.3

Caveat: ESMFold2 remains ahead on mean R-precision in these categories; the useful signal is smaller gap and higher MarinFold per-protein win fraction.

Notes

Generated statistics, parquet tables, LDA outputs, and plots are not checked into git. The branch keeps the fixed manifests plus scripts/notebook snippets. Rebuild the feature/performance table from the tracked scripts and durable S3 outputs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    agent-generatedOpened by an AI agentexperimentResearch experiment tracked under experiments/kind/evalsExperiment kind: runs evals on trained models

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions