Where the pre-extracted embeddings live, how to load them, and how to re-run extraction from scratch.
See TECHNICAL.md for the full pipeline context.
- Echo embeddings (use these first)
- ECG embeddings
- Model weights
- Encoder configs
- Reproducing extraction from scratch
All EchoJEPA / V-JEPA2 embeddings for MIMIC-IV-Echo (~525K clips) are published as a gated HuggingFace dataset:
MITCriticalData/mimic-iv-echo-jepa-embeddings
The dataset card is the source of truth for available checkpoint variants, the Parquet schema, and access requirements — it is not duplicated here. Access is gated behind an active PhysioNet credentialed account with a signed MIMIC-IV-Echo DUA.
from datasets import load_dataset
ds = load_dataset(
"MITCriticalData/mimic-iv-echo-jepa-embeddings",
data_dir="vjepa2.1-vitl-mimic-pt-100", # variant used for the reported results
)Each row carries subject_id, study_id, dicom_id, file_path, acquisition/study timestamps,
and a mean-pooled embedding (1024-d for ViT-L, 768-d for ViT-B). Join on study_id or
subject_id to align with the paired cohort from build_cohort.py.
Use
vjepa2.1-vitl-mimic-pt-100for PRIMED-AI probe experiments — it is the variant behind every reported echo number.echo-vitl-mimic117(117-epoch fine-tune) is the highest-epoch ViT-L checkpoint if you are starting a new sweep.
The same Parquet shards are mirrored on the MIT ORCD pool under
jepa-embeddings-mimiciv-echo/mimic-iv-echo-jepa-embeddings/ for jobs running on the cluster, plus
raw per-folder .pt dumps keyed by relative MP4 path (p{subject_id}/{study_id}/{dicom_id}.mp4):
import torch
# Dict[str, torch.Tensor], one entry per clip; p10–p19 folders, or *_all.pt merged (~2.1 GB for ViT-L)
embeddings = torch.load("vitl_embeddings_p10.pt", map_location="cpu")
X = torch.stack(list(embeddings.values())) # [N, embed_dim]The reported ECG numbers come from pre-pooled HuBERT-ECG embeddings (768-d), read directly from
Parquet — path in configs/encoder/ecg_fm.yaml under
hubert_ecg_parquet. The model is
Edoardo-BS/hubert-ecg-base; its model card
covers loading and the pretraining corpus.
encoders/ecg_fm.py wraps
ECG-FM as an alternative extraction path, driven by
scripts/extract_ecg_embeddings.py. It downloads weights via
huggingface_hub and needs the fairseq_signals stack. It is implemented and unit-tested but has
not produced any reported number. Both paths are 768-d after pooling, so probes accept either.
Echo checkpoints live on ORCD under weights/. Access requires ORCD cluster membership — contact
the PRIMED-AI leads.
| File | Model alias | embed_dim | Source |
|---|---|---|---|
vitl.pt |
vitl |
1024 | dl.fbaipublicfiles.com/vjepa2/vitl.pt |
vith.pt |
vith |
1280 | dl.fbaipublicfiles.com/vjepa2/vith.pt |
vitg.pt |
vitg |
1408 | dl.fbaipublicfiles.com/vjepa2/vitg.pt |
vitg-384.pt |
vitg-384 |
1408 | dl.fbaipublicfiles.com/vjepa2/vitg-384.pt |
| File | Model alias | embed_dim | Description |
|---|---|---|---|
vitl-scratch-pt-210-c25.pt |
echo-vitl-scratch |
1024 | ViT-L trained from scratch on echo, 210 epochs |
vjepa21_vitl_mimic_pt100.pt |
echo-vitl-mimic100 |
1024 | V-JEPA2.1 ViT-L fine-tuned on MIMIC-IV-Echo, 100 epochs |
vjepa21_vitl_mimic_pt117.pt |
echo-vitl-mimic117 |
1024 | V-JEPA2.1 ViT-L fine-tuned on MIMIC-IV-Echo, 117 epochs |
vitl-vmix22m-pt220-c55.pt |
echo-vitl-vmix22m |
1024 | ViT-L pretrained on VMix-22M (220 epochs) + MIMIC fine-tune |
vjepa2_1_vitb_mimic_pt169_c60.pt |
echo-vitb-mimic169 |
768 | V-JEPA2.1 ViT-B fine-tuned on MIMIC-IV-Echo, 169 epochs |
Hydra configs live in configs/encoder/: echojepa.yaml (checkpoint path,
embed_dim, frame count, pooling) and ecg_fm.yaml (ECG-FM wrapper settings plus the
hubert_ecg_parquet path the probes actually read). Both set frozen: true — foundation-model
weights are never fine-tuned.
Only needed for a checkpoint variant not already published. Scripts live in
scripts/embedding_extraction/.
3-stage pipeline:
| Stage | Script | Output |
|---|---|---|
| 1 | convert_dicom.py |
~525K MP4 files, 256×256 |
| 2 | extract_embeddings.py (SLURM array via extract_echo_slurm.sh) |
.pt per folder |
| 2b | merge_embeddings.py |
merged *_all.pt |
| 3 | to_parquet.py |
sharded Parquet with MIMIC metadata joined |
module load miniforge/24.3.0-0
conda activate vjepa2-312
sbatch scripts/embedding_extraction/extract_echo_slurm.sh echo-vitl-mimic117
python scripts/embedding_extraction/merge_embeddings.py --model echo-vitl-mimic117
python scripts/embedding_extraction/to_parquet.py --model echo-vitl-mimic117Model aliases: vitl, vith, vitg, vitg-384, echo-vitl-scratch, echo-vitl-mimic100,
echo-vitl-mimic117, echo-vitl-vmix22m, echo-vitb-mimic169.
| Field | Value |
|---|---|
| Partition | mit_preemptable |
| GPU | gpu:l40s:1 per task |
| Array | 0–9 (folders p10–p19) |
| Max concurrent | 4 tasks (QOSMaxGRESPerUser cap) |
| Job time limit | 2 days (preemptable; jobs auto-resume via cache) |