Skip to content

fix(model): load NeuMF pretrain checkpoints with weights_only=False - #2219

Open
feiiiiii5 wants to merge 1 commit into
RUCAIBox:masterfrom
feiiiiii5:fix/neumf-pretrain-weights-only
Open

feiiiiii5 wants to merge 1 commit into
RUCAIBox:masterfrom
feiiiiii5:fix/neumf-pretrain-weights-only

Conversation

@feiiiiii5

Copy link
Copy Markdown

Fixes #2212.

Root cause

NeuMF.load_pretrain calls torch.load(path, map_location="cpu") without weights_only. Since torch 2.6 the default flipped to weights_only=True, which rejects the full pickles this library writes for GMF/MLP pretraining (state dicts wrapped with metadata) — stage 3 fails with _pickle.UnpicklingError: Weights only load fail while stages 1–2 work.

Fix

Pass weights_only=False explicitly on both loads, with a TypeError fallback to the plain call: setup.py allows torch>=1.10.0, and torch before 1.13 has no weights_only flag (where the old default already loads fully). Diff: 1 file, +9/-2.

Test

  • No existing test covers load_pretrain (searched). The failure mode is the documented torch-2.6 default flip plus the reporter's stage-3 traceback; a full 3-stage NeuMF run is heavy for a unit test — full verification left to CI.

torch>=2.6 defaults torch.load to weights_only=True, which rejects the
full pickles this library writes (state dicts wrapped with metadata).
Pass weights_only=False explicitly, with a TypeError fallback keeping
torch<1.13 working (no weights_only flag there; setup.py allows
torch>=1.10.0). Fixes RUCAIBox#2212.

Signed-off-by: fei <204683769+feiiiiii5@users.noreply.github.com>
@feiiiiii5

Copy link
Copy Markdown
Author

Checking in on this one, since it has been quiet for a couple of weeks.

The fix is one line in the NeuMF loader: torch.load(..., weights_only=False),
because the released checkpoints pickle the embedding tensors and
weights_only=True (the PyTorch 2.6 default) refuses to load them.

One question I could not answer from the issue tracker: whether you would rather
this be handled centrally, so every checkpoint load in RecBole is covered, rather
than per call site. I kept it to the one loader because that is the only path
that fails today, and a central switch would change behaviour for models that
currently load fine under the default. If you prefer the central change, I will
redo it that way.

Also glad to add a test that loads a released checkpoint if you want the coverage
in this PR rather than a follow-up.

@feiiiiii5

Copy link
Copy Markdown
Author

Checking in on this before it goes stale.

The change: NeuMF.load_pretrain calls torch.load(path, map_location="cpu") without weights_only. Since torch 2.6 the default flipped to True, which rejects the full pickles this library writes for GMF/MLP pretraining — stage 3 dies with _pickle.UnpicklingError: Weights only load fail while stages 1–2 pass, which is what makes it look like a training bug rather than a load bug.

The fix: weights_only=False explicitly on both loads, with a TypeError fallback to the plain call, because setup.py allows torch>=1.10.0 and torch before 1.13 has no weights_only flag.

What I would like to know: whether you would rather see the alternative — declaring torch>=2.6 and letting the default stand — or keeping the explicit flag. I went with the explicit flag because the repo's own floor is torch 1.10, but if you are planning to move that floor anyway this becomes unnecessary.

I am not asking you to re-read the diff to answer; it is a one-line policy question about where the compatibility floor should live.

Happy either way, and no rush if this is simply sitting lower in the queue — I will not ping again.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[🐛BUG] NeuMF.load_pretrain() incompatible with PyTorch ≥ 2.6 due to weights_only default change

1 participant