Conversation
Returns the trained alpha and lam, the q(f) marginals, the stationary targets score and info (nu_bar and lambda_bar in Opper & Archambeau 2009, Eqs. 13-14), and the negative-ELBO gradients g_nu and g_lambda (Eqs. 11-12). score and info depend only on the marginals, so they stay well defined when K is ill-conditioned; the raw alpha - score and lam - info differences do not, which is why convergence is read from g_nu and g_lambda.
bwengals
added this pull request to stack #59
September 29, 2026 20:52
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #42 (base is
vgp; GitHub retargets this tomainwhen #42 merges).Closes #43.
What it adds
pg.gp.get_vgp_point_diagnostics(vgp, fit, X, y)returns aVGPPointDiagnosticsnamedtuple aligned to the training rows:alpha,lam: trained variational parameters.fmean,fvar: q(f) marginals (fvarisdiag(S)).score,info:E_q[d log p / df]andE_q[-d^2 log p / df^2], the stationary targetsnu_barandlambda_barin Opper & Archambeau (2009), Eqs. 13-14. Computed by autodiff of the variational expectation with respect to the marginal mean and variance, so they work for any likelihood, includingCompositeLikelihood.g_nu,g_lambda: gradients of the negative ELBO with respect toalphaandlam, which equalK (alpha - score)and0.5 (S o S)(lam - info)(Eqs. 11-12).converged:fit.result.success.Also adds a
docs/agents/ptgp-vgp/skill with a per-likelihood reading table and the workflow below.Differences from the plan in #43
Checking against Opper & Archambeau (2009) and Khan, Mohamed & Murphy (NeurIPS 2012) changed a few things:
scoreandinfo, notalphaandlam. The raw residualsalpha - scoreandlam - infoonly enter the gradients throughKandS o S, so they are weakly identified when either is ill-conditioned, which is the usual case. On 40 points in [0, 5] with a lengthscale of 1 (cond(K) ~ 3.5e10),alpha - scorestayed O(1) at a converged fit whileK (alpha - score)was about 1e-5. Convergence is read fromg_nu/g_lambdainstead.lambdaand slow under L-BFGS. VGP fits here took 500 to 8000 iterations and often ended withsuccess=False. Follow-up algorithms are noted on Training techniques: natural gradients, hybrid optimizers, and more #13.lam > 0is a ptgp restriction. O-A only needK^-1 + diag(lam)to be positive definite. Add VGP model and CompositeLikelihood #42 uses softplus, so for non-log-concave likelihoodslamis pinned near zero whereinfo < 0.y = 1,infopeaks nearf = -1and turns negative for badly misclassified points (aboutf < -2.3). The boundary peak in Expose and document VGP variational parameters for fit diagnostics #43 holds only forinvlink=pt.sigmoid.Tests
g_nu == K (alpha - score)andg_lambda == 0.5 (S o S)(lam - info), withScomputed in NumPy (agree to 1e-8).lam ~ exp(fmean + fvar / 2)andscore == y - rate.info < 0andlam < 0.05.📚 Documentation preview 📚: https://ptgp--58.org.readthedocs.build/en/58/