Skip to content

Add VGP point diagnostics and ptgp-vgp skill - #58

Open
bwengals wants to merge 2 commits into
vgpfrom
vgp-diagnostics
Open

bwengals wants to merge 2 commits into
vgpfrom
vgp-diagnostics

Conversation

@bwengals

@bwengals bwengals commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Stacked on #42 (base is vgp; GitHub retargets this to main when #42 merges).

Closes #43.

What it adds

pg.gp.get_vgp_point_diagnostics(vgp, fit, X, y) returns a VGPPointDiagnostics namedtuple aligned to the training rows:

  • alpha, lam: trained variational parameters.
  • fmean, fvar: q(f) marginals (fvar is diag(S)).
  • score, info: E_q[d log p / df] and E_q[-d^2 log p / df^2], the stationary targets nu_bar and lambda_bar in Opper & Archambeau (2009), Eqs. 13-14. Computed by autodiff of the variational expectation with respect to the marginal mean and variance, so they work for any likelihood, including CompositeLikelihood.
  • g_nu, g_lambda: gradients of the negative ELBO with respect to alpha and lam, which equal K (alpha - score) and 0.5 (S o S)(lam - info) (Eqs. 11-12).
  • converged: fit.result.success.

Also adds a docs/agents/ptgp-vgp/ skill with a per-likelihood reading table and the workflow below.

Differences from the plan in #43

Checking against Opper & Archambeau (2009) and Khan, Mohamed & Murphy (NeurIPS 2012) changed a few things:

  • Read score and info, not alpha and lam. The raw residuals alpha - score and lam - info only enter the gradients through K and S o S, so they are weakly identified when either is ill-conditioned, which is the usual case. On 40 points in [0, 5] with a lengthscale of 1 (cond(K) ~ 3.5e10), alpha - score stayed O(1) at a converged fit while K (alpha - score) was about 1e-5. Convergence is read from g_nu / g_lambda instead.
  • Slow convergence is normal. KMM show the O-A parameterization is non-concave in lambda and slow under L-BFGS. VGP fits here took 500 to 8000 iterations and often ended with success=False. Follow-up algorithms are noted on Training techniques: natural gradients, hybrid optimizers, and more #13.
  • lam > 0 is a ptgp restriction. O-A only need K^-1 + diag(lam) to be positive definite. Add VGP model and CompositeLikelihood #42 uses softplus, so for non-log-concave likelihoods lam is pinned near zero where info < 0.
  • Bernoulli correction. The default is probit, not logit. With ptgp's 0.001 label-flip floor it is also not log-concave: for y = 1, info peaks near f = -1 and turns negative for badly misclassified points (about f < -2.3). The boundary peak in Expose and document VGP variational parameters for fit diagnostics #43 holds only for invlink=pt.sigmoid.

Tests

  • Eqs. 11-12 at a non-stationary point: g_nu == K (alpha - score) and g_lambda == 0.5 (S o S)(lam - info), with S computed in NumPy (agree to 1e-8).
  • Poisson at convergence: lam ~ exp(fmean + fvar / 2) and score == y - rate.
  • Student-t outlier: info < 0 and lam < 0.05.

📚 Documentation preview 📚: https://ptgp--58.org.readthedocs.build/en/58/

Returns the trained alpha and lam, the q(f) marginals, the stationary
targets score and info (nu_bar and lambda_bar in Opper & Archambeau 2009,
Eqs. 13-14), and the negative-ELBO gradients g_nu and g_lambda (Eqs. 11-12).
score and info depend only on the marginals, so they stay well defined when
K is ill-conditioned; the raw alpha - score and lam - info differences do
not, which is why convergence is read from g_nu and g_lambda.
@bwengals
bwengals added this pull request to stack #59 September 29, 2026 20:52

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant