Add real data examples - #696
Merged
Merged
Conversation
…rness tests/integration_tests.py runs each example via exec(contents, globals, locals) with separate globals/locals dicts. Under that scheme, module-level assignments land in `locals`, but nested scopes (def bodies and comprehensions) resolve free variables against `globals`, which only holds `jnp` and the injected `gpx`. The Barcelona-solar notebook tripped this twice: - helper functions (`hours_to_standardised`, `unstandardise_*`) closed over `hour_mean`/`irradiance_*` and used `np` -> NameError. Bind the constants as default arguments and use `jnp` (guaranteed present) in the bodies. - the empirical hourly mean/std list comprehensions closed over `irradiance`/ `hour_of_day` -> NameError. Rewrite as an explicit top-level loop. Numerics are unchanged (the `history` trajectory is identical); verified the notebook execs cleanly through the exact harness code path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FGBs2UN5iP8kBKuwqZwqKk
The docs build converts notebooks with `docs/_examples` as the working directory and copies `examples/data/` alongside as `data/`, so the hard-coded `examples/data/max_tempeature_switzerland.csv` path raised FileNotFoundError during nbconvert. Fall back to the `data/` path, matching the pattern used by the other real-data notebooks. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FGBs2UN5iP8kBKuwqZwqKk
thomaspinder
temporarily deployed
to
docs-preview
July 12, 2026 07:50 — with
GitHub Actions
Inactive
Three docs follow-ups:
1. poisson.py — the ragged lower credible band pinned near zero was an artifact
of the sampler never being adapted: `num_adapt` was defined but unused, so
NUTS ran with a fixed 1e-3 step size and identity mass matrix (acceptance
~0.9996, near-zero mixing), occasionally throwing latent rates ~0. Run
blackjax window adaptation, split the predictive PRNG keys per draw, and plot
a smooth rate-lambda credible band plus a lighter count-predictive band. The
remaining mild asymmetry is the genuine exp-link effect (lambda floored at 0).
2. Restore the original synthetic `spatial_linear_gp.py` and move the Swiss
MeteoSwiss kriging example into a new `semiparametric_kriging.py` ("Semi-
Parametric Kriging"), extending the docs rather than replacing the tutorial.
3. Fix misrendering LaTeX degree units in the kriging prose by using Unicode °C
(matches the other notebooks and the existing f-string output).
Wired the new notebook into mkdocs nav and the gen_examples CTA set. Verified
poisson and semiparametric_kriging execute end-to-end.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FGBs2UN5iP8kBKuwqZwqKk
The docs job executes every example notebook via `gen_examples.py --execute`,
serially. This PR adds three new notebooks (temperature_forecasting,
climate_extrapolation, semiparametric_kriging) to an already near-budget build,
pushing total execution past the 30-minute cap (the job was cancelled mid-run,
not errored).
Run the executions in parallel across the runner's 4 vCPUs. Also fix the
`--parallel` argument, which used `type=bool` (so `bool("False")` was truthy);
it is now a proper `store_true` flag. Applied to both the PR docs check and the
main docs deploy, since the latter builds the same notebooks.
Verified locally: all 25 notebooks build successfully with `--parallel
--max_workers 4` (exit 0, no OOM).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FGBs2UN5iP8kBKuwqZwqKk
Reverts the parallel-execution attempt: on a 4-vCPU runner, four concurrent notebook processes each let JAX/XLA grab all cores, so they oversubscribe and thrash rather than speed up — the build still hit the 30-minute wall. The suite genuinely grew: this PR adds two new notebooks and restores the (heavier, synthetic) spatial_linear_gp alongside the new Swiss semiparametric_kriging, so there are now two spatial NUTS notebooks where the 26.5-minute passing build had one. Fix it directly: - Raise the docs-job timeout 30 -> 45 minutes (PR check and main deploy). - Trim semiparametric_kriging: NUTS warm-up/samples 1000/1500 -> 500/1000 for both models and grid 45 -> 35. Runtime roughly halves locally (27.5s -> 14.9s) with unchanged results (lapse rate 7.05 C/km, RMSE 1.49 -> 0.885). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FGBs2UN5iP8kBKuwqZwqKk
thomaspinder
temporarily deployed
to
docs-preview
July 12, 2026 10:14 — with
GitHub Actions
Inactive
Clip the posterior mean/std temperature fields to the national border and draw the outlines of Switzerland and its neighbours for geographic context. Uses a matplotlib-only overlay (stdlib json + matplotlib Path/PathPatch) of a bundled Natural Earth 1:50m boundary (examples/data/switzerland_neighbours.geojson, 48 KB, public domain) — no geopandas/shapely dependency, and the data path resolves under both the repo root and the docs build directory. The grid now spans Switzerland's bounding box so the clipped field fills the country, with a latitude-corrected aspect ratio. Results unchanged (lapse rate 7.05 C/km, RMSE 1.49 -> 0.885). Verified end-to-end through the jupytext -> nbconvert docs path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FGBs2UN5iP8kBKuwqZwqKk
thomaspinder
temporarily deployed
to
docs-preview
July 12, 2026 11:38 — with
GitHub Actions
Inactive
On the mean map the station markers were coloured by raw t_max (spanning -11..16.7 C across 205-3582 m of elevation) while the field shows t_max at the mean reference elevation (~4-9 C) — the same colormap on two different, auto- scaled ranges, so markers and field did not line up and read as inconsistent. Colour the markers by elevation-adjusted t_max (subtracting the fitted altitude effect) and share a single colour normalisation with the field. Markers now match the surface they sit on, so the plot honestly shows the GP interpolating the adjusted observations, with departures visible as residuals. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01FGBs2UN5iP8kBKuwqZwqKk
thomaspinder
temporarily deployed
to
docs-preview
July 12, 2026 12:45 — with
GitHub Actions
Inactive
thomaspinder
temporarily deployed
to
docs-preview
July 12, 2026 14:16 — with
GitHub Actions
Inactive
thomaspinder
temporarily deployed
to
docs-preview
July 13, 2026 20:07 — with
GitHub Actions
Inactive
thomaspinder
temporarily deployed
to
docs-preview
July 13, 2026 22:05 — with
GitHub Actions
Inactive
thomaspinder
temporarily deployed
to
docs-preview
July 14, 2026 19:03 — with
GitHub Actions
Inactive
thomaspinder
temporarily deployed
to
docs-preview
July 14, 2026 20:17 — with
GitHub Actions
Inactive
thomaspinder
added a commit
that referenced
this pull request
Aug 5, 2026
The newly-loud harness exposed pre-existing drift in all four examples: the collapsed/uncollapsed goldens predated the real-data example swap (#696) and the regression/heteroscedastic goldens predated subsequent behaviour fixes (#707/#708/#713 and dependency bumps). The toothless harness never noticed. Re-pinned so the net measures the v1.0 refactor, not history. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
thomaspinder
added a commit
that referenced
this pull request
Sep 28, 2026
* test(oracle): pin conjugate MLL, predict, and LOOCV to closed form
conjugate_mll had no value-level test anywhere in the suite, yet serves as
the oracle for the Kalman MLL and collapsed_elbo. These closed-form pins,
computed through an independent jnp.linalg path, give the reference frame
its ground truth ahead of the v1.0 conditioning refactor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* test(equivalence): pin cross-derivation agreement; xfail two-owner jitter bug
collapsed_elbo(z=X) vs conjugate_mll and whitened-vs-unwhitened predicts at
matched parameters now guard the five independent derivations of the
conjugate conditioning algebra. The strict xfail documents that at
non-default jitter the derivations factorise different matrices — the bug
the v1.0 conditioning module removes.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* test(integration): fail loudly when golden values drift
_compare previously swallowed AssertionError with a print, so the harness
could never fail. Failures are now collected per-example and raised at the
end of test(), making the four golden-value pins a real no-behaviour-change
net for the v1.0 refactor.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* test(integration): re-pin golden values to current-main behaviour
The newly-loud harness exposed pre-existing drift in all four examples: the
collapsed/uncollapsed goldens predated the real-data example swap (#696) and
the regression/heteroscedastic goldens predated subsequent behaviour fixes
(#707/#708/#713 and dependency bumps). The toothless harness never noticed.
Re-pinned so the net measures the v1.0 refactor, not history.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* feat(linalg): stabilised_cholesky — single stabilise-and-factor seam
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* feat(dataset): static n_total metadata; get_batch stamps full size
The full-dataset size a minibatch ELBO needs now travels on the one object
that knows it, as static pytree aux_data, instead of being smuggled through
likelihood constructors.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* feat(likelihoods)!: pure conditional families — drop num_datapoints and noise_prior
Of ~245 occurrences of num_datapoints, only four were real reads: the ELBO
minibatch scale (now served by Dataset.n_total) and latent sizing (moves to
data-contact time in the JointModel rewrite). No likelihood used the value
internally, and nothing validated it — a wrong value silently mis-scaled
the ELBO. noise_prior moves to the model layer, where priors live.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* feat!: v1.0 conditioning architecture — JointModel, Posterior, condition
The API now mirrors the maths: prior * likelihood -> JointModel (the joint
p(f,y), the trainable object); model.condition(D) — sugar: model | D —
returns an immutable Posterior pytree caching the Cholesky factor and
representer weights. The predictive, log_marginal_likelihood, loo, and
pathwise sample_approx are views of that one factorisation, deleting the
eleven independent derivations and the two-owner jitter split
(prior.jitter is now the single knob, applied once inside conditioning).
- gpjax/conditioning.py: deep module (Posterior, ExactPosterior,
LatentPosterior); MO validation moves to condition time; sample_approx
refuses multi-output loudly instead of silently broadcasting wrong.
- gps.py: Prior (AbstractPrior folded in), ConjugateModel,
NonConjugateModel (lazy latent, sized at data contact),
HeteroscedasticModel (owns noise_prior — likelihoods are pure
conditionals again, killing the likelihoods->gps circular import).
Deleted: AbstractPrior, AbstractPosterior, LatentPosterior marker,
ChainedPosterior marker, construct_posterior (now construct_model).
- objectives: conjugate_mll/conjugate_loocv/log_posterior_density are
one-line views of the conditioned posterior.
- fit: _prepare_model hook sizes lazily-initialised state from data.
- predict(t, D) survives as documented one-line sugar everywhere.
- return_covariance_type kwarg renamed to covariance.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* refactor!: rename sweep — tests, benchmarks, examples onto the v1.0 API
Mechanical: ConjugatePosterior->ConjugateModel and friends,
construct_posterior->construct_model, return_covariance_type->covariance,
num_datapoints/noise_prior constructor ceremony deleted (~240 sites).
Semantic: heteroscedastic tests build HeteroscedasticModel directly;
non-conjugate tests size the latent via init_latent; docs/index.md
quickstart shows condition(); regression example narrates the
condition API; StateSpaceConjugatePosterior renamed
StateSpaceConjugateModel.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* style: ruff format after the rename sweep
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* docs: v1.0 migration guide, CONTEXT.md glossary, ADR-0001
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* docs: self-contained migration snippets
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* test(equivalence): pin predict/MLL same-matrix consistency at non-default jitter
The model-side two-owner jitter bug is fixed; the strict xfail narrows to
the family-side knob, which unifies in the variational stack PR.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* docs: fix Sphinx build for the v1.0 API
- reference/gps.md lists the JointModel hierarchy and the conditioning
module; state_space.md and linalg.md updated for renamed/new symbols
- stale glossary/sharp_bits/classification xrefs renamed
- poisson example initialises the lazy non-conjugate latent before MCMC
- ADR directory excluded from the docs site (in-repo records for now)
- codeautolink match_block warnings suppressed on every path: a matcher
limitation on doctest-SKIP blocks, predating this stack — the docs
workflow had not run cold since the Sphinx migration, so tonight's PR
pushes surfaced it for the first time
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* feat(natural-gradients): add fit_natgrads and exponential-family machinery
Implements Salimbeni et al. 2018 (arXiv:1803.09151) natural-gradient VI for
VariationalGaussian and WhitenedVariationalGaussian.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* fix(natural-gradients): address review findings on PR#1
Adversarial review of the natural-gradient core turned up two correctness
defects, one performance defect, three contract mismatches and a set of
documentation and test gaps. Fixes, in order of severity:
* `_first_valid_trial` leaked float64 into the scan carry. Under x64 the
exponent `jnp.arange(K + 1)` is int64, so `backoff ** arange` was a non-weak
float64 that promoted a float32 model and made `lax.scan` reject the carry.
The trial ladder is now cast to the dtype of Theta_2, so `fit_natgrads` is no
longer strictly narrower than `fit`.
* The backoff replicated the whole theta -> xi map across all K+1 trials, at a
measured 13% of total training wall clock at M=200 -- not the "negligible"
cost impl-plan 1.2.5 assumed. Only the admissibility probe is replicated now;
the inversion, the X^T X product and the second Cholesky run once, at the
accepted step size. Measured overhead at M=200 falls to 4.6%.
* `fit_natgrads`' signature rejected everything `_check_natgrad_lr` blessed:
`natgrad_lr=1`, `map_jitter=0`, `backoff=1` and a 0-d array all raised under
the beartype import hook. Annotations widened, the validator now accepts 0-d
arrays and rejects bool, and the entry point (not just the validator) is
tested with each.
* `_reject_frozen_coordinates` matched coordinates against top-level dataclass
fields by identity, so a future nested registration would have passed the
guard silently. It now re-walks the tree with the selector as the `is_leaf`
predicate and reports the full key path. Its message also pluralises and
points at a remedy that exists -- the old one recommended freezing the whole
family, which re-raises the same error.
* `fit_natgrads` now calls the guard under `safe=True` as impl-plan 1.2.4
prescribes, and forwards `log_rate` to `vscan` instead of documenting a knob
that did nothing.
Tests: `test_natgrad_backoff_recovers_from_large_step` never exercised the
backoff (the k=0 trial was already admissible at natgrad_lr=100), so it now
starts from a tight S_0, asserts the un-shrunk step genuinely leaves the cone,
and checks the accepted step size. The Cholesky-budget test could not see the
vmap width; it is now an exact count parametrised over max_backoff, plus a
lowered-IR test asserting the batched Cholesky has leading dimension K+1. Added
dtype-preservation tests, and moved the duplicated conjugate oracle into
`tests/_reference/conjugate_svgp.py` so the two transcriptions cannot drift.
Docs: all seven exported functions gained runnable `Example:` blocks (impl-plan
section 6), the map_jitter bias on `history` and the eta -> xi cancellation
regime are now documented on the public surface, and two docstrings became raw
strings.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* refactor(natgrads): adapt to the v1.0 conditioning API
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* docs(fit): replace Markdown code fences with RST doctest block in fit_natgrads
The ```pycon fences render fine under MkDocs but are invalid RST under
Sphinx/napoleon, producing docutils warnings that fail the -W docs-ci
gate. Drop the fences; the indented doctest block matches fit()'s style.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* docs(examples): add dual/t-SVGP tutorial notebook and wire docs nav
Adds examples/dual_svgp.py, the second of the two natural-gradient tutorial
notebooks, and wires both of them into the documentation.
The notebook derives the dual parameterisation for a reader who has been
through examples/natgrads.py: the additive split eta = eta_0(theta) + lambda,
the EP-style likelihood sites and their tying to inducing space, the two
convention traps (flanked vs un-flanked storage, and the -1/2 on Lambda_2),
and the tied natural-gradient update, which is an affine convex combination on
the stored sites because grad_mu KL == lambda exactly, so the KL is never
differentiated and no theta <-> eta round trip is needed.
Measured in the executed run:
* one rho = 1 full-batch step on a conjugate model with a non-zero mean
function reproduces the Titsias optimum to 1.7e-12 (mean) and 1.2e-13
(covariance), a second step moves nothing, and dual_elbo matches the
collapsed bound up to exactly N * jitter / (2 sigma^2);
* rho is gamma: matched dual and Salimbeni E-steps agree in (m, S) to 3.1e-15
over six full-batch steps at rho in {0.3, 0.8, 1.0}, and two frozen-
hyperparameter fit_natgrads runs overlay to 4.3e-14 over 50 iterations;
* the banana benchmark (N = 2000, M = 50, B = 256, 1000 iterations, the same
make_banana and jr.key(42) as the natural-gradients notebook) with t-SVGP,
the Salimbeni natgrad and Adam alone, per iteration and per wall-clock
second -- both natural-gradient runs reach Adam's 1000-iteration bound at
iteration 109, and the dual iteration is not cheaper at this scale;
* hyperparameter learning: the dual/standard hyper-gradient gap falls from
3.9e1 to 8.5e-15 as the E-step converges; dual dominance is not uniform when
sparse (negative gaps at M = 5 and M = 10) but holds by 1.5e5 nats at Z = X;
and a 40-round VEM loop ends 0.50 nats ahead on dual_elbo with equal
held-out NLPD.
Docs wiring: both notebooks added to the mkdocs.yml Tutorials nav and to
CTA_NOTEBOOKS in docs/scripts/gen_examples.py, plus an adam2021dual entry in
docs/refs.bib.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* docs(examples): address dual notebook review findings
Corrections to examples/dual_svgp.py, all re-verified against a fresh
end-to-end execution:
* the intro no longer implies the two hyperparameter gradients agree at
theta_t; they agree only at a converged E-step, which is what the
notebook's own gradient table measures;
* added a notation-reconciliation note bridging the natural-gradients
notebook's (theta, eta, lambda) to this one's (eta, mu, lambda), and
restated the borrowed identity and H_2 in these letters;
* the c(theta) remark now names the site convention it holds under
(normalised projected) and fixes the sign apposition;
* the bound-slice prose is now asymmetric, as the data are: l collapses
on the long-lengthscale side while l-bar barely moves, but both
collapse together on the short side, where the sparse approximation
itself has failed. Added edge diagnostics to back it;
* the banana benchmark now states which run ends ahead and bounds what
that comparison can mean;
* the VEM panel plots the round-by-round bound lead rather than two
indistinguishable traces. That exposed a false claim: the dual M-step
is behind for the first seven rounds, crosses at round 8 and holds a
sub-nat lead thereafter. Prose corrected and the crossing printed;
* order-of-magnitude claims restated from the printed values (Titsias
agreement, the jitter residual now printed to twelve digits, the
banana condition-number ratio, the M = 20 crossing-point noise);
* the dominance row now names conjugacy as well as Z = X;
* the roadmap names the two sections it had omitted;
* the banana-copy rationale no longer overstates cross-notebook
comparability, and the wall-clock explanation leads with the
O(M^3)-vs-O(BM^2) argument rather than an evaluation XLA folds away;
* ruff-format clean, no code line over 88 characters.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* refactor(variational)!: remove NaturalVariationalGaussian and ExpectationVariationalGaussian
Both classes were parameterisation-only: they stored the natural or
expectation coordinates of q(u) but shipped no way to take a
natural-gradient step in them, so they bought nothing over the standard
families.
Natural-gradient geometry belongs to the optimiser, not the family. The
Fisher matrix is exactly the Jacobian dn/dt, so the natural gradient with
respect to the natural parameters equals the ordinary gradient with
respect to the expectation parameters, in any parameterisation. fit_natgrads
(PR#1, gpjax/natural_gradients.py) therefore computes the transforms
on the fly and operates directly on VariationalGaussian and
WhitenedVariationalGaussian, which store constraint-respecting
coordinates. Users of the removed classes should switch to
VariationalGaussian with gpjax.fit_natgrads.
Also drops the now-dead _psd helper and the cholesky_factor import, whose
only call sites lived inside the deleted classes, and the "natural" and
"expectation" arms of the VariationalParametrisationSuite ASV benchmark.
BREAKING CHANGE: NaturalVariationalGaussian and ExpectationVariationalGaussian
are removed from gpjax.variational_families.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* feat(variational): add DualVariationalGaussian and dual_elbo (t-SVGP)
Implements the dual/site parameterisation of sparse variational GPs from
Adam, Chang, Khan and Solin (2021), "Dual Parameterization of Sparse
Variational Gaussian Processes", NeurIPS 2021 (arXiv:2111.03412).
`DualVariationalGaussian` stores an unnormalised Gaussian site on the
*centred* inducing outputs -- `dual_vector` is the site's first natural
parameter and `dual_matrix` its precision, both `Real` and both defaulting
to zero, so q(u) = p(u) at initialisation. Neither carries a constraining
bijection: PSD-ness of the site precision comes from the convex-combination
structure of the natural-gradient update, and a bijection would destroy that
affine step. Moments, marginals, prior KL and predictions all route through
the working matrix R = Kzz + Kzz L2 Kzz, which dominates Kzz and is therefore
always factorisable even when the site precision is rank deficient; exactly
two Cholesky factorisations are taken per call and nothing is inverted.
The centred convention is the correction to the reference implementation's
mean-function bug, which shifts by `predict_f(Z)` and so is wrong for any
non-zero mean function. `test_dual_natgrad_handles_non_zero_mean_function`
pins the correct behaviour and asserts that the uncentred variant misses.
`marginals` adds the family's jitter to every marginal variance. This is
load-bearing, not cosmetic: `VariationalGaussian.predict` runs `add_jitter`
on its output covariance, so the per-point marginals `elbo` sees carry the
same offset, and without it `dual_elbo` would miss `elbo` at matched moments
by N*eps/(2 sigma^2).
`dual_elbo` is the same functional as `elbo` but evaluated as a function of
the sites and the hyperparameters. Its value matches `elbo` at the implied
moments (measured 5.7e-14 absolute, 2.8e-16 relative at random PSD sites)
while its hyperparameter gradient differs, because q moves with theta through
Kzz while the sites stay frozen. Kzz is deliberately not detached and no
moments are cached on the module; a cached implementation would pass every
value assertion and fail only `test_dual_elbo_hyper_gradients_differ_away_
from_optimum`.
`natural_gradient_step` and `variational_coordinates` gain dual
registrations. Because grad_mu KL = lambda exactly, the KL is never
differentiated and the step is a convex combination towards a closed-form
target built from one `jax.grad` of `expected_log_likelihood` (Bonnet and
Price), with N/B scaling, a trace-safe floor on beta and a symmetrise. That
makes it the Salimbeni step at gamma = rho: matched initialisations agree to
1.4e-15 in (m, S) over six Bernoulli steps at every rate tested. One rho = 1
full-batch conjugate step lands on the Titsias optimum to 4.9e-15.
`fit_natgrads` rejects a numeric step size above one on this family, since
the update is a convex combination; schedules cannot be checked statically
and are documented as the caller's responsibility. Plain `fit` on the family
also works and is documented as gradient descent in the dual coordinates.
The `VariationalParametrisationSuite` benchmark gains a `dual` arm, in both
`benchmarks/objectives.py` and the hardcoded parametrize list in
`tests/test_benchmarks_smoke.py`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* docs(examples): add natural gradients tutorial notebook
Adds examples/natgrads.py, the first of two tutorial notebooks for the
natural-gradient stack. It derives the exponential-family view of q(u),
shows that the Fisher information is the Jacobian d(eta)/d(theta) (checked
numerically to 2.9e-15), reads the update as mirror descent, and then runs
two demos:
* a conjugate 1D regression where one gamma=1 full-batch natural-gradient
step recovers the Titsias optimum to 1.2e-13 while Adam on the same
problem is still 4.6 nats short after 2000 iterations;
* a mini-batched 2D banana Bernoulli benchmark (N=2000, M=50, B=256, 1000
iterations) comparing natural gradients + Adam against Adam alone, per
iteration and per wall-clock second.
Closes with the negative-definite cone result, a gamma sweep reproducing
its boundary, and a demonstration of the step-size backoff.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* fix(examples): avoid markdown-katex crash in the dual notebook
`mkdocs build` aborted with `IndexError: string index out of range` while
rendering `_examples/dual_svgp.md`. The trigger is a markdown-katex parser
bug: `iter_inline_katex` reads `line[end + 1]` without a bounds check, so any
line whose final characters are a backtick code span immediately preceded by
`$` crashes the build. The prose read
... that is $-$`sparsity_gap`
above, and ...
where the closing `$` of `$-$` abuts the code span and the span ends the line.
Reword to "the negated `sparsity_gap` computed above", which removes the
`$`-adjacent code span entirely rather than relying on a particular line wrap.
The meaning is unchanged, and the sentence still says c(theta) is minus the
Titsias trace term. A scan of all generated `docs/_examples/*.md` confirms this
was the only occurrence of the pattern in the docs tree.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* fix(variational): address PR#2 review findings
CHANGELOG: add the missing "### Added" entry for fit_natgrads and
gpjax.natural_gradients. PR#1 shipped both without a changelog entry, so the
Removed entry added here forward-referenced an API the changelog never
announced. Bringing it forward from PR#3 keeps any release cut mid-stack
self-consistent; PR#3 appends the dual entries to the same section.
CHANGELOG: correct the justification prose. "the natural gradient with respect
to theta equals the ordinary gradient with respect to eta, in any
parameterisation" is false as literally written -- for a reparameterisation xi
with J = dtheta/dxi, the natural gradient in xi is J^-1 grad_eta L, not
grad_eta L. The identity is specific to the natural/expectation pair of an
exponential family. Reworded to state that pairing and the point it supports:
either coordinate system is recoverable on the fly, so no dedicated class is
needed.
benchmarks: drop diff-relative wording from the
VariationalParametrisationSuite docstring. "surviving" only means something to
someone reading this commit's diff, and "Both" is a count PR#3 invalidates when
it re-adds the dual arm. The module docstring keeps its explicit
"(standard, whitened)" list, which PR#3 must extend regardless.
tests: split the _psd guard out of test_removed_families_are_gone. The _psd arm
asserted `"_psd" not in __all__`, vacuous for a helper that was never exported,
under a failure message about superseded parameterisations. It is now its own
test with a docstring saying what it actually guards.
CLAUDE.md: "Three optimisers" -> four. fit_natgrads landed in gpjax/fit.py in
PR#1; the sentence has been stale since.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* fix(variational): address PR#3 review findings
Factor R in the Kzz basis. Forming R = Kzz + Kzz Lambda_2 Kzz explicitly
carries a rounding error of order ||Kzz||^2 ||Lambda_2|| eps, which for a
large-variance kernel (or in float32) exceeds lambda_min(R) ~= jitter, so
chol(R) returned NaN and poisoned every later fit_natgrads iterate --
measured at RBF variance 1e4, M=80, default jitter, float64, where the
matched VariationalGaussian run stayed finite. R = Lk (I + Lk^T Lambda_2 Lk)
Lk^T is factorised instead, giving a lower-triangular Lr = Lk chol(I + G)
whose inner matrix has lambda_min >= 1 - O(||G|| eps). Same two Choleskys
per call, plus one M x M product. The "chol(R) never fails" and "no Cholesky
a backoff could rescue" claims are softened to match.
Also: split _gram_and_root off _working_matrices so the dual natgrad step
stops discarding a chol(R); take tr(R^-1 Kzz) as ||Lr^-1 Lk||_F^2 rather
than a full cho_solve; bound-check an optax schedule against rho <= 1 over
the whole num_iters horizon for the dual family, which previously returned
a silent all-NaN history; hoist the _fmt_Kzt_Ktt/_fmt_inducing_inputs hooks
to AbstractVariationalGaussian and keep one typed _symmetrise; convert the
dual family's numpydoc sections to the Google style the file and mkdocs use,
and document marginals' inputs argument.
Doc corrections, all measured: elbo on a DualVariationalGaussian returns the
same value and the same gradients as dual_elbo (bit-identical value, 1.7e-14
on gradients), so the CHANGELOG's gradient claim now names the matched
VariationalGaussian as the comparison; and vmap does not repeat the
unbatched factorisations per datum, so neither elbo nor the benchmark arm
pays 2N Choleskys -- eager counts are 4 (dual), 2 (standard), 1 (whitened),
and the compiled dual_elbo and dual step are 2 potrf each.
Tests: regression for chol(R) at variance 1e4/M=80 and 1e3/M=50 with the
default jitter, an Lr Lr^T = R reconstruction check, the schedule guard, a
half-batch arm on the dual/elbo equivalence plus a direct pin on the N/B
factor, and the triplicated dual fixtures moved to tests/_dual_helpers.py.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* docs(examples): address natgrads notebook review findings
Correct the mini-batch ramp argument (the target is q-dependent outside
conjugacy, so gamma=1 lands on a moving fixed-point target, not the
mini-batch optimum), replace the unmeasured calibration claim after the
banana contours with the metrics the cell actually prints, and attribute
the cone sweep's gamma=2 failure to the over-confident S_0 rather than to
gamma=2 itself.
Smaller corrections: the Fisher solve is O((M + M(M+1)/2)^3) in the vec_s
coordinates, not O((M + M^2)^3); the one-step demo agrees to ~1e-13, not
fourteen decimal places; the sparse/exact predictive deviation is
quantified and located outside the data range; the Adam ELBO-gap
description now matches the shape of the log-log panel; the roadmap says
"exact variational optimum" where the notebook later reserves "exact
posterior" for the non-sparse GP; K=100 is explained against the paper's
dataset-dependent K.
Code: derive the ELBO figure's y-limits from the smoothed histories so
neither curve is clipped, guard the crossing report against never
reaching the target, draw the held-out points in the categorical palette
with per-class markers instead of the contour colourmap, use banana_data
in the split report, and run ruff format (the pinned make_banana block is
byte-identical afterwards).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* fix(natural-gradients): address whole-stack review findings
The dual/Salimbeni E-step divergence on the banana demo was attributed to
floating-point conditioning. It is not: `inv_probit` clips its output into
[1e-3, 1-1e-3], so the computed Bernoulli log-likelihood is not log-concave in
the tails (positive second derivative for f < -2.44). A confidently mislabelled
point then yields beta_i < 0, the dual branch's beta_floor clips it, and the two
branches diverge. With the clip disabled the same six steps agree to 6.2e-13
instead of 5.1e-3.
- Re-attribute the mechanism in the dual notebook and add a diagnostic cell that
measures it, and qualify the "identical iterates" claim wherever it is stated
(natural_gradients.py, fit.py, CHANGELOG, both notebooks): it holds provided
the computed beta stays non-negative.
- Note in the natgrads notebook that the cone discussion assumes a log-concave
*computed* likelihood, which the clipped probit violates in the far tails.
- Reword the Salimbeni registration docstrings: dispatch covers
GraphVariationalGaussian, but the standard elbo path is broken upstream
(MatrixLinearOperator dimensionality error out of gram), on main too.
- Extend _check_natgrad_schedule to reject non-positive rates for every family,
not just the dual upper bound.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
* refactor(natgrads): adapt vestigial-family removal to the v1.0 API
The Sphinx docs arrived with v1.0, after this branch was cut, so the
deletion of NaturalVariationalGaussian and ExpectationVariationalGaussian
now has to reach three doc files the original commits could not know
about: drop the two classes from the variational-families autosummary
page, and repoint the glossary's "natural parameters" entry from the
removed classes to fit_natgrads on the surviving families. fit_natgrads
is added to the fit reference page so that glossary link resolves —
an omission from PR#1, which shipped the function without a reference
entry.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* refactor(natgrads): adapt DualVariationalGaussian and dual_elbo to the v1.0 API
The v1.0 likelihoods are pure conditionals, so the minibatch ELBO scale in
dual_elbo and the dual natural-gradient step now derives from
Dataset.n_total instead of likelihood.num_datapoints, the constructor
annotation follows the AbstractPosterior -> JointModel split, and the
docstring examples plus the dual test fixtures drop num_datapoints. The
minibatch tests stamp n_total onto their batch views to preserve the N/B
factor they assert.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* refactor(natgrads): adapt natural-gradients notebook to the v1.0 API
Likelihoods lose num_datapoints, prior * likelihood is now narrated as the
joint model (variables renamed *_posterior -> *_model), and the exact-GP
comparison uses model.condition(D) followed by a posterior query instead of
predict(x, train_data). Sphinx-migration adaptations: the notebook moves to
docs/examples/, imports utils from the notebook directory, uses {cite:t}
citations and a relative link to the uncollapsed-VI notebook, gains the
nb-download line, and is wired into the docs/index.md toctree. Verified by
executing the notebook end to end; all printed diagnostics match the
narration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* refactor(natgrads): adapt dual notebook to the v1.0 API
Likelihoods lose num_datapoints, and the prior * likelihood factories are
renamed to say what they build (conjugate_model, logit_model, banana_model,
vem_joint_model) with the num_datapoints parameter dropped from the family
helpers. Sphinx-migration adaptations: the notebook lives in docs/examples/,
imports utils from the notebook directory, uses {cite:t} citations, links to
the natgrads notebook relatively, gains the nb-download line, and is wired
into the docs/index.md toctree in place of the removed mkdocs.yml nav; the
natgrads notebook's closing pointer becomes a live relative link. Verified by
executing the notebook end to end; the printed diagnostics match the
narration (Titsias optimum to ~1e-12, rho = gamma to 3e-15, beta < 0 first at
step five, both natural-gradient runs crossing Adam's bound at iteration 109).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* feat(conditioning): add sparse and collapsed conditioning modes
Two new internal Posterior implementations extend the deep conditioning
module to variational families:
- SparsePosterior: the one sparse-predictive derivation, covering both the
whitened and unwhitened parameterisations behind a static `whitened`
flag, with a new diagonal-covariance path the families never had.
- CollapsedPosterior: caches the Titsias factorisation at condition time;
the predictive, the collapsed evidence bound (`elbo_bound`), and the
optimal-q KL (`prior_kl`) are views of it. The trace term uses
`kernel.diagonal(x)` rather than a vmap of `kernel.__call__`, which is
correct for every kernel.
Both classes stabilise K_zz through `stabilised_cholesky` with the model's
single `Prior.jitter` knob and stay out of `gpjax/__init__`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* refactor(variational)!: families condition through the conditioning module
Every Gaussian-output variational family now conditions through
gpjax/conditioning.py:
- The five copied predictive derivations (VariationalGaussian, Whitened,
Graph via its hooks, Dual, Collapsed) are deleted; `predict` is one-line
sugar over the new `condition()`, which returns a Posterior like any
other. The dual family converts its sites to moments in `condition()`,
preserving the implicit hyperparameter dependence under differentiation.
- BREAKING: the family field and constructor kwarg `posterior` is renamed
`model`, and the family-side `jitter` knob is removed — conditioning uses
the model's single `Prior.jitter`.
- `AbstractVariationalFamily` now declares both `predict` AND `prior_kl`
as abstract; the collapsed family gains `prior_kl(train_data)` as a view
of its conditioned posterior.
- `variational_expectation` drops the vmapped per-point q_moments hack for
the conditioned posterior's diagonal path; as a side effect the graph
family's `elbo` now works (it previously raised through the per-point
vmap of `kernel.gram`).
- `collapsed_elbo` becomes `q.condition(data).elbo_bound`; the duplicated
Titsias derivation dies with it.
- HeteroscedasticVariationalFamily takes the rename and the jitter removal
only; its predict internals are untouched.
Reference predict moments, prior_kl, elbo, collapsed_elbo, and dual_elbo
on small fixtures at default jitter match the pre-refactor values to
1e-10 (max observed drift 1.1e-13).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* test(variational): rename sweep, jitter relocation, xfail flip, conformance
- Mechanical `posterior` -> `model` field/kwarg rename across the test suite
and benchmarks; fixtures move their non-default jitters onto the Prior,
the single knob conditioning now reads.
- Flip the strict xfail: collapsed_elbo vs conjugate_mll at z = x now holds
at non-default jitter too, with the tolerance scaled linearly in the knob
(the Titsias bound's intrinsic O(n * jitter / (2 sigma^2)) gap).
- New conformance suite (tests/test_variational_conditioning.py) over the
five Gaussian-output families: condition() returns a conditioning
Posterior, predict is exact sugar over it, prior_kl is a finite scalar,
and the new diagonal path matches the dense marginals to 1e-10.
- The abstract-family test now also pins that a family without prior_kl is
not instantiable.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* docs(examples): posterior -> model rename and jitter relocation
Mechanical sweep of every notebook that builds a variational family:
constructor kwarg `posterior=` becomes `model=`, family field reads become
`.model...`, and the non-default jitters (natgrads' and dual_svgp's 1e-8)
move onto the Prior — the single knob conditioning now reads. Prose that
described the family-side jitter now names `Prior.jitter`.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* docs(records): variational universalisation in CONTEXT, ADR-0001, migration
- CONTEXT.md gains the "variational family" term (a trainable approximate
posterior over inducing values; to sparse GPs what JointModel is to
exact ones; `.condition()` yields a Posterior like any other) and the
condition entry drops its future tense.
- ADR-0001 consequences record the landing: one jitter knob for families
too, the flipped equivalence xfail, the posterior->model rename; the
follow-ups list picks up the skipped lz-sharing and the heteroscedastic
NamedTuple debt.
- migration.md's 1.0.0 guide documents the family changes: field/kwarg
rename, jitter removal, the condition API, and the prior_kl contract.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* style: ruff-clean the safety-net test files
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
* refactor!: complete the conditioning contract — uniform signature, Kalman posterior
Addresses the structural findings from the two-axis review of the v1.0 stack.
- `condition` becomes an abstract method on `AbstractVariationalFamily`, with
`__or__` delegating to it. Nothing previously enforced the contract.
- The three divergent signatures are unified: every conditionable object now
takes `condition(train_data)`, required. Families whose maths does not
consume the data accept it and say so.
- `StateSpaceConjugateModel` gains a real `condition` returning a Kalman-backed
`StateSpacePosterior`. It previously inherited `ConjugateModel.condition`, so
`model | D` silently built the dense O(N^3) posterior the state-space path
exists to avoid, and disagreed with its own `predict`.
- `HeteroscedasticModel` no longer stores `noise_prior` alongside `noise_model`.
The same parameters appeared at two pytree paths (9 leaves, 6 duplicated), so
an optimiser step desynchronised them while inference read only one. The
noise model is now the single owner and `noise_prior` is a property.
- The dense/diagonal predictive assembly, copied four times across the
`Posterior` modes, is extracted to a single `_assemble_predictive`.
- Cholesky sites in the variational families route through
`stabilised_cholesky`, restoring the single-jitter-seam invariant.
- `Dataset.full_size` replaces the `n_total if ... else n` idiom.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(api): Google-convention docstrings and RST across the natgrads surface
`pyproject.toml` sets ruff `convention = "google"` and `docs/conf.py` sets
`napoleon_numpy_docstring = False`, but `natural_gradients.py`, `fit.py` and
`objectives.py` carried NumPy-style sections and Markdown fences, which
napoleon renders as broken RST. Ruff's `D` rules are not selected, so nothing
caught them. Follows the pattern established for `fit_natgrads` in d53c8081.
Also narrows a bare `except Exception: return` that silently swallowed
learning-rate schedule errors, and clears docstring references to classes this
stack deleted.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: close the verification gaps the review found
- `tests/test_conditioning.py`, specified by the implementation plan and never
written: pytree round-trips for all four `Posterior` modes, jit/grad through
`condition`, the multi-output `sample_approx` refusal (previously untested),
and a conformance check that `predict` is sugar rather than a second
implementation.
- `tests/test_state_space/test_conditioning.py`: pins that `model | D` yields a
`StateSpacePosterior` rather than a dense one, that the sugar identity holds
exactly, and that `covariance="dense"` refuses.
- The flipped xfail no longer asserts a magic `atol=2e-1`. The collapsed-ELBO /
MLL identity is exact only as jitter -> 0 (measured gap 1.86e-4 at 1e-6,
1.10e-1 at 1e-3), so a fixed tolerance at high jitter cannot discriminate.
It now pins the O(jitter) scaling, which is what the original two-owner bug
actually violated.
- `CollapsedPosterior.prior_kl` gets an oracle: ~35 lines of new mathematics
whose only assertion was `isfinite`. The oracle derives the KL independently
from the Gaussian definition rather than re-arranging the implementation.
- An elbo value fixture, required by the spec and never added; `test_elbo`
asserted only shape, jit-ability and gradient flow.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(records): state what "universal" means as shipped
ADR-0001 gains a mode -> object -> Posterior table covering exact, latent,
sparse, collapsed and state-space conditioning, and names the heteroscedastic
path as the one exclusion (two latent processes, no single conditioned
process). CONTEXT.md and the migration guide are corrected to the uniform
`condition(D)` signature; both still documented the no-argument form.
`gpjax.models.oilmm` is recorded as sitting outside the contract.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* refactor(oilmm)!: bring OILMM onto the conditioning contract
`OILMMModel` gains `condition(train_data)` and `__or__`, and `OILMMPosterior`
becomes an `eqx.Module` subclassing `gpjax.conditioning.Posterior`. It is the
last public model that stood outside the v1.0 contract.
`OILMMModel` is not a `JointModel` and does not become one — it is not built
from `prior * likelihood`. It does not need to be: contract membership is
defined by exposing `condition`, not by ancestry, exactly as it is for the
variational families.
Two things fell out of the change rather than motivating it:
- **The old posterior was not conditioned.** `condition_on_observations` built
`prior * likelihood` per latent — M *unconditioned* `ConjugateModel`s — and
stored them beside M datasets, deferring the real conditioning to `predict`.
Every prediction therefore re-factorised M Choleskys. Holding
`ExactPosterior`s caches them at `condition` time: ~1.8x faster repeated
prediction at n=300, and `predict` becomes sugar rather than the place the
work happens.
- **The stated reason for the plain class was untrue.** The docstring claimed
it could not be an `eqx.Module` because it "holds Dataset objects which are
not JAX pytree nodes". `Dataset` is a registered pytree, and `ExactPosterior`
has held one as a field since it was written.
`log_marginal_likelihood` moves onto the posterior with the Prop. 9 projection
correction cached alongside it, so `oilmm_mll` is now a one-line view like
every other objective. `condition_on_observations` and
`predict(return_full_cov=)` survive as deprecated aliases.
For this multi-output process `covariance="dense"` means the joint (NP, NP)
covariance across test inputs and outputs; `covariance="diagonal"` asks the
same of each latent process, so dense latent covariances are never formed.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test(oilmm): migrate callers and pin contract conformance
Moves the repo's own callers off the deprecated spellings, and adds conformance
tests mirroring the other conditioning modes: `condition` returns a
`gpjax.conditioning.Posterior`, `|` matches `condition`, `predict` is sugar
rather than a second implementation, evidence agrees with `oilmm_mll`, the
diagonal path matches the dense diagonal, the posterior survives a pytree
round-trip, conditioning is jittable, and both deprecated spellings warn.
Three existing tests iterated `zip(latent_posteriors, latent_datasets)` to
re-condition each latent by hand; the latent processes now arrive conditioned,
so they read straight through.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(records): OILMM is on the contract
ADR-0001 gains a multi-output row and replaces the paragraph that recorded
OILMM as an exclusion, including the correction that its stated pytree
obstacle never existed. The migration guide gains a runnable OILMM section and
notes the multi-output meaning of `covariance="dense"`.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* docs(oilmm): migrate the latent-space cell onto the conditioning contract
The refactor removed `OILMMPosterior.latent_datasets` and changed
`latent_posteriors` to hold conditioned `ExactPosterior`s, each carrying its
own projected training set. The "Latent space after optimisation" cell was the
one caller left behind, so `sphinx-build -W` turned its `AttributeError` into a
docs build failure — and, because the cell aborted, nothing downstream of it
was rendered or checked either.
Read the projected data through the latent that owns it and drop the
`train_data=` re-pass, which the conditioned process now ignores.
The migration guide covered the `condition_on_observations` and
`return_full_cov` renames but not the removed attribute, i.e. the only part of
the break that actually had a caller. Documented.
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): scan every PR, not only those targeting main
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.
The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.
Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.
`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.
(cherry picked from commit a8d66032fdabc70411ec47b0bc53553bf76d64ce)
* ci(security): exclude TruffleHog's Lob detector, which flags pytest names
Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1
stack for the first time, and it immediately failed three PRs (#714, #716,
#748) on four "verified" secrets. All four are pytest function names:
test_fit_natgrads_accepts_optax_schedule
test_dual_natgrad_matches_salimbeni_step
test_dual_working_matrices_reconstruct_r
test_prior_kl_matches_textbook_reference
TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and
each of those names is exactly 40 characters long. Worse, the Lob verifier
reports them as *verified*, so `--only-verified` does not filter them: a fresh
repository containing nothing but one of those names reproduces
`Lob verified=true` under trufflehog 3.96.0, the same version the action pins.
`Secrets Detection` is a required context on `main`, so this is not cosmetic —
it means an arbitrary subset of PRs, selected by test-name length, can never
merge. GPJax has no Lob (direct-mail API) integration, so excluding the
detector costs no coverage. Verified that the flag is parsed rather than
silently ignored: an unrecognised detector name makes trufflehog exit with
"unrecognized detector type", and the four findings disappear with the flag
and return without it.
(cherry picked from commit 83ad296ae767ccd826c161033636ec222934246d)
* ci(security): exclude TruffleHog's Lob detector, which flags pytest names
Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1
stack for the first time, and it immediately failed three PRs (#714, #716,
#748) on four "verified" secrets. All four are pytest function names:
test_fit_natgrads_accepts_optax_schedule
test_dual_natgrad_matches_salimbeni_step
test_dual_working_matrices_reconstruct_r
test_prior_kl_matches_textbook_reference
TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and
each of those names is exactly 40 characters long. Worse, the Lob verifier
reports them as *verified*, so `--only-verified` does not filter them: a fresh
repository containing nothing but one of those names reproduces
`Lob verified=true` under trufflehog 3.96.0, the same version the action pins.
`Secrets Detection` is a required context on `main`, so this is not cosmetic —
it means an arbitrary subset of PRs, selected by test-name length, can never
merge. GPJax has no Lob (direct-mail API) integration, so excluding the
detector costs no coverage. Verified that the flag is parsed rather than
silently ignored: an unrecognised detector name makes trufflehog exit with
"unrecognized detector type", and the four findings disappear with the flag
and return without it.
(cherry picked from commit 83ad296ae767ccd826c161033636ec222934246d)
* ci(security): exclude TruffleHog's Lob detector, which flags pytest names
Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1
stack for the first time, and it immediately failed three PRs (#714, #716,
#748) on four "verified" secrets. All four are pytest function names:
test_fit_natgrads_accepts_optax_schedule
test_dual_natgrad_matches_salimbeni_step
test_dual_working_matrices_reconstruct_r
test_prior_kl_matches_textbook_reference
TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and
each of those names is exactly 40 characters long. Worse, the Lob verifier
reports them as *verified*, so `--only-verified` does not filter them: a fresh
repository containing nothing but one of those names reproduces
`Lob verified=true` under trufflehog 3.96.0, the same version the action pins.
`Secrets Detection` is a required context on `main`, so this is not cosmetic —
it means an arbitrary subset of PRs, selected by test-name length, can never
merge. GPJax has no Lob (direct-mail API) integration, so excluding the
detector costs no coverage. Verified that the flag is parsed rather than
silently ignored: an unrecognised detector name makes trufflehog exit with
"unrecognized detector type", and the four findings disappear with the flag
and return without it.
(cherry picked from commit 83ad296ae767ccd826c161033636ec222934246d)
* ci(security): exclude TruffleHog's Lob detector, which flags pytest names
Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1
stack for the first time, and it immediately failed three PRs (#714, #716,
#748) on four "verified" secrets. All four are pytest function names:
test_fit_natgrads_accepts_optax_schedule
test_dual_natgrad_matches_salimbeni_step
test_dual_working_matrices_reconstruct_r
test_prior_kl_matches_textbook_reference
TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and
each of those names is exactly 40 characters long. Worse, the Lob verifier
reports them as *verified*, so `--only-verified` does not filter them: a fresh
repository containing nothing but one of those names reproduces
`Lob verified=true` under trufflehog 3.96.0, the same version the action pins.
`Secrets Detection` is a required context on `main`, so this is not cosmetic —
it means an arbitrary subset of PRs, selected by test-name length, can never
merge. GPJax has no Lob (direct-mail API) integration, so excluding the
detector costs no coverage. Verified that the flag is parsed rather than
silently ignored: an unrecognised detector name makes trufflehog exit with
"unrecognized detector type", and the four findings disappear with the flag
and return without it.
(cherry picked from commit 83ad296ae767ccd826c161033636ec222934246d)
* ci(security): exclude TruffleHog's Lob detector, which flags pytest names
Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1
stack for the first time, and it immediately failed three PRs (#714, #716,
#748) on four "verified" secrets. All four are pytest function names:
test_fit_natgrads_accepts_optax_schedule
test_dual_natgrad_matches_salimbeni_step
test_dual_working_matrices_reconstruct_r
test_prior_kl_matches_textbook_reference
TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and
each of those names is exactly 40 characters long. Worse, the Lob verifier
reports them as *verified*, so `--only-verified` does not filter them: a fresh
repository containing nothing but one of those names reproduces
`Lob verified=true` under trufflehog 3.96.0, the same version the action pins.
`Secrets Detection` is a required context on `main`, so this is not cosmetic —
it means an arbitrary subset of PRs, selected by test-name length, can never
merge. GPJax has no Lob (direct-mail API) integration, so excluding the
detector costs no coverage. Verified that the flag is parsed rather than
silently ignored: an unrecognised detector name makes trufflehog exit with
"unrecognized detector type", and the four findings disappear with the flag
and return without it.
(cherry picked from commit 83ad296ae767ccd826c161033636ec222934246d)
* ci(security): exclude TruffleHog's Lob detector, which flags pytest names
Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1
stack for the first time, and it immediately failed three PRs (#714, #716,
#748) on four "verified" secrets. All four are pytest function names:
test_fit_natgrads_accepts_optax_schedule
test_dual_natgrad_matches_salimbeni_step
test_dual_working_matrices_reconstruct_r
test_prior_kl_matches_textbook_reference
TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and
each of those names is exactly 40 characters long. Worse, the Lob verifier
reports them as *verified*, so `--only-verified` does not filter them: a fresh
repository containing nothing but one of those names reproduces
`Lob verified=true` under trufflehog 3.96.0, the same version the action pins.
`Secrets Detection` is a required context on `main`, so this is not cosmetic —
it means an arbitrary subset of PRs, selected by test-name length, can never
merge. GPJax has no Lob (direct-mail API) integration, so excluding the
detector costs no coverage. Verified that the flag is parsed rather than
silently ignored: an unrecognised detector name makes trufflehog exit with
"unrecognized detector type", and the four findings disappear with the flag
and return without it.
(cherry picked from commit 83ad296ae767ccd826…
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Checklist
uv run poe formatbefore committing.Description
Replaces simulated data in four example notebooks with real climate data from the Open-Meteo API, and adds two new applied notebooks. All data is pulled once into checked-in CSV snapshots so poe docs-build never touches the network.
Existing notebooks wherein synthetic data replaced with real data
spatial_linear_gp- 152 Swiss weather stations with linear mean in elevation + spatial GP residualheteroscedastic_inference- Barcelona hourly solar irradiance vs. hour-of-daypoisson- Madrid annual hot-day counts (Tmax ≥ 30 °C), 1960–2023multioutput- North-Atlantic wave components (ICM, rank 2)New notebooks
2.40 °C, 98.4% coverage) where a plain RBF reverts to the mean. Motivates real-world kernel design.
between-model ensemble spread (12.4–13.8 °C) - a lesson in epistemic vs. structural/scenario uncertainty.
Supporting changes
examples/data/ so gen_examples.py copies but never executes it.