Repository navigation
[DO NOT MERGE] natgrads stack aggregate preview (1/5–5/5) - #731
Closed
thomaspinder wants to merge 47 commits into
Closed
thomaspinder wants to merge 47 commits into
thomaspinder wants to merge 47 commits into
Conversation
`Zero()` was trainable and drifted towards the data mean during `fit` (0.0 -> 5.09 on a dataset with mean 5), contradicting its own docstring and silently changing the posterior mean of every model using the default mean function. Regression of #330, fixed once in #500. The cause is a changed trainability contract rather than a lost line. Under nnx, `fit` optimised only `Parameter` instances, so `Zero`'s bare array was inert by construction -- which is why #530 could drop the `Static` wrapper and stay correct. Under Equinox, `fit` partitions on `eqx.is_array`, making every array leaf trainable, and `Zero` was still relying on the old meaning of a bare array. Wrap the constant in `paramax.non_trainable` so the invariant holds by construction rather than by the ambient filter semantics. The guard for this was weakened rather than removed: #614 replaced `test_zero_mean_remains_zero` with a test of the initial value only, and `test_zero_mean_function_uses_raw_value` asserted the defect outright. Restore the end-to-end fit assertion and invert the unit test. Fixes #712 Claude-Session: https://claude.ai/code/session_01Bj9k5fnAZ8JzD4Rg3HMDMj Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
thomaspinder
temporarily deployed
to
docs-preview
July 27, 2026 04:26 — with
GitHub Actions
Inactive
thomaspinder
temporarily deployed
to
docs-preview
July 27, 2026 05:29 — with
GitHub Actions
Inactive
Bumps [absl-py](https://github.com/abseil/abseil-py) from 2.3.1 to 2.5.0. - [Release notes](https://github.com/abseil/abseil-py/releases) - [Changelog](https://github.com/abseil/abseil-py/blob/main/CHANGELOG.md) - [Commits](abseil/abseil-py@v2.3.1...v2.5.0) --- updated-dependencies: - dependency-name: absl-py dependency-version: 2.5.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [chex](https://github.com/google-deepmind/chex) from 0.1.91 to 0.1.92. - [Release notes](https://github.com/google-deepmind/chex/releases) - [Commits](google-deepmind/chex@v0.1.91...v0.1.92) --- updated-dependencies: - dependency-name: chex dependency-version: 0.1.92 dependency-type: direct:production update-type: version-update:semver-patch ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
…#720) Bumps the ml-tools group with 2 updates in the / directory: [scipy](https://github.com/scipy/scipy) and [numpy](https://github.com/numpy/numpy). Updates `scipy` from 1.17.0 to 1.17.1 - [Release notes](https://github.com/scipy/scipy/releases) - [Commits](scipy/scipy@v1.17.0...v1.17.1) Updates `numpy` from 2.4.1 to 2.4.6 - [Release notes](https://github.com/numpy/numpy/releases) - [Changelog](https://github.com/numpy/numpy/blob/main/doc/RELEASE_WALKTHROUGH.rst) - [Commits](numpy/numpy@v2.4.1...v2.4.6) --- updated-dependencies: - dependency-name: numpy dependency-version: 2.4.6 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: ml-tools - dependency-name: scipy dependency-version: 1.17.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: ml-tools ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps [tqdm](https://github.com/tqdm/tqdm) from 4.67.1 to 4.69.1. - [Release notes](https://github.com/tqdm/tqdm/releases) - [Commits](tqdm/tqdm@v4.67.1...v4.69.1) --- updated-dependencies: - dependency-name: tqdm dependency-version: 4.69.1 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Thomas Pinder <tompinder@live.co.uk>
Dependabot opened one PR per dependency for anything outside the four named pattern groups. Replace those with catch-all `*` groups so each ecosystem produces a single PR per run, plus a second group with `applies-to: security-updates` since grouping defaults to version updates only. Groups cannot span `updates` entries, so the floor is two PRs: Python and GitHub Actions. Also switch the Python ecosystem from `pip` to `uv`. All dev tooling is declared in `[tool.uv].dev-dependencies`, which `pip` does not read, so ruff, pytest, asv and friends were never being updated. The `uv` ecosystem also keeps `uv.lock` in step with `pyproject.toml`. The removed `testing` and `dev-tools` groups were matching nothing, and `dev-tools` still listed black, isort and mypy, which ruff replaced. Finally, fix the Actions label. Custom labels replace the defaults and missing ones are ignored, so `github-actions` + `automated` meant those PRs landed with no labels at all. The repo label is `github_actions`.
Bumps the python-dependencies group with 23 updates: | Package | From | To | | --- | --- | --- | | [jax](https://github.com/jax-ml/jax) | `0.9.0` | `0.10.2` | | [optax](https://github.com/google-deepmind/optax) | `0.2.6` | `0.2.8` | | [numpyro](https://github.com/pyro-ppl/numpyro) | `0.19.0` | `0.21.0` | | [jaxtyping](https://github.com/patrick-kidger/jaxtyping) | `0.3.6` | `0.3.11` | | [equinox](https://github.com/patrick-kidger/equinox) | `0.13.5` | `0.13.8` | | [lineax](https://github.com/google/lineax) | `0.1.0` | `0.1.1` | | [scipy](https://github.com/scipy/scipy) | `1.17.0` | `1.17.1` | | [numpy](https://github.com/numpy/numpy) | `2.4.1` | `2.4.6` | | [rich](https://github.com/Textualize/rich) | `14.3.1` | `15.0.0` | | [mkdocs-material](https://github.com/squidfunk/mkdocs-material) | `9.7.1` | `9.7.7` | | [mkdocs-jupyter](https://github.com/danielfrg/mkdocs-jupyter) | `0.25.1` | `0.26.3` | | [mkdocs-gen-files](https://github.com/oprypin/mkdocs-gen-files) | `0.6.0` | `0.6.1` | | [mkdocs-literate-nav](https://github.com/oprypin/mkdocs-literate-nav) | `0.6.2` | `0.6.3` | | [matplotlib](https://github.com/matplotlib/matplotlib) | `3.10.8` | `3.11.1` | | [ipython](https://github.com/ipython/ipython) | `9.9.0` | `9.15.0` | | [ipykernel](https://github.com/ipython/ipykernel) | `6.31.0` | `7.3.0` | | [blackjax](https://github.com/blackjax-devs/blackjax) | `1.3` | `1.6.2` | | [pymdown-extensions](https://github.com/facelessuser/pymdown-extensions) | `10.20.1` | `11.0.1` | | [nbconvert](https://github.com/jupyter/nbconvert) | `7.16.6` | `7.17.1` | | [scikit-learn](https://github.com/scikit-learn/scikit-learn) | `1.8.0` | `1.9.0` | | [markdown-it-py](https://github.com/executablebooks/markdown-it-py) | `4.0.0` | `4.2.0` | | [pygments](https://github.com/pygments/pygments) | `2.19.2` | `2.20.0` | | [typing-extensions](https://github.com/python/typing_extensions) | `4.15.0` | `4.16.0` | Updates `jax` from 0.9.0 to 0.10.2 - [Release notes](https://github.com/jax-ml/jax/releases) - [Changelog](https://github.com/jax-ml/jax/blob/main/CHANGELOG.md) - [Commits](jax-ml/jax@jax-v0.9.0...jax-v0.10.2) Updates `optax` from 0.2.6 to 0.2.8 - [Release notes](https://github.com/google-deepmind/optax/releases) - [Commits](google-deepmind/optax@v0.2.6...v0.2.8) Updates `numpyro` from 0.19.0 to 0.21.0 - [Release notes](https://github.com/pyro-ppl/numpyro/releases) - [Commits](pyro-ppl/numpyro@0.19.0...0.21.0) Updates `jaxtyping` from 0.3.6 to 0.3.11 - [Release notes](https://github.com/patrick-kidger/jaxtyping/releases) - [Commits](patrick-kidger/jaxtyping@v0.3.6...v0.3.11) Updates `equinox` from 0.13.5 to 0.13.8 - [Release notes](https://github.com/patrick-kidger/equinox/releases) - [Commits](patrick-kidger/equinox@v0.13.5...v0.13.8) Updates `lineax` from 0.1.0 to 0.1.1 - [Release notes](https://github.com/google/lineax/releases) - [Commits](patrick-kidger/lineax@v0.1.0...v0.1.1) Updates `scipy` from 1.17.0 to 1.17.1 - [Release notes](https://github.com/scipy/scipy/releases) - [Commits](scipy/scipy@v1.17.0...v1.17.1) Updates `numpy` from 2.4.1 to 2.4.6 - [Release notes](https://github.com/numpy/numpy/releases) - [Changelog](https://github.com/numpy/numpy/blob/main/doc/RELEASE_WALKTHROUGH.rst) - [Commits](numpy/numpy@v2.4.1...v2.4.6) Updates `rich` from 14.3.1 to 15.0.0 - [Release notes](https://github.com/Textualize/rich/releases) - [Changelog](https://github.com/Textualize/rich/blob/main/CHANGELOG.md) - [Commits](Textualize/rich@v14.3.1...v15.0.0) Updates `mkdocs-material` from 9.7.1 to 9.7.7 - [Release notes](https://github.com/squidfunk/mkdocs-material/releases) - [Changelog](https://github.com/squidfunk/mkdocs-material/blob/master/CHANGELOG) - [Commits](squidfunk/mkdocs-material@9.7.1...9.7.7) Updates `mkdocs-jupyter` from 0.25.1 to 0.26.3 - [Changelog](https://github.com/danielfrg/mkdocs-jupyter/blob/main/CHANGELOG.md) - [Commits](https://github.com/danielfrg/mkdocs-jupyter/commits) Updates `mkdocs-gen-files` from 0.6.0 to 0.6.1 - [Release notes](https://github.com/oprypin/mkdocs-gen-files/releases) - [Commits](oprypin/mkdocs-gen-files@v0.6.0...v0.6.1) Updates `mkdocs-literate-nav` from 0.6.2 to 0.6.3 - [Release notes](https://github.com/oprypin/mkdocs-literate-nav/releases) - [Commits](oprypin/mkdocs-literate-nav@v0.6.2...v0.6.3) Updates `matplotlib` from 3.10.8 to 3.11.1 - [Release notes](https://github.com/matplotlib/matplotlib/releases) - [Commits](matplotlib/matplotlib@v3.10.8...v3.11.1) Updates `ipython` from 9.9.0 to 9.15.0 - [Release notes](https://github.com/ipython/ipython/releases) - [Commits](ipython/ipython@9.9.0...9.15.0) Updates `ipykernel` from 6.31.0 to 7.3.0 - [Release notes](https://github.com/ipython/ipykernel/releases) - [Changelog](https://github.com/ipython/ipykernel/blob/main/CHANGELOG.md) - [Commits](ipython/ipykernel@v6.31.0...v7.3.0) Updates `blackjax` from 1.3 to 1.6.2 - [Release notes](https://github.com/blackjax-devs/blackjax/releases) - [Commits](blackjax-devs/blackjax@1.3...1.6.2) Updates `pymdown-extensions` from 10.20.1 to 11.0.1 - [Release notes](https://github.com/facelessuser/pymdown-extensions/releases) - [Commits](https://github.com/facelessuser/pymdown-extensions/commits/11.0.1) Updates `nbconvert` from 7.16.6 to 7.17.1 - [Release notes](https://github.com/jupyter/nbconvert/releases) - [Changelog](https://github.com/jupyter/nbconvert/blob/main/CHANGELOG.md) - [Commits](jupyter/nbconvert@v7.16.6...v7.17.1) Updates `scikit-learn` from 1.8.0 to 1.9.0 - [Release notes](https://github.com/scikit-learn/scikit-learn/releases) - [Commits](scikit-learn/scikit-learn@1.8.0...1.9.0) Updates `markdown-it-py` from 4.0.0 to 4.2.0 - [Release notes](https://github.com/executablebooks/markdown-it-py/releases) - [Changelog](https://github.com/executablebooks/markdown-it-py/blob/master/CHANGELOG.md) - [Commits](executablebooks/markdown-it-py@v4.0.0...v4.2.0) Updates `pygments` from 2.19.2 to 2.20.0 - [Release notes](https://github.com/pygments/pygments/releases) - [Changelog](https://github.com/pygments/pygments/blob/master/CHANGES) - [Commits](pygments/pygments@2.19.2...2.20.0) Updates `typing-extensions` from 4.15.0 to 4.16.0 - [Release notes](https://github.com/python/typing_extensions/releases) - [Changelog](https://github.com/python/typing_extensions/blob/main/CHANGELOG.md) - [Commits](python/typing_extensions@4.15.0...4.16.0) --- updated-dependencies: - dependency-name: jax dependency-version: 0.10.2 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: optax dependency-version: 0.2.8 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: numpyro dependency-version: 0.21.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: jaxtyping dependency-version: 0.3.11 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: equinox dependency-version: 0.13.8 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: lineax dependency-version: 0.1.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: scipy dependency-version: 1.17.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: numpy dependency-version: 2.4.6 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: rich dependency-version: 15.0.0 dependency-type: direct:production update-type: version-update:semver-major dependency-group: python-dependencies - dependency-name: mkdocs-material dependency-version: 9.7.7 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: mkdocs-jupyter dependency-version: 0.26.3 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: mkdocs-gen-files dependency-version: 0.6.1 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: mkdocs-literate-nav dependency-version: 0.6.3 dependency-type: direct:production update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: matplotlib dependency-version: 3.11.1 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: ipython dependency-version: 9.15.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: ipykernel dependency-version: 7.3.0 dependency-type: direct:production update-type: version-update:semver-major dependency-group: python-dependencies - dependency-name: blackjax dependency-version: 1.6.2 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: pymdown-extensions dependency-version: 11.0.1 dependency-type: direct:production update-type: version-update:semver-major dependency-group: python-dependencies - dependency-name: nbconvert dependency-version: 7.17.1 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: scikit-learn dependency-version: 1.9.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: markdown-it-py dependency-version: 4.2.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: pygments dependency-version: 2.20.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: typing-extensions dependency-version: 4.16.0 dependency-type: direct:production update-type: version-update:semver-minor dependency-group: python-dependencies ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com>
Bumps the actions group with 2 updates: [actions/labeler](https://github.com/actions/labeler) and [actions/setup-python](https://github.com/actions/setup-python). Updates `actions/labeler` from 6 to 7 - [Release notes](https://github.com/actions/labeler/releases) - [Commits](actions/labeler@v6...v7) Updates `actions/setup-python` from 6 to 7 - [Release notes](https://github.com/actions/setup-python/releases) - [Commits](actions/setup-python@v6...v7) --- updated-dependencies: - dependency-name: actions/labeler dependency-version: '7' dependency-type: direct:production update-type: version-update:semver-major dependency-group: actions - dependency-name: actions/setup-python dependency-version: '7' dependency-type: direct:production update-type: version-update:semver-major dependency-group: actions ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Thomas Pinder <tompinder@live.co.uk>
) Dependabot's uv parser never read `[tool.uv].dev-dependencies`, so ruff, pytest, asv, hypothesis and poethepoet were never updated. Move the 19 entries to `[dependency-groups].dev` (PEP 735), which the parser does read and which uv installs by default. `uv lock` reproduces uv.lock unchanged; `uv sync --frozen` in CI is unaffected. Also fix the documented install command: `uv sync --extra dev` appears in five files but no `dev` extra exists, so it fails today. Correct form is plain `uv sync`.
`test_init` in tests/test_kernels/test_nonstationary.py is a fixture
consumed by `test_gram` and `test_cross_covariance`, but it also carried
two `@pytest.mark.parametrize` decorators. Marks on a fixture have never
done anything; pytest silently ignored them. Both consumers already
parametrize `kernel`, `params` and `variance` themselves, so the marks
were pure dead weight.
pytest 9.1 promotes this from a silent no-op to a hard collection error:
Failed: Marks cannot be applied to fixtures.
That aborts collection for the whole file, which is why #736 (which
bumps pytest 9.0.2 -> 9.1.1) fails every matrix job.
Remove the two marks. Verified against pytest 9.1.1: the file collects
and all 297 cases pass, and the full suite is green, so nothing else in
the tree trips the new pytest.
`test_init` in tests/test_kernels/test_nonstationary.py is a fixture
consumed by `test_gram` and `test_cross_covariance`, but it also carried
two `@pytest.mark.parametrize` decorators. Marks on a fixture have never
done anything; pytest silently ignored them. Both consumers already
parametrize `kernel`, `params` and `variance` themselves, so the marks
were pure dead weight.
pytest 9.1 promotes this from a silent no-op to a hard collection error:
Failed: Marks cannot be applied to fixtures.
That aborts collection for the whole file, which is why #736 (which
bumps pytest 9.0.2 -> 9.1.1) fails every matrix job.
Remove the two marks. Verified against pytest 9.1.1: the file collects
and all 297 cases pass, and the full suite is green, so nothing else in
the tree trips the new pytest.
…#736) Bumps the python-dependencies group with 7 updates: | Package | From | To | | --- | --- | --- | | [asv](https://github.com/airspeed-velocity/asv) | `0.6.5` | `0.6.6` | | [codespell](https://github.com/codespell-project/codespell) | `2.4.1` | `2.4.3` | | [pytest](https://github.com/pytest-dev/pytest) | `9.0.2` | `9.1.1` | | [pytest-cov](https://github.com/pytest-dev/pytest-cov) | `7.0.0` | `7.1.0` | | [coverage](https://github.com/coveragepy/coveragepy) | `7.13.2` | `7.15.2` | | [xdoctest](https://github.com/Erotemic/xdoctest) | `1.3.0` | `1.3.2` | | [poethepoet](https://github.com/nat-n/poethepoet) | `0.40.0` | `0.48.0` | Updates `asv` from 0.6.5 to 0.6.6 - [Release notes](https://github.com/airspeed-velocity/asv/releases) - [Changelog](https://github.com/airspeed-velocity/asv/blob/main/CHANGES.rst) - [Commits](airspeed-velocity/asv@v0.6.5...v0.6.6) Updates `codespell` from 2.4.1 to 2.4.3 - [Release notes](https://github.com/codespell-project/codespell/releases) - [Commits](codespell-project/codespell@v2.4.1...v2.4.3) Updates `pytest` from 9.0.2 to 9.1.1 - [Release notes](https://github.com/pytest-dev/pytest/releases) - [Changelog](https://github.com/pytest-dev/pytest/blob/main/CHANGELOG.rst) - [Commits](pytest-dev/pytest@9.0.2...9.1.1) Updates `pytest-cov` from 7.0.0 to 7.1.0 - [Changelog](https://github.com/pytest-dev/pytest-cov/blob/master/CHANGELOG.rst) - [Commits](pytest-dev/pytest-cov@v7.0.0...v7.1.0) Updates `coverage` from 7.13.2 to 7.15.2 - [Release notes](https://github.com/coveragepy/coveragepy/releases) - [Changelog](https://github.com/coveragepy/coveragepy/blob/main/CHANGES.rst) - [Commits](coveragepy/coveragepy@7.13.2...7.15.2) Updates `xdoctest` from 1.3.0 to 1.3.2 - [Release notes](https://github.com/Erotemic/xdoctest/releases) - [Changelog](https://github.com/Erotemic/xdoctest/blob/main/CHANGELOG.md) - [Commits](Erotemic/xdoctest@v1.3.0...v1.3.2) Updates `poethepoet` from 0.40.0 to 0.48.0 - [Release notes](https://github.com/nat-n/poethepoet/releases) - [Commits](nat-n/poethepoet@v0.40.0...v0.48.0) --- updated-dependencies: - dependency-name: asv dependency-version: 0.6.6 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: codespell dependency-version: 2.4.3 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: pytest dependency-version: 9.1.1 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: pytest-cov dependency-version: 7.1.0 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: coverage dependency-version: 7.15.2 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: python-dependencies - dependency-name: xdoctest dependency-version: 1.3.2 dependency-type: direct:development update-type: version-update:semver-patch dependency-group: python-dependencies - dependency-name: poethepoet dependency-version: 0.48.0 dependency-type: direct:development update-type: version-update:semver-minor dependency-group: python-dependencies ... Signed-off-by: dependabot[bot] <support@github.com> Co-authored-by: dependabot[bot] <49699333+dependabot[bot]@users.noreply.github.com> Co-authored-by: Thomas Pinder <tompinder@live.co.uk>
* docs: migrate from MkDocs to Sphinx, mirroring impulso's build Replaces the MkDocs Material stack with the Sphinx/MyST-NB setup used by impulso, so both repositories share one documentation build idiom. The old build had no notebook caching, so editing a line of prose re-executed all 22 example notebooks. Four of them fetched data from NOAA, GitHub and UCI while executing, making builds non-hermetic and non-reproducible. Nothing failed the build: link validation was set to warn, so broken references accumulated silently. Publishing spanned two mechanisms plus a Netlify configuration that existed only in a vendor web UI, and a release tag overwrote main's docs with a snapshot. Build: - Sphinx + MyST-NB reads the jupytext py:percent notebooks directly; the jupytext -> nbconvert conversion script is deleted. - Notebook execution is cached, with separate cache directories for smoke and full renders because jupyter-cache keys on cell source alone. - GPJAX_DOCS_CI=1 shrinks the sampling budgets of the five expensive notebooks; GPJAX_DOCS_RESILIENT=1 lets the deploy survive one failed notebook. - PR builds are strict (-W --keep-going); the scheduled Sunday job renders at full fidelity and opens an issue on failure. Reference: - Pages are driven by each module's __all__. objectives, parameters and citation had none and gained one, so their public symbols now appear. - The old filesystem walk published two private state_space modules; it cannot happen now by construction. - sphinx-codeautolink links API names in executed notebook cells to the reference and injects backreferences onto each documented object. Docstrings: - 15 fenced maths blocks converted to .. math:: and 15 fences around doctests removed, since reStructuredText has no fenced-block concept. - 28 docstrings standardised on Google style; ruff and CLAUDE.md corrected to match, resolving a three-way contradiction with the old mkdocs config. - Seven pre-existing reStructuredText defects fixed, which is what gets the strict build to zero warnings. - $-maths is unchanged in source: sphinx-math-dollar renders it, so no translation layer and no rewrite of ~195 sites. Hermetic builds: - The four remote datasets are vendored into docs/examples/data (88 KB) behind a manual pull script, matching the existing Open-Meteo pattern. Note this freezes the Mauna Loa series, whose output previously drifted every build. Publishing: - One mechanism: upload-pages-artifact + deploy-pages, gated on main. The tag trigger is removed. PR builds upload a plain artifact and have contents:read only, so they cannot reach Pages - the failure mode impulso hit, where Pages preview deployments silently overwrite production. - miniconda (installed only for pandoc) and Node 16 (only for npm katex) are both gone; Sphinx needs neither. - 79 redirects preserve every /api/* and /_examples/* URL. - polyfill.io is removed with mkdocs.yml; that domain was sold in 2024 and served malware. The only custom build code is impulso's ~10-line render-mode stamp. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * ci: restore the "Build docs" job name required by branch protection main's branch protection requires a status check literally named "Build docs", matched on the job's display name. The migration renamed the job to "Build docs (smoke render)", so the required check had nothing reporting it and every PR sat in "Expected — waiting for status to be reported". The step inside the job still says "smoke render", which is where that detail belongs. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * ci: publish a Netlify deploy preview for each docs PR Mirroring impulso left PRs with no docs preview at all — it uploads a plain artifact and nothing else. That was the deliberate trade for killing the old deploy-docs-preview job, which used actions/deploy-pages with preview:true and silently republished the production site with a PR's smoke render. This restores a preview without that failure mode. `netlify deploy --alias` creates a branch deploy at pr-<n>--<site>.netlify.app on the same site that serves production; without `--prod` it cannot touch the production deploy. The URL is posted as a sticky PR comment. Both secrets are read at job level so the step's own `if` can test them — a step cannot see env vars it declares itself. The step is skipped on forks and whenever NETLIFY_AUTH_TOKEN is unset, so a missing preview never fails the docs gate. Requires NETLIFY_AUTH_TOKEN and NETLIFY_SITE_ID repository secrets. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * docs: stop nesting equation environments inside display maths MathJax rendered "Erroneous nesting of equation structures" in place of the equation on the homepage, and the same failure was latent on five reference pages. The build was green throughout: Sphinx writes LaTeX into the HTML verbatim and MathJax parses it in the browser, so -W cannot see this class of defect at all. Prose and notebooks (25 blocks across 5 files): the `$$` wrapper around an amsmath environment is removed and the environment starred. The mechanism is not the `\[...\]` wrapper, which MyST always emits and which renders fine — it is Sphinx's default displaymath writer, which wraps `$$` content containing `\\` in `\begin{split}` or `\begin{align}\begin{aligned}`, producing align inside align. Dropping `$$` routes the block to the amsmath renderer instead. Starred because the live site renders these unnumbered. Docstrings (7 blocks in gps.py, objectives.py, variational_families.py): these sit inside a `.. math::` directive, which is already display maths, so `align` nested the same way. They now use `aligned`, the inner variant intended to nest inside display maths. Also fixes `$$` mis-pairing where an opener was glued to the preceding prose line: delimiters shifted and paragraphs were rendered AS mathematics. The old build contained `\[where\]` and `\[Further, the log of the marginal likelihood...\]`. That is also why intro_to_gps carried 7 spurious equation numbers. One cross-reference is converted to the MyST `{eq}` role; `\eqref` was rendering as literal text. Verified by loading the built pages in a browser and querying for mjx-merror nodes — the only way to observe this, since the error is generated client-side. Zero errors on index, sharp_bits, and the three affected reference pages. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * ci: fail the docs gate on MathJax rendering errors Sphinx writes LaTeX into the HTML verbatim and MathJax parses it in the browser, so `sphinx-build -W` reports zero warnings while a page shows "Erroneous nesting of equation structures" to every reader. That shipped to a preview earlier today. Nothing in the existing gate can see it. docs/scripts/check_mathjax.py serves the built site, loads every page containing maths in Chromium, waits on MathJax.startup.promise and reports three distinct failure modes — only the first of which announces itself: 1. mjx-merror nodes; 2. blocks MathJax left entirely untypeset, which is what an unbalanced brace produces — no container, no error, just raw \[...\] on the page; 3. undefined macros, which the noundefined package renders as red literal text rather than an error, so a missing \bm ships silently. Detecting only (1) was the first implementation, and review rejected it: a one-character brace typo in a docstring passed with Sphinx green AND the checker green. Detecting only `class="math` also skipped MyST amsmath pages entirely, whose class is `amsmath math ...` — acute, since the migration just converted every display block to that form. The extract-and-retypeset shortcut is deliberately not used. It is much faster and gives false negatives: it reports OK for markup that visibly fails in situ, because the failure depends on the page's own MathJax config and on how Sphinx wrapped the block. The script says so, at length, to stop it being "optimised" back. Runs on the strict paths only — the PR gate and the scheduled full render. Deliberately absent from build_docs.yml, which is resilient by design so a deploy is never blocked. It found two real defects on its first clean run: intro_to_gps.py had two `$$ \begin{alignat} ... $$` blocks still rendering the nesting error, which the earlier manual pass missed because it only looked for `align`. Adds playwright to the docs extra; CI caches the browser on the resolved playwright version. ~18s over 55 pages. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * docs: number equations per page, fix dark mode, and tidy the sidebar Four rendering problems visible to readers, plus two mathematical errors and a gap in the redirect map. Equations were inconsistent — some numbered, some not, and four carried hand-rolled tags like `\qquad (\dagger)`. All 107 display blocks across 19 notebooks and two prose pages now use one form: $$ \begin{aligned} ... \end{aligned} $$ (eq-descriptive-label) `aligned` rather than `align`, since `align` inside `$$` is the nesting bug fixed earlier in this branch. One number per block, not per row. The manual tags become real numbers and their prose mentions become `{eq}` links, as do two hard-coded "Equation (N)" references that would have gone stale. Numbering is per page. Sphinx's default `math_numfig = True` routes numbering through `env.toc_fignumbers`, which counts across the whole project in toctree order — `sharp_bits` opened at (105). With it False the number comes from `env.new_serialno('eqno')`, documented as "unique in the current document". Equations only: section and toctree numbering come from `:numbered:`, unused here. Dark mode: code blocks kept light-mode Pygments colours against a dark background. The accent `#7a2e2a` is built for a white page and measured 2.02:1 on `#111113`, under both the 4.5:1 AA text and 3:1 UI thresholds. Dark mode now uses `#d07b76` — same hue and saturation, lightness 0.32 -> 0.64 — at 6.09:1. Light keeps `#7a2e2a` at 9.33:1. Sidebar: the two migration guides are one page with a `##` per release rather than separate documents. shibuya's fold state is uniform by depth and cannot open some groups while folding others, so a folded parent would have needed custom JavaScript; one page gives the same tidy single line for free. Groups are regrouped so `design` and `sharp_bits` sit under Background rather than in with governance. Two mathematical errors, both reader-visible: - intro_to_gps: the marginal log-likelihood opened with a bare `& =` — the left-hand side `\log p(\mathbf{y})` was missing, though the prose either side names it. - classification: `\cos(2 * + \epsilon)` had a missing operand AND the wrong coefficient; the code cell below it uses `jnp.cos(3 * x + ...)`. Redirects: the map covered `api/*` and `_examples/*` but no prose page. MkDocs ran with `use_directory_urls`, so `/installation/`, `/design/`, `/sharp_bits/` and the rest were published as directory URLs and would all have 404'd after the cutover. Eight added; all 87 verified to resolve to files that exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * docs: put every citation through the {cite} role Three styles coexisted: {cite} roles, `<strong data-cite="key">` left over from the MkDocs bibtex plugin, and plain hyperlinks to papers. The data-cite tags were not merely inconsistent, they were invisible. Sphinx has no handler for them and most had empty bodies, so they rendered as nothing and left sentences broken mid-clause — barycentres read "It was shown in that the barycentre", uncollapsed_vi read "see Sections 3.1 and 4.1 of the excellent review paper ." All ten are now {cite:t} or {cite:p} depending on whether the citation is the grammatical subject, with the surrounding prose repaired. Ten hyperlinks that were really citations follow suit. Ordinary links do not: of the ~130 external links in the docs only these were scholarly references, so JAX and NumPyro docs, blog posts, Wikipedia, the UCI and NOAA dataset pages and Duvenaud's Kernel Cookbook site all stay as links. Five new refs.bib entries, each taken from a fetched source rather than written from memory: alvarez2012kernels (DOI 10.1561/2200000036, cross-checked against arXiv:1106.6251), bruinsma2020scalable and lu2022additive (BibTeX copied from their PMLR pages), duvenaud2014automatic (thesis title page plus the Cambridge repository record), and matern1960SpatialV lifted from gpjax/citation.py, which already carried a hand-verified copy. Corrects one pre-existing entry: hensman2013gaussian claimed `journal = {Artificial intelligence and statistics}`, but "Gaussian Processes for Big Data" is a UAI paper — arXiv:1309.6835 gives report number UAI-P-2013-PG-282-290. It is now an @inproceedings with the right venue and pages. references.md gains `:filter: cited` so the page reflects what the docs actually cite; three uncited entries stay in the .bib. The build is the fence: sphinxcontrib-bibtex warns on an unknown key and the docs gate runs -W, so a dangling {cite} fails rather than shipping. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * fix(docs): stop a third-party outage from failing the docs gate CI run 30694917815 failed the PR docs build with: WARNING: failed to reach any of the inventories with the following issues: intersphinx inventory 'https://docs.kidger.site/equinox/objects.inv' not fetchable due to 415 Client Error: Unsupported Media Type The same URL returns 200 from a laptop — bot protection reacting to a GitHub runner, not a broken link. Because the gate runs `-W`, an unreachable inventory is a hard failure, so any of the six upstream sites could break the build at any time for reasons entirely outside this repository. It also cannot reproduce locally: intersphinx caches inventories, so only a cold build refetches. Suppressing it is not an option — the warning is logged with no warning type (sphinx/ext/intersphinx/_load.py:326), so `suppress_warnings` cannot single it out, leaving a choice between a flaky gate and no `-W` at all. Each mapping now lists the upstream inventory followed by a vendored copy under docs/_inventories/ (411 KB total). intersphinx stops at the first that loads, and the log levels make this work: all locations failing warns, but some failing while another succeeds logs only info, so the fallback keeps the build green and silent. Verified by pointing the equinox remote at an unreachable host and confirming a cold build still exits 0. This is the same reasoning that vendored the notebook datasets: a documentation build should not depend on a third party being up. Also fixes three bibliography defects: - salimbeni2018 was @misc while carrying booktitle, series, volume, pages and publisher. @misc ignores all of them, so the AISTATS venue was silently dropped from the rendered entry. Now @inproceedings; metadata confirmed against PMLR v84 (84:689-697). - higham2022accuracy is dated 2002 — Crossref on DOI 10.1137/1.9780898718027 confirms 2002, SIAM. The key was the wrong part; renamed higham2002accuracy and the empty `address` filled. - lao2020tfp was @Article with `journal = {arXiv preprint arXiv:2002.01184}`. The arXiv API reports no journal_ref, so it is a preprint; now @misc with proper eprint fields. And the quote in sharp_bits.md attributed the numerical-analysis line to "Nicholas Highman". Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ * docs: group the examples by purpose instead of behind one "Examples" node The Examples group rendered a caption "EXAMPLES", then a page node "Examples", then the notebooks beneath it — the caption and the node said the same thing, and the notebooks sat a level deeper than everything else in the sidebar. The 22 notebooks now sit directly under four captions that say what they are for: Getting started installation and the five notebooks that teach the basics — intro to GPs and kernels, then regression, classification and count data Accelerating GPs the four about making inference scale: collapsed and uncollapsed sparse VI, state-space GPs, OILMM Applied modelling eight worked applications — barycentres, graph kernels, heteroscedastic noise, multi-output, OAK, ocean currents, spatial modelling, UCI benchmarking Guides for customisation the five about extending GPJax — kernels, likelihoods, deep kernels, NumPyro, backend design Every notebook appears exactly once and none is unlisted, asserted against the files on disk rather than by eye. docs/examples/index.md is deleted. It was the page behind the duplicated node, and a page in no toctree is an orphan, which fails the -W build. Its "What you'll learn" table is lost with it; the group names now carry that orientation, and the two entry points it recommended (New to Gaussian Processes? and Regression) are linked from the landing page instead. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01QXcSotTJspF11tgaJfrXrZ --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
conjugate_mll had no value-level test anywhere in the suite, yet serves as the oracle for the Kalman MLL and collapsed_elbo. These closed-form pins, computed through an independent jnp.linalg path, give the reference frame its ground truth ahead of the v1.0 conditioning refactor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…tter bug collapsed_elbo(z=X) vs conjugate_mll and whitened-vs-unwhitened predicts at matched parameters now guard the five independent derivations of the conjugate conditioning algebra. The strict xfail documents that at non-default jitter the derivations factorise different matrices — the bug the v1.0 conditioning module removes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
_compare previously swallowed AssertionError with a print, so the harness could never fail. Failures are now collected per-example and raised at the end of test(), making the four golden-value pins a real no-behaviour-change net for the v1.0 refactor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
The newly-loud harness exposed pre-existing drift in all four examples: the collapsed/uncollapsed goldens predated the real-data example swap (#696) and the regression/heteroscedastic goldens predated subsequent behaviour fixes (#707/#708/#713 and dependency bumps). The toothless harness never noticed. Re-pinned so the net measures the v1.0 refactor, not history. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
The full-dataset size a minibatch ELBO needs now travels on the one object that knows it, as static pytree aux_data, instead of being smuggled through likelihood constructors. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…nd noise_prior Of ~245 occurrences of num_datapoints, only four were real reads: the ELBO minibatch scale (now served by Dataset.n_total) and latent sizing (moves to data-contact time in the JointModel rewrite). No likelihood used the value internally, and nothing validated it — a wrong value silently mis-scaled the ELBO. noise_prior moves to the model layer, where priors live. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
The API now mirrors the maths: prior * likelihood -> JointModel (the joint p(f,y), the trainable object); model.condition(D) — sugar: model | D — returns an immutable Posterior pytree caching the Cholesky factor and representer weights. The predictive, log_marginal_likelihood, loo, and pathwise sample_approx are views of that one factorisation, deleting the eleven independent derivations and the two-owner jitter split (prior.jitter is now the single knob, applied once inside conditioning). - gpjax/conditioning.py: deep module (Posterior, ExactPosterior, LatentPosterior); MO validation moves to condition time; sample_approx refuses multi-output loudly instead of silently broadcasting wrong. - gps.py: Prior (AbstractPrior folded in), ConjugateModel, NonConjugateModel (lazy latent, sized at data contact), HeteroscedasticModel (owns noise_prior — likelihoods are pure conditionals again, killing the likelihoods->gps circular import). Deleted: AbstractPrior, AbstractPosterior, LatentPosterior marker, ChainedPosterior marker, construct_posterior (now construct_model). - objectives: conjugate_mll/conjugate_loocv/log_posterior_density are one-line views of the conditioned posterior. - fit: _prepare_model hook sizes lazily-initialised state from data. - predict(t, D) survives as documented one-line sugar everywhere. - return_covariance_type kwarg renamed to covariance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Mechanical: ConjugatePosterior->ConjugateModel and friends, construct_posterior->construct_model, return_covariance_type->covariance, num_datapoints/noise_prior constructor ceremony deleted (~240 sites). Semantic: heteroscedastic tests build HeteroscedasticModel directly; non-conjugate tests size the latent via init_latent; docs/index.md quickstart shows condition(); regression example narrates the condition API; StateSpaceConjugatePosterior renamed StateSpaceConjugateModel. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…ault jitter The model-side two-owner jitter bug is fixed; the strict xfail narrows to the family-side knob, which unifies in the variational stack PR. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
thomaspinder
force-pushed
the
natgrads-05-notebook-dual
branch
3 times, most recently
from
August 6, 2026 09:46
1e50a85 to
d713106
Compare
- reference/gps.md lists the JointModel hierarchy and the conditioning module; state_space.md and linalg.md updated for renamed/new symbols - stale glossary/sharp_bits/classification xrefs renamed - poisson example initialises the lazy non-conjugate latent before MCMC - ADR directory excluded from the docs site (in-repo records for now) - codeautolink match_block warnings suppressed on every path: a matcher limitation on doctest-SKIP blocks, predating this stack — the docs workflow had not run cold since the Sphinx migration, so tonight's PR pushes surfaced it for the first time Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…inery Implements Salimbeni et al. 2018 (arXiv:1803.09151) natural-gradient VI for VariationalGaussian and WhitenedVariationalGaussian. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
Adversarial review of the natural-gradient core turned up two correctness defects, one performance defect, three contract mismatches and a set of documentation and test gaps. Fixes, in order of severity: * `_first_valid_trial` leaked float64 into the scan carry. Under x64 the exponent `jnp.arange(K + 1)` is int64, so `backoff ** arange` was a non-weak float64 that promoted a float32 model and made `lax.scan` reject the carry. The trial ladder is now cast to the dtype of Theta_2, so `fit_natgrads` is no longer strictly narrower than `fit`. * The backoff replicated the whole theta -> xi map across all K+1 trials, at a measured 13% of total training wall clock at M=200 -- not the "negligible" cost impl-plan 1.2.5 assumed. Only the admissibility probe is replicated now; the inversion, the X^T X product and the second Cholesky run once, at the accepted step size. Measured overhead at M=200 falls to 4.6%. * `fit_natgrads`' signature rejected everything `_check_natgrad_lr` blessed: `natgrad_lr=1`, `map_jitter=0`, `backoff=1` and a 0-d array all raised under the beartype import hook. Annotations widened, the validator now accepts 0-d arrays and rejects bool, and the entry point (not just the validator) is tested with each. * `_reject_frozen_coordinates` matched coordinates against top-level dataclass fields by identity, so a future nested registration would have passed the guard silently. It now re-walks the tree with the selector as the `is_leaf` predicate and reports the full key path. Its message also pluralises and points at a remedy that exists -- the old one recommended freezing the whole family, which re-raises the same error. * `fit_natgrads` now calls the guard under `safe=True` as impl-plan 1.2.4 prescribes, and forwards `log_rate` to `vscan` instead of documenting a knob that did nothing. Tests: `test_natgrad_backoff_recovers_from_large_step` never exercised the backoff (the k=0 trial was already admissible at natgrad_lr=100), so it now starts from a tight S_0, asserts the un-shrunk step genuinely leaves the cone, and checks the accepted step size. The Cholesky-budget test could not see the vmap width; it is now an exact count parametrised over max_backoff, plus a lowered-IR test asserting the batched Cholesky has leading dimension K+1. Added dtype-preservation tests, and moved the duplicated conjugate oracle into `tests/_reference/conjugate_svgp.py` so the two transcriptions cannot drift. Docs: all seven exported functions gained runnable `Example:` blocks (impl-plan section 6), the map_jitter bias on `history` and the eta -> xi cancellation regime are now documented on the public surface, and two docstrings became raw strings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…_natgrads The ```pycon fences render fine under MkDocs but are invalid RST under Sphinx/napoleon, producing docutils warnings that fail the -W docs-ci gate. Drop the fences; the indented doctest block matches fit()'s style. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…tionVariationalGaussian Both classes were parameterisation-only: they stored the natural or expectation coordinates of q(u) but shipped no way to take a natural-gradient step in them, so they bought nothing over the standard families. Natural-gradient geometry belongs to the optimiser, not the family. The Fisher matrix is exactly the Jacobian dn/dt, so the natural gradient with respect to the natural parameters equals the ordinary gradient with respect to the expectation parameters, in any parameterisation. fit_natgrads (PR#1, gpjax/natural_gradients.py) therefore computes the transforms on the fly and operates directly on VariationalGaussian and WhitenedVariationalGaussian, which store constraint-respecting coordinates. Users of the removed classes should switch to VariationalGaussian with gpjax.fit_natgrads. Also drops the now-dead _psd helper and the cholesky_factor import, whose only call sites lived inside the deleted classes, and the "natural" and "expectation" arms of the VariationalParametrisationSuite ASV benchmark. BREAKING CHANGE: NaturalVariationalGaussian and ExpectationVariationalGaussian are removed from gpjax.variational_families. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
CHANGELOG: add the missing "### Added" entry for fit_natgrads and gpjax.natural_gradients. PR#1 shipped both without a changelog entry, so the Removed entry added here forward-referenced an API the changelog never announced. Bringing it forward from PR#3 keeps any release cut mid-stack self-consistent; PR#3 appends the dual entries to the same section. CHANGELOG: correct the justification prose. "the natural gradient with respect to theta equals the ordinary gradient with respect to eta, in any parameterisation" is false as literally written -- for a reparameterisation xi with J = dtheta/dxi, the natural gradient in xi is J^-1 grad_eta L, not grad_eta L. The identity is specific to the natural/expectation pair of an exponential family. Reworded to state that pairing and the point it supports: either coordinate system is recoverable on the fly, so no dedicated class is needed. benchmarks: drop diff-relative wording from the VariationalParametrisationSuite docstring. "surviving" only means something to someone reading this commit's diff, and "Both" is a count PR#3 invalidates when it re-adds the dual arm. The module docstring keeps its explicit "(standard, whitened)" list, which PR#3 must extend regardless. tests: split the _psd guard out of test_removed_families_are_gone. The _psd arm asserted `"_psd" not in __all__`, vacuous for a helper that was never exported, under a failure message about superseded parameterisations. It is now its own test with a docstring saying what it actually guards. CLAUDE.md: "Three optimisers" -> four. fit_natgrads landed in gpjax/fit.py in PR#1; the sentence has been stale since. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
The Sphinx docs arrived with v1.0, after this branch was cut, so the deletion of NaturalVariationalGaussian and ExpectationVariationalGaussian now has to reach three doc files the original commits could not know about: drop the two classes from the variational-families autosummary page, and repoint the glossary's "natural parameters" entry from the removed classes to fit_natgrads on the surviving families. fit_natgrads is added to the fit reference page so that glossary link resolves — an omission from PR#1, which shipped the function without a reference entry. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Implements the dual/site parameterisation of sparse variational GPs from Adam, Chang, Khan and Solin (2021), "Dual Parameterization of Sparse Variational Gaussian Processes", NeurIPS 2021 (arXiv:2111.03412). `DualVariationalGaussian` stores an unnormalised Gaussian site on the *centred* inducing outputs -- `dual_vector` is the site's first natural parameter and `dual_matrix` its precision, both `Real` and both defaulting to zero, so q(u) = p(u) at initialisation. Neither carries a constraining bijection: PSD-ness of the site precision comes from the convex-combination structure of the natural-gradient update, and a bijection would destroy that affine step. Moments, marginals, prior KL and predictions all route through the working matrix R = Kzz + Kzz L2 Kzz, which dominates Kzz and is therefore always factorisable even when the site precision is rank deficient; exactly two Cholesky factorisations are taken per call and nothing is inverted. The centred convention is the correction to the reference implementation's mean-function bug, which shifts by `predict_f(Z)` and so is wrong for any non-zero mean function. `test_dual_natgrad_handles_non_zero_mean_function` pins the correct behaviour and asserts that the uncentred variant misses. `marginals` adds the family's jitter to every marginal variance. This is load-bearing, not cosmetic: `VariationalGaussian.predict` runs `add_jitter` on its output covariance, so the per-point marginals `elbo` sees carry the same offset, and without it `dual_elbo` would miss `elbo` at matched moments by N*eps/(2 sigma^2). `dual_elbo` is the same functional as `elbo` but evaluated as a function of the sites and the hyperparameters. Its value matches `elbo` at the implied moments (measured 5.7e-14 absolute, 2.8e-16 relative at random PSD sites) while its hyperparameter gradient differs, because q moves with theta through Kzz while the sites stay frozen. Kzz is deliberately not detached and no moments are cached on the module; a cached implementation would pass every value assertion and fail only `test_dual_elbo_hyper_gradients_differ_away_ from_optimum`. `natural_gradient_step` and `variational_coordinates` gain dual registrations. Because grad_mu KL = lambda exactly, the KL is never differentiated and the step is a convex combination towards a closed-form target built from one `jax.grad` of `expected_log_likelihood` (Bonnet and Price), with N/B scaling, a trace-safe floor on beta and a symmetrise. That makes it the Salimbeni step at gamma = rho: matched initialisations agree to 1.4e-15 in (m, S) over six Bernoulli steps at every rate tested. One rho = 1 full-batch conjugate step lands on the Titsias optimum to 4.9e-15. `fit_natgrads` rejects a numeric step size above one on this family, since the update is a convex combination; schedules cannot be checked statically and are documented as the caller's responsibility. Plain `fit` on the family also works and is documented as gradient descent in the dual coordinates. The `VariationalParametrisationSuite` benchmark gains a `dual` arm, in both `benchmarks/objectives.py` and the hardcoded parametrize list in `tests/test_benchmarks_smoke.py`. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
Factor R in the Kzz basis. Forming R = Kzz + Kzz Lambda_2 Kzz explicitly carries a rounding error of order ||Kzz||^2 ||Lambda_2|| eps, which for a large-variance kernel (or in float32) exceeds lambda_min(R) ~= jitter, so chol(R) returned NaN and poisoned every later fit_natgrads iterate -- measured at RBF variance 1e4, M=80, default jitter, float64, where the matched VariationalGaussian run stayed finite. R = Lk (I + Lk^T Lambda_2 Lk) Lk^T is factorised instead, giving a lower-triangular Lr = Lk chol(I + G) whose inner matrix has lambda_min >= 1 - O(||G|| eps). Same two Choleskys per call, plus one M x M product. The "chol(R) never fails" and "no Cholesky a backoff could rescue" claims are softened to match. Also: split _gram_and_root off _working_matrices so the dual natgrad step stops discarding a chol(R); take tr(R^-1 Kzz) as ||Lr^-1 Lk||_F^2 rather than a full cho_solve; bound-check an optax schedule against rho <= 1 over the whole num_iters horizon for the dual family, which previously returned a silent all-NaN history; hoist the _fmt_Kzt_Ktt/_fmt_inducing_inputs hooks to AbstractVariationalGaussian and keep one typed _symmetrise; convert the dual family's numpydoc sections to the Google style the file and mkdocs use, and document marginals' inputs argument. Doc corrections, all measured: elbo on a DualVariationalGaussian returns the same value and the same gradients as dual_elbo (bit-identical value, 1.7e-14 on gradients), so the CHANGELOG's gradient claim now names the matched VariationalGaussian as the comparison; and vmap does not repeat the unbatched factorisations per datum, so neither elbo nor the benchmark arm pays 2N Choleskys -- eager counts are 4 (dual), 2 (standard), 1 (whitened), and the compiled dual_elbo and dual step are 2 potrf each. Tests: regression for chol(R) at variance 1e4/M=80 and 1e3/M=50 with the default jitter, an Lr Lr^T = R reconstruction check, the schedule guard, a half-batch arm on the dual/elbo equivalence plus a direct pin on the N/B factor, and the triplicated dual fixtures moved to tests/_dual_helpers.py. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
…e v1.0 API The v1.0 likelihoods are pure conditionals, so the minibatch ELBO scale in dual_elbo and the dual natural-gradient step now derives from Dataset.n_total instead of likelihood.num_datapoints, the constructor annotation follows the AbstractPosterior -> JointModel split, and the docstring examples plus the dual test fixtures drop num_datapoints. The minibatch tests stamp n_total onto their batch views to preserve the N/B factor they assert. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Adds examples/natgrads.py, the first of two tutorial notebooks for the natural-gradient stack. It derives the exponential-family view of q(u), shows that the Fisher information is the Jacobian d(eta)/d(theta) (checked numerically to 2.9e-15), reads the update as mirror descent, and then runs two demos: * a conjugate 1D regression where one gamma=1 full-batch natural-gradient step recovers the Titsias optimum to 1.2e-13 while Adam on the same problem is still 4.6 nats short after 2000 iterations; * a mini-batched 2D banana Bernoulli benchmark (N=2000, M=50, B=256, 1000 iterations) comparing natural gradients + Adam against Adam alone, per iteration and per wall-clock second. Closes with the negative-definite cone result, a gamma sweep reproducing its boundary, and a demonstration of the step-size backoff. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
Correct the mini-batch ramp argument (the target is q-dependent outside conjugacy, so gamma=1 lands on a moving fixed-point target, not the mini-batch optimum), replace the unmeasured calibration claim after the banana contours with the metrics the cell actually prints, and attribute the cone sweep's gamma=2 failure to the over-confident S_0 rather than to gamma=2 itself. Smaller corrections: the Fisher solve is O((M + M(M+1)/2)^3) in the vec_s coordinates, not O((M + M^2)^3); the one-step demo agrees to ~1e-13, not fourteen decimal places; the sparse/exact predictive deviation is quantified and located outside the data range; the Adam ELBO-gap description now matches the shape of the log-log panel; the roadmap says "exact variational optimum" where the notebook later reserves "exact posterior" for the non-sparse GP; K=100 is explained against the paper's dataset-dependent K. Code: derive the ELBO figure's y-limits from the smoothed histories so neither curve is clipped, guard the crossing report against never reaching the target, draw the held-out points in the categorical palette with per-class markers instead of the contour colourmap, use banana_data in the split report, and run ruff format (the pinned make_banana block is byte-identical afterwards). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
Likelihoods lose num_datapoints, prior * likelihood is now narrated as the
joint model (variables renamed *_posterior -> *_model), and the exact-GP
comparison uses model.condition(D) followed by a posterior query instead of
predict(x, train_data). Sphinx-migration adaptations: the notebook moves to
docs/examples/, imports utils from the notebook directory, uses {cite:t}
citations and a relative link to the uncollapsed-VI notebook, gains the
nb-download line, and is wired into the docs/index.md toctree. Verified by
executing the notebook end to end; all printed diagnostics match the
narration.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
Adds examples/dual_svgp.py, the second of the two natural-gradient tutorial
notebooks, and wires both of them into the documentation.
The notebook derives the dual parameterisation for a reader who has been
through examples/natgrads.py: the additive split eta = eta_0(theta) + lambda,
the EP-style likelihood sites and their tying to inducing space, the two
convention traps (flanked vs un-flanked storage, and the -1/2 on Lambda_2),
and the tied natural-gradient update, which is an affine convex combination on
the stored sites because grad_mu KL == lambda exactly, so the KL is never
differentiated and no theta <-> eta round trip is needed.
Measured in the executed run:
* one rho = 1 full-batch step on a conjugate model with a non-zero mean
function reproduces the Titsias optimum to 1.7e-12 (mean) and 1.2e-13
(covariance), a second step moves nothing, and dual_elbo matches the
collapsed bound up to exactly N * jitter / (2 sigma^2);
* rho is gamma: matched dual and Salimbeni E-steps agree in (m, S) to 3.1e-15
over six full-batch steps at rho in {0.3, 0.8, 1.0}, and two frozen-
hyperparameter fit_natgrads runs overlay to 4.3e-14 over 50 iterations;
* the banana benchmark (N = 2000, M = 50, B = 256, 1000 iterations, the same
make_banana and jr.key(42) as the natural-gradients notebook) with t-SVGP,
the Salimbeni natgrad and Adam alone, per iteration and per wall-clock
second -- both natural-gradient runs reach Adam's 1000-iteration bound at
iteration 109, and the dual iteration is not cheaper at this scale;
* hyperparameter learning: the dual/standard hyper-gradient gap falls from
3.9e1 to 8.5e-15 as the E-step converges; dual dominance is not uniform when
sparse (negative gaps at M = 5 and M = 10) but holds by 1.5e5 nats at Z = X;
and a 40-round VEM loop ends 0.50 nats ahead on dual_elbo with equal
held-out NLPD.
Docs wiring: both notebooks added to the mkdocs.yml Tutorials nav and to
CTA_NOTEBOOKS in docs/scripts/gen_examples.py, plus an adam2021dual entry in
docs/refs.bib.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
Corrections to examples/dual_svgp.py, all re-verified against a fresh end-to-end execution: * the intro no longer implies the two hyperparameter gradients agree at theta_t; they agree only at a converged E-step, which is what the notebook's own gradient table measures; * added a notation-reconciliation note bridging the natural-gradients notebook's (theta, eta, lambda) to this one's (eta, mu, lambda), and restated the borrowed identity and H_2 in these letters; * the c(theta) remark now names the site convention it holds under (normalised projected) and fixes the sign apposition; * the bound-slice prose is now asymmetric, as the data are: l collapses on the long-lengthscale side while l-bar barely moves, but both collapse together on the short side, where the sparse approximation itself has failed. Added edge diagnostics to back it; * the banana benchmark now states which run ends ahead and bounds what that comparison can mean; * the VEM panel plots the round-by-round bound lead rather than two indistinguishable traces. That exposed a false claim: the dual M-step is behind for the first seven rounds, crosses at round 8 and holds a sub-nat lead thereafter. Prose corrected and the crossing printed; * order-of-magnitude claims restated from the printed values (Titsias agreement, the jitter residual now printed to twelve digits, the banana condition-number ratio, the M = 20 crossing-point noise); * the dominance row now names conjugacy as well as Z = X; * the roadmap names the two sections it had omitted; * the banana-copy rationale no longer overstates cross-notebook comparability, and the wall-clock explanation leads with the O(M^3)-vs-O(BM^2) argument rather than an evaluation XLA folds away; * ruff-format clean, no code line over 88 characters. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
`mkdocs build` aborted with `IndexError: string index out of range` while
rendering `_examples/dual_svgp.md`. The trigger is a markdown-katex parser
bug: `iter_inline_katex` reads `line[end + 1]` without a bounds check, so any
line whose final characters are a backtick code span immediately preceded by
`$` crashes the build. The prose read
... that is $-$`sparsity_gap`
above, and ...
where the closing `$` of `$-$` abuts the code span and the span ends the line.
Reword to "the negated `sparsity_gap` computed above", which removes the
`$`-adjacent code span entirely rather than relying on a particular line wrap.
The meaning is unchanged, and the sentence still says c(theta) is minus the
Titsias trace term. A scan of all generated `docs/_examples/*.md` confirms this
was the only occurrence of the pattern in the docs tree.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
The dual/Salimbeni E-step divergence on the banana demo was attributed to floating-point conditioning. It is not: `inv_probit` clips its output into [1e-3, 1-1e-3], so the computed Bernoulli log-likelihood is not log-concave in the tails (positive second derivative for f < -2.44). A confidently mislabelled point then yields beta_i < 0, the dual branch's beta_floor clips it, and the two branches diverge. With the clip disabled the same six steps agree to 6.2e-13 instead of 5.1e-3. - Re-attribute the mechanism in the dual notebook and add a diagnostic cell that measures it, and qualify the "identical iterates" claim wherever it is stated (natural_gradients.py, fit.py, CHANGELOG, both notebooks): it holds provided the computed beta stays non-negative. - Note in the natgrads notebook that the cone discussion assumes a log-concave *computed* likelihood, which the clipped probit violates in the far tails. - Reword the Salimbeni registration docstrings: dispatch covers GraphVariationalGaussian, but the standard elbo path is broken upstream (MatrixLinearOperator dimensionality error out of gram), on main too. - Extend _check_natgrad_schedule to reject non-positive rates for every family, not just the dual upper bound. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD
Likelihoods lose num_datapoints, and the prior * likelihood factories are
renamed to say what they build (conjugate_model, logit_model, banana_model,
vem_joint_model) with the num_datapoints parameter dropped from the family
helpers. Sphinx-migration adaptations: the notebook lives in docs/examples/,
imports utils from the notebook directory, uses {cite:t} citations, links to
the natgrads notebook relatively, gains the nb-download line, and is wired
into the docs/index.md toctree in place of the removed mkdocs.yml nav; the
natgrads notebook's closing pointer becomes a live relative link. Verified by
executing the notebook end to end; the printed diagnostics match the
narration (Titsias optimum to ~1e-12, rho = gamma to 3e-15, beta < 0 first at
step five, both natural-gradient runs crossing Adam's bound at iteration 109).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
thomaspinder
force-pushed
the
natgrads-05-notebook-dual
branch
from
August 6, 2026 10:41
d713106 to
c6d2bce
Compare
|
📖 Docs preview: https://pr-731--endearing-crepe-c2d5fe.netlify.app Smoke render — the expensive notebooks run with reduced budgets, so |
Collaborator
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This draft PR exists so the entire natural-gradients stack can be reviewed as one diff. Merge the stack in order instead: #714 → #715 → #716 → #729 → #730 (each auto-retargets to
natgradsas its predecessor merges and its branch is deleted). This PR will close itself once #730 lands.What the stack delivers
Natural-gradient variational inference for GPJax, implementing:
fit_natgrads()+gpjax/natural_gradients.py(feat(natural-gradients): fit_natgrads + exponential-family machinery [stack 3/10] #714), with the vestigial parameterization-only families removed (refactor(variational)!: remove vestigial Natural/Expectation families [stack 4/10] #715).DualVariationalGaussian+dual_elbo+ the tied-site natural-gradient update, dispatched through the samefit_natgrads()entry point (feat(variational): DualVariationalGaussian + dual_elbo (t-SVGP) [stack 5/10] #716).Every stack point passed
uv run poe all-tests; the tip passed the fullpoe docs-buildwith both notebooks executing end-to-end.Issue Number: N/A
🤖 Generated with Claude Code
https://claude.ai/code/session_01VHxYz2P7JCSqpax5RmDwWD