Skip to content

test: v1.0 safety net — closed-form oracles, equivalence pins, loud integration failures [stack 1/10] - #743

Merged
thomaspinder merged 7 commits into
v1.0from
v1-01-safety-net
Aug 7, 2026
Merged

thomaspinder merged 7 commits into
v1.0from
v1-01-safety-net

Conversation

@thomaspinder

@thomaspinder thomaspinder commented Aug 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Stack PR 1/10 for the v1.0 conditioning architecture (design: 2026-08-06 grilling session; plan: plans/2026-08-06-v1-conditioning-implementation.md). This PR is the safety net that must be green before any math moves in PR 2 — it is independently mergeable and changes no library behaviour.

1. Closed-form oracles (tests/test_conditioning_oracle.py)

conjugate_mll had no value-level test anywhere in the suite, yet serves as the oracle for the Kalman MLL and collapsed_elbo — the reference frame had no ground truth. Three closed-form pins (MLL, predict, LOOCV) on a tiny fixed problem, computed through an independent jnp.linalg path.

2. Cross-derivation equivalence pins (tests/test_conditioning_equivalence.py)

  • collapsed_elbo(z = X) ≈ conjugate_mll (to 2e-4; the identity is exact only in the jitter→0 limit since the family's jitter enters Kzz while the model's enters Σ)
  • whitened vs unwhitened variational predicts at matched parameters
  • strict xfail: the same identity at non-default jitter — documents the two-owner jitter bug (Prior.jitter vs Posterior.jitter vs q.jitter factorise different matrices). PR 2 resolves the model-side split.

3. Integration harness now fails loudly (tests/integration_tests.py)

_compare swallowed AssertionError with a print, so the harness could never fail. Failures are now collected and raised.

⚠️ Hardening immediately exposed pre-existing golden drift on main — the pinned values dated from before the real-data example swaps (#696) and subsequent behaviour fixes, and the toothless harness never noticed:

example variable old golden actual on main
regression predictive_mean 36.24 37.91
regression predictive_std 197.05 202.37
collapsed-vi history 1924.76 1851.12
collapsed-vi predictive_mean -8.40 +1.37
collapsed-vi predictive_std 255.75 248.32
uncollapsed-vi history -2678.41 +59440.08
uncollapsed-vi meanf -54.15 -55.19
uncollapsed-vi sigma 121.43 555.41

The collapsed/uncollapsed magnitudes are consistent with the examples having been rewritten onto real Open-Meteo data — different dataset, different scale — not with numerical regression. Goldens are re-pinned to current-main behaviour so the net measures the refactor, not history. Please sanity-check the re-pinned values in review.

Test plan

  • uv run pytest tests/test_conditioning_oracle.py tests/test_conditioning_equivalence.py → 5 passed, 1 xfailed
  • uv run python tests/integration_tests.py → all four examples green against re-pinned values

Part of the v1.0 stack: PR 2 (conditioning architecture) is stacked on this branch. Do not merge yet — stack lands together for v1.0.

🤖 Generated with Claude Code

https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB

thomaspinder and others added 4 commits August 6, 2026 01:27
conjugate_mll had no value-level test anywhere in the suite, yet serves as
the oracle for the Kalman MLL and collapsed_elbo. These closed-form pins,
computed through an independent jnp.linalg path, give the reference frame
its ground truth ahead of the v1.0 conditioning refactor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…tter bug

collapsed_elbo(z=X) vs conjugate_mll and whitened-vs-unwhitened predicts at
matched parameters now guard the five independent derivations of the
conjugate conditioning algebra. The strict xfail documents that at
non-default jitter the derivations factorise different matrices — the bug
the v1.0 conditioning module removes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
_compare previously swallowed AssertionError with a print, so the harness
could never fail. Failures are now collected per-example and raised at the
end of test(), making the four golden-value pins a real no-behaviour-change
net for the v1.0 refactor.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
The newly-loud harness exposed pre-existing drift in all four examples: the
collapsed/uncollapsed goldens predated the real-data example swap (#696) and
the regression/heteroscedastic goldens predated subsequent behaviour fixes
(#707/#708/#713 and dependency bumps). The toothless harness never noticed.
Re-pinned so the net measures the v1.0 refactor, not history.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
@github-actions

github-actions Bot commented Aug 5, 2026

Copy link
Copy Markdown

📖 Docs preview: https://pr-743--endearing-crepe-c2d5fe.netlify.app

Smoke render — the expensive notebooks run with reduced budgets, so
figures are not publication fidelity. /render-mode.txt says smoke.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts
on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A
stacked PR targets its predecessor rather than `main`, so neither job ever
ran on one: PRs in the current v1 stack report 16 check runs where a
main-targeting PR reports 18.

The effect is that a stack gets reviewed, approved and declared green with no
security gate at any point — the scans first fire when each PR auto-retargets
to `main` at merge time, which is the worst moment to discover a verified
secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had
ever been scanned.

Drop the filter so the trigger follows the code rather than the target branch.
Both jobs are already branch-agnostic — TruffleHog scans `./` over full
history and the dependency scan reads the checkout — so neither needs a base
ref to work off `main`.

`push` stays pinned to `main`: post-merge scanning of every branch push would
duplicate the PR run for no extra signal.

(cherry picked from commit a8d6603)
…ames

Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1
stack for the first time, and it immediately failed three PRs (#714, #716,
#748) on four "verified" secrets. All four are pytest function names:

  test_fit_natgrads_accepts_optax_schedule
  test_dual_natgrad_matches_salimbeni_step
  test_dual_working_matrices_reconstruct_r
  test_prior_kl_matches_textbook_reference

TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and
each of those names is exactly 40 characters long. Worse, the Lob verifier
reports them as *verified*, so `--only-verified` does not filter them: a fresh
repository containing nothing but one of those names reproduces
`Lob verified=true` under trufflehog 3.96.0, the same version the action pins.

`Secrets Detection` is a required context on `main`, so this is not cosmetic —
it means an arbitrary subset of PRs, selected by test-name length, can never
merge. GPJax has no Lob (direct-mail API) integration, so excluding the
detector costs no coverage. Verified that the flag is parsed rather than
silently ignored: an unrecognised detector name makes trufflehog exit with
"unrecognized detector type", and the four findings disappear with the flag
and return without it.

(cherry picked from commit 83ad296)
@thomaspinder thomaspinder changed the title test: v1.0 safety net — closed-form oracles, equivalence pins, loud integration failures [v1 stack 1/N] test: v1.0 safety net — closed-form oracles, equivalence pins, loud integration failures [stack 1/10] Aug 7, 2026
@thomaspinder thomaspinder changed the title test: v1.0 safety net — closed-form oracles, equivalence pins, loud integration failures [stack 1/10] test: v1.0 safety net — closed-form oracles, equivalence pins, loud integration failures [stack 1/10] Aug 7, 2026
@thomaspinder
thomaspinder changed the base branch from main to v1.0 August 7, 2026 13:38
@thomaspinder
thomaspinder merged commit 74a9106 into v1.0 Aug 7, 2026
22 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant