test: v1.0 safety net — closed-form oracles, equivalence pins, loud integration failures [stack 1/10] - #743
Merged
Merged
Conversation
conjugate_mll had no value-level test anywhere in the suite, yet serves as the oracle for the Kalman MLL and collapsed_elbo. These closed-form pins, computed through an independent jnp.linalg path, give the reference frame its ground truth ahead of the v1.0 conditioning refactor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
…tter bug collapsed_elbo(z=X) vs conjugate_mll and whitened-vs-unwhitened predicts at matched parameters now guard the five independent derivations of the conjugate conditioning algebra. The strict xfail documents that at non-default jitter the derivations factorise different matrices — the bug the v1.0 conditioning module removes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
_compare previously swallowed AssertionError with a print, so the harness could never fail. Failures are now collected per-example and raised at the end of test(), making the four golden-value pins a real no-behaviour-change net for the v1.0 refactor. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
The newly-loud harness exposed pre-existing drift in all four examples: the collapsed/uncollapsed goldens predated the real-data example swap (#696) and the regression/heteroscedastic goldens predated subsequent behaviour fixes (#707/#708/#713 and dependency bumps). The toothless harness never noticed. Re-pinned so the net measures the v1.0 refactor, not history. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
|
📖 Docs preview: https://pr-743--endearing-crepe-c2d5fe.netlify.app Smoke render — the expensive notebooks run with reduced budgets, so |
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB
This was referenced Aug 6, 2026
`Dependency Vulnerability Scan` and `Secrets Detection` are required contexts on `main`, but the workflow filtered `pull_request` to `branches: [main]`. A stacked PR targets its predecessor rather than `main`, so neither job ever ran on one: PRs in the current v1 stack report 16 check runs where a main-targeting PR reports 18. The effect is that a stack gets reviewed, approved and declared green with no security gate at any point — the scans first fire when each PR auto-retargets to `main` at merge time, which is the worst moment to discover a verified secret. Of the 45 commits in the v1 stack only the 5 on the bottom branch had ever been scanned. Drop the filter so the trigger follows the code rather than the target branch. Both jobs are already branch-agnostic — TruffleHog scans `./` over full history and the dependency scan reads the checkout — so neither needs a base ref to work off `main`. `push` stays pinned to `main`: post-merge scanning of every branch push would duplicate the PR run for no extra signal. (cherry picked from commit a8d6603)
…ames Unfiltering the `pull_request` trigger made `Secrets Detection` run on the v1 stack for the first time, and it immediately failed three PRs (#714, #716, #748) on four "verified" secrets. All four are pytest function names: test_fit_natgrads_accepts_optax_schedule test_dual_natgrad_matches_salimbeni_step test_dual_working_matrices_reconstruct_r test_prior_kl_matches_textbook_reference TruffleHog's Lob key pattern is `(live|test)_` plus 35 more characters, and each of those names is exactly 40 characters long. Worse, the Lob verifier reports them as *verified*, so `--only-verified` does not filter them: a fresh repository containing nothing but one of those names reproduces `Lob verified=true` under trufflehog 3.96.0, the same version the action pins. `Secrets Detection` is a required context on `main`, so this is not cosmetic — it means an arbitrary subset of PRs, selected by test-name length, can never merge. GPJax has no Lob (direct-mail API) integration, so excluding the detector costs no coverage. Verified that the flag is parsed rather than silently ignored: an unrecognised detector name makes trufflehog exit with "unrecognized detector type", and the four findings disappear with the flag and return without it. (cherry picked from commit 83ad296)
This was referenced Aug 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Stack PR 1/10 for the v1.0 conditioning architecture (design: 2026-08-06 grilling session; plan:
plans/2026-08-06-v1-conditioning-implementation.md). This PR is the safety net that must be green before any math moves in PR 2 — it is independently mergeable and changes no library behaviour.1. Closed-form oracles (
tests/test_conditioning_oracle.py)conjugate_mllhad no value-level test anywhere in the suite, yet serves as the oracle for the Kalman MLL andcollapsed_elbo— the reference frame had no ground truth. Three closed-form pins (MLL, predict, LOOCV) on a tiny fixed problem, computed through an independentjnp.linalgpath.2. Cross-derivation equivalence pins (
tests/test_conditioning_equivalence.py)collapsed_elbo(z = X) ≈ conjugate_mll(to 2e-4; the identity is exact only in the jitter→0 limit since the family's jitter enters Kzz while the model's enters Σ)Prior.jittervsPosterior.jittervsq.jitterfactorise different matrices). PR 2 resolves the model-side split.3. Integration harness now fails loudly (
tests/integration_tests.py)_compareswallowedAssertionErrorwith aprint, so the harness could never fail. Failures are now collected and raised.main— the pinned values dated from before the real-data example swaps (#696) and subsequent behaviour fixes, and the toothless harness never noticed:The collapsed/uncollapsed magnitudes are consistent with the examples having been rewritten onto real Open-Meteo data — different dataset, different scale — not with numerical regression. Goldens are re-pinned to current-
mainbehaviour so the net measures the refactor, not history. Please sanity-check the re-pinned values in review.Test plan
uv run pytest tests/test_conditioning_oracle.py tests/test_conditioning_equivalence.py→ 5 passed, 1 xfaileduv run python tests/integration_tests.py→ all four examples green against re-pinned valuesPart of the v1.0 stack: PR 2 (conditioning architecture) is stacked on this branch. Do not merge yet — stack lands together for v1.0.
🤖 Generated with Claude Code
https://claude.ai/code/session_019d7TF7oQt2Du4EQ74bMBQB