Repository navigation
docs(tests): audit every test that does not run, assert, or exist - #263
Conversation
Answers four questions with per-test evidence: 5 tests are @disabled (not the 7 a raw grep returns - one hit is javadoc, one is a @DisabledOnOs platform guard), 1 has an empty body, 4 are placeholders, and 4 more were deleted rather than implemented by confluentinc#493 while the branch that would have written them never merged. Only 1 of the 5 disabled tests records why. The two core ones were both disabled by c1fefbc "Create and commit offset map" in 2020 and have been dark since; the real gap is the end-to-end per-CommitMode assertion that offset commits respect key-order blocking. Also records two categories nobody asked about because an annotation grep cannot see them: 15 of 289 tests assert nothing, and OffsetEncodingTests reports green with most of its assertions branched away by a helper named assumeWorkingCodec that is not an assumption. Absorbs docs/test-hardening/disabled-and-weakened-tests-audit-2026-04-22.md, which existed only on the unmerged refactor/test-hardening branch, and corrects the two reasons its own git history refutes. Findings are keyed by class and method rather than line number - that predecessor's line numbers had drifted ~22 lines in four months. Records only. No test behaviour, assertion, timeout or volume is changed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SizQDD2hVUjD7EhESe9Hkb
A 455-line audit of disabled, kneecapped and weakened tests sat unread on
refactor/test-hardening for four months. Not because it was wrong - because
docs/refactoring.md described that branch by its OTHER commit ("OOM diagnostics
for LargeVolumeInMemoryTests at 1M") and never mentioned the audit, while
docs/inflight/branch-stale-and-diagnostic.md filed the branch under Superseded,
i.e. safe to delete. The only copy of unique content was on the delete list,
indexed under a description that did not describe it.
Both entries now say what the branch actually holds and where the content went.
The branch is genuinely safe to delete once its 1M/OOM commits are salvaged.
Also corrects the origin/refactor/empty-tests entry: its removal half landed on
master via confluentinc#493, so only the implement half is still open, and the
four tests it would restore are now named.
Adds the deferred work the audit found to the refactoring backlog - the
@timeout(60000L) unit bug, the assumeWorkingCodec misnomer, the JUnit 4
Assume on the Jupiter classpath, and the key-order commit coverage gap - plus
an AGENTS.md pointer so the next agent finds docs/test-hardening/ without
knowing it exists.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SizQDD2hVUjD7EhESe9Hkb
An independent falsification pass against the audit found the evidentiary spine sound - every git attribution survived, including the two hardest - but the counting was not publishable. Corrected here: - confluentinc#493 deleted TEN tests, not four. The six omitted were the most substantive intent in the commit: dead-letter queue, retry ordering, and in-flight backpressure. Also names confluentinc#494, the second follow-up that never merged, and records that only 3 of the 10 ever got NotImplementedException stubs on the restore branch. - Denominator was 289; it is 292 (301 with @archtest). @CartesianTest contributes 3 methods in TransactionAndCommitModeTest and is meta-annotated @testtemplate, so "zero @testtemplate" was true only of the literal string. - LargeVolumeInMemoryTests was recorded as fixed by PR #49. It is NOT fixed: master still runs 500 messages with 1_000_000 commented out, and PR #49 never touched the file. The predecessor audit carried a caveat to re-check this before flagging, and dropping that caveat is how an open defect became a closed one. This was the worst error in the document. - Quarantine script subclasses hold 27 tests, not 29; assumeWorkingCodec has 5 call sites, not 6; assumption aborts are 5 sites, not 6; the example modules have 6 test methods, not 9. - 9.4 claimed ~15s of unfailable sleep; ProgressBarTest is @disabled and costs nothing, so it is ~5s. - 8.5 refuted a "Short superseded Long" claim the predecessor never made - that was a hypothesis raised during investigation, now labelled as such. Three of the wrong counts came from one habit: counting raw grep hits instead of enumerating what matched, which is the exact error the document opens by warning about. The reproduction command has been replaced with one that enumerates annotation shapes rather than trusting a hard-coded whitelist. Also stops underselling existing key-order coverage: TransactionAndCommitModeTest runs the full CommitMode x ProcessingOrder cross-product including KEY. The central gap claim survives anyway - those tests assert counts, not committed offsets, and run on one partition. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SizQDD2hVUjD7EhESe9Hkb
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
✅ Duplicate Code ReportTwo engines run in parallel for cross-validation. Each has its own thresholds tuned to its baseline - the real safety net is the per-engine "max increase vs base" check. ✅ PMD CPD
No new clones introduced by this PR. ✅ jscpd (language-agnostic)
No new clones introduced by this PR. Powered by astubbs/duplicate-code-cross-check |
|
Claude finished @astubbs's task in 5m 55s —— View job Review complete
This is a documentation-only PR (audit doc + plan doc + two ledger edits + one Nearly everything checked out. All disabled-test attributions, commit hashes/dates, code snippets (empty test bodies, commented-out assertions, One confirmed finding, posted as an inline comment:
Inline comment: #263 (comment) |
✅ SpotBugs ReportNo bugs found (new bugs only — baseline from base branch excluded). |
The issue-reference gate caught three bare numbers below #1000 on added lines, and it is right to: the fork's numbering sits inside upstream's range, so a bare number is a coin flip. - "PR #49" (twice, in the audit and the backlog) is a fork PR: #49. - "(#190)" was quoted verbatim from a squashed commit subject, which is exactly how an upstream number leaks in looking innocent. Reworded so the quote stops at the subject and the number is cited as confluentinc#190 outside it. bin/check-issue-refs.sh now passes on this branch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01SizQDD2hVUjD7EhESe9Hkb
|
Claude finished @astubbs's task in 1m 43s —— View job Review complete
This PR was synchronized with one new commit, The one finding from my previous review is still open — nothing in this new commit touches
No other issues found — the rest of the diff (ledger corrections, the 640-line audit doc, the 409-line plan doc) checks out as it did in the previous pass. |
…y PR The data is hand-written and a documentation generator reads it, so a malformed file or a missing required field produces a broken generated page rather than a loud error - and the failure reaches a reader rather than the author. bin/check-docs-data.sh parses every file, checks it declares a kind the schema knows, checks the fields that kind requires are present, and checks any readme_anchor resolves to a real anchor in the template. Structure only, deliberately. Nothing can check whether the claims are true, and a gate that implied otherwise would be worse than no gate. Verified by negative control rather than by assumption: emptying a required field turns it red with a named finding, restoring it turns it green. It also passes the repo's own script gates for sigpipe and copyright headers. Also records the disabled-test release gate as an inflight entry pointing at #263, which is auditing every test that does not run, assert or exist. No separate work needed there.
0.6.0.0 does not ship while any test is disabled. The four carrying @disabled all predate the fork and one is a stub whose body asserts false, so it is inherited debt rather than a rule being broken - but the testing data asserts flake discipline, and a reader who greps for @disabled a minute later is exactly who it is written for. Quarantining does not clear the gate, since a release is separately blocked while the registry is non-empty. Tracked by #263, so no separate work is needed. maxFailureHistory is settable and read nowhere in the tree. Its feature record was written and then removed rather than shipped, because a page for it would tell a user to configure something inert. Recorded as a defect with the decision it needs: implement the retention, or remove the option, the latter being an API change belonging with 1.0 settlement. The Connect and Streams records are held until their modules exist, rather than shipping a Maven coordinate that will not resolve. And the 1.0 release train issue needs grooming: it has not been touched in years, and once the roadmap data owns 1.0 it becomes a stale second account on the public tracker. Its big-picture items moved to the data; the shared-nothing refactor and removing the streaming interfaces are backlog rather than gates, since the thread complexity the first was raised for has largely been fixed.
…e-tests # Conflicts: # AGENTS.md
🧪🔒 Quarantine Lane Report
🔴 expected while the owner PR is open · 🟡🎲 flapper, pass proves nothing · 🚨 a deterministic quarantined test passing means its fix landed: delete its |
Both branches now have master merged. Records the merge order, the AGENTS.md and TODO_INDEX rename conflicts already resolved (so they are not re-litigated), the three units blocked on #260, and two things a reader would otherwise get wrong: the audit's quarantine count has drifted since it was written, and the core unit suite was not re-run locally after the merge. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01L6PnsPmbwjazSu13M8JYm5
The plan's R19 asked for a pointer to `docs/test-hardening/` inside AGENTS.md's `## Testing` section, "in the style of the existing todo-index.md and quarantine pointers, so the next agent finds it without knowing it exists". What shipped was only the docs-map table row under "Where things live" - the Testing section, which is where an agent debugging a dark test actually reads, said nothing about the audit. The plan's Definition of Done nonetheless claimed the pointer was added. That is the exact discoverability failure this audit exists to fix: its 2026-04-22 predecessor rotted unread for four months because nothing pointed at it from the path an agent walks. Adds the bullet alongside the quarantine one, and names the current dated audit so the reader lands on a file rather than a directory listing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The handoff note added in 2d5c5b3 cited both PRs as bare `#263`/`#264`, which is exactly what the PR Checklist issue-ref gate forbids: this fork's numbers sit inside confluentinc's range, so a bare number is a coin flip on which repo it means. The gate flagged 13 of them and went red on this PR - a document about the stack's merge readiness was itself the thing blocking the merge. Qualifies every reference as `astubbs#NN`, the form the gate strips. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…hanges) Directory move only, so every path is 100% similar and git's exact-rename detection cannot fail on it. The content edits follow in the next commit. Generated by bin/rename-packages.sh.
Text edits only. No file moves in this commit, so it cannot dilute the rename detection in its parent. Generated by bin/rename-packages.sh.
…e-tests # Conflicts: # README.adoc # bin/rename-packages.sh # parallel-consumer-core/src/test/java/bz/stub/parallelconsumer/TestConventionRules.java # src/docs/README_TEMPLATE.adoc
…ediation Ancestry only - this merge changes no file. `git diff` against the previous commit is empty, so the tree is byte-identical to the one verified above. The PR read CONFLICTING against its own base because #263 and #264 had each run the io.confluent -> bz.stub rename independently, giving the same move two unrelated histories with no common ancestor to reconcile them. Merging master (#294) into this branch supplied that ancestor: master's rename is now in both lines, so the duplicate moves reconcile and the base merges clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…remediation #263 was squash-merged, so GitHub retargeted this PR onto master and the content this branch already carried came back as three add/add and content conflicts - the squash gave the same documents a second, unrelated history. master also gained #260 and #277. All three conflicts are the same shape: master's copy is #263's merged state, and this branch's copy is that state plus the corrections made here afterwards. Verified rather than assumed - master's audit is byte-identical to #263's tip, and nothing was lost on master. - inactive-tests-audit-2026-08-08.md: kept this branch's copy. master's is the earlier draft; this one carries the "Corrected 2026-08-08" pass (the nine restated claims, the §4 rewrite, the disposition of all ten deleted stubs). - refactoring.md: kept this branch's copy. Taking "both sides" would have been wrong here - master still lists the three `@Timeout(60000L)` annotations as work, which this branch deliberately moved to "Not listed as work" because #206 owns them, recording that `@Timeout(60)` would have been wrong (two of those tests wait 45s and 50s internally). master also still says the OOM diagnostics are unsalvaged; they are salvaged. - inflight/branch-stale-and-diagnostic.md: same - master's copy predates the OOM salvage this branch records as done. Verified on Temurin 17 after the merge, not deferred to CI: full reactor test-compile clean, core unit suite 338 tests 0 failures under -Pci (up 10, from #260's new KafkaTestUtils and ParallelEoSStreamProcessor tests), and all four repo gates pass. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…me it #263 is merged and GitHub retargeted this PR onto master, so the top of the note was false where it mattered most: it still said #264 was stacked, that #263 must merge first, and that `Check PR Dependencies` fails by design. That check passes. `docs/inflight/AGENTS.md` says what to do with the rest: when something closes, do not rewrite it into a FIXED/DONE narrative - shrink the file to the open follow-ups and rename it. The "Settled" section was exactly that wrong move, and it duplicated the commit messages. 80 lines to 38, carrying the two things that are actually open: the undecided LoadTest auto-run question, and the audit's known drift. Every surviving claim re-checked against the tree and the API. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…un, assert or exist now do (#264) Acts on the audit that landed in #263: every test that did not run, did not assert, or was never written. OffsetEncodingTests is the substantive change. Five OffsetEncoding values that `assumeWorkingCodec` branched around now assert the degraded contract - work is repeated, nothing is lost - instead of reporting green while skipping the assertions. Four dead tests are deleted: a stub whose body was `assertThat(true).isFalse()` behind @disabled, an empty `{}` body, a commented-out @test over an infinite Stream.generate, and a non-asserting diagnostic (plus the utility it orphaned). Manual procedures the audit found deleted as "dead code" are recovered as runnable knobs, gating values unchanged: LoadTest's 40k/80k/400k ladder, the TransactionAndCommitModeTest concurrency ladder, VeryLargeMessageVolumeTest's 2M aspiration, and MultiInstanceHighVolumeTest's 10M rung - the last of which needed its hard-coded 60s wait derived from the volume before the rung was reachable at all. JavaEnvTest's environment dump is automated into AmbientProbeExtension's failure autopsy rather than restored as a test that asserts nothing, and the Vert.x gap a deleted stub was named for is now characterized: a 5xx is a delivered response, so the offset commits, while a transport failure does not. Three core tests are finished rather than deferred. `processInKeyOrder` and `offsetsAreNeverCommittedForMessagesStillInFlightLong` had been @disabled since 2020 and failed 100% deterministically - the library was right and the tests were wrong, on two counts: a committed offset is exclusive (finishing records 0-2 commits 3, not 2), and partition 1's base offset is 4 because record creation uses a global counter. `userSucceedsButProduceToBrokerFails` is new, covering a produce-failure path that was reachable and untested. Both re-enabled tests assert the committed-offset FRONTIER - the highest offset per partition - rather than the exact commit history, because the exact form asserts where the wall-clock commit tick fell. Measured: it failed 3 of 10 runs, as [1, 3] where [3] was expected and [3, 4, 5, 6] where [3, 4, 6] was, in two different commit modes. Both are correct PC behaviour. #260 fixed this same class of defect for repeat commits days earlier; the general rule was written only in the javadoc of the helpers implementing the narrow case, so it did not reach the next test. It is now written down in docs/solutions/test-flakiness/assert-the-commit-frontier-not-the-tick-path.md. The new 40,000-message LoadTest case runs automatically in the required Performance Tests leg. That was measured before being left automatic, per the AGENTS.md rule to separate contention from a concurrency bug: 5/5 green on an uncontended broker at ~52s against a derived 600s ceiling, and green on the real lane since. What the measurement could not clear is recorded at the site and in docs/inflight/test-required-perf-lane-scope.md. Also corrected, each because it would have sent the next reader wrong: - Two documented knob invocations selected ZERO tests and exited BUILD SUCCESS - both classes are @tag("performance"), the default excluded.groups contains performance, and exclusion beats inclusion. Both now use bin/performance-test.sh. - The failsafe comment cited a performance.yml workflow on dedicated hardware. Neither exists; the lane is maven.yml's required leg on ubuntu-latest. - AmbientProbeExtension's autopsy dumps every system property to CI logs; values under credential-looking keys, and credentials embedded inside values, are now masked. Masked by key name rather than an allowlist, which would silently drop the next knob somebody adds. - PartitionStateManager's javadoc claimed to truncate offsets on commit. It does not. - docs/todo-index.md carried a marker count - a derived number stored beside the data it is derived from. Removed from the generator, not just the file. Three landed plan documents are deleted (~1,525 lines) per the AGENTS.md rule that a plan goes stale once its work lands; the measurements that outlived them were salvaged first rather than deleted with them. Verification: core unit suite 347 tests, 0 failures, 8 skipped under -Pci; the previously-flaky class 12/12 clean on full-class runs; all five repo gates pass. 19 review threads resolved, and a seven-reviewer pass found four P1s - two reproduced by running the tests rather than reading them - all fixed here. Deferred deliberately, with the reasoning recorded rather than lost: docs/inflight/test-inactive-test-review-followups.md (chiefly hoisting awaitFrontier into the shared test base, and giving its negative checkpoints a hold rather than a sample) and docs/inflight/test-required-perf-lane-scope.md. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…d-waits Brings in #264 (now merged), which was this PR's declared dependency, plus #277, #260 and #263. One conflict, in docs/todo-index.md - a generated file. Master's bin/todo-index.sh changed the header format (the marker count moved out of the prose) while this branch still carried the old shape, so both sides had edited the same generated line. Resolved by regenerating from the merged tree rather than picking a side: 84 markers, and `bin/todo-index.sh --check` reports it current. Master did not touch BlockedThreadAsserter or junit-platform.properties in these four commits, so this PR's two central changes had no semantic collision to resolve. Verified after merging: junit-platform.properties still deleted, core's pom still carries the surefire/failsafe configurationParameters, and the asserter rewrite plus its negative-control test are intact. Local verification: core 354 tests and vertx 18, 0 failures, 0 errors. bin/check-quarantine-registry.sh and bin/check-copyright-headers.sh both clean.
…etire the two notes this PR overrode A ce-doc-review pass over this branch's three commits falsified six factual claims they made and found two inflight notes the PR had silently overridden. Each is a claim that read as verified and was not, which is the class this repo's own rules call a false green: none of them would have gone red anywhere. 1. The `@Disabled` grep claim was wrong in two places - the new test-hardening entry and release-0.6.0.0.md's gate section both said `grep -rn "@disabled" --include="*.java" .` returns nothing but historical prose. Running it returns a live `@DisabledOnOs(OS.WINDOWS)` on AbstractQuarantineScriptTest and two `@Disabled` string literals built into TransactionalClaimCoverageTest's assertion messages, on top of the prose. Both sentences now describe what the command actually returns and why the gate still holds - no live bare `@Disabled` on a test class or method - and name the command to re-run rather than a count that will drift. 2. The new ChaosRevokeUnderWorkTransactionalIT sighting said the PR "touches no Java" in the same sentence that named the Java file it deletes. Replaced with ground the deletion does not undermine: nothing under `src/main` and nothing in the `chaostests` package changed, and `.github/workflows/maven.yml` gives each chaos shard a hardcoded `scenarios:` list (verified - Suite 2/4 is ChaosRevokeUnderWorkTransactionalIT,ChaosRevokeUnderWorkKeyOrderIT, passed as CHAOS_SCENARIOS), so removing a sanity-package class cannot reshuffle a shard. The opening parenthetical's artefact counts become the shape plus `git diff --name-status <base>..<head>`, and the bare `AbstractRevokeUnderWorkScenario.java:211` becomes the seed banner's own greppable literal, matching how its sibling citation two lines up already cites. 3. The blockers recheck said "every occurrence is already conditional" of module-maturity.yaml's production-use wording. Half true: the `support_posture` lines are conditional, the `maturity: production-use` field values are bare. The record now says exactly that, so the second recheck starts from the state rather than from the first pass's summary of it. No yaml value changed - whether an unqualified maturity value is a claim the open confluentinc#857 family falsifies is a release call for the maintainer. 4. release-0.6.0.0.md's opening block said #80 emptied the quarantine registry so release.yml's gate now passes, while the section this PR rewrote records MultiInstanceRebalanceTest.largeNumberOfInstances as quarantined and blocking. docs/quarantined-tests.md confirms the entry is live and unowned, so rule 5 still bites; the opening block now says the gate does NOT pass and points at the registry as the enforced copy. 5. Retired docs/inflight/test-progressbar-width-needs-a-machine-assertion.md. It argued for splitting the test into a machine assertion plus a tagged demo rather than deleting it, and this PR deleted it without naming that argument. Per AGENTS.md - record the reasoning you are overriding, where you override it - the note's alternative and why deletion won (the test asserted nothing, so the split describes a test still to be written rather than one being preserved; ProgressBarUtils.getNewMessagesBar keeps its other in-tree callers; and reinstating a tagged demo stays a maintainer's call this deletion does not block) are migrated into the test-hardening entry before the note is removed. 6. Retired docs/inflight/test-disabled-tests-before-v6.md. Its stated delete-when condition is met - #263 landed and no test carries a live bare `@Disabled` - and its four-name list is the same stale one this PR already replaced. Its one framing not owned elsewhere, that the disabled tests were inherited debt rather than a rule broken here, is migrated to release-0.6.0.0.md's gate section, which points at the 2026-08-08 audit for the per-test provenance behind it. No inbound links to fix - `grep -rn` over the tree returned none for either note. 7. docs/refactoring.md described ProgressBarTest.width in the present tense as a deliberate manual check. Past tense now, citing the deletion entry. 8. Reconciled "three of four races refound" (the new testing-evidence entry) with "Lincheck refound four real races unaided" (release-0.6.0.0.md, untouched by the earlier commits). docs/plans/2026-08-25-001-test-lincheck-poc-plan.md's verdict table supports the first: three of four found by the stress strategy plus one nobody had named, with the fourth half-found once by model checking and not reproducible. The release note now says that, and says to use that wording in the announcement. Left alone deliberately, because they are decisions rather than errors: the lincheck entry's merge_gate field versus the CI leg, and the jcstress anomaly disposition. Gates: bin/check-all.sh (16 passed, 0 failed), plus check-issue-refs.sh, check-docs-data.sh and check-file-refs.sh individually. Docs only - no code, no test behaviour, no changelog entry. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019SVm2cT6ZgUyPMim9CukYK
Description
grep @Disabledreturns 7. Five of those are real disabled tests, and four of the five record no reason at all. Nothing on master said so, and the last person to write this down put it on a branch nobody could find.This lands a per-test audit of every test that does not run, does not assert, or was never written - with each finding traced to evidence rather than suspicion - and fixes the two ledger entries that buried its predecessor.
The answers, up front:
@Quarantined; one is@DisabledOnOs(OS.WINDOWS)on an abstract harness that skips nothing on any machine we build on.ProgressBarTest.widthcarries"For reference sanity only". The other four have no annotation message and no comment.Two categories nobody asked about, because an annotation grep cannot see them and they are the ones that actually mislead: 15 of 292 tests assert nothing, and
OffsetEncodingTestsreports green with most of its assertions branched away by a helper namedassumeWorkingCodecthat is not an assumption.What the evidence turned up
The two long-dark core tests were both disabled by one commit -
c1fefbc64"Create and commit offset map", 2020-08-27 - which rewrote commit-assertion semantics across that whole file.git blamegives the wrong answer here (a 2021 reformat moved the lines); it tookgit log -Sto find. An abandoned branch,bugs/turn-on-commit-tests, names the cause outright: "Turn back on offset commit tests which were dibbled when the offset map feature was added". Somebody already tried the cheap un-disable and stopped at WIP.The real coverage gap is narrower than "key ordering is untested" and worse than it looks: nothing asserts, end-to-end and per-
CommitMode, that offset commits respect key-order blocking across partitions.Why the ledger changes are part of this
A 455-line audit of adjacent scope has existed since 2026-04-22 and has never been read. It lives only on
refactor/test-hardening, has no PR anywhere, and is referenced in exactly two places that do not mention it -docs/refactoring.mddescribes that branch by its other commit, anddocs/inflight/branch-stale-and-diagnostic.mdfiles it under Superseded, on the safe-to-delete list.It did not rot because its numbers drifted. It rotted because it was filed under a description that did not describe it, on a branch queued for deletion. This absorbs its contents, corrects the two reasons its own git history refutes, and fixes both entries.
On this document's own accuracy
The audit was fact-checked adversarially and eight of its numbers were wrong, including one that reported an open defect as fixed. All corrected in a follow-up commit here, with the cause named in the commit message: counting raw grep hits instead of enumerating what matched - the exact error the document opens by warning about. The reproduction commands now enumerate annotation shapes rather than trusting a hard-coded whitelist.
Deliberately not built
A generated
docs/INACTIVE_TESTS.mdwith a--checkstaleness gate, in the shape ofbin/todo-index.sh. It matches repo convention, but it does not address why the predecessor was lost, and its gate would fail the PR Checklist job on any open PR touching a test annotation. Recorded as follow-up.No test behaviour, assertion, timeout or volume changes in this PR. Records only.
Checklist
N/AN/A- documentation only; this PR deliberately changes no testN/A- touches no CI runners or workflows