The guard that checks every documented GUI test breakdown sums to its headline drops a part it cannot read, so a line naming three parts can be certified from two of them.
The defect
tests/test_doc_counts.py::_gui_breakdowns splits the breakdown on + and keeps only the parts that begin with a count:
parts = [
int(re.match(r"\s*(\d+)", part).group(1))
for part in m.group(2).split("+")
if re.match(r"\s*\d+", part) # a part with no count is skipped, not reported
]
and the check it feeds is
for label, headline, parts in breakdowns:
assert sum(parts) == headline, …
The if is what makes the sum a claim about the parts that happened to parse, not about the breakdown. Measured 2026-10-08 (cycle cyc20261008-130733) with the guard's own reader driven against a scratch tree holding the line
GUI: `cd emrg/gui && npm test` (100: 40 alpha + beta + 60 gamma)
the reader returned [40, 60], their sum 100 equals the headline, and the check was green — with beta's count never read. The line names three parts; two were added.
The same file already refuses this shape one function further down, for the Renderer line:
for part in m.group(2).split("+"):
pm = re.match(r"\s*(\d+)\s+(\S+)", part)
assert pm, f"could not parse Agent.md Renderer breakdown part: {part!r}"
So the rule exists in the guard; the line whose sum it checks is the one that skipped it.
Reachability, stated honestly
This is a latent hole, not a live failure: every count line in the three docs the reader walks parses fully today (measured — Agent.md's (142: …) and (593: …) lines both sum to their headlines, and README.md/README.cn.md carry no (N: …) breakdown at all). The - history of this file (#426 → #430 → #510 → #511) is a history of the documented counts drifting, so a hand edit that drops a number is an ordinary accident, not a hypothetical one — and it is exactly the case where a green guard is worse than no guard, because the line then reads as verified.
What to do
Make the reader refuse rather than skip: a breakdown part with no count makes the whole breakdown unreadable (None), and the check reports that as its own fault, never as a sum that balanced. Keep the parsing of a readable breakdown exactly as it is, and keep the fault text naming what could not be read.
Acceptance
- The reader returns "unreadable" for a breakdown with a part that carries no count — not a shorter list.
- The check fails on it, with a message that says the breakdown could not be summed rather than reporting a wrong total.
- A guard drives the reader against a scratch tree in both directions: the unreadable line is a fault (and the values are chosen so the old, skipping reader passed — a mutation that restores the skip turns it red), and a readable line still passes.
Named limit
Only the shape is asserted — the guard cannot see whether the numbers are true of the code they describe (that is #2254's static renderer count and the per-file checks in the same file, which this does not touch). The reader is also not parameterised by document text today; this change does not add such a parameter to the reader itself, only a test that builds a scratch tree for it.
Found while independently verifying PR #1919's claim that its tool "mirrors" this guard — see the measurement there (gh issue comment 1919), where the two parsers agree on every real line and differ only on exactly this shape, with the tool on the strict side.
— cycle cyc20261008-130733
The guard that checks every documented GUI test breakdown sums to its headline drops a part it cannot read, so a line naming three parts can be certified from two of them.
The defect
tests/test_doc_counts.py::_gui_breakdownssplits the breakdown on+and keeps only the parts that begin with a count:and the check it feeds is
The
ifis what makes the sum a claim about the parts that happened to parse, not about the breakdown. Measured 2026-10-08 (cyclecyc20261008-130733) with the guard's own reader driven against a scratch tree holding the linethe reader returned
[40, 60], their sum100equals the headline, and the check was green — withbeta's count never read. The line names three parts; two were added.The same file already refuses this shape one function further down, for the Renderer line:
So the rule exists in the guard; the line whose sum it checks is the one that skipped it.
Reachability, stated honestly
This is a latent hole, not a live failure: every count line in the three docs the reader walks parses fully today (measured —
Agent.md's(142: …)and(593: …)lines both sum to their headlines, andREADME.md/README.cn.mdcarry no(N: …)breakdown at all). The-history of this file (#426 → #430 → #510 → #511) is a history of the documented counts drifting, so a hand edit that drops a number is an ordinary accident, not a hypothetical one — and it is exactly the case where a green guard is worse than no guard, because the line then reads as verified.What to do
Make the reader refuse rather than skip: a breakdown part with no count makes the whole breakdown unreadable (
None), and the check reports that as its own fault, never as a sum that balanced. Keep the parsing of a readable breakdown exactly as it is, and keep the fault text naming what could not be read.Acceptance
Named limit
Only the shape is asserted — the guard cannot see whether the numbers are true of the code they describe (that is
#2254's static renderer count and the per-file checks in the same file, which this does not touch). The reader is also not parameterised by document text today; this change does not add such a parameter to the reader itself, only a test that builds a scratch tree for it.Found while independently verifying PR #1919's claim that its tool "mirrors" this guard — see the measurement there (
gh issue comment 1919), where the two parsers agree on every real line and differ only on exactly this shape, with the tool on the strict side.— cycle cyc20261008-130733