Repository navigation
feat(inflight): rank the open backlog across every ref, and diff it against the standing ranking - #438
Conversation
…t it refuses to say The design for `bin/inflight.mjs rank` - a read-only subcommand that reads open in-flight notes from every ref, groups them the way the session index groups them, and reports where that picture disagrees with `process-candidate-ranking.md`. The working note lands with this first commit rather than the last, per docs/inflight/AGENTS.md, and it carries the parts a later reader would otherwise take for an oversight: - The gather step is already solved. `docs list inflight <impact>` reads every ref, keeps open notes, groups them by impact and marks off-baseline ones with their branch. `rank` reuses `corpusIndex`, `classifyNote` and `inflightGroupOf` rather than rebuilding any of it, and adds the carrying branch's pull request, the live-versus-archival split, the filename number, the register delta, and an accounting of the open notes no impact bucket claimed. - It never says a branch fixes a note. A note travels on the branch that produced it, which makes carriage cheap and ownership unavailable; conflating them is the expensive wrong answer this command exists to stop. - A third `candidate` relation - a branch whose name encodes a number matching the note's - was designed, measured and dropped rather than shipped behind a caveat. Some matches are cross-namespace by construction, because a note filename carries a fork number while branch names here encode upstream ones. Deferred, not rejected: a pull request body naming the note path would be a signal that earns it. - The open filter follows `classifyNote` rather than a list of state words, because `open` is set by the presence of any state marker. A note declaring `inflight-state: open - <reason>` is not open by that rule, and one is on the baseline today. Planning round found seven further problems that shaped the requirements, the largest being that keeping only the impact buckets would silently drop the corpus's largest group of open notes.
…gainst the standing ranking `bin/inflight.mjs rank` reads open in-flight notes from every ref, groups them the way the session index groups them, annotates each row with the relations it can prove, and reports where that picture disagrees with `docs/inflight/process-candidate-ranking.md`. The delta is the deliverable: it turns the next ranking pass from "read every note" into "look at these disagreements". The gather step is not rebuilt. `docs list inflight <impact>` already walks every ref, keeps open notes and groups them by impact, so `corpusIndex`, `classifyNote` and `inflightGroupOf` are imported rather than re-derived - the group rule in particular has to stay one owner, or two surfaces would place the same note differently and nothing would report the disagreement. WHAT IT REFUSES TO SAY. A note travels on the branch that produced it, so carriage is cheap to know and ownership is not available at all. Rows say CARRIES. The worked case is a data-loss note carried by one branch whose own text says the bug predates that branch's pull request; a row reading "fixed by" it would be confidently wrong. Three defects were found by running it against this repository, each now with a negative control: - A note's versions can DISAGREE. The data-loss note above is open on the branch that owns the bug and closed on a branch that fixed something adjacent, and reading the first sorted live ref dropped it from the backlog entirely. It is now placed by a version that is still open work, and the row names the refs that disagree - the axis `ci-inflight-next-commands.md` states as "flow with git, do not suppress it": one status is a summary that has thrown away the disagreement. - A number can resolve to SEVERAL notes. Filenames get recycled and renamed here, so a note and its own dead predecessor carry the same positional number; a map keeping the last writer reported the register's live entry as stale while naming the dead copy. The entry is now satisfied when any candidate is open work, and a genuine collision names every candidate. - The pull-request snapshot is keyed on the bare branch name, so looking it up by the full ref returned null for every `origin/*` branch - the whole corpus reading as pull-request-less while the snapshot said it had answered. Other decisions the tests pin: openness follows the marker's PRESENCE rather than the words inside it, so a note declaring `inflight-state: open - <reason>` is not open and one such note is on the baseline; open notes no impact bucket claimed are emitted rather than dropped, because that is the largest group in the corpus and the register names notes in it; a filename number is never attributed to a repository, since pre-convention names carry confluentinc numbers and `pr-` carries a pull request; and the bulk pull-request snapshot is used alone, because the per-branch fall-through `branchView` uses for one branch would be one untimed `gh` subprocess per pull-request-less branch here, with absences deliberately uncached. A register that could not be read is a failed run reported after everything that did run - exit 2, the shape `refactor-window` already uses - because "the delta was empty" and "the delta never ran" are different answers.
…ys it refuses docs/inflight-tool.md gets the section help text cannot carry - what the answer looks like on a real question, and why the working-tree version of it is wrong. The worked example is the carriage-is-not-ownership case, because that is what the annotation step buys: a pull request sitting on the branch that carries a note, which does not fix the bug the note describes. It also states the two rules a reader would otherwise have to infer from output: a filename's number is printed without a repository, since pre-convention names carry confluentinc numbers and `pr-` carries a pull request; and a note whose versions disagree is reported as disagreeing rather than resolved to whichever ref sorted first. docs/features/cross-ref-repository-queries.yaml already records what this tooling declines to do, so the ownership refusal and the disagreement-over-resolution rule go there beside the others rather than only in the command's own help.
…itory, and four vacuous controls Local simplify and code-review passes over the new command. The findings that mattered were not style; three were defects that only surfaced by running the thing against the real corpus. CRASHES - `numberFor(path, text)` read `text` from outside its scope, so the command threw `ReferenceError: text is not defined` on any note carrying a number - which the corpus has plenty of - and exited 1, neither of the tool's documented codes. The deciding version now carries its own text forward. - A note closed on every LIVE ref but still open on a preserved tag made the chosen version's refs and the path's live refs disjoint sets, so the read ref resolved to nothing and the row threw on `readRef.replace(...)`. Reproduced against a purpose-built corpus before the fix. The read ref now comes from the chosen version's own refs. A REFERENCE THAT RESOLVES TO THE WRONG THING `numberFor` printed `gh issue view N -R astubbs/parallel-consumer` for every non-`pr-` number, while `docs/inflight/AGENTS.md` names the pre-convention notes whose number is confluentinc's - `bug-857-family.md` is on the baseline and its own title reads "The confluentinc#857 family". Its own docstring claimed it named no repository. Attribution now comes from the note's own text: a qualified mention of its own number decides it, and a note that says neither gets BOTH lookups and is labelled unattributable. AGENTS.md's rule is that a wrong reference which resolves is worse than a broken one. A NOTE THAT COULD NOT BE READ `blobContents` can return ok overall while one blob comes back `missing` - a partial clone, a gc race - even though the ref listing just named it. That note was dropped with no accounting, which is the failure-rendering-as-an-empty-result shape this file is written against. Unreadable paths are now named and the run reports that the answer is incomplete. FOUR CONTROLS THAT ASSERTED NOTHING Each was found by its mutant staying green, which is what the negative-control arm is for: - Two mutation anchors patched an EARLIER identical occurrence in the file, leaving the code under test untouched - `patch` replaces the first match, and `registerBlob` carries the same guard shape as `rank`. - One assertion read a field that is only populated when a group scopes the call, so it was vacuously true on the unscoped call it was making. - One anchor stopped being load-bearing when the crash fix moved the lookup it pointed at. DUPLICATION REMOVED `INFLIGHT_GROUPS` is exported from its owner rather than re-authored here (four of its five entries had already drifted), and `refsText`/`scopeLine` moved into `views.mjs` so the tool's core disclaimer - what was searched, and that it was not the working tree - has one owner rather than two copies. Also: dead result fields no caller read, a duplicated `refKind` pass, a shadowed variable, a dead conditional, a repeated `--impact` silently taking the first, and the version-selection logic extracted and named.
…he PR The working note gains the three findings whose value survives this branch: the filename number is attributed by the note's own text rather than by the convention (a pre-convention name on the baseline carries a confluentinc number, and the fork-qualified lookup for it resolves to the wrong issue); the two crashes a real corpus produced and a fixture did not; and the note that ls-tree named while cat-file returned missing, which was being dropped with no accounting. The lesson worth keeping is the second crash's shape: this command has an ok/reason contract and an exit-code contract, and neither protects against a ReferenceError. Nothing asserts the front door cannot throw, so the guard is the self-test running the real command end to end.
…or ever
Every negative control copies bin/ plus .claude/hooks into a temp directory to mutate, and nothing
ever removed one. A run therefore leaked one copy per check permanently, and the leak grew with the
suite - so adding checks made it worse.
Measured on the machine that found it: 20 GB of inflight-selftest-* directories from accumulated
runs, on a disk the repo's own pre-commit hook was already warning was 94% full. Reproduce with
du -sh "${TMPDIR}"/inflight-selftest-* | tail -1, and count a single run's contribution by listing
that glob before and after.
The cleanup is a finally rather than a line at the end of the loop body, because the
could-not-be-built path takes a continue and leaked too. Verified by running the suite and counting
the glob either side: unchanged, where it previously grew by the number of checks.
Found while finishing another change in this file rather than by a gate - nothing measures the
suite's own footprint, and a temp directory nobody lists is invisible until the disk fills.
…ays so Three findings from the agent-native review, all about whether the output teaches an agent what to do next - which is the whole reason this front door exists. - **Every level prints the next level's command**, and this level did not. `docs list inflight <impact>` prints `docs show <path>` beside each row; `rank` printed the path and left the reader to know that command exists and retype it. Rows now carry it. - **A scoped group with no rows printed nothing at all**, which is indistinguishable from a section that was dropped - the silence this command is organised against. It now says the group is empty and that this is a result about every ref, not about the checkout. - **The exclusion counts are whole-corpus and did not say so.** They are identical whether the call is scoped or not, so sitting them unqualified among scope-limited lines read as though they described the group. Also closes a coverage gap the same review named: nothing asserted on `formatRank`'s rendered text at all - every existing check drove the `rank()` data function - so none of the three defects above could have been caught by the suite. One check now renders both a populated scope and an empty one.
…ility defects hid The note gains the lesson from the agent-native review: this command's own help text says every level of the front door prints the next level's commands, and its rows did not - it printed the path and left the reader to know docs show exists. The reusable half is why the suite could not have caught it. Every check drove the rank() data function; nothing rendered a view, so formatRank's contract was unenforced however carefully the query underneath it was tested. A view no check renders is worth naming for whoever adds the next command here.
…ead prose as a ranking The correctness review found the deliverable itself was wrong on the real register, both ways at once. `parseRegister` scanned the whole document for `astubbs#<n>` and for any hyphenated `.md` token, so nothing distinguished a ranking from a sentence about one. OVER-REPORTING. Every number the delta reported as "resolves to no note on any ref" was a false positive. Most were prose - continuation lines, cross-references, and the register's own "What is NOT on this list" paragraph. Two were real entries whose notes the SAME RUN had already resolved by filename, because the two halves of an entry were asked independently. A citation of a doc outside `docs/inflight/` became a phantom entry reported as `absent` - an instruction to delete a live cross-reference, indistinguishable from the real case that row exists for. UNDER-REPORTING. A live note was suppressed from the unranked half on the strength of the sentence "fixing #177 does not close it". Prose about a note, read as a ranking of it. So an entry is now a LIST ITEM INCLUDING ITS CONTINUATION LINES, and the citations on it belong to it. That last part matters: this register routinely opens an item with the number and names the note two lines down, so a line-scoped parse split both of the entries it got wrong. An entry is satisfied when EITHER half resolves to open work, which is what stops a number being called unresolvable while the filename beside it resolves - and collapses the two halves into one finding per entry rather than two. Filenames count only when bare or under `docs/inflight/`, and both spellings of the number are read, since AGENTS.md mandates the qualified form and the register already uses it. Three more from the same review: - **`docs/inflight/CLAUDE.md` was ranked as open work.** This walked the corpus index raw without the membership guard `docsShape` applies, so the two surfaces disagreed about what the corpus even contains - the failure this module's header claims importing `inflightGroupOf` prevents. The guard is now imported rather than restated. - **A register the parse could not read rendered as a clean delta.** Zero recognised entries printed the all-clear and then listed the whole corpus as unranked: the found-nothing / could-not-look collapse, moved from the git walk to the parse, at exit 0. The output now says how many entries it recognised and, at zero, that the delta is the parse not reaching the register. - **`byNumber` ignored the attribution `numberFor` had just computed**, so a fork-qualified entry could be satisfied by a note whose own text says its number is confluentinc's. And from the testing review, the coverage that let all of this hide: nothing asserted on `disagreement` - the field this command exists for - nor on the CLI's unknown-option, missing-value or unknown-group behaviour. Those now have controls, alongside ones for the prose parse, the continuation-line entry, and the zero-entry register.
…scusses the thing The register-parse defect is the most reusable thing this branch learned, so the note carries it rather than only the commit that fixed it. These documents are written for people. A scan for the vocabulary finds it in the sentences ABOUT a ranking as readily as in the ranking - the register's own "What is NOT on this list" paragraph read as entries, and "fixing #177 does not close it" marked that note as ranked and hid it. Anchor to the structure, not the vocabulary: the list item, its continuation lines, the marker. And say in the output how much of the document the parse recognised, because a parse that reached nothing renders identically to a document everything agrees with.
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
✅ Duplicate Code ReportTwo engines run in parallel for cross-validation. Each has its own thresholds tuned to its baseline - the real safety net is the per-engine "max increase vs base" check. ✅ PMD CPD
No new clones introduced by this PR. ✅ jscpd (language-agnostic)
No new clones introduced by this PR. Powered by astubbs/duplicate-code-cross-check |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #438 +/- ##
============================================
+ Coverage 81.84% 82.43% +0.59%
- Complexity 1462 1471 +9
============================================
Files 95 95
Lines 5133 5142 +9
Branches 500 501 +1
============================================
+ Hits 4201 4239 +38
+ Misses 732 707 -25
+ Partials 200 196 -4
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
[superseded - a quarantined test changed outcome] 🧪🔒 Quarantine Lane Report
🔴 expected while the owner PR is open · 🟡🎲 flapper, pass proves nothing · 🚨 a deterministic quarantined test passing means its fix landed: delete its Updated for Superseded by a newer quarantine lane report. |
🟢 Throughput — OKThis branch measured about 23% slower than master, on the one test this measures. That is larger than this test's own run-to-run spread of about 17%, so it is worth looking at.
Allowable range 🟢 ≥ 0.70 · 🟡 0.50–0.70 (about a 30% loss) · 🔴 < 0.50 (about a 50% loss) What the numbers mean, and what they cannot tell youThe one that gets misread. Why a shape and not a rate. A rate depends on which runner you drew. A shape does not: every test here processes a fixed number of records, so a runner twice as slow doubles the subject and the controls together and leaves their ratio alone. That is the whole trick, and it is why the reported rate is shown last and labelled as this machine only. Reading the comparison. By conservation, not by correction. Every test in this lane processes a fixed number of records, so within one run the ratio of one test's time to another's is invariant under machine speed — a runner twice as slow doubles both terms and leaves the ratio alone. There is no machine-index correction to be wrong, because nothing needed correcting. Per-method times, not class times. A class time is Reference is the median of 10 recent What this still cannot do. It removes machine-to-machine variance. It does not remove this test's own run-to-run variance, measured at about 30% on a single unchanged commit while its controls stayed within 5%. That is a property of the test, not of the comparison, and no arithmetic here can touch it — which is why the reference is a median and the bounds are deliberately coarse. 🟡 means look at this; only 🔴 is outside the measured spread. Runs used: 9999144, cc36b64, 1941cdf, 05c02bb, 5706686, c668acb, fb5ea93, 6573781, c30aaee, ca9c21b Since the previous push: ratio 0.973 -> 0.77, share 1.654 -> 2.101, rate 65316 -> 64712 (-0.9%). One push of difference sits inside this test's measured spread - read it as movement, not as a result. Updated for |
…the note was open The adversarial review found the command committing the exact failure it exists to prevent, one layer up from the version-selection bug already fixed on this branch - and worse, because it does not merely omit a row, it prescribes the wrong disposition to the register on the baseline's word. `chooseVersion` preferred the baseline's version unconditionally. So a note the baseline calls deferred while a live branch carries it OPEN was dropped from its group entirely, and the delta told the register it was deferred, with no ref named anywhere. Reproduced on this corpus before the fix: `core-auto-scaling.md` is `deferred` on the baseline and carries `inflight-impact: throughput` with no state marker on three live refs, and it is one of the register's own ready picks. A still-open version now beats the baseline's, and the baseline only breaks ties among those. The row then names the ref it actually read, says the baseline's copy is not open work, and emits the disagreement rather than suppressing it. Three more from the same review, each re-verified before fixing: - **An entry could be satisfied by a note no live ref carries.** `isOpenWork` asked the group and never reachability, so the delta could report nothing it ranks has stopped being open work a few lines above its own row saying that note survives only on a tag. - **`recognised` was a numerator with no denominator.** "11 entries recognised" reads as a complete reading of the register and is not - the ready-picks half cites bare and upstream numbers this parse does not read. It now states the denominator and what the unread items look like. This had already cost something: the bare count was quoted to a human as a complete reading. - **The group partition was hand-maintained beside a taxonomy another module owns.** A group in neither set made `buckets.get(key)` undefined and `.push` throw, and with no top-level try/catch the process exited 1 - neither documented code, so a caller testing for 2 reads it as a successful run. The complement is now derived, with a guard for anything outside the taxonomy entirely. Also: the disagreement line named tags as though they were branches, breaking the rule stated three lines above the code printing it; `index.unreadableRefs` was computed by the corpus index and read by nobody, so a ref whose listing failed was a could-not-look reported as a found-nothing; and the delta headline printed "11 entrys". One claim did not reproduce and is not fixed: an unknown `--impact` group does not crash. The CLI guard and its check already cover it - `rank --impact nope` exits 0 and names the valid groups - and the library does not throw either. The crash the finding describes is the taxonomy-partition one above, which is real and is fixed. The control for the openness rule had to move twice. Its first anchor collided with an identical guard earlier in the file; its second was absorbed by the new `!buckets.has(key)` belt, which made the mutant green while asserting nothing - the same fallback-defeats-the-control shape as the `?? chosen.refs[0]` removal earlier on this branch. It now anchors on what actually widens the rankable set, with a comment saying why, so the next reader does not simplify it back.
… will be wrong The command committed its own cardinal error twice, one layer apart, and the second was found only after the first was fixed and reviewed. Both were preference rules that looked like sensible defaults and quietly outranked the evidence: first-sorted-live-ref, then the baseline preference. The second is worse than a missing row. A missing row is silence; the delta actively prescribed deferred to the register on the baseline's word while a live branch carried the note open. The lesson for whatever reads this corpus next is in the note: a default that picks one version is not neutral, and it is wrong exactly when the versions disagree - which is the only case anyone is asking about.
[superseded - a quarantined test changed outcome] 🧪🔒 Quarantine Lane Report
🔴 expected while the owner PR is open · 🟡🎲 flapper, pass proves nothing · 🚨 a deterministic quarantined test passing means its fix landed: delete its No quarantined test changed outcome since the previous push. Updated for Superseded by a newer quarantine lane report. |
|
@codex review Steer, from this PR's own author and from what earlier review passes actually caught: Push hardest on the version-disagreement path in
Three finds of the same class after review is not converging evidence of correctness - it is a signal the case is hard. It is also the only case anyone runs this command for. Two named residuals, unfixed and stated in the description: the register parse assumes a continuation line is an indented line, which fits this register and is enforced by nothing; and Worth knowing about the test suite: mutants are the controls here, and one control was already found to be vacuous because a later guard absorbed its mutation and left it green while asserting nothing. If a check's mutant dies by crashing rather than by the assertion it claims to test, that check is not proving what it says. This is a cross-model pass specifically - every review so far has been in-process, and the two that found the most landed last. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 79407c9653
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
The rank branch note carries three general lessons, none of them about rank. A prior-art run over every ref (`inflight prior-art`, contents not filenames) returned nothing from docs/solutions/ for any of them, so as written they die when that note is trimmed at merge - which is exactly the "not a place knowledge goes to die" case docs/inflight/AGENTS.md names. Two write-ups rather than three. The version-selection and the parse-a-human- document lessons are one lesson with two worked cases: a source that is not a database, a reader that resolved the ambiguity silently, and a fix with the same two halves both times - anchor to what actually decides, then report what could not be resolved. Splitting them would have produced two documents each missing the half that makes the other memorable. best-practices/a-source-that-can-disagree-with-itself-needs-a-reader-that-says-so.md The three version-selection defects of one class, each found after the previous was fixed and reviewed, with the placeholder-beats-impact one that a cross-model review found last; and the register scan that read its own "What is NOT on this list" paragraph as rankings while a sentence about a note suppressed it. Sibling to a-guard-that-greps-java-must-read-what-javac-decided.md, which owns the case where a parser exists - a markdown register has none, so the document's own structure is all there is to anchor to. test-issues/a-view-no-check-renders-has-an-unenforced-contract.md Three legibility defects that hid because every check drove the query layer and none rendered the view - and the sharper second case, a control whose mutant died on a data field the renderer never reads, so a view rewritten to say "fixed by" would have left it green. Coverage that appeared to exist. Both are cross-referenced from the neighbours they sit beside, and cite the worked cases rather than stating the lesson abstractly.
…the two controls that were proving nothing The Codex pass steered at the version-disagreement path returned eleven findings. Each was verified before it was fixed; none was refuted, and two were confirmed against the live corpus with a probe that predicted the magnitude before the fix ran. Three were of the class the steer named, which is now four separate defects in `chooseVersion` and what reads it - the case is hard, and it is where the next reviewer should still push hardest. THE VERSION-DISAGREEMENT PATH, AGAIN An impact-bearing version now beats a placeholder, and the baseline only breaks ties inside that pool. `feature` and `unmatched` are open work and belong in RANKED_GROUPS, so a baseline copy carrying no impact tag TIED with a live copy carrying one and the baseline tie-breaker took the placeholder. Two silent costs: the row landed in the catch-all rather than its impact bucket, and `delta` accepts only the impact scale, so a register entry naming it read as stale while a live ref carried it as ranked work. Prediction before the fix: 85 notes on this corpus are in that state and the unmatched bucket should fall by exactly that. It went 202 -> 117. Tagging a note on the branch that works it is the ordinary flow. The delta is now told whether the CHOSEN version is reachable, not whether the path is. Every live copy closed and the open one surviving on a tag made the path-level answer true, so the delta could print the all-clear a few lines above its own row saying the version it read is archival. That case gets its own stale reason - `open only on an archive` - because printing the group would have said `stall` is the reason a stall note is not open work. A disagreement names a LIVE carrier wherever one exists, the rule `readRef` already followed. `refs` is sorted, so a version carried by both a branch and an archive named whichever sorted first: 30 rows here presented a disagreement as preserved history while a live branch carried that exact state, hiding the only half a reader can act on. READS, NUMBERS AND EXIT CODES Any missing version is an incomplete read, not only all of them. Comparing against zero meant a path whose other version came back was reported as complete - and the version that went missing is exactly the one that could have carried the disagreement. The row is still built from what did read; the path is named, which is what makes the run exit 2. A number named only as `#370` is now attributed. AGENTS.md mandates that spelling for anything posted to GitHub and notes use it, so the short-form-only test called such a number unattributable and printed a confluentinc lookup beside the fork one - a reference that RESOLVES to an unrelated upstream issue. `parseRegister` already read both spellings; this half had not been brought with it. A ref whose listing failed now fails the run. It carries an unknown number of notes and the renderer said so, while the exit code returned the documented "ran successfully" status - and the exit code is what a caller tests. The three failure depths are one exported `runFailure`, so each clause has a control and the end-to-end run proves the front door calls it. `rank stall` is refused rather than answered. Validating only `--`-prefixed tokens meant it ran the unscoped query, and `rank --impact stall extra` ran the scoped one with a token dropped. WHAT THE READER SEES Every off-baseline read names its pull request. The row for a note deferred on the baseline and open on a branch returned early and rendered with no pull request at all - the row where naming it matters most, since the branch is the only place that work is live. The unranked half no longer makes a claim the parse cannot support. A note named by a list item this parse does not read was counted as "NOT named by the register" three lines under a sentence saying those items are outside the delta entirely. The counts stay, labelled as the upper bound they are. TWO CONTROLS THAT WERE PROVING NOTHING The ownership refusal - the most consequential claim this command makes - serialised the row objects and asserted on `row.relation`. The user-facing sentence is hard-coded in the renderer and never reads that field, so a view rewritten to say "fixed by" would have left the control green while its mutant died on an otherwise unused field. It now asserts on `formatRank`'s text and mutates the renderer. The register-failure control never invoked the front door and expected `ok: true`, so deleting the front door's failure branch left it green. It now drives `bin/inflight.mjs rank` against a baseline with no register and asserts exit 2, with the library assertion kept as the second half. Two findings from the review are deliberately NOT taken, and stay named in the PR body: the register parse assuming a continuation line is indented, and `carryingRefs` counting refs carrying any version while `readRef` names the chosen one. Both are now in the note that outlives this branch.
…e it before its name lands docs/inflight/AGENTS.md gives four outcomes in order and deleting is the last of them. The general lessons went to docs/solutions/ in the commit before this one; what is left here is the three things nobody has decided. RENAMED, and deliberately. `branch-` means work sitting on a branch with no pull request, which stopped being true when #438 opened. Renaming a note breaks every citation of it, which is why the directory doc weighs the move rather than recommending it - but this note has never been on the baseline, and exactly one document cites it. It is free now and never again, so the call is made now rather than left to rot as a prefix that says the wrong thing. `ci-` because what remains is about the agent harness, which is that prefix's stated area. No "delete this when it merges" marker: the merge is exactly when nobody is looking here. WHAT STAYS, and two of the three were only in the pull request body: - the `candidate` relation, recorded as deferred rather than rejected, with what would earn it - a signal binding the branch to the note rather than to a number; - the register parse's assumption that a continuation line is an indented line, which fits this register, is permitted-otherwise by markdown, and is enforced by nothing - the failure is a wrong answer, not an error; - `carryingRefs` counting refs carrying any version while `readRef` names the chosen one, so "carried by N refs, read from X" overstates carriage of the version actually reported. The design, the alternatives and the reasoning are already owned elsewhere - the plan, `bin/lib/rank.mjs`'s header beside the code, `docs/inflight-tool.md`'s worked example - so none of it is restated. The plan now points back at the note's new name and at each migrated lesson's owner, so a reader arriving from either side finds it. Also here, since it was measured in the same pass: a `docs/refactoring.md` line for `views.mjs`. It grew by a third with `formatRank` and is now second-largest in bin/lib, but its boundary is intact and `refsText`/`scopeLine` are shared with docsShape on purpose. The trigger for splitting is a SECOND rank-family command, not the size - `docs-views.mjs` earned its split by being several formatters, while every other command's single formatter still lives in views.mjs.
…ne on purpose AGENTS.md's merge-prep rule: once the defect CLASS is understood, grep for its shape and say what was found, including where nothing was. Three live instances outside this work, each recorded rather than fixed, because each changes a command this PR has no business changing. The sharpest is that `docsShape` still makes BOTH version choices rank had to abandon - the unconditional baseline preference and the first-sorted-live-ref - and then groups what it read with the same imported group rule. So a note deferred on the baseline and open on a branch is grouped as deferred in the index injected into every session. It is also the case rank exists to report, which makes it the first thing the tool can now say about its own family. Changing it moves what every agent sees at session start, so it wants a measurement, not a rider on this PR. `note find` and `stranded` ignore `unreadableRefs` at exit 0, where rank now fails the run - `stranded` is the one that matters, since an unlisted ref can make a cluster look preserved when a live ref carries it. `docs show` takes the first positional and drops the rest, the shape `rank stall` had; `docs for-branch` already refuses extras and is the model. Ruled out, and worth saying so: `bin/lib/terms.mjs` already reads all three issue-reference spellings, so the short-form-only regex was one instance rather than a pattern; and among the self-test's controls only the ownership refusal asserted on a data field whose wording is written out independently.
…o be deleted The reproduce line anchored the by-construction argument to origin/feats/inflight-rank-cli. That branch is deleted when #438 merges, at which point the command errors and the claim it supports becomes unverifiable - with nothing going red to say so, because no gate checks a branch ref. AGENTS.md makes the same point about branch names generally: a branch name says nothing about whether the work landed, and nobody comes back to upgrade it. Asked of the pull request instead, the answer is identical and outlives the branch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MWLkoGdvM2CHQCsYpgUMAR
…ion (#447) RegistrationRaceStaleResidentIT.freshArrivalCollidingWithStaleShardResidentMustStillGetProcessed failed on #438's CI with the same mid-loop pause-point signature every other sighting in that row carries - its saturation setup guard, not the confluentinc#909 assertion the test exists to make. The row already argued master-state and test-side by comparison: the same test passing on the same branch at a neighbouring commit, and two documentation-only heads on #442. This sighting reaches the same conclusion by construction instead. Those documentation-only heads still sit on branches whose earlier commits changed Java, so the tree that was built carries those changes; #438's whole branch diff is bin/ and docs/, compiling no Java and changing no test, so the tree built there is master's exactly. A branch that could not have caused an integration failure produced one, and no neighbouring commit is needed to see it. The reproduce line is asked of the pull request rather than of the branch, because the branch is deleted on merge and the check has to outlive it - the answer is identical either way. Recorded before the observing PR merges, per AGENTS.md: seeds and job logs expire, and while `bin/inflight.mjs codecov test` holds the per-commit outcome independently, no command produces the interpretation. Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
🧪🔒 Quarantine Lane Report
🔴 expected while the owner PR is open · 🟡🎲 flapper, pass proves nothing · 🚨 a deterministic quarantined test passing means its fix landed: delete its No quarantined test changed outcome since the previous push. Updated for |
|
@claude review this |
|
Claude finished @astubbs's task in 3m 28s —— View job Review
What I checkedThis PR already went through an unusually thorough cycle: 11 Codex findings, all confirmed by astubbs in One further instance of the disagreement-ref classThe fixed finding ( const disagreement = [...new Set(seen.map((v) => v.group))]
.filter((g) => g !== key)
.map((g) => {
const refs = seen.find((v) => v.group === g).refs // <-- only the FIRST matching version
const ref = refs.find((r) => !archival.get(r)) ?? refs[0]
return { group: g, ref, archival: archival.get(ref) === true }
})When two different blobs (distinct content, e.g. two branches that independently closed the note with different edits) both classify into the same disagreeing group, Given 317 live refs in this corpus and how common independent branch divergence has proven to be here (the 30-row count for the already-fixed case), this seems plausible to hit rather than theoretical. It only affects the annotation line ( Everything elseNo other issues found. The two documented residuals (continuation-line indentation assumption, |
Inherited #440, #447 and #435. Two of them touch this branch's ground. #440 QUARANTINES the registration-race test this branch has been recording sightings of. Nothing here needs to change for that - the ledger row and the quarantine registry are different instruments, and master keeps both - but it means a future red from that test is no longer this branch's problem to attribute. THE LEDGER ROW, CONFLICTED A THIRD TIME AND UNIONED AGAIN Both sides said "12 seen" and neither was a superset: they count DIFFERENT seventh sightings on 2026-09-03 - master's is #438, this branch's is #433 - so the union is thirteen. A count that matches on both sides is the easiest kind of conflict to resolve wrongly, because taking either side whole looks like a no-op and silently drops one real sighting. Master's #438 control is the stronger of the two and is kept as such: that branch's whole diff is `bin/` and `docs/`, compiling no Java, so the tree built there IS master's - branch content is excluded by construction rather than by comparison. This branch's within-branch control is kept beside it because it establishes something that one cannot: one tree, both outcomes, which excludes the tree itself and leaves only the runner. They are different arguments, not two tellings of one. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01C3KnLcX4KibVMCP6sMfNgM
…s chosen The silent no-post recurred on 2026-09-06 - run 34066111691, dispatched with a steer, concluded success, posted nothing. Roughly a month after the first measurement, so it is a recurrence, and the reviewer job's own guard passed exactly as ci-review-agent.md predicts it must. WHAT THE SECOND OCCURRENCE ADDS is an isolation the first could not. That PR had no earlier reviewer comment, so nothing existed for the gate to rest on and claude-review correctly stayed red. The no-post and the false-green are therefore separable: the no-post fires every time, the false-green only follows when a prior reviewer comment exists. A first-time review fails safe; a re-review does not, and that is the case where a reader most expects the check to be about the current head. IT RECURRED TO SOMEONE WHO HAD THE WRITE-UP AND DID NOT READ IT. The route was chosen from docs/ci.md, whose automated-review section opens by telling you to dispatch and never mentioned the risk; the measurement was filed only under docs/solutions/. A finding that lives where nobody is standing when they choose does not change the choice. So the consequence and the verify-a-comment-arrived rule now sit beside the dispatch command itself, with the evidence still owned by the solutions write-up rather than restated. The inflight note's deferred state is deliberately unchanged, with the reasoning recorded rather than left as a non-decision: deferral is right for that file, and what the recurrence argues for is a split - the guard is one post-condition step calling bin/check-review-posted.sh, which already exists and which the gate already uses. That split is its own work and does not belong to the PR that noticed it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MWLkoGdvM2CHQCsYpgUMAR
…named the archive The fifth defect of the version-disagreement class, and the first found in the annotation rather than in the chosen version. Raised by the cross-model review pass on #438, against the fix that landed one round earlier for the same rule one level down. `rank` deduped the disagreeing groups and then reached for `seen.find(v => v.group === g)` - the FIRST version classifying into that group - before preferring a live ref among its refs. When two different blobs share one group, which is what two branches closing the same note in different words produces, only one of them was ever consulted. `for-each-ref` orders by full refname, so `refs/backup/**` is enumerated ahead of `refs/remotes/**`: the archival closure won the lookup, and the row said `DISAGREEMENT: on <a backup ref> (an ARCHIVE)` while a live branch carried that exact state. Same cost as the round before it - the half a reader can go and act on is the half that goes unnamed - and the repository has 44 `refs/backup` refs for this to fire on. The fix pools the refs of every version in the group before applying the live-first rule, so the two levels now follow one rule rather than two. It changes nothing for a group carried by a single version, which is why the existing control stays green. Control: `rank-names-a-live-carrier-when-two-versions-share-one-disagreeing-group` drives the fixture note `bug-two-refs-closed-it-differently` - open on the baseline, closed by two distinct blobs, one on `refs/backup/closed-differently` and one on `origin/z-closes-it-live` - and asserts the disagreement names the live branch and that the backup ref appears nowhere in the rendered output. Its mutant restores the `find`, and was verified to die BY THAT ASSERTION rather than by a crash: with the mutation applied the entry reads `{"ref":"backup/closed-differently","archival":true}` and the check returns false, no throw. `seen.find` cannot return undefined here - the group came from `seen` - so a crash-death was the plausible way this control could have been vacuous, and it is not. 320 self-tests pass (160 checks and their mutants); `bin/check-all.sh` clean. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MWLkoGdvM2CHQCsYpgUMAR
Fixed now, in e1596d0, rather than recorded - it is the same rule the code already states two lines above the defect ("a live carrier names the disagreement wherever one exists"), so this was a hole in an existing rule rather than a new one, and the two residuals are decisions that are still open where this was not. The finding holds exactly as described, and the ordering is worse than "array order" suggests: The fix pools the refs of every version in the group before applying the live-first rule, so the two levels follow one rule instead of two. A group carried by a single version is unaffected, which is why Control: On the vacuous-mutant point in the steer above: the mutant was checked to die by that assertion, not by a crash. Self-tests all pass; |
A rank row read off the baseline prints its pull request when one exists, and [PR state UNKNOWN] when the lookup could not answer. The third case - no suffix at all - was the one a reader could not tell apart from nobody having asked, which is the false-negative shape this command exists to refuse. Codex asked for a per-row snapshot-miss marker. Declined, and the reasoning is kept rather than discarded: prKnown already carries the unknown-versus-absent distinction, and most carrying refs have no pull request at all, so a per-row absence marker would land on the majority of rows and crowd out the informative suffix on the minority that carry one. Operator ruling: state it once in the scope header instead. The ambiguity goes away at no per-row cost, and the header is where a reader is already being told what the run did and did not cover. Pinned by a check on the RENDERED text, not on the row objects - this branch's own lesson is that a view no check renders has an unenforced contract, and the three legibility defects it already ate all hid behind exactly that gap. The mutant blanks the line and the check goes red on its assertion rather than by crashing. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MWLkoGdvM2CHQCsYpgUMAR
…this branch's sighting #433 and #443. One conflict, the RegistrationRaceStaleResidentIT ledger row, and both sides had reached "13 seen" while counting different sets - so either taken whole would have dropped a sighting under an authoritative-looking number, for the third time on this file. Master's side is the base and deserves to be: it carries the #442 and #438 sightings and, more importantly, a WITHIN-branch control this branch had no way to know about - #433 ran the same content twice, green then red after a re-cut that changed only the commit split, so one tree produced both outcomes and the tree itself is ruled out, leaving the runner. That is a stronger claim than the cross-branch controls it joins. Added to it: this branch's 2026-09-04 #444 sighting with its same-head re-run, and one clause noting that sighting is another of the same-day cross-branch kind. 14 seen. Recorded because the first attempt at this resolution asserted on an anchor that did not exist, wrote nothing, and the file was then staged WITH its conflict markers - caught by grepping the markers rather than trusting the exit code. The second attempt carries an anchor-free fallback that announces itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WErnxQd9Ew57F9SsqzdPU5
Description
Adds
bin/inflight.mjs rank- a read-only subcommand that reads open in-flight notes from everyref, groups them the way the session index groups them, annotates each row with the relations it
can prove, and reports where that picture disagrees with
docs/inflight/process-candidate-ranking.md.The delta is the deliverable. It turns the next ranking pass from "read every note" into "look at
these few disagreements". Run
node bin/inflight.mjs rankfor the current set.Read "defects found by running it" below before trusting any output quoted here. Earlier
descriptions of this PR cited the delta as evidence the feature worked. It was not: review found it
both inventing disagreements that were not there and hiding real ones, and both were live on this
repository. They are fixed and pinned by controls, but the honest summary is that this command's own
output was the least reliable thing about it until very late.
The gather step is not rebuilt, and this PR says so up front.
docs list inflight <impact>already walks every ref, keeps open notes, groups them by impact and names the branch an off-baseline
note was read from.
corpusIndex,classifyNoteandinflightGroupOfare imported rather thanre-derived; the group rule in particular has to keep one owner, or two surfaces would place the same
note differently and nothing would report the disagreement. What
rankadds is the carrying branch'spull request, the live-versus-archival split, the number in the filename, an accounting of the open
notes no impact bucket claimed, and the register delta.
What it refuses to say, and the case that forced it
A note travels on the branch that produced it, so carriage is cheap to know and ownership is not
available at all. Every row says
CARRIES. This is the live output today:Two things are happening there, and both are the point.
The pull request sits on the branch that carries the note and does not fix the bug the note
describes - the note's own text says the bug predates it. A row reading "owned by" that pull request
would be confidently wrong, so no row here ever says it.
And the
DISAGREEMENTline is the facade doing its job on real state rather than a constructedcase: a branch has settled that note while the baseline still shows it open. Reading a single
arbitrary version would have answered for whichever ref sorted first and silently dropped the note
from the backlog - which is exactly what the first cut did.
A
candidaterelation (a branch whose name encodes a number matching the note's) was designed,measured and dropped rather than shipped behind a caveat: some matches are cross-namespace by
construction, because a note filename carries a fork number while branch names here encode upstream
ones. Deferred, not rejected - a pull request body naming the note path would be a signal that earns
it.
Defects found by running it, not by reading it
Each is now pinned by a check whose mutant restores the defect.
from the backlog entirely.
refs and the path's live refs disjoint sets, so the read ref resolved to nothing and the row threw.
Reproduced against a purpose-built corpus before the fix.
numberForread itstextargument from outside its scope, so the command threw aReferenceErroron any note carrying a number - most of them - and exited 1, which is neither ofthe tool's two documented codes.
own dead predecessor carry the same positional number; a map keeping the last writer reported the
register's live entry as stale while naming the dead copy.
gh issue view 600for a release version.release-0600-blockers.mdis on the baseline andits
0600is the release 0.6.0.0 - a reference that resolves to the wrong issue, whichAGENTS.mdrates worse than a broken one. A leading zero is now refused.docs/inflight/AGENTS.mdnames the pre-conventionnotes whose number is confluentinc's, and
bug-857-family.mdis on the baseline with"Paused consumption across multiple consumers confluentinc/parallel-consumer#857" in its own title. Attribution now comes from the note's own text; a note that
names neither repository gets both lookups and is labelled unattributable.
missingfrom the batch read - a partial clone, a gc race - wasdropped with no accounting while the run still reported success.
origin/*ref read as pull-request-less, because the snapshot is keyed on bare branchnames and the lookup used the full ref.
so prose, continuation lines and the register's own "What is NOT on this list" paragraph became
ranked entries - every number reported as resolving to nothing was a false positive, and two were
real entries whose filename the same run resolved. A citation of a doc outside
docs/inflight/became a phantom entry reported as
absent. In the other direction a live note was suppressedfrom the unranked half by the sentence "fixing confluentinc#833: ParallelConsumer would run for a while and then exit due to InternalRuntimeException(Timeout) #177 does not close it". An entry is now a
list item including its continuation lines - this register opens items with the number and names
the note two lines down - and is satisfied when either half resolves.
docs/inflight/CLAUDE.mdwas ranked as open work, because this walked the corpus index withoutthe membership guard
docsShapeapplies. Two surfaces disagreeing about what the corpus containsis the failure importing
inflightGroupOfwas supposed to prevent.the all-clear and then listed the whole corpus as unranked. The same found-nothing / could-not-look
collapse, moved from the git walk to the parse.
to prevent, committed by the command, one layer up from the version-selection bug above and worse.
chooseVersionpreferred the baseline's version unconditionally, so a note the baseline callsdeferred while three live refs carry it OPEN vanished from its group and the delta told the
register it was deferred, naming no ref.
core-auto-scaling.mdis the worked case and one of theregister's own ready picks. A still-open version now beats any preference.
had stopped being open work a few lines above its own row saying that note survives only on a tag.
register cites bare unqualified numbers and upstream ones this parse does not read, so the count
read as a complete reading of the register and was not - it is 11 of 17 list items. The line now
says so, and what the unread six look like. This one had already been quoted to a human as a
complete reading before the denominator existed.
neither set made
buckets.get(key)undefined and threw - and with no top-level try/catch theprocess exited 1, neither documented code, which a caller testing for 2 reads as a successful run.
The complement is now derived from the owning list.
Smaller, from the same round: the disagreement line named tags as though they were branches;
index.unreadableRefswas computed by the corpus index and read by nobody, so a ref whose listingfailed was a could-not-look reported as a found-nothing; and the delta headline printed
11 entrys.One reported defect did not reproduce and is deliberately not fixed. An unknown
--impactgroupwas reported as crashing with a bare
TypeErrorand exiting 1. It does not:rank --impact nopeexits 0 and prints the valid groups, the library does not throw, and both are covered by a check.
The real crash in that finding is the taxonomy-partition one above, which is fixed.
Other decisions the tests pin
inflight-state: open - <reason>is not open, and one such note is on the baseline today.feature,unmatched) are emitted rather than dropped: thatis the largest group in the corpus, and the register names notes in it.
branchViewuses for onebranch would be one untimed
ghsubprocess per pull-request-less branch here, with absencesdeliberately uncached.
(exit 2), the shape
refactor-windowalready uses - "the delta was empty" and "the delta neverran" are different answers.
A test-harness fix that rides along
Not a change to the feature, and worth reviewing separately:
bin/test-inflight.mjsleaked a fullcopy of
bin/plus.claude/hooksfor every check, permanently. Nothing ever removed the mutantdirectories, so each run left one copy per check behind and the leak grew every time a check was
added.
Measured on the machine that found it: 20 GB of
inflight-selftest-*directories from accumulatedruns, on a disk the repo's own pre-commit hook was already warning was 94% full. Reproduce with
du -sh "${TMPDIR}"/inflight-selftest-* | tail -1, and count a single run's contribution by listingthat glob before and after.
The cleanup is a
finallyrather than a line at the end of the loop body, because thecould-not-be-built path takes a
continueand leaked too. Verified by counting the glob either sideof a run: unchanged, where it previously grew by the number of checks. Found while finishing other
work in the file rather than by a gate - nothing measures the suite's own footprint, and a temp
directory nobody lists is invisible until the disk fills.
Review coverage, stated honestly
Applied: coherence, feasibility, scope-guardian and adversarial on the plan; reuse, quality and
efficiency on the code; project-standards, maintainability, reliability, agent-native, correctness
and testing on the code.
The last two are why this section is worth reading.
correctnessfound the delta - the stateddeliverable - wrong in both directions on the real register, after every other lens had passed over
it;
testingfound thatdisagreement, the field this command exists for, was asserted nowhere.Both landed after an earlier version of this description had already claimed the delta's output as
evidence the feature worked. The claim was wrong and has been corrected above.
The
adversariallens ran last, in-process, after an independent cross-model pass against Codextimed out without producing an artifact. It found the baseline-preference P1 that every earlier
lens had passed over, including on code they had already reviewed.
That is the shape of this PR's review history worth knowing: each of the last three lenses to return
found a defect the previous ones missed, and the most serious one came last. The corroboration to
draw from that is not "it has been reviewed a lot" but that a note whose versions disagree is the
case this command keeps getting wrong, and it is where a reviewer should push hardest.
The cross-model reading, which did arrive
An earlier version of this description ended "still not obtained: a genuinely cross-model reading.
The Codex route never produced one." That is no longer true. A Codex pass was requested with a steer
naming the version-disagreement path as where to push hardest, and it returned findings across P1 and
P2. Each was verified before being acted on; every reply is in its own thread.
It found a fourth defect of the class the steer named.
featureandunmatchedare open workand belong in the ranked set, so a baseline copy carrying no impact tag TIED with a live copy
carrying one, and the baseline tie-breaker took the placeholder - the row landed in the catch-all
instead of its impact bucket, and the delta told the register the entry was stale while a live ref
carried it as ranked work. Predicted before the fix and confirmed after: the
unmatchedbucket fallsby exactly the count of notes in that state, 202 -> 117 on this corpus. That is the third such defect
found after an earlier one was fixed and reviewed, and it should be read as the case being hard
rather than as the review record being thorough.
Two more it found are live here and were not hypothetical: a disagreement carried by both a branch
and an archive named whichever ref sorted first, presenting current state as preserved history on
dozens of rows; and a number named only in the fully qualified
astubbs/parallel-consumer#NNspellingwas called unattributable and printed a confluentinc lookup beside the fork one - a reference that
resolves to an unrelated upstream issue. The same file's register parser already read both spellings,
so the module disagreed with itself.
And it found both remaining vacuous controls. The ownership refusal - the most consequential claim
this command makes - asserted on a data field the renderer never reads, so a view rewritten to say
"fixed by" would have left it green. The register-failure control was named for a process exit code
and never started a process. Both now assert on the behaviour in their own names, with their mutants
moved to the layer that produces it.
One finding is half declined, and its thread is deliberately left open for the author: rendering a
successful pull-request-snapshot miss as an explicit absence.
prKnownalready carriesunknown-is-not-absent, and most carrying refs here have no pull request, so an absence marker would
appear on the majority of rows. The accepted half - every off-baseline read naming its pull request -
is fixed.
Two residuals, still deliberately unfixed
Both are now recorded where they outlive this branch, in
docs/inflight/ci-inflight-rank-relations-and-parse-assumptions.mdrather than only here:is permitted otherwise by markdown, and is enforced by nothing - the failure is a wrong answer, not
an error;
carryingRefscounts refs carrying any version whilereadRefnames the chosen one, so"carried by N refs, read from X" overstates carriage of the version actually reported.
The
candidaterelation stays deferred rather than rejected, in that same note.origin/masterwas merged in before this was opened; the merge was clean and touched no file thisbranch changes.
Checklist
docs/inflight-tool.mdgains a worked-example section, anddocs/features/cross-ref-repository-queries.yamlgains ause_this_whenline plus aboundariesentry for the ownership refusaldocs/features/- the existingcross-ref-repository-queries.yamlalready describes this tooling and was extended rather thanduplicated; a second file describing the same surface would be the drift this repo treats as a
defect
bin/test-inflight.mjs, each with a negative controlverified to go red; the suite passes both arms.
grep -c "^ id: '" bin/test-inflight.mjscounts themdocs/inflight/working note started at the PR's first commit, and prepared for merge -opened under a
branch-name in the first commit, and renamed todocs/inflight/ci-inflight-rank-relations-and-parse-assumptions.mdat merge prep:branch-means work on a branch with no pull request, which stopped being true when this opened. It has
never been on the baseline and exactly one document cited it, so the rename is free now and
never again. Its general lessons migrated to
docs/solutions/; what remains is the threethings nobody has decided
ce-simplifyandce-code-reviewlocally, plus a cross-model Codex pass - the in-processfindings are the "defects found by running it" section above and the cross-model ones are in
their own section. The automated
@claudereview has deliberately not been requested, so ared
claude-reviewon this draft is the expected state, as is a redreview: human LGTM.Design, alternatives and the full requirement set:
docs/plans/2026-09-03-002-feat-inflight-rank-backlog-view-plan.md.Two things added after review, and after this description was first written
The Codex thread on
views.mjsis settled, by the middle option. A rank row read off the baselineprints its pull request when one exists and
[PR state UNKNOWN]when the lookup could not answer - sothe third case, no suffix at all, was the one a reader could not tell apart from nobody having asked.
The per-row absence marker Codex asked for was declined:
prKnownalready carries theunknown-versus-absent distinction, and most carrying refs have no pull request, so that marker would
land on the majority of rows and crowd out the informative suffix on the minority that have one. It is
stated once in the scope header instead, and pinned by a check on the rendered text - this branch's
own lesson is that a view no check renders has an unenforced contract.
A review-route finding rode along, because this PR is what surfaced it. Requesting the automated
review by workflow dispatch produced a run that concluded
successand posted nothing - the secondoccurrence of a defect measured once before and deferred since. What the recurrence adds is an
isolation the first could not: this PR had no earlier reviewer comment, so nothing existed for the gate
to rest on and
claude-reviewcorrectly stayed red. The silent no-post and the false-green aretherefore separable, and a first-time review fails safe where a re-review does not. It recurred to
someone who had the write-up and did not read it, because
docs/ci.mddescribed the dispatch route atthe point of choice without carrying the risk - so the consequence and the verify-a-comment-arrived
rule now sit beside the dispatch command, with the evidence still owned by the solutions write-up.