Repository navigation
test(soak) astubbs#119: any retry-forever instance eventually stops fetching, at a threshold you can compute - #487
Conversation
…k reports its freeze #471's two thirty-minute runs found a total intake stall and could not name what stopped intake, because the only figures on the progress line were the counters that had already stopped moving. Its arm 1 - "re-run either arm with WorkManager at DEBUG and read the isSufficientlyLoaded= line at the moment successes freeze" - was blocked on instrumentation that did not exist. This is that instrumentation, plus the accounting gap the hypothesis rests on, pinned as a unit test. THE FLAG THAT WOULD NOT HAVE REACHED THE RUN. The gate's DEBUG line lives on bz.stub.parallelconsumer.state.WorkManager, which sits under the bare bz.stub.parallelconsumer pin in both test logging profiles - so -Dpc.log.level=debug does NOT raise it, and a run started that way would have produced a confident "no gate line, so not the gate". It now has a logger of its own, on its own property (-Dpc.loadgate.log.level=debug), because it fires once per control-loop tick and nothing else pc.log.level raises wants that volume. The logger goes in logback-test.xml, which is the file an integration or soak run actually reads - the integration profile's header records that nothing selects it - and in logback-integration-test.xml as well, each defaulting to its own file's level, so neither can lose it by being the one selected. It replaces the commented-out logger that was already there for exactly this investigation: a flag cannot be committed by accident and a commented-in logger can, which is the argument logback-test.xml's own header makes about pc.log.level. WHAT THE SOAK NOW REPORTS. Every progress line carries the gate's operands, read through the same ShardManager#getWorkableRecords() accessor the gate itself decides on, plus Kafka's paused-partition count - which is what a latched gate actually does. A frozen success count beside `loaded=true pausedPartitions=20` is the gate holding the poller down; the same freeze beside `loaded=false pausedPartitions=0` is something else entirely, and after the fact those two used to be indistinguishable. At INFO, so it is in the log of every soak that ever runs, rather than behind the DEBUG flag above. THREE KNOBS, EACH SO AN ARM CAN CHANGE EXACTLY ONE TERM. -Dsoak.ordering=UNORDERED is the control for the head-of-line half of the hypothesis: no shard head can block anything behind it, so inShards becomes an honest count of selectable records. -Dsoak.messageBufferSize=N raises the gate's threshold and nothing else - it pins the load factor so target*factor is the size asked for, leaving maxConcurrency, the worker count and the retry throughput untouched, which raising maxConcurrency would not. -Dsoak.progressInterval=PT10S, because the 60s interval that suits a thirty-minute run puts the whole freeze inside a single sample. THE ACCOUNTING GAP, PINNED. WorkManagerTest#theLoadGateCountsRecordsQueuedBehindABlockedKeyHeadAsWorkable is a characterisation, not an assertion that the reading is right: with three records of one key and a failing head, the gate reads two workable while nothing at all is selectable, and three while one is. The gate's own javadoc excludes retry-parked records because "no amount of worker capacity can advance them"; a record behind a blocked KEY head meets that description and is counted anyway. The test deliberately also records what this is NOT evidence of - a permanently failing head is itself workable and counts on its own account, so a buffer full of poisoned heads latches the gate whether or not anything is queued behind them, and only a soak arm separates the two. Co-authored-by: Claude Opus <noreply@anthropic.com>
…found The note this PR earns at its first commit, per docs/inflight/AGENTS.md. It carries what the unit test alone cannot: what the accounting gap is NOT evidence of, why "count only what is selectable" is the wrong repair, the three directions that are available, and the second consumer of the same over-read that drain() gates on. Named for #119, the fork mirror of confluentinc#857, because a note filename carries the FORK number - bug-857-family.md predates that rule and is left alone. Co-authored-by: Claude Opus <noreply@anthropic.com>
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
… than reasoning about it The sweep for other readers of "records in shards" as "available work" named drain() - it gates the transition to closing on isRecordsAwaitingProcessing(), which sums the per-shard selection-claim counters, and a record queued behind a blocked KEY head still holds its claim. That was an argument from reading the code; it is now an assertion in the same test that pins the gate itself, so the claim in the note is measured. The consequence differs from the gate's - a close that waits out its drain timeout, not a poller that stays paused - which is why the note records it beside the gate question rather than inside it. Co-authored-by: Claude Opus <noreply@anthropic.com>
✅ Duplicate Code ReportTwo engines run in parallel for cross-validation. Each has its own thresholds tuned to its baseline - the real safety net is the per-engine "max increase vs base" check. ✅ PMD CPD
No new clones introduced by this PR. ✅ jscpd (language-agnostic)
No new clones introduced by this PR. Powered by astubbs/duplicate-code-cross-check |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## master #487 +/- ##
============================================
+ Coverage 82.50% 82.95% +0.45%
- Complexity 1563 1582 +19
============================================
Files 96 96
Lines 5383 5421 +38
Branches 537 544 +7
============================================
+ Hits 4441 4497 +56
+ Misses 747 727 -20
- Partials 195 197 +2
Flags with carried forward coverage won't be shown. Click here to find out more. ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
🟢 Throughput — OKThis branch measured about 27% faster than master, on the one test this measures. That is larger than this test's own run-to-run spread of about 17%, so it is worth looking at.
Allowable range 🟢 ≥ 0.70 · 🟡 0.50–0.70 (about a 30% loss) · 🔴 < 0.50 (about a 50% loss) What the numbers mean, and what they cannot tell youThe one that gets misread. Why a shape and not a rate. A rate depends on which runner you drew. A shape does not: every test here processes a fixed number of records, so a runner twice as slow doubles the subject and the controls together and leaves their ratio alone. That is the whole trick, and it is why the reported rate is shown last and labelled as this machine only. Reading the comparison. By conservation, not by correction. Every test in this lane processes a fixed number of records, so within one run the ratio of one test's time to another's is invariant under machine speed — a runner twice as slow doubles both terms and leaves the ratio alone. There is no machine-index correction to be wrong, because nothing needed correcting. Per-method times, not class times. A class time is Reference is the median of 9 recent What this still cannot do. It removes machine-to-machine variance. It does not remove this test's own run-to-run variance, measured at about 30% on a single unchanged commit while its controls stayed within 5%. That is a property of the test, not of the comparison, and no arithmetic here can touch it — which is why the reference is a median and the bounds are deliberately coarse. 🟡 means look at this; only 🔴 is outside the measured spread. Runs used: eb9fdb0, 51d9bb2, 0ca787c, b654cb2, a055248, 65e11e3, e05b399, ee3d8a9, 890aa53 Since the previous push: ratio 0.979 -> 1.265, share 1.876 -> 1.46, rate 64008 -> 124172 (+94.0%). One push of difference sits inside this test's measured spread - read it as movement, not as a result. Updated for |
|
…PR that carries it The three lines the release note still has to name now read like the rest of the burn-down: a box, the PR in front, and what closes the box. A ticked box means the v6 action for that line is done, not that the defect is closed - the revoke wait is ticked because its release-note sentence is written and its fix is #408 after v6; the INSTANCE_STALL line waits on #488, whose load arm has now measured the starvation reading from the direction the idle replays never could; the intake stall waits on #487, whose soak arms are instrumented and predicted but not yet run. Claude-Session: 460f7df9-dcc2-4b00-a9f9-62f3a2c6d5e4 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
… blocking is not why #471 named the record-intake load gate as the untested candidate for the stall its soak found and said one run reading the gate's own DEBUG line would settle it. Three ran. The hypothesis is half right, and the half that is wrong is the mechanism it named. THE ARMS, each differing from the first by exactly ONE term, six minutes rather than thirty because the stall is reached in the first second. Same seed 3747722682837130843, failureFraction 0.5, 1000 keys over 20 partitions, maxConcurrency 14, 100ms user function, 1000 records every 20s, the suite's Testcontainers Kafka, macOS arm64 workstation at 5-13 load average. WorkManager at DEBUG throughout, and the config that carried it VERIFIED in the log by OnConsoleStatusListener rather than assumed - which mattered, because the gate's logger sits under a bare package pin and -Dpc.log.level=debug does not reach it. ARM 1 KEY, the experimental arm. PREDICTED: successes freeze inside 60s; the gate reads true with workable above the target and STAYS true; every partition paused; failures climb at ~14 workers / 100ms = 140/s. OUTCOME: confirmed on every clause. Succeeded froze at 451 - #471's own thirty-minute number, reproduced in six minutes - failed 47,977 (133/s). The gate read true on 37,356 of 37,360 evaluations; the four false ones are the first 800ms, before the first fetch. It latched at inShards=500 vs target(14)*loadingFactor(2)=28 - on the FIRST fetch, 0.8s in - and never unlatched. 20 partitions paused at every sample. ARM 2 UNORDERED, the control on ORDERING. No shard head can block anything behind it. PREDICTED (and recorded before running, in the form expected to be right rather than the form that would be convenient): it stalls too, which refutes the head-of-line half while leaving the gate half standing. OUTCOME: as predicted. Gate true on 39,689 of 39,693 evaluations, 20 partitions paused throughout, successes crawling 232 -> 314 on records already in the buffer. ARM 3 -Dsoak.messageBufferSize=20000, the control on the GATE ITSELF. Threshold 42 -> 20,006, maxConcurrency and therefore the worker count and retry throughput untouched. PREDICTED: the outcome flips - the gate stays false, partitions stay unpaused, records keep arriving and successes keep rising. OUTCOME: flipped. Gate false on all 35,652 evaluations, ZERO partitions paused at any sample, inShards climbing monotonically 549 -> 17,103 with the producer, successes 897. VERDICT. Arm 3 is the positive control, so the gate is what stops intake. Arm 2 kills the stated mechanism. Arm 1's own arithmetic kills it more directly: the 549 records it held came from ONE burst over 1000 distinct keys, so there was at most one record per key and NOTHING was queued behind any blocked head. What latches the gate is records that are themselves perfectly workable - retried continuously, saturating every worker - and that never retire. AND ARM 3 EXPOSES A SECOND BOUND, which is why no gate change fixes this. Lifting the intake bound did not restore throughput: successes doubled and then plateaued by minute three while the held population kept climbing linearly. The poisoned records re-offer themselves every retry delay and consume the whole worker budget, so raising the threshold converts a hard stall into an unbounded-memory slow starve. "Count only what is selectable" is worse: under KEY or PARTITION ordering at most one record per shard is ever selectable and a shard whose head is at a worker has none, so a healthy loaded instance would call itself under-loaded and fetch without bound. The distinguishing property is liveness of the shard head, not a count, and that is not decidable from the shard's state. The fix has to bound the FAILURES - #149's dead letter queue, which roadmap.yaml already describes in these terms - and the cheapest thing available before then is to make the latch loud rather than silent. Still eliminated, re-measured on all three arms: offset-encoding back pressure. Neither of PartitionState#updateBlockFromEncodingResult's two messages appears once in any of the three logs. RECORDED WHERE IT WILL BE FOUND. CommitResponseTimeoutSoakIT's Calibration status gains the three arms and the verdict, and its earlier block is dated so the two read as a sequence. docs/inflight/bug-119-load-gate-counts-blocked-work-as-available.md carries the operator-visible symptom, the decision and the sweep. bug-177-commit-response-timeout-unreproduced.md's arm 1 is marked done and hands the gate question to that note, keeping #175's own question where it belongs. Co-authored-by: Claude Opus <noreply@anthropic.com>
…o gate fix #487's three arms settle the last replay-shaped unknown. The load gate latches on the first fetch under a poison-record workload and never unlatches; UNORDERED stalls the same way and the KEY arm's arithmetic leaves nothing behind any head, so the head-of-line half of #471's hypothesis is refuted while the load-gate half is confirmed by a positive control on the threshold alone. Lifting the threshold trades the stall for unbounded memory, so the fix is a retry bound, not a gate change - #149's dead-letter queue, after v6. The burn-down records the verdict, the release-note sentence it implies, and the one interim the owner may want: a warning when the gate latches with nothing retiring. Claude-Session: 460f7df9-dcc2-4b00-a9f9-62f3a2c6d5e4 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…high-failure-rate case Reviewer feedback on #487, relayed by the owner: under retry-forever the poison population only grows, workers saturate at about concurrency times retry delay over function duration, and only then does the gate latch - so any long-lived instance with any poison and no terminal handling stalls eventually. The burn-down's release-note line now says that, names the saturation figure and its inputs, ties it to the two upstream flat-counter reports, and records that the low-rate arm which measures it is queued. Claude-Session: 460f7df9-dcc2-4b00-a9f9-62f3a2c6d5e4 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…ll-under-always-failing-key
…or a busy pool - it is an eventual certainty Review feedback on #487 said the previous close understated the finding: under retry-forever the held poison population only grows, so the latch is not a possibility to be weighed but something any long-lived instance reaches. Checked against the existing logs first, then measured with a fourth arm rather than asserted. The reviewer's conclusion is CONFIRMED and their arithmetic is a correct UPPER BOUND; the ordering claim inside it is refuted, in the direction that makes the defect worse. THE ARITHMETIC, CHECKED AGAINST ARMS 1-3 BEFORE ADOPTING IT. parkedForRetry is not a property of the population - by Little's law it is retry-throughput times retryDelay. Three independent confirmations in logs already taken: parkedForRetry has median 135 and hard max 140 in all three arms while inShards ranges 549 to 17,103, a 31x population change with an unchanged parked count; the failure rate is 133 / 133.5 / 129.5 per second in those same three arms against a C/U ceiling of 140/s; and arm 1's predicted unparked of 549-140=409 is exactly its observed minimum. So unparked = P - (retry throughput * retryDelay) latch when unparked > targetAmountOfRecordsInFlight * loadingFactor and since retry throughput cannot exceed maxConcurrency/userFnDuration, the subtracted term is bounded by maxConcurrency*retryDelay/userFnDuration = 140 here. P only grows. The latch is therefore unavoidable, with a computable ceiling of 140+42 = 182 held records at these defaults. ARM 4 - failureFraction 0.01, ten minutes, everything else arm 1's. PREDICTED, written before the run: the gate oscillates early and then stops permanently (the discriminator, since arms 1 and 2 never oscillated at all); successes flow for minutes; parked tracks P then pins at 140; the latch arrives late but arrives, at P between 168 and 182, around 5.6-6.1 minutes. OUTCOME: the shape is confirmed, the numbers are not, and the miss matters. The gate DID oscillate - 264 false readings against exactly four in each of arms 1 and 2, all of those at startup - and then stopped: the last false is inShards=71 - parkedForRetry=29 = 42 vs target(14)*loadingFactor(3) =42, the boundary exactly, after which the final unbroken true run is 8m57s. Successes rose 992 -> 1,971 -> 2,946 -> 3,902 and then froze at 3,902 for the remaining nine minutes. But it latched at P=98, about 64 seconds in - not 182 at six minutes - because its retry throughput settled at 30.6/s, so its parked term was ~30 rather than 140. A SLOWER RETRY SERVICE LATCHES THE GATE SOONER, since fewer records are in back-off and more therefore read as workable. 182 is a ceiling, not an estimate. SATURATION IS NOT A PRECONDITION, and this is the correction worth carrying. At arm 4's latch the pool was doing 30.6 failures/s - about 3 of its 14 workers, 22% utilisation, against arm 1's 95%. The instance stopped fetching from the broker while 78% IDLE. What limits the retry cadence to ~3.2s per record against a static 1s delay is NOT measured - confirmed static, no retryDelayProvider is set and WorkContainer#computeRetryDueAt has no progressive backoff - and it is now the first arm in the scenario's list, because the latch point is a function of that number. HEAD-OF-LINE BLOCKING GAINS THE ROLE IT ACTUALLY HAS. It is still not what latches the gate (arm 2 latches with no ordering constraint; arm 1 held one record per key). It is what stops the residue draining afterwards: arm 4's inShards fell 103 -> 98 and then sat at exactly 98 for 6,201 consecutive evaluations, nothing retiring for nine minutes. Those 98 are ~40 poison plus ~58 healthy records queued behind poisoned heads on their own keys - eleven idle workers beside 58 deliverable records they were not allowed to reach. TWO MORE MECHANISMS RULED OUT WITH EVIDENCE, both things a reader would otherwise suspect. A dead poll thread (#477's class): the soak reported no finding in any of the four arms, so the poller is alive and paused. A lost wakeup from a stale pause cache: maybeWakeupPoller() is gated on !isSufficientlyLoaded(), so with the gate continuously true no wakeup is ever attempted - its "Found Poller paused" line appears zero times in every arm, and nothing can lose a wakeup never sent. AND A DISTINCTION THE MERGE MADE NECESSARY. #478 withdraws an eager-sync stall as "a bound the processor count crosses, not a defect". Do not collapse this note into that one: that symptom DRAINED and needed a processor-count cap to appear, while this one never drains and is reached at 22% utilisation with no cap at all. The failureFraction=0 arm is withdrawn as superseded - arm 4 answers what it was aimed at with a live workload rather than an empty one. Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xoi3HYae8pjsEatuNFKieD
…turation is not a precondition #487's fourth arm, at a low poison rate with the prediction written first, latched the gate with three of fourteen workers busy. The parked term the gate subtracts is throughput times retry delay, not a population property, so the unparked count crosses the threshold as soon as the poison population outgrows what the retry service holds in back-off - and a slower retry service latches sooner. The earlier line said saturation had to come first; it was wrong, and this replaces it with what was measured, plus the one thing the arms left open: the retry cadence is about three times the configured delay. Claude-Session: 460f7df9-dcc2-4b00-a9f9-62f3a2c6d5e4 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
|
@claude review this |
|
Claude finished @astubbs's task in 3m 7s —— View job Review
This is a test/diagnostics-only PR (no Overall: solid. The investigative methodology (four single-variable arms, a positive control that flips the outcome, a derived latch-point bound that's then checked against measurement) is exactly what A few things worth a look, none blocking:
No test-weakening, no scope creep beyond what the PR body describes, and the two |
…ll-under-always-failing-key
The commit-response-timeout note handed the load-gate question to the load-gate note after three arms. A fourth arm at a low poison rate ran afterwards and changed the model - the gate latches on an idle instance at a computable threshold, saturation is not a precondition - so the handoff sentence now says so, and no longer calls what remains a product decision. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xoi3HYae8pjsEatuNFKieD
|
@claude review this |
|
Claude finished @astubbs's task in 4m 0s —— View job Review
Confirmed test/diagnostics-only: no changes under any module's Verified, not just read:
Two things worth a second look, neither blocking:
Not a finding: No test-weakening, no product-code changes, and the two |
…ll-under-always-failing-key
…namic factor cannot grow out of the latch The review on #487 asked twice for the same clause: `140 + 42 = 182` sums two different units, and a reader who takes 42 for a second measured population re-derives the whole thing before believing it. So say what it is - the gate's own threshold term, `targetAmountOfRecordsInFlight(14) * loadingFactor(3)`, read off the arms' gate lines. Arm 1 latched against 28, because `DynamicLoadFactor` starts at 2 and had not stepped 0.8s in, which is the same fact from the other side. Naming `loadingFactor` as dynamic raises the question the note did not answer: if the threshold term can climb toward `DEFAULT_MAX_LOADING_FACTOR`, the gate would unlatch on its own and the "eventual certainty" would be an eventual inconvenience. It cannot, and the reason is in the source rather than in these runs. `checkPipelinePressure()` steps the factor only when `isPoolQueueLow() && lastWorkRequestWasFulfilled`, and `retrieveAndDistributeNewWork` sets that second term as `gotWorkCount >= delta`, where `delta` is the shortfall of dispatched records against the loaded target. Once the shards have nothing selectable left, every pass hands back less than the shortfall, the flag stays false, and the factor is pinned wherever it stood when the latch arrived. That correction replaces a wrong reason with the right one. The note previously said the factor barely moved "since the pool was never starved" - which is false of arm 4, whose pool sat at 22% utilisation with an empty executor queue, so `isPoolQueueLow()` read true throughout. It is the fulfilment term that holds the factor down, not the pressure term, and arm 4 is the arm that proves it: `loadingFactor(3)` unchanged across an unbroken `true` run of nearly nine minutes with eleven workers idle. The soak scenario's `Calibration status` block carries the naming clause too, since it states the same arithmetic, and points at the note for the step-up argument rather than restating it. No product code changes, and no measurement is revised - both arms' recorded numbers are unchanged. Co-Authored-By: Claude Opus <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xoi3HYae8pjsEatuNFKieD
Review responseBoth automated reviews (the 00:47 one and the 01:03 one at 1.
|
…y-key removal are on master The intake-stall measurement landed with the retry-cadence arm and the interim latch warning as what remains, and its box is ticked. The last by-key shard removal the #483 sweep reported is fixed on master, red first with an ablation arm per leg of the guard. The known unknown that named it is struck through with the outcome, and the retry queue's same-shaped removal is recorded as that queue's keying model rather than a remaining instance. Claude-Session: 460f7df9-dcc2-4b00-a9f9-62f3a2c6d5e4 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
… named, and five stale lines are corrected Owner's decisions, 2026-09-09: the merge queue is closed as of today, with later finds 0.6.0.x unless data loss on a default configuration; the poisoned-transaction wedge is the second named exception beside #44, and the release-note draft now carries it; the gate-latch warning #487 argued for is v6-sized and joins tier 1 as the last item; the upstream flat-counter reporters are not asked. Corrections from the owner's read of the note: the release page body is posted by hand on the day with gh release edit, because release.yml's exact heading match misses the unreleased heading on master - so #199 follows the tag rather than gating it, and the two lines that said the workflow already publishes the curated section are fixed; the #468 line no longer asks the reader to check a PR body for two by-key removals that #468 dismissed and #492 fixed; the vetting sweep's opening claim that the quarantine registry is non-empty is struck as the sweep's dated reading; and the disposition list names the three deferred bug notes it omitted. Claude-Session: 460f7df9-dcc2-4b00-a9f9-62f3a2c6d5e4 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…a heading gets its blank line The automated review read "Merged" after "joins tier 1 as its own item" as the gate-latch warning having merged; it was #487 that merged, and the sentence now says so. The data-loss heading gains the blank line every other section boundary has. Claude-Session: 460f7df9-dcc2-4b00-a9f9-62f3a2c6d5e4 Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
This branch is docs-only - two markdown files differ from master - so neither red it drew can be its own. Both are recorded against the notes that own them, and reading each occurrence turned up a correction the note needed. INTEGRATION TESTS, on this PR's own head. The shape in ci-broker-container-exit-126-is-undiagnosable.md, exactly: one class fell slowly at the container-start timeout and every other broker class fell in milliseconds with NoClassDefFoundError on BrokerIntegrationTest. Codecov renders that as "20 Tests Failed"; it is one failure. What is new is that the cause was in the log all along. Testcontainers prints the failed container's own output at GenericContainer#tryStart, one line below the "Wait strategy failed" line the note's signature block quotes, and it reads "sh: /tmp/testcontainers_start.sh: Text file busy" - ETXTBSY, exec refused because the starter script was still open for writing. The container command waits for that script to EXIST and then executes it, so a file the daemon has created but not finished extracting is executable-shaped and not executable, and the shell reports the refusal as exit 126. A Testcontainers start race, widened by a busy runner; nothing in the product, the image or the Kafka configuration. Refetching #347's 2026-08-25 job, the run this note was written from, shows the identical two lines. So the note's premise under item 1 - "the container's stdout is nowhere in the job log" - was false of its own founding evidence. The instrument was fine; the triage stopped one line short. That correction is proposed in the note's vetting marker rather than applied, because the note's impact is misdirection and those are the owner's to close. CHAOS PAIN SUITE 4/4, on #495 - also docs-only, one README paragraph. ChaosChurnStormIT NO_PROGRESS at 96632/100000 for 30s against a 30s bound, seed 3717713223451201639. It goes in test-no-progress-window-may-not-transfer-to-w1.md as one appended row, with the part that makes it worth having: the fleet KEPT CONSUMING, reaching 99569 by the settle summary, so the outstanding count fell from 3368 to 431 - inside the TAIL_SLACK of 500. That is the "drains" branch of the deciding experiment the note states. It is the weak form and the row says so: no recovery diagnostic, so the counter compared is the ledger's rather than the probe's, and the conductor's churn ended 10s after the firing, so it is recovery-once-churn-stops. The bigger finding is that the deciding experiment had already been answered twice and this note never took delivery. test-857-churn-storm-async-stalls.md drained six for six on seed 9086872209853284830 with the diagnostic engaged, and its 2026-09-08 sighting drained seed 5650361238717170909 from 93487 to 101070/100000 at an outstanding count of 6513 - larger than every row in the table. That sighting says outright that this note owns the question; the pointer was written and nobody followed it. Proposed in the vetting marker for the same reason as above. RULED OUT, with a control arm rather than an argument. The six merges that landed on master today - #480, #487, #488, #491, #492 and #493 - are the obvious suspects for a chaos red, and #491 does touch ProgressProbe.java. Its diff does not touch the NO_PROGRESS path at all - it adds the UNCOMMITTED_COMPLETIONS detector, edits javadoc, and refactors the finding sink - and the same test PASSED on two heads that carry every one of those merges, four and six minutes either side of the failing run. A deterministic regression is excluded; a rate change is not, and one failure could not establish one. Nothing quarantined. The container fault has no test to quarantine and the exit-126 note says a re-run is the correct response there. The chaos firing has no rate that rule 1 would accept, and docs/quarantined-tests.md is empty - which is the state to preserve. Co-Authored-By: Claude Opus <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xoi3HYae8pjsEatuNFKieD
… it turned up THE CONFIRMATION. One run of CommitResponseTimeoutSoakIT's KEY arm against the shipped WARN - same seed 3747722682837130843, failureFraction 0.5, same workstation, shortened to PT4M because the latch arrives in the first second and this is a confirmation rather than a re-derivation, with the gate logger at info so the report is visible without the per-tick DEBUG equation. The arm reproduced: succeeded=451, which is arm 1's number and #471's thirty-minute number. Exactly ONE WARN in the whole run, about ten seconds after the banner - a hundred passes at the latched cadence #487 measured, as designed - and no clear line, because the latch never cleared. Its operands agree with the arms that derived them: inShards=549 (arm 1's pinned population), parkedForRetry=138 (inside the band arms 1-3 measured), all twenty partitions paused. Written into the scenario's own Calibration status block, which owns these numbers. THE SIGHTING. The Lincheck lane cancelled at its 20-minute budget twice on this branch, with every other job in both runs green. It is the lane, not the branch, and the control arm is a documentation-only branch that succeeded at 19m58s the same afternoon - it cannot change what the model checker explores. A third branch took 12m55s. Against a recorded 7m42s baseline, which side of the bound a branch lands on is the runner it drew. Also ruled out directly: no harness in the lane reaches the changed code - WorkManagerLincheckTest's operations are handleFutureResult and the revoke/reassign pair, and no harness mentions the intake gate. Recorded in the lane's own note because a CI log expires and the evidence that a red was infrastructure does not survive it. The decision it needs - a larger budget, fewer iterations on the arm that dominates the wall clock, or a split - is left open there. Co-Authored-By: Claude Opus <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xoi3HYae8pjsEatuNFKieD
… not "any poison" Owner review point, 2026-09-09. "Retry-forever plus any poison at all" was inherited from #487 and overstates the result. A SINGLE record that never succeeds does not latch the gate and cannot. The gate is inShards minus parkedForRetry against target times loading factor: one held record, minus one parked while it waits out its back-off, is nowhere near a threshold of tens. Its offset map encodes a single gap compactly, the commit sits below it, and the instance runs indefinitely with that record retrying beneath a healthy stream that keeps retiring. What latches the gate is a non-zero FRACTION of a live stream that never succeeds, and both properties that make it inevitable are about the population rather than any one record: healthy records retire and leave the shards, these do not, so their share of what is held rises monotonically while the stream keeps arriving, and the parked term subtracted from it is bounded by throughput rather than population. "Any fraction" is the correct claim and still a strong one - the measured arm reached it at 1%. Corrected in place in the note, which is the live record rather than a dated one, with a dated section saying what was corrected and why. #487's own text on master is left alone; the soak scenario's Calibration status block carries the correction as a note against its verdict rather than a rewrite of it, because a dated record of runs may not be edited to match today's reading - none of its measurements change, only the claim drawn from them narrows. Every arm it ran was a fraction (0.5, then 0.01). Co-Authored-By: Claude Opus <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xoi3HYae8pjsEatuNFKieD
… nothing retiring (#497) #487 measured the record-intake stall and left the cheap mitigation written up but not built. An instance that retries forever, while a non-zero fraction of its stream never succeeds, ends with every partition paused, its workers churning on records that never retire, and nothing whatsoever in the log - the state was exported only as the pc.partitions.paused gauge. In 487's low-poison arm the instance reached it with eleven of its fourteen workers idle, looking healthy from outside. This is that mitigation: one observation and one log line. IT TAKES A FRACTION, NOT ONE BAD RECORD. A single record that never succeeds is one held minus one parked against a threshold of tens, so it never crosses; its offset map encodes one gap compactly, the commit sits below it, and the instance runs indefinitely with that record retrying underneath a healthy stream. What latches the gate is the share of held records that never leave rising while the stream keeps arriving - healthy ones retire, these do not - and the term subtracted from that share is bounded by throughput rather than by population, so it crosses eventually. 487's own wording, "retry-forever and any poison at all", overstates it and is corrected here; its measurements are untouched, because every arm it ran was a fraction. NO SEMANTIC CHANGE. The gate's decision, the poller's pausing, the retry service and every counter are untouched. isSufficientlyLoaded() returns exactly what it returned before, and the broker-poll thread's route into it through shouldThrottle() observes nothing at all. The control loop's once-per-pass call in maybeWakeupPoller becomes isSufficientlyLoadedReportingLatch(pausedPartitions), which takes ONE reading of the shards and uses it for both the wakeup decision and the report - so the report can never print an equation the decision was not made on, the same trap ShardManager#getWorkableRecords already exists to close for the DEBUG line. CONSECUTIVE PASSES, NOT ELAPSED TIME. Both operands were already there: the gate's own reading, and RecordPopulation#getRetiredTotal, which is monotonic and is exactly "no record left a shard, by any route". Nothing new is counted and no clock is read. A timing bound would be a threshold argument nobody can win. WHY A HUNDRED, AND WHERE IT IS WRONG. Chosen from the loop's two cadences rather than a target wall-clock time, because when nothing retires the pass rate differs by two orders of magnitude between the state being reported and the state that must not be. Latched, every held record fails on every attempt, so results arrive in the mailbox continuously and each pass returns at once - 487's low-poison arm measured 6,201 gate evaluations across a nine-minute unbroken latch, about 87ms a pass, so a hundred passes is about nine seconds, and the state is permanent so a few seconds late costs nothing. Healthy inside a long user function, the mailbox is empty and each pass blocks for the commit interval, five seconds by default - so a hundred passes is over eight minutes in which not one record anywhere retired. The pass count gives the healthy case roughly fifty times the grace it gives the latch, which no single elapsed-time bound can do. PC puts no ceiling on a user function, which is why the report is a WARN and not an exception. Eight minutes is the best case, not the bound, which the first draft of the derivation got wrong and two review passes caught between them. getTimeToBlockFor has two branches. On one, the healthy bound is the commit interval - and under the transactional commit mode that default is two orders of magnitude shorter, so the grace collapses to the same order as the latched cadence. On the other, taken whenever dispatch sits below full concurrency (the ordinary state under KEY or PARTITION ordering with fewer active keys than maxConcurrency allows) and any record is in retry back-off, the pass blocks for the retry delay instead - a one-second cadence on stock defaults, about a hundred seconds of grace rather than eight minutes, with no unusual configuration at all. Both are now named where the constant is defined and in the note, the second marked as a static trace of the two branches rather than a measured arm. The trigger is deliberately NOT narrowed on parkedForRetry, which would separate a latched instance from a merely-slow one: the line already prints that operand so a human can separate them, and changing what fires wants its own measured arm, because a transient zero in the parked count would suppress a real latch. Left open in the note with both candidates. ONCE, THEN QUIET, AND THE CLEAR SAYS WHICH CLEAR IT WAS - WITHOUT CLAIMING RECOVERY. The WARN fires on the pass that reaches the count and not again, naming inShards, parkedForRetry, workable, target times loading factor and the paused-partition count as of the last poll - scalars only, so the line cannot be truncated past the point where it stops identifying the event. The clear is two statements, because the two ways it clears are not the same news: a gate that has gone unloaded means the poller can fetch again and says so, while records leaving the shards under a still-loaded gate leaves the poller paused. Either clear re-arms the WARN. Neither says processing recovered, and that is the correction review forced. The trigger reads getRetiredTotal, which rises when a record leaves a shard by ANY route - success, revocation, or a stale container being swept. Right for the trigger: a revocation really does drain the shards and unlatch the gate. Wrong for a message, because revoking a stalled instance is exactly what an operator or the group coordinator does TO one, so "intake has resumed" would have been an all-clear delivered at the moment somebody was intervening, with nothing having succeeded. Both lines now report what is measured and name the three routes, pinned by an arm that latches, revokes with zero successes, and asserts the clear does not claim recovery. THREAD MODEL. The three counters are written and read only by observeLoadGateLatch, whose one production caller is maybeWakeupPoller on the control thread. They deliberately carry no @ThreadConfined and no owning-thread assertion, which is this repo's usual pairing: the harness drives controlLoop from more than one thread inside a single test, so the assertion would fire on a legitimate caller, and a diagnostic that can kill the consumer is a worse trade than one that can miscount. No decision reads these fields. The ledger entry points at the javadoc rather than restating it. RED FIRST, ON EVERY CLAUSE. Four sabotage arms, each restored. With the report removed entirely - master's behaviour - the two positive tests fail on their assertions rather than on compilation, and the control test still passes, which is what a negative control should do. Dropping only the "nothing retired" clause flips exactly the other two: the healthy-instance test sees a WARN naming 300 held records, and the recovery line never arrives. Deleting the count reset from the recovery branch fires the second WARN one pass after the recovery instead of a hundred, which the re-arm arm now catches. Dropping the honest tail from the all-clear branch reddens the revocation arm. CONFIRMED AGAINST A REAL BROKER. One run of CommitResponseTimeoutSoakIT's KEY arm at 487's seed and failure fraction reproduced its number exactly (succeeded=451) and emitted exactly one WARN, about ten seconds in - the hundred passes at the measured latched cadence - with operands agreeing with the arms that derived them: inShards=549, parkedForRetry=138, all twenty partitions paused. No clear line in four minutes, because the latch never clears. Recorded in the scenario's own Calibration status block. REVIEWED FIVE WAYS. Codex found five things; four taken. A local simplification pass found one - sharing the operand tail across the log statements - rejected, because it would trade per-field arguments for one opaque token and make each line unsearchable in source. A local review pass found three more, all fixed: the false all-clear above; a mutant that survived every test, deleting the count reset from the recovery branch while leaving the flag reset, which the re-arm arm could not see because it asserted a cumulative count after a full batch rather than the boundary; and the wrong bound in the threshold derivation. The automated review on the PR then found the half of the threshold correction the local pass had missed - the retry-delay branch above. One finding is left open for the owner rather than fixed: the observation runs only while the state is RUNNING, so a graceful shutdown hanging in DRAINING - a plausible shape of this defect - is never reported, and widening that guard is a decision about what the diagnostic is for. The Codex finding rejected, with the argument recorded as a dated cleared suspicion: moving the observation after the mailbox drain. The observation sits at the same point in every pass, so the window between two readings spans a whole pass, drain included, and no retirement falls between them; what the position costs is one pass of latency out of a hundred, self-correcting on the next pass. Both remedies are worse - a second gate read per pass buys it with an O(n) fair-lock acquisition on the hottest path and reintroduces the equation-the-decision-was-not-made-on trap, and gating on an empty mailbox would suppress the WARN in exactly the state it exists for, because the latched workload's failure results flow through that same mailbox continuously. Also here: BrokerPollSystem gains getPausedPartitionCountForBackPressure(), with isSubscriptionsPausedForBackPressure() expressed in terms of it; an unposted draft response for the issue mirror, scoped to append to its Fork status rather than replace it; and a sighting of the Lincheck lane cancelling at its twenty-minute budget - this branch on both sides of the bound, three cancels and one 19m57s pass with nothing changed that the lane can see. That sighting is now corroboration rather than a diagnosis of its own: #499 established the cause across every branch on the day and raised the cap, and the note keeps 499's section whole with this branch's data pointing at it. Serves #119 (confluentinc#857); closes neither. The fix that bounds the failures rather than the buffer is #149's dead-letter queue, and the note stays open on it. Co-authored-by: Claude Opus (1M context) <noreply@anthropic.com>
The gate-latch warning is on master, which ticks the last tier 1 box. The line records what the reviews settled: the trigger is consecutive passes rather than elapsed time, the soak confirmation at #487's seed, the two design calls left for the owner, and the correction that the latch takes a fraction of never-succeeding records rather than one. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xoi3HYae8pjsEatuNFKieD
…ut it (#475) 0.6.0.0 is a bugs-only stability release, and it is overdue: the fork has carried the fixes for upstream's most-reported defects for months while the release waited on features. This note is the source of truth for cutting it - the owner's decisions, the merge queue from those decisions to the tag, every open question, and the checks that make the published artefacts true on the day. #197 is the tracking handle and its body points here; nothing is maintained on the issue. THE DECISIONS, 2026-09-07 and confirmed since. The bar is the stability release and nothing else; Streams and Connect move to the next-0x horizon in the roadmap data; the producer-recovery stack is outside v6. The release claim carries two named exceptions rather than waiting on them: the transactional revoke wait (#44, bounded since #466, not yet declined) and, from 2026-09-09, the poisoned-transaction wedge, both in the transactional producer mode only. The merge queue closed on 2026-09-09; later finds are 0.6.0.x unless they are data loss on a default configuration. THE BURN-DOWN, recorded as each merge landed. Tier 1, the self-contained fixes, is complete: the last two to join were the batchSize bound (#496) and the gate-latch warning (#497), both decided v6-sized on the day the queue closed. Tier 3, the plumbing, has the changelog section finalised as the release notes and the claim amended (#498) and the release page body posted verbatim from CHANGELOG.md by release.yml (#501, closing #199); what remains is the tag-day checks, the drafted issue responses, and the tag. A can-follow list names what is deliberately not v6. WHAT THE RELEASE NOTE SAYS ABOUT THE confluentinc#857 FAMILY, each line with the PR that settled it: the revoke-path deadlock proven by control arm, the eager stall withdrawn as a timing bound that flips with the processor count, the fifth item measured as the consumer-group protocol under churn rather than PC, the poller death fixed, the instance-stall sightings classified as worker saturation from the load side. The intake stall #471 found has its verdict from #487: the record-intake load gate is what stops intake, head-of-line blocking is not why, and any instance that retries forever while a fraction of its stream never succeeds latches eventually at a computable threshold, idle or not. There is no gate fix; the fix bounds the failures (#149's dead-letter queue), and until then #497 makes the state visible. One arm stays unattributed and is named as such. DATA LOSS AND DUPLICATES: the bug-162 replay branch refuted and the false truncation warning fixed (#494, closing #162). KNOWN UNKNOWNS, split in two so nothing is papered over: what is still unknown at the cut - the shard half of the per-shard liveness blind spot, the flake rows kept open with reasons, the maturity claim - and, under its own heading, the unknowns made known on 2026-09-08 and how each was settled. TAG-DAY CHECKS, folded in from the retired blockers note: master green with the lanes known to lie named, the churn scenario's no-progress window settled by replay and widened in #499 with the rebalance-dwell bound named as that class's survivor, the Lincheck lane's timeout raised against runner-speed variance, the rename named in both groupId and packages, the README's trademark wording claiming nothing it does not have (#495), and the changelog section as the release notes since #498, posted as the release body by release.yml since #501. ONE CHANGELOG EDIT, on the owner's decision of 2026-09-10: the "size of this release" table of merged-PR and line counts is removed. Measured on a branch, carrying its own re-measure instruction, stale from the next merge on; the notes make their claim through the fixes they name. Also here: a ci- note from this PR's own last review round - the file-refs gate reads a token as a path only with two segments, so the changelog rename left this branch-only note naming the old file with nothing to go red, and the note records the allow-list that would close it; the vetting sweep's reading and every open bug note's disposition, moved into the ranking note where the tiers override them; a dated survey of upstream items with no fix and no response as its own deferred note; the refactoring registry's codec entry corrected for what #480 did and did not change; and the confluentinc#546 manifest entry marked merged. Two notes retired with their content migrated: the blockers register and the merge-order plan for a far larger v6. The question this note began as, "when is v6 good enough?", was answered on 2026-09-08 and the file renamed. Serves #197; closes nothing. The tracker closes when the tag is cut. Co-authored-by: Claude Fable 5.1 (1M context) <noreply@anthropic.com>
…data pattern, not a count of failing records The Known limitations bullet said any instance holding records that never succeed will eventually stop fetching. That is wrong on its own source: #497 says a single failing record, or a bounded set, never latches the intake gate, and the soak in #487 measured the gate at a 42-record target where a larger messageBufferSize removed it. What bounds the system is the offset-encoding contract: the committed offset is the lowest incomplete one, the completions above it are encoded into 4 KB of commit metadata, and regular patterns - a contiguous run of failures, or every second offset failing - compress to almost nothing, so only an irregular scatter of permanent failures among successes outgrows it. The bullet now says that, keeps the gate as the separate and configuration- dependent thing it is, and says the gate is slated for removal by the direct-pull engine (#361). The published release page is edited to match. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019XS64Xttx4vF5datYh7fmk
…data pattern, not a count of failing records The Known limitations bullet said any instance holding records that never succeed will eventually stop fetching. That is wrong on its own source: #497 says a single failing record, or a bounded set, never latches the intake gate, and the soak in #487 measured the gate at a 42-record target where a larger messageBufferSize raised the ceiling. What bounds the system is the offset-encoding contract: the committed offset is the lowest incomplete one, the completions above it are encoded into 4 KB of commit metadata, and regular patterns - a contiguous run of failures, or every second offset failing - compress to almost nothing, so only an irregular scatter of permanent failures among successes outgrows it. The bullet now says that, as a nested list, and keeps the intake gate as the separate, configuration-dependent thing it is. The published release page is edited to match. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_019XS64Xttx4vF5datYh7fmk
…inds no commit-response timeout (#518) Records three arms of CommitResponseTimeoutSoakIT on the released code (v0.6.0.0 plus two docs commits, 2e6f13e) in the scenario's Calibration status and in docs/inflight/bug-177-commit-response-timeout-unreproduced.md. No product or test code changes - the record is the deliverable. The question since #471: does a BARE "Timeout waiting for commit response" - the poller wedged but alive, the one mechanism nobody has characterised - ever appear on the released code when the workload is arranged so the intake gate does not latch and commits keep happening for the whole soak. #471's two 30-minute runs could not answer it: the retry-forever poison latched the record-intake gate inside the first minute and no commit was attempted after that, which #487 measured and #497 made loud. Common to all arms: seed 3747722682837130843, failureFraction 0.5, PERIODIC_CONSUMER_SYNC at 1s, 1000 keys over 20 partitions, maxConcurrency 14, 100ms user function, 1000 records every 20s, retry-forever; Testcontainers confluentinc/cp-kafka:7.9.0 on Docker; a Linux x86_64 workstation with every JVM pinned to 8 processors; no other Maven JVM on the box at the start of any arm; -Dpc.loadgate.log.level=info and -Dsoak.progressInterval=PT30S. Commit activity was read from the broker's __consumer_offsets with the broker's own formatter, because every commit-path line in the product is DEBUG. Arm A - control, unchanged, 6 minutes. The latch reproduces on the released code to the number: succeeded frozen at 451 from the first 30s sample, the #497 WARN 13.2s after the banner (inShards=549 parkedForRetry=140 workable=409 vs 42, pausedPartitions=20), every partition paused at all 11 samples, failed 49,710 (138/s). The broker saw 8 commit instants, all inside the first 7.3s, and none in the remaining 5m50s. Load average 2.93 -> 1.54. Arm B - 30 minutes, -Dsoak.messageBufferSize=20000, KEY, everything else the reporter's. No findings. The gate held open for 7m06s (loaded=false, zero paused partitions, at every sample to inShards=20,103), then latched at inShards=20993 vs target(14)*loadingFactor(1429)=20006 and stayed latched for the remaining 22m54s - so the #487 arm-3 value buys about seven minutes of this producer, not thirty, because under KEY nearly everything produced is held. But the assertion had already stopped being falsifiable at 3m06s, four minutes before the latch, and the gate is not why: successes went 680, 784, 869, 884, 893, 895, 897 and froze with the gate open, and the broker's last commit landed 185.7s after the first (27 commit instants). Under KEY each key retires records until its first poisoned one and never again, so permanent poison bounds successes by the key space, and dirty is derived from completions, so once successes stop commitAndWait is never entered. No buffer size changes that under KEY. Succeeded 897, failed 250,067 (139/s). Load average 1.17 -> 0.13. Arm C - 30 minutes, -Dsoak.messageBufferSize=100000 -Dsoak.ordering=UNORDERED: the only combination of the existing knobs that holds both conditions. UNORDERED lifts the head-of-line bound so healthy records keep retiring; 100,000 clears the ~90,000 a thirty-minute producer can have held. Both held for the whole run: loaded=false and zero paused partitions at all 59 samples, inShards peaking at 51,949 against 100,002, no WARN, no back-pressure line; succeeded 37,641 and still rising at the end, failed 213,533 (119/s); a commit in every one of the thirty minutes - 619 commit instants, the last 1794.8s after the first, offset-map metadata growing to 768 base64 chars per partition. No findings: no bare timeout, no poller death, no thread dump written. Load average 0.12 -> 1.46. What it establishes, as a rate under conditions: on v0.6.0.0, one instance, a quiet box, zero bare commit-response timeouts and zero poller deaths across 654 acknowledged commit instants, 619 of them in thirty minutes of a continuously exercised commit path. One run at one shape; it says the shape did not reproduce once, not that it cannot. What it does not establish: anything about #175's own configuration (128 partitions, concurrency 64, a user function of minutes - still unrun), a loaded box, a rebalance, more than one instance, durations past thirty minutes, or KEY ordering with the commit path alive - which on this evidence needs the per-attempt-failure arm rather than a buffer size. Arm C departs from #177's reporter's ordering deliberately, and says so in both records. No new knobs were added: -Dsoak.duration, -Dsoak.messageBufferSize and -Dsoak.ordering already existed. The only thing genuinely missing from the harness - evidence that commits were being attempted - was taken from the broker rather than by adding a log line, and the command is recorded in the javadoc so the next run can do the same. Co-authored-by: Claude Fable 5.1 (1M context) <noreply@anthropic.com>
Serves #175 (confluentinc#809) and #119 (confluentinc#857).
Description
#471's soak found that a single instance under
KEYordering, with recordsthat throw on every attempt, stops taking new work inside the first minute: successes freeze while
the failure rate holds exactly constant. It named one untested candidate - the record-intake load
gate, on the theory that
inShardscounts records queued behind a blocked shard head - and said onerun reading the gate's own DEBUG line would settle it. Three ran. The hypothesis is half right, and
the half that is wrong is the mechanism it named.
The arms - each differing from the first by exactly one term
Six minutes rather than thirty (ten for arm 4, whose whole question is when), because the stall is reached in the first second. Same seed
3747722682837130843,failureFraction0.5, 1000 keys over 20 partitions,maxConcurrency14, 100msuser function, 1000 records every 20s, the suite's Testcontainers Kafka, macOS arm64 workstation at
5-13 load average.
WorkManagerat DEBUG throughout, and the logging config that carried it verifiedin the log by
OnConsoleStatusListenerrather than assumed - which mattered, because the gate'slogger sits under a bare package pin that
-Dpc.log.level=debugdoes not reach.trueand stays; all partitions paused; failures at ~140/strueon 37,356 of 37,360 evaluations (the 4falseare the first 800ms). Latched atinShards=500 vs target(14)*loadingFactor(2)=28on the first fetch, 0.8s in, never unlatched. 20 partitions paused at every sample.UNORDEREDorderingtrueon 39,689 of 39,693, 20 partitions paused throughout, successes crawling 232 → 314 on records already in the buffer.messageBufferSize=20000(threshold 42 → 20,006)false, partitions unpaused, records keep arriving, successes keep risingfalseon all 35,652 evaluations, zero partitions paused at any sample,inShardsclimbing 549 → 17,103 with the producer, successes 897.failureFraction=0.01, 10 minfalsevs four in arms 1-2, all of those at startup) and stopped: lastfalseisinShards=71 - parkedForRetry=29 = 42 vs target(14)*loadingFactor(3)=42, the boundary exactly, then an unbrokentruerun of 8m57s. Successes 992 → 1,971 → 2,946 → 3,902, frozen for the last nine minutes. But it latched at P=98, ~64s in, not 182 at six minutes.Verdict
Arm 3 is the positive control, so the gate is what stops intake. Arm 2 kills the stated mechanism.
Arm 1's own arithmetic kills it more directly: the 549 records it held came from one burst over
1000 distinct keys, so there was at most one record per key and nothing was queued behind any
blocked head.
This is not a possibility to weigh - it is where every long-lived instance ends up
parkedForRetryis not a property of the population. By Little's law it isretry throughput × retryDelay, so withPpermanently-failing records held:Under retry-forever
Ponly grows while the subtracted term is bounded - retry throughput cannotexceed
maxConcurrency / userFunctionDuration, so the parked term cannot exceedmaxConcurrency × retryDelay / userFunctionDuration= 140 at these defaults. The latch istherefore an eventual certainty, at a computable ceiling of
140 + 42 = 182held records. The twoaddends are different units on purpose: 140 bounds the parked share, and the 42 is the gate's own
threshold term,
target(14) * loadingFactor(3), not a second measured population.loadingFactoris
DynamicLoadFactor#getCurrentFactor, which starts at 2 and steps up one at a time - arm 1 latchedagainst 28 before it had stepped at all - and it cannot grow back out of the latch, because
checkPipelinePressure()steps it only whenisPoolQueueLow() && lastWorkRequestWasFulfilledand thelatch is precisely what stops a work request being fulfilled.
Measured three ways before being adopted, on logs already taken:
parkedForRetryhas median 135and hard max 140 across arms 1-3 while
inShardsranges 549 → 17,103 (a 31x population change,unchanged parked count); the failure rate is 133 / 133.5 / 129.5 per second against a
C/Uceiling of140/s; and arm 1's predicted
unparkedof549 − 140 = 409is exactly its observed minimum.But 182 is a ceiling, not an estimate, and arm 4 beat it in the dangerous direction. It latched at
P=98because its retry throughput settled at 30.6/s, making its parked term ~30 rather than 140. Aslower retry service latches the gate sooner, since fewer records are in back-off and more therefore
read as workable.
Saturation is not a precondition - the instance stalls while idle. At arm 4's latch the pool was at
22% utilisation (3 of 14 workers) against arm 1's 95%. It stopped fetching from the broker while
78% idle, looking healthy the whole time. What limits the retry cadence to ~3.2s per record against a
static 1s delay (confirmed static: no
retryDelayProvider, no progressive backoff inWorkContainer#computeRetryDueAt) is not measured, and is now the first arm in the scenario'slist, because the latch point is a function of it.
This is the best explanation anyone has produced for confluentinc#809 and
confluentinc#833, whose reporter showed
pc_processed_records_totalflat acrossthe window their timeout fired in - which is this state, not a busy one.
Head-of-line blocking gains the role it actually has
Not what latches the gate (arm 2 latches with no ordering constraint at all). It is what stops the
residue draining afterwards: arm 4's
inShardsfell 103 → 98 and then sat at exactly 98 for 6,201consecutive evaluations, nothing retiring for nine minutes - ~40 poison plus ~58 healthy records
queued behind poisoned heads on their own keys. Eleven idle workers beside 58 deliverable records they
were not allowed to reach.
Why no gate change fixes it
Arm 3 lifts the intake bound and throughput still dies: successes doubled, then plateaued by minute
three while the held population climbed linearly. Raising the threshold converts a hard stall into an
unbounded-memory slow starve; "count only what is selectable" is worse, because under
KEY/PARTITIONordering at most one record per shard is ever selectable and a shard whose head is at a worker has
none. The distinguishing property is liveness of the shard head, not a count, and that is not decidable
from the shard's state. The fix has to bound the failures - #149's dead
letter queue, which
roadmap.yamlalready describes in exactly these terms. The interim is to make thelatch loud: it is exported only as
pc_num_paused_partitionsand logged nowhere.Ruled out with evidence, not by reading
Offset-encoding back pressure (absent from all four logs). A dead poll thread
(#477's class - the soak reported no finding in any arm; the poller is alive
and paused). A lost wakeup from a stale pause cache (
maybeWakeupPoller()is gated on!isSufficientlyLoaded(), so with the gate continuouslytrueno wakeup is ever attempted - itsFound Poller pausedline appears zero times in every arm). And this is not#478's withdrawn eager-sync stall: that one drained and needed a
processor-count cap to appear; this one never drains and needs no cap.
What is here
-Dpc.loadgate.log.level=debugraisesWorkManager'sper-tick gate line, on its own property, in
logback-test.xml(the file an integration or soak runactually reads) and in the integration profile, each defaulting to its own file's level.
frozen success count says in the same place why it froze.
-Dsoak.ordering,-Dsoak.messageBufferSize,-Dsoak.progressInterval.WorkManagerTest#theLoadGateCountsRecordsQueuedBehindABlockedKeyHeadAsWorkable- the accountinggap pinned as a characterisation, with its javadoc stating what it is not evidence of. Its last
assertion measures the second consumer of the same over-read:
drain()gates onisRecordsAwaitingProcessing(), which reads true with nothing selectable.docs/inflight/bug-119-load-gate-counts-blocked-work-as-available.md- the operator-visiblesymptom, the verdict, the design decision, and the sweep for the same shape elsewhere.
CommitResponseTimeoutSoakIT'sCalibration statusgains the three arms(its earlier block is now dated so the two read as a sequence), and
bug-177-commit-response-timeout-unreproduced.md's arm 1 is marked done and hands the gate questiononward, keeping confluentinc#809: Sporadic timeouts from ConsumerOffsetCommitter.CommitRequest #175's own question where it belongs.
No product code changes. The measurement says a gate change is the wrong fix, so none is proposed.
Checklist
docs/inflight/notes, and the two logging profilesdocs/features/- N/A - test instrumentation and a diagnostic flag; no user-facing featuredocs/inflight/working note started at the PR's first commit -bug-119-load-gate-counts-blocked-work-as-available.mdce-simplifyandce-code-reviewlocally - N/A - asking for@claude review thison the PR instead, which is the only route that can open inline threads