Skip to content

feat(core) astubbs#227: auto scaling / adaptive concurrency - the engine discovers its own admission target - #333

Draft
astubbs wants to merge 141 commits into
perf/engine-concurrencyfrom
feats/ideate-distributed-throttling
Draft

astubbs wants to merge 141 commits into
perf/engine-concurrencyfrom
feats/ideate-distributed-throttling

Conversation

@astubbs

@astubbs astubbs commented Aug 22, 2026 •

Copy link
Copy Markdown
Owner

depends on #335 - the atomic work claim
depends on #336 - the poller load gate by conservation
depends on #358 - dispatch parity counted, not timed
depends on #359 - residence time, the first latency metric
depends on #360 - virtual threads for the user function
depends on #361 - direct pull finished, scan made O(1)
depends on #362 - the bench harness and its results
depends on #363 - the campaign records and the integration branch
depends on #367 - the strategy-conversation notes this work inherits (the delta vote's WHY encoding, the shared-execution-resources design, the admission model)

Serves #227, and closes both halves of #311 along the way.

Base branch: this targets perf/engine-concurrency, not master. That branch is not in master
and has no PR of its own, and this work is grounded on the tree it produced - the direct-pull
engine, virtual threads, the conservation-derived load gate and the residence-time instrumentation
all changed the seam this feature governs. Targeting master would show that branch's entire history
as though it were part of this change. This cannot merge before its base does; retargeting to
master is a one-flag change if that is preferred.

Description

maxConcurrency is a compile-time guess about a runtime quantity. Too low silently strands
throughput nobody files an issue about; too high floods downstream services. The right value depends
on the data currently in the instance's assigned partitions, differs between instances in the same
group, and shifts with the workload - so no static number stays correct.

This adds an opt-in adaptive controller that discovers it. When enabled, the controller governs the
admission target - the records the engine allows in flight - from measured service time, outcome
signals and sustained achieved in-flight, stepping up while performance holds and contracting when it
degrades, always beneath an effective maximum. The thread pool stays fixed at the ceiling; admission
is the control variable, which is the one knob that works identically on a platform-thread pool and
under virtual threads.

The mode is off by default and this is staged deliberately. OBSERVE computes and reports every
decision without acting - it is the diagnostic an operator can turn on to see what concurrency the
engine would pick and why it is not moving, at zero behavioural change. ENFORCE lets it act.
Promoting one to the other is a restart, and the system property exists for the bench harness and CI
matrix but may select at most OBSERVE - an ambient JVM-wide flag must never be able to hand an
experimental controller the admission target of a production instance.

What the design review changed

A five-persona review plus an architecture pass ran against the plan before implementation and found
several things that would have shipped as bugs. Each of these is now a test, most of them
sabotage-proven red before restore:

  • The pinned load factor would have doubled the enforced target. The factor means keep the
    workers N deep in buffered work
    , which held only while pool size equalled the target. With the
    pool at the ceiling and a live target below it, target x factor records would have run, not
    buffered. Dispatch now consumes the target un-multiplied; the factor applies only to the poller
    gate's buffer arithmetic.
  • Starting unseeded at the effective maximum would have raised concurrency on opt-in - exactly
    the flooding shape the feature exists to prevent - because the ceiling substitutes upward when the
    user left maxConcurrency alone. The unseeded start is now the target today's configuration
    derives, and the controller ramps into the headroom from there.
  • A fast-rejecting overloaded downstream reads as improved latency, so a latency gradient alone
    grows the target and accelerates the flood. A failure-fraction inhibitor freezes growth
    independently of what latency says.
  • A contraction narrows the buffer and can manufacture its own starvation evidence, which would
    then freeze the controller in a one-way ratchet. A starved window below the ceiling earns a bounded
    probe up instead.
  • The drain release had nowhere to fire from: an edge action on the draining transition runs on
    the caller's closing thread while the tick is already gated off. It is a state-derived seam read, so
    a contracted target cannot stretch a drain past its timeout - including when close arrives before
    the control loop ever ran.
  • A cooperative rebalance fires callbacks on instances whose partitions did not move, so resetting
    on any callback would starve the controller of history in exactly the churning groups where
    per-instance adaptation matters most. The reset is gated on a real delta to this instance's
    assignment.

The control law, and what measurement settled

The math is Gradient2 with the upstream anti-drift fixes, ported from Netflix/concurrency-limits
(Apache-2.0, attributed per file) rather than taken as a dependency - the core stays
zero-dependency, and the upstream interfaces fit neither min-composed ceilings nor the sampling this
engine needs. Upstream ships no tests for that algorithm, so the port brought its own.

Two deviations are measured rather than preferred. The utilization term is the window median of
in-flight with a spread, not the maximum: under a skewed key distribution this engine sustains one
record in flight of a configured twenty-four while a maximum reads full width, so a maximum reports
health exactly when the engine is starved. And the port's own gate test settled a question the plan
left open - the ported law provably cannot descend when every sample it has ever seen comes from
an already-flooded operating point, so a bounded probe-down is implemented rather than deferred as a
known residual.

Scope

Core engine, pre-loaded-queue path only. The async engines follow their timing fix; the direct-pull
engine consumes no admission target at all and refuses the mode with a warning, as do the external
engines - refusals are loud, never silent. Instance-count recommendation, predictive scaling and any
distributed coordination substrate are explicitly out of scope.

Where ordering starves a workload, admission is not the binding constraint and no admission
controller can raise throughput. The controller's job there is to report rather than adapt against a
constraint it cannot move.

Same-defect sweep for #311

The defect class is rounding up to a multiple using the wrong operand. Every modulo in main code
was searched, untruncated: eight raw hits, of which one other is real ceiling arithmetic - the
buffer-size division in the module's load-factor construction - and it is a correct single-pair
ceiling division, not the wrong-third-variable shape. Single site, not a pattern.

Still to come

The adaptive benchmark arm, which lands on its own branch cut from the arrival-controlled harness -
the value claim is only measurable below saturation, and the roadmap stage is gated on that result
rather than on this PR. That is also why the roadmap entry here moves to in-progress rather than
implemented: that ladder reserves implemented for work proven in use.

Checklist

  • Docs updated - operator-facing README section (regenerated from the template) and the metrics table
  • User-facing feature documentation data added under docs/features/
  • Tests added/updated
  • Title & body reflect the final content of this PR

Stack position

This PR targets the perf/engine-concurrency integration branch, whose content the perf train is landing on master wagon by wagon (docs/inflight/branch-engine-concurrency-pr-stack.md is the wagon ledger). This PR retargets onto master once the train has landed; the depends-on list at the top is every cut wagon, kept current as the train moves.

astubbs and others added 30 commits August 17, 2026 23:23
Ideation document for the distributed rate limiting ask (#228,
mirror of confluentinc#24), grounded in the linked production report
confluentinc#766 (400K RPS in, 50K RPS downstream, 113 nodes on 100
partitions) and in a code-verified map of the dispatch seam, retry
deferral, backpressure, and packaging conventions.

46 raw ideas from five parallel framing passes were deduped, verified by
a fresh-context basis check (8/8 code spot-checks held), and cut to
seven survivors spanning five axes:

  1. Partition-share budget - distributed limiting with zero new
     dependencies (the consumer group itself as division substrate)
  2. Dispatch-seam throttle via availableAt deferral - shape, don't
     police; reuses the RetryQueue calendar for wakeups
  3. Little's Law controller - an RPS target becomes the in-flight
     target PC already enforces at dispatch AND poller
  4. The throttling constitution - liveness beats the limit; epoch
     fencing; two verified corrections to #228's own
     architecture argument (ShardKey reach, batch cost timing)
  5. Socket-and-plug packaging - SPI in core, bucket4j/Redis backends
     as opt-in artifacts per the -vertx/-reactor/-mutiny precedent
  6. Adaptive AIMD on rate - no configured number, no coordination
  7. Decorrelated jitter for the deterministic retry path -
     independently shippable

The rejection table records the roads not taken with reasons, notably:
assignor-userData lease allocation (feasible, deferred on adoption
friction), offset-metadata stigmergy (contends with a near-capacity
payload), and broker-side KIP-124 quotas (bytes-in is not requests-out).

Also re-checks the 2020 library shortlist from confluentinc#24:
bucket4j alive and grown (with a per-permit round-trip caveat),
ratelimitj dormant, resilience4j#350 still unresolved - evidence the
gap is real. Autopsy of the abandoned features/rate-limiting branch
(e9f49d3) included; it policed instead of deferring and never used
the distributed dependency it declared.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
The ideation doc names seven directions but no chosen one; this entry
records the two gating decisions (enforcement fork, who owns the number)
and the independently shippable jitter item, so a future session restarts
from the artifact instead of re-deriving it.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…nk branch recovery

Extends the throttling ideation with idea 8: one adaptive controller
(#227 / confluentinc#21, TCP-congestion-theory concurrency) fed by
latency gradient, failures, and a structured rate-limit exception the
user function throws to teach the engine per-service ceilings at runtime.
Capacity limits are discoverable; contractual quotas are not - explicit
ceilings and adaptive discovery compose by min(), not compete.

Also: idea 5 gains the curated strategy-menu framing (each strategy a
distinct situation, min-composition, enforcement not user-selectable);
docs/refactoring.md idea bank gains the three missing flow-control
branches (dynamic-concurrency-control @6f85eac41, auto-tuning-pressure
@f4aa09788, rate-limiting @e9f49d321) plus upstream draft PR
confluentinc#22; the inflight entry now records the standalone-vs-
self-scaling gating decision.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
Too low silently wastes headroom, too high floods the downstream
(confluentinc#766) - completes the case for adaptive control in idea 8.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
Split from next-distributed-throttling.md: the two efforts will reach
mergeable PRs at different times and carry distinct prototype trails.
Records the two staged dimensions - per-instance adaptive concurrency
first (no coordination substrate, design-to-optimal downstream feedback),
instance-count recommendation later (metric + PC-owned cool-down +
rebalance as acknowledgement, capped at partition count, HPA/KEDA as the
consumer). Ranked into next-candidates.md; idea 8 card cross-references.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
… scaling

Records the conflicting-recommendation problem (instances only see their
own performance) with candidate resolutions ranked simplest-first: local
delta-votes / headroom gauges that HPA aggregates natively, the assignor
userData leader channel if a single number is ever needed, median-never-
max if raw totals are aggregated anyway. Earmarks the partition-cap
interaction with per-topic functions (#254) and share groups
(KIP-932), and flags the STRATEGY.md positioning follow-up.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…al number

Each instance votes +1/0/-1 from its own state; infrastructure sums (or
HPA averages a headroom gauge). Post-rebalance cooldown per instance
doubles as the acknowledgement loop: act -> rebalance -> cool down ->
fresh votes. Assignor-leader coordination demoted to likely-unnecessary.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
Records what had only lived in conversation: plateau-cause-agnostic
control (the reason runtime discovery beats configuration), the vote
clamp rationale (bounded steps converge AIMD-style), dynamic cooldowns
as the rule with fixed values only as fallback floor, the resolved
relationship to rate limiting (ceilings as inputs, no substrate; shared
SPI, independent shipping), the depersonalised engine-layer positioning
argument for the strategy run, and the branch plan.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
Target problem gains the static-configuration crux (deploy-time guesses
for runtime-dependent quantities, wrong in both directions). Approach
gains the second half of the bet: the client is where ground truth
lives, so the engine measures and decides what configuration used to
guess - vs black-box lag-based autoscalers. New Self-tuning track
(priority raised 2026-08-18) with the client-side-vantage moat argument.
The achieved-fan-out-vs-configured-max metric becomes discovered
concurrency vs sustainable ceiling, configured-max kept as interim
proxy. One-liner unchanged. Also earmarks the #255 pairing:
auto-scaling under a Kafka Streams topology.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…ches

Self-contained brief for a future session: one PR extracting stranded
planning docs from unmerged branches onto master, sequenced on
#305's audit, with provenance headers, three-shape triage, and a
per-branch verdict as the definition of done.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…ension 1)

Requirements-only unified plan from the ce-brainstorm session: the
adaptive controller governs the admission target under min-composed
ceilings, core engine first, opt-in with an optional seed. Records the
seven session-settled decisions with provenance, the rip-out criteria as
testable guardrails, the DynamicLoadFactor coexistence requirement, and
the corrected #155 status (stall fixed upstream; residual
mis-firing warning). CONCEPTS.md gains the admission-target term.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
Merging master brought #312 (the idea-extraction sweep this
branch's handoff note commissioned - executed by another session) and
#305 (the branch audit). Reconciliations git cannot do: the
sweep handoff note retires as satisfied work; the orphans note's
observability guess for feature/auto-tuning-pressure is corrected to
the #227 flow-control family (settled by reading its commits);
the idea-bank group gains its manifest cross-ref (sweep-2023-long-tail)
per #305's convention.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…entry

Review finding on #308: CONCEPTS.md entries
stand alone - no file paths, class names or config values - so the
glossary does not rot when the surface changes. Rephrased in prose.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…s - the NO_PROGRESS arm, twice in one night

Two ProgressProbe fleet-stall kills of ChaosChurnStormIT four hours
apart on unrelated docs-only branches (#310 and #308),
seeds 3086917415748208232 and 8603691233664838594 captured before log
expiry, each with its passing revoke-under-work control arm. Records
the truncated-console-log trap that misattributed the first of them to
the passing cooperative test (the circulated seed 4087023100803854645
is the control arm - do not replay it expecting a failure) and the
run-logs-archive retrieval route that recovers what --job log cuts off.
Also logs the untracked ReactorBatchTest.simpleBatchTest awaitility
flake seen on #308's CI at d930ca9.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…n with git -C

Third recorded occurrence of the class (2026-08-06 checkout/rebase,
2026-08-10 merges on local master, 2026-08-18 merge on the rename
branch); the middle one lived only in session memory, which is likely
why it recurred. Records the merge-shaped tell (conflicts in files your
branch never touched) alongside AGENTS.md's checkout-shaped one, and the
discipline that held: git -C on every history-changing command.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
… endpoints for diagnosis

Promotes the general retrieval mechanism out of the bug-857 ledger's
bug-scoped notes: three routes ranked by completeness (report artifact,
run-logs archive zip + attempts endpoint, console log as convenience
only), the completeness check before any diagnosis, and the incident
where 1654 of 5948 lines misattributed a chaos failure to a passing
test in a circulated handoff.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Fq6AcSoq6nMqbRHQxw25wx
…buted-throttling

# Conflicts:
#	STRATEGY.md
#	docs/ideation/2026-08-17-distributed-throttling-ideation.html
#	docs/inflight/branch-audit-orphans.md
#	docs/inflight/bug-857-family.md
#	docs/inflight/next-candidates.md
#	docs/inflight/test-untracked-ci-flakes.md
#	docs/plans/2026-08-18-001-feat-self-scaling-concurrency-plan.md
#	docs/solutions/workflow-issues/gh-run-view-log-truncation.md
…y, and fold in the design review

The requirements-only plan becomes the implementation plan: Planning Contract
(13 KTDs), 11 implementation units, verification contract, definition of done,
grounded against the merged perf/engine-concurrency tree.

A five-persona design review plus an architecture pass produced 25 accepted
corrections, the load-bearing ones being: the pinned DynamicLoadFactor must
never multiply the dispatch quantity (it would run 2x the published target);
the mode system property may select OBSERVE at most, never ENFORCE; the
unseeded start is the static-configuration-derived target, never the
substituted ceiling (start-at-ceiling reproduced the confluentinc#766 flooding
shape on opt-in); a failure-fraction growth inhibitor guards the case where a
fast-rejecting overloaded downstream reads as improved latency; the drain
release is a state-derived seam read, not an unreachable edge action; rebalance
resets gate on an actual assignment delta so group churn cannot starve the
controller of history. U1 now closes both halves of #311.

Delivery is split into two releasable milestones: OBSERVE-only first (the
trust-building diagnostic, zero behavior change), enforcement second, with
ENFORCE documentation gated on closed-loop bench results.
…ored demo registered

Merging master into the perf/engine-concurrency stack produced a tree neither
branch had tested: the perf stack tightened bin/check-copyright-headers.sh
while master grew files the older checker never enforced. Six fork-original
files gain the fork header (four .github/scripts gates, two bench/threads
ceiling probes); Demo.java - the 2021 upstream demo restored verbatim before
modification - is registered in EXTRACTED_FROM_UPSTREAM so its legally-correct
Confluent + Modifications header stops reading as a fork-original violation;
*.css joins ENFORCED_TYPES (the landing-page stylesheet already carried its
header and was failing only as unclassified).
…t, and bound batchSize

calculateQuantityToRequest rounded a batched request up by target - modulo
where the shortfall to the next whole batch is batchSize - modulo, so nearly
every pass under batching requested almost a full extra target of work and
in-flight settled at roughly twice the configured target. The characterization
captured it before the fix: batch 5, target 24, 7 in flight requested 39 where
20 is right, and a real WorkManager handed over 30 records where 15 is right.
The corrected delta also makes lastWorkRequestWasFulfilled honest without
further change, which un-jams the load-factor step-up it silently suppressed.

The second half of the issue: batchSize was never validated, and a zero or
null failed three different ways depending on unrelated settings (silent
no-work stall; ArithmeticException at construction under messageBufferSize;
NPE unboxing). validate() now rejects anything below 1 with a field-named
message.

Release-Note: Batched consumption no longer over-fetches to ~2x the configured
concurrency target, and a batchSize below 1 is rejected at construction.
…th the suite upstream never had

New internal/admission package: the Gradient2 long/short-EWMA gradient with
the PR-88 anti-drift fixes, ported from Netflix/concurrency-limits (Apache-2.0,
attributed per-file) and amended where this engine measured the upstream
defaults wrong for its workloads. Deliberate deviations, each carrying its
rationale and a mutation-proven test: the utilization term is the window
MEDIAN of in-flight with a p90-p10 spread (a maximum reads healthy exactly
when the engine is starved - tail experiment, 2026-08-22); an AIMD backoff arm
answers overload drops (Gradient2 ignores them entirely); a failure-fraction
inhibitor freezes growth when non-success exceeds 20% of a window, because a
fast-rejecting overloaded downstream LOWERS measured latency and the gradient
alone reads that as headroom; a starved-below-ceiling window earns one bounded
probe up so a contraction cannot manufacture the evidence that freezes it; and
windows close on their time bound, holding APP_LIMITED below 10 samples, so a
slow handler cannot stall recovery.

The contaminated-baseline gate test settled the plan's open conditional: the
ported law provably cannot descend when every sample it has ever seen comes
from an already-flooded operating point (bit-pinned at cap, 100 windows). The
bounded probe-down (x0.9 every 5 at-cap-flat windows, ending when a probe
stops improving) is therefore implemented, and the gate asserts descent from a
saturated start without collapse. Flat-latency gating applies only to the
initial trigger - gating the recovery probes on flatness too was tried and
observed to stall descent at the fourth probe.

Pure math only: no engine wiring, no metrics, no threads, every time value
injected. 35 deterministic tests including a seeded simulation whose latency
curve is to be re-fitted from arrival-harness measurements when they land.
…seed, and loud refusal

New options: adaptiveConcurrencyMode (DISABLED default / OBSERVE / ENFORCE)
and adaptiveConcurrencyInitialTarget (0 = unseeded). The pc.adaptiveConcurrency
system property exists for the bench harness and CI matrix, and it may select
at most OBSERVE: a property value of ENFORCE resolves to OBSERVE with a WARN,
because an ambient JVM-wide flag must never be able to hand an experimental
controller the admission target of a production instance - enforcement takes
an explicit line of code. An unrecognized property value fails construction
naming the property and the raw value, never a silent fallback.

Validation rejects a seed outside [1, maxConcurrency] and a seed set while
the mode is DISABLED. Engine gating is a capability method evaluated once in
the constructor before the worker pool is forced: ExternalEngine refuses by
override, and the direct-pull engine refuses through the option read - direct
pull is an option on the core engine, not a subclass, which is exactly how a
class-level mirror of supportsDirectPull() would have gotten it wrong. A
requested-but-unsupported mode logs one WARN naming the reason and the
instance runs static; adaptiveConcurrencyActive is the single downstream
truth.

No dispatch, pool, or target arithmetic changes yet - this unit is the
switch, its guards, and its tests (16 new, each validation and downgrade
branch red-proven by sabotage).
…get and its ceiling

PCModule-wired component holding one instance of the ported window and one of
the control law, adding no arithmetic of its own beyond ceiling resolution and
the publish clamp. Ceiling resolution implements the one-knob rule: a user-set
maxConcurrency is pool size and effective maximum; leaving it at the library
default under ENFORCE substitutes ADAPTIVE_DEFAULT_CEILING (64, explicitly
calibration-pending - the constant is a memory decision as much as a thread
decision, and its javadoc carries the worked buffering example). The
substitution applies under ENFORCE only: OBSERVE keeps the configured value
for pool and admission, computing its would-be target against the ceiling
ENFORCE would use.

The unseeded start is the static-configuration-derived target, never the
substituted ceiling - enabling the mode never raises t=0 admission above
today's behavior - and OBSERVE moves only the would-be value; both rules are
sabotage-proven red before restore. DISABLED constructs an inert controller so
downstream reads never null-check.
…, and never times the factor

One seam: PCModule#admissionTargetRecords() - the controller target in slots
times batchSize under active ENFORCE, byte-for-byte the static derivation
under DISABLED, OBSERVE, a refused engine, or a bare-WorkManager test env.
Dispatch (getPoolLoadTarget and its chain) and the poller gate threshold
(isSufficientlyLoaded, which keeps its load-factor multiplication - that is
buffer arithmetic) both read it.

The KTD10 rule lands here and is mutation-proven: under ENFORCE the dispatch
chain consumes the target UN-multiplied by the pinned load factor. The
factor's meaning - keep the workers N deep in buffered work - held only while
pool size equaled the target; with the pool at the ceiling and a live target
below it, target x factor records would run, not buffer, and the engine would
execute double what the controller published.

getTimeToBlockFor becomes timeToBlockFor (the Truth-generator no-get
convention) and, under active ENFORCE only, its retry branch is bounded by
the remaining time to the next commit check - the retry branch measures
against the full commit interval, and the ceiling pin on
isWorkInFlightMeetingTarget would otherwise push the loop past a due commit.
DISABLED and OBSERVE arithmetic is pinned byte-identical by tests, including
the pre-existing full-interval quirk, which is preserved deliberately for
parity.
…one of them lie

Service time: one sample per user-function invocation, the batch duration
divided by the records actually in it - batch fill is downstream of the
control variable, and an unnormalized batch duration would let the controller
react to its own actuation. A retry anywhere in the batch voids the whole
sample (a fast-failing retry reads as improvement; under-sampling is covered
by the APP_LIMITED hold, pollution is not recoverable), and retries still
count fully as outcome signal. A throwing invocation contributes no latency
sample for the same reason.

Outcomes classify in WorkManager#handleFutureResult on both verdict branches:
success, or IGNORE for every failure in v1 - the package-private classifier
is the documented socket where the structured rate-limit exception and
timeout shapes will map to OVERLOAD_DROP.

In-flight: one sample per control-loop pass reading the conservation-derived
getNumberRecordsOutForProcessing - never the completion path, which samples
just after a decrement and biases the median low. All taps gate on the
active flag (OBSERVE records; DISABLED and refused engines record nothing),
and none of it touches Micrometer: the intake works with no registry
configured. Window access gained one uncontended lock because service-time
samples arrive on worker threads.

Batch normalization and retry exclusion are both sabotage-proven red.
…forgets, and what pins the factor

The tick rides the control loop, gated to RUNNING and to an active mode, with
the window boundary on the controller's own injected clock rather than the
loop's block time - which is target-derived, so a contracted target would
otherwise slow the controller's own recovery. A tick that grows the target
wakes a paused poller so new headroom is actually used. No new sender was
added to the interrupt-based wake channel.

Rebalance handling is gated on a real delta to THIS instance's assignment:
the callbacks do pure set bookkeeping on the broker-poll thread and the reset
decision is taken on the control thread at the next tick. A cooperative
rebalance that moved nothing for this instance keeps its history - otherwise
group churn would starve the controller of the samples it needs, in exactly
the churning groups where per-instance adaptation matters most. A real delta
discards the window and the law's baseline and freezes the target, carried
over as the best available prior, for a cooldown.

The drain release is a state-derived seam read, not an edge action: leaving
RUNNING makes the seam return the effective-maximum derivation, so
close(DRAIN) with a contracted target dispatches at full width on the next
pass - including when close() arrives before the control loop ever ran.
transitionToDraining runs on the caller's closing thread while the tick is
already gated off, so an edge action there would have been unreachable.

Pausing poisons the in-progress window; the first post-resume window carries
no pre-pause samples. DynamicLoadFactor is pinned static under ENFORCE only -
OBSERVE must stay byte-for-byte today's construction - read from the options
because the factory can run before the capability flag exists, and the
messageBufferSize branch now divides by the ceiling-derived in-flight figure
so the buffer is sized for the widest dispatch the controller may publish.
…moving

Four meters under the processor subsystem: pc.admission.target,
pc.admission.would.be.target, pc.admission.constraint and
pc.admission.movements. The constraint gauge follows the PC_STATUS pattern -
hand-assigned values, never ordinals, with the mapping rendered into the
description - because the operator's real questions in steady state are all
"why is it NOT moving": at cap, at floor, app-limited, failure-limited and
post-rebalance cooldown are five different answers that a target gauge alone
cannot tell apart.

Observability does not depend on the operator having wired a MeterRegistry.
The default registry is a null sink, so the same states also surface through
a rate-limited log line naming the mode, the constraint, the live and
would-be targets and the effective maximum - otherwise observe-only mode,
whose entire purpose is to report, would report nothing in the default
configuration.

Gauges are held in fields with strong references and reclaimed through the
existing close path. Registration is gated on the mode rather than the
processor capability flag, because under ENFORCE the controller is
constructed while the load factor is being resolved - before capability
exists; the trade is documented at the registration site.
… why the branch is not the suspect

A full core unit run on this branch failed four times, all point checks in
ParallelEoSStreamProcessorTest; the class then passed 58/58 in isolation on
the same build, on a box running several agent builds at once. Two of the
four are new to the ledger. The change in flight registers meters and one
rate-limited log line, both gated on a mode that defaults to DISABLED and
that nothing in the suite or the poms sets, so it neither registers a meter
nor ticks in those tests - recorded as a sighting, not a quarantine.
astubbs and others added 14 commits August 26, 2026 12:39
Its token regex requires two path segments, so `](sibling.md)` was never checked - a whole class of
link structurally invisible to the gate whose entire job is links. Both `foo.md` and `./foo.md` have
exactly one segment, so neither ever fired.

Relaxed ONLY inside `](...)`, where a target is unambiguously a path and never prose. The
two-segment rule exists to stop running text containing a dot - `Set.removeAll`, `check-all.sh:
message` - being misread as a citation, and none of that prose sits inside a markdown link target,
so relaxing it there carries none of the risk it guards against. A bare one-segment token outside a
link is still not a citation.

It found a real dangling link immediately, in the same sweep that fixed it: this branch renamed a
note and left `bug-857-family.md` pointing at the old name. That link had been invisible to the gate
for as long as it existed.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…axis so this work is findable

THE LEDGER had accumulated fourteen CLASS2_STALL sightings read as evidence of the confluentinc#857
family. The replays settle them as one timing proxy, recorded with the reasoning, the limits, and a
prediction that can falsify it.

Why the ~154s constant was never corroboration: several entries read the tight clustering of peaks as
a signature. It is arithmetic - the probe samples every 5s and the scenario fail-fasts on the first
crossing, so the peak is always bound + detection latency and encodes no severity.

WHAT THE CHANGE COST, which its own write-up initially understated. An independent cross-model pass
found the demotion REDUCED per-shard liveness coverage rather than relocating it: INSTANCE_STALL
carries that property at INSTANCE granularity, so one partition's committed offset freezing while the
owning instance's other shards keep completing fires nothing that gates - and the correctness ledger
does not close it either, counting records PROCESSED rather than offsets durably COMMITTED.
ProgressProbe had written that gap down; the demotion removed the sentence's premise and left the
sentence. Every surface that said otherwise now says so, and the correlated gate is tracked with the
bar it must clear first: a red control proving it fires on an injected commit-freeze, THEN both
replay seeds proving it does not fire on the old false positive. The bound being replaced was itself
green-calibrated, argued for, and wrong for three months.

A THIRD CANDIDATE, recorded not fixed. NO_PROGRESS has fired - `stuck at 98804/100000 for 30s
(bound 30s)` on ChaosChurnStormIT, which unlike W4 does not widen that window. The sweep first called
it "not found", which was wrong: it searched only local logs. Not fixed because nobody has replayed
the seed, and the alternative reading - a genuine fleet-wide stall - would be the most interesting
result in the family.

What this does NOT claim: the wedge is real and #29 still fixes a real
deadlock. The narrower finding is that the chaos suite has never reproduced it. Honest limit:
INSTANCE_STALL's silence is one storm run on a days-old detector, so "never cried wolf" and "never
had a wolf" are not yet distinguishable.

A LABELS AXIS, because neither existing axis can answer "show me the concurrency work". The filename
prefix says the AREA, the impact says the CONSEQUENCE, and concurrency is neither - those notes live
under bug-, core-, static- and test- alike, and their consequences are already spread across stall,
data-loss, crash and reliability. So `bug-concurrency` as a prefix cannot work and a label can.

The first measurement argued AGAINST it and was wrong: an unqualified grep for race|deadlock|volatile
|lock matches most of the corpus, because `lock` matches "blocked" and "blocker" in schedule prose.
By title it is a small minority - the band where a label partitions usefully. Closed set, validated
per value so one typo cannot found a group of one, applied where concurrency is the MECHANISM rather
than a mentioned word - including bug-857-family, which the title grep missed. Adding a label is a
schema commit, which is the discipline that keeps it from becoming tag soup.

A RULE GOT WIDENED BECAUSE IT FAILED HERE. "Never write down what a command can answer" was
illustrated only with git facts, so a table of analyser counts read as out of scope - and a register
written around a total went stale the moment three findings were fixed, leaving it asserting open
work that was closed. docs/inflight/AGENTS.md now names counts explicitly and says why the rule fires
hardest while writing up a measurement you just took: the number feels like the finding, which is
exactly when it is most likely to change. docs/merge-checklist.md gains a read-and-judge item, not a
grep gate, because legitimate figures exist.

The probe-critique note is DELETED rather than rewritten into a FIXED narrative, and its citations
repaired both ways - live docs to the successor, dated records via `git show`, since those may not be
reworded. The same rule shrinks the Phase 2 roster to its one open follow-up. Five copies of the
"RetryQueue.closed is still plain" fact collapse to one owner plus cross-references.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Brings the static-analysis wave, the collapsed repo-hygiene lane and the
inflight labels axis onto the adaptive-concurrency branch. Five conflicts,
and the analysis wave then found real work in the proxy modules master has
never seen.

CONFLICTS, and which side won:

- AbstractParallelEoSStreamProcessor - master extracted forceWorkerPoolConstruction()
  and lookupManagedResource() out of code this branch had rewritten for virtual
  threads. Took master's extractions (its javadoc says what our inline comment did,
  better) and kept OUR signatures: setupWorkerPool returns ExecutorService, not
  ThreadPoolExecutor, because a virtual-thread pool is not one - and
  requireRejectionIsVisible stays ExecutorService-typed for the same reason.
- .github/workflows/repo-hygiene.yml - master collapsed nine enumerated jobs into
  one discovered lane (`bin/check-all.sh --with-tests`) precisely so a gate added
  to bin/ cannot run nowhere. Took that wholesale; our six new bin/check-*.sh are
  found by its glob. Kept ONE step it would have dropped: the PyYAML install two
  of those gates need, which no glob can know about.
- .gitignore, bug-857-family.md, test-chaos-phase2.md - both sides appended; kept
  both, with master's superseding seam-coverage bullet replacing our stale copy.

WHAT THE MERGE THEN SURFACED, none of it conflict-shaped:

- Error Prone crashes on EVERY record. 2.42.0 throws `invalid replacement: [0, -1)`
  building a fix over a record's generated toString. Three checks disabled with the
  reason and the re-enable trigger recorded; the trigger is the existing EP pin.
  MEASURED, not inferred: pasting a record into an unrelated core test file
  reproduced it, then reverted. @desugar cannot be dropped to dodge it - Jabel
  refuses the record without it, tested.
- requireUpperBoundDeps (new from master) fails on modules master does not have.
  protobuf-java declared BELOW what grpc-protobuf already brings, and two versions
  managed to the highest in the tree rather than excluded.
- fb-contrib meets the proxy clients. Two findings fixed outright (an identity
  lambda, a default-encoding getBytes); four scoped to the client packages with
  per-pattern reasoning, because TWO OF THE OFFERED FIXES WOULD BE BUGS - making
  dispatchQueueDepth static shares one client's negotiated concurrency with every
  other in the JVM, and the "overwritten" executor is assigned exactly once behind
  a guard that throws.
- The inflight labels axis arrived from master with a one-value vocabulary while
  41 notes on this side were already labelled against a documented five. Widened,
  with the relaxation of master's mechanism-only rule stated where it is relaxed
  rather than quietly broken.
- Every self-reference to this branch and PR rewritten as it will read after the
  merge, per the new gate; the PR's own note takes the documented exempt-file.
- Two config files from master were unclassified by this branch's newer copyright
  scanner - headers added and the .txt named in ENFORCED_TYPES.
- javac.*.args is now ignored: a crashing compile drops an argument dump in the
  module root, and the copyright gate caught four of them staged by a `git add -A`.

Verified: full reactor test-compile green, bin/check-all.sh 18/18 with only
check-shell-lint CANNOT (shellcheck is not installed on this box; it passes in CI).
The clone had silently re-shallowed mid-merge - `git merge-base` was answering
`exit 1` where it had answered a commit an hour earlier - so it was unshallowed
before the final sweep, or half these gates would have judged a truncated graft.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HwDcZdfv1Ne4u3g6ThrG9
…351)

The most frequent tracked flake (4/45) was quarantined undiagnosed and blocked every
PR. It is a test-design defect, not a product defect: the assertion demanded an offset
that back pressure exists to stop advancing.

THE MECHANISM. The high-water mark encoded into commit metadata is the highest
SUCCEEDED offset, not the highest polled one - `encodeOffsetsCompressed` takes its
range top from `getOffsetHighestSucceeded()`. Offset-encoding back pressure does not
pause the poller; it gates `PartitionState#couldBeTakenAsWork`, which once the
partition is blocked refuses every record at or above the highest succeeded offset. So
the instant back pressure engages, the succeeded frontier FREEZES. The old
`expectedHighestSeen` - the last offset sent - was therefore UNREACHABLE, not late,
which is why no wait ever closed it and why the actual varied run to run. It passed 41
times in 45 only because the mock consumer usually hands the whole extra batch over in
one poll, so those records were claimed as work before the block fired.

THE EXPERIMENT, because a fix that works is not evidence of the cause. A deterministic
probe with no threads, replaying the exact config, shows the payload crossing the size
threshold at offset 136 - exactly the reported actual. The real test was then made to
fail deterministically by splitting the extra send 37 + 3, reproducing the CI signature
character-for-character: `expected: 139 but was : 136 within 30 seconds`. Moving the
claim boundary and the threshold moves the frozen frontier with them - 136 / 124 / 129
/ 136 across four configurations, never 139. The competing explanation, slowness, is
dead: the frozen state is reached with zero time pressure.

THE FIX STRENGTHENS RATHER THAN WEAKENS, which is the bar for touching a test that
fails under stress. The constant moves to where it is true - the partition's
`getOffsetHighestSeen()`, awaited, since record REGISTRATION is not gated by back
pressure - and that check tightens from `isGreaterThan(numberOfRecordsToPrimeWith)` to
an exact equality. Quiescence is then awaited so the succeeded frontier is read as a
still value rather than one of two moving ones, and the committed payload's high-water
mark is asserted to equal that settled frontier.

TWO REVIEWER PROPOSALS WERE IMPLEMENTED, MUTATION-TESTED, AND REJECTED. Both asked for
the lower bound tightened so that back pressure engaging early would go red - a numeric
floor at the measured block point, and the committed payload asserted over the pressure
threshold. Mutating `updateBlockFromEncodingResult` from
`metaPayloadLength > getPressureThresholdValue()` to `... - 2` leaves the test green
under both, logging the identical `Payload size 32 higher than threshold 30.0` as the
control. The succeeded frontier is decided by the CLAIM BOUNDARY, not by where the
threshold sits, so no assertion on it can observe the threshold moving. Adding an
assertion that kills no mutant another assertion does not already kill is what
`docs/testing.md` exists to prevent. The bound is restated as the band it describes, at
identical strength. The gap is real, belongs to the deterministic probe, and is left
open in `docs/inflight/test-back-pressure-engage-point-is-unasserted.md` with the
mutation evidence.

RULED OUT, so it is not re-derived. The torn-read family does not explain this:
#344's tear widens the encoded range and pushes the decoded value UP, while this
is a shortfall - wrong direction; and the test passes base 0 explicitly on deserialise,
so #337's committed-base/payload tear cannot shift what it reads. The
retry-delay sleep was ruled out previously: it runs after the failing assertion.

DEFECT-CLASS SWEEP - none found elsewhere, and here is where I looked.
`OffsetEncodingBackPressureUnitTest` carries the identical expression but asserts
`getOffsetHighestSeen()` and succeeds every extra record synchronously before encoding,
which is exactly why it never flaked. Every other frontier assertion in the offsets
tests - `OffsetEncodingTests`, `WorkManagerOffsetMapCodecManagerTest`,
`BitSetEncodingTest`, `RunLengthEncoderTest` - compares against `highestSucceeded` by
name or against a fixed literal with no back pressure in play.

RESIDUAL, stated rather than buried. The ledger's other recorded actual, 132, is below
what today's constants allow - the frozen frontier sits at or above the 136 block
point. It implies a lower effective threshold in that run, most likely an older tree;
the failing sha was never recorded. The fix removes the sensitivity either way.

PROVENANCE CARRIED ACROSS. The retired ledger entry was the reason quarantine rule 1
changed from "no quarantine without diagnosis" to accepting a sighting ledger. That
fact would have been lost when the entry was deleted, so it now lives in the solution
write-up's Prevention section.

ALSO CARRIED, deliberately: a confluentinc#857 chaos sighting that is NOT this PR's -
the only code file changed is a surefire unit test the chaos lane does not run.
`ChaosRevokeUnderWorkIT`, seed 8584935079849032188, failing a CORRECTNESS assertion
rather than a timing bound: a commit request timing out after PT10S from a broker-poll
thread the runtime has positively established is alive and not throwing. #354
independently recorded the same signature from the Integration lane the same day, on
all three preconditions for the AB-BA cycle - so this is a second lane corroborating,
not a tally mark. Neither identifies the deadlock; both want the same discriminator, a
thread dump at the moment of the timeout. Three CLASS2_STALL entries written up earlier
on this branch are dropped: #354 ran the discriminator and closed that line as
one timing proxy, and the ~154s constant is arithmetic - the probe samples every 5s and
the scenario fail-fasts, so the peak is always bound plus detection latency. Their two
seeds are kept so the pair is not rediscovered and re-argued.

With #57 and #262, this empties the quarantine registry and clears the
rule-5 release gate.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…and read as a shutdown bug

`executorThreadsInterruptedOnShutdownTimeout` asserts that a worker parked in
the user function is interrupted when close() times out. It established that
worker's existence with `awaitForSomeLoopCycles(2)` and hoped: if the primed
record had not reached a worker yet, there was no blocked thread to interrupt,
`interrupted` stayed false, and the failure pointed at shutdown.

MEASURED, not guessed. After merging master the full core suite failed it twice
running, same parameter, same 2.09s - while the class ALONE passed on the merged
tree, on master, and on the pre-merge commit (58/58 each). Four pre-merge
full-suite samples never failed it; three merged samples failed it twice. That
gap is suggestive and not significant (p is about 0.14), which is exactly why it
was not worth buying more samples: the test could not say which of the two
possible causes it had met, so more of it would have measured the same ambiguity.

It now awaits an `entered` latch before closing, so the outcomes are
distinguishable rather than merged: a timeout at the new await says the record
never reached a worker, and a failure at the original assertion says the
interrupt genuinely did not arrive. Nothing was loosened - no timeout widened,
no assertion weakened, no retry - the test asserts strictly more than before.
That is the fix docs/inflight/test-untracked-ci-flakes.md already prescribes for
this file's failure shape, "a point check taken after an await on something
else", rather than a quarantine.

Two full-suite runs since: the shutdown test passes both. The suite is still not
clean - `JStreamParallelEoSStreamProcessorTest` failed once here and once on the
pre-merge side, which puts it outside this change and outside the merge.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HwDcZdfv1Ne4u3g6ThrG9
…weep

`bin/check-quarantine-owners.sh` reached the owning PR's base and merge
preview with `git fetch --quiet --depth=1`. A depth-limited fetch writes
the `shallow` file, and that file lives in the shared `--git-common-dir`
- so one run truncated history for EVERY worktree of the clone at once,
not just the caller's. `bin/check-all.sh` globs `bin/check-*.sh`, which
made the mandated pre-push sweep an instruction to corrupt the clone.

WHAT IT COST. Nothing goes red; the damage lands on other commands.
Observed three times in one session: `git merge-base` returned empty,
ahead/behind counts read 295 and 835 against true values in the tens,
and commits that had demonstrably landed reported "NOT an ancestor of
master". Read naively that says master was rewritten and a day's work is
gone, and two sessions nearly acted on it. One repair failed with
"shallow file has changed since we read it" because a sibling worktree
was fetching concurrently.

THE FIX. Ref previews are now fetched into a throwaway git dir
(`git --git-dir=<mktemp -d> fetch --depth=1 ...`) and read back through
`FETCH_HEAD` there. The working clone is never a fetch target, so its
depth is neither read nor written. Measured cost of the isolation on
this repository: ~1.4s and ~2.3MB for the first fetch, with the dir
reused for the rest of the run; the live gate run takes 3.5s total.

REJECTED: choosing the depth from `git rev-parse --is-shallow-repository`
and only passing `--depth=1` when the clone is already shallow. It is
the obvious repair and it still SAMPLES shared state that a sibling
worktree can change between the sample and the fetch. Fetching elsewhere
needs no sample, so there is nothing to race on; it also cannot deepen a
clone that arrived shallow on purpose, and an interrupted run leaves the
clone exactly as it was.

The scratch dir also replaces gh_query's per-call `mktemp`, which leaked
a file whenever a run was interrupted. Two things surfaced while proving
that, both fixed here and both worth knowing: `exit` from inside a
signal handler is documented to run the EXIT trap and measurably does
not always - instrumented, the TERM handler ran in the main shell and
the EXIT trap did not follow, on roughly one run in five - so the
handlers call the cleanup themselves; and a second signal arriving
during teardown re-enters the handler and abandons a half-finished
`rm -rf`, so cleanup disarms INT/TERM as its first act.

`bin/test-check-quarantine-owners.sh` is new and asserts the invariant
rather than describing it: a full clone stays full, a deliberately
shallow clone is neither deepened nor re-shallowed, an interrupted run
leaves both the clone and the temp dir clean - and the gate still
reaches its OK and its "does NOT yet remove the quarantine" verdicts, so
the side effect cannot be removed by removing the check. It is hermetic:
a fake `gh` on PATH and a local file:// origin carrying real
refs/pull/N/merge refs. Verified red against the unfixed script first -
6 of 13 arms failed, including all three depth arms - while both
behaviour arms stayed green, which is what shows this preserves what the
gate does.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Fixing the one script removes today's instance. The class - a
depth-limited fetch into a clone whose `shallow` file every worktree
shares - has three entrances, and none of them can see the others, so
each gets its own guard.

A SCRIPT DOES IT. `bin/check-shell-hazards.sh` grows a second category,
`shared-git-state`, which is what its generic shape was built for: `git
fetch --depth` is portable and correct on every platform, and what it
does silently is truncate history for every worktree of the clone. Same
signature as the GNU-vs-BSD rows - no error, a wrong answer somewhere
else, unreachable by a linter - so it is a table row rather than new
control flow. It had two live findings on arrival, both in the script
fixed by the previous commit.

`git clone --depth` is deliberately NOT the hazard: a clone owns its own
depth. `git pull --depth` is, and is matched, because it is the obvious
next spelling. The pattern steps over git's own global options,
including the value-taking ones - the first version read `core.pager=cat`
as the subcommand and missed `git -c k=v fetch --depth=1` entirely,
which is the same walk-past the shallow-history hook's header documents
for `git -C DIR rev-list`. Both forms now have arms.

AN AGENT OR HUMAN TYPES IT. `.claude/hooks/check-shallow-history.sh`
already denied the depth-dependent QUERIES that answer wrongly from a
truncated graft; it now also denies the fetch that does the truncating,
and names the throwaway-git-dir alternative in the deny message. The two
directions are gated in opposite senses on purpose: a query is wrong
only when the clone is ALREADY shallow, while a shallowing fetch is only
worth stopping while it is NOT - denying it in an intentionally shallow
CI clone would be noise, and noise is what gets a hook switched off.
`--git-dir=` elsewhere is exempt, since that names another repository.

Both guards verified red against the previous code first: 5 new hook
arms failed with no false positives among the 33 existing ones, and the
two walk-past hazard arms failed against the looser pattern.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The story behind the two guards had no durable home: the only write-up
was an inflight note living on another branch, which is deleted when
that branch's work lands. This gives the `shared-git-state` hazard row
and the hook's new deny message something to cite.

Records what the class is, why a conditional depth is not enough, the
three vectors and which guard closes each, the measured cost of the
isolation, and the two bash findings that came out of proving it - `exit`
from a signal handler not reliably reaching the EXIT trap, and a second
signal abandoning a half-finished cleanup.

It also states what was checked and RULED OUT, since a sweep is only
worth reading if it says where it looked: the workflows' `fetch-depth`
checkouts, pr-checklist.yml's own depth-1 fetch on a CI runner,
ci-mutation-test.sh's undepthed fetch, the `git clone --depth=1` fixtures
in two self-tests, and every other git mutation in bin/ - all of which
target a scratch repository or a per-job workspace rather than the
working clone.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
… count

The `highcpu-box-exclusive` concurrency group was discarding most of the work it
was meant to schedule. Removing it, and with it every workflow-side claim about
how much the self-hosted box can do at once.

WHAT WAS WRONG. A GitHub `concurrency` group is a deduplication primitive, not a
queue: it keeps one run in progress and at most ONE pending, and discards
whatever arrives behind that. Used as a mutex over a shared machine, that means
a third arrival evicts the run already waiting - so with several branches active,
each push deleted somebody else's pending measurement.

MEASURED, in the 50 minutes after the group landed (2026-08-26, 01:03Z-01:53Z,
16 runs across 9 branches): 26 of 32 box jobs never executed a single step -
Chaos Pain Suite 12 of 16 evicted while pending, Performance 14 of 16 - while
five of the six runners sat idle. Runs were cancelled roughly a minute after
creation, never having started. The evicted job is whoever queued last, which is
as often chaos as Performance, so the tripwire built to hunt confluentinc#857
was the suite most often not running.

WHAT REPLACES IT: nothing. Jobs queue on the runners like every other GitHub
Actions job, and every queued job eventually runs. How many run at once is the
box's own decision, made by how many runner processes it runs - six today. That
is not a workflow concern, and encoding it here is what produced this defect;
if six is too many, the lever is on the box and needs no change in this repo.
The only cancellation left is the ordinary one: a new push to a PR supersedes
that PR's previous run.

WHY IT IS SAFE TO DROP THE EXCLUSION. The co-residency reds that motivated the
mutex were ~154s lagStagnation against a 150s bound - the bound meeting the
load, not a defect. That detector was demoted to non-gating in the same pull
request that added the mutex: ProgressProbe.recordLagStagnation now calls
observe() rather than recording a violation. The failure was already fixed in
the instrument, where it belonged.

PREDICTION TESTED AND REFUTED. The design hinged on how many runners serve the
`highcpu` label: one runner would have made the group redundant, since a single
runner already serialises. Six are registered and online, all advertising the
same labels - so the group was not redundant, and simply deleting it does return
up to six-way co-residency. That is accepted deliberately rather than by
oversight, and whether six is too many is now an open runner-count question
recorded in docs/inflight/.

ALSO REMOVED: `max-parallel: 1`, which existed only to stop one run's two suites
displacing each other inside the mutex, and the `box-exclusive` matrix key that
selected into it. `chaos-pain.yml` and `mutation-full-sweep.yml` shared the same
group and lose it too - the sweep's own comment recorded that its 360-minute
timeout could starve PRs of chaos measurements for hours, which is this defect
at its widest.

REJECTED: keeping the group but making eviction loud, which leaves the work
discarded; moving chaos to a scheduled or post-merge lane, which docs/ci.md
forbids for anything a PR gate already covers ("does time alone change the
answer?"); and a dedicated single-slot runner label, which needs provisioning
and carries the silent-queue trap of a label nothing serves - reducing the
runner count achieves the same thing with neither cost.

A cancelled check still renders as a failure in `gh pr checks`, so docs/ci.md
keeps the rule that a cancelled chaos check means NOT MEASURED - neither a pass
nor a failure - and the job summary still names the commit it measured.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SFRbSkHwaRw6YYFHHBWZ7p
…e redirect

CheckQuarantineOwnersScriptTest caught a real regression from the
previous commit. It stubs `git` on PATH with a `case "$1" in
fetch|show|cat-file)` dispatcher, so `git --git-dir=X show ...` matched
nothing, the stub printed nothing, and three of its twelve assertions
failed while the script itself behaved perfectly.

The stub is naive, but the property it depends on is worth keeping for
free: `GIT_DIR=<dir> git <subcommand>` means exactly what `git --git-dir
=<dir> <subcommand>` means, and leaves the subcommand where every wrapper
looks for it. This is the third instance of the same walk-past in a week
- `git -C DIR rev-list` past the shallow-history hook, `git -c k=v fetch`
past the new hazard row, and now a global option past a test stub - so
the script now says why the spelling matters rather than leaving the next
editor to rediscover it.

`.claude/hooks/check-shallow-history.sh` had to learn the same idiom, or
it would deny the very alternative its own deny message recommends: a
`GIT_DIR=` assignment prefix now counts as a redirect alongside
`--git-dir`, with arms for it and for an unrelated env prefix, which must
still be denied.

Verified: all 12 CheckQuarantineOwnersScriptTest cases pass locally
(`./mvnw -o -pl :parallel-consumer-core -am test -Dtest=...`), the three
shell self-test suites are green, and shellcheck is clean at
`--severity=error` over the whole corpus through the cached image.

Also records the hook's two deliberate imprecisions, both raised in
review and both erring towards denial: `-C <dir>` is not read as a
redirect (treating it as one would be a bypass, since `-C .` names this
repository), and once the clone is already shallow this arm stands down,
so a fetch that CHANGES the depth still rewrites the shared graft. The
second residual is bounded rather than missed - in an already-shallow
clone the query arm is armed, so no truncated answer reaches anyone
silently.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…ther it needs the box at all

Adds an advisory `Chaos Pain Suite (hosted, experimental)` lane to maven.yml's
per-PR suite matrix, on ubuntu-latest, alongside the existing self-hosted one.
Chaos runs twice per PR on purpose while the trial lasts, so the same commit
produces a hosted result and a self-hosted one and any disagreement between them
is the measurement.

WHAT IS UNDER TEST. That the suite needs many real cores to provoke anything is
an assumption nobody has measured - it is why chaos lives on `highcpu`. If it is
wrong, the co-residency problem disappears rather than being managed: a hosted
runner gives every job its OWN VM, so there is no shared box to contend for and
no scheduling to get right. Every scheduling contortion that lane has
accumulated - the per-suite group keys, the repo-wide mutex that discarded most
of its jobs, the co-residency watch - is the cost of where chaos runs, not of
what chaos does.

Supporting evidence that the assumption is weak: bin/chaos-test.sh needed no
change to run here. It passes no forkCount and no -Dparallel-tests, so the suite
was never actually configured to exploit the extra cores it was placed there
for.

ADVISORY, DELIBERATELY. `optional: "true"` makes the entry continue-on-error so
the trial cannot gate a merge before we know it is neither flaky nor timing out.
The job carries a 60-minute cap, matching integration and performance.

HOW IT RESOLVES. A green that selected no scenarios is not a pass - read the
job's own "Chaos suite timing" summary and its zero-tests-selected warning. If
it works, the self-hosted chaos entry goes and that lane is left carrying only
the full mutation sweep, which is the one job that genuinely wants every core it
can get. If it does not, record which way it failed before reverting - wall-clock
against the cap, an OOM, or scenarios that never fire without real parallelism
are three different findings, and only the third actually justifies the box.

Tracked in docs/inflight/ci-chaos-on-hosted-runners-experiment.md.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SFRbSkHwaRw6YYFHHBWZ7p
…-the-clone' into feats/ideate-distributed-throttling
365 is the fix for something this branch hit twice today: a gate was truncating
the shared clone's history on every sweep, so `git merge-base` answered exit 1
where it had answered a commit an hour earlier, and half the gates would have
judged a truncated graft while reporting green. Taking it here stops that
recurring mid-merge.

366 brings the hosted-chaos trial and, with it, #351 - which diagnosed
`OffsetEncodingBackPressureTest` (the 4/45 flake, the most frequent tracked one)
as a test that froze the offset it then asserted.

FOUR CONFLICTS, and the two that needed a decision:

- docs/quarantined-tests.md - 366 removes the back-pressure entry because #351
  released it. Taken. But the PCMetricsTest entry that stays cited it as "the
  entry above" for the shared shape of their failures, so that cross-reference is
  repaired rather than left pointing at nothing: the rhyme is now stated in the
  past tense and names #351 as the first place to look for whether the two are one
  phenomenon or two. Registry re-verified, 2 entries.
- docs/inflight/test-untracked-ci-flakes.md - 366 retires the whole "NOT
  diagnosed - quarantined anyway" section. Taken whole rather than re-adding a
  section its author deliberately removed, after confirming the durable part - why
  quarantine rule 1 changed - survives in the solutions write-up that replaced it.
- docs/ci.md and maven.yml - both sides edited the same paragraph and the same
  matrix. Kept both edits: the proto-breaking gate this branch gained from master
  stays listed beside 366's hosted-chaos job, and `continue-on-error` now honours
  both `advisory` (the execution-mode axis) and `optional` (366's experiment).
  Two names for one meaning, deliberately not unified - renaming either would edit
  the other branch's vocabulary and re-conflict when it lands.

Also fixes two SpotBugs findings the merge surfaced in SessionEndTest: an ignored
`CountDownLatch.await(timeout)` return in each of two user functions. A discarded
await cannot be told from a call whose result nobody needs, and false there means
the release never came inside the budget. Named per the repo convention rather
than thrown from, since throwing on that thread would fail the record for a reason
the test is not measuring.

Verified on the merged tree: full reactor test-compile green, bin/check-all.sh
18/18 with only check-shell-lint CANNOT (shellcheck is not installed on this box).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HwDcZdfv1Ne4u3g6ThrG9
…said zero

`includeTests` is on in the root pom with a comment arguing that test code is main
code for a concurrency library. The execution that consumes it was bound to
`process-classes`, which Maven runs BEFORE `test-compile` - so on a clean build
target/test-classes does not exist yet, and the gate printed "BugInstance size is 0"
for test code it had never read. bin/ci-unit-test.sh runs `clean test` and
bin/ci-build.sh runs `clean verify`, so that was every CI build since the setting
landed.

Incrementally it was worse than blind: the directory still held the PREVIOUS
compilation, so the verdict described the source as it was one edit ago. That is how
it surfaced - a finding stayed red at its old line number after the code was fixed,
which reads as a stubborn tool rather than a stale one.

MEASURED BOTH WAYS, and the red-before arm is the load-bearing half. With an ignored
`CountDownLatch.await(timeout)` restored in SessionEndTest: at `process-classes` a
clean build reported 0 findings and BUILD SUCCESS; at `process-test-classes` the same
source reported the finding and failed. Without that arm, moving the phase and seeing
green would have proved nothing - green is what the broken state produced too.

`process-test-classes` rather than `test-compile`: it is the phase Maven provides for
exactly this, immediately after test compilation, and both CI goals pass through it.
The cost is that a bare `./mvnw test-compile` no longer runs the gate, because the
lifecycle stops one phase short - which is honest, since that invocation never had
test classes to analyse either.

ENFORCED, NOT DOCUMENTED. bin/check-analysis-phase.sh refuses a test-including
analysis execution bound earlier than test-compile, and refuses one with no phase
written down at all. Its self-test carries the shipped shape as a negative control,
because a self-test that only proves the gate passes on a good tree proves nothing
here: the broken state also looked like a pass. A walk that finds nothing in scope
exits 2 rather than 0, so check-all counts it as CANNOT rather than a pass.

Recorded as instance 6 in the write-up that owns this class,
docs/solutions/workflow-issues/an-inert-analysis-config-reads-as-a-clean-codebase.md,
whose status line said the effective-pom check "is a habit, not a gate" - which is
what a missing gate cost here: the setting was added deliberately, reviewed, and
silently did nothing.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_019HwDcZdfv1Ne4u3g6ThrG9
@astubbs astubbs changed the title feat(core) astubbs#227: adaptive concurrency - the engine discovers its own admission target feat(core) astubbs#227: auto scaling / adaptive concurrency - the engine discovers its own admission target Aug 27, 2026
astubbs added a commit that referenced this pull request Aug 27, 2026
…ls that mean most of it is assembly

Architecture from Antony, recorded with the parts that are load-bearing argued
rather than listed.

Constrained worker containers - one CPU, 4GB - are the point rather than a
compromise: a big host hides thread starvation and GC interacting with poll
timeouts, and this family already has a signature that replays red on a contended
box and green on an uncontended one. A user function whose latency follows a
schedule is what makes #333's adaptive concurrency evaluable at all, since an
admission target cannot be judged without a load trajectory to track. External
scaling driven by the same indicators produces workload-driven rebalances, which is
a different and more realistic shape than the injected ones the suite uses now.

The note puts the weight on REAPING rather than looping, because that is what
decides whether a multi-day run is worth anything: split log streams at the logback
level, the autoscaling track record as structured data rather than prose, time
series in preference to log lines for anything continuous, a ring buffer flushed
only on failure - three days of DEBUG is unusable and the chaos log already reaches
126MB in minutes - and an automatic thread dump at the moment of failure, since the
six dumps that identified the revoke deadlock are the only reason it stopped being a
signature and became a mechanism.

Two things stated so they are not misread later: a seed reproduces the SCHEDULE and
not the outcome, because real brokers and containers are not deterministic; and the
harness must report what a run REACHED, not only whether it passed - a soak with no
notion of reaching the interesting state manufactures the same false green this repo
has already paid for three times, at much greater length.

And a section on not reinventing wheels, because most of this is assembly. Kafka's
Trogdor is built for long-running fault-injection soaks; its verifiable
producer/consumer give an oracle for loss and duplication INDEPENDENT of PC's own
counters, which every correctness claim here currently relies on; Toxiproxy adds the
network fault class nothing exercises today and has a Testcontainers module;
Prometheus and Grafana turn the autoscaling question into a chart, and PC already
exposes micrometer; logback already has SiftingAppender and CyclicBufferAppender for
the routing and the ring buffer. Jepsen is the wrong shape but its
generator/nemesis/checker structure is exactly conductor, chaos schedule and
correctness ledger - all three of which this repo already has and which the harness
should drive rather than replace.

Corrected while writing: GitHub-hosted runners cap a job at six hours, but
self-hosted has no GitHub-imposed limit beyond a raisable timeout-minutes, so the
highcpu rig can host a run of days.
astubbs added a commit that referenced this pull request Aug 27, 2026
…ighting, and inflight the next logging iterations

The ledger's conclusions were reachable only from a summary at the top. A reader
landing on the eleventh sighting saw a detailed, confident write-up and nothing
saying it had been withdrawn. Every sighting now carries a STATUS line directly
under its heading.

The verdicts are the ones this file already reached elsewhere - nothing new is
decided here, it is moved to where it is read. Roughly half the ledger is
superseded: every CLASS2_STALL entry falls with the 2026-08-25 demotion, the
seventh was a test defect and is explicitly not a family sighting, the fourth was a
harness double-start race. The eleventh is marked withdrawn in its own words. What
remains open is the ASYNC NO_PROGRESS line - the ninth, tenth and their siblings -
which is now tagged as tractable, because that line finally has a seed that replays
on unmodified master.

Doing this makes the file usable as what it is: a place agents log a sighting and
move on. Without it the next agent inherits twenty confident entries with no way to
tell which are evidence, which is how the CLASS2 reading survived for a year.

Also inflights the three logging iterations the first cut deliberately skipped, each
with a trigger rather than a date: a ring buffer for the verbose stream before the
first long soak, with errors and warnings staying permanently outside it; error
classification so an unforced error becomes a finding, which needs the close and
revoke paths to stop swallowing everything into one warn; and a structured
autoscaling decision stream once #333 lands, since a controller is judged on
a trajectory and prose is the wrong shape for one.
astubbs added a commit that referenced this pull request Aug 31, 2026
…333 is dimension 1

Owner correction to the previous commit: the "~500 useful concurrent operations" in the follow-up
review is not a configured fiction - it is the discovered admission target that #333
already implements (Gradient2 port, OBSERVE -> ENFORCE staging, on the perf/engine-concurrency
stack; assume it merges soon). The arbitration note's caveat is rewritten around that: what
survives is only that the arbiter bites when the sum of per-function profitable concurrency
exceeds the discovered process-wide target, plus the ordering-starvation edge #333 itself
names.

Also fixes core-auto-scaling.md, whose "next step: brainstorm dimension 1 into requirements" was
stale against reality - dimension 1 is an open implementation PR, not future work.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PFA3CXkGc2m32cBJMAWTK
astubbs added a commit that referenced this pull request Aug 31, 2026
…hy am I not going faster?"

Second bit of the follow-up Codex strategy conversation (2026-08-29/30). The genuinely new feature
idea: promote the limiting-regime classification the adaptive controller already computes
internally (LOCAL_CPU / DOWNSTREAM_SATURATION / ORDERING_PARALLELISM) into a per-function,
operator-facing diagnosis - and its differentiator over every observability product is that the
evidence is experimental: the engine did not infer that concurrency 300 was worse, it tried 300.
Scoped as surfacing, not invention: #333 already recognises the ordering regime and its
OBSERVE mode already reports why the target is not moving.

Two compositions recorded with it: causal propagation through a Streams topology (a limited
function has deep runnable work and a flat probe, a starved one has an empty queue - so PC can
name the causal bottleneck instead of the busiest operator), and bottleneck-directed scaling
(scale out FOR a named function; the new instance's allocator preferentially feeds it). One
tension flagged rather than resolved: magnitude ("~4,000 records/sec exploitable") belongs in the
diagnosis, while the dimension-2 vote stays deliberately clamped to +1/0/-1 - upgrading the vote
would reopen the oscillation question the clamp answered.

The arbitration note is sharpened with the conversation's second pass: the allocator question
("where does the next unit of concurrency produce the greatest marginal benefit"), scale-out as
the consequence of failing to satisfy profitable internal demand, and the decoupling endpoint -
functions and their ordering domains become the scheduling entities; process, pod, partition and
language all stop being units of parallelism. It also now names a Streams topology (#271)
as the nearest existing multi-function process, ahead of the per-topic-functions route.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PFA3CXkGc2m32cBJMAWTK
astubbs added a commit that referenced this pull request Aug 31, 2026
…up review

Third bit of the follow-up Codex strategy conversation (2026-08-29/30), two feature ideas.

core-partition-advisor.md: tell the operator, from discovered evidence, when partition count has
ACTUALLY become the processing constraint - and when adding partitions buys nothing. Nearly free
once its parents exist: dimension 2 already caps the instance recommendation at partition count,
and the advisor is that cap's contrapositive (the cap binding while profitable parallelism
remains IS the finding). Scoped to processing capacity only, with the key-remapping cost named -
an advisor that recommends repartitioning without pricing the ordering-transition and Streams
state-locality break would cause the incident it exists to prevent.

core-slo-objective-api.md: subordinate the #333 controller to a declared objective -
"p99 under 500ms, maximize throughput" - the natural completion of #227's "stop making
users pick maxConcurrency". SLO violation becomes the scale signal and its converse the strongest
do-not-scale signal; latency budgets propagate through a topology; the per-function allocator
becomes importance-aware (#236 is the static ancestor). The load-bearing caveat is
Little's law: residence under backlog measures the backlog, so the controller must separate
service-time-dominated from queue-dominated residence or the first Monday-morning catch-up
discredits the feature. The cost-to-SLO benchmark note already names the same gap from the
measurement side; the two notes are cross-linked as twins.

Content series gains the partition lines ("Choose partitions for Kafka. Let PC choose parallelism
for your application"), and the attribution note names ownership as the fourth regime.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PFA3CXkGc2m32cBJMAWTK
astubbs added a commit that referenced this pull request Aug 31, 2026
…eatures, filed and bounded

Fifth bit of the follow-up Codex strategy conversation (2026-08-29/30): the deliberate move away
from adaptive-concurrency territory, mining what PC uniquely knows where Kafka semantics meet
application execution. Eight new notes, each kept to the idea, its prior-art hook, and the caveat
that decides whether it is honest:

- core-record-semantic-tracing.md - the per-record "why did it wait" timeline; #359 stamps
  only the endpoints, and the per-gate stamps are the opportunity model's own instrumentation.
- core-ordering-profiler.md - ordering tax, ordering-scope discovery, per-record critical path:
  one architectural profiler in three stages, near-free on the shard state #361 added.
- core-retry-economics.md - blast-radius quarantine, per-function retry amplification, and storm
  detection (#333's failure-fraction inhibitor promoted to an operator warning).
- core-capacity-fingerprinting.md - persist what the controller learns: warm starts, empirical
  workload models, semantic regression detection; open question is where the fingerprint lives.
- perf-workload-replay-simulator.md - sanitised trace capture feeding the #362 harness as
  a capacity-planning simulator over real workloads.
- release-certified-execution-semantics.md - publish #293's conformance matrix as a
  certification claim; a table that cannot show a cross certifies nothing.
- core-function-manifest.md - split the idea on the positioning line: adopt the language-neutral
  manifest, leave "pc deploy" to platforms (embedded-not-cluster).
- core-scheduler-canarying.md - A/B a scheduler over 1% of ordering domains; stratify or the
  comparison reports the sampling.

Amendments: web-control-plane gains true lag as the fifth instrument (broker lag vs effective
lag); the facades note gains the migration advisor with its honesty bound (observation mode sees
keys and poll cadence, not handler time); the research program gains broker portability as
question 6 with the reproduce-at-own-operating-point trap named; and the opportunity model gains
the standing task the conversation closed on - inventory the boundary knowledge on purpose.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PFA3CXkGc2m32cBJMAWTK
astubbs added a commit that referenced this pull request Aug 31, 2026
…sion stays explainable

Thirteenth bit of the follow-up conversation (2026-08-30 ~5:31pm). Two genuinely new items.

core-scale-in-proof.md: run the counterfactual before Kubernetes acts - constrain the fleet's
controllers to 11/12ths of capacity and prove the SLO holds before removing an instance, then
iterate downward. Entirely a recombination of existing machinery (admission constraint, residence
time, backlog trajectory as the low-demand guard, the controller's experiment discipline). The
deliverable is the metric: current 12, proven safe 8, overprovisioning ~33% - with "proven"
carrying the same experimental-evidence weight as attribution, and per-function
instance-equivalents giving cost attribution by Kafka function, plausibly more commercially
valuable than another throughput benchmark. Caveat kept: a proof is valid for the traffic it ran
under; seasonality decides what it licenses.

The attribution note gains the design rule the exchange closed on, which should bind #333
as it evolves: every adaptive decision retains the probe history that produced it - one ledger
becomes autoscaling evidence, diagnostics, GUI content, research data and promotional material,
and turns the product surface from a metrics page into a conclusions page. The
machinery-as-features note gains the phenomenon's name (compositional reinforcement - other
capabilities emerge by reusing a primitive's semantics rather than adding mechanisms), and the
content series gains the origin-story line: we tried to add global rate limiting and discovered
the correct solution was a distributed execution scheduler.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016PFA3CXkGc2m32cBJMAWTK
astubbs added a commit that referenced this pull request Sep 2, 2026
…persona

A second customer for the same credit machinery, from the owner (2026-09-02). The
persona already on the record reaches for us knowing they need a rate limiter.
This one never thinks that sentence: they need throughput, and in the owner's
framing they are not handed credits at all. They register work and receive a
trigger - go now - and never read a body, because to them it is a technical
signal, not a business record. Speed-limit sign versus pacemaker; in house
terms, the scheduler's dispatch loop extended to a foreign caller.

WHY IT IS NOT A SECOND BUILD: it is the language-proxy's worker protocol with
the payload removed. The sidecar sends records to foreign workers under
credit-on-ack flow control; this sends permission to foreign callers under the
same, and the per-record delivery token the sidecar already carries is exactly
what a trigger carries - identity, nothing else. The delivery is the message.

What differs is mechanical, and one part corrects this session's own earlier
draft. A paced client has no local spend, so the grant dial does not apply to
it: batch-of-one by construction, delivered by push, overshoot impossible, with
throughput from pipelining outstanding registrations rather than delegation.
The earlier draft had placed the throughput persona at the delegation end - it
sits at the other end, and batch-of-one now has a customer where before it had
only a semantics argument. It still owes one verb back (done - without it the
flow control has no feedback and the scheduler paces blind), and being thin has
a liveness price: a credit-holding client keeps going when the shard owner is
unreachable, a paced client stops.

THE COMPETITIVE LINE MOVES. The paced persona compares us to adaptive-
concurrency libraries, not rate limiters, and the reference is Netflix's
concurrency-limits, whose Gradient2 #333 already ported. Its usage is
described concretely (wrap a call site; acquire, then report the outcome; learns
its ceiling locally from RTT; sheds when over) so the shape distinction is
visible: consult-a-gate versus be-driven. Positioning settled with the owner as
"the next layer up", not "more powerful" - the controller consumes ceilings as
inputs, so concurrency-limits is what one instance discovers alone and this is
what happens when ceilings start being agreed.

TWO CORRECTIONS THE OWNER MADE TO THE FIRST DRAFT OF THAT LINE. The cost of
"coordination" was overstated: the plane is embedded and sharded, failover
follows Kafka ownership, a credit-holding client makes zero hops on the hot path
(identical to the local library) and a paced client one pipelined hop (identical
to any remote call). What is genuinely left is one endpoint for a non-Kafka
caller to reach, which every non-local limiter also costs - concurrency-limits
is the only zero-dependency option in the set, and a team with one downstream to
protect is right to pick it. And the failure mode of the global half IS the local
library, by dimension 1's design, for participants that have a local controller;
the paced client's floor is stop. The "beats concurrency-limits" claim is
recorded as the hypothesis the falsification staircase's rung 1 tests, not as a
positioning line.

The standalone note's embedded-not-cluster line is refined from "per-call
decision violates it" to FORCING VERSUS OFFERING: a caller who asks to be paced
has handed over the hot path on purpose, the way a consumer depends on its
broker, and the batch dial becomes a choice between two customers rather than a
positioning hazard. The default stays at delegation because it is the shape that
costs the caller no dependency. Law 3 in the vision doc gains the connection
that its failure mode is the law with the global half removed.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CEwH7RcimXni5JmU2mBobA
astubbs added a commit that referenced this pull request Sep 3, 2026
…oth directions

The branch sweep found the ideas; this connects them to where the decision now lives,
so neither side has to be rediscovered again.

WHAT THE LINKS ESTABLISHED. Two of them answered a question worth asking - did our
current work know? Both did. #333's `core-auto-scaling.md` already cites
`features/dynamic-concurrency-control` by SHA and summarises it, and #392's
`core-distributed-throttling.md` already names `features/rate-limiting` among related
abandoned branches. So the ideas were not lost; they were catalogued in
`docs/refactoring.md`'s idea bank and found from there. What was missing was the REVERSE
link - a branch had no way to point at the work that superseded it - and a record in the
ledger, which is why the same branches still had to be rediscovered by reading commits.

`client-factory` -> #420. Same shape for the consumer that #420 builds for the
producer, and more relevant now than when written: master already enforces exclusive
consumer ownership at RUNTIME through `ThreadConfinedConsumer`, so a factory makes
structural what is currently a guard.

`refactor/control-loop` -> the God-class section, with the five names it cuts along:
`ControlLoop` / `Controller` / `StateMachine` / `PCWorkerPool` / `WorkMailbox`. That is
the useful part - the section argues about WHETHER to split without recording what a
working split actually cut along, so anyone starting fresh re-derives seams that were
already shown to compile.

`features/partial-batch-failure` -> `core-189-batch-failure-granularity.md`, which
describes "per-record result correlation was started and abandoned" WITHOUT naming the
branch that holds it. It is named now, with its sibling and what they carry -
`TerminalFailureReaction` and the rest, none of which reached master.

FOUR OF THE SIX NEEDED NO NEW NOTE, and that is the point rather than a shortcut: their
owners already exist, and a second note would be the duplicate this repository's rules
forbid. `features/consumer-interface`'s owner - `core-alternate-api-facades.md` - is in
flight on #367 across a dozen refs, so writing a copy here would collide when that
lands; the ledger links both instead. Only the consumer-side client factory was genuinely
unowned, and it gets `core-pc-owns-the-clients-it-uses.md`.

A FALSIFIED ARGUMENT, found while removing a stale line count. The actor-collection
ideation reasons from "the file has moved one line in 3.5 years, which is the clearest
evidence that branch-shaped goals don't move it", and proposes tracking progress by that
count. Measured: the God class was 1534 lines at that ideation's own date and is over a
thousand larger now - it moved in a fortnight what the argument said it had not moved in
three and a half years. The dated record is left as written per `docs/citations.md`; the
correction sits with the live decision. It does not weaken the case for decomposing, it
inverts the reason: the class is accreting, not inert.

Sizes are no longer written down in the section that kept getting them wrong. The command
is there instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01RVXBFaG4YXFP7Gjrj6pbVM
astubbs added a commit that referenced this pull request Sep 4, 2026
…e primitive, and the composition test

An external repo-level read of the current tree, mined for what it adds rather than recorded whole.
Its factual load-bearing claims were checked first: #392 is open and is the
navigator micro-MVP for soft shared-resource credits, and #333's own title
confirms the sharpest observation in the piece - the engine discovers its own ADMISSION TARGET, not a
worker-pool size.

FOUR THINGS WORTH KEEPING.

The primitive is probably not a rich mutable Execution object. Read against what #333 and #392
actually built, the deeper primitive looks like a work candidate plus an admission/eligibility
decision state - identity, position and incarnation hang off it, but the operation the engine
performs is evaluate the predicates, then atomically claim an eligible candidate. Smaller and more
faithful than an object graph with behaviour, and it matches where the implementation went without
anyone designing it that way. Recorded on the work-identity note as an assessment, not a ruling.

Blocking in user code destroys information, stated sharply enough to use. A worker that starts and
then blocks leaves the engine with one fact - worker busy - having lost the one it needed: this
record cannot proceed because a named resource is exhausted, while thirty thousand others could run.
So admission is not a performance optimisation; it is what preserves the scheduler's information
advantage, and the loss is irreversible at the moment of blocking. The operational form is better
than the buffet metaphor: first determine what is admissible, then select among the admissible.

The composition test replaces "connect the lighthouse boxes" as the decisive threshold. Do adaptive
admission, semantic eligibility, shared-resource authority and deep work knowledge compose into one
scheduler WITHOUT bespoke paths for each? It is falsifiable early, fails loudly - the tell is a
special case added for one of the four - and tests the laws rather than the feature list. Two of the
four are already in flight, which makes it live. The vision doc gains it, plus the build rule it
should have carried already: build the smallest mechanism that proves a law, and preserve the future
seam.

STRATEGY.md now has a concrete proposal rather than only a complaint. The gap is no longer breadth
but level of abstraction, and #392 is what changed the urgency: deferring was reasonable while this
was a captured conversation, less so now that a law is running machinery and the root document
answers "what are we building" differently from the branches. The proposal is deliberately smaller
than absorbing the vision doc - keep the separation, add the engine thesis in two sentences, the four
decisions, the programming/execution/control separation and embedded-rather-than-cluster, and leave
everything speculative linked out. Still the owner's call; STRATEGY.md stays untouched here.

The structural observation underneath is worth keeping on its own: the corpus has settled into laws
-> claims and notes -> implementations and experiments, which is why absorbing the vision doc into
the root strategy would be the wrong move.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012jaa7GTSsURafgg5HrGXG8
astubbs added a commit that referenced this pull request Sep 4, 2026
The open-questions register's most urgent item resolves without a decision. It read as a build
running ahead of the decisions that gate it: core-distributed-throttling.md says its gating decisions
stay open until adopted, w2-vision.md agrees the micro-MVP is gated on them, and the navigator
micro-MVP is in flight.

But "in flight" and "ahead of" are not the same thing. #392's description
already declares a dependency on #367 (and on #333), so the PR-dependency
gate refuses to merge it until the PR carrying those decisions has landed. The build proceeds on a
branch; the merge waits on the decisions. That is the sequencing working as designed. The vision
doc's composition-test section now says so where it names the micro-MVP.

What it does not resolve is recorded beside it: whether the code has taken the enforcement-fork and
standalone-versus-controller decisions implicitly. If it has, the notes should record them as taken
before #367 merges rather than after, so corpus and code agree on the day they land together.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012jaa7GTSsURafgg5HrGXG8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant