Repository navigation
feats(demos) astubbs#242: the ten per-language demos, and the evidence behind the seed's conventions - #331
feats(demos) astubbs#242: the ten per-language demos, and the evidence behind the seed's conventions#331astubbs wants to merge 90 commits into
Conversation
…ss count Two arms, per the cross-language contract: confluent_kafka.Consumer one record at a time, and this module's client library over a real sidecar it spawns as a child process. Nothing here speaks the protocol by hand - the seed's first version did, and it proved the engine worked while saying nothing about the client library, which is the artifact a user actually touches. Both entry points exist and both were run: demo/run.sh natively (with no arguments at all, and with flags), demo/run.sh --docker, and the plain `docker compose up` a reader who has never seen this repository would type. The broker is a compose sibling on both paths and the demo container is never granted the host Docker socket; the sidecar is a child process, not a compose service. THE ONE DEFAULT THAT IS NOT THE SEED'S IS --concurrency, 16 rather than 100, and it is forced rather than chosen. The proxy's executor-count function is identity today, so max_concurrency IS the executor count - and in Python an executor is a worker PROCESS. A hundred interpreters measures a laptop's memory pressure rather than the engine. Fixed rather than derived from os.cpu_count(), so two readers' fingerprints compare. The formula itself stays an open owner decision (blocker-executor-count-formula). THE CONTRACT'S PYTHON RULE IS AIMED AT THE WRONG THING, and the demo says so rather than quietly reshaping itself. It asks Python for a "non-occupying wait"; time.sleep already is one - it releases the GIL and parks on the kernel's timer, and the busy loop the rule excludes is what would pin a core per record. What a Python wait occupies is a whole process, and no wait primitive changes that: the client hands a worker one record and takes one outcome back, so an event loop inside the worker cannot overlap a second record. So the divergence lands on the default concurrency instead. Recorded in docs/inflight/clients/python.md for the integrator, not edited into the shared contract. Both arms are timed from the first record to the last, not from before consumption: poll() forks the pool, spawns the sidecar, handshakes and starts consuming in one call, so the seed's line cannot be drawn in the same place. Keeping its rule - no arm charges itself for construction - is what that costs. Measured: charging the JVM boot would have reported roughly a quarter of the arm's rate. The demo never starts a broker itself. run.sh brings up the compose broker its container path already needs, because the alternative is a Docker client library inside the demo of a Kafka client library. dependency:build-classpath needs -DincludeScope=runtime, and the failure is quiet: the default scope pulls in core's test jar, whose logback-test.xml prints logback's status report to STDOUT ahead of the `port: <n>` line the client library scans for. confluent-kafka arrives in its own `demo` extra rather than in `dev`: the library never imports it and neither the suite nor the conformance runner needs it, so the CI row's critical path is unchanged. make demo-build installs it, and run.sh and the container both go through that rather than repeating the pip line.
…rgence is the contract's fault The wave may not edit parallel-consumer-proxy/demo/README.md, so the four places this demo departs from it are recorded here for the integrator, each with the reasoning that decides it rather than the fact that it happened. One of them is not Python's to fix. The contract singles Python out for a "non-occupying wait" because a hundred sleeping processes is not a hundred sleeping threads - true - but the conclusion does not reach the wait primitive. time.sleep releases the GIL and costs no CPU; the process is the cost, and no primitive changes it, because the client hands a worker one record and takes one outcome back. The rewording is suggested here: Python's divergence is the default concurrency, not the wait. TypeScript's entry in the same list is sound and is left alone. Also recorded: that bin/ci-demo-test.sh is Java-only and hard-codes that module's paths and arm names, so the contract's "both entry points are tested" clause is met by the Java demo alone - which matters more than it looks, given that script's own header is about exactly this defect class. Generalising it needs bin/, outside this wave's scope. And the quiet one the next JVM-sidecar-spawning language will hit: dependency:build-classpath without -DincludeScope=runtime puts core's test jar on the sidecar's classpath, whose logback-test.xml then writes logback's status report to stdout ahead of the port line a client scans for. It survives only because that scan tolerates preceding lines.
…nt cannot fake
Two arms, per the shared contract: Kotlin's own Kafka client serially, and Kotlin over the
sidecar through THIS module's client library - the spawn, the handshake, the dispatch, the
report and the reap are all code in src/main/kotlin, called the way an application calls it.
The Java seed spoke gRPC by hand once and proved the engine worked while saying nothing about
the artifact users touch; that mistake is not repeated here.
It keeps the contract exactly: the seven flags, the PC_DEMO_* variables, flags beat environment
beats defaults, the effective-configuration fingerprint printed first and never the bootstrap
address, the same two tables in the same order, and no latency anywhere.
THE ONE DIVERGENCE, AND IT IS NOT COSMETIC. The simulated work is `delay`, not `Thread.sleep`.
The contract's rule is "the simulated work must use that language's non-occupying wait" and its
list then says a blocking sleep is fine in Kotlin. The list is written for a thread-per-record
client; this one runs each record as a coroutine on Dispatchers.IO, whose default parallelism is
64. A blocking sleep there caps in-flight records at 64 however high --concurrency is set, while
the fingerprint goes on printing the number the reader asked for - a throughput figure reported
against settings that did not apply, which is the exact failure the fingerprint exists to
prevent. Measured rather than argued: at --delay-ms 50 --concurrency 200, 4000 records finished
in 2.7s, which is 200 record-seconds of work in 2.7s, so about 74 records in flight - above the
ceiling a blocking sleep could ever reach. Contention can only push that figure down, never
above the thread ceiling, so the conclusion survives a loaded machine even though the rate does
not. The AK core arm still blocks, because a serial loop has nothing to interleave.
This may be a defect in the shared contract rather than a Kotlin exception: the real predicate
is not the language but whether the client is thread-per-record, which puts any coroutine-,
fiber- or async-native client in the same position. Recorded in docs/inflight/clients/kotlin.md,
not edited into the shared contract file, which is not this wave's to change.
NOT A MAVEN MODULE, AND THAT WAS FORCED. A new module needs a line in the clients aggregator
pom, a file every parallel language wave shares. So the demo lives inside this module under
demo/, compiled by a kotlin-demo profile behind -Dpc.kotlinDemo into target/test-classes - the
published surface is guarded by -Xexplicit-api=strict and a demo is not part of it. The profile
is what keeps the module's standing invariant intact: the demo hands its spawned sidecar this
JVM's classpath, so it needs the engine, and an unconditional dependency would give the module a
permanent reactor edge to parallel-consumer-proxy, whose jar `bin/build.sh`'s opening `clean`
would then delete out from under every other language's conformance test. Re-measured after the
change: `./mvnw -pl :parallel-consumer-proxy-client-kotlin -am validate` still does not print
parallel-consumer-proxy. Scala should copy this arrangement rather than re-derive it.
The logging file is called pc-kotlin-demo-logback.xml and is named explicitly in
-Dlogback.configurationFile by both entry points. parallel-consumer-core's test output is on
this classpath and carries a logback-test.xml, which logback prefers over any logback.xml
whatever the classpath order - so named the obvious thing, the demo's own configuration was
silently ignored and the first run printed core's INFO levels over the tables.
RUN, ALL FOUR WAYS, on a machine at load average 23 of 12 cores - so the absolute throughput
below is contended and should be re-measured on an idle box before it is quoted:
- native, minimal (--records 20 --replay-factor 1): both arms, exit 0
- native, NO ARGUMENTS - the double-click case that has broken before: AK core 2000 in 6.7s,
kotlin-sidecar 2000 in 1.3s then 40000 in 3.5s, exit 0
- container via run.sh --docker: image built from the repository context (9.17MB uploaded, so
.dockerignore is working), broker a compose sibling, sidecar spawned inside the demo
container, no host Docker socket anywhere, exit 0
- container via plain `docker compose up`, no arguments and no environment: AK core 2000 in
7.6s, kotlin-sidecar 2000 in 1.1s then 40000 in 5.0s, demo container exited 0
Argument handling checked without a broker too: an unknown flag exits 2 with usage rather than
reporting numbers nobody asked for, and PC_DEMO_RECORDS=7 with --records 11 printed records = 11.
STILL OPEN: nothing runs this demo's two entry points in CI. bin/ci-demo-test.sh does that for
the Java demo only, and it is a shared bin/ script rather than this module's to extend. That is
the largest gap and it is recorded in docs/inflight/clients/kotlin.md.
The Go copy of the reference demo: the same records through franz-go one at a
time, and through the Go client library over a real sidecar the library spawns
as a child process. Two arms, which is the whole contract outside Java.
It keeps parallel-consumer-proxy/demo/README.md exactly - the seven flags with
their PC_DEMO_ variables and flags-beat-environment-beats-defaults precedence,
the same defaults, the effective-configuration fingerprint printed first and
never carrying the bootstrap address, the two tables with the same columns in
the same order, the big replay over the parallel arms only, and no latency
anywhere.
THROUGH THE CLIENT LIBRARY, NEVER HAND-WRITTEN gRPC. The Java seed learned this
the expensive way: an arm that speaks the protocol by hand proves the engine
works and says nothing about the client library, which is the artifact users
actually touch. This arm calls Open/Poll/Close the way a user's program would.
THE DEMO BINARY NEVER STARTS A BROKER - run.sh does, and that is the one place
the shape differs from Java. Java's DemoBroker falls back to Testcontainers;
Go has no comparable dependency here, and adding one would put a Docker client
library into a demo whose point is that the application needs no
infrastructure. So natively run.sh starts a cp-kafka container on a random free
port and removes it on exit; in the container the compose sibling is the
broker; the binary always receives an address. The user-facing promise is
unchanged - omit --bootstrap and a broker appears - and the rule that produced
the container half is untouched: A DEMO CONTAINER IS NEVER GRANTED THE HOST
DOCKER SOCKET, and the sidecar is a child process rather than a compose
service.
Two things measured rather than reasoned about, both of which cost a run:
- Kafka refuses 0.0.0.0 as a KRaft CONTROLLER listener's bind address, not
only as an advertised one. The format step aborts with "advertised.listeners
cannot use the nonroutable meta-address" before the broker starts, so the
native broker binds CONTROLLER://localhost.
- GOTOOLCHAIN=local plus an older Go cannot even PARSE the client library's
go.mod ("unknown block type: tool"), so run.sh probes the module with
`go list -m`, retries once with GOTOOLCHAIN=auto, says out loud that it is
doing so, and only then falls back to the container.
THE DEMO IS A NESTED GO MODULE. Go's module graph propagates requirements to
every consumer, so franz-go in the library's go.mod would hand a Kafka client
library to applications whose entire reason for using the proxy is not needing
one. The consequence is worth knowing: `go build ./...` in the parent does not
descend here, so Maven never compiles the demo - run.sh and the Dockerfile do,
and both were run.
Run, not merely written. Native and container entry points both exit 0 with
both arms reporting rows; the no-arguments path - the case this family has
broken on before - was exercised through run.sh with no arguments at all and
through the container with an empty PC_DEMO_ARGS; the big-replay table was
exercised at --replay-factor 2. Throughput figures were deliberately not
measured: ten agents were on the machine, so any number would have described
the load rather than the arms.
… found in the contract
The Go demo's own divergences are in its README, where a reader meets them.
This records the half a later session needs and no command can answer: why each
was taken, and the five places the shared contract at
parallel-consumer-proxy/demo/README.md left the next nine languages to
rediscover something.
Report only, per KTD23 - the contract is the integrator's to change:
1. "Omit --bootstrap to start one" inherits Java's Testcontainers fallback
without saying WHO starts the broker. Nine languages will each decide it
alone, some with a heavyweight dependency, some by moving it into run.sh
as Go did.
2. Every non-JVM demo container is a TWO-TOOLCHAIN image, and the contract
does not say so. The sidecar is a JVM application spawned as a child of
the running demo, so it cannot be a build-stage artifact the way the
language's own binary can.
3. The flag/environment table has no room for the sidecar's location, which
every non-JVM demo needs in some spelling. Go's is
PC_DEMO_SIDECAR_CLASSPATH, under the prefix the contract reserves for
flags - an integrator diffing variables cannot tell from the contract
whether it is a divergence or a shared necessity.
4. bin/ci-demo-test.sh runs the Java demo only, while the contract says both
entry points are tested on every pull request and that a per-language demo
inherits this. Ten demos can ship untested while the contract says
otherwise.
5. The big replay's table is a single row everywhere except Java, and its
"vs AK core*" column then compares the only arm against one not in the
table. That is what the contract asks for; it should say the case is
expected rather than a bug.
Also recorded: the two measured facts that cost this wave a run each - Kafka
refusing 0.0.0.0 as a KRaft controller listener's BIND address, and
GOTOOLCHAIN=local making the client module unparseable on an older Go - and one
open follow-up, that staticcheck does not yet cover the nested demo module.
…that works The crate's build.rs resolves protoc in a documented order - $PROTOC, then PATH, then the copy the protocol module's protobuf-maven-plugin downloads into the local Maven repository. The Maven fallback selected a candidate on is_file() alone, and Maven stores that artifact 0644: the plugin chmods the copy it extracts for its own use, never the one in ~/.m2. Because that fallback is consulted BEFORE PATH, a machine with a perfectly good protoc installed still failed - with "Permission denied" naming a path under ~/.m2, which reads as a corrupt download rather than as a permission bit. Found by building the Rust demo on exactly such a machine. The candidate filter now requires the executable bit, so an unexecutable copy falls through to PATH instead of being preferred over it. Where else the class could live, checked: the Go client's scripts/generate-proto.sh resolves the same ~/.m2 directory the same way and already chmod +x'es what it finds, so it does not have the defect - it fixes it by mutating the user's local Maven repository, which this change declines to do. No other client resolves that directory.
The same records through kafkajs one at a time, and through this module's client library over a sidecar it spawns as a child process - with the contract's seven flags, its PC_DEMO_ environment variables and their precedence, its fingerprint without the bootstrap address, its two tables, and no latency anywhere. THE SANCTIONED DIVERGENCE, AND WHAT IT IMPLIES. Node is a single event loop, so a blocking sleep would stop the transport, the executors and the timers at once and the "parallel" arm would run exactly as serially as the serial one. The work is an awaited timer, which means the concurrency this demo shows is promise concurrency: Configured.executor_count concurrent awaits, not threads. That is the wave-one concurrency decision exercised end to end for the first time, and the demo says out loud that a synchronously CPU-bound processor would collapse the arm - it does not pretend otherwise. TWO DIVERGENCES THAT WERE NOT SANCTIONED. run.sh starts the broker rather than the demo, because the containerised path is always "an address was supplied" anyway - a demo container is never granted the host Docker socket - and the alternative was a 47 MB @testcontainers/kafka dependency in a package tree whose npm ci sits on the CI matrix's critical path. And the demo is its own npm package reaching the library through file:.., so it loads the built dist/ a user would install, and kafkajs is the demo's dependency and never the library's: a client library that pulled in a Kafka client would contradict the arm it exists to demonstrate. RUN, NOT JUST WRITTEN. Both arms completed at 100 records and again with no arguments at all under PC_DEMO_ variables, exit 0 both times, with the small and big replays printing. The figures are start-up dominated at that volume and are recorded in the inflight note as explicitly not throughput measurements; ten agents were on the machine, so measuring was not the job. The run found two things reading would not have. A KRaft broker started with listeners on 0.0.0.0 exits 1 in preflight, because the CONTROLLER listener has no advertised entry and is taken from `listeners` - the compose sibling already binds to its service name for this reason, and the standalone docker run now does too. And kafkajs 2.2.4 prints TimeoutNegativeWarning on Node 25; cosmetic, but it is the shape of problem an unmaintained client keeps producing. eslint.config.mjs gains the demo's tsconfig as a typed-lint project rather than ignoring the directory: `eslint .` is the CI matrix row's exact command and goes red without it, and no-floating-promises is precisely what an arm racing a countdown against a session's end needs. It immediately found an unused parameter. Still open, and recorded rather than hidden: nothing runs this demo in CI - bin/ci-demo-test.sh was not this agent's to edit - and there is no ReferenceDemoIT equivalent.
…e it blocks
The same records through Rust's own Kafka client and through Rust over the
sidecar, keeping the contract in parallel-consumer-proxy/demo/README.md: same
seven flags, same PC_DEMO_* variables with flags beating environment beating
defaults, same defaults, the effective-configuration fingerprint first and
without the bootstrap address, the two replays, the two tables, no latency.
demo/run.sh # native or container, it picks and says which
cd demo && docker compose up # the plain container path
THE SIDECAR ARM GOES THROUGH THE CLIENT LIBRARY, not through hand-rolled gRPC.
This crate has no protobuf dependency and opens no socket: it hands
ClientOptions an absolute path and arguments, and the library spawns the child,
reads its port line and supervises it. The Java seed hand-rolled the protocol
and had to be rewritten, because it proved the engine worked and said nothing
about the artifact users actually touch.
A CRATE OF ITS OWN, BESIDE THE CLIENT RATHER THAN INSIDE IT. The demo needs
rdkafka for its AK core arm, for creating the topic and for seeding; the client
library must never need a Kafka client at all, because an application on the
sidecar path does no Kafka I/O. A src/bin/ target would have shared the
library's dependency list and quietly contradicted the claim the demo exists to
make.
WHAT IS SPECIFIC TO RUST, MEASURED RATHER THAN ASSERTED. The contract says a
blocking sleep is fine in Rust, unlike Python and TypeScript. It is - and where
it blocks decides whether the engine's ceiling means anything, because the
library's executors are tasks on an async runtime. One term changed, everything
else identical (12 cores, 40,000 records, --delay-ms 2, --concurrency 100):
blocking(|r| thread::sleep(..)) 10,341 msg/s
async move { thread::sleep(..) } 3,518 msg/s
The low figure is about the core count divided by the delay - the predicted
ceiling for a runtime whose workers are all asleep, stated before the run. So
the demo uses the library's blocking() entry point, and the README says why.
THE BROKER IS A COMPOSE SIBLING ON BOTH PATHS, AND THE CONTAINER NEVER GETS THE
HOST DOCKER SOCKET. The reference starts its native broker from inside the JVM
with Testcontainers; Rust's equivalent would put a Docker API client in the demo
binary's dependency tree for the one path that already has Docker, so the native
path starts the `broker` service from this directory's own compose file and
reaches it over a host listener. One broker definition serves both paths. In a
container the demo refuses to start a broker at all rather than reaching for a
socket. The sidecar is not a compose service either: the client library spawns
it as a child process, in the container exactly as natively.
Run natively at the defaults - no arguments, the case that has broken before -
on a 12-core machine: AK core 2,000 records in 6.4s (314 msg/s), rust-grpc the
same 2,000 in 0.9s (2,292 msg/s, 7.3x), and the big replay's 40,000 records in
3.9s (10,341 msg/s). Those are start-up-inclusive figures from a loaded machine,
recorded because they were observed, not as a benchmark.
… test one The container run of the demo printed "No SLF4J providers were found" and the native run did not, which is not a difference between containers - it is the demo's two entry points handing the sidecar two different classpaths and calling the result the same demo. build-classpath with no includeScope writes EVERY scope, and logback-classic is test scope repository-wide, so the native path was quietly flattered by a logging provider the product does not ship. run.sh now scopes it to runtime. Both entry points run bz.stub.parallelconsumer.proxy.Main on the classpath a user would deploy, and both print the warning. THE WARNING IS A FINDING, NOT NOISE TO SILENCE: the sidecar a user deploys has no SLF4J provider at all, so every log.info in it goes nowhere. That belongs to the proxy module rather than to a demo wave, and it is recorded for its owners in the inflight note. The demo's README lists it, with the other two expected noise lines, so a reader is not left guessing which of them means something. Re-run afterwards at --records 80 --replay-factor 2: both arms, both replays, exit 0. The container path had already run both arms and both replays to exit 0 before this change, and the inflight note now carries those figures - marked, as the rest are, as not being measurements.
…emo people will run Two things the first container run taught, one of them predicted beforehand. THE .git WALK. The demo locates its own compose file by walking up to the repository root, and the repository-root .dockerignore excludes .git from the build context - so inside the demo's own image that walk cannot succeed. It was predicted by reading the two against each other, and the image built from the pre-fix source then confirmed it exactly: The demo failed: no git working tree above this process's working directory printed after the fingerprint and before the first Kafka call, on a container whose sidecar classpath and JVM had both resolved perfectly well from the environment. Only the branch that starts a broker natively needs the compose file, so that resolution became lazy in the commit that added the demo. What is here is the other half: repository_root() now takes what the caller was looking for and says so. It has two callers, and its old message named the environment variable belonging to the OTHER one - which in the container was set, so the message pointed at the one thing that was not the problem. THE IMAGE. Cargo's target directory for this crate is gigabytes of intermediates and the registry a few hundred megabytes more, none of it reachable once the binary is copied to /app. They are now removed in the same layer that produces them, because removing them in a later one would save nothing. This is not tidiness: exporting the image took longer than building it on a loaded machine, and every gigabyte is time a reader spends before seeing a table. The Maven downloads stay - the sidecar's classpath points into /root/.m2. Any language whose demo derives paths from the repository root has the first trap; any that builds a compiled toolchain in its image has the second.
…as no demo deps Putting demo/tsconfig.json in eslint's typed-lint projects is the version that should be right - no-floating-promises is exactly what an arm racing a countdown against a session's end needs, and while it was wired up it immediately caught an unused processor parameter. It is still wrong here, and only running it both ways showed why. The demo is a separate npm package, so the typed rules can resolve kafkajs only when demo/node_modules exists - and the CI matrix row installs the library's dependencies alone. Measured: with the demo's project listed and its node_modules present, `npx --no-install eslint .` exits 0; with it absent, 62 no-unsafe-* errors and exit 1. That is a lint green on a developer's machine and red in CI, which is worse than one that does not run, and making the project list conditional on the directory existing would be a gate that silently disappears - the failure mode this repo already has a rule about. So demo/** is ignored, the measurement is written into eslint.config.mjs where the next person to try this will read it, and the demo's gate is tsc under the same strict settings. If an integrator wants the demo linted, the lever is the CI row installing demo/'s dependencies, not the config.
…t sits outside The per-language note now carries what the demo wave learned, because a later session picking Rust up will otherwise rediscover all of it. The finding worth carrying to the wave sync is the blocking-sleep carve-out. The contract exempts Python and TypeScript from a blocking sleep and names Rust among the languages where one is fine. It is - and where it blocks decides whether the engine's ceiling means anything, because this client's executors are tasks on an async runtime. One term changed, everything else identical: 10,341 msg/s through the library's blocking() entry point against 3,518 msg/s for the same sleep inside the executor task, the low figure being about the core count divided by the delay. The same question is owed to C#, TypeScript and asyncio, so the contract may want "through whatever the client library offers for blocking user code" rather than "a blocking sleep is fine". Also recorded: the sidecar's logs are invisible to a Rust application, because the proxy has no logback configuration and logback's default appender writes to stdout - which is the lifecycle channel the client drains and discards; the two Debian packages and the .git absence a container build meets; and the protoc defect, with the Go client checked for the same class and found already immune by a route this fix declines to take. Two things are left open and are outside this branch's ownership: no CI row runs this demo, and no gate reaches the demo crate at all - the module's pom runs clippy and cargo test in the client crate's directory, one level up.
The by-hand build steps read as one command in the demo directory, which skips the step that matters: the file:.. link resolves to the library's dist/, so the library has to be installed and compiled first or the demo cannot import it. run.sh and the Dockerfile already do them in that order; the README now says so.
… to copy The contract in parallel-consumer-proxy/demo/README.md, rendered in Scala: the seven flags with the same defaults, the same PC_DEMO_* variables with flags beating environment beating defaults, the effective-configuration fingerprint printed first and never carrying the bootstrap address, the two tables in the same order with the same columns, and no latency reported anywhere. TWO ARMS, AND DECLINING THE OTHER FOUR IS THE DECISION THIS COMMIT MAKES. AK core is a plain KafkaConsumer driven from Scala, one record at a time; scala-grpc is this module's own ParallelConsumerClient over a sidecar the client library spawns as a child process. Scala is a JVM language, so it could have run pc-core, java-direct, java-grpc-uds and java-raw-grpc against the same broker in the same process exactly as the Java seed does - which is precisely why the temptation had to be named and declined. The contract's whole value is that a reader who has run one language's demo has run them all, and a six-row Scala table beside a two-row Ruby table would not be that. Those arms are the seed's alone. IT GOES THROUGH THE CLIENT LIBRARY, NOT THE WIRE. Nothing under demo/src/main/scala names a protobuf message, a channel or a token - it opens a ParallelConsumerClient, hands it a ClientOptions, and returns a Future[Outcome] per record, which is the whole of what a Scala user writes. An earlier version of the Java seed spoke the protocol by hand and had to be rewritten, because it proved the engine worked and said nothing about the artifact users actually touch. THE DEMO IS BEHIND A MAVEN PROFILE, AND THAT IS LOAD-BEARING. The sidecar arm has to hand its child process a classpath carrying parallel-consumer-proxy. Declared unconditionally that is a permanent reactor edge to the engine, and bin/build.sh opens with clean - which would delete the sidecar jar every other language's conformance test spawns. So the demo's sources are a test source root added only by the scala-demo profile, and its dependencies live there too. Verified with a control arm rather than asserted: the ordinary reactor is 8 modules with no Language Proxy in it, and adding -Dpc.scalaDemo makes it 9 with the engine present. One term changed, the outcome flips. demo/logback.xml is not cosmetic. With no logging configuration anywhere on the classpath, logback falls back to root at DEBUG: the first run of this demo buried both its tables under more than four thousand lines of Netty frames and docker-java headers. One of the levels it sets is a rule rather than a preference - every Kafka client logs its full effective configuration at INFO when constructed, bootstrap.servers included, which would print the address the fingerprint deliberately omits. It is pointed at with -Dlogback.configurationFile rather than shipped as a resource, because target/test-classes is also what scripts/conformance-runner runs from and the demo's logging preferences must not silently become the shared suite's. The container keeps both rules that are not negotiable: the broker is a compose sibling and the host Docker socket is never mounted, and the sidecar is a child process rather than a compose service. Run natively, under heavy concurrent load, so no throughput figure from these runs means anything and none is recorded: --records 20 --replay-factor 2 completed both arms and rendered both tables at exit 0; the demo with NO arguments, configured entirely through PC_DEMO_*, did the same, which is the case bash 3.2 under set -u has broken before; --help and a misspelled flag both reach the usage text and the misspelled flag exits 2. The container path is written but was NOT run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
… that are not Scala's Records the demo against the Scala client note, with what was actually run and - more usefully - what was not: the container path is written and unproven, and the demo has no row in bin/ci-demo-test.sh, which is a shared file this branch does not own. The contract is explicit that a demo with one tested entry point has an untested entry point, so that gap is stated as unproven rather than assumed fine. Three findings, and the first two belong to every JVM demo including the Java seed rather than to Scala: - With no logback configuration on the classpath, logback's fallback is root at DEBUG. Measured, not reasoned about: a fifty-record run emitted over four thousand lines of Netty frame dumps, docker-java HTTP headers and Kafka client configuration, with both tables buried in the middle. - Every Kafka client logs its full effective configuration at INFO when constructed, and bootstrap.servers is in it. The contract's rule that a demo never prints the bootstrap address exists because own-cluster mode puts a user's real broker there; the fingerprint honours it while the client's own dump prints it anyway. The rule may be worth restating as binding the whole run rather than only the fingerprint block - recorded here rather than edited into the shared contract. - SidecarCommand requires an absolute path to an executable and a JVM sidecar is a jar, so "this JVM plus a -cp argument" is now written in three places. That reinforces the existing note that spawning belongs to the Java lifecycle unit rather than to each client. The profile arrangement is recorded with its control arm and with the reason the documented verification command currently fails for something unrelated - a reader who greps for BUILD SUCCESS would otherwise conclude the check is broken. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
… could not file where it belongs Running the Kotlin module's default lane on a ten-agent box (load average 83 of 12 cores) failed BlockedThreadAsserterTest.functionThatReturnsOnItsOwnScheduleIsRejected in parallel-consumer-core. Established as contention with a control rather than assumed: the class re-run alone passed 7/7, and this module's own tests are green. It is a NEW signature on a helper that already has prior art - the existing entry in test-untracked-ci-flakes.md is about assertUnblocksAfter measuring a window two milliseconds short (owned by #262), whereas this is the helper's own self-test failing to reject a timer-returning function. It belongs in that file. It is written into clients/kotlin.md instead because the Kotlin demo wave owns only that file, and a sighting recorded in the wrong place is still worth more than a sighting that dies with the terminal scrollback.
…need demo/run.sh in the .NET client module: the same records through Confluent.Kafka one at a time (AK core), and through this module's client library over a sidecar the library spawns itself (dotnet-grpc). It mirrors the contract in parallel-consumer-proxy/demo/README.md exactly - same flags and defaults, same PC_DEMO_* variables with the same precedence, the effective configuration printed first and never the bootstrap address, the same two tables, no latency anywhere. IT GOES THROUGH THE CLIENT LIBRARY, NOT THE PROTOCOL. That is the whole point of the Java seed's rewrite: a demo that speaks gRPC by hand proves the engine works and says nothing about the artifact users actually touch. THE ONE DIVERGENCE, AND IT CONTRADICTS THE CONTRACT'S OWN LIST. The contract names C# among the languages where a blocking sleep is fine as the simulated work, and exempts only Python (worker processes) and TypeScript (one event loop). For this client that list is wrong: the library's executors are Tasks on the thread pool, so Thread.Sleep occupies a pool thread, and at the contract's default --concurrency 100 the sidecar arm would report the pool's thread-injection rate rather than the engine's throughput. Both arms use await Task.Delay instead - both, because the serial arm is the denominator of every ratio and must not differ from its numerator by the wait primitive as well as by the transport. The contract is recorded as wrong, not edited; the rule it should state is about the client's concurrency shape rather than the language's name, which puts Kotlin and Swift under the same question. THE BOOTSTRAP ADDRESS IS NORMALISED, AND FINDING OUT WHY COST A WHOLE RUN. Testcontainers for .NET returns its address as a URI - measured here as plaintext://127.0.0.1:62347/, lower-cased scheme and trailing slash. librdkafka accepts that verbatim, so the AK core arm worked with it untouched; the Java client behind the sidecar rejects it, and the trailing slash survives naive scheme stripping. The arm died at session construction with a deliberately reason-free error, because R48 withholds the cause - a Kafka ConfigException embeds property values and those may be credentials. Two findings against the product came out of that hour and are recorded in the inflight note: the sidecar's own log is unreachable from any client (logback's default appender writes to stdout, which is the lifecycle channel the client library drains and discards), and R48 could name the offending property KEY without its value and give nothing away. THE CONTAINER TAKES TWO TOOLCHAINS, which no other client's image needs: the demo is .NET and the sidecar it spawns is a JVM program in this repository rather than a shipped binary. Broker as a compose sibling, never the host Docker socket (U35); the sidecar spawned as a child process, never a compose service (KTD41). .dockerignore covers target/ and not .NET's bin/ and obj/, so the image removes them itself rather than editing a file ten concurrent worktrees share. WHAT WAS ACTUALLY RUN: the native entry point, twice - once with no arguments at all (the case that has broken before, scaled down by PC_DEMO_* variables) and once with explicit flags and a big replay. Both arms completed, both tables printed, exit 0. Twenty records, on a machine running ten agents: enough to prove the machinery, nowhere near enough to measure anything, and nothing in this branch should be read as a measurement. The JVM lookup the demo and the conformance harness both need is now one file, compiled into both projects by link. The demo project joined the module's one solution rather than adding a second, so the ordinary dotnet build - and the clients CI row - keep it compiling under the same analyzers-as-errors lint as the library.
…, in one container The contract's two arms, in C++: AK core is librdkafka consuming one record at a time, and cpp-grpc is this application as a foreign client, going through THIS MODULE'S client library, which spawns the sidecar as a child process. No hand-written gRPC anywhere - the Java seed was written that way first and had to be rewritten, because it proved the engine worked and said nothing about the client library, which is the artifact users actually touch. Everything else is mirrored deliberately: the same seven flags with the same defaults, the same PC_DEMO_* variables with flags beating the environment beating the defaults, the effective-configuration fingerprint printed first and never carrying the bootstrap address, the same two tables in the same order, and no latency reported. The broker is a compose sibling and the demo container is never given the host Docker socket; the sidecar is a child process rather than a compose service. WHAT C++ COULD NOT MIRROR, AND WHY There is no native mode. gRPC, protobuf and librdkafka arrive as system dev packages rather than as a versioned toolchain - the same fact that makes bin/build-client.sh build this language in an image - so the container IS the toolchain here. `run.sh --native` answers with that reason rather than failing as an unknown flag, because a reader arriving from the Java demo will type it. It follows that the broker is always supplied from outside: Java starts one with Testcontainers when --bootstrap is absent, and a demo container that is never granted the Docker socket could not do that even if C++ had Testcontainers. A missing address prints an explanation and exits 2 rather than running something different. The sidecar is a JVM program and a foreign client library spawns a binary by absolute path, so the image installs a four-line launcher and PC_DEMO_SIDECAR names it. `exec` in that launcher is load-bearing: the client holds the write end of the child's stdin and EOF there is the proxy's parent-death signal, so a wrapper that forked would leave a live JVM holding group membership behind a dead application. It is an environment variable and not an eighth flag, because the contract fixes the flag list at seven and where the binary lives is a property of the image. TWO THINGS THE IMAGE DOES THAT ARE WORTH KNOWING The runtime stage is built FROM its own toolchain stage, which is the opposite of ../Dockerfile: that one exports statically linked artifacts to a host and runs nothing, so it pays for the static link's absl archive group to make a portability claim. Nothing is extracted here, so there is no claim to prove and the binary is dynamically linked against the libraries it was compiled with. The Maven local repository is a BuildKit cache mount, which the Java demo's Dockerfile records that it could not have: that image computes a classpath file pointing into /root/.m2, so a cache mount - not part of the resulting image - would leave it naming jars the container does not have. This one copies the jars out and uses a wildcard classpath, which has no such coupling. It also adds an slf4j-simple binding, because every binding in this reactor is test-scoped and a sidecar that logs nothing silences exactly the diagnostics the client library inherits its stderr to preserve. Measured, not assumed: without it the spawned proxy printed `No SLF4J providers were found` and then nothing. Tested by running it: the image builds (which for C++ is what "it compiles" means), the demo's own option tests run inside that build, and one end-to-end run at CI's own demo volume put records through both arms.
…hing runs Records the three places C++ had to diverge from the reference demo and why each was forced rather than chosen - no native mode, a broker that always arrives from outside, and a sidecar that needs a launcher script because a foreign client library spawns binaries and the sidecar is a JVM program. Also records what is NOT wired up, which matters more than the divergences: neither entry point is in bin/ci-demo-test.sh, because bin/ and .github/ were outside this wave's ownership. The contract says a demo with one tested entry point has an untested entry point, so this is a gap rather than deferred polish - and C++ is the cheapest language to add there, having no native path to test separately.
…ar, one command The Ruby transcription of the contract in parallel-consumer-proxy/demo/README.md: same flags, same PC_DEMO_ environment variables with the same precedence, same two tables in the same order, the effective configuration printed first and never the bootstrap address, no latency reported. Two arms, which is the whole contract outside Java. AK core is the rdkafka gem, one record at a time. ruby-grpc is THIS MODULE'S CLIENT LIBRARY over a sidecar it spawns - not hand-written gRPC, which is the mistake the Java seed had to be rewritten out of: it proved the engine worked and said nothing about the artifact users touch. WHY rdkafka AND NOT ruby-kafka. librdkafka behind FFI is what a Ruby application consumes Kafka with today - Karafka is built on it - and ruby-kafka has been archived by its authors since 2023. A comparison whose serial arm is an unmaintained gem flatters the sidecar for a reason that has nothing to do with Parallel Consumer. It also ships precompiled for linux-gnu, so the container builds it without a C toolchain. THE BLOCKING SLEEP WAS CHECKED, NOT INHERITED. The contract lists Ruby among the languages where a blocking sleep is honest and names Python and TypeScript as the exceptions. It holds here for a reason that had to be confirmed against this client's own design rather than assumed from the list: the executors are threads, and MRI releases the GVL around sleep, so N executors sleeping is N records in flight. Had this wave copied Python's worker processes, the contract's list would have been wrong for Ruby. THE CONTAINER NEVER GETS THE HOST DOCKER SOCKET (U35). The broker is a compose sibling, and the sidecar is a child process rather than a compose service (KTD41) - a service would teach a deployment the product does not ask for. The image is two stages: a JDK stage that builds the sidecar out of the reactor with the Maven repository as a CACHE MOUNT, and a Ruby stage carrying a JRE. Only jars cross. The cache mount is possible here and was not in the Java demo's image, because this one copies the jars it needs out of the cache rather than running a classpath that points into it. THREE THINGS THIS DEMO DOES THAT THE SEED DOES NOT, each because Ruby forced the question: - The BROKER IS STARTED BY THE ENTRY POINT, not by the demo. Ruby has no Testcontainers equivalent worth depending on, so run.sh starts the same compose broker the container path uses - one broker definition, not two - and hands the address in. Natively it reaches its HOST listener on 29092, chosen so a broker already on 9092 is left alone. - COMPOSE FORWARDS EVERY PC_DEMO_ VARIABLE, not just BOOTSTRAP and ARGS. Compose forwards nothing it is not told to, so without this the contract's "environment beats the defaults" silently stops being true on the container path - which is the path a reader without Ruby always takes. Recorded against the shared contract in docs/inflight/clients/ruby.md rather than edited into it. - The ENTRYPOINT passes "$@" through, so `docker compose run demo --help` reaches the demo's own parser instead of landing in $0 and being discarded. PC_DEMO_SIDECAR_CLASSPATH and PC_DEMO_SIDECAR_JAVA are plumbing, not dials: they have no flags on purpose, because a flag would invite pointing the demo at an arbitrary binary - the decision the client library's SidecarCommand deliberately refuses to make. The module's own checks now cover the demo: rake syntax parses it (named files, not demo/**, which would parse every gem in demo/vendor), and RuboCop lints it with demo/vendor excluded. RUN, on a machine shared with nine other demo builds: the container path at 20 records with the big replay skipped and again with it running, the no-argument path configured entirely through PC_DEMO_ variables, and --help and an unknown flag through both run.sh and the image (exit 0 and exit 2). Both arms completed every time. No figure from those runs is a measurement - twenty records at 1ms is group-join cost. The native path is unrun; this machine has Ruby 2.6.10 and the library's floor is 3.2.
…does to the contract The same records through Swift's own Kafka client and through Swift over the sidecar, mirroring the Java seed's interface exactly: the seven flags with their defaults, a PC_DEMO_ environment variable for each with flags beating environment beating defaults, the effective-configuration fingerprint printed first and never carrying the bootstrap address, the two tables in order, and no latency anywhere. The sidecar arm goes through the client library in this module, never the protocol by hand - that was the seed's original mistake and it proved the engine while saying nothing about the artifact users touch. Container-only, because there is no Swift toolchain on a developer box here. The broker is a compose sibling, the demo container is never given the host Docker socket, and the sidecar is a child process rather than a compose service. THREE DIVERGENCES, all recorded rather than quietly taken. The simulated work is Task.sleep, not the blocking sleep the contract permits in Swift. poll() starts executorCount Swift concurrency TASKS on the cooperative pool, whose width is the core count, so a blocking sleep in the user function would cap the arm at core-count in flight however large a ceiling it asked for - the table would report the pool while appearing to report the engine. This module already recorded the same mechanism biting the conformance runner's ceiling barrier. It is reasoned, not measured: no blocking-sleep control arm has been run. --partitions reaches the broker's num.partitions through docker-compose.yml rather than a CreateTopics call, because swift-kafka-client has no admin client and no public metadata API. The flag works on the supported path; what it cannot do is verify, and with --bootstrap pointing at someone else's cluster it has no effect at all. There is no committed Package.resolved for the demo package, since resolving needs a toolchain that does not exist outside the image. The image copies the one it produced to /app/Package.resolved. The demo is its own SwiftPM package with a path dependency on the client, not a target of it: its AK core arm drags in librdkafka, zstd, OpenSSL and SwiftNIO, which have no business in the dependency graph of every consumer of the client library. The path dependency is also what makes it legal, since the library's targets use unsafeFlags. RUN, not just written. Both arms end to end in the container twice - once with the big replay skipped and once with it running, so both tables and the across-replays footnote were exercised - plus the image's no-argument, --help and misspelled-flag paths, and run.sh's own parsing against a stubbed docker. Exit codes 0, 1 and 2 all observed where intended. Those volumes prove the machinery and no number: the sidecar arm read 0.9x over 20 records and 2.7x over 60, which is the same code saying two different things. A default-scale run on an unloaded machine is still owed, as is wiring this demo into bin/ci-demo-test.sh, which today runs the Java one only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
…thing has run Records the demo's five decisions where a later session will look for them rather than in the diff: rdkafka over the archived ruby-kafka, the blocking sleep confirmed against this client's thread executors rather than inherited from the contract's list, the broker started by the entry point because Ruby has no Testcontainers equivalent, the two sidecar variables that deliberately have no flags, and the two toolchains a native run needs. TWO THINGS AN OWNER HAS TO DECIDE, both filed here rather than acted on: The contract promises that every flag has a PC_DEMO_ variable and that the environment beats the defaults. Compose forwards nothing it is not told to, and the Java seed's compose file declares only BOOTSTRAP and ARGS - so on the container path every other variable is silently dropped, and that is the path a reader without the language's toolchain always takes. This module forwards all of them; either the seed should too, or the contract should say the environment is native-only. The same gap is in whichever of the ten demos copied the seed's compose file literally. blocker-executor-count-formula.md does not block Ruby. The identity function turns --concurrency 100 into 100 executors, which in Python is 100 worker processes and here is 100 threads. Nothing needs capping in this demo, and the run confirmed the granted count is exactly what was asked for. WHAT NOTHING HAS RUN, stated so it is not mistaken for coverage: the native path, because this machine ships Ruby 2.6.10 and the library's floor is 3.2 - which leaves the Maven classpath step, the compose broker's host listener and a source build of rdkafka unexercised; the demo at its own defaults, deliberately deferred to an unloaded machine; and CI, because bin/ci-demo-test.sh is Java-only and outside this wave's file scope. Whoever owns bin/ should settle how the fan-out's ten demos get their both-entry-points run, once rather than ten times.
…, so the demo's image could not exist Grpc.Tools 2.71.0 ships a bundled linux_arm64 protoc that SEGFAULTS - MSB6006, "exited with code 139" - when MSBuild spawns it inside a container on an Apple Silicon host. That is the whole module, not just the demo: no image containing this client library could be built on such a machine, and the demo's container is the first thing that ever tried. ESTABLISHED WITH A CONTROL, NOT BY UPGRADING AND HOPING. One container, one four-line .proto, one project, only the Grpc.Tools version changed: 2.71.0 died at MSB6006, 2.83.0 generated the stubs. Two further facts place the defect outside this repository. The same protoc, run BY HAND with MSBuild's exact command line - captured with -consoleLoggerParameters:ShowCommandLine, flags bisected one at a time - succeeded at both versions; and the same minimal project built under --platform linux/amd64 at 2.71.0. So it is protoc-spawned-by-MSBuild on linux/arm64, which is why CI's amd64 runners never saw it and only a developer on Apple Silicon would. The bump is otherwise inert: ProtobufVersion stays at 3.31.1, the library builds warning-free under the same analyzers-as-errors lint, dotnet format --verify-no-changes is clean, and the module's conformance test still passes against the real wire. The control is written beside the version in Directory.Build.props, because the next person to see a version pin four releases behind the latest will want to know whether it may move down. With this, demo/run.sh --docker runs end to end: both arms, both tables, demo container exit 0, broker as a compose sibling.
…issue-ref gate TWO THINGS, BOTH FOUND BY VERIFYING RATHER THAN ASSERTING. 1. ReactorPCTest.concurrencyTest failed locally at "expected < 1200 but was 1266" while the full Maven reactor ran alongside two Docker image builds on a 12-core box. Diagnosed as contention WITH A CONTROL ARM, not by assumption, and recorded because no entry for it existed anywhere: not the inflight ledger, the quarantine registry, docs/solutions, docs/plans or the hardening audits. branch, uncontended: 5 pass / 0 fail 1000, 1073, 1000, 1062, 1000 master, uncontended: 5 pass / 0 fail 1000, 1000, 1076, 1000, 1000 Indistinguishable, both well under the threshold, run SEQUENTIALLY because running the arms together would reintroduce the variable under test. The control was warranted rather than skipped, and the note says why: this branch's stack modifies ShardManager, WorkManager, ProcessingShard and WorkContainer - the classes that bound concurrency - adding the language proxy's abandonment path. There is even a route from that change to this exact symptom: onAbandonedResult decrements numberRecordsOutForProcessing, and a DOUBLE decrement would under-count in-flight work, let extra records out, and read as concurrency exceeding its maximum. The author already guarded it, with an early return placed before an unconditional decrement. It is also unreachable from this test - nothing in core or reactor calls the abandon path - which is the half the master baseline confirms empirically. The note says what not to do: do not raise MAX_CONCURRENCY_OVERFLOW_ALLOWANCE and do not quarantine. Quarantine needs a diagnosis and is master-state; this fails on neither master nor an idle box. Loosening the allowance would delete the only signal that the bound is real, in the library whose whole purpose is bounded concurrency. 2. The gate was RED on a commit already pushed, and I had not re-run it after writing that file. It flagged five "unqualified issue references" that are JVM thread ids in a quoted jstack (#98, #99, #118, #119, #188). Wrapped in the gate's own documented exempt-begin/exempt-end markers for quoted source material - which is what that mechanism exists for, not a way around a red check. Gate now exits 0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
… breakage and are not Running the full local suite produced three failures, none of them code defects, and none of them documented anywhere. Each looks exactly like something is broken in the project. 1. JDK. Build Requirements said "JDK 17" and stopped there. On JDK 21 the build dies with "Unable to delombok: InvocationTargetException" in parallel-consumer-core - a module the change never touched - because lombok-maven-plugin 1.18.20.0 predates 21. Verified by re-running on 17: delombok completes and the reactor gets 0 ERROR lines through the same modules. Also records that the JDK should be set per command rather than by repointing a shared default, since every other session on the machine inherits that. 2. Go, in the conformance module. It reports `go.mod: unknown block type: tool` when `go env GOTOOLCHAIN` is `local` - a persisted Go setting, not an env var - because a `tool` block needs 1.24+ and Go's own default (`auto`) would have fetched it. Controlled: same command, same directory, only GOTOOLCHAIN differs -> exit 1 pinned, exit 0 with auto. 3. Ruby, same module: a missing bundler, because macOS's system 2.6 is on PATH and the floor is 3.2. The conformance module fails the build when a runner cannot be built, which is correct and worth keeping - "a language nobody could build is not a language that passed" is its own message, and the alternative is a silent skip. But an undocumented hard requirement that surfaces as a bundler error costs an hour before anyone suspects their PATH. Consequence worth stating plainly: the full suite CANNOT be green on a machine without Go 1.24+ and Ruby 3.2+, so "all green locally" is not available here and its absence is not evidence of a defect. CI is unaffected - Linux runners carry both. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
…t a list of versions Corrects the entry added an hour ago, which named "Go 1.24+ and Ruby 3.2+" as build requirements. That was misframed, and I reached it by diagnosing from symptoms without checking prior art first - the thing this repo's own investigation rules put before forming a hypothesis. The prior art exists and answers the question directly. docs/inflight/parked-containerised-toolchains-and-runtime.md records the decision: every toolchain the fan-out was missing is in mise's registry, mise is the mechanism, AGENTS MAY INSTALL IT AND RUN `mise use -g` THEMSELVES, and - the part that matters for anyone tempted to extend the container route - "do not build an image fleet for problems mise has already solved". bin/build-client.sh already implements exactly that split in its header: "CONTAINER route: swift and cpp. These are the only two toolchains mise cannot serve here", everything else native. Its container route never even runs a container - a multi-stage Dockerfile with `--output type=local` lifts the artefacts out. So the two failures I recorded were not missing project requirements. mise is simply not installed on this machine, so Maven fell through to whatever was on PATH: a Go pinned to GOTOOLCHAIN=local and macOS's system Ruby 2.6. Installed mise and confirmed the mechanism rather than assuming it - `mise use -g go@1.25` gives go1.25.14 and the Go client then builds exit 0 with no GOTOOLCHAIN override at all. Ruby 3.4 is installing (it compiles from source, which the parked note already warned is the slow one). Today's run is also the empirical case FOR that split, which is worth keeping next to it: the two containerised languages built fine on a machine with neither toolchain, while both native-route failures were host PATH problems. The criterion is sound; it just needs mise present to hold up. Nothing in .zshrc or any shared default was touched - the installer only printed its activation line as advice, and it stays that way. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
… and CI cannot see it The first full local run of the conformance suite - the first time every prerequisite was actually satisfied - got all ten runners built and then failed 20 of 20 scenarios for exactly the two container-route languages, cpp and swift, each with exit 126, "found but cannot execute". The extracted artifact is an `ELF 64-bit LSB pie executable, ARM aarch64` on a macOS arm64 host. The parked note already carried a caveat for this and it understates the problem. It frames the risk as linking discipline - static by default, so a glibc mismatch does not bite on a different host. That is right for Linux-to-Linux and irrelevant here: no amount of static linking makes a Linux ELF run on Darwin. Build-and-extract is sound only when the host OS matches the image OS. WHY THIS MATTERS BEYOND THE TWO LANGUAGES. It is the deciding constraint on the "fall back to containers for anyone without mise" idea. That fallback would be broken precisely where it is needed: macOS developers are the ones who cannot get Swift and C++ from mise, and they are the ones for whom an extracted Linux binary will not execute. A container BUILD fallback has to be a container RUN fallback as well, which is a much larger change than pointing an exec at Docker - and it is the parked note's section 2, still open, rather than its section 1, which is settled. CI IS UNAFFECTED AND STAYS GREEN THROUGH THIS, because its runners are Linux and the extracted binary is native there. That is the trap worth recording: the gap is invisible to every check that exists and appears only on a developer machine - the same shape as the toolchain drift in ci-toolchain-versions-declared-twice.md, added in this branch. None of this argues against containers for BUILDING. They built both languages cleanly on a host that has neither toolchain, which is exactly what they are for. It argues that "built" and "runnable here" are different claims, and only the first is currently true off Linux. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
…t <lang>-grpc The demo exists to show Parallel Consumer and the arm that RAN Parallel Consumer never said so. The table read `AK core (rdkafka)` against `ruby-grpc (this client)`, where the second name looks like a transport detail. It was also internally inconsistent: the Java seed already printed `pc-core (ParallelEoSStreamProcessor)`, so `pc-` already meant Parallel Consumer in this very output. Every arm that runs Parallel Consumer is now `pc-` prefixed. `AK core` is the only arm that is not Parallel Consumer, so it keeps its name. KOTLIN WAS ALREADY OFF-CONVENTION and nobody had noticed: it called its arm `kotlin-sidecar` while the other ten said `<lang>-grpc`. Nothing caught it because the conformance harness's normaliser accepts `(grpc|sidecar)` either way, and its javadoc gave no reason for the divergence - so this overrides it to `pc-kotlin-grpc` rather than preserving an unexplained difference. THE HARNESS HAD TO MOVE WITH IT, or the rename would have been silently self-defeating: the arm normaliser matched `[a-z0-9]+-(grpc|sidecar)`, with no hyphen in the class, so every `pc-` row would have stopped matching and the drift check would have gone back to comparing nothing - the exact failure fixed earlier on this branch. Now `(pc-)?[a-z0-9-]+-(grpc|sidecar)`, verified by running the real `skeleton()` and `normalise_arms()` over output in the new shape. SEVEN ARM COLUMNS WERE NOW TOO NARROW, because the prefix adds three characters to every label. Only Go asserts its own column width, so only Go went red - the other six would have printed a ragged table with nothing to say so, since column width is deliberately not part of the cross-language contract and no shared check covers it. Widened to longest-label-plus-two in go, rust, ruby, cpp, dotnet, typescript, scala and python; kotlin, swift and java already had room (java sizes its column from the widest label at runtime). TWO SELF-INFLICTED BUGS, both from the same wrong assumption, both caught by building rather than by reading: `\b` does not protect a hyphenated compound. In `protoc-gen-go-grpc` the character before `go` is a hyphen, so `\bgo-grpc\b` matched inside Google's plugin name and rewrote it in `go.mod`, `scripts/generate-proto.sh` and generated code. The same class renamed five Maven MODULE names (`<module>...-client-java-grpc</module>`), which broke the Java build outright. Both reverted, then every introduced token was grouped and audited rather than spot-checked. Also adds a `.gitignore` for the Swift module. SwiftPM's `.build/` holds a full git checkout of every dependency, and nothing ignored it - a native `swift build` followed by `git add -A` stages 1,100+ vendored files, which is how this was found. Checked every other language for the same hole: none. Verified: go build + go test, cargo check, dotnet build, ruby -c, python compile, clang++ syntax, and the conformance skeleton over the renamed output. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
…ery version once Two guards for the same class of failure: a toolchain that is present but wrong, which never says so and instead surfaces as an error inside a module the change never touched. 1. BUILD-TIME ASSERTION, in bin/build-client.sh before it hands over to Maven. Maven shells out to whatever go, ruby, node, python3, rustc or dotnet is on PATH, and the two real cases each cost an hour before anyone suspected their PATH: Go 1.23 against a 1.25 module reports "go.mod: unknown block type: tool", and macOS's system Ruby 2.6 reports a missing bundler. Neither says "your toolchain is the wrong version", which is the only useful sentence. It compares MAJOR.MINOR, not the exact patch. The declarations are pinned exactly and the gate below holds them identical, but asserting the patch here would fail every developer whose mise resolved one release along - this was written on a box with go 1.25.14 against a pinned 1.25.13 - while catching nothing, because every failure it exists for was a major or minor gap. A patch difference prints a line and continues. A version it CANNOT read is a failure, not a pass. Proven able to fire rather than assumed: rust on 1.94.1 against the pinned 1.88.0 dies with both numbers and the mise command that fixes it. 2. ONE DECLARATION, mise.toml, plus bin/check-toolchain-versions.sh asserting the clients.yml matrix mirrors it. The two INSTALL differently on purpose - mise locally, setup-* actions on the runner, because those carry the client matrix's caching and ruby/setup-ruby is pinned by commit SHA, a supply-chain posture worth keeping. What was not on purpose is that they were two sources of truth for the same numbers with nothing comparing them, and they had already drifted by WHOLE MAJOR VERSIONS: dotnet 8.0.404 against 9.0.101, node 22.17.0 against 25.9.0. The gate also asserts the four languages with NO host toolchain are absent from both files - swift and cpp build in containers, kotlin and scala on the Maven reactor - because "nobody declared it" and "it deliberately has none" are otherwise indistinguishable. Its self-test has twelve cases and every red one is a distinct way the two can disagree, including the one that matters most: a workflow whose format changed parses zero languages, so every comparison would trivially succeed. That case asserts the gate FAILS rather than passing vacuously - a green run that checked nothing is worse than no gate. Wired into Repo Hygiene as its own job, self-test first, matching the pattern the other gates use. An immediate payoff worth recording: with mise.toml in the repo, mise's shims now resolve the PINNED go and ruby here rather than the machine-global ones - the local build and CI moved into agreement the moment the file landed, without anyone running anything. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
…wn image, so macOS can drive them Before this, C++ and Swift could not be conformance-tested on a macOS machine at all. Their runners are built in a container and extracted, which is sound on Linux and impossible anywhere else: the artifact is a Linux ELF, so every scenario failed with exit 126 - "found, but cannot execute" - 20 times, with empty stdout and nothing anywhere saying "wrong operating system". Now: cpp 5 scenarios driven 10 tests 0 failures BUILD SUCCESS swift 5 scenarios driven 10 tests 0 failures BUILD SUCCESS NO PRODUCT CHANGE WAS NEEDED - no client API, no protocol change, nothing touching R29's authority allowlist. That is the whole point, and it came from the owner's correction. I had been designing a containerised RUNNER, which is what created every "boundary problem" I then had to solve: the shim file to mount, the loopback port to reach. The client spawns its sidecar as its own child and dials 127.0.0.1 (KTD41), and SidecarShim announces a bare port with NO host deliberately, so the client exercises its real spawn-and-reap path rather than gaining a connect-to-a-port option that would bind ten languages. Client, shim and engine therefore have to share one loopback. Move the whole run inside - suite, in-JVM engine, shim, runner - and that loopback is just the container's own. Measured first, since the alternative sounds plausible: from a container on Docker Desktop 27.4.0, `--network host` cannot reach the host's 127.0.0.1 at all, while host.docker.internal can - an address no client can be told to use without changing the protocol. HOW IT IS PUT TOGETHER. A `conformance` stage in each module's existing Dockerfile: FROM toolchain, plus a JDK, plus the runner COPIED FROM THE SAME `build` STAGE CI EXTRACTS FROM - so this path and CI drive one binary rather than two builds that could drift. LanguageRunners picks it up via PC_CONFORMANCE_PREBUILT_RUNNERS and carries an empty build command, the shape Kotlin and Scala already use for "the reactor built it". That also avoids a trap: the ordinary build command is bin/build-client.sh, which needs Docker, and there is none inside a container. FOUR FAILURES ON THE WAY, each of which moved it a stage further and two of which were mine: - `-am` ran parallel-consumer-core's whole 363-test suite inside the container and hit three contention flakes, one already in the ledger. Now scoped to ConformanceSuiteTest; -am stays, because the suite genuinely needs those jars and test-jars built. - `--network none` looked like a tidy proof that nothing crosses the boundary. It made the tool unusable from a cold cache: a macOS ~/.m2 holds no protoc:exe:linux-aarch_64, so Maven died in the protocol module before the suite ran. Now PC_CONFORMANCE_OFFLINE=1 restores the proof once warm. - My own prebuilt-runner guard threw for SWIFT inside the CPP image, because LanguageRunners.all() constructs every language's runner to build the selectable list, long before selection filters it. Relaxed to fall through, which keeps the loudness the throw was protecting: build-client.sh refuses by name inside a container, "docker is not installed ... this is NOT a pass". STILL OPEN, and recorded rather than glossed: CI will never exercise this path, because its runners are Linux where the extracted artifact is native. It is a macOS-developer path only, so it rots silently - the risk the owner named before it was built. Nothing schedules a run yet. The demo runtime needed nothing: each demo already runs in its own container, where binary and OS match by construction. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
Dependency ReviewThe following issues were found:
|
❌ Duplicate Code ReportTwo engines run in parallel for cross-validation. Each has its own thresholds tuned to its baseline - the real safety net is the per-engine "max increase vs base" check. ✅ PMD CPD
No new clones introduced by this PR. ❌ jscpd (language-agnostic)
|
🧪🔒 Quarantine Lane Report
🔴 expected while the owner PR is open · 🟡🎲 flapper, pass proves nothing · 🚨 a deterministic quarantined test passing means its fix landed: delete its |
|
… is not the UI Design for the eleven-language grid, on a branch stacked on the per-language demos. No implementation. THE PRIOR ART CHANGED THE DESIGN, and I had started without reading it. next-polyglot-demo-app.md points at ideas 21-27 in the language-proxy interaction-model ideation - seven ranked directions for "a demo that runs all eleven language bindings at once and proves in one view that they read the same records" - and says explicitly that the grid belongs to whichever idea wins, while the per-language demos already on #331 are a deliberate narrow cut of a different idea. Idea 21, ranked first at 85% confidence, is what the owner's framing describes: run the consumption work and report one cohesive table, not eleven existing demos in sequence. Its key move is not the table - it is that each worker's Success report carries a receipt (record_id, language, offset, attempt, payload_checksum) onto a demo-observations topic, so KAFKA IS THE AGGREGATION BUS and sameness becomes computed rather than eyeballed. One TUI consumes that topic and can end with "11/11 languages converged at offset N". MOST OF ITS PILLARS NOW EXIST. The runners are there for all eleven and already speak production gRPC; cpp and swift became runnable off-Linux on the stacked branch; there is a shared broker; the table shape is fixed and its records/keys columns are deterministic across languages, which is what makes any cross-language claim checkable at all; and arms now name the product, so a row identifies itself. THE REAL GAP IS THE SIDECAR, NOT THE UI. Conformance runners are driven with a SHIM: it announces a bare port and holds stdin, because the engine lives in the suite's JVM where dispatch can be observed. The demo needs each runner to spawn a REAL sidecar against a REAL broker - idea 21's "real-sidecar flag" - which is a change to eleven runners. Anyone starting with the TUI will produce a beautiful view of one language. AND A MEASUREMENT TRAP WORTH SETTLING BEFORE ANY NUMBER IS RENDERED: eleven runtimes on one host at once cannot produce quotable throughput. This project already discarded a whole fan-out's figures for that reason. So either the grid runs in parallel and shows only the deterministic columns - records, keys, offsets, checksum agreement, convergence, which is idea 21's actual claim and is immune to contention - or it runs sequentially on a quiet host and stops being one live view. The perf track owns measurement semantics by prior agreement, so the speed table is its artifact and the grid should show sameness. The note records the smallest first step (two languages, one native and one container-only, receipts converging before any UI exists) and three things not to do, including the audience-as-workload keynote mode that was already rejected for violating the no-visitor-input posture. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
…e sidecar, and 21 is not 22 Rewrites the design with two corrections found by reflecting it back before building anything, and adds the three decisions that have to land first. Each one changes what gets built, so the note says not to start until they are settled. CORRECTION 1 - I NAMED THE WRONG GAP. The first version said the expensive part was giving the eleven conformance runners a "real-sidecar flag". That is stale: idea 21 was written before #331, and every language now HAS a demo binary that already spawns a real sidecar against a real broker. The costume exists; it is just called a demo. So the base-binary choice is a real fork - demos, where the expensive per-language half is already paid for but which carry seeding and their own tables as baggage; or conformance runners, which are frozen and machine-checked but would need both a real sidecar and receipts in eleven languages. CORRECTION 2 - I CONFLATED TWO DIFFERENT DEMOS. The convergence claim depends entirely on consumer-group topology, and the first version asserted one without saying so. One topic with ELEVEN SEPARATE GROUPS means every language reads every record, per-record checksums are comparable, and "11/11 converged at offset N" is a real assertion - idea 21, claiming they read the same records identically. One topic with ONE SHARED GROUP means the eleven split the records and rebalance across languages - idea 22, claiming eleven runtimes cooperate in one group. Both are good; they are not variants of one demo, and only the first supports the convergence line. Either way something must own seeding, because eleven demos against one topic currently means eleven backlogs. VERIFIED RATHER THAN ASSUMED: the receipts route is real. proxy.proto carries `repeated ProduceRecord produce` in the worker's report, commented as the sanctioned route for worker output, at-least-once. So receipts ride the protocol's own output path, Kafka is the aggregation bus, and nothing parses eleven demos' stdout - the orchestration shape the owner ruled out. The third decision is whether the grid shows a rate at all. Eleven runtimes on one host cannot produce quotable throughput, and this project has already binned a whole fan-out's figures for that. The perf track owns measurement semantics by prior agreement, so the speed table is its artifact. Also records the smallest first step - two languages, one native and one container-only, receipts converging before any UI exists - and cross-references every adjacent note, each checked to exist. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01WcdtMdmdhyYug3krZRWkZT
Receives master through the parent chain, so this branch's PR diff shows only its own work. Three conflicts, all in the agent-hook wiring, and all of the same kind: this branch adds the low-disk hook while master added three of its own, and each side edited the same list. Both sides kept in every case - the merged settings.json declares eight hooks and still parses, and the self-test keeps both sections with this branch's first so the parent's header stays immediately above the tests it introduces. docs/agent-harness.md matched neither side: one said four hooks and the other five, and the merged file has eight. The sibling sentence had already dropped its count on master, which is the more durable phrasing given it has now gone stale twice. Verified: full reactor builds; the hook self-test reports 156 passing and 14 failing, and those 14 are the pre-existing set a pristine master worktree also reports, untouched here; copyright and docs-data gates green.
…sing a rung Absorbs the uber-demo design note and retires the branch that carried it. That branch held one file of 118 lines - a design for a demo that runs all eleven language bindings against one workload and asserts in a single view that they read the same records. It had its own branch and its own draft PR (#332), and that was over-engineering. A design nobody needs to review separately from the work it follows does not need a rung, and it belongs with the demos PR it designs the successor to. The cost of the rung was not cosmetic. Two PRs had stacked above it - the FFI work and the Kafka Streams proof of concept - so shipped, tested code was gated behind a draft design note whose own first paragraph says three decisions are open and not to start. Neither of those branches references the uber demo, the eleven-lane grid or its observations topic; the stacking was chronology, not dependency. The note's header is repointed in the same commit. It named its own branch as its home and cited the now-closed PR, which a rename or deletion would otherwise have left dangling - the check the merge checklist asks for. It now names the branch that actually carries it, and records why the separate PR was a mistake, so the same instinct produces a different result next time. Verified: copyright, docs-data and issue-reference gates green; no reference to the retired branch remains anywhere in the tree.
The note filed one commit ago pinned the collision on #340. That is the branch stacked ON TOP of it, and the mistake matters more than a wrong number would: `gh pr diff #340` shows a ONE-LINE change to bin/check-copyright-headers.sh, because #340 is based on feats/polyglot-demos rather than master and inherits the rewrite from it. A reader following the note to the PR it named would have seen a trivial diff and concluded the note was overblown - the same false negative this class of finding is about, reproduced inside the note warning about it. The review pass hit exactly that and had to go find the blob by hand. Corrected, measured against the right branch: the rewrite is #331 (feats/polyglot-demos), master to it is +476/-40, it to #340 is +1/-0 (the `*.options|hash` row), and its copy differs from this branch's by +203/-89. Verified #331's copy is missing the same five fixes, so the asymmetry argument is unchanged - only its subject is. Two other corrections from the same pass: - Dropped "then re-apply only genuinely new entries". There are none: `*.options|hash` is the only row the stacked branch adds and this branch's copy already carries it, so the take-whole is complete rather than a first step. - Named the two fixes the bullet list had left to the summary sentence - grandfathered() failing open on an unreadable blob, and plain `git ls-files` instead of `-c core.quotePath=false` - so all five are checkable from the note rather than four plus a claim. Also recorded that bin/test-check-copyright-headers.sh is untouched by the other branches, so the one script is the whole of the conflict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…sure the same thing twice Owner's call: the benchmarking work (#362) and the demonstration work in the language-proxy descendants (#328, #331, #332) independently built workload generators, engine arms and result rendering. Two work models will disagree, and the bench's took four defect-fix rounds to become honest - a second implementation will rediscover those silently, in front of an audience. Records why it matters and the candidate directions, without picking one: post-merge reconciliation, not a blocker on either side, but due before the demos are used for anything public. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01MedgsxqrM8vjSt5ncuAo8g
…culum, one demo to converge on Eleventh bit of the follow-up Codex strategy conversation (2026-08-29). New note docs-executable-progression.md: structure the flagship example as numbered runnable stages and the learning path stops being authored - it is encoded in the transitions, so an agent diffing stage 05 -> 06 can generate the quickstart, tutorial, manual chapter, workshop, talk and the "coming from KafkaConsumer / Streams / Python" entries from one running application that cannot drift from compiling code. Generalises the trust the repo already places in generated docs (the tagged-source README) and builds on #266's parcel example and the #208 docs site. Per the owner: the note reconciles rather than adds a fourth demo artefact - the uber demo (#332, collapsed into #331), the three-reveal demo, and the realistic-domain Streams benchmark branch are cross-linked both directions, with the frame that the progression's domain IS the realistic domain, its later stages ARE the reveals, and any stage can feed the uber demo's language matrix. The owner also ruled this convergence strategy-level, so the engine thesis's briefing list gains it with its merge-gate-shaped consequence: no new standalone demo codebase when a stage of the progression could carry it. The engine-thesis note's "keep your code; replace the runtime underneath it" also gains the follow-up conversation's own qualification: a design pressure, not an absolute promise. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016PFA3CXkGc2m32cBJMAWTK
…sorb the toolchain gate Forward-merges `feats/java-sidecar-demo` (#328) into this branch, the second link of the descendant-chain half of the plan's retarget move - `docs/plans/2026-08-31-001-process-god-branch-decomposition-plan.md`. Ordinary merge: no history rewritten, nothing force-pushed, this PR's base unchanged. The pre-merge tip is preserved as `origin/backup/pre-forward-merge-331`. The ten per-language demos, the container conformance path, `mise.toml` and the toolchain gate are untouched - they are this rung's reason to exist. What conflicted is the agent harness, which this branch built a first version of and `master` then rebuilt. ## Nine conflicts, and the pattern behind seven of them **This branch wrote the low-disk hook; `master` shipped it better.** That is the shape of most of what follows, and in every case the incoming text is not merely newer - it says, at the site, which claim of the earlier version turned out to be false. - **`.claude/hooks/warn-low-disk.sh`** - `master`'s. Established rather than assumed: an add/add with no merge base, so both sides were compared function by function. `fs_free_gib` and `fs_id` are subsumed by `fs_probe`, `file_blocks`/`file_mtime` by `stat_field`, and every capability this branch's version had - the `Docker.raw` sparse-file branch, the settings-store lookup, the `docker system df` TTL cache, the re-warn throttle - is present in the incoming one. Nothing in ours was content theirs lacked. - **`docs/agent-harness.md`** - three hunks, all `master`'s, and each one is a correction this branch's text invited: a hook count that had already drifted, a "matches `gh` in command position" claim the incoming text records as disproved by running it, and the "two stat-class syscalls, measured at 7ms" figure that was neither. Audited for loss afterwards: every heading and every distinctive line of this branch's version survives, several with wider coverage. - **`bin/test-check-agent-hooks.sh`** - the union merged cleanly and then **duplicated the entire `warn-low-disk.sh` self-test section**, once per side. The incoming copy is a superset - it has the realistic multi-key ceiling parse and the critical-to-warn downgrade case the earlier one lacks, and it already contains every cross-platform assertion the earlier one had - so the second copy is deleted and its section divider restored. A union merge that produces two of something is the failure mode to look for whenever both sides added the same section. - **`.claude/settings.json`** - the same duplication one level up: the union left `warn-low-disk.sh` registered **twice**, so it would have run twice on every Bash call. The conflicted registration becomes `check-upstream-map-merged.sh`, which is what the incoming side put there; the disk hook keeps the single registration further down. - **`.github/workflows/repo-hygiene.yml`** - `master`'s single `hygiene` job, and **this branch loses nothing by it**, which is checked rather than hoped: that job runs `bin/check-all.sh --with-tests`, and every script this branch ran as its own job - including `check-toolchain-versions.sh` and its self-test, which are this rung's - matches `bin/check-*.sh` or `bin/test-*.sh` and is therefore globbed. The glob absorbing a new gate with no edit anywhere is exactly the property `bin/AGENTS.md` says it was designed for; this is the first time it has had to. - **`AGENTS.md`** - the routing table is a union of all four rows. This branch's `bin/AGENTS.md` row wins because its `bin/AGENTS.md` really did grow to cover `.githooks/` and `.claude/hooks/`; the incoming `docs/inflight/AGENTS.md` wording wins because it is the later one; and both the clients row and the core-engine row are kept. ## The two that are this branch's, adapted rather than taken - **`parallel-consumer-proxy-client-rust/build.rs`** - this branch's `is_executable`, which carries the finding behind it: Maven leaves the `.exe` classifier artefact at `0644`, so selecting on `is_file()` alone picks a `protoc` that fails `Permission denied` in preference to a working one on `PATH`. The incoming one-liner says none of that. Only the parameter type is taken from the incoming side, to match the `Path` it imports. - **`parallel-consumer-proxy-client-python/README.md`** and the **dotnet `ConformanceHarness.cs`** - both are additive collisions. The README keeps this branch's `make demo-build` line and the incoming `make test` wording, which names the handshake the rung below added; the harness keeps this branch's two `using`s, which its shared `JvmToolchain` needs. ## A flake sighting, recorded rather than swallowed `ParallelEoSStreamProcessorTest.processInKeyOrder[3]` failed its own sanity check once here, 1 in 543. It is the known `master`-state flake with an existing ledger, and this sighting is appended to it rather than written up again - `docs/inflight/test-processinkeyorder-sanity-check-races-the-first-poll.md`. **The control arm is the strongest that ledger has.** `parallel-consumer-core` is byte-identical between this merge's result and the rung below it (`git diff --cached feats/java-sidecar-demo -- parallel-consumer-core` is empty), and the rung below ran the identical full-reactor command five times on the same box within the same few hours, 543 tests and 0 failures every time. One tree red once, a byte-identical tree green five times, and no code difference to appeal to. ## How it was verified JDK 17, macOS/arm64, toolchains through `mise`. - Full default reactor `./mvnw --fail-at-end test`. - `bin/check-all.sh` on its real exit code: 1, three gates failing and all three inherited. Every finding not already present one rung down names a `docs/inflight/` note belonging to this branch - the gates arrived with `master` and are seeing those notes for the first time - and none names a file this merge's resolutions touched. The two proto gates, which report CANNOT RUN without `buf`, both pass under `mise`. **Not run, and named rather than skipped quietly:** the `cpp` and `swift` conformance cells. This rung is precisely where `bin/run-conformance-in-container.sh` fixes them - on macOS their extracted runners are Linux ELF binaries and every scenario fails exit 126 - but exercising that path means building both containers, and this box is under a low-disk warning that names a container build as what would tip it over. They are deselected with `-Dpc.conformance.language` and left to CI's per-language rows, which is the one place that lane runs on Linux anyway.
… engine back Forward-merges `feats/proxy-requirements` (#293) into this branch after that branch took the Wagon A stack (merge 8a7dc3e) and was retargeted onto `feats/proxy-dispatch-clients`. This is the descendant-chain half of the plan's retarget move - `docs/plans/2026-08-31-001-process-god-branch-decomposition-plan.md` - and it is an ordinary merge: no history was rewritten, nothing was force-pushed, and this PR's base is unchanged. The pre-merge tip is preserved as `origin/backup/pre-forward-merge-328`. Most of the volume is content this branch was simply behind on: the ten foreign clients, the conformance framework, the protocol module's documents, current `master`, and the five records that became plain final classes because Error Prone crashes on a Jabel-desugared record. Only four files conflicted. ## The one that mattered: two `Main`s, and only one of them can be the sidecar `Main` was an add/add conflict, and the two sides are the same class written for opposite worlds. - This branch's is **unit U10**: it hosts `ConfigureHandler`, drains records held in a foreign process on shutdown, accepts `--socket` to listen on a Unix domain socket, and reserves exit 3 for a drain that timed out. `ReferenceDemo` - the whole point of this rung, and the 27x headline - spawns it through `SidecarProcess`, so a `Main` without an engine is a demo that cannot run. - The stack's is #384's engine-less shell: `NoEngineSessionService` behind a named `sessionServiceFactory()` seam, `EXIT_BIND_FAILED`, and a package-private `run` overload that lets a test name a port so the bind failure is reachable. The result is this branch's semantics on the stack's structure. `Main` keeps the seam, the bind exit code and the testable port; `sessionServiceFactory()` now returns the engine; the drain runs against whatever the session actually reached. Both usage texts, both sets of exit codes and both transports survive. ## Landing the engine there was not free, and the rung below said exactly what it would cost `docs/inflight/proxy-the-production-entry-point-still-hosts-no-engine.md` was written for this moment. It named the substitution as one call site, and then named the three things that had to happen with it - all three are done here, and the note is deleted rather than left describing a world that no longer exists: - **`NoEngineSessionService` is demoted to a test fixture, not deleted**, which is the outcome the note predicted. Eight cross-language `SidecarHandshakeTest`s (Kotlin, Scala, Go, Python, TypeScript, Rust, Ruby, C#) spawn a sidecar and assert the refusal arrives as `UNIMPLEMENTED` *specifically* - `PERMISSION_DENIED` and `RESOURCE_EXHAUSTED` are what the two interceptors raise before the service method runs, so the status code is the assertion rather than "it failed". An engine behind `Main` makes their subject vanish. - **`NoEngineMain` is the entry point they now spawn.** It is in the proxy module's test tree, so it ships in the test jar every client's sidecar classpath already includes for `TestModeMain` - nothing about how any of them spawns changed except the class name. It reimplements no lifecycle: it calls `Main`'s own seam with the no-engine supplier, so the binary those tests drive is the production lifecycle with one substitution, not a second copy that can drift. - **Nine call sites were re-pointed**: the eight foreign constants plus the Java client's own handshake test, which runs in-JVM and so calls `NoEngineMain.run` rather than spawning. The prose around each was corrected too - six of them described the thing they spawn as "the production `Main`", which is now the wrong binary rather than a wrong adjective. **`MainTest` is one class again.** Both sides grew one for the same class, in different packages, overlapping on arguments-refused, port-announced and two-sidecars. They are merged in `Main`'s own package keeping every assertion either side had, with the `UNIMPLEMENTED` test re-pointed at the no-engine service through the seam. `theProductionEntryPointHostsTheEngine` is new and is the substitution made assertable: without it the two suppliers could quietly swap back and every other test here would still pass, because the lifecycle is identical either way. ## The other three conflicts - **`.github/workflows/maven.yml`** - both sides added a matrix entry and neither displaced the other: this branch's `Demo (both entry points)` lane and master's `Chaos Pain Suite`. Union. - **`bin/check-copyright-headers.sh`** - master's refactor wins for the code (`check_derived`, `check_always_derived`, `fork_point_path`; this branch's copy still called `fp_path_of`, which no longer exists). The `Demo.java` recovery comment is this branch's, because master's version says the entry is ahead of its file and cites no ledger path *because none resolved on master* - both the file and `docs/inflight/branch-classic-comparison-demo.md` are here, so the path is cited. - **`ParentDeathWatchdog`** - a javadoc line only. This branch's wording wins because it is the one that is now true: the response to death is `DrainCoordinator`'s, and master's "once there is an engine it is the shutdown drain's" describes a future that has arrived. ## What was corrected because this merge falsified it `AGENTS.md`'s module table said the proxy module "hosts no engine, so a session is answered UNIMPLEMENTED", and the Python client's README called the engine-less sidecar "the production `Main`". Both are now wrong by one commit, so both were changed rather than left to rot. ## What was NOT corrected, and is a finding rather than an omission `docs/data/testing-evidence.d/parallel-consumer-proxy.yaml` says the module has no `src/test-integration` at all and that "there is no production entry point yet". Both were already false on this branch's own pre-merge tip - U10 landed `SidecarLifecycleIT` and `Main` without updating the record - so this is inherited rather than introduced, and rewriting an evidence record is the author's editorial call rather than a merge resolution. ## Two gates that the merge, not the code, turned red Both are 328's own files meeting analysers that arrived from `master` after this branch forked, so neither existed as a finding on either parent alone. - **SpotBugs now runs over this module** (`includeTests`, plus fb-contrib), and reported thirteen findings in the demo: three SLF4J format rules, class-envy and an over-concrete parameter on `ReferenceDemo`, a non-owned lock in its raw-gRPC arm, five exception-softening findings in `DemoBroker`, and two more in the test tree. **They arrive in two batches depending on whether `target/test-classes` exists when the gate runs** - the gate binds at `process-classes`, which is before `test-compile`, so a build that stops at `test-compile` sees the main-code findings only and a full `test` on a warm tree sees both. Do not read the first count as the whole set. Each was read at the site and dismissed in this module's own `spotbugs-exclude.xml`, class by class and pattern by pattern with the reason beside it - the shape the Java client aggregator's pom already prescribes and the api module already uses. None is a mute: a demo's output *is* pre-rendered text, a table renderer *does* read the rows it renders, and a `main()` whose only caller is a person reading a terminal is the one place `EXS_EXCEPTION_SOFTENING_NO_CONSTRAINTS` is the wrong rule. - **An unreadable exclude filter is a WARNING SpotBugs continues past.** The first version of that file quoted two command-line flags in an XML comment, and a double hyphen may not appear inside one - so the filter loaded as nothing, the gate kept failing with the same count, and the only evidence was one line reading `Unable to read filter` buried in the build log. Worth knowing before writing the next one. ## How it was verified JDK 17, macOS/arm64, toolchains through `mise` (`ruby 3.4.10`, `buf 1.72.0`). - **Full default reactor `./mvnw --fail-at-end test`.** Every module green, including the conformance suite over eleven of the thirteen bindings - so `java-grpc` runs against a live engine, and the Java, Kotlin and Scala handshake tests exercise the re-pointed no-engine entry point end to end. - **`bin/check-all.sh` on its real exit code: 1**, three gates failing and all three inherited. Established rather than assumed: the same gates on the incoming tip report the same `check-file-refs` and `check-branch-self-reference` findings, `check-inflight-tags` reports the same problems, and every *additional* finding names a file this merge did not touch - each verified with `git diff --quiet HEAD -- <file>`. `check-proto-lint` and `check-proto-breaking`, which report CANNOT RUN without `buf`, both pass under `mise`. - **`COPYRIGHT_CHECK_REQUIRE_FORK_POINT=1 bin/check-copyright-headers.sh`** - 0 violations against a real fork point. **Two things this box cannot answer, both established as inherited rather than assumed:** - **The `cpp` and `swift` conformance cells fail all ten of their scenarios with exit 126**, "found but cannot execute". Their runners are built in a container and extracted, so on macOS the artefact is a Linux binary - `file` reports `ELF 64-bit ... GNU/Linux` for both on a `Darwin arm64` host, which is the whole diagnosis and has nothing to do with anything in this merge. It is the condition #331 documents and fixes with `bin/run-conformance-in-container.sh`, one rung up. They are deselected here with `-Dpc.conformance.language`, named rather than silently skipped. - **Six of the eight foreign `SidecarHandshakeTest`s were not run** - Go, Python, TypeScript, Rust, Ruby and C# sit behind `-Dpc.foreignClients`, and C++ and Swift build inside containers on a box under a low-disk warning. Kotlin and Scala are JVM modules in the default reactor and did run, as did the Java client's in-JVM equivalent. What changed in all of them is one constant, and the process they now spawn is the same lifecycle with one supplier swapped.
… its exemption Forward-merges `feats/polyglot-demos` (#331) into this branch, the third link of the descendant-chain half of the plan's retarget move - `docs/plans/2026-08-31-001-process-god-branch-decomposition-plan.md`. Ordinary merge: no history rewritten, nothing force-pushed, this PR's base unchanged. The pre-merge tip is preserved as `origin/backup/pre-forward-merge-340`. **Zero textual conflicts** - the FFI work and the ten per-language demos touch almost disjoint files. What the merge did surface is a guard doing exactly its job. ## The registry guard found the C probe, which is the point of it `LanguageRunnerRegistryTest.everyClientModuleOnDiskHasARunnerRegisteredAndTheReverse` arrives with the rung below. It reads the client modules **off disk** rather than from a list, precisely so that "an eleventh language fails this test on the day its directory appears, rather than on the day somebody remembers". This branch adds a twelfth directory - `parallel-consumer-proxy-client-c` - and the guard went red naming it. That is the guard working, not a collision. `c` genuinely has no conformance runner, and cannot have one: its own README calls it *a reach probe, not a supported client* - it links the engine in through GraalVM's shared library and speaks no gRPC at all, so there is no sidecar for a runner to spawn and nothing for the shared scenarios to drive. It has no `pom.xml` either. **The exemption was re-keyed rather than doubled.** The test already had one, `NOT_A_SPAWNED_RUNNER = "java"`, a bare `String`. The second case has a completely different justification, so a second constant would have recorded *that* two things are exempt and nothing about *why* either is - which is how an exemption list rots into a mute. It is now a `Map<String, String>` of language to reason, with both reasons written out, and the javadoc says that being on it is a claim about the module rather than a way to quiet the test. ## The `--shared` build, which is this rung's whole proposition, still works after the merge #385 predicted this rung would inherit both of the obstacles that make a native image hard - the Kafka client and a logging backend - and warned that **neither fails at build time**. Run rather than assumed, on the merged tree: ``` ==> GraalVM: native-image 23 2024-09-17 (Oracle GraalVM 23+37.1) ==> native-image --shared (libpc) Finished generating 'libpc' in 1m 34s. 76,802,760 bytes, libpc.dylib ``` and then the Go smoke against that library: ``` ok isolate created ok pc_session_open -> handle 1 ok a malformed frame is rejected as ERR_BAD_FRAME ok pc_send(Configure) accepted, broker=localhost:19092 topic=pc-ffi-smoke ok Configured: max_concurrency=4 executor_count=4 ok session closed PARALLEL CONSUMER RAN INSIDE THIS GO PROCESS - no sidecar, no gRPC, no JVM ``` **The obstacles are real and were already paid for on this branch, and the merge did not disturb the payment.** `build-shared-library.sh` carries `--initialize-at-build-time=ch.qos.logback,org.slf4j,...`, and the proxy module ships `META-INF/native-image/reachability-metadata.json`, which `native-image` finds on the classpath by itself. Both survived the merge - so the prediction holds as a description of the difficulty and is answered here, rather than arriving as a surprise. ## A defect of the class #385 named, found by looking for it That PR recorded a silent trap: `dependency:build-classpath` takes `-DincludeScope`, **not** `-Dmdep.includeScope`, and the ignored spelling yields the TEST classpath. Swept for other instances across this branch: the two demo `run.sh` files already use the correct spelling; **`ffi/build-shared-library.sh` is the one remaining instance**, and it is live - `build/proxy-classpath.txt` holds **18 test-scope artefacts out of 86**, including JUnit 4 and 5, Mockito, Testcontainers, ArchUnit and Truth, plus `parallel-consumer-core`'s test-jar. **Not fixed here, and the reason is that it is not this merge's to fix**: it is a one-word change to a file this merge does not touch, and it changes the classpath every measurement on this branch was taken against. Its worst named consequence does not land, which is worth recording so nobody re-derives it - that test-jar carries no `logback.xml`, so the shipped library is not picking up this repository's test logging config. What it costs today is a needlessly large closed-world analysis surface. Left for #340's author, with the evidence above. ## A flake sighting, and the day's tally `ParallelEoSStreamProcessorTest` went red on the second run of the merged tree - 5 failures in 543, `processInKeyOrder[1]` together with `inFlightMessagesCommittedIfProcessedDuringShutdown` and `executorThreadsInterruptedOnShutdownTimeout`, all on the same `[sanity check input data]` assertion. Appended to the existing ledger rather than written up again: `docs/inflight/test-processinkeyorder-sanity-check-races-the-first-poll.md`. The sighting is worth more than usual because the whole chain ran today on one box over three trees whose `parallel-consumer-core` is **byte-identical** at every link: green 5 of 5 two rungs down, red once then green one rung down, green then red here. Same code, same box, same afternoon, outcome flipping run to run - so the remaining variable is the run, not the tree. ## How it was verified JDK 17, macOS/arm64, toolchains through `mise`; GraalVM 23 for the shared library only, reached by path so the project's JDK 17 stays `JAVA_HOME`. - Full default reactor `./mvnw --fail-at-end test`, with the conformance suite over eleven of the thirteen bindings. - The `--shared` build and the Go smoke above. - `bin/check-all.sh` on its real exit code. `check-issue-refs` is red on **the PR body**, not the tree - #340's description carries three bare `#NNN` references, which the gate reads and correctly refuses; that is a body fix, not a code one. **Not run, and named rather than skipped quietly:** the `cpp` and `swift` conformance cells, whose runners are Linux ELF binaries on this host and whose container path this box's disk headroom does not permit exercising. They are deselected with `-Dpc.conformance.language`.
…ce-row correction Cascade link 1 of 3 for the two owner-ordered fixes on this chain. Brings #328's `bafbb3c51` down: the proxy module's testing-evidence row no longer claims the module has no `src/test-integration` and no production entry point, both of which stopped being true before the forward merge and stayed written afterwards. Ordinary merge, first parent this branch's own tip `0d6541a2f`. Zero conflicts - the two links share history from today's forward merge and only one file differs.
…e-row correction Cascade link 2 of 3. Brings #328's evidence-row correction down through #331 onto this branch, which already carries the other owner-ordered fix of this pass - `69366a57f`, the FFI classpath resolved at runtime scope rather than silently at test scope. Ordinary merge, first parent this branch's own tip `69366a57f`. Zero conflicts, one file changed. The two fixes do not touch each other: one is a shell script under the Go client, the other a docs-data yaml.
Serves #242.
depends on #328
Description
The ten non-Java demos, built by agents working in parallel - one worktree and one branch each -
against the seed and contract that PR #328 establishes. Every language runs
the same two arms over the same records: its own Kafka client one record at a time, and itself as a
foreign client over the sidecar.
This PR is also the evidence for several of the conventions #328 states. That seed's contract now
says a blocking sleep is governed by a predicate - is the client thread-per-record? - rather than by
a list of languages. The list came first, and four languages on it turned out to be wrong for their
client. Rust ran the control and stated its prediction before running it: 10,341 msg/s through the
library's blocking adapter against 3,518 with a raw thread sleep. Ruby, asked to check rather than
assume, confirmed the rule genuinely is right for Ruby - the negative result that makes the
correction certain instead of plausible. Reading #328 alone shows the conclusions; this is where the
argument is.
The same is true of the fixed column order. One document produced three different orders across
eleven implementations, every one defensible, two with written reasoning. An unstated rule does not
produce consistency - it produces a vote that nothing goes red to report.
What else is here
implementations can be compared where elapsed time and throughput never could be. The deterministic
records/keyspair is what makes that possible.bin/run-conformance-in-container.sh. Theirrunners are built in a container and extracted, which works only on Linux - elsewhere the artifact
is a Linux binary and every scenario failed with exit 126, "found but cannot execute". Containerising
only the runner cannot fix it: the client spawns its sidecar as a child and dials loopback, and the
suite's shim announces a bare port with no host, deliberately, so the client exercises its real
spawn-and-reap path. Measured: from a container,
--network hostcannot reach the host's 127.0.0.1at all. So the whole run moves inside one container, where that loopback is the container's own -
with no client, protocol or API change.
mise.tomlis now the single place versions are stated, andbin/check-toolchain-versions.shfails if the CI matrix disagrees. They install differently onpurpose - mise locally,
setup-*actions on the runner, which carry the matrix's caching and aSHA-pinned
ruby/setup-ruby- but they had already drifted by whole major versions, with nothingcomparing them.
goorrubysays so instead offailing later inside a module the change never touched.
under 9 GiB free with nothing watching.
Known-open, recorded rather than glossed
close()can report "Close complete" and leave non-daemon threads running, hanging the Java demo onthe container path. Root-caused to
parallel-consumer-core, not the demo - a user's applicationwould hang the same way. See
docs/inflight/bug-engine-close-leaves-non-daemon-threads.md.silently. It needs a scheduled run.
machinery and refuse to report numbers, and every one complied; the box ran at load 20-113
throughout.
Checklist
what stays open
docs/features/- N/A - demos and testinfrastructure; no user-facing library feature changes here
fail, the cross-language drift check, and a self-test for the new toolchain gate