PackageImport.SubmitAsync's cleanup filter at application/src/Ashlar.CLI/Commands/PackageImport.cs:148 is too narrow: it only cleans up after the exception types it anticipates, so any other fault leaves a parked forge-queue entry behind.
Ordinarily that is one leaked entry per manual invocation. But pkg pull iterates a whole store, and the round-3 hardening (correctly) widened the per-row try so a failing row is refused and the pass continues — which means a systemic fault now leaks one entry per package file, per pass, forever, and multiplies the gate-lock deadline from once to 15s per row.
Two things worth separating:
- The leak itself is pre-existing and lives in the narrow cleanup filter, not in the per-row
catch. Continuing past a bad row is the right behaviour and should stay.
- The amplification is new, and is the reason this is worth fixing now rather than later.
There is a second-order reporting problem in the same area: because the per-row catch is deliberately broad (it has to be — a narrow IOException/UnauthorizedAccessException pair is exactly what let an OutOfMemoryException abort a whole pass), a genuine internal fault in the gate store now renders as an ordinary × REFUSED row. An operator sees a refusal, not a bug. Worth distinguishing "this package was refused" from "this node failed while judging it" in the row wording.
Fix: broaden the cleanup to run on any failure (a finally, or a filter that is not exception-type-specific), and consider a distinct row shape for internal faults.
Found by the round-3 adversarial review (availability lens); three independent skeptics confirmed the mechanism and agreed the underlying defect is the cleanup filter rather than the per-row catch.
PackageImport.SubmitAsync's cleanup filter atapplication/src/Ashlar.CLI/Commands/PackageImport.cs:148is too narrow: it only cleans up after the exception types it anticipates, so any other fault leaves a parked forge-queue entry behind.Ordinarily that is one leaked entry per manual invocation. But
pkg pulliterates a whole store, and the round-3 hardening (correctly) widened the per-rowtryso a failing row is refused and the pass continues — which means a systemic fault now leaks one entry per package file, per pass, forever, and multiplies the gate-lock deadline from once to 15s per row.Two things worth separating:
catch. Continuing past a bad row is the right behaviour and should stay.There is a second-order reporting problem in the same area: because the per-row catch is deliberately broad (it has to be — a narrow
IOException/UnauthorizedAccessExceptionpair is exactly what let anOutOfMemoryExceptionabort a whole pass), a genuine internal fault in the gate store now renders as an ordinary× REFUSEDrow. An operator sees a refusal, not a bug. Worth distinguishing "this package was refused" from "this node failed while judging it" in the row wording.Fix: broaden the cleanup to run on any failure (a
finally, or a filter that is not exception-type-specific), and consider a distinct row shape for internal faults.Found by the round-3 adversarial review (availability lens); three independent skeptics confirmed the mechanism and agreed the underlying defect is the cleanup filter rather than the per-row catch.