stop dying when a container disappears mid-stats - #17
Merged
Merged
Conversation
This was referenced Aug 21, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes swarmpit/swarmpit#736.
Root cause
ContainerUsagedeclaresvar v *types.StatsJSONand then:The error branch is a no-op left over from Docker CLI's stats loop, which retries — here there is no loop, so
vstaysniland the very next line dereferences it.stats.go:157is exactly the frame in the reported trace.nullis the nastier case: it decodes successfully and leavesvnil, so an error check alone isn't enough.Verified against the pre-fix logic — all four bodies panic with the reported error:
This happens routinely: a container that exits while its stats are being read returns an empty or truncated body.
Why it takes the whole agent down
ContainerUsageruns in a goroutine (ContainersUsage, stats.go:122) and there is norecover()anywhere in the agent, so one dying container panics the process. That makes it a second, independent cause of "Statistics not ready" in the UI — the agent logsStats collector started, then goes silent.Changes
nilContainerUsagenow returns(status, ok)so failed reads aren't reported as zero-valued containersrecover()per collection goroutine, so a future panic in one container degrades that container instead of killing the agentioimportTests
New
swarmpit/task/stats_test.godrivesContainerUsage/ContainersUsageagainst anhttptestfake of the daemon stats endpoint: empty,null, truncated, non-object and HTML bodies all returnok=falseinstead of panicking; a well-formed body still parses; and a mixed run reports the healthy container while skipping the bad one.go build ./...,go vet ./...andgo test ./swarmpit/task/...pass ongolang:1.12(matching the Dockerfile). Onlystats.goand the new test are touched —go.modis unchanged, and I left the repo's pre-existing gofmt drift in other files alone.