Background & Problem
When benchmarking Agent Substrate under realistic agent workloads (running test suites, compilation, and git operations across concurrent actor sandboxes), diagnosing whether tail latency spikes and cold-start regressions stem from CPU throttling, memory thrashing, or disk I/O contention requires Linux kernel Pressure Stall Information (PSI).
Currently, cAdvisor in benchmarking/monitoring.yaml drops PSI metrics and leaves the root cgroup id="/" unlabeled, making it difficult to isolate host-level resource starvation from actor container pressure.
Proposed Changes
- Relabel the root cgroup
id="/" to container="node" in the cAdvisor scrape job in benchmarking/monitoring.yaml.
- Expand the metric allowlist to capture
container_pressure_{cpu,memory,io}_waiting_seconds_total, along with network throughput and filesystem usage.
References
cc Max Smythe (@maxsmythe) Haowei Cai (Roy) (@roycaihw) - happy to submit a PR for this!
Background & Problem
When benchmarking Agent Substrate under realistic agent workloads (running test suites, compilation, and git operations across concurrent actor sandboxes), diagnosing whether tail latency spikes and cold-start regressions stem from CPU throttling, memory thrashing, or disk I/O contention requires Linux kernel Pressure Stall Information (PSI).
Currently, cAdvisor in
benchmarking/monitoring.yamldrops PSI metrics and leaves the root cgroupid="/"unlabeled, making it difficult to isolate host-level resource starvation from actor container pressure.Proposed Changes
id="/"tocontainer="node"in the cAdvisor scrape job inbenchmarking/monitoring.yaml.container_pressure_{cpu,memory,io}_waiting_seconds_total, along with network throughput and filesystem usage.References
cc Max Smythe (@maxsmythe) Haowei Cai (Roy) (@roycaihw) - happy to submit a PR for this!