Skip to content

Trend tools return every raw point: get_file_io_trend sends 12,451 points / 1.4 MB for one server at defaults, more than an LLM client's context #3897

Description

@erikdarlingdata

Summary

The trend tools return every raw point in the window. At default arguments, get_file_io_trend sends 12,451 points and 1.4 MB of JSON for one server. That is roughly 350k tokens at ~4 characters per token: more than a 200k-context model can take in, for a single tool call.

The store answers in about 100 ms. The slowness users feel is the client: transferring, parsing and feeding that payload to the model, or truncating it and reasoning over a fragment.

The size grows with objects per server (database files, databases, counters), so it is invisible on our servers and severe on large ones. Measured at d21bbc9a against nightly 458 on DARLING01.

Measured (tool text payload at default arguments)

Tool Server Payload Points
get_file_io_trend SQL2022 (27 files) 1,379 KB 12,451
get_store_metrics (store) 1,246 KB —
get_pg_io_trend PG18 449 KB 1,429
get_lock_wait_trend SQL2022 323 KB 3,269
get_pg_database_trend PG18 294 KB 1,429
get_query_duration_trend SQL2022 162 KB 1,247
get_query_heatmap SQL2022 100 KB 344 cells

get_file_io_trend is files × one point per 1-minute collection. SQL2022 has 27 files. A 300-database server (about 600 files) would return about 22× this, roughly 30 MB for one call.

My own client spilled a single 87 KB get_collection_log result to disk rather than put it in context. Clients with smaller budgets fare worse.

Existing mitigations, and why they do not reach this

Neither changes the number of points, and the point count is what scales.

Fix shape

  • Downsample on the server to a bounded number of points per series by default: a time_bucket sized from the window so each series lands at or under about 200–300 points. Carry min/avg/max per bucket (or an LTTB-style selection) so a spike survives.
  • Bound the series count for per-object trends. For example, get_file_io_trend defaults to the top N files by the trended measure, with a remaining_series count and totals for the rest, so a 600-file server does not return 600 series.
  • get_store_metrics: 1.2 MB is mostly per-object history. Default to a summary plus the top movers, with the full series behind a parameter.
  • A census pin: every trend tool's default call stays under a byte budget on a fixture sized like a large server (hundreds of files and databases), so this cannot regrow.

Where the web viewer charts these readers through its /api/read/* mirror, server-side bucketing also cuts browser payload and render time on large servers.

Activity

  1. erikdarlingdata commented on Sep 22, 2026

    @erikdarlingdata
    OwnerAuthor

    Scope note: #3903 (production measurements, 2026-09-22) now owns get_store_metrics' payload, with a sharper proposal (summary-first default, object filter, bounded page). The get_store_metrics bullet here defers to it; this issue stays on the trend tools (get_file_io_trend, get_pg_io_trend, get_lock_wait_trend, get_pg_database_trend, get_query_duration_trend), whose size scales with objects per server × points.

  2. added
    in-progressActively being worked by a local session or its agents (PR open or in flight)
    on Sep 23, 2026
  3. erikdarlingdata commented on Sep 23, 2026

    @erikdarlingdata
    OwnerAuthor

    Closed by the watcher: delivered in PR #3968, merged to dev.

  4. removed
    in-progressActively being worked by a local session or its agents (PR open or in flight)
    on Sep 23, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions