Skip to content

The top-CPU drill-down resolves statement text for the five queries it prints, not every plan-cache row in the window (#3959) - #3979

Merged
erikdarlingdata merged 1 commit into
devfrom
feature/3959-topcpu-drilldown-text
Sep 23, 2026
Merged

erikdarlingdata merged 1 commit into
devfrom
feature/3959-topcpu-drilldown-text

Conversation

@erikdarlingdata

Copy link
Copy Markdown
Owner

Closes #3959.

Why

The top-CPU drill-down (behind CPU_SQL_PERCENT and CPU_SPIKE findings) projected query_text inside its window. On Darling, v_query_stats resolves that column with a LEFT JOIN to query_text_dim, the fleet-wide text dimension. The read therefore looked up text for every plan-cache row in the analysis window and carried it through the ROW_NUMBER sort, all to print LEFT(MAX(query_text), 500) for five groups. This is the pattern #3956 removed from the parameter-sensitivity drill-down. I filed it separately there because this read is pinned byte-identical to Lite's (#3648), and its MAX(query_text) output needs care to keep identical.

The frozen Dashboard's twin did the same with MAX() over every window row's compressed varbinary(max) statement.

What changes

Test plan

Measured before and after, same data, interleaved:

read before after
Darling, rig (PG 18.4 / TSDB 2.28.1, C collation, 400K-text query_text_dim, 72K-row query_stats window, work_mem 31MB), median of 5 260.6 ms (259-270) 132.2 ms (118-137)
Lite, DuckDB 1.4.4 (Lite-shaped query_stats, 72K rows of ~1.2 KB inline text), median of 3 156 ms 149 ms
Dashboard, SQL Server 2016 scratch database (72K rows of ~1.2 KB compressed text, the install's IX_query_stats_hash_lookup), warm 1,140-1,360 ms 32-47 ms
Darling, DARLING01 real data (SQL2025, SQL2022: about 1K CPU rows per window) 7.9-8.8 / 6.1-6.6 ms 8.5-8.9 / 7.2-7.8 ms

At DARLING01's size the extra subqueries cost about 1 ms. The win arrives with the production-sized window, where the old read resolved and sorted text for every row.

  • Identity, oracle = the old SQL verbatim, every column of every row. The seed is built to break a restriction that is almost right:
    • three of the five printed groups have a NULL key (NULL database, NULL hash, both NULL);
    • a group's winning text is the inline legacy text beside a digest;
    • a digest has no dimension row;
    • a text runs past the 500 cut;
    • trap rows carry a text that would win MAX but lie outside the read's own filter: zero CPU, before the window, another server.
    • Tests: TopCpuQueriesTextLiveTests (Darling, live), TopCpuQueriesTextTests (Lite, DuckDB). Both also pin each text case by value.
  • Identity, md5 over the full ordered output: the rig (5 rows) and DARLING01's SQL2025 and SQL2022 (5 rows each, read-only) match.
  • Dashboard identity: old and new side by side on a scratch database on SQL Server 2016 (compatibility level 130; dropped afterward). Identical rows on the 72K-row set, and on a set with two NULL-hash groups (different databases) at the top of the ranking.
  • Read shape (Darling, live): the plan resolves text for 20 rows, the printed groups' rows. The old shape resolves 140, every window row, and serves as the positive control.
  • Shape pins, ungated: TopCpuQueriesTextShapeTests (Darling), the shape fact in TopCpuQueriesTextTests (Lite, Lite's inline text), and TopCpuQueriesTextMirrorTests (Dashboard.Tests, source idiom). The Top Cpu Queries' Max Dop is a cross-plan cumulative maximum with no provenance: it said 16 on a MAXDOP-1 instance whose stored plan is serial, and drove a wrong tuning recommendation #3648 byte-identity pin (DrillDownDopProvenanceParityTests) passes unchanged.
  • Checked red:
    • Each pin fails against the old sources.
    • The Darling live test fails on the old SQL: "resolved text for 140 row(s) to print 5 groups of 4".
    • Two mutations fail the identity assertions in Darling, and the first also in Lite:
      • with the NULL-safe branch disabled, a NULL-key group prints "";
      • with the text read's CPU filter dropped, the trap text wins.
  • Full Darling.Tests with DARLING_TEST_PG: 12,702 total, 0 failed, 24 skipped (the runtime-gated tests).
  • Full Lite.Tests: 5,134 total, 0 failed.
  • Full Dashboard.Tests: 817 total, 0 failed.

🤖 Generated with Claude Code

https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif

…t prints (#3959)

The CPU_SQL_PERCENT / CPU_SPIKE drill-down projected query_text inside its
window. On Darling, v_query_stats resolves that column from the fleet-wide
query_text_dim for every row it returns, so the read resolved and sorted
the whole window's text to print LEFT(MAX(query_text), 500) for five
groups. It now ranks and cuts to five without the text, then reads the
same MAX over the same rows for the five that print.

Equality serves the normal case through the (server_id, query_hash,
collection_time) index. GROUP BY puts NULL keys in one group that
equality cannot match, so only a group with a NULL key reads NULL-safe.
The SQL stays byte-identical to Lite's (the #3648 pin), and Lite runs the
same text. The frozen Dashboard mirrors it: its MAX over every window
row's compressed varbinary(max) becomes an OUTER APPLY per printed group.

Output is identical: an old-vs-new oracle over seeded data in all three
SKUs, with NULL-key groups at the top, several texts per group, inline
legacy text, a digest with no dimension row, text past the 500 cut, and
trap rows outside the read's own filter. Plus md5 on the production-scale
rig and on DARLING01's real data.

Measured (medians):
- Darling rig (400K-text dimension, 72K-row window, C collation): 260.6 ms to 132.2 ms.
- DuckDB (Lite-shaped, 72K rows of inline text): 156 ms to 149 ms.
- Dashboard (SQL Server 2016, 72K compressed rows): 1,140-1,360 ms warm to 32-47 ms.

The EXPLAIN join-probe counter the #3902 PSP test used is now shared
(ExplainJoinProbe), so both drill-down text tests count the same way.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
@erikdarlingdata
erikdarlingdata enabled auto-merge (squash) September 23, 2026 02:29
@erikdarlingdata
erikdarlingdata merged commit 4c0a5bc into dev Sep 23, 2026
15 of 16 checks passed
@erikdarlingdata
erikdarlingdata deleted the feature/3959-topcpu-drilldown-text branch September 23, 2026 02:36
erikdarlingdata added a commit that referenced this pull request Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989)

The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev.


Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant