Repository navigation
The top-CPU drill-down resolves statement text for the five queries it prints, not every plan-cache row in the window (#3959) - #3979
Merged
Conversation
…t prints (#3959) The CPU_SQL_PERCENT / CPU_SPIKE drill-down projected query_text inside its window. On Darling, v_query_stats resolves that column from the fleet-wide query_text_dim for every row it returns, so the read resolved and sorted the whole window's text to print LEFT(MAX(query_text), 500) for five groups. It now ranks and cuts to five without the text, then reads the same MAX over the same rows for the five that print. Equality serves the normal case through the (server_id, query_hash, collection_time) index. GROUP BY puts NULL keys in one group that equality cannot match, so only a group with a NULL key reads NULL-safe. The SQL stays byte-identical to Lite's (the #3648 pin), and Lite runs the same text. The frozen Dashboard mirrors it: its MAX over every window row's compressed varbinary(max) becomes an OUTER APPLY per printed group. Output is identical: an old-vs-new oracle over seeded data in all three SKUs, with NULL-key groups at the top, several texts per group, inline legacy text, a digest with no dimension row, text past the 500 cut, and trap rows outside the read's own filter. Plus md5 on the production-scale rig and on DARLING01's real data. Measured (medians): - Darling rig (400K-text dimension, 72K-row window, C collation): 260.6 ms to 132.2 ms. - DuckDB (Lite-shaped, 72K rows of inline text): 156 ms to 149 ms. - Dashboard (SQL Server 2016, 72K compressed rows): 1,140-1,360 ms warm to 32-47 ms. The EXPLAIN join-probe counter the #3902 PSP test used is now shared (ExplainJoinProbe), so both drill-down text tests count the same way. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif
erikdarlingdata
enabled auto-merge (squash)
September 23, 2026 02:29
erikdarlingdata
added a commit
that referenced
this pull request
Sep 23, 2026
…3920, #3927, #3931, #3932, #3940, #3942, #3946, #3947, #3950, #3952, #3955, #3956, #3957, #3964, #3965, #3966, #3968, #3972, #3975, #3979, #3980, #3981, #3983, #3984, #3985) (#3989) The wave's fix PRs deliberately carried no CHANGELOG edits (parallel-agent hot-spot protocol); each agent reported its entry and this commit lands them together, byte-verified against origin/dev. Claude-Session: https://claude.ai/code/session_01GdmA4ND1wLSqA91ax1m4xv Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #3959.
Why
The top-CPU drill-down (behind CPU_SQL_PERCENT and CPU_SPIKE findings) projected
query_textinside its window. On Darling,v_query_statsresolves that column with a LEFT JOIN toquery_text_dim, the fleet-wide text dimension. The read therefore looked up text for every plan-cache row in the analysis window and carried it through theROW_NUMBERsort, all to printLEFT(MAX(query_text), 500)for five groups. This is the pattern #3956 removed from the parameter-sensitivity drill-down. I filed it separately there because this read is pinned byte-identical to Lite's (#3648), and itsMAX(query_text)output needs care to keep identical.The frozen Dashboard's twin did the same with
MAX()over every window row's compressedvarbinary(max)statement.What changes
query_text.top_queriesCTE ranks and cuts to five exactly as before.LEFT(MAX(query_text), 500)for each printed group, over the same rows the single pass aggregated: that group's window rows under the window's own filter (server, time bounds,delta_worker_time > 0).(server_id, query_hash, collection_time)index serves.CASEtherefore sends a group with a NULLdatabase_nameorquery_hashto anIS NOT DISTINCT FROMread, and only such a group; PostgreSQL never runs that branch otherwise.SELECT TOP 5 ... MAX(query_text)becomes:TOP (5)CTE without text;OUTER APPLYthat takes the sameMAX(qs.query_text)per printed group, thenDECOMPRESS;(qs.query_hash = t.query_hash OR (qs.query_hash IS NULL AND t.query_hash IS NULL)), becauseIS NOT DISTINCT FROMneeds SQL Server 2022.database_nameis NOT NULL there.total_cpu_us, which one makes the top five was arbitrary before and still is.ExplainJoinProbe, so both drill-down text tests count the same way.Test plan
Measured before and after, same data, interleaved:
query_text_dim, 72K-rowquery_statswindow, work_mem 31MB), median of 5query_stats, 72K rows of ~1.2 KB inline text), median of 3IX_query_stats_hash_lookup), warmAt DARLING01's size the extra subqueries cost about 1 ms. The win arrives with the production-sized window, where the old read resolved and sorted text for every row.
MAXbut lie outside the read's own filter: zero CPU, before the window, another server.TopCpuQueriesTextLiveTests(Darling, live),TopCpuQueriesTextTests(Lite, DuckDB). Both also pin each text case by value.TopCpuQueriesTextShapeTests(Darling), the shape fact inTopCpuQueriesTextTests(Lite, Lite's inline text), andTopCpuQueriesTextMirrorTests(Dashboard.Tests, source idiom). The Top Cpu Queries' Max Dop is a cross-plan cumulative maximum with no provenance: it said 16 on a MAXDOP-1 instance whose stored plan is serial, and drove a wrong tuning recommendation #3648 byte-identity pin (DrillDownDopProvenanceParityTests) passes unchanged."";Darling.TestswithDARLING_TEST_PG: 12,702 total, 0 failed, 24 skipped (the runtime-gated tests).Lite.Tests: 5,134 total, 0 failed.Dashboard.Tests: 817 total, 0 failed.🤖 Generated with Claude Code
https://claude.ai/code/session_018Z4T51J4E6L8fjriPHpsif