Repository navigation
cuda: BATCH_INVARIANT: occupancy-independent FA KV split, PTQ1_0 mat-vec up to 8 columns - #322
Open
professorpalmer wants to merge 1 commit into
Open
professorpalmer wants to merge 1 commit into
professorpalmer wants to merge 1 commit into
Conversation
…vec up to 8 columns Under GGML_CUDA_BATCH_INVARIANT: - the non-stream-k flash-attention KV split is sized from a fixed blocks-per-SM value instead of the occupancy of the template instance that runs; the 1-query and multi-query instances differ in registers and shared memory, so their splits (and combine order) differed. - PTQ1_0 batches up to MMVQ_MAX_BATCH_SIZE stay on the PT mat-vec, whose per-column arithmetic does not depend on the column count; a 5-column speculative verify on MMQ did not match the same tokens decoded alone. Measured cost: pp5 172 -> 162 tok/s, only on 5-8 column batches. GGML_CUDA_PTQ1_MMVQ_MAX overrides the mat-vec / MMQ crossover (default 4, confirmed: MMQ wins from 5). Not covered: the stream-k split of the MMA kernel follows the padded KV length (and the tile instance follows the query count), so attention in a verify batch is not bit-identical to single-token decode past ~32k, or at any depth on the MMA decode route. Measured at 4k/20k/40k: the weight path matches, a few continuations diverge at the rounding level. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit dc0cd6c)
Open
2 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
One of the five pieces of #285, split per bri-prism's request on #317. Single commit on current
prism(eaecb50c7), applies without conflicts. Only active underGGML_CUDA_BATCH_INVARIANT=1.What it does
GGML_CUDA_PTQ1_MMVQ_MAXoverrides the crossover; the default of 4 is confirmed by measurement (pp5: 172 tok/s on MMQ vs 162 on the mat-vec, so 4 is where MMQ starts to win and where the verify columns of a draft of 4 still match).Receipts
test-backend-ops -b CUDA0withGGML_CUDA_BATCH_INVARIANT=1:FLASH_ATTN_EXT2994/2994,MUL_MAT1516/1516.