Repository navigation
Run batch-1 Krea 2 forwards on the valid prompt rows (0.11.1) - #18
Merged
Merged
Conversation
Krea2Pipeline pads every prompt to 512 text tokens and masks the padding. The fused blocks skipped the padded rows with a gather and scatter in every block; a batch-1 forward now drops them once, before the text fusion, and runs without a mask. The trimmed embeddings and positions are reused for the same prompt, so the text-fusion cache stays warm across steps. The padded rows got no attention weight and never reached the output, so the image rows see the same keys. RTX 4060 Ti, 1248x832, 8 steps, through the pipeline with zero-copy offload: a denoising step takes 1.03 s, as with the OrbitQuant 0.10 runner, instead of about 1.09 s.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
orbitquant.fusedKrea 2: a batch-1 forward with a padded prompt drops the padded text rows once (embeddings and the matching RoPE positions) and runs without a mask, instead of gathering and scattering the valid rows in every block. The trimmed tensors of the last two prompts are reused, so the per-prompt text-fusion cache keeps hitting across steps (a CFG pipeline alternates two prompts). Batch > 1 and grad mode keep the per-block path.The padded rows were never attended to and were dropped at the output, so the image rows attend to the same keys; the result differs from 0.11.0 only by kernel rounding (the masked and unmasked attention take different kernels).
Measurements
RTX 4060 Ti 16 GB,
WaveCut/Krea-2-Turbo-OrbitQuant-W4A4fused revision, 1248×832, 8 steps, the draw worker's zero-copy offload:Krea2FastRunner(previous revision)Per-step CUDA kernel time of the 0.10 runner and the 0.11.1 pipeline is identical (1.027 / 1.025 s per profiled step).
Checks
ruff check, CPUpytest: pass.tests/test_fused_family_kernels.py tests/test_fused_dit_kernels.py tests/test_fused_layout.pypass, including the newtest_krea2_batch_one_forward_runs_on_the_valid_text_rows.