Skip to content

sd: sync with master-956-1b0ba10 - #2524

Open
wbruna wants to merge 3 commits into
LostRuins:concedo_experimentalfrom
wbruna:kcpp_sd_update_202610_1
Open

wbruna wants to merge 3 commits into
LostRuins:concedo_experimentalfrom
wbruna:kcpp_sd_update_202610_1

Conversation

@wbruna

@wbruna wbruna commented Oct 10, 2026 •

Copy link
Copy Markdown

master-945-228c707:

  • perf: avoid redundant permute+cont in non-interleaved RoPE
  • fix: avoid f16 overflow in Z-Image quantized matmuls on CUDA
  • feat: add Z-Image L2P support
  • feat: support for merging LoRA weights on model conversion
  • fix: correct step count adjustment for custom sigmas list
  • fix: correct non-circular tile placement and blending
  • refactor: shorten log levels and move source locations to the end
  • refactor: unify circular and non-circular tiling
  • fix: keep MiniMax-H3 VAE weights resident across temporal chunks
  • refactor: escape newlines in prompt logs and clarify source separators
  • chore: resolve MSVC warnings in ggml extensions

master-951-f89d9b1:

  • fix: align SD3 encoder token chunks before conditioning
  • perf: use ggml rope apply op on supported backends
  • fix: preserve singleton output dimension when merging LoRA
  • feat: add Iris-3B text-to-image support

On top of #2518 . edit: it's independent now

@wbruna
wbruna force-pushed the kcpp_sd_update_202610_1 branch 3 times, most recently from 7eb4c8d to 821fdd2 Compare October 10, 2026 17:56
* perf: avoid redundant permute+cont in non-interleaved RoPE
* fix: avoid f16 overflow in Z-Image quantized matmuls on CUDA
* feat: add Z-Image L2P support
* feat: support for merging LoRA weights on model conversion
* fix: correct step count adjustment for custom sigmas list
* fix: correct non-circular tile placement and blending
* refactor: shorten log levels and move source locations to the end
* refactor: unify circular and non-circular tiling
* fix: keep MiniMax-H3 VAE weights resident across temporal chunks
* refactor: escape newlines in prompt logs and clarify source separators
* chore: resolve MSVC warnings in ggml extensions
* fix: align SD3 encoder token chunks before conditioning
* perf: use ggml rope apply op on supported backends
* fix: preserve singleton output dimension when merging LoRA
* feat: add Iris-3B text-to-image support
* fix: halve VAE decode tile size on OOM retry instead of clamping to 256px
* fix: demote resident params to disk residency when memory reclamation fails
* feat: run conditional and unconditional CFG in one batched UNet forward
@wbruna

wbruna commented Oct 11, 2026

Copy link
Copy Markdown
Author

master-956-1b0ba10:

  • fix: halve VAE decode tile size on OOM retry instead of clamping to 256px
  • fix: demote resident params to disk residency when memory reclamation fails
  • feat: run conditional and unconditional CFG in one batched UNet forward

@wbruna wbruna changed the title sd: sync with master-951-f89d9b1 sd: sync with master-956-1b0ba10 Oct 11, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant