Skip to content

Add Support for Xing4.0 (cleanup) - #29141

Closed
fairydreaming wants to merge 19 commits into
ggml-org:masterfrom
fairydreaming:xing4_0-port
Closed

fairydreaming wants to merge 19 commits into
ggml-org:masterfrom
fairydreaming:xing4_0-port

Conversation

@fairydreaming

Copy link
Copy Markdown
Contributor

Overview

This is an experimental PR where I'm trying to clean up #29012 by simplifying/removing duplicate code.

Requirements

shixq7 and others added 13 commits September 16, 2026 16:51
…en.py -> conversion/xing.py (class Xing4_0Model, keeps XingChen4ForCausalLM registered for legacy HF checkpoints)- src/models/xingchen4.cpp -> src/models/xing4_0.cpp- ggml/src/ggml-cuda/xc4-hc.cu/.cuh -> xing4_0-hc.cu/.cuh- GGUF arch string: xingchen4 -> xing4_0 (old GGUFs will not load)- ggml ops: GGML_OP_XC4_HC_* -> GGML_OP_XING4_0_HC_*, ggml_xc4_hc_* -> ggml_xing4_0_hc_*- cparams/graph fused-op flags and enums renamed accordingly- tests updated (test-llama-archs, test-backend-ops)
@github-actions github-actions Bot added model Model specific testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend conversion labels Sep 19, 2026
@fairydreaming

Copy link
Copy Markdown
Contributor Author

@ggerganov @am17an I managed to reuse DSv4 mHC OPs in Xing4.0 impl while maintaining bit-exact results, but I had to add another variant of ggml_dsv4_hc_comb() - ggml_dsv4_hc_comb_clamp() with added float limit argument (and one additional ggml_transpose() call in the model implementation).

Internally the clamp variant works a little different from the original. In DeepSeek V4 positivity of the comb matrix before Sinkhorn transform was enforced by performing softmax and adding eps to all matrix elements, while in the variant with clamping positivity of the matrix is enforced by the choice of the limit value (so that exp() of clamped values - max is non-zero when expressed in float) then the Sinkhorn loop can perform all norm iterations with eps added only in the denominator.

I wanted to know if this solution is acceptable or you have some other idea before I mark this ready for full review.

@am17an am17an left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

if you introduce new params you need to also take care of other backend which support this op to return false in supports_op. In terms of the actual impl for this model it looks ok to me to do this.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants