Skip to content

Add Support for Xing4.0 - #29012

Open
shuxiaoqiong wants to merge 12 commits into
ggml-org:masterfrom
shuxiaoqiong:xing4_0-port
Open

shuxiaoqiong wants to merge 12 commits into
ggml-org:masterfrom
shuxiaoqiong:xing4_0-port

Conversation

@shuxiaoqiong

@shuxiaoqiong shuxiaoqiong commented Sep 17, 2026 •

Copy link
Copy Markdown

Overview

I opened this PR to support TeleAI's new model: Xing4.0-29B-A4B. The Xing4.0-29B-A4B model is a domestically developed LLM from China Telecom AI Technology Co., Ltd. (TeleAI), originally named "telechat,"or "xingchen4" and is scheduled to be open-sourced soon.

Since Xing4.0-29B-A4B is primarily designed for personal PC deployment, and the 4-bit quantized model still weighs 19GB, a pure CPU deployment approach has little practical value. This PR implements both CPU backend and GPU backend support.

So far, I have completed the following work:

  • Registered the Xing4.0-29B-A4B model architecture in the llama.cpp framework with parameter alignment.
  • Implemented conversion and quantization of the model's safetensors weights to GGUF format.
  • Added CPU backend inference support.
  • Added GPU backend inference support, along with specific GGML operator implementations.
  • Wrote and tested operator test cases.
  • Conducted llama-perplexity and llama-bench evaluations.

Additional information

I have opened an issue and initiated a discussion regarding this, you can check it at #28153

I have tested the perplexity performance of the Xing4.0-29B-A4B model under this scheme, as well as its llama-bench performance, etc. The test data are as follows:

llama-bench:
Model | Quant | Size | Params | ngl | pp512 (t/s) | tg128 (t/s)
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | -1 | 4196.77 ± 130.33 | 106.43 ± 0.57
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | 99 | 4174.66 ± 147.84 | 106.46 ± 0.62
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | -1 | 5196.35 ± 612.33 | 151.65 ± 0.65
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | 99 | 5212.54 ± 608.44 | 151.69 ± 0.65

llama-perplexity:
Model | Quant | Size | Params | ngl | PPL
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | -1 | 8.6697 +/- 0.09061
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | 99 | 8.6697 +/- 0.09061
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | -1 | 8.7699 +/- 0.09188
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | 99 | 8.7699 +/- 0.09188

So difference = 0.1002, which is a reasonable degradation for 4-bit quantization with only ~1% PPL increase.

In addition, the newly added operators – XING4_0_HC_PRE, XING4_0_HC_COMB, andXING4_0_HC_POST – have all passed the test-backend-ops tests.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES
  • AI helped me complete the steps of model file loading, conversion, quantization, and initial operator implementation, among others.

@ggml-gh-bot

ggml-gh-bot Bot commented Sep 17, 2026

Copy link
Copy Markdown

Hi @shuxiaoqiong, thanks for your contribution!

Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:

  • Multiple backend changes in one PR: When adding support for a new model or feature, focus on CPU support only in the initial PR. Add support for other backends like CUDA in follow-up PRs. If you have a good reason to modify multiple backends in one PR, please explain it.

  • Large PR: Large changes require prior discussion (e.g. an issue or RFC) and maintainers may not be able to review this PR as-is. Consider splitting it into smaller, focused PRs.


Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below.

@github-actions github-actions Bot added model Model specific testing Everything test related ggml changes relating to the ggml tensor library for machine learning CUDA Related to the CUDA backend conversion labels Sep 17, 2026
@bricss

bricss commented Oct 4, 2026

Copy link
Copy Markdown

Looks like it needs a rebase ♻️

shixq7 added 12 commits October 8, 2026 09:57
…en.py -> conversion/xing.py (class Xing4_0Model, keeps XingChen4ForCausalLM registered for legacy HF checkpoints)- src/models/xingchen4.cpp -> src/models/xing4_0.cpp- ggml/src/ggml-cuda/xc4-hc.cu/.cuh -> xing4_0-hc.cu/.cuh- GGUF arch string: xingchen4 -> xing4_0 (old GGUFs will not load)- ggml ops: GGML_OP_XC4_HC_* -> GGML_OP_XING4_0_HC_*, ggml_xc4_hc_* -> ggml_xing4_0_hc_*- cparams/graph fused-op flags and enums renamed accordingly- tests updated (test-llama-archs, test-backend-ops)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants