Repository navigation
Add Support for Xing4.0 - #29012
Add Support for Xing4.0#29012shuxiaoqiong wants to merge 12 commits into
Conversation
|
Hi @shuxiaoqiong, thanks for your contribution! Per our contribution guidelines, the automated PR checker found the following issue(s) that need your attention:
Please note that maintainers reserve the right to make final decisions on PRs. If you believe there is a mistake, please comment below. |
|
Looks like it needs a rebase ♻️ |
…en.py -> conversion/xing.py (class Xing4_0Model, keeps XingChen4ForCausalLM registered for legacy HF checkpoints)- src/models/xingchen4.cpp -> src/models/xing4_0.cpp- ggml/src/ggml-cuda/xc4-hc.cu/.cuh -> xing4_0-hc.cu/.cuh- GGUF arch string: xingchen4 -> xing4_0 (old GGUFs will not load)- ggml ops: GGML_OP_XC4_HC_* -> GGML_OP_XING4_0_HC_*, ggml_xc4_hc_* -> ggml_xing4_0_hc_*- cparams/graph fused-op flags and enums renamed accordingly- tests updated (test-llama-archs, test-backend-ops)
63c16fb to
8c37919
Compare
Overview
I opened this PR to support TeleAI's new model: Xing4.0-29B-A4B. The Xing4.0-29B-A4B model is a domestically developed LLM from China Telecom AI Technology Co., Ltd. (TeleAI), originally named "telechat,"or "xingchen4" and is scheduled to be open-sourced soon.
Since Xing4.0-29B-A4B is primarily designed for personal PC deployment, and the 4-bit quantized model still weighs 19GB, a pure CPU deployment approach has little practical value. This PR implements both CPU backend and GPU backend support.
So far, I have completed the following work:
Additional information
I have opened an issue and initiated a discussion regarding this, you can check it at #28153
I have tested the perplexity performance of the Xing4.0-29B-A4B model under this scheme, as well as its llama-bench performance, etc. The test data are as follows:
llama-bench:
Model | Quant | Size | Params | ngl | pp512 (t/s) | tg128 (t/s)
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | -1 | 4196.77 ± 130.33 | 106.43 ± 0.57
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | 99 | 4174.66 ± 147.84 | 106.46 ± 0.62
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | -1 | 5196.35 ± 612.33 | 151.65 ± 0.65
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | 99 | 5212.54 ± 608.44 | 151.69 ± 0.65
llama-perplexity:
Model | Quant | Size | Params | ngl | PPL
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | -1 | 8.6697 +/- 0.09061
Xing4.0-29B-A4B | F16 | 59 GiB | 29.51 B | 99 | 8.6697 +/- 0.09061
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | -1 | 8.7699 +/- 0.09188
Xing4.0-29B-A4B | IQ4_NL | 19 GiB | 29.51 B | 99 | 8.7699 +/- 0.09188
So difference = 0.1002, which is a reasonable degradation for 4-bit quantization with only ~1% PPL increase.
In addition, the newly added operators – XING4_0_HC_PRE, XING4_0_HC_COMB, andXING4_0_HC_POST – have all passed the test-backend-ops tests.
Requirements