Repository navigation
vocab : add ufakzeka pre-tokenizer - #29033
Conversation
There was a problem hiding this comment.
🟢 Approval recommended
The new pre-tokenizer type is consistently wired through enum, regex selection, GGUF pre-tokenizer string mapping, and conversion hash mapping without introducing unresolved behavioral or integration gaps.
Pull request overview
Adds a new pre-tokenizer variant to support the HuggingFace model ufakai/ufakzeka-1 by introducing a dedicated ufakzeka pre-tokenizer identifier and matching regex behavior in llama.cpp and the HF->GGUF conversion tooling.
Changes:
- Add
LLAMA_VOCAB_PRE_TYPE_UFAKZEKAand wire it into the BPE pre-tokenizer regex switch. - Recognize
tokenizer.ggml.pre = "ufakzeka"during vocab load and select the new pre-tokenizer type (withclean_spaces = false). - Extend conversion tooling to map the tokenizer hash to
"ufakzeka"and include it in the pre-computed hash list.
File summaries
| File | Description |
|---|---|
| src/llama-vocab.h | Adds a new pre-tokenizer enum value for ufakzeka. |
| src/llama-vocab.cpp | Implements the ufakzeka pre-tokenizer regex and maps "ufakzeka" to the new enum during vocab load. |
| conversion/base.py | Maps the ufakzeka tokenizer hash to the "ufakzeka" tokenizer.ggml.pre value. |
| convert_hf_to_gguf_update.py | Adds ufakzeka to the pre-computed tokenizer hash list used to generate/update conversion mappings. |
Review details
- Files reviewed: 4/4 changed files
- Comments generated: 0
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Why not in the regular |
Hello @CISC, the model repo was private when I opened the PR, it's public now. I moved it to the models list and regenerated the mapping with the script. Thank you |
* vocab : add ufakzeka pre-tokenizer * vocab : move ufakzeka to the models list and regenerate the hash mapping
* vocab : add ufakzeka pre-tokenizer * vocab : move ufakzeka to the models list and regenerate the hash mapping
* vocab : add ufakzeka pre-tokenizer * vocab : move ufakzeka to the models list and regenerate the hash mapping
Overview
Adds a pre-tokenizer type for ufakai/ufakzeka-1, a 151M Turkish model on the Qwen3 architecture (https://huggingface.co/ufakai/ufakzeka-1).
The tokenizer is a byte-level BPE trained on Turkish. Its pre-tokenizer regex is the Qwen2 pattern without the English contraction alternative ('s, 'd, 'll, ...). Turkish attaches suffixes after an apostrophe (Ankara'da, Ali'nin), and the Qwen2 rule splits them differently from how the tokenizer was trained. On a short apostrophe-heavy Turkish sample this changed perplexity from 15.4 to 19.0 and changed two of six greedy answers, so the exact regex is needed.
Changes:
src/llama-vocab.h: new enum valueLLAMA_VOCAB_PRE_TYPE_UFAKZEKAsrc/llama-vocab.cpp: the regex for the type, and the mapping fromtokenizer.ggml.pre = "ufakzeka"conversion/base.py: the tokenizer hash mapped to"ufakzeka"convert_hf_to_gguf_update.py: the model added to the pre-computed hash listAdditional information
Tested on master at b49650a:
convert_hf_to_gguf.pyon the HF model produces a GGUF withtokenizer.ggml.pre = "ufakzeka"llama-tokenizeon an apostrophe-heavy Turkish text gives the same 91 token ids as transformersllama-climatches the transformers outputGGUF files built with this change are at https://huggingface.co/ufakai/ufakzeka-1-GGUF.
Requirements