Skip to content

vocab : add ufakzeka pre-tokenizer - #29033

Merged
ggerganov merged 2 commits into
ggml-org:masterfrom
stfurkan:ufakzeka-pretok
Sep 18, 2026
Merged

ggerganov merged 2 commits into
ggml-org:masterfrom
stfurkan:ufakzeka-pretok

Conversation

@stfurkan

Copy link
Copy Markdown
Contributor

Overview

Adds a pre-tokenizer type for ufakai/ufakzeka-1, a 151M Turkish model on the Qwen3 architecture (https://huggingface.co/ufakai/ufakzeka-1).

The tokenizer is a byte-level BPE trained on Turkish. Its pre-tokenizer regex is the Qwen2 pattern without the English contraction alternative ('s, 'd, 'll, ...). Turkish attaches suffixes after an apostrophe (Ankara'da, Ali'nin), and the Qwen2 rule splits them differently from how the tokenizer was trained. On a short apostrophe-heavy Turkish sample this changed perplexity from 15.4 to 19.0 and changed two of six greedy answers, so the exact regex is needed.

Changes:

  • src/llama-vocab.h: new enum value LLAMA_VOCAB_PRE_TYPE_UFAKZEKA
  • src/llama-vocab.cpp: the regex for the type, and the mapping from tokenizer.ggml.pre = "ufakzeka"
  • conversion/base.py: the tokenizer hash mapped to "ufakzeka"
  • convert_hf_to_gguf_update.py: the model added to the pre-computed hash list

Additional information

Tested on master at b49650a:

  • the hash in the patch equals the value computed by the update script method for this tokenizer
  • convert_hf_to_gguf.py on the HF model produces a GGUF with tokenizer.ggml.pre = "ufakzeka"
  • llama-tokenize on an apostrophe-heavy Turkish text gives the same 91 token ids as transformers
  • greedy generation with llama-cli matches the transformers output

GGUF files built with this change are at https://huggingface.co/ufakai/ufakzeka-1-GGUF.

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: YES. The patch was drafted with an AI assistant. I reviewed each change, built llama.cpp with it and ran the tests listed above.

Copilot AI lite review requested due to automatic review settings September 17, 2026 14:39
@stfurkan
stfurkan requested a review from CISC as a code owner September 17, 2026 14:39

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟢 Approval recommended

The new pre-tokenizer type is consistently wired through enum, regex selection, GGUF pre-tokenizer string mapping, and conversion hash mapping without introducing unresolved behavioral or integration gaps.

Pull request overview

Adds a new pre-tokenizer variant to support the HuggingFace model ufakai/ufakzeka-1 by introducing a dedicated ufakzeka pre-tokenizer identifier and matching regex behavior in llama.cpp and the HF->GGUF conversion tooling.

Changes:

  • Add LLAMA_VOCAB_PRE_TYPE_UFAKZEKA and wire it into the BPE pre-tokenizer regex switch.
  • Recognize tokenizer.ggml.pre = "ufakzeka" during vocab load and select the new pre-tokenizer type (with clean_spaces = false).
  • Extend conversion tooling to map the tokenizer hash to "ufakzeka" and include it in the pre-computed hash list.
File summaries
File Description
src/llama-vocab.h Adds a new pre-tokenizer enum value for ufakzeka.
src/llama-vocab.cpp Implements the ufakzeka pre-tokenizer regex and maps "ufakzeka" to the new enum during vocab load.
conversion/base.py Maps the ufakzeka tokenizer hash to the "ufakzeka" tokenizer.ggml.pre value.
convert_hf_to_gguf_update.py Adds ufakzeka to the pre-computed tokenizer hash list used to generate/update conversion mappings.
Review details
  • Files reviewed: 4/4 changed files
  • Comments generated: 0
  • Review effort level: Lite

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

@CISC

CISC commented Sep 17, 2026

Copy link
Copy Markdown
Member
  • convert_hf_to_gguf_update.py: the model added to the pre-computed hash list

Why not in the regular models list?

@stfurkan

Copy link
Copy Markdown
Contributor Author
  • convert_hf_to_gguf_update.py: the model added to the pre-computed hash list

Why not in the regular models list?

Hello @CISC, the model repo was private when I opened the PR, it's public now. I moved it to the models list and regenerated the mapping with the script. Thank you

@CISC CISC added the merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge. label Sep 17, 2026
@ggerganov
ggerganov merged commit dc85f89 into ggml-org:master Sep 18, 2026
27 of 31 checks passed
fencerJP pushed a commit to fencerJP/llama-apu that referenced this pull request Sep 19, 2026
* vocab : add ufakzeka pre-tokenizer

* vocab : move ufakzeka to the models list and regenerate the hash mapping
@stfurkan
stfurkan deleted the ufakzeka-pretok branch September 26, 2026 18:52
LadislavSopko pushed a commit to 0ics-srls/llama.cpp that referenced this pull request Oct 5, 2026
* vocab : add ufakzeka pre-tokenizer

* vocab : move ufakzeka to the models list and regenerate the hash mapping
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* vocab : add ufakzeka pre-tokenizer

* vocab : move ufakzeka to the models list and regenerate the hash mapping
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion merge ready A maintainer can use this label to indicate that they consider the changes final and ready to merge.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants