Skip to content

common : add LLM-jp-4.1 Harmony dialect handler (draft for review) - #1

Draft
e-mon wants to merge 28 commits into
masterfrom
llmjp-harmony-handler
Draft

e-mon wants to merge 28 commits into
masterfrom
llmjp-harmony-handler

Conversation

@e-mon

@e-mon e-mon commented Sep 12, 2026 •

Copy link
Copy Markdown
Owner

Summary

Adds a dedicated chat-format handler for LLM-jp-4.1 (llm-jp/llm-jp-4.1-*-thinking), a model that uses the GPT-OSS (Harmony) format with two dialect differences:

  1. Its SentencePiece-style tokenizer emits a leading space after every special token, so the model's output reads <|channel|> analysis<|message|> ... / <|start|> assistant to=functions.X<|channel|> commentary <|constrain|> json<|message|> {...} instead of the unspaced GPT-OSS form. With the GPT-OSS handler, llama-server returns 500 "The model produced output that does not match the expected peg-native format" for tool requests and 400 "Failed to initialize samplers" for tool_choice: required.
  2. Parallel tool calls are rendered as consecutive assistant messages, every call but the last closed by <|end|> and the last one by <|call|>.

The template declares the dialect on its first line ({#- chat_format=llm-jp-harmony-v1 -#}), which is what the handler selection keys on, so no other template is affected.

Design

  • common/parsers/llm-jp-harmony.cpp: a self-contained handler. Prompt handling (message delimiters, thinking tags, continuation, the <|return|> replacement) is that of common_chat_params_init_gpt_oss; the PEG is the GPT-OSS one with the two differences built in. The gpt-oss-20b-specific rules (stray commentary prefix, unsolicited builtin calls) are not carried over: they never occur in this model's output (0 of 400 BFCL generations). The GPT-OSS handler is not modified; the only additions outside the new file are the declaration in parsers.h, the source list entry, the selection branch in common/chat.cpp, the template, tests and docs.
  • Special tokens in headers are followed by [ ]*; <|message|> by at most one space, so intentional leading whitespace in a message body survives (" padded" → " padded"). Plain spaces are used rather than p.space() because the GBNF space rule allows a single space, while <|constrain|> json carries the template's space plus the tokenizer's.
  • Parallel calls: tool_choice (end start tool_choice)* inside the trigger rule when parallel_tool_calls is set, so the lazy grammar covers every call; with the flag off the grammar constrains the model to one call, as for other formats.
  • Lazy-grammar trigger patterns are the GPT-OSS ones with \s* after <|start|> / <|channel|>.
  • Selection in common/chat.cpp before the GPT-OSS check; template added to models/templates/; docs rows added.

Tests

  • tests/test-chat.cpp: new LLM-jp-4.1 block (final channel, one-space body rule, reasoning + final, partial reasoning, tool call in role/channel header with <|constrain|> json, parallel calls with <|end|>, structured output). ./build/bin/test-chat → [chat] All tests passed! (the peg_tester also validates that the generated grammar accepts each input).
  • E2E with the released llm-jp/llm-jp-4.1-8b-thinking-gguf (Q4_K_M) on llama-server (temperature 0, template embedded in the GGUF, no --chat-template-file):
case result
reasoning + final reasoning_content and content separated; no leading space in content
single tool call (tool_choice: auto) get_weather {"city": "Tokyo"} (was 500)
parallel tool calls (parallel_tool_calls: true) 3 calls returned for a prompt asking for three cities at once
tool_choice: required tool call returned (was 400)
response_format: json_schema constrained JSON returned
streaming reasoning_content deltas, then content
tool-result continuation (single, and 2 parallel results attributed by position) final answer uses the results

Run on Linux (Ubuntu 22.04 on WSL2, RTX 3090, CUDA 12.6 build of this commit): test-chat passes and all cases pass. The same cases also pass on macOS (Metal) with the pre-rebase build of the handler (identical parser code).

Notes

Differential check against the official parser on BFCL prompts

400 BFCL v3 prompts (simple / multiple / parallel / parallel_multiple, 100 each, drawn at random with a fixed seed from the 1,000 non-live AST prompts) were run through llama-server (this branch, CUDA, temperature 0, 4096-token limit, -v + return_tokens: true so that the chat response exposes the generated token ids). The very same token ids were then parsed by the official reference parser (HarmonyMessageParser from the llm-jp-vllm plugin) and compared with what the handler returned:

agree
tool calls (names + JSON arguments) 400 / 400
reasoning_content 400 / 400
content 400 / 400
special-token leaks / leading spaces 0
invalid tool names 0 (the grammar enforces the declared names)

Outputs covered single calls, 2–4 parallel calls separated by <|end|>, text answers without a call, and 11 generations cut off by the token limit. With the official bfcl-eval AST checker the same outputs score 83.0% (86 / 85 / 86 / 75 per category; 95% CI 79.0–86.4%); every miss is a model-side argument or count error, a text answer, or a truncation, none is a parsing error.

Precedents

Self-contained handlers added for one model family without touching the others: ggml-org#19931 (GigaChat v3/3.1, by the model vendor), ggml-org#24615 (Cohere2MoE), ggml-org#26210 (MiniMax M3), ggml-org#21418 (Gemma 4); the per-model file layout comes from ggml-org#27764, and ggml-org#20393 is the PEG rework of the GPT-OSS handler this one mirrors.

@e-mon
e-mon force-pushed the llmjp-harmony-handler branch 3 times, most recently from 66b90f6 to d71fd74 Compare September 17, 2026 11:28
IMbackK and others added 25 commits September 28, 2026 13:47
…ims on x86 (ggml-org#29423)

* ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86

* add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail

* ggml-cpu: fix FA softcap handling for padded KV tiles
* tests : use llama_context_ptr in test-recurrent-state-rollback

Replace raw llama_context pointers with llama_context_ptr and drop the
manual llama_free calls and cleanup lambda.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : run test-recurrent-state-rollback over all dummy models

Add a --models DIR mode that mirrors test-save-load-state: iterate every
dummy model, report PASS/FAIL/SKIP in a table and fail only when a model
fails. Register a single ctest entry with ARGS --models instead of the four
per-model registrations.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* cont : fix typo

* metal : allow fusing 0-element nodes to keep graph packing shape-independent

The fusion packing in ggml_metal_fusion_max excluded 0-element tensors and
the topk_moe/moe_reduce checks rejected n_tokens == 0, so graphs decoding
batches with no outputs packed differently from the worst-case reserved
graph. The Metal optimizer then reordered the nodes differently and
ggml_gallocr_needs_realloc failed on the layout mismatch, forcing an
unexpected graph re-reserve (caught by GGML_SCHED_DEBUG_REALLOC).

Treat empty tensors like their non-empty counterparts: match them in the
pattern sequence and only reject genuinely malformed shapes. Fused kernels
dispatch zero threadgroups for empty graphs, which is a legal no-op.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : run test_multi_seq_split_replay as a separate test

test_multi_seq_split_replay was invoked at the end of test_rollback,
so its result was folded into the rollback status and it only ran when
the rollback part passed.

Give it its own test_status return, run both tests independently over
both cache fills via a shared run_tests helper, and report them as
separate rollback / split replay columns in the --models table with
per-test summaries. The exit code fails when either test fails.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : loosen the split replay nmse bound to 1e-4

test-generate-models seeds its weights from std::random_device, and some
generated lfm2 models drift up to ~1.7e-5 nmse on the split replay due to
rounding noise, tripping the previous 1e-5 bound. Raise the bound to 1e-4
so the random generations stop flaking.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL

* tests : reuse run_tests_for_model in single-model mode

The single-model path duplicated the model init and the non-recurrent
check from run_tests_for_model; route it through the shared helper
instead. Model load failures now return FAIL rather than SKIP so that
--model with a broken file still exits non-zero, and the helper loads
with model_only like the --models loop does since the tests create
their own contexts.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
…sor (ggml-org#29471)

* Fix:  Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor

* Clang formatting
Supersedes ggml-org#29158

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
)

* adapt common

* add common_batch

* wip

* wip: spec

* cont

* common_speculative_process

* server_batch to use common_batch

* rm some stale calls

Assisted-by: Claude Fable 5.1

* migrate mtmd

* handle imrope, handle return val of add()/add_embd()

* add spec zeros vector

* add warning on zero fill path
* models: pad on the left with ggml_pad_ext

The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders
build a left padding as a right pad followed by a roll, and DFlash2
concatenates a zero filled block in front of the previous tokens.
ggml_pad_ext does both in one node now that every backend supports a
left padding. The Gemma 4 audio embeddings are bit identical.

* models: skip the DFlash2 taps that only read padding

A tap at or past block_size shifts every row out of the block, so its
term is zero. The loop runs min(kernel_size, block_size) taps.
Fixes the compile error: `no template named 'function' in namespace 'std'`
…eddings endpoint (ggml-org#29556)

* server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding)

Accept the OpenAI-style wrapped content array format for multimodal
embedding requests. Each {"content": [...]} object is one input that
produces one embedding; text parts are concatenated and image_url parts
are decoded via handle_media then spliced with process_mtmd_prompt.

The legacy formats (plain string, token arrays, mixed arrays, and the
{prompt_string, multimodal_data} object) continue to work unchanged via
tokenize_input_prompts. Bare content arrays (the unwrapped shape) are
rejected with a migration message.

Also disables KV prefix reuse for stateless embedding/rerank tasks so
that repeated inputs do not incorrectly share cached KV across requests.

Assisted-by: Opencode Qwen3.8 27B

* clean up comments and docs

* refactor

* add tests

* support video and audio inp

---------

Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com>
Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
* ci : update the oneAPI toolkit to 2026.1

oneDNN is removed from Intel Deep Learning Essentials in 2026.0, so
staying on the deep-learning-essentials path would silently lose oneDNN
support when the toolkit version is updated. Switch both the Ubuntu and
Windows CI jobs to the new unified Intel oneAPI Toolkit installer,
which still includes oneDNN (until 2027.0) and keeps the component IDs
unchanged for the Windows install script.

Measured with the same code (b10899) built with oneAPI 2026.1 vs the
2025.3-based release build on Arc B570: prompt processing 1331 vs 434
t/s (3.1x), token generation 50.1 vs 45.3-48.0 t/s.

Assisted-by: GLM (z-ai/glm-5.3-flash)

* docs : update the SYCL backend build requirements for oneAPI 2026.1

With the 2026.0 release the Base toolkit and the HPC toolkit are
combined into the oneAPI Toolkit, and oneDNN is removed from the Deep
Learning Essentials package. Update the install instructions, the
verified release table and the news section accordingly.

Assisted-by: GLM (z-ai/glm-5.3-flash)

* ci : update the release workflow for oneAPI 2026.1 and Level Zero SDK 1.33.1

Align the release package build with the CI build update:
- oneAPI toolkit 2025.3.3 -> 2026.1 (the unified oneAPI Toolkit)
- Level Zero SDK 1.28.2 -> 1.33.1, and the Debian package names
  (level-zero/level-zero-devel -> libze1/libze-dev)
- The Windows DLL copy list for the 2026.1 runtime: sycl9.dll and the
  .6/.3 MKL library versions

Assisted-by: GLM (z-ai/glm-5.3-flash)

* ci : remove the removed .spv fallback files from the Windows DLL copy list

oneAPI 2026.1 no longer ships libsycl-fallback-bfloat16.spv and
libsycl-native-bfloat16.spv (the OpenCL fallback mechanism changed), so
the copy step failed with exit 1.

Assisted-by: GLM (z-ai/glm-5.3-flash)

* devops : update the oneAPI toolkit image in the Intel Dockerfile

Assisted-by: GLM (z-ai/glm-5.3-flash)

---------

Co-authored-by: Asahi-Prv <Asahi-Prv@users.noreply.github.com>
- Avoid useless string conversions on Windows.
- No need for BSD or emscripten special cases.

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…inja (ggml-org#29615)

* chat : fix Muse Glimmer ignoring response_format json_schema with --jinja

Fixes ggml-org#29613

* chat : accept json fences and clean up

* chat : fix choice parenthesis

---------

Co-authored-by: Alde Rojas <hello@alde.dev>
…9280)

* vulkan : reuse descriptor sets when bindings are constant

* vulkan : bump buffer_destroy_count before destroying the buffer
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* ggml : speed up model loading

A crafted model could hang the server for a very long time, try with:

    llama-cli -hf angt/test-gguf-1Mkv -hff model.gguf

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

* Avoid empty keys

Signed-off-by: Adrien Gallouët <angt@huggingface.co>

* Fix

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

---------

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
ggml-org#29624)

* musa: build the docker image and CI container from the MUSA SDK images

Use registry.mthreads.com/mcconline/musa_sdk:5.2.0-{devel,runtime}-ubuntu22.04-s5000
instead of registry.mthreads.com/mcconline/inference/pytorch:2.9.1.post1-py3.10-musa5.2.0-mp31-devel-ubuntu22.04-amd64
for the MUSA docker image and the MUSA CI container, and let the runtime stage use the
runtime image instead of reusing the devel one, which drops the MUSA toolchain from the
published images.

* musa: install the MUSA headers and loader path the SDK images omit

musa_sdk:5.2.0-*-s5000 does not ship the cub and thrust headers that the MUSA
backend builds against, and its runtime image does not register
/usr/local/musa/lib with the dynamic loader.

Install both header packages in the build stage and in the MUSA CI container,
and write the loader path in the runtime stage.

* musa: install libmthreads-compute for the MUSA runtime library

The MUSA SDK images do not install libmthreads-compute, which provides
libmusa.so.1 in /usr/lib/x86_64-linux-gnu, so linking anything against the
MUSA backend fails.

* musa: install libmthreads-compute in the runtime stages

The MUSA runtime image does not install libmthreads-compute, so the published
images would have no libmusa.so.1 at run time.

---------

Co-authored-by: yeahdongcn <yeahdongcn@users.noreply.github.com>
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
@e-mon
e-mon force-pushed the llmjp-harmony-handler branch from d71fd74 to 5d4a53f Compare September 29, 2026 10:19
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 29, 2026
ggerganov and others added 2 commits September 29, 2026 13:37
graph_inputs was populated while splitting the graph, so it only
contained the inputs that are used as srcs of some node. With pipeline
parallelism (n_copies > 1), each graph input contributes n_copies leafs
to graph_copy, so switching between batches that consume different
inputs (e.g. token batches that do not use the embeddings input vs
image batches that do) changed the graph composition. This shifted the
input copies in graph_copy, making the backend ids comparison report
spurious changes and forcing the scheduler to re-reserve. The
re-reserve could then record smaller input sizes (e.g. out_ids with
n_outputs = 0) and abort later on a graph with an unchanged size via
GGML_SCHED_DEBUG_REALLOC.

Collect the inputs after the split instead, from all input leafs of the
graph, so that the graph composition depends only on which inputs
exist, not on which inputs are used.

Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
- Check for buffered write errors when closing downloaded files.
- Use UTF-8 paths when writing ETag files on Windows.
- Write in binary mode on Windows.

Signed-off-by: Adrien Gallouët <angt@huggingface.co>
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space
after every special token and parallel tool calls are separated by
<|end|>. The GPT-OSS handler rejects this output, so add a dedicated
handler, selected by the chat_format=llm-jp-harmony-v1 declaration in
the chat template.

Assisted-by: Claude Fable 5.1
@e-mon
e-mon force-pushed the llmjp-harmony-handler branch from b546eee to 0213322 Compare September 29, 2026 12:28
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.