Repository navigation
Conversation
e-mon
force-pushed
the
llmjp-harmony-handler
branch
3 times, most recently
from
September 17, 2026 11:28
66b90f6 to
d71fd74
Compare
…ims on x86 (ggml-org#29423) * ggml-cpu: enable tiled flash attention for non-vector-multiple head dims on x86 * add AVX2 support for masked loading and storing in simd_gemm_ukernel_tail * ggml-cpu: fix FA softcap handling for padded KV tiles
* tests : use llama_context_ptr in test-recurrent-state-rollback Replace raw llama_context pointers with llama_context_ptr and drop the manual llama_free calls and cleanup lambda. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : run test-recurrent-state-rollback over all dummy models Add a --models DIR mode that mirrors test-save-load-state: iterate every dummy model, report PASS/FAIL/SKIP in a table and fail only when a model fails. Register a single ctest entry with ARGS --models instead of the four per-model registrations. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * cont : fix typo * metal : allow fusing 0-element nodes to keep graph packing shape-independent The fusion packing in ggml_metal_fusion_max excluded 0-element tensors and the topk_moe/moe_reduce checks rejected n_tokens == 0, so graphs decoding batches with no outputs packed differently from the worst-case reserved graph. The Metal optimizer then reordered the nodes differently and ggml_gallocr_needs_realloc failed on the layout mismatch, forcing an unexpected graph re-reserve (caught by GGML_SCHED_DEBUG_REALLOC). Treat empty tensors like their non-empty counterparts: match them in the pattern sequence and only reject genuinely malformed shapes. Fused kernels dispatch zero threadgroups for empty graphs, which is a legal no-op. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : run test_multi_seq_split_replay as a separate test test_multi_seq_split_replay was invoked at the end of test_rollback, so its result was folded into the rollback status and it only ran when the rollback part passed. Give it its own test_status return, run both tests independently over both cache fills via a shared run_tests helper, and report them as separate rollback / split replay columns in the --models table with per-test summaries. The exit code fails when either test fails. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : loosen the split replay nmse bound to 1e-4 test-generate-models seeds its weights from std::random_device, and some generated lfm2 models drift up to ~1.7e-5 nmse on the split replay due to rounding noise, tripping the previous 1e-5 bound. Raise the bound to 1e-4 so the random generations stop flaking. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL * tests : reuse run_tests_for_model in single-model mode The single-model path duplicated the model init and the non-recurrent check from run_tests_for_model; route it through the shared helper instead. Model load failures now return FAIL rather than SKIP so that --model with a broken file still exits non-zero, and the helper loads with model_only like the --models loop does since the tests create their own contexts. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
…sor (ggml-org#29471) * Fix: Handle unaligned writes in ggml_backend_webgpu_buffer_set_tensor * Clang formatting
Supersedes ggml-org#29158 Signed-off-by: Adrien Gallouët <angt@huggingface.co>
) * adapt common * add common_batch * wip * wip: spec * cont * common_speculative_process * server_batch to use common_batch * rm some stale calls Assisted-by: Claude Fable 5.1 * migrate mtmd * handle imrope, handle return val of add()/add_embd() * add spec zeros vector * add warning on zero fill path
* models: pad on the left with ggml_pad_ext The Parakeet, LFM2-Audio, Granite Speech and Gemma 4 audio encoders build a left padding as a right pad followed by a roll, and DFlash2 concatenates a zero filled block in front of the previous tokens. ggml_pad_ext does both in one node now that every backend supports a left padding. The Gemma 4 audio embeddings are bit identical. * models: skip the DFlash2 taps that only read padding A tap at or past block_size shifts every row out of the block, so its term is zero. The loop runs min(kernel_size, block_size) taps.
Fixes the compile error: `no template named 'function' in namespace 'std'`
…eddings endpoint (ggml-org#29556) * server : support multimodal input for /v1/embeddings (Qwen3-VL-Embedding) Accept the OpenAI-style wrapped content array format for multimodal embedding requests. Each {"content": [...]} object is one input that produces one embedding; text parts are concatenated and image_url parts are decoded via handle_media then spliced with process_mtmd_prompt. The legacy formats (plain string, token arrays, mixed arrays, and the {prompt_string, multimodal_data} object) continue to work unchanged via tokenize_input_prompts. Bare content arrays (the unwrapped shape) are rejected with a migration message. Also disables KV prefix reuse for stateless embedding/rerank tasks so that repeated inputs do not incorrectly share cached KV across requests. Assisted-by: Opencode Qwen3.8 27B * clean up comments and docs * refactor * add tests * support video and audio inp --------- Co-authored-by: timothywang21 <timothywang21@users.noreply.github.com> Co-authored-by: Xuan Son Nguyen <son@huggingface.co>
* ci : update the oneAPI toolkit to 2026.1 oneDNN is removed from Intel Deep Learning Essentials in 2026.0, so staying on the deep-learning-essentials path would silently lose oneDNN support when the toolkit version is updated. Switch both the Ubuntu and Windows CI jobs to the new unified Intel oneAPI Toolkit installer, which still includes oneDNN (until 2027.0) and keeps the component IDs unchanged for the Windows install script. Measured with the same code (b10899) built with oneAPI 2026.1 vs the 2025.3-based release build on Arc B570: prompt processing 1331 vs 434 t/s (3.1x), token generation 50.1 vs 45.3-48.0 t/s. Assisted-by: GLM (z-ai/glm-5.3-flash) * docs : update the SYCL backend build requirements for oneAPI 2026.1 With the 2026.0 release the Base toolkit and the HPC toolkit are combined into the oneAPI Toolkit, and oneDNN is removed from the Deep Learning Essentials package. Update the install instructions, the verified release table and the news section accordingly. Assisted-by: GLM (z-ai/glm-5.3-flash) * ci : update the release workflow for oneAPI 2026.1 and Level Zero SDK 1.33.1 Align the release package build with the CI build update: - oneAPI toolkit 2025.3.3 -> 2026.1 (the unified oneAPI Toolkit) - Level Zero SDK 1.28.2 -> 1.33.1, and the Debian package names (level-zero/level-zero-devel -> libze1/libze-dev) - The Windows DLL copy list for the 2026.1 runtime: sycl9.dll and the .6/.3 MKL library versions Assisted-by: GLM (z-ai/glm-5.3-flash) * ci : remove the removed .spv fallback files from the Windows DLL copy list oneAPI 2026.1 no longer ships libsycl-fallback-bfloat16.spv and libsycl-native-bfloat16.spv (the OpenCL fallback mechanism changed), so the copy step failed with exit 1. Assisted-by: GLM (z-ai/glm-5.3-flash) * devops : update the oneAPI toolkit image in the Intel Dockerfile Assisted-by: GLM (z-ai/glm-5.3-flash) --------- Co-authored-by: Asahi-Prv <Asahi-Prv@users.noreply.github.com>
Assisted-by: pi:llama.cpp/Qwen3.8-27B
- Avoid useless string conversions on Windows. - No need for BSD or emscripten special cases. Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…inja (ggml-org#29615) * chat : fix Muse Glimmer ignoring response_format json_schema with --jinja Fixes ggml-org#29613 * chat : accept json fences and clean up * chat : fix choice parenthesis --------- Co-authored-by: Alde Rojas <hello@alde.dev>
…9280) * vulkan : reuse descriptor sets when bindings are constant * vulkan : bump buffer_destroy_count before destroying the buffer
Assisted-by: Claude Code
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
* ggml : speed up model loading
A crafted model could hang the server for a very long time, try with:
llama-cli -hf angt/test-gguf-1Mkv -hff model.gguf
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* Avoid empty keys
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
* Fix
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
---------
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
ggml-org#29624) * musa: build the docker image and CI container from the MUSA SDK images Use registry.mthreads.com/mcconline/musa_sdk:5.2.0-{devel,runtime}-ubuntu22.04-s5000 instead of registry.mthreads.com/mcconline/inference/pytorch:2.9.1.post1-py3.10-musa5.2.0-mp31-devel-ubuntu22.04-amd64 for the MUSA docker image and the MUSA CI container, and let the runtime stage use the runtime image instead of reusing the devel one, which drops the MUSA toolchain from the published images. * musa: install the MUSA headers and loader path the SDK images omit musa_sdk:5.2.0-*-s5000 does not ship the cub and thrust headers that the MUSA backend builds against, and its runtime image does not register /usr/local/musa/lib with the dynamic loader. Install both header packages in the build stage and in the MUSA CI container, and write the loader path in the runtime stage. * musa: install libmthreads-compute for the MUSA runtime library The MUSA SDK images do not install libmthreads-compute, which provides libmusa.so.1 in /usr/lib/x86_64-linux-gnu, so linking anything against the MUSA backend fails. * musa: install libmthreads-compute in the runtime stages The MUSA runtime image does not install libmthreads-compute, so the published images would have no libmusa.so.1 at run time. --------- Co-authored-by: yeahdongcn <yeahdongcn@users.noreply.github.com>
Signed-off-by: Aaron Teo <aaron.teo1@ibm.com>
e-mon
force-pushed
the
llmjp-harmony-handler
branch
from
September 29, 2026 10:19
d71fd74 to
5d4a53f
Compare
graph_inputs was populated while splitting the graph, so it only contained the inputs that are used as srcs of some node. With pipeline parallelism (n_copies > 1), each graph input contributes n_copies leafs to graph_copy, so switching between batches that consume different inputs (e.g. token batches that do not use the embeddings input vs image batches that do) changed the graph composition. This shifted the input copies in graph_copy, making the backend ids comparison report spurious changes and forcing the scheduler to re-reserve. The re-reserve could then record smaller input sizes (e.g. out_ids with n_outputs = 0) and abort later on a graph with an unchanged size via GGML_SCHED_DEBUG_REALLOC. Collect the inputs after the split instead, from all input leafs of the graph, so that the graph composition depends only on which inputs exist, not on which inputs are used. Assisted-by: pi:llama.cpp/MiMo-V2.6-Flash-RL
- Check for buffered write errors when closing downloaded files. - Use UTF-8 paths when writing ETag files on Windows. - Write in binary mode on Windows. Signed-off-by: Adrien Gallouët <angt@huggingface.co>
e-mon
force-pushed
the
llmjp-harmony-handler
branch
2 times, most recently
from
September 29, 2026 12:19
e94f5fb to
8d86d74
Compare
e-mon
force-pushed
the
llmjp-harmony-handler
branch
from
September 29, 2026 12:25
8d86d74 to
b546eee
Compare
LLM-jp-4.1 uses the GPT-OSS format, but its tokenizer decodes a space after every special token and parallel tool calls are separated by <|end|>. The GPT-OSS handler rejects this output, so add a dedicated handler, selected by the chat_format=llm-jp-harmony-v1 declaration in the chat template. Assisted-by: Claude Fable 5.1
e-mon
force-pushed
the
llmjp-harmony-handler
branch
from
September 29, 2026 12:28
b546eee to
0213322
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a dedicated chat-format handler for LLM-jp-4.1 (
llm-jp/llm-jp-4.1-*-thinking), a model that uses the GPT-OSS (Harmony) format with two dialect differences:<|channel|> analysis<|message|> .../<|start|> assistant to=functions.X<|channel|> commentary <|constrain|> json<|message|> {...}instead of the unspaced GPT-OSS form. With the GPT-OSS handler,llama-serverreturns500 "The model produced output that does not match the expected peg-native format"for tool requests and400 "Failed to initialize samplers"fortool_choice: required.<|end|>and the last one by<|call|>.The template declares the dialect on its first line (
{#- chat_format=llm-jp-harmony-v1 -#}), which is what the handler selection keys on, so no other template is affected.Design
common/parsers/llm-jp-harmony.cpp: a self-contained handler. Prompt handling (message delimiters, thinking tags, continuation, the<|return|>replacement) is that ofcommon_chat_params_init_gpt_oss; the PEG is the GPT-OSS one with the two differences built in. The gpt-oss-20b-specific rules (stray commentary prefix, unsolicited builtin calls) are not carried over: they never occur in this model's output (0 of 400 BFCL generations). The GPT-OSS handler is not modified; the only additions outside the new file are the declaration inparsers.h, the source list entry, the selection branch incommon/chat.cpp, the template, tests and docs.[ ]*;<|message|>by at most one space, so intentional leading whitespace in a message body survives (" padded"→" padded"). Plain spaces are used rather thanp.space()because the GBNFspacerule allows a single space, while<|constrain|> jsoncarries the template's space plus the tokenizer's.tool_choice (end start tool_choice)*inside the trigger rule whenparallel_tool_callsis set, so the lazy grammar covers every call; with the flag off the grammar constrains the model to one call, as for other formats.\s*after<|start|>/<|channel|>.common/chat.cppbefore the GPT-OSS check; template added tomodels/templates/; docs rows added.Tests
tests/test-chat.cpp: new LLM-jp-4.1 block (final channel, one-space body rule, reasoning + final, partial reasoning, tool call in role/channel header with<|constrain|> json, parallel calls with<|end|>, structured output)../build/bin/test-chat→[chat] All tests passed!(the peg_tester also validates that the generated grammar accepts each input).llm-jp/llm-jp-4.1-8b-thinking-gguf(Q4_K_M) onllama-server(temperature 0, template embedded in the GGUF, no--chat-template-file):reasoning_contentandcontentseparated; no leading space incontenttool_choice: auto)get_weather {"city": "Tokyo"}(was 500)parallel_tool_calls: true)tool_choice: requiredresponse_format: json_schemareasoning_contentdeltas, thencontentRun on Linux (Ubuntu 22.04 on WSL2, RTX 3090, CUDA 12.6 build of this commit):
test-chatpasses and all cases pass. The same cases also pass on macOS (Metal) with the pre-rebase build of the handler (identical parser code).Notes
models/templates/is the releasedllm-jp/llm-jp-4.1-8b-thinkingchat template (identical to the one embedded in the released GGUFs); thechat_format=llm-jp-harmony-v1declaration on line 1 is part of the official template (hf/ver4.0: declare the chat format in alpha_1.1's template llm-jp/llm-jp-tokenizer#32, benchmarks? ggml-org/llama.cpp#34)./completionoutput keeps the tokenizer's spaces.Differential check against the official parser on BFCL prompts
400 BFCL v3 prompts (simple / multiple / parallel / parallel_multiple, 100 each, drawn at random with a fixed seed from the 1,000 non-live AST prompts) were run through
llama-server(this branch, CUDA, temperature 0, 4096-token limit,-v+return_tokens: trueso that the chat response exposes the generated token ids). The very same token ids were then parsed by the official reference parser (HarmonyMessageParserfrom thellm-jp-vllmplugin) and compared with what the handler returned:Outputs covered single calls, 2–4 parallel calls separated by
<|end|>, text answers without a call, and 11 generations cut off by the token limit. With the officialbfcl-evalAST checker the same outputs score 83.0% (86 / 85 / 86 / 75 per category; 95% CI 79.0–86.4%); every miss is a model-side argument or count error, a text answer, or a truncation, none is a parsing error.Precedents
Self-contained handlers added for one model family without touching the others: ggml-org#19931 (GigaChat v3/3.1, by the model vendor), ggml-org#24615 (Cohere2MoE), ggml-org#26210 (MiniMax M3), ggml-org#21418 (Gemma 4); the per-model file layout comes from ggml-org#27764, and ggml-org#20393 is the PEG rework of the GPT-OSS handler this one mirrors.