Repository navigation
llama.cpp bump to 2026-10-03, Kolibri 1 (experimental), launcher flag fix - #101
Merged
Merged
Conversation
The vendored fork moves from the 2026-08-22 pin to upstream 836d571 and gains the kolibri1 architecture (Aleph Alpha Kolibri-1) as patch 0006. Rebase - All five fork patches rebased with a 3-way merge instead of git apply. - Upstream split the monolithic Metal shader into one library per kernels/<name>.metal. The TurboQuant kernels are redistributed into that layout: shared tables in kernels/common.h, quantizers and dequantizers in the shared headers, and three new kernel files (turbo, fa_turbo, fa_vec_turbo) registered in CMake and the GGML_METAL_LIBS X-macro. - Flash-attention pipelines keep upstream's single-type names for symmetric K/V; only mixed turbo pairs encode both types. - 0002: the training-context clamp moved into n_ctx_slot(); removed there. - 0005: the Gemma 4 preserved-token fix follows the template code into common/parsers/gemma4.cpp; two API drifts fixed. - 0006: kolibri1 (sigmoid-logit-add expert routing, one shared expert, sliding-window / full attention interleave, sandwich norms). Community patch; upstream tracking issue ggml-org/llama.cpp#29922. Bump script - -DLLAMA_OPENSSL=OFF and the deployment target, matching turboquant.yml. Without it the self-containment gate rejected a binary linked against Homebrew libssl. - Asserts the three turbo kernel files exist and are registered twice. - Fails if the binary embeds a developer home path. - The sync honours the documented vendoring exclusions and no longer deletes PATCHES.md. Verification - Catalog gate on the committed binary: 11/11 (load, generate, tool call), including the turbo4 KV path, Muse Glimmer, North Mini Code, DiffusionGemma. - KV quantization drift vs f16 is unchanged within noise against the previous binary (table in llama-cpp-turboquant/PATCHES.md). - Binary links only system libraries and embeds no developer paths. Known gaps are listed in PATCHES.md: the Metal TQ-weight rotated fast path is removed, CUDA turbo flash-attention takes the F16 fallback, and turbo2 as the K type is broken on both the old and new binaries.
Catalog - New entry `kolibri` (Kolibri-1 Q4_K_M, 47.5 GB), pinned to an exact revision and sha256 because it is a community quant. Architecture `kolibri1`, text-only, default reasoning policy, never auto-recommended. - Model group for browsed quants; Qwen-style family adapter (ChatML, <think>, Hermes tool calls). Launcher - llama.cpp replaced --mmap / --no-mmap / --mlock with --load-mode. The old flag is a hard argument error on the new bundled server, so every launch would have failed. The model gate did not catch it because it starts the server with its own short flag list. - New test asks the bundled binary's --help whether it accepts every flag the launcher emits, for each runtime mode. Verified on the bundled server with the launcher's own generated command: Kolibri (generation, tool loop, reasoning on/off, exact recall at 6K / 24K / 61.8K prompt tokens, prefix cache reuse, q8_0 and turbo4 KV) and five existing catalog models (generation + tool call, vision sidecars loading).
bernardladenthin
pushed a commit
to bernardladenthin/java-llama.cpp
that referenced
this pull request
Oct 6, 2026
Upstream llama.cpp does not support the kolibri1 architecture yet (ggml-org/llama.cpp#29922). Patch 0016 carries it until it does and is dropped then. It is derived from the two community ports and improves on both: they write incompatible GGUFs (Eliasfpv28: gating 2 + pre-tokenizer qwen2; the Qwen3-MoE port behind Hob-forge's GGUFs, mjwsolo/localcode#101: gating 5 + pre-tokenizer kolibri1), and this patch loads both dialects. The router selects top-k on logits + expert_bias and weights by the unbiased sigmoid(logits), built outside build_moe_ffn so upstream's shared MoE code stays untouched. The patch header lists every known implementation with its license; REUSE annotates it MIT AND Apache-2.0. test_kolibri1.cpp (5 tests) writes tiny random GGUFs in both dialects and compares every logit, batched and token by token through the iSWA cache, with an independent double-precision reference written from Aleph Alpha's vLLM implementation. Verified red against a DeepSeek-style router and against RoPE on the full-attention layers. 598/598 C++ tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
836d571(2026-10-03). All five fork patches are rebased and a sixth adds thekolibri1architecture. Closes upstream-bump: patch 0001-turboquant.patch no longer applies #71.kolibri: 78B total, about 3.5B active, 47.5 GB Q4_K_M community quant, pinned to an exact revision and sha256. Never auto-recommended.--mmapwith--load-mode mmap; the old flag is a hard argument error on the new server, so every launch would have failed.The rebase
Upstream split the monolithic Metal shader into one library per
kernels/<name>.metal. The TurboQuant kernels are redistributed into that layout (shared tables inkernels/common.h, three new kernel files registered in CMake and theGGML_METAL_LIBSX-macro). Details, the 3-way rebase recipe, and the accepted gaps are inllama-cpp-turboquant/PATCHES.md.kolibri1is a community patch (upstream issue ggml-org/llama.cpp#29922). It was read in full before adoption. Delete patch0006when upstream ships the architecture.Bump script
-DLLAMA_OPENSSL=OFFand the deployment target, matchingturboquant.yml. Without it the script built a binary linked against Homebrewlibssl; its own gate rejected it.PATCHES.md, and installs the verified binary.Verification
Known gaps
--mmapremoval. The new unit test covers the flags; the gate itself should move to the launcher's generated command.No version bump in this PR; the changelog entry sits under Unreleased.