Repository navigation
feat: llama.cpp b11320 -> b11457, Kolibri-1 support (patch 0016) - #478
Merged
Merged
Conversation
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…temone) b11361 (#29818) adds the /v1/systemone decision-model API. 0007 breaks on exactly that commit (one new route-table line) and is refreshed; server-decision.cpp joins the jllama and jllama_test sources. LlamaModel.handleSystemOne forwards to upstream's post_systemone handler through a server_routes the context now holds for its lifetime. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11368 (#27694) adds --spec-draft-sampling {greedy,probabilistic}; exposed
as ModelParameters.setDraftSampling(DraftSampling). All patches apply.
Also fixes the README badge label, stale at b11320 since b11327.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
llama.cpp #29792 (b11327..b11337) made CCCL configurable (GGML_CUDA_CCCL_VERSION) and pins v3.4.3 in upstream's CUDA release jobs: CUB DeviceTopK needs >= 3.4.3 and falls back to a sort below it, and CUDA 13.4 bundles an older 3.4. Both CUDA builds pass the same flag now; drop it with CUDA 13.5+. The runbook lists the vendor settings that follow upstream's CI. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11374 (#29852) moves ggml-openvino and upstream's release jobs to OpenVINO 2026.4.1; both OpenVINO build jobs follow. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Upstream added a Windows arm64 Vulkan build to its release set at b11395 (#29954). New build job build-windows-arm64-vulkan (windows-11-arm, clang-cl, the LunarG installer with the arm64 component as upstream installs it), its natives.csv row, the generated pom execution, the README row; package waits for it, and the all-windows-aarch64 fat jar picks it up from natives.csv. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11401 (#29895) splits router-child commands from logs and makes log colours self-contained. 0007 and 0014 break on exactly that commit (context only: a new llama_server overload and server_child local; a new colors member and get_colors()); both refreshed, +/- lines unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11411 (#28498) bumps LLAMA_STATE_SEQ_VERSION 3 -> 4: slot state files written by an earlier release no longer restore. Documented in saveSlot's Javadoc and the CHANGELOG. All patches apply. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
#29987 (4d60b4d08, right before b11429) reports input/output modalities in GET /models and breaks 0008 on exactly that commit (moved update_args, renamed update_caps in the trailing context); refreshed, + lines unchanged. RouterModel and ModelMeta gain getInputModalities(), getOutputModalities() and isDecisionModel(); getModelMeta() carries the modalities, built with upstream's helper. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…ENSOR b11450 (#26610) adds -sm tensor over RPC and bumps the RPC protocol to 8. 0015 breaks on exactly that commit (context only) and is refreshed, +/- lines unchanged. GpuSplitMode gains TENSOR. The new server-to-server comm (0.0.0.0 listener, unstoppable accept) is on file in TODO.md. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
The header is maintained by hand and is what gives LlamaModel's JNI entry points C linkage; without the declaration handleSystemOne would be exported under its C++-mangled name and every call would throw UnsatisfiedLinkError. CLAUDE.md no longer says mvn compile generates it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
SpotBugs OCP_OVERLY_CONCRETE_COLLECTION_PARAMETER on the new constructor (it only copies the values); spotless:apply on the three test classes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Upstream llama.cpp does not support the kolibri1 architecture yet (ggml-org/llama.cpp#29922). Patch 0016 carries it until it does and is dropped then. It is derived from the two community ports and improves on both: they write incompatible GGUFs (Eliasfpv28: gating 2 + pre-tokenizer qwen2; the Qwen3-MoE port behind Hob-forge's GGUFs, mjwsolo/localcode#101: gating 5 + pre-tokenizer kolibri1), and this patch loads both dialects. The router selects top-k on logits + expert_bias and weights by the unbiased sigmoid(logits), built outside build_moe_ffn so upstream's shared MoE code stays untouched. The patch header lists every known implementation with its license; REUSE annotates it MIT AND Apache-2.0. test_kolibri1.cpp (5 tests) writes tiny random GGUFs in both dialects and compares every logit, batched and token by token through the iSWA cache, with an independent double-precision reference written from Aleph Alpha's vLLM implementation. Verified red against a DeepSeek-style router and against RoPE on the full-attention layers. 598/598 C++ tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
The docs said test_kolibri1.cpp "must keep passing" once upstream supports Kolibri-1. That promised too much: the numerical comparison against the reference must stay green, but the GGUF format each test writes (gating function, pre-tokenizer, optional keys, the rejection of gating 1) is upstream's converter's decision. A red format row is a real signal too -- that dialect's published GGUFs stop loading -- and is to be decided, then moved to upstream's format with the reference untouched. Said so in CLAUDE.md, TODO.md and next to the tests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
bernardladenthin
had a problem deploying
to
startgate
October 6, 2026 23:45 — with
GitHub Actions
Error
bernardladenthin
had a problem deploying
to
maven-central
October 6, 2026 23:45 — with
GitHub Actions
Failure
bernardladenthin
had a problem deploying
to
maven-central
October 6, 2026 23:45 — with
GitHub Actions
Failure
|
Review of PR 478: llama.cpp b11320 to b11457 upgrade with Kolibri-1 support STRENGTHS:
CODE QUALITY:
MINOR OBSERVATIONS:
VERDICT: Well-structured PR with comprehensive documentation and testing. Ready for merge pending CI verification. |
|
6 of 7 tasks
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.



Summary
0007at #29818 (b11361) and #29895 (b11401)0014at #298950008at #29987 (just before b11429)0015at #26610 (b11450)docs/history/llama-cpp-breaking-changes.md.LlamaModel.handleSystemOne(/v1/systemone, decision models; also served byNativeServer)ModelParameters.setDraftSampling(--spec-draft-sampling)GpuSplitMode.TENSORRouterModelandModelMetavulkan-windows-aarch64natives jar (27 natives jars in total)CHANGELOG.md.kolibri1) ahead of upstream (Feature Request: Aleph-Alpha/Kolibri-1 ggml-org/llama.cpp#29922), as the temporary patch0016-model-kolibri1.patch. It is dropped once upstream registers the architecture.qwen2kolibri1logits + biasand weights by the unbiasedsigmoid(logits), as in Aleph Alpha's reference. The selection is built outsidebuild_moe_ffn, so upstream's shared MoE code is untouched.REUSE.tomlannotates the fileMIT AND Apache-2.0.Test plan
ctestlocally (fresh configure with all 16 patches applied at b11457).test_kolibri1.cpp(5 tests) checks Kolibri-1 numerically:NativeLibraryLoadSmokeTest,LlamaLoggerTest,NativeServerSmokeTest,LoggingSmokeTest) against the new library.reuse lint, thebuildcheckunit tests andcheck-natives.py(27 natives jars, 0 disagreements).README.md,CLAUDE.md,TODO.md,CHANGELOG.md).Not verified: the real 78B Kolibri-1 model. HuggingFace is unreachable from the session that wrote this, and the smallest GGUF is 28.6 GB.
Related issues / PRs
Refs ggml-org/llama.cpp#29922, mjwsolo/localcode#101
Checklist
CONTRIBUTING.mdandCODE_OF_CONDUCT.md🤖 Generated with Claude Code
https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Generated by Claude Code