Skip to content

feat: llama.cpp b11320 -> b11457, Kolibri-1 support (patch 0016) - #478

Merged
bernardladenthin merged 28 commits into
mainfrom
claude/hopeful-pascal-9jlbqb
Oct 6, 2026
Merged

bernardladenthin merged 28 commits into
mainfrom
claude/hopeful-pascal-9jlbqb

Conversation

@bernardladenthin

Copy link
Copy Markdown
Owner

Summary

  • Upgrade the pinned llama.cpp from b11320 to b11457 in 21 small steps, each ending at a tag.
    • Every carried patch that broke is traced to the one upstream commit that broke it:
      • 0007 at #29818 (b11361) and #29895 (b11401)
      • 0014 at #29895
      • 0008 at #29987 (just before b11429)
      • 0015 at #26610 (b11450)
    • Each refresh moved context only. All nine patches are still needed.
    • Per-step record: docs/history/llama-cpp-breaking-changes.md.
  • New upstream features exposed:
    • LlamaModel.handleSystemOne (/v1/systemone, decision models; also served by NativeServer)
    • ModelParameters.setDraftSampling (--spec-draft-sampling)
    • GpuSplitMode.TENSOR
    • modalities on RouterModel and ModelMeta
    • a new vulkan-windows-aarch64 natives jar (27 natives jars in total)
  • Build changes: CUDA builds pin CCCL v3.4.3 as upstream's release jobs do, and OpenVINO moves to 2026.4.1.
  • Breaking:
    • RPC protocol 7 → 8: RPC peers must be upgraded together.
    • Slot state files written by earlier releases no longer restore (#28498).
    • Details in CHANGELOG.md.
  • Aleph Alpha Kolibri-1 (kolibri1) ahead of upstream (Feature Request: Aleph-Alpha/Kolibri-1 ggml-org/llama.cpp#29922), as the temporary patch 0016-model-kolibri1.patch. It is dropped once upstream registers the architecture.
    • It is derived from the two community ports, whose GGUFs are mutually incompatible:
    • Unlike either port, this patch loads both.
    • The router selects on logits + bias and weights by the unbiased sigmoid(logits), as in Aleph Alpha's reference. The selection is built outside build_moe_ffn, so upstream's shared MoE code is untouched.
    • The patch header lists every known implementation with its license. REUSE.toml annotates the file MIT AND Apache-2.0.

Test plan

  • C++: 598/598 ctest locally (fresh configure with all 16 patches applied at b11457).
  • test_kolibri1.cpp (5 tests) checks Kolibri-1 numerically:
    • It writes tiny random GGUFs in both dialects.
    • It compares every logit, batched and token by token through the iSWA cache, against an independent double-precision reference written from Aleph Alpha's vLLM code.
    • It goes red with a DeepSeek-style router or with RoPE on the full-attention layers (both mutations tried and reverted).
  • Java: model-free smokes (NativeLibraryLoadSmokeTest, LlamaLoggerTest, NativeServerSmokeTest, LoggingSmokeTest) against the new library.
  • clang-format 23.1.3, reuse lint, the buildcheck unit tests and check-natives.py (27 natives jars, 0 disagreements).
  • CI is green on this branch. GPU backends and model-backed Java tests are covered only by CI.
  • Docs / CHANGELOG updated (README.md, CLAUDE.md, TODO.md, CHANGELOG.md).

Not verified: the real 78B Kolibri-1 model. HuggingFace is unreachable from the session that wrote this, and the smallest GGUF is 28.6 GB.

Related issues / PRs

Refs ggml-org/llama.cpp#29922, mjwsolo/localcode#101

Checklist

  • I have read CONTRIBUTING.md and CODE_OF_CONDUCT.md
  • My commits follow Conventional Commits: they use this repo's "Upgrade llama.cpp from … to …" style instead.
  • No security-sensitive changes

🤖 Generated with Claude Code

https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2


Generated by Claude Code

claude added 28 commits October 6, 2026 22:38
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…temone)

b11361 (#29818) adds the /v1/systemone decision-model API. 0007 breaks on
exactly that commit (one new route-table line) and is refreshed;
server-decision.cpp joins the jllama and jllama_test sources.

LlamaModel.handleSystemOne forwards to upstream's post_systemone handler
through a server_routes the context now holds for its lifetime.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11368 (#27694) adds --spec-draft-sampling {greedy,probabilistic}; exposed
as ModelParameters.setDraftSampling(DraftSampling). All patches apply.
Also fixes the README badge label, stale at b11320 since b11327.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
llama.cpp #29792 (b11327..b11337) made CCCL configurable
(GGML_CUDA_CCCL_VERSION) and pins v3.4.3 in upstream's CUDA release jobs:
CUB DeviceTopK needs >= 3.4.3 and falls back to a sort below it, and
CUDA 13.4 bundles an older 3.4. Both CUDA builds pass the same flag now;
drop it with CUDA 13.5+. The runbook lists the vendor settings that
follow upstream's CI.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11374 (#29852) moves ggml-openvino and upstream's release jobs to
OpenVINO 2026.4.1; both OpenVINO build jobs follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Upstream added a Windows arm64 Vulkan build to its release set at
b11395 (#29954). New build job build-windows-arm64-vulkan (windows-11-arm,
clang-cl, the LunarG installer with the arm64 component as upstream
installs it), its natives.csv row, the generated pom execution, the README
row; package waits for it, and the all-windows-aarch64 fat jar picks it
up from natives.csv.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11401 (#29895) splits router-child commands from logs and makes log
colours self-contained. 0007 and 0014 break on exactly that commit
(context only: a new llama_server overload and server_child local; a new
colors member and get_colors()); both refreshed, +/- lines unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
b11411 (#28498) bumps LLAMA_STATE_SEQ_VERSION 3 -> 4: slot state files
written by an earlier release no longer restore. Documented in saveSlot's
Javadoc and the CHANGELOG. All patches apply.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
#29987 (4d60b4d08, right before b11429) reports input/output modalities
in GET /models and breaks 0008 on exactly that commit (moved
update_args, renamed update_caps in the trailing context); refreshed,
+ lines unchanged.

RouterModel and ModelMeta gain getInputModalities(),
getOutputModalities() and isDecisionModel(); getModelMeta() carries the
modalities, built with upstream's helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
…ENSOR

b11450 (#26610) adds -sm tensor over RPC and bumps the RPC protocol to 8.
0015 breaks on exactly that commit (context only) and is refreshed,
+/- lines unchanged. GpuSplitMode gains TENSOR. The new server-to-server
comm (0.0.0.0 listener, unstoppable accept) is on file in TODO.md.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
The header is maintained by hand and is what gives LlamaModel's JNI
entry points C linkage; without the declaration handleSystemOne would be
exported under its C++-mangled name and every call would throw
UnsatisfiedLinkError. CLAUDE.md no longer says mvn compile generates it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
SpotBugs OCP_OVERLY_CONCRETE_COLLECTION_PARAMETER on the new constructor
(it only copies the values); spotless:apply on the three test classes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Upstream llama.cpp does not support the kolibri1 architecture yet
(ggml-org/llama.cpp#29922). Patch 0016 carries it until it does and is
dropped then.

It is derived from the two community ports and improves on both: they
write incompatible GGUFs (Eliasfpv28: gating 2 + pre-tokenizer qwen2;
the Qwen3-MoE port behind Hob-forge's GGUFs, mjwsolo/localcode#101:
gating 5 + pre-tokenizer kolibri1), and this patch loads both dialects.
The router selects top-k on logits + expert_bias and weights by the
unbiased sigmoid(logits), built outside build_moe_ffn so upstream's
shared MoE code stays untouched. The patch header lists every known
implementation with its license; REUSE annotates it MIT AND Apache-2.0.

test_kolibri1.cpp (5 tests) writes tiny random GGUFs in both dialects
and compares every logit, batched and token by token through the iSWA
cache, with an independent double-precision reference written from
Aleph Alpha's vLLM implementation. Verified red against a DeepSeek-style
router and against RoPE on the full-attention layers. 598/598 C++ tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
The docs said test_kolibri1.cpp "must keep passing" once upstream
supports Kolibri-1. That promised too much: the numerical comparison
against the reference must stay green, but the GGUF format each test
writes (gating function, pre-tokenizer, optional keys, the rejection of
gating 1) is upstream's converter's decision. A red format row is a real
signal too -- that dialect's published GGUFs stop loading -- and is to be
decided, then moved to upstream's format with the reference untouched.
Said so in CLAUDE.md, TODO.md and next to the tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
@bernardladenthin
bernardladenthin merged commit d775c4c into main Oct 6, 2026
11 of 15 checks passed
@bernardladenthin
bernardladenthin deleted the claude/hopeful-pascal-9jlbqb branch October 6, 2026 23:46
@claude

claude Bot commented Oct 6, 2026

Copy link
Copy Markdown

Review of PR 478: llama.cpp b11320 to b11457 upgrade with Kolibri-1 support

STRENGTHS:

  • Kolibri-1 patch (0016) is well-designed with comprehensive reference validation
  • Router implementation correctly differentiates Kolibri-1 vs DeepSeek-V3 style routing
  • All patch refreshes are context-only with proper upstream change handling
  • CCCL v3.4.3 pinning follows upstream approach and is well-documented
  • Test coverage includes both batch and token-by-token decoding through iSWA KV cache
  • Two GGUF dialects (AFMoE vs Qwen3-MoE) handled with appropriate fallbacks

CODE QUALITY:

  • Router selection correctly uses (logits + bias) for indices while passing original logits to build_moe_ffn
  • Sliding-window RoPE application properly gated to SWA-only layers
  • CLAUDE.md and CHANGELOG.md thoroughly updated with breaking change notices

MINOR OBSERVATIONS:

  • Windows arm64 Vulkan job is build-only (no GPU testing) which matches documented policy
  • Consider adding clarification on GPT-2 byte tokenization in test_kolibri1.cpp for future maintainers

VERDICT: Well-structured PR with comprehensive documentation and testing. Ready for merge pending CI verification.

@sonarqubecloud

sonarqubecloud Bot commented Oct 6, 2026

Copy link
Copy Markdown

This branch had an error being deployed

1 failed deployment
startgate — 66a4e2fe Deployed Oct 6, 2026 by bernardladenthin via Start gate (abort window) #1094
maven-central — 66a4e2fe Deployed Oct 6, 2026 by bernardladenthin via Verify GPG signing key (no secrets printed) #1094
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants