Skip to content

llama.cpp bump to 2026-10-03, Kolibri 1 (experimental), launcher flag fix - #101

Merged
mjwsolo merged 3 commits into
mainfrom
feat/kolibri
Oct 4, 2026
Merged

mjwsolo merged 3 commits into
mainfrom
feat/kolibri

Conversation

@mjwsolo

@mjwsolo mjwsolo commented Oct 4, 2026

Copy link
Copy Markdown
Owner

What

  • llama.cpp bump. The vendored fork moves from the 2026-08-22 pin to upstream 836d571 (2026-10-03). All five fork patches are rebased and a sixth adds the kolibri1 architecture. Closes upstream-bump: patch 0001-turboquant.patch no longer applies #71.
  • Kolibri 1 (Aleph Alpha), experimental. New catalog entry kolibri: 78B total, about 3.5B active, 47.5 GB Q4_K_M community quant, pinned to an exact revision and sha256. Never auto-recommended.
  • Launcher fix. Upstream replaced --mmap with --load-mode mmap; the old flag is a hard argument error on the new server, so every launch would have failed.

The rebase

Upstream split the monolithic Metal shader into one library per kernels/<name>.metal. The TurboQuant kernels are redistributed into that layout (shared tables in kernels/common.h, three new kernel files registered in CMake and the GGML_METAL_LIBS X-macro). Details, the 3-way rebase recipe, and the accepted gaps are in llama-cpp-turboquant/PATCHES.md.

kolibri1 is a community patch (upstream issue ggml-org/llama.cpp#29922). It was read in full before adoption. Delete patch 0006 when upstream ships the architecture.

Bump script

  • -DLLAMA_OPENSSL=OFF and the deployment target, matching turboquant.yml. Without it the script built a binary linked against Homebrew libssl; its own gate rejected it.
  • Asserts the three turbo kernel files exist and are registered twice.
  • Fails if the binary embeds a developer home path.
  • The sync honours the documented vendoring exclusions, keeps PATCHES.md, and installs the verified binary.

Verification

Check Result
Official bump script, pristine clone: 6 patches replay, static build, self-contained pass
Catalog gate on the committed binary (load, generate, tool call) 11 / 11
Five existing models launched with the launcher's own generated command 5 / 5
KV quantization drift vs f16, old binary vs new (q8_0, turbo4, turbo3, mixed) unchanged within noise
Kolibri: generation, German, tool loop, reasoning on/off pass
Kolibri: exact recall at 6K / 24K / 61.8K prompt tokens, prefix cache reuse pass
Kolibri with turbo4 KV pass
Unit suite, including the new emitted-flags test pass
Binary links only system libraries, no developer paths pass

Known gaps

  • turbo2 as the K type is broken on both the old and the new binary (pre-existing).
  • CUDA turbo flash-attention takes upstream's F16 fallback. Not built or tested; only Metal ships.
  • The Metal TQ3_1S / TQ4_1S rotated fast path is removed; those weight types use the generic kernels. No catalog model uses them.
  • Kolibri coding quality is not benchmarked.
  • The model gate starts the server with its own short flag list. That is why it missed the --mmap removal. The new unit test covers the flags; the gate itself should move to the launcher's generated command.

No version bump in this PR; the changelog entry sits under Unreleased.

The vendored fork moves from the 2026-08-22 pin to upstream 836d571 and gains
the kolibri1 architecture (Aleph Alpha Kolibri-1) as patch 0006.

Rebase
- All five fork patches rebased with a 3-way merge instead of git apply.
- Upstream split the monolithic Metal shader into one library per
  kernels/<name>.metal. The TurboQuant kernels are redistributed into that
  layout: shared tables in kernels/common.h, quantizers and dequantizers in
  the shared headers, and three new kernel files (turbo, fa_turbo,
  fa_vec_turbo) registered in CMake and the GGML_METAL_LIBS X-macro.
- Flash-attention pipelines keep upstream's single-type names for symmetric
  K/V; only mixed turbo pairs encode both types.
- 0002: the training-context clamp moved into n_ctx_slot(); removed there.
- 0005: the Gemma 4 preserved-token fix follows the template code into
  common/parsers/gemma4.cpp; two API drifts fixed.
- 0006: kolibri1 (sigmoid-logit-add expert routing, one shared expert,
  sliding-window / full attention interleave, sandwich norms). Community
  patch; upstream tracking issue ggml-org/llama.cpp#29922.

Bump script
- -DLLAMA_OPENSSL=OFF and the deployment target, matching turboquant.yml.
  Without it the self-containment gate rejected a binary linked against
  Homebrew libssl.
- Asserts the three turbo kernel files exist and are registered twice.
- Fails if the binary embeds a developer home path.
- The sync honours the documented vendoring exclusions and no longer deletes
  PATCHES.md.

Verification
- Catalog gate on the committed binary: 11/11 (load, generate, tool call),
  including the turbo4 KV path, Muse Glimmer, North Mini Code, DiffusionGemma.
- KV quantization drift vs f16 is unchanged within noise against the previous
  binary (table in llama-cpp-turboquant/PATCHES.md).
- Binary links only system libraries and embeds no developer paths.

Known gaps are listed in PATCHES.md: the Metal TQ-weight rotated fast path is
removed, CUDA turbo flash-attention takes the F16 fallback, and turbo2 as the
K type is broken on both the old and new binaries.
Catalog
- New entry `kolibri` (Kolibri-1 Q4_K_M, 47.5 GB), pinned to an exact
  revision and sha256 because it is a community quant. Architecture
  `kolibri1`, text-only, default reasoning policy, never auto-recommended.
- Model group for browsed quants; Qwen-style family adapter (ChatML,
  <think>, Hermes tool calls).

Launcher
- llama.cpp replaced --mmap / --no-mmap / --mlock with --load-mode. The old
  flag is a hard argument error on the new bundled server, so every launch
  would have failed. The model gate did not catch it because it starts the
  server with its own short flag list.
- New test asks the bundled binary's --help whether it accepts every flag the
  launcher emits, for each runtime mode.

Verified on the bundled server with the launcher's own generated command:
Kolibri (generation, tool loop, reasoning on/off, exact recall at 6K / 24K /
61.8K prompt tokens, prefix cache reuse, q8_0 and turbo4 KV) and five existing
catalog models (generation + tool call, vision sidecars loading).
@mjwsolo
mjwsolo merged commit 4a0ce17 into main Oct 4, 2026
12 checks passed
bernardladenthin pushed a commit to bernardladenthin/java-llama.cpp that referenced this pull request Oct 6, 2026
Upstream llama.cpp does not support the kolibri1 architecture yet
(ggml-org/llama.cpp#29922). Patch 0016 carries it until it does and is
dropped then.

It is derived from the two community ports and improves on both: they
write incompatible GGUFs (Eliasfpv28: gating 2 + pre-tokenizer qwen2;
the Qwen3-MoE port behind Hob-forge's GGUFs, mjwsolo/localcode#101:
gating 5 + pre-tokenizer kolibri1), and this patch loads both dialects.
The router selects top-k on logits + expert_bias and weights by the
unbiased sigmoid(logits), built outside build_moe_ffn so upstream's
shared MoE code stays untouched. The patch header lists every known
implementation with its license; REUSE annotates it MIT AND Apache-2.0.

test_kolibri1.cpp (5 tests) writes tiny random GGUFs in both dialects
and compares every logit, batched and token by token through the iSWA
cache, with an independent double-precision reference written from
Aleph Alpha's vLLM implementation. Verified red against a DeepSeek-style
router and against RoPE on the full-attention layers. 598/598 C++ tests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AytmJF9faEiQEVt6eetQS2
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

upstream-bump: patch 0001-turboquant.patch no longer applies

1 participant