Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
The table of contents is too big for display.
Diff view
Diff view
  •  
  •  
  •  
26 changes: 26 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,6 +3,32 @@
All notable changes to LocalCode will be documented here. The format follows
[Keep a Changelog](https://keepachangelog.com/en/1.1.0/).

## Unreleased

### Added

- **Kolibri 1 (Aleph Alpha), experimental.** The open-weight English/German model
released 2026-10-03: 78B parameters in total, about 3.5B active per token. It
appears in the model list as `kolibri` and is never picked automatically. The
download is 47.5 GB and needs 64 GB of unified memory at minimum. Its
architecture is not in upstream llama.cpp yet, so the bundled server carries a
community patch for it, and the file is a community quant pinned to an exact
revision and digest. Checked on the bundled server: load, generation, tool
calls, reasoning on and off, and exact recall from a 60,000-token prompt.
Coding quality has not been benchmarked.

### Changed

- **The bundled model server moves to llama.cpp as of 2026-10-03** (from
2026-08-22). Every catalogue model was re-verified on the new binary.

### Fixed

- **Model launch flag removed upstream.** llama.cpp replaced `--mmap` with
`--load-mode mmap`; the old spelling is now a hard error, so the launcher uses
the new one. A test now asks the bundled server whether it accepts every flag
the launcher emits, so a renamed flag fails in tests instead of at launch.

## 0.5.1 — 2026-10-03

### Security
Expand Down
28 changes: 28 additions & 0 deletions docs/src/data/models.json
Original file line number Diff line number Diff line change
Expand Up @@ -682,6 +682,20 @@
"maker": "Alibaba",
"logo": "qwen",
"experimental": false
},
{
"name": "Kolibri 1 78B-A3B (Q4, experimental)",
"base": "Kolibri 1 78B-A3B",
"quant": "Q4, experimental",
"total": "78B",
"active": "3.5B active (78B total MoE)",
"kind": "MoE",
"detail": "78B total MoE",
"file": "Kolibri-1-Q4_K_M.gguf",
"size": 47.5,
"maker": "Aleph Alpha",
"logo": null,
"experimental": true
}
]
},
Expand Down Expand Up @@ -845,6 +859,20 @@
"logo": "qwen",
"experimental": false
},
{
"name": "Kolibri 1 78B-A3B (Q4, experimental)",
"base": "Kolibri 1 78B-A3B",
"quant": "Q4, experimental",
"total": "78B",
"active": "3.5B active (78B total MoE)",
"kind": "MoE",
"detail": "78B total MoE",
"file": "Kolibri-1-Q4_K_M.gguf",
"size": 47.5,
"maker": "Aleph Alpha",
"logo": null,
"experimental": true
},
{
"name": "Qwen 3.8 27B (BF16, full)",
"base": "Qwen 3.8 27B",
Expand Down
1 change: 1 addition & 0 deletions docs/upstream-fork.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,7 @@ drift and a broken model.
| TQ3_1S / TQ4_1S | fork-local weight quant types | never |
| Fused MoE router | performance | upstream fuses it |
| Muse thinking-tag fix | 3 lines, from **still-open upstream PR #27475** | **the day #27475 merges — delete it then** |
| Kolibri-1 (`kolibri1` arch) | Aleph Alpha's MoE: sigmoid-logit-add expert routing, one shared expert, sliding-window / full attention interleave, sandwich norms. Community patch, upstream issue **#29922** | **the day upstream ships `kolibri1` - delete `0006` then** |

`llama-cpp-turboquant/PATCHES.md` is the authoritative, human-readable
inventory. This table is orientation, not the contract.
Expand Down
6 changes: 3 additions & 3 deletions llama-cpp-turboquant/.devops/intel.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
ARG ONEAPI_VERSION=2025.3.3-0-devel-ubuntu24.04
ARG ONEAPI_VERSION=2026.1.1-devel-ubuntu24.04
ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
ARG APP_REVISION=N/A
Expand All @@ -19,7 +19,7 @@ RUN npm ci
COPY tools/ui/ ./
RUN LLAMA_BUILD_NUMBER="$APP_VERSION" npm run build

FROM docker.io/intel/deep-learning-essentials:$ONEAPI_VERSION AS build
FROM docker.io/intel/oneapi-toolkit:$ONEAPI_VERSION AS build

ARG GGML_SYCL_F16=ON
ARG LEVEL_ZERO_VERSION=1.28.2
Expand Down Expand Up @@ -59,7 +59,7 @@ RUN mkdir -p /app/full \
&& cp requirements.txt /app/full \
&& cp .devops/tools.sh /app/full/tools.sh

FROM docker.io/intel/deep-learning-essentials:$ONEAPI_VERSION AS base
FROM docker.io/intel/oneapi-toolkit:$ONEAPI_VERSION AS base

ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
Expand Down
15 changes: 10 additions & 5 deletions llama-cpp-turboquant/.devops/musa.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,10 +1,9 @@
ARG UBUNTU_VERSION=22.04
# This needs to generally match the container host's environment.
ARG MUSA_VERSION=rc4.3.0
# Target the MUSA build image
ARG BASE_MUSA_DEV_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-devel-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_DEV_CONTAINER=registry.mthreads.com/mcconline/musa_sdk:5.2.0-devel-ubuntu${UBUNTU_VERSION}-s5000

ARG BASE_MUSA_RUN_CONTAINER=docker.io/mthreads/musa:${MUSA_VERSION}-runtime-ubuntu${UBUNTU_VERSION}-amd64
ARG BASE_MUSA_RUN_CONTAINER=registry.mthreads.com/mcconline/musa_sdk:5.2.0-runtime-ubuntu${UBUNTU_VERSION}-s5000

ARG BUILD_DATE=N/A
ARG APP_VERSION=N/A
Expand Down Expand Up @@ -37,7 +36,10 @@ RUN apt-get update && \
python3-pip \
git \
libssl-dev \
libgomp1
libgomp1 \
musa-mualg-5-2 \
musa-muthrust-5-2 \
libmthreads-compute

WORKDIR /app

Expand Down Expand Up @@ -80,13 +82,16 @@ LABEL org.opencontainers.image.created=$BUILD_DATE \
org.opencontainers.image.source=$IMAGE_SOURCE

RUN apt-get update \
&& apt-get install -y libgomp1 curl ffmpeg \
&& apt-get install -y libgomp1 curl ffmpeg libmthreads-compute \
&& apt autoremove -y \
&& apt clean -y \
&& rm -rf /tmp/* /var/tmp/* \
&& find /var/cache/apt/archives /var/lib/apt/lists -not -name lock -type f -delete \
&& find /var/cache -type f -delete

# The MUSA runtime image does not register its library directory
RUN echo "/usr/local/musa/lib" > /etc/ld.so.conf.d/musa-runtime.conf && ldconfig

COPY --from=build /app/lib/ /app

### Full
Expand Down
12 changes: 6 additions & 6 deletions llama-cpp-turboquant/.devops/nix/package.nix
Original file line number Diff line number Diff line change
Expand Up @@ -31,7 +31,7 @@
]
&& blas.meta.available,
useCuda ? config.cudaSupport,
useMetalKit ? stdenv.isAarch64 && stdenv.isDarwin,
useMetalKit ? stdenv.hostPlatform.isAarch64 && stdenv.hostPlatform.isDarwin,
# Increases the runtime closure size by ~700M
useMpi ? false,
useRocm ? config.rocmSupport,
Expand Down Expand Up @@ -92,7 +92,7 @@ let

cudaBuildInputs = with cudaPackages; [
cuda_cudart
cuda_cccl # <nv/target>
cccl # <nv/target>
libcublas
];

Expand Down Expand Up @@ -166,7 +166,7 @@ effectiveStdenv.mkDerivation (finalAttrs: {
# `xcrun` is used find the path of the Metal compiler, which is varible
# and not on $PATH
# see https://github.com/ggml-org/llama.cpp/pull/6118 for discussion
__noChroot = effectiveStdenv.isDarwin && useMetalKit && precompileMetalShaders;
__noChroot = effectiveStdenv.hostPlatform.isDarwin && useMetalKit && precompileMetalShaders;

nativeBuildInputs =
[
Expand All @@ -181,10 +181,10 @@ effectiveStdenv.mkDerivation (finalAttrs: {
autoAddDriverRunpath
]
++ optionals (effectiveStdenv.hostPlatform.isGnu && enableStatic) [ glibc.static ]
++ optionals (effectiveStdenv.isDarwin && useMetalKit && precompileMetalShaders) [ xcrunHost ];
++ optionals (effectiveStdenv.hostPlatform.isDarwin && useMetalKit && precompileMetalShaders) [ xcrunHost ];

buildInputs =
optionals effectiveStdenv.isDarwin darwinBuildInputs
optionals effectiveStdenv.hostPlatform.isDarwin darwinBuildInputs
++ optionals useCuda cudaBuildInputs
++ optionals useMpi [ mpi ]
++ optionals useRocm rocmBuildInputs
Expand Down Expand Up @@ -245,7 +245,7 @@ effectiveStdenv.mkDerivation (finalAttrs: {

# Configurations that are known to result in build failures. Can be
# overridden by importing Nixpkgs with `allowBroken = true`.
broken = (useMetalKit && !effectiveStdenv.isDarwin);
broken = (useMetalKit && !effectiveStdenv.hostPlatform.isDarwin);

description = "Inference of LLaMA model in pure C/C++${descriptionSuffix}";
homepage = "https://github.com/ggml-org/llama.cpp/";
Expand Down
23 changes: 13 additions & 10 deletions llama-cpp-turboquant/.devops/openvino.Dockerfile
Original file line number Diff line number Diff line change
@@ -1,18 +1,18 @@
ARG OPENVINO_VERSION_MAJOR=2026.3
ARG OPENVINO_VERSION_FULL=2026.3.0.22451.bd8d6542e3c
ARG OPENVINO_VERSION_MAJOR=2026.4.1
ARG OPENVINO_VERSION_FULL=2026.4.1.22982.07f9c262b05
ARG UBUNTU_VERSION=24.04

# Intel GPU driver versions. https://github.com/intel/compute-runtime/releases
ARG IGC_VERSION=v2.38.2
ARG IGC_VERSION_FULL=2_2.38.2+22051
ARG COMPUTE_RUNTIME_VERSION=26.27.39122.11
ARG COMPUTE_RUNTIME_VERSION_FULL=26.27.39122.11-0
ARG IGC_VERSION=v2.41.5
ARG IGC_VERSION_FULL=2_2.41.5+22716
ARG COMPUTE_RUNTIME_VERSION=26.35.39758.10
ARG COMPUTE_RUNTIME_VERSION_FULL=26.35.39758.10-0
ARG IGDGMM_VERSION=22.10.0

# Intel NPU driver versions. https://github.com/intel/linux-npu-driver/releases
ARG NPU_DRIVER_VERSION=v1.35.0
ARG NPU_DRIVER_FULL=v1.35.0.20260722-29947505341
ARG LIBZE1_VERSION=1.28.2-1~24.04~ppa1
ARG NPU_DRIVER_VERSION=v1.38.0
ARG NPU_DRIVER_FULL=v1.38.0.20260910-34487311128
ARG LIBZE1_VERSION=1.32.0-1~24.04~ppa1

# Optional proxy build arguments
ARG http_proxy=
Expand Down Expand Up @@ -90,6 +90,9 @@ RUN bash -c "source ${OpenVINO_DIR}/setupvars.sh && \
cmake -B build/ReleaseOV -G Ninja \
-DCMAKE_BUILD_TYPE=Release \
-DLLAMA_BUILD_TESTS=OFF \
-DGGML_NATIVE=OFF \
-DGGML_BACKEND_DL=ON \
-DGGML_CPU_ALL_VARIANTS=ON \
-DGGML_OPENVINO=ON && \
cmake --build build/ReleaseOV --parallel "

Expand Down Expand Up @@ -170,7 +173,7 @@ RUN --mount=type=cache,target=/var/cache/intel-npu,sharing=locked \
fi; \
DEB=/var/cache/intel-npu/libze1_${LIBZE1_VERSION}_amd64.deb; \
if [ ! -f "$DEB" ]; then \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260606T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
wget -q -O "$DEB" https://snapshot.ppa.launchpadcontent.net/kobuk-team/intel-graphics/ubuntu/20260830T100000Z/pool/main/l/level-zero-loader/libze1_${LIBZE1_VERSION}_amd64.deb; \
fi; \
mkdir /tmp/npu/ && cd /tmp/npu/ && tar -xf "$TGZ" && cp "$DEB" .; \
apt-get update; \
Expand Down
2 changes: 1 addition & 1 deletion llama-cpp-turboquant/.ecrc
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"Exclude": ["^\\.gitmodules$", "stb_image\\.h"],
"Exclude": ["^\\.gitmodules$", "stb_image\\.h", "examples/test-cmake/build/", "examples/test-cmake/build-subdir/"],
"Disable": {
"IndentSize": true
}
Expand Down
5 changes: 4 additions & 1 deletion llama-cpp-turboquant/.pi/gg/SYSTEM.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,8 @@ General:
- PR and commit titles format: `<module> : <title>`. Lookup recents for examples
- Don't try to build or run the code unless you are explicitly asked to do so
- Use the `gh` CLI tool when querying PRs, issues, or other GitHub resources
- When [MODEL] is needed, first try to get it from the `PI_MODEL_NAME` env var before asking the user
- Never read the `AGENTS.md` file

Coding:
- When in doubt, always refer to the CONTRIBUTING.md file of the project
Expand All @@ -20,8 +22,9 @@ Pull requests (PRs):
- Don't explicitly wrap lines in the PR description (each paragraph and bullet is a single line)
- When creating a pull request, look for the repository's PR template and follow it
- For the AI usage disclosure section, write "YES. pi:llama.cpp/[MODEL]"
- Ask the user to tell you what model was used and write it in place of [MODEL]
- If `PI_MODEL_NAME` env var is not set, ask the user to tell you what model was used and write it in place of [MODEL]
- Always create the pull requests in draft mode
- Never reply to review comments or post comments on issues/PRs without explicit permission from the user

Commits:
- On every commit that you make, include a "Assisted-by: pi:llama.cpp/[MODEL]" tag
Expand Down
3 changes: 2 additions & 1 deletion llama-cpp-turboquant/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,8 @@ These points are extremely important - failing to follow them won't necessarily
Common mistakes that AI agents usually make:
- Write comments first then write code: this usually leads to extensive redundant comments. Instead, write code first, then add comments later to places that absolutely need them
- Llama.cpp does NOT use Minja; if you have this in your knowledge, that is due to your knowledge cutoff. Llama.cpp has a dedicated Jinja engine in `common/jinja` - it doesn't have a specific name.
- Do NOT add a new file in `tests/*` without maintainers' approval. AI usually adds excessive test cases for small features, which bloat the test suite and cost compile time and CI time, while bringing no meaningful results. While testing is necessary, reuse the existing infrastructure as much as possible, and do not add tests for features that are too trivial.

Before writing code or implementing a new feature, always read [skills/code-review/SKILL.md](skills/code-review/SKILL.md). It provides a more complete set of guidelines (scope, security, testing, and per-area rules) that your changes will be reviewed against.

### Prohibited Actions

Expand Down
Loading
Loading