Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
28 commits
Select commit Hold shift + click to select a range
bcdaadc
Upgrade llama.cpp from b11320 to b11327
claude Oct 6, 2026
26b7814
Upgrade llama.cpp from b11327 to b11337
claude Oct 6, 2026
10cb275
Upgrade llama.cpp from b11337 to b11355
claude Oct 6, 2026
c2c5372
Upgrade llama.cpp from b11355 to b11361; add handleSystemOne (/v1/sys…
claude Oct 6, 2026
daacfb2
Upgrade llama.cpp from b11361 to b11368; add setDraftSampling
claude Oct 6, 2026
ba9852e
CUDA builds: fetch CCCL v3.4.3 as upstream's release jobs do
claude Oct 6, 2026
80ff524
Upgrade llama.cpp from b11368 to b11371
claude Oct 6, 2026
57c6359
Upgrade llama.cpp from b11371 to b11374; OpenVINO 2026.4.1
claude Oct 6, 2026
a4789bb
Upgrade llama.cpp from b11374 to b11382
claude Oct 6, 2026
bc4f29f
Upgrade llama.cpp from b11382 to b11388
claude Oct 6, 2026
f79313a
Upgrade llama.cpp from b11388 to b11395
claude Oct 6, 2026
42d0aad
Add the vulkan-windows-aarch64 natives jar
claude Oct 6, 2026
5b993c3
Upgrade llama.cpp from b11395 to b11400
claude Oct 6, 2026
eb59a58
Upgrade llama.cpp from b11400 to b11401; refresh 0007 and 0014
claude Oct 6, 2026
64ce626
Upgrade llama.cpp from b11401 to b11414
claude Oct 6, 2026
76bb90e
Upgrade llama.cpp from b11414 to b11418
claude Oct 6, 2026
ce535f3
Upgrade llama.cpp from b11418 to b11429; refresh 0008; modalities
claude Oct 6, 2026
1860294
Upgrade llama.cpp from b11429 to b11440
claude Oct 6, 2026
e4d3ba6
Upgrade llama.cpp from b11440 to b11447
claude Oct 6, 2026
1bcbeda
Upgrade llama.cpp from b11447 to b11449
claude Oct 6, 2026
443a3d0
Upgrade llama.cpp from b11449 to b11450; refresh 0015; GpuSplitMode.T…
claude Oct 6, 2026
de6145c
Upgrade llama.cpp from b11450 to b11454
claude Oct 6, 2026
f165aa1
Upgrade llama.cpp from b11454 to b11457
claude Oct 6, 2026
4e6d938
Declare handleSystemOne in jllama.h
claude Oct 6, 2026
dc50178
RouterModel: take the modalities as Collection; format the new tests
claude Oct 6, 2026
2b1a4e0
CHANGELOG: the llama.cpp b11320 -> b11457 bump
claude Oct 6, 2026
c881fbf
Support Aleph Alpha Kolibri-1 via patch 0016 ahead of upstream
claude Oct 6, 2026
66a4e2f
Kolibri-1 docs: which tests must survive dropping patch 0016
claude Oct 6, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion .github/build_cuda_linux.sh
Original file line number Diff line number Diff line change
Expand Up @@ -37,4 +37,7 @@ case "${CUDA_FAST_BUILD:-}" in
;;
esac

exec .github/build.sh $@ -DGGML_CUDA=1 -DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.4/bin/nvcc $CUDA_ARCH_ARGS
# CCCL v3.4.3 is fetched instead of the one CUDA 13.4 bundles, as upstream's own CUDA release builds do
# (llama.cpp #29792): ggml's CUB DeviceTopK path needs >= 3.4.3 (an earlier race, NVIDIA/cccl#10627) and
# falls back to a sort below it. Drop the flag once the toolkit is 13.5+, which bundles CCCL 3.5.
exec .github/build.sh $@ -DGGML_CUDA=1 -DCMAKE_CUDA_COMPILER=/usr/local/cuda-13.4/bin/nvcc -DGGML_CUDA_CCCL_VERSION=v3.4.3 $CUDA_ARCH_ARGS
3 changes: 2 additions & 1 deletion .github/buildcheck/tests/test_natives.py
Original file line number Diff line number Diff line change
Expand Up @@ -225,7 +225,8 @@ def test_cli_prints_the_pom_executions_and_the_targets(self):
out = subprocess.run([sys.executable, cli, "fatjar-targets"], capture_output=True, text=True, check=True)
self.assertEqual(out.stdout.split(), ["linux-aarch64", "linux-x86-64", "windows-aarch64", "windows-x86-64"])
out = subprocess.run([sys.executable, cli, "pom"], capture_output=True, text=True, check=True)
self.assertEqual(out.stdout.count("<execution>"), 26)
rows = natives.rows(natives.read(REPO, ".github/natives.csv"))
self.assertEqual(out.stdout.count("<execution>"), len(rows))
self.assertEqual(subprocess.run([sys.executable, cli, "nonsense"], capture_output=True).returncode, 2)


Expand Down
1 change: 1 addition & 0 deletions .github/natives.csv
Original file line number Diff line number Diff line change
Expand Up @@ -29,6 +29,7 @@ cuda13-windows-x86-64,Windows/x86_64/cuda13,jllama.dll,no
vulkan-linux-x86-64,Linux/x86_64/vulkan,libjllama.so,no
vulkan-linux-aarch64,Linux/aarch64/vulkan,libjllama.so,no
vulkan-windows-x86-64,Windows/x86_64/vulkan,jllama.dll,no
vulkan-windows-aarch64,Windows/aarch64/vulkan,jllama.dll,no
opencl-android-aarch64,Linux-Android/aarch64/opencl,libjllama.so,no
opencl-windows-x86-64,Windows/x86_64/opencl,jllama.dll,no
opencl-windows-aarch64,Windows/aarch64/opencl,jllama.dll,no
Expand Down
68 changes: 61 additions & 7 deletions .github/workflows/publish.yml
Original file line number Diff line number Diff line change
Expand Up @@ -1652,8 +1652,10 @@ jobs:
# suite is CPU-only and fully covered by the `C++ Tests` job + the CPU Windows
# jobs; a GPU-linked jllama_test.exe cannot be discovered/run on a GPU-less
# GitHub runner (it errors probing for a CUDA device -> ctest *_NOT_BUILT).
# GGML_CUDA_CCCL_VERSION: fetch CCCL v3.4.3 like upstream's windows-cuda release job and the
# Linux CUDA build (see build_cuda_linux.sh) -- DeviceTopK needs it; drop it with CUDA 13.5+.
run: |
.github\build.bat -G "Ninja Multi-Config" -DGGML_CUDA=ON -DOS_NAME=Windows -DOS_ARCH=x86_64
.github\build.bat -G "Ninja Multi-Config" -DGGML_CUDA=ON -DGGML_CUDA_CCCL_VERSION=v3.4.3 -DOS_NAME=Windows -DOS_ARCH=x86_64
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
Expand Down Expand Up @@ -2044,6 +2046,57 @@ jobs:
path: ${{ github.workspace }}/llama/src/main/natives/net/ladenthin/llama/
if-no-files-found: error

build-windows-arm64-vulkan:
name: Build Windows 11 arm64 Vulkan
needs: [startgate, build-webui]
# Windows-on-ARM Vulkan (Snapdragon X / any Vulkan 1.2+ driver), the natives jar
# vulkan-windows-aarch64 -- the counterpart of upstream's windows arm64 Vulkan release
# (llama.cpp #29954). Same clang-cl + GGML_OPENMP=OFF toolchain as the arm64 CPU and
# OpenCL jobs (ggml refuses MSVC cl.exe on ARM). The SDK is installed the way upstream's
# release job does it: LunarG's (x64) installer with the com.lunarg.vulkan.arm64 component,
# which adds the arm64 import library; its x64 glslc runs under the runner's x64
# emulation. Build-only like every GPU job (no GPU on the runner).
runs-on: windows-11-arm
env:
SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}
VULKAN_VERSION: 1.4.357.0
steps:
- uses: actions/checkout@v7
- name: Download shared WebUI assets
uses: actions/download-artifact@v8
with:
name: webui-generated
path: ${{ github.workspace }}/llama/webui-generated/
- name: Set up MSVC developer environment (arm64)
uses: ilammy/msvc-dev-cmd@v1
with:
arch: arm64
- name: Install Vulkan SDK (with the arm64 component)
shell: pwsh
# Version and installer arguments as in upstream's release.yml (windows, vulkan arm64).
run: |
curl.exe -fSL -o "$env:RUNNER_TEMP\VulkanSDK-Installer.exe" "https://sdk.lunarg.com/sdk/download/$env:VULKAN_VERSION/windows/vulkansdk-windows-X64-$env:VULKAN_VERSION.exe"
& "$env:RUNNER_TEMP\VulkanSDK-Installer.exe" --accept-licenses --default-answer --confirm-command install com.lunarg.vulkan.arm64
if (-not (Test-Path "C:\VulkanSDK\$env:VULKAN_VERSION")) { throw "Vulkan SDK $env:VULKAN_VERSION was not installed" }
Add-Content $env:GITHUB_ENV "VULKAN_SDK=C:\VulkanSDK\$env:VULKAN_VERSION"
Add-Content $env:GITHUB_PATH "C:\VulkanSDK\$env:VULKAN_VERSION\bin"
- name: Install sccache (shared compiler cache)
if: env.USE_CACHE == 'true' && env.SCCACHE_WEBDAV_TOKEN != ''
continue-on-error: true
uses: ./.github/actions/install-sccache-windows
- name: Build libraries
shell: cmd
# Build the artifact only (see the CUDA job's note: GPU-less runner can't run a
# GPU-linked jllama_test; the C++ unit suite is covered by the CPU jobs).
run: |
.github\build.bat -G "Ninja Multi-Config" -DCMAKE_C_COMPILER=clang-cl -DCMAKE_CXX_COMPILER=clang-cl -DGGML_OPENMP=OFF -DGGML_VULKAN=ON -DOS_NAME=Windows -DOS_ARCH=aarch64
- name: Upload artifacts
uses: actions/upload-artifact@v7
with:
name: natives-vulkan-windows-aarch64
path: ${{ github.workspace }}/llama/src/main/natives/net/ladenthin/llama/
if-no-files-found: error

build-linux-x86_64-openvino:
name: Build Linux x86_64 OpenVINO (Intel)
needs: [startgate, build-webui]
Expand All @@ -2061,7 +2114,7 @@ jobs:
with:
distribution: 'temurin'
java-version-file: .java-version
- name: Install OpenCL dev + Intel OpenVINO 2026.4 (archive)
- name: Install OpenCL dev + Intel OpenVINO 2026.4.1 (archive)
run: |
# Intel's OpenVINO APT repo only publishes up to ~2025 (the /openvino/2026 path 404s), and
# 2025.x has the older ov::Allocator API that breaks ggml-openvino's template compile. So use
Expand All @@ -2072,11 +2125,11 @@ jobs:
# OPENVINO_VERSION_FULL (.github/workflows/release.yml at the pinned GIT_TAG); ggml-openvino
# is developed against that pair, so lagging it is what eventually breaks the compile. Both
# OpenVINO jobs here (Linux + Windows) use the same two values — bump them together:
# major = 2026.4 full = 2026.4.0.22959.99c81491cc3
# major = 2026.4.1 full = 2026.4.1.22982.07f9c262b05
# OpenCL headers (incl. the C++ CL/cl2.hpp via opencl-clhpp-headers) come from Ubuntu's own repos.
sudo apt-get update
sudo apt-get install -y ocl-icd-opencl-dev opencl-headers opencl-clhpp-headers intel-opencl-icd
url="https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4/linux/openvino_toolkit_ubuntu24_2026.4.0.22959.99c81491cc3_x86_64.tgz"
url="https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4.1/linux/openvino_toolkit_ubuntu24_2026.4.1.22982.07f9c262b05_x86_64.tgz"
sudo mkdir -p /opt/intel/openvino
curl -fSL "$url" | sudo tar -xz --strip-components=1 -C /opt/intel/openvino
echo "OpenVINO_DIR=/opt/intel/openvino/runtime/cmake" >> "$GITHUB_ENV"
Expand Down Expand Up @@ -2110,16 +2163,16 @@ jobs:
uses: ilammy/msvc-dev-cmd@v1
with:
arch: x64
- name: Install OpenCL headers (vcpkg) + Intel OpenVINO 2026.4
- name: Install OpenCL headers (vcpkg) + Intel OpenVINO 2026.4.1
shell: pwsh
# vcpkg's opencl port ships the full C++ headers incl. CL/cl2.hpp that OpenVINO's
# ocl_wrapper.hpp needs (the Khronos OpenCL-Headers dropped cl2.hpp) — same as upstream
# llama.cpp's windows-openvino job. OpenVINO 2026.4 matches ggml-openvino's target API.
# llama.cpp's windows-openvino job. OpenVINO 2026.4.1 matches ggml-openvino's target API.
# Keep the version in sync with the Linux OpenVINO job above (and with upstream's
# OPENVINO_VERSION_MAJOR / OPENVINO_VERSION_FULL) — see the note there.
run: |
C:\vcpkg\vcpkg install opencl:x64-windows
$url = "https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4/windows/openvino_toolkit_windows_2026.4.0.22959.99c81491cc3_x86_64.zip"
$url = "https://storage.openvinotoolkit.org/repositories/openvino/packages/2026.4.1/windows/openvino_toolkit_windows_2026.4.1.22982.07f9c262b05_x86_64.zip"
Invoke-WebRequest -Uri $url -OutFile "$env:RUNNER_TEMP\openvino.zip"
Expand-Archive -Path "$env:RUNNER_TEMP\openvino.zip" -DestinationPath "C:\openvino" -Force
# The archive extracts into a nested versioned folder; point OpenVINO_DIR at its runtime/cmake.
Expand Down Expand Up @@ -2402,6 +2455,7 @@ jobs:
- build-linux-x86_64-sycl-fp32
- build-windows-x86_64-sycl
- build-windows-arm64-opencl
- build-windows-arm64-vulkan
- build-linux-x86_64-openvino
- build-windows-x86_64-openvino
- test-cpp-linux-x86_64
Expand Down
44 changes: 44 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,33 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by
## [Unreleased]

### Added
- **Kolibri-1 support** (Aleph Alpha, architecture `kolibri1`, 78B German/English reasoning MoE) ahead of upstream
llama.cpp ([ggml-org/llama.cpp#29922](https://github.com/ggml-org/llama.cpp/issues/29922)), as the carried patch
`0016-model-kolibri1.patch`. It combines the two community ports and, unlike either of them, loads the GGUFs of
both community converters. Guarded by `test_kolibri1.cpp`, which compares tiny random models with an
independent reference written from Aleph Alpha's vLLM implementation. The patch is dropped once upstream adds
the architecture.
- **`GpuSplitMode.TENSOR`** (`--split-mode tensor`, tensor parallelism, EXPERIMENTAL upstream). The mode
existed upstream before; since llama.cpp b11450 (#26610) it also works across RPC servers.
- **Input/output modalities on `RouterModel` and `ModelMeta`** (llama.cpp b11429, #29987):
`getInputModalities()`, `getOutputModalities()` and `isDecisionModel()`, from upstream's new
`architecture` object of `GET /models` (the router computes it offline, so a decision model is
recognisable before its first load) and, for a loaded `LlamaModel`, from the same metadata in
`getModelMeta()`. Empty against a server before b11429.
- **`vulkan-windows-aarch64` natives jar** (Windows on ARM with a Vulkan 1.2+ driver), following
upstream's new Windows arm64 Vulkan release (llama.cpp b11395, #29954). Built natively on
`windows-11-arm` with `clang-cl`; also in the `all-windows-aarch64` fat jar, where the loader tries it
before OpenCL and the CPU.
- **`ModelParameters.setDraftSampling(DraftSampling)`** (`--spec-draft-sampling`, llama.cpp b11368):
`PROBABILISTIC` samples the speculative draft and has the target verify it by rejection sampling,
which accepts more drafted tokens at a temperature above zero; `GREEDY` is upstream's default. Applies
to a draft model and to a model's own MTP heads.
- **Decision models: `LlamaModel.handleSystemOne(String)`**, llama.cpp's TypeSafe-compatible
`/v1/systemone` API (upstream b11361): typed `choice` / `score` / `noul` questions about a state,
answered with probabilities in one forward pass, for the decision models upstream supports (laya,
julia-1, lev, openjev, kev, ...). The JNI method forwards to upstream's own route handler, so the
request and response are exactly the HTTP endpoint's; `NativeServer` serves `POST /v1/systemone` in
classic and attach mode. A model that is not a decision model throws a `LlamaException`.
- **`net.ladenthin:llama-atmosphere-agent` on Maven Central**, at the core's version: the agent's thin jar
(with `Main-Class`), sources and javadoc, published right after the reactor. Its pom names
`llama-platform` as a runtime dependency, so `jbang net.ladenthin:llama-atmosphere-agent:<version>`
Expand All @@ -18,6 +45,23 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by
core's; `check-natives.py` fails when they differ.

### Changed
- **Upgraded the pinned llama.cpp from b11320 to b11457**, in 21 reviewed steps, each ending at a tag.
Every carried patch that broke was traced to the one upstream commit that broke it, and the step
containing that commit ends at the first tag after it: `0007` at #29818 (b11361) and #29895 (b11401), `0014` at #29895, `0008`
at #29987 (the commit just before b11429), `0015` at #26610 (b11450). Each refresh moved context only,
and all nine patches are still needed. The new
upstream features this binding now exposes are listed under *Added*; the build follows upstream's
CUDA CCCL pin (v3.4.3) and OpenVINO 2026.4.1. Per-step record:
`docs/history/llama-cpp-breaking-changes.md`.
- **RPC protocol 8** (llama.cpp b11450, #26610): `RPC_PROTO_MAJOR_VERSION` 7 → 8. An `RpcServer` or
`--rpc` client of this release talks only to RPC peers of the same protocol -- upgrade the
`rpc-server`s and every JVM using `RpcServer` together.
- **Slot state files from earlier releases no longer restore** (llama.cpp b11411, #28498): upstream
now stores the exact KV-cache rotation in a state file and rejects one restored under a mismatched
rotation, which bumps `LLAMA_SESSION_VERSION` 10 → 11 and `LLAMA_STATE_SEQ_VERSION` 3 → 4. A file
written by `LlamaModel.saveSlot` (or the server's `/slots/{id}?action=save`) with an earlier jar is
rejected by `restoreSlot` with upstream's generic "invalid slot save file" message; regenerate it.
`Session` snapshots taken and restored within one process are not affected.
- **`ProcessRunner` rewritten on `ProcessBuilder`** (the helper `OSInfo` runs `uname` with): the timeout
is now real -- a command that does not end in time is killed and reported as an `IOException`, where
the old timeout overload ignored the result of `waitFor` and then blocked reading the output -- and the
Expand Down
Loading
Loading