Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,11 +45,11 @@ from version 5.0.0 onward. Pre-fork releases (`1.x`–`4.2.0`) were authored by
core's; `check-natives.py` fails when they differ.

### Changed
- **Upgraded the pinned llama.cpp from b11320 to b11457**, in 21 reviewed steps, each ending at a tag.
- **Upgraded the pinned llama.cpp from b11320 to b11462**, in 23 reviewed steps, each ending at a tag.
Every carried patch that broke was traced to the one upstream commit that broke it, and the step
containing that commit ends at the first tag after it: `0007` at #29818 (b11361) and #29895 (b11401), `0014` at #29895, `0008`
at #29987 (the commit just before b11429), `0015` at #26610 (b11450). Each refresh moved context only,
and all nine patches are still needed. The new
and every patch is still needed. The new
upstream features this binding now exposes are listed under *Added*; the build follows upstream's
CUDA CCCL pin (v3.4.3) and OpenVINO 2026.4.1. Per-step record:
`docs/history/llama-cpp-breaking-changes.md`.
Expand Down
14 changes: 10 additions & 4 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@ This file provides guidance to Claude Code (claude.ai/code) when working with co

Java bindings for [llama.cpp](https://github.com/ggerganov/llama.cpp) via JNI, providing a high-level API for LLM inference in Java. The Java layer communicates with a native C++ library through JNI.

Current llama.cpp pinned version: **b11457**
Current llama.cpp pinned version: **b11462**

## Natives jars: one directory per backend (`.github/natives.csv`)

Expand Down Expand Up @@ -214,6 +214,12 @@ To change the CUDA version, update the following places:
`DeviceTopK` path needs CCCL >= 3.4.3 and falls back to a sort below it, and CUDA 13.4 bundles an
older 3.4. **Drop both flags once the toolkit is 13.5 or newer** (it bundles CCCL 3.5); follow
upstream's `release.yml` matrix comment, which says the same.
**The pin makes ggml fetch CCCL, and CCCL calls `include(CTest)` unconditionally**, which creates
`BUILD_TESTING` as a cache variable defaulting to ON. That is why `llama/CMakeLists.txt` declares
`option(BUILD_TESTING ... OFF)` before its first `FetchContent_MakeAvailable()`: declared after,
the option was a no-op, both CUDA jobs built `jllama_test`, and its gtest discovery failed on the
GPU-less runners (no `libcuda.so.1` / `nvcuda.dll`) -- the first Publish run after the bump
(37550451678). Keep the option there.
5. **`CLAUDE.md`** — the "Current CUDA version" line above.

Available CUDA versions for RHEL8/Manylinux_2_28 can be browsed at:
Expand Down Expand Up @@ -704,7 +710,7 @@ needs no extra step here, `build-webui` re-reads the tag and rebuilds the matchi
ships no UI):
```bash
# needs node/npm + network for the asset build; the embed step is plain cmake -P
git clone --depth 1 --branch b11457 https://github.com/ggml-org/llama.cpp /tmp/lc
git clone --depth 1 --branch b11462 https://github.com/ggml-org/llama.cpp /tmp/lc
( cd /tmp/lc/tools/ui && npm ci && npm run build )
mkdir -p webui-generated /tmp/ui-gen
cmake -DUI_SOURCE_DIR=/tmp/lc/tools/ui -DUI_BINARY_DIR=/tmp/ui-gen \
Expand Down Expand Up @@ -744,7 +750,7 @@ cache lives in **Depot Cache** over sccache's **WebDAV** backend:
- `SCCACHE_WEBDAV_TOKEN: ${{ secrets.DEPOT_TOKEN }}` — a Depot **organization** token, stored
as the repo secret **`DEPOT_TOKEN`**.

Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11457`), the
Because `sccache` is **content-addressed** and llama.cpp is pinned (`GIT_TAG b11462`), the
~280 upstream object files are byte-identical every run, so a warm cache recompiles only the
*changed* files. Depot's cache is **shared across all branches** (unlike GitHub's
per-branch `actions/cache`), so every branch builds incrementally; a `b<nnnn>` version bump
Expand Down Expand Up @@ -1867,7 +1873,7 @@ ctest --test-dir build --output-on-failure -R "ResultsToJson"

#### Upstream source location (in CMake build tree)

llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11457`.
llama.cpp is fetched via CMake FetchContent, pinned to `GIT_TAG b11462`.

**GoogleTest** is a separate `BUILD_TESTING`-only FetchContent (`GIT_TAG v1.18.0`), used solely
by the `jllama_test` C++ unit-test binary — not by the shipped library, and not coupled to the
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,7 +11,7 @@
**Build:**
![Java 8+](https://img.shields.io/badge/Java-8%2B-informational)
![Platform](https://img.shields.io/badge/Platform-Linux%20%7C%20macOS%20%7C%20Windows%20%7C%20Android-lightgrey)
[![llama.cpp b11457](https://img.shields.io/badge/llama.cpp-%23b11457-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11457)
[![llama.cpp b11462](https://img.shields.io/badge/llama.cpp-%23b11462-informational)](https://github.com/ggml-org/llama.cpp/releases/tag/b11462)
[![JPMS](https://img.shields.io/badge/JPMS-modular%20JAR-25A162)](https://openjdk.org/projects/jigsaw/)
![JUnit](https://img.shields.io/badge/tested%20with-JUnit6-25A162)
[![JSpecify](https://img.shields.io/badge/JSpecify-1.0.0%20%40NullMarked-25A162)](https://jspecify.dev)
Expand Down
4 changes: 4 additions & 0 deletions docs/history/llama-cpp-breaking-changes.md
Original file line number Diff line number Diff line change
Expand Up @@ -828,3 +828,7 @@ Used during `llama.cpp` version bumps: when upgrading, scan this file from the r
| b11450–b11454 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11454 with `git apply`. Patch-target files in the range: `src/llama-model.{cpp,h}` (`0012`; the two new architectures are registered away from `load_tensors`' split block and the helper declarations). Drop-checks against pristine b11454 unchanged: `0001` override present and `common_params_parse_main` absent, `0002` assignment unconditional, no `0003` getter, no `0006`-`0008` counterpart, no `0012` guard (`split_sum` still unguarded), no `0014` hook, no `0015` stop function. `tools/server/` unchanged in the range. |
| b11454–b11457 | Three commits, 9 files, ~122/27 lines. **#30045 adds the PLaMo-3 tokenizer pre-segmentation**, a new `LLAMA_VOCAB_TYPE_PLAMO3 = 8` (`include/llama.h`, appended to the enum, so no existing value moves). `ModelMeta.getVocabType()` returns the number as is (`8` for such a model); the project has no Java enum of vocab types to extend. #28782 uses a per-thread CUDA stream for the buffer-init padding memset; #29955 adds BF16 to CUDA's `XIELU`. **No project source change.** This is the target of the bump: b11457 was the newest llama.cpp release when it was done (b11458 had no release assets). |
| b11454–b11457 | patches + upstream verification | **All nine patches apply unchanged**, verified in order against pristine b11457 with `git apply`. No patch-target file in the range; nothing under `common/` or `tools/server/` changed. Drop-checks against pristine b11457 -- the final state of the bump: `common_params_parse` still carries the `#ifdef _WIN32` `argv = utf8.ptrs.data()` override and `common_params_parse_main` is absent (`0001` needed); `load_progress_callback` still assigned unconditionally (`0002`); no slot-similarity getter (`0003`); no `llama_server_attach`, `LLAMA_SERVER_WORKER_CMD` or embedded-shutdown counterpart (`0006`-`0008`); no `split_sum == 0` guard (`0012`); no callback hook in `common/log.h` (`0014`); no `stop_server` and a `GGML_ABORT("Failed to connect` still present (`0015`). **Every carried patch is still needed.** Across the whole bump (b11320 → b11457) four patches broke, each at exactly one upstream commit: `0007` at #29818 (b11361) and #29895 (b11401), `0014` at #29895 (b11401), `0008` at #29987 (`4d60b4d08`, before b11429), `0015` at #26610 (b11450); every refresh moved context only. |
| b11457–b11458 | One commit, 9 files, ~978/733 lines, **#30067 only**: the Hexagon backend's CPY/CONCAT/CONT/DUP overhaul (DMA/HVX for every path, `htp/hex-cpy-dma.h` renamed to `dma-copy.h`) plus its Snapdragon developer guide. No natives jar builds Hexagon, so **nothing in it reaches a shipped library**. At 112 KiB it is over the 100 KiB step threshold, but a single commit cannot be split further. **No project source change.** |
| b11457–b11458 | patches + upstream verification | **All ten patches apply unchanged** (`0001`-`0016`, now including the Kolibri-1 carry `0016`), verified in order against pristine b11458 with `git apply`. No patch-target file and nothing under `src/`, `common/`, `tools/server/` or `include/` in the range. |
| b11458–b11462 | Four commits, 11 files, ~119/192 lines, backends only. **SYCL:** #27689 removes the separate flash-attention KV buffers (`fattn-buffers.{cpp,hpp}` deleted; `ggml-sycl/CMakeLists.txt` globs its sources, so nothing to wire) and #29500 adds an IQ3_S multi-column MMVQ kernel. **Vulkan:** #30049 skips the direct host read of uncached host-visible memory on AMD UMA devices (iGPUs), which is write-combined and slow to read, and takes the device-to-host copy path instead. **WebGPU:** #27069 (not built here). **No project source change.** This is the target of the bump: b11462 is the newest llama.cpp release. |
| b11458–b11462 | patches + upstream verification | **All ten patches apply unchanged**, verified in order against pristine b11462 with `git apply`. No patch-target file and nothing under `src/`, `common/`, `tools/server/` or `include/` in the range, so every drop-check of the b11457 row still holds unchanged; `git grep -i kolibri` finds nothing under `src/` at b11462, so `0016` is still needed too. No `.github/` change upstream either: OpenVINO, CCCL (`v3.4.3`) and the ROCm wheels stay where b11457 left them. |
13 changes: 10 additions & 3 deletions llama/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -36,6 +36,14 @@ set(LLAMA_CURL OFF)

option(LLAMA_VERBOSE "llama: verbose output" OFF)

# Declared here, before any FetchContent_MakeAvailable(), and not next to the tests below: a
# subproject that calls include(CTest) creates BUILD_TESTING as a cache variable defaulting to ON,
# after which this option() would be a no-op. CCCL does exactly that, unconditionally, and ggml
# fetches it whenever GGML_CUDA_CCCL_VERSION is set (our CUDA builds pin v3.4.3) -- the CUDA job
# then built jllama_test, whose gtest discovery needs libcuda.so.1 and failed the build on the
# GPU-less runner. -DBUILD_TESTING=ON still enables the tests.
option(BUILD_TESTING "Build C++ unit tests for jni_helpers / json_helpers / utils" OFF)

#################### json ####################

FetchContent_Declare(
Expand Down Expand Up @@ -182,7 +190,7 @@ set(LLAMA_BUILD_APP OFF CACHE BOOL "" FORCE)
FetchContent_Declare(
llama.cpp
GIT_REPOSITORY https://github.com/ggerganov/llama.cpp.git
GIT_TAG b11457
GIT_TAG b11462
PATCH_COMMAND ${CMAKE_COMMAND}
-DPATCH_DIR=${CMAKE_CURRENT_SOURCE_DIR}/patches
-DLLAMA_SRC=<SOURCE_DIR>
Expand Down Expand Up @@ -537,8 +545,7 @@ endif()

#################### C++ unit tests ####################

option(BUILD_TESTING "Build C++ unit tests for jni_helpers / json_helpers / utils" OFF)

# BUILD_TESTING is declared at the top of this file, before any subproject (see there).
if(BUILD_TESTING)
FetchContent_Declare(
googletest
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -9,28 +9,28 @@
* library was compiled against, exposed as a compile-time constant so callers can render a badge or
* emit a startup log line without loading the native library.
*
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11457"}) that mirrors the
* <p>{@link #LLAMA_CPP_VERSION} is a pure-Java string ({@code "b11462"}) that mirrors the
* {@code GIT_TAG} in {@code llama/CMakeLists.txt}. It is available even when {@code libjllama} is
* absent (pure-Java checkout, before {@code System.load}), which is what makes it suitable for a
* lightweight version badge in Android or other UIs.</p>
*
* <p>For the <em>authoritative</em> value that is baked into the native binary — the build number
* plus the resolved upstream commit, e.g. {@code "b11457-<commit>"} — call
* plus the resolved upstream commit, e.g. {@code "b11462-<commit>"} — call
* {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} instead; that reads llama.cpp's own
* {@code build-info} through JNI and therefore cannot drift from the compiled library (but requires
* the native library to be loaded).</p>
*/
public final class LlamaCppVersion {

/**
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11457"}.
* The pinned llama.cpp release tag this library was built against, e.g. {@code "b11462"}.
*
* <p>Kept in lockstep with {@code GIT_TAG} in {@code llama/CMakeLists.txt} — see the
* "Upgrading/Downgrading llama.cpp Version" checklist in {@code CLAUDE.md}. This is the
* compile-time pin; use {@link net.ladenthin.llama.LlamaModel#getLlamaCppBuildInfo()} for the
* value actually linked into the native binary.</p>
*/
public static final String LLAMA_CPP_VERSION = "b11457";
public static final String LLAMA_CPP_VERSION = "b11462";

// Constants holder — not instantiable.
private LlamaCppVersion() {}
Expand Down
Loading