Repository navigation
llama.cpp : bump version to 0.3.0 - #27696
Merged
Merged
Conversation
danbev
approved these changes
Aug 25, 2026
ggerganov
marked this pull request as ready for review
August 25, 2026 09:41
ggerganov
force-pushed
the
llama-rc-v0.3.0
branch
from
August 25, 2026 09:41
1ae324b to
088a8ba
Compare
Member
|
It's cool that we're doing semver releases now! |
thecodacus
pushed a commit
to thecodacus/llama.cpp
that referenced
this pull request
Sep 7, 2026
* llama.cpp : bump version to 0.3.0 * ci : update release default desc * scripts : add prompt for generating release summary
zbrad
pushed a commit
to zbrad/llama.cpp
that referenced
this pull request
Sep 10, 2026
* llama.cpp : bump version to 0.3.0 * ci : update release default desc * scripts : add prompt for generating release summary
pl752
pushed a commit
to pl752/llama.cpp
that referenced
this pull request
Sep 15, 2026
* llama.cpp : bump version to 0.3.0 * ci : update release default desc * scripts : add prompt for generating release summary
zsogitbe
pushed a commit
to zsogitbe/llama.cpp
that referenced
this pull request
Sep 17, 2026
* llama.cpp : bump version to 0.3.0 * ci : update release default desc * scripts : add prompt for generating release summary
frostyautumnleaf
pushed a commit
to frostyautumnleaf/llama.cpp
that referenced
this pull request
Oct 5, 2026
* llama.cpp : bump version to 0.3.0 * ci : update release default desc * scripts : add prompt for generating release summary
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
llama.cpp 0.3.0 introduces the dots3-note multimodal model (with a new DSA-ISWA KV cache), MTP support for GLM-4.5-Air, and tensor-split (
-sm tensor) plus multi-sequence rollback fixes for DeepSeek 4. ggml is bumped to v0.22.0 (meta-backend tensor split, per-op Metal kernels with parallel compilation, non-in-placeggml_clamp), while mtmd gains dots3-note vision/audio, WebP decoding and a Pillow-accurate resize. The server adds aLLAMA_SERVER_SLOTS_N_DIFFdebug knob, and the web UI gets tabbed chat navigation.New models
Core changes
-sm tensor(#26490)ggml_rope_set_offsetin deepseek2/4, dflash, minicpm3 and plm (#27382)\-in char classes as a literal hyphen (#27591)json.habstraction (#27511) with a clang LTO fix (#27575)fitmoved out of the server and now takesn_streamsinto account (#27496)-no-cnvCLI option (#27542)Multi-modality changes
resize_algofor all models (#27594)ggml_rope_set_offsetin the CLIP graph (#27521)Server changes
LLAMA_SERVER_SLOTS_N_DIFFenv var to widen the slot debug diff window (#27600)fit, now accounting forn_streams(#27496)json.habstraction (#27511)UI changes
ggml changes
ggml_clampto be a proper non-in-place op. It also brings new ops(
POOL_1D,PAD_REFLECT_1D), Q2_K SYCL kernels, MoE bias fusion on OpenCL, and assorted fixes across the CUDA, Metal, SYCL, Vulkan, OpenCL and WebGPU backends.