Skip to content

llama.cpp : bump version to 0.3.0 - #27696

Merged
ggerganov merged 3 commits into
masterfrom
llama-rc-v0.3.0
Aug 25, 2026
Merged

ggerganov merged 3 commits into
masterfrom
llama-rc-v0.3.0

Conversation

@ggerganov

@ggerganov ggerganov commented Aug 25, 2026 •

Copy link
Copy Markdown
Member

Overview

llama.cpp 0.3.0 introduces the dots3-note multimodal model (with a new DSA-ISWA KV cache), MTP support for GLM-4.5-Air, and tensor-split (-sm tensor) plus multi-sequence rollback fixes for DeepSeek 4. ggml is bumped to v0.22.0 (meta-backend tensor split, per-op Metal kernels with parallel compilation, non-in-place ggml_clamp), while mtmd gains dots3-note vision/audio, WebP decoding and a Pillow-accurate resize. The server adds a LLAMA_SERVER_SLOTS_N_DIFF debug knob, and the web UI gets tabbed chat navigation.

New models

  • Add dots3-note model with a new DSA-ISWA KV cache type (#27060)

Core changes

  • DeepSeek 4: add tensor-split mode via -sm tensor (#26490)
  • DeepSeek 4: fix rollback with multiple sequences (#26756)
  • Fix meta tensor split state propagation for tensor parallel (#27574)
  • GLM-4.5-Air: add MTP (multi-token prediction) support (#26534)
  • bailingmoe3: support DSpark (#27508)
  • mamba2: flatten in/out projections to dispatch GEMM instead of GEMV (#27513)
  • Models: use ggml_rope_set_offset in deepseek2/4, dflash, minicpm3 and plm (#27382)
  • Grammar: parse \- in char classes as a literal hyphen (#27591)
  • Common: add json.h abstraction (#27511) with a clang LTO fix (#27575)
  • Common: fit moved out of the server and now takes n_streams into account (#27496)
  • Common: fix draft-mtp with embeddings (#27400)
  • Arg: remove the -no-cnv CLI option (#27542)

Multi-modality changes

  • Support dots3-note vision and audio (#27524)
  • Support WebP images via ffmpeg (#27520)
  • Fix loading videos with the moov atom at the end of the file (#27596)
  • Use a Pillow-accurate resize algorithm and correct resize_algo for all models (#27594)
  • Use ggml_rope_set_offset in the CLIP graph (#27521)

Server changes

  • Add LLAMA_SERVER_SLOTS_N_DIFF env var to widen the slot debug diff window (#27600)
  • Slot fitting logic moved to the common fit, now accounting for n_streams (#27496)
  • Adopt the common json.h abstraction (#27511)

UI changes

  • Tabbed navigation for chat conversations (#27263)
  • Fix keyboard shortcuts for the chat tabs navigation (#27609)

ggml changes

  • ggml bumped to v0.22.0 (ggml/1607):
    • This release adds tensor-split support to the multi-backend (meta) backend with improved split-state propagation, reworks the Metal kernels into per-op sources with parallel compilation, and fixes ggml_clamp to be a proper non-in-place op. It also brings new ops
      (POOL_1D, PAD_REFLECT_1D), Q2_K SYCL kernels, MoE bias fusion on OpenCL, and assorted fixes across the CUDA, Metal, SYCL, Vulkan, OpenCL and WebGPU backends.

@github-actions github-actions Bot added build Compilation issues devops improvements to build systems and github actions labels Aug 25, 2026
@ggerganov
ggerganov marked this pull request as ready for review August 25, 2026 09:41
@ggerganov
ggerganov requested a review from a team as a code owner August 25, 2026 09:41
@ggerganov
ggerganov merged commit c1d0e7a into master Aug 25, 2026
12 checks passed
@ggerganov
ggerganov deleted the llama-rc-v0.3.0 branch August 25, 2026 09:42
@nikwen

nikwen commented Aug 25, 2026

Copy link
Copy Markdown
Member

It's cool that we're doing semver releases now!

thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
* llama.cpp : bump version to 0.3.0

* ci : update release default desc

* scripts : add prompt for generating release summary
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
* llama.cpp : bump version to 0.3.0

* ci : update release default desc

* scripts : add prompt for generating release summary
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
* llama.cpp : bump version to 0.3.0

* ci : update release default desc

* scripts : add prompt for generating release summary
zsogitbe pushed a commit to zsogitbe/llama.cpp that referenced this pull request Sep 17, 2026
* llama.cpp : bump version to 0.3.0

* ci : update release default desc

* scripts : add prompt for generating release summary
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
* llama.cpp : bump version to 0.3.0

* ci : update release default desc

* scripts : add prompt for generating release summary
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build Compilation issues devops improvements to build systems and github actions

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants