Skip to content

mtmd: support Qwen3-TTS (note: breaking change to llama-tts binary) - #26254

Merged
ngxson merged 56 commits into
masterfrom
xsn/qwen3-tts
Aug 4, 2026
Merged

ngxson merged 56 commits into
masterfrom
xsn/qwen3-tts

Conversation

@ngxson

@ngxson ngxson commented Jul 28, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

Target support https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-Base

Qwen3-TTS

Available params:

  • --tts-lang can be zh, en, de, it, pt, es, ja, ko, fr, ru (default: en)
  • --tts-speaker-file should point to a speaker reference audio file (wav, mp3)

Example usage:

llama-tts -hf ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF \
    -p "Hello world" \
    --tts-lang english \
    --tts-speaker-file speaker.mp3 \
    --output out.wav

API design choices

This Qwen3-TTS works mostly the same as Sesame CSM, see PR #12648 :

  • Audio reference is encoded to text embd space via an audio encoder (speaker_encoder for qwen; mimi encoder for sesame)
  • Audio reference and prompt are decoded via a causal "backbone" model (talker.model)
  • Backbone model sample the first codebook entry (aka semantic audio code) and expose the output embedding
  • "Code predictor" model take the sampled semantic code & backbone embedding and generate the next 15 acoustic tokens (causal, similar to MTP model)
  • Sum 16 codebook entries, then decode the summed value via backbone to generate the next audio code

These design choices are made to adapt this model to existing llama.cpp infrastructure:

  • For talker.model:
    • codec_embedding is concat to the text embedding table, vocab is extended, example:
      • codec_bos_id(2149) --> "<|codec_bos|>"
      • codec_eos_token_id(2150) --> "<|codec_eos_token|>"
      • codec_language_id.chinese(2055) --> "<|codec_language_chinese|>"
      • other rows --> "<|codec_0|>", "<|codec_1|>", ..., "<|codec_1023|>"
    • output tensor codec_head is smaller than vocab, so logits will be padded at inference time
    • suppress_tokens is used to limit the backbone to only sample either semantic or EOS (stop) token
  • For speaker_encoder: treat it as a normal audio encoder, using existing mtmd_audio infrastructure
  • For code_predictor:
    • I don't reuse libllama or MTP infra for it, because (1) it will introduce quite a lot of hacky patches and (2) libllama doesn't support generating N tokens in-graph; it still needs to sync at every token which makes it quite slow (memory bound) for very small model like code_predictor (ref: tts : implement sesame CSM + Mimi decoder #12648 (comment))
    • The model runs 15 steps + sampling to generate the 15 acoustic codes; everything is done on one single graph (one forward output 15 codes)
    • To make that possible, I implemented (1) backend sampling for top_k/p and (2) minimal on-graph kv management; kv is alloc once per forward call (we don't reuse it afterward anyway)

Develooment

TODO - easy items (can reuse the existing infra):

  • convert talker backbone
  • convert speaker encoder
  • load speaker encoder via mtmd and impl the encoder pass

TODO - complex items (need to write new components):

  • handle codec_embeddings and codec control tokens --> append them to the vocab, see conversion script
  • handle codec_head --> use it as output head for backbone, pad logits output to match vocab; we also masked away codec control tokens from sampling via suppress_tokens
  • convert & pack the rest into mmproj
  • create mtmd_gen interface
  • impl code_predictor
  • impl code2wav in mtmd
  • add helper
  • revamp llama-tts
  • update dev docs

Follow-up PRs:

  • add support to llama-server
  • add warmup pass, make sure --fit handles it correctly

Requirements

  • I have read and agree with the contributing guidelines
  • AI usage disclosure: About 5-10% of core code (API design and some boilerplates) are human-coded, the rest is AI-generated to fit the target design

@ngxson ngxson changed the title convert text model mtmd: support Qwen-TTS Jul 28, 2026
@ngxson ngxson changed the title mtmd: support Qwen-TTS mtmd: support Qwen3-TTS Jul 28, 2026
@github-actions github-actions Bot added the mtmd Related to multimodal functionality (video/image/audio) label Jul 28, 2026
@github-actions github-actions Bot added the model Model specific label Jul 29, 2026
Co-authored-by: Pascal <admin@serveurperso.com>
@woheller69

Copy link
Copy Markdown

Generates frames and then fails...

ggml-cpu/ops.cpp:4886: GGML_ASSERT(i01 >= 0 && i01 < ne01) failed

llama.cpp b10276 and Q8 files from here: https://huggingface.co/ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF/tree/main

@ServeurpersoCom

ServeurpersoCom commented Aug 5, 2026 •

Copy link
Copy Markdown
Contributor

Generates frames and then fails...

ggml-cpu/ops.cpp:4886: GGML_ASSERT(i01 >= 0 && i01 < ne01) failed

llama.cpp b10276 and Q8 files from here: https://huggingface.co/ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF/tree/main

Please test this : master...ServeurpersoCom:llama.cpp:ggml/build-forward-order

It's the tested fix for CPU

@woheller69

Copy link
Copy Markdown

yes, that fixes it :-)

@quine00

quine00 commented Aug 6, 2026

Copy link
Copy Markdown

usually that's mentioned on the model card, it seems to be 3s:

I'm not sure what to expect, but I tried cloning my own voice with reference audio, using invocations like,

./build/bin/llama-tts -hf ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF \
  -p "the quick brown fox jumped over the lazy dog" \
  --tts-lang english --tts-speaker-file /tmp/out.wav --output out.wav

Where /tmp/out.wav contained various reference texts spoken by myself ranging from 3s to 30s in length. In all cases, the resulting speech was no where near my original voice, which is a fairly standard British English accent. Is this a known limitation? The model card / description makes this feature sound like it can do a better job than I've managed to reproduce.

localai-org-maint-bot pushed a commit to mudler/LocalAI that referenced this pull request Aug 9, 2026
Picks up ggml-org/llama.cpp#26254 (Qwen3-TTS via mtmd) and #26536 (the
short-input audio chunk fix). Adds 0002-add-server-task-type-tts.patch,
the server-side half of the still-draft #26603, so TTS runs through the
slot scheduler instead of racing it. Remove that patch when #26603 merges.

The patch is rebased on top of the score patch: its tokenize-switch hunk
collided with the SERVER_TASK_TYPE_SCORE case, and its lone SRV_WRN call
passes no variadic argument, which the macro cannot expand. The score
patch itself needed no refresh.

Also fixes fallout from the bump in grpc-server.cpp: upstream dropped the
per-slot n_ctx argument from server_schema::eval_llama_cmpl_schema. Only
the schema branch loses it, since forks predating the server-schema split
still expect the old argument list.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
mudler added a commit to mudler/LocalAI that referenced this pull request Aug 10, 2026
* fix(config): do not read a TTS speaker-encoder mmproj as vision support

Qwen3-TTS on llama-cpp ships an mmproj holding the speaker encoder and
code predictor. VisionSupported() treated any non-empty MMProj as proof
of image input, so every such model would be advertised as vision-capable.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(llama-cpp): add TTS request option parsing helper

Validates text and speaker reference presence and strictly parses the
top_k / top_p per-request params, in a header with no llama.cpp or gRPC
dependencies so the standalone C++ unit test gate picks it up.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(llama-cpp): range-check the TTS top_k and top_p request params

Format validation alone let NaN, infinity and out-of-range values through.
The consumer copies both values into the audio generation input
unconditionally and only guards its separate sampler assignment with
"> 0", a test NaN also fails, so a NaN reached llama.cpp with the guard
never firing. top_k must now be >= 0 and top_p must fall within 0.0 to 1.0
inclusive, with the bound written as a negated in-range test so NaN is
rejected rather than silently accepted.

Also cover the two checks the suite could not previously kill: the
whole-string check in the float parser and the int32 range check.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(llama-cpp): bump pin to f9e832c10 and carry the TTS server task

Picks up ggml-org/llama.cpp#26254 (Qwen3-TTS via mtmd) and #26536 (the
short-input audio chunk fix). Adds 0002-add-server-task-type-tts.patch,
the server-side half of the still-draft #26603, so TTS runs through the
slot scheduler instead of racing it. Remove that patch when #26603 merges.

The patch is rebased on top of the score patch: its tokenize-switch hunk
collided with the SERVER_TASK_TYPE_SCORE case, and its lone SRV_WRN call
passes no variadic argument, which the macro cannot expand. The score
patch itself needed no refresh.

Also fixes fallout from the bump in grpc-server.cpp: upstream dropped the
per-slot n_ctx argument from server_schema::eval_llama_cmpl_schema. Only
the schema branch loses it, since forks predating the server-schema split
still expect the old argument list.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(llama-cpp): implement the TTS and TTSStream RPCs

Both were declared in backend.proto but unimplemented. They now submit a
SERVER_TASK_TYPE_TTS task and drain the response reader, the same shape
PredictStream uses.

The streaming path emits a leading sample_rate message and then raw PCM,
because ModelTTSStream builds the WAV header itself; the non-streaming
path emits a complete WAV to the requested dst.

The streamed samples are converted from the pipeline's float32 to signed
16-bit first. MTMD_HELPER_GEN_AUDIO_OUTTYPE_PCM hands back floats, while
the header ModelTTSStream writes announces 16-bit samples, so shipping
the floats verbatim would decode as noise.

prepare.sh and CMakeLists.txt now stage tts_request_options.h alongside
the other grpc-server helpers, and register its standalone test with
ctest the way passthrough_options_test is registered.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(llama-cpp): mask non-codec tokens for Qwen3-TTS generation

The Qwen3-TTS gen-audio pipeline maps a sampled backbone token to a
codebook row with an unchecked subtraction, in mtmd-helper-gen.cpp:

    inp.code0 = sampled - codec_0;

For ggml-org/Qwen3-TTS-12Hz-1.7B-Base-GGUF the vocab is 155008 tokens,
<|codec_0|> is 151936 and the codec codes end at 153983. The model's own
tokenizer.ggml.suppress_tokens holds 1023 ids covering 153984..155007,
every special above the codec range except <|codec_eos_token|> (154086)
which stays reachable as the stop token. Nothing masks the text range
0..151935, so the backbone can sample a text token at any step, the
subtraction goes negative, and ggml_compute_forward_get_rows aborts the
whole backend process on GGML_ASSERT(i01 >= 0 && i01 < ne01).

Complete the mask upstream started: bias every token below <|codec_0|>
to -INFINITY for TTS tasks so only codec codes and the codec EOS remain
reachable. The biases are appended to task.params.sampling.logit_bias,
which common_sampler_init already merges with the model's suppress
tokens into one llama_sampler_init_logit_bias, so no sampler is added to
the chain. Measured cost is 0.082 ms per sampled token and 1.16 MB, set
against a forward pass in the multi-millisecond range.

It lands in launch_slot_with_task rather than in a route handler so that
llama.cpp's own POST /tts and LocalAI's TTS/TTSStream RPCs are both
covered, and <|codec_0|> is resolved from the vocab rather than
hardcoded so a model without it is left alone.

This is reproducible with upstream's own llama-tts and no LocalAI code
loaded, aborting at frame 55 on Q4_K_M and frame 71 on Q8_0, so it is
neither a quantization artifact nor an artifact of the gRPC adapter.
Two further defects in the same draft pipeline still prevent end-to-end
audio; they are independent of this one and are recorded in the task
report for an upstream bug report.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* chore(llama-cpp): bump pin to 9de0fcf2b and drop the TTS codec mask

Upstream fixed the Qwen3-TTS abort in ggml-org/llama.cpp c8e03ce81
("mtmd/ggml: add ggml_build_forward_order", #26649), landed one hour
after the previous pin. ggml_build_forward_expand marks a tensor and all
its ancestors for compute, so using it as a pure ordering hint defeated
ggml_build_forward_select and made GEN_WAV calls execute the GEN_CODE
branch against a stale inp_code0, hitting the get_rows bound assert in
ggml_compute_forward_get_rows.

That single defect accounts for every abort seen on this model, so
0003-mask-non-codec-tokens-for-tts.patch is removed rather than rebased.
The mask changed the observed behavior, but it was perturbing a graph
ordering bug rather than fixing a sampling one: at the new pin the whole
path works without it. Keeping it would have meant carrying a 152k-entry
logit bias, and rebasing it on every pin bump, for no benefit.

Verified at 9de0fcf2b with only 0001 and 0002 applied, which both apply
clean with no fuzz and needed no rebase:

  non-streaming  HTTP 200, 410924 bytes, 8.56 s
                 RIFF (little-endian) data, WAVE audio, Microsoft PCM,
                 16 bit, mono 24000 Hz
  streaming      HTTP 200, 560684 bytes, 11.68 s, exactly one RIFF at
                 byte 0, same format, which also exercises the
                 float32-to-s16 conversion at runtime for the first time

Pristine unpatched llama-tts at the same pin now also completes, 130
frames to a valid WAV, where it aborted at frame 55 before.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(llama-cpp): clear the TTS slot sequence between requests

Only the first TTS request in a backend process succeeded. Every later
one failed instantly, in about 0.13 s, with "TTS prompt processing
failed" from step_prompt, regardless of streaming or non-streaming and
regardless of the text. With LOCALAI_SINGLE_ACTIVE_BACKEND=true the
process is kept alive between requests, so a deployment would have
served exactly one utterance per backend start.

The cause is missing KV hygiene, not anything in the gRPC adapter. TTS
slots never enter the shared batch: pre_decode() returns early for them
and process_tts_slots() drives them instead, so they skip the
prompt-cache bookkeeping that clears a slot's sequence between requests.
Nothing in the gen-audio path makes up for it: mtmd_helper_gen_audio_reset
only clears host-side buffers, and the pipeline always decodes from
position 0 into the sequence identified by slot.id. So the second task
on a slot writes positions 0..N over the first task's tokens and
llama_decode fails.

Fix is one call to slot.prompt_clear(), the same helper the normal path
uses, in the SERVER_TASK_TYPE_TTS branch of launch_slot_with_task before
set_input. It goes into 0002 rather than a new patch file because it is
a defect in the code that patch introduces, and the header now records
it as ours so we know whether it still needs carrying if #26603 merges
without it.

Verified in one backend process, different text on every request:
three consecutive non-streaming requests, three consecutive streaming
requests, and an interleaved non-streaming, streaming, non-streaming,
streaming run. All ten returned HTTP 200 with
RIFF ... WAVE audio, Microsoft PCM, 16 bit, mono 24000 Hz, the streamed
ones carrying exactly one RIFF header at byte 0, and every output
measured as real speech rather than silence or a truncated fragment.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(llama-cpp): expose max_frames for TTS requests

The Qwen3-TTS backbone does not always emit <|codec_eos_token|>, and
when it does not, generation runs to upstream's 512-frame n_predict
default. At the model's 12.5 Hz frame rate that is 40.96 s of audio,
which a short input can trigger: one request in this session produced
40.96 s for a ten-word sentence. prepareTTSTask hardcoded n_predict to
-1, so callers had no way to bound it.

Add a max_frames key alongside top_k and top_p, parsed with the same
strict whole-string parsing so a typo is an error rather than a silently
truncated value, and rejected with a field-naming message when negative.
0 keeps the existing sentinel convention and means unset, so a request
that omits it behaves exactly as before.

Named max_frames rather than n_predict because frames are what the
parameter means at a TTS endpoint: one frame is 0.08 s of audio.

The 512-frame default is deliberately unchanged. Lowering it would
truncate legitimately long inputs, which is a worse failure than an
occasionally overlong one.

Verified end to end on one text of thirty words:

  max_frames=25    HTTP 200,  96044 bytes,  2.00 s, exactly 25 frames
  max_frames=50    HTTP 200, 192044 bytes,  4.00 s, exactly 50 frames
  no max_frames    HTTP 200, 572204 bytes, 11.92 s, stopped at its own
                   codec EOS after 149 frames, unchanged behavior

  max_frames=-1    InvalidArgument "max_frames must be >= 0, got \"-1\""
  max_frames=many  InvalidArgument "max_frames must be an integer, got \"many\""

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(llama-cpp): send the TTS sample rate up front, and tidy three review items

Four items from the Task 4 review.

Streaming first-byte latency. TTSStream sent the sample-rate reply only
once the first audio result arrived, and a chunk needs a whole 72-frame
window, roughly 5.8 s of audio and far longer in wall time on CPU. The
Go side blocks on that reply before it can emit the WAV header, so a
streaming client sat at zero bytes for the whole stretch. The rate is a
property of the loaded model and is available synchronously from
mtmd_gen_audio_get_info, so it now goes out immediately after post_task
and the rate_sent bookkeeping is gone. Measured on a warm model, first
byte drops from 30.48 s to 0.014 s, and the output is still a valid WAV
with exactly one RIFF header at byte 0.

Unchecked close. The non-streaming path ignored ofstream::close(), so a
failure that only surfaces on flush was reported as success while
leaving a truncated file at dst. It now returns INTERNAL like the other
write failures.

Wrong comment on set_lang. gen_audio::inp::get() already maps a stored
blank to nullptr, so our guard is behavior-preserving, not
behavior-fixing. The comment claimed otherwise; the code was right.

Repetition penalty. penalty_last_n = -1 is inert at this pin, because
llama_sampler_init_penalties clamps it with std::max(penalty_last_n, 0)
and then builds a disabled sampler, so the 1.05 penalty never applies.
Upstream's README attributes looping to a missing repeat_penalty, so it
was worth testing as a root-cause fix for the model running to the frame
cap. Dropping the line lets the sampling default of 64 apply, which was
confirmed in the sampler chain trace as penalty_last_n = 64 with
repeat_penalty = 1.050. Over 15 uncapped short requests each way it did
not help: 0 of 15 ran to the cap with the penalty inert, 1 of 15 with it
active. Both lines are therefore kept for parity with upstream's draft,
and a comment now records that the pair is inert and why, so the next
reader does not believe a penalty is applied. max_frames remains the way
to bound output.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* build(llama-cpp): let unpatched forks opt out of the TTS task

turboquant and bonsai copy grpc-server.cpp into llama.cpp forks that do
not carry our patches. disable-tts-task.sh injects the same kind of
preprocessor switch disable-score-task.sh already uses, so those builds
answer UNIMPLEMENTED rather than failing to compile.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): keep a TTS speaker-encoder projector out of vision detection

Task 1 exempted a declared-TTS model's mmproj from VisionSupported, but the
first real gallery entry with an mmproj still came back vision-capable through
two paths the earlier fix did not close.

GuessUsecases has no FLAG_VISION branch, so it falls through to true for any
chat-ish model. That is not just a wrong answer at the call site:
syncKnownUsecasesFromString rewrites KnownUsecaseStrings from HasUsecases, and
the loader calls it more than once per config file, so the guessed FLAG_VISION
is written out and parsed back into KnownUsecases as if the operator had
declared it. Give GuessUsecases a FLAG_VISION branch that defers to the same
explicit signals VisionSupported uses.

Second, llama.cpp builds an mtmd context for the speaker-encoder projector and
reports its media marker on the first chat probe, which resurrected vision
after the model had been used once. Apply the same declared-TTS exemption to
MediaMarker that the mmproj check already had.

Verified against the qwen3-tts-llamacpp-q4 gallery entry: no vision capability
and no image input modality, before load, after a TTS request, and after a chat
probe.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* feat(gallery): add Qwen3-TTS entries for the llama-cpp backend

Two entries over upstream's own GGUF conversion, Q8_0 and Q4_K_M, each
pairing a backbone with the Q8_0 projector. Named to sit alongside the
existing qwen3-tts-cpp entries rather than replace them.

Also tags the llama-cpp backend text-to-speech / TTS so the backend browser
surfaces the capability.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* docs: cover Qwen3-TTS on the llama-cpp backend

Adds the gallery variants, the two-file mmproj configuration, the
required voice reference, and the language and sampling knobs. Also
corrects the streaming-support list, which named only voxcpm.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(config): register llama-cpp as a TTS and voice-cloning backend

The branch taught the llama-cpp backend to serve Qwen3-TTS and shipped two
gallery entries for it, but never told the capability table. llama-cpp still
declared only the text RPCs and usecases, so:

- VoiceCloningForModel returned nil at the capability check, before it ever
  reached the model's own tts.voice_cloning override, and /tts answered 400
  "selected model does not support reference-audio voice cloning" for any
  localai://voice-profiles/... voice. No model YAML could opt back in.
- GET /api/backends/usecases did not list tts for llama-cpp, so the gallery
  greyed out the TTS filter for the entries this branch adds.
- The React TTS page saw voice_cloning: null and kept both models out of the
  Voice Library.

Add the TTS RPCs and usecase, and the reference-audio contract.

The contract needs narrowing, because the per-backend switch in
VoiceCloningForModel ends in a permissive default: an unnarrowed entry would
have advertised reference-audio cloning on every GGUF chat model in the
gallery. Narrow on the declared TTS usecase rather than the model name. The
TTS checkpoints are the only llama-cpp models carrying known_usecases: [tts];
name matching would have to guess at third-party repacks, and "base", the
substring the neighbouring Qwen and vLLM cases key on, is a routine word in
text-model names. The check reads the declared bit directly instead of going
through HasUsecases, which falls through to GuessUsecases and would hand the
decision to a heuristic that never had a llama.cpp TTS model in mind.

DefaultUsecases stays [chat]: a bare GGUF served by llama.cpp is a chat model,
and both the gallery filter and the importer read that field.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(gallery): declare what nemotron-3-nano-omni actually accepts

The entry is backend: vllm-omni with known_usecases: [chat, completion], no
mmproj and no media marker, so it used to report vision only through the
blanket GuessUsecases fallthrough that the vision branch in this branch
removed. Nemotron 3 Nano Omni is a multimodal understanding model: image,
video and audio in, text out. Declaring that is what the sibling
vllm-omni-qwen3-omni-30b already does.

known_usecases gains vision only. FLAG_VIDEO is video GENERATION, an output
modality, and this model generates none; video and audio input belong in
known_input_modalities, which is where AudioInputSupported and
VideoInputSupported read them from.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(importers): import a Qwen3-TTS GGUF repo as TTS, not chat

The llama-cpp importer hardcodes known_usecases: [chat] and assigns any
mmproj-matching file as a vision projector, so ggml-org/Qwen3-TTS-12Hz-1.7B-
Base-GGUF imported as a chat model with vision. Both fields were wrong, and
the model was unreachable from /tts and from the Voice Library.

Filenames cannot fix this. A Qwen3-TTS repo has the exact shape of a vision
repo, one backbone GGUF plus one mmproj-*.gguf, so the projector's own header
is the only honest signal: mtmd writes clip.has_gen_audio_encoder for the
projectors it can drive as a speech pipeline and refuses to build one without
it. Probe the selected mmproj for that flag, reusing the range-fetch the MTP
detection already does, and declare tts when it is set. The mmproj assignment
then stops reading as vision on its own, since a declared-TTS model already
exempts its projector from vision detection.

The probe is best-effort like the MTP one: a network blip leaves the chat
default in place rather than failing the import.

Verified against the real artifacts on disk: the Qwen3-TTS projector reports
gen-audio, its backbone does not.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

* fix(llama-cpp): stop non-TTS models crashing on the new pin

Two regressions, both hit every ordinary llama-cpp model and neither was
caught locally because every test on this branch loaded a TTS model.

The first is a null dereference. server_slot::tts_ctx::reset() called
mtmd_helper_gen_audio_reset() unconditionally, but the gen-audio pipeline
is only allocated for models carrying a gen-audio mmproj, and upstream's
implementation reads ctx->pipeline before null-checking anything. Since
server_slot::reset() runs during slot initialization for every model, any
non-TTS model segfaulted the backend the moment it loaded. Guard the call
on the is_supported() predicate already defined beside it, and keep the
plain field resets unconditional.

The second is unrelated to TTS and came in with the pin bump.
PredictOptions.Penalty is a bare proto float, so a caller that names no
repetition penalty sends 0 rather than omitting the field. Since
9de0fcf2b, common_sampler_init() rejects a non-positive penalty_repeat
outright because it would divide logits by zero, turning every such
request into "Failed to initialize samplers". Treat 0 as unset and leave
llama.cpp's own neutral default in place.

Verified with the same suite CI runs, which is what caught both:
tests/e2e-backends passes 6 of 6 including the load and predict specs
that were red. Qwen3-TTS still synthesises on both paths, 24 kHz mono
16-bit WAV with exactly one RIFF header on the streamed output.

Assisted-by: Claude:claude-fable-5 [Claude Code]
Signed-off-by: Ettore Di Giacinto <mudler@localai.io>

---------

Signed-off-by: Ettore Di Giacinto <mudler@localai.io>
Co-authored-by: Ettore Di Giacinto <mudler@localai.io>
satindergrewal pushed a commit to satindergrewal/llama.cpp that referenced this pull request Aug 12, 2026
…gml-org#26254)

* convert text model

* main model load ok

* convert encoder ok

* speaker encoder loading ok

* speaker enc graph

* adapt vocab for backbone (with some tricks)

* add suppress_tokens

* poc new mtmd gen api

* convert code_predictor to gguf

* load gen_code model ok

* add clip_encode

* wire up

* code gen cgraph init version

Co-authored-by: Pascal <admin@serveurperso.com>

* code2wav convert to gguf

* code2wav graph ok

* wire up in/out

* (wip) subgraph

* wire up

* wip, correct code2wav

* demo (to be removed)

* code2wav preserve kv between calls

* demo voice clone

* llama: add llama_model_get_tok_embd

* mtmd_helper_gen_audio API

* fix clamp cold prefix

Co-authored-by: Pascal <admin@serveurperso.com>

* fuse snake op

Co-authored-by: Pascal <admin@serveurperso.com>

* demo: use proper sampling

* update dev docs

* polymorphism helper

* revamp llama-tts binary

* update docs

* fix compile

* fix lint

* nits

* add guide + docs

* more timings info

* clean up code comments

* security fixes

* update docs

* use ggml_build_forward_select, clean up comments

* fix ci

* use ISO 639-1 language code

* rename CODE2WAV --> GEN_WAV, update docs

* clean up

* clean up tts.cpp

* add seq_id

* add step_prompt()

* mtmd_helper_model_can_chat

* clean up comments

---------

Co-authored-by: Pascal <admin@serveurperso.com>
brittlewis12 pushed a commit to brittlewis12/llama.cpp that referenced this pull request Aug 17, 2026
…gml-org#26254)

* convert text model

* main model load ok

* convert encoder ok

* speaker encoder loading ok

* speaker enc graph

* adapt vocab for backbone (with some tricks)

* add suppress_tokens

* poc new mtmd gen api

* convert code_predictor to gguf

* load gen_code model ok

* add clip_encode

* wire up

* code gen cgraph init version

Co-authored-by: Pascal <admin@serveurperso.com>

* code2wav convert to gguf

* code2wav graph ok

* wire up in/out

* (wip) subgraph

* wire up

* wip, correct code2wav

* demo (to be removed)

* code2wav preserve kv between calls

* demo voice clone

* llama: add llama_model_get_tok_embd

* mtmd_helper_gen_audio API

* fix clamp cold prefix

Co-authored-by: Pascal <admin@serveurperso.com>

* fuse snake op

Co-authored-by: Pascal <admin@serveurperso.com>

* demo: use proper sampling

* update dev docs

* polymorphism helper

* revamp llama-tts binary

* update docs

* fix compile

* fix lint

* nits

* add guide + docs

* more timings info

* clean up code comments

* security fixes

* update docs

* use ggml_build_forward_select, clean up comments

* fix ci

* use ISO 639-1 language code

* rename CODE2WAV --> GEN_WAV, update docs

* clean up

* clean up tts.cpp

* add seq_id

* add step_prompt()

* mtmd_helper_model_can_chat

* clean up comments

---------

Co-authored-by: Pascal <admin@serveurperso.com>
truecharts-admin added a commit to trueforge-org/truecharts that referenced this pull request Aug 20, 2026
#51645)

> ℹ️ **Note**
> 
> This PR body was truncated due to platform limits.

This PR contains the following updates:

| Package | Update | Change |
|---|---|---|
|
[docker.io/localai/localai](https://redirect.github.com/mudler/LocalAI)
| minor | `df29190` → `d78cd11` |

---

> [!WARNING]
> Some dependencies could not be looked up. Check the [Dependency
Dashboard](../issues/18710) for more information.

Add the preset `:preserveSemverRanges` to your config if you don't want
to pin your dependencies.

---

### Release Notes

<details>
<summary>mudler/LocalAI (docker.io/localai/localai)</summary>

###
[`v4.9.0`](https://redirect.github.com/mudler/LocalAI/releases/tag/v4.9.0)

[Compare
Source](https://redirect.github.com/mudler/LocalAI/compare/v4.8.2...v4.9.0)

### 🎉 LocalAI 4.9.0 Release! 🚀

<h1 align="center">
  <br>
<img height="300"
src="https://raw.githubusercontent.com/mudler/LocalAI/refs/heads/master/core/http/static/logo.png">
  <br>
  <br>
</h1>

LocalAI 4.9.0 is out!

Thirteen days and 146 pull requests, spent on the parts of LocalAI you
touch every day rather than on new engines. Authentication is now
deny-by-default, chat gained end-to-end context compression, models and
backends each have one canonical page instead of three, and `vllm-cpp`
grew a video modality serving MiniMax-H3 with a real audio track.

**Highlights:**

- 🔐 **Authentication is deny-by-default** - every HTTP route requires
credentials unless it appears in an explicit public registry. This
closes a class of bypass in which unprefixed aliases such as
`/moderations`, `/models`, `/backends` and `/mcp/chat/completions` fell
outside the old protected-prefix list. Reported by [Naor
Yaacov](https://www.linkedin.com/in/naor-yaacov/).
- 🗜️ **Chat context compression** - opt-in per model, older complete
turns are compressed through a LocalAI model before inference,
preserving system prompts, the newest messages and whole tool-call
units. Ratio and duration come back as response metadata and metrics.
- 🖥️ **One page per resource** - `/app/models` now owns Explore and
Installed, `/app/backends` owns Catalog and Installed, and the nested
Host view is gone. Old `/app/manage` bookmarks still work.
- 🎬 **MiniMax-H3 video generation** - `vllm-cpp` opens a second engine
handle for the H3 checkpoint set and renders video and audio jointly, so
the MP4 arrives with a real AAC track. Ask for speech in the prompt and
the model lip-syncs it.
- 🗣️ **Qwen3-TTS on llama.cpp** - text-to-speech on the full accelerator
matrix already shipped for text generation (CUDA, ROCm, SYCL, Vulkan,
Metal, L4T), using upstream's own GGUF conversion.
- 🧭 **KNN as a first-class router** - similarity-weighted voting over a
curated, persisted corpus of labelled prompts. No classifier model, and
a prompt unlike anything labelled is treated as undecidable rather than
guessed.
- 📊 **Global admission control and live backend traces** - process-wide
HTTP admission bounds, in-flight backend operations are represented
while they run, and the UI links straight to their logs.
- 🕵️ **Reversible PII pseudonyms** - masked values become request-scoped
deterministic pseudonyms (`EMAIL_001`) and are restored if the backend
echoes them, across JSON and SSE tokens split over writes.
- 📦 **Parallel Hugging Face downloads** - snapshot materialization runs
up to N whole-file transfers at once, so a repository split into many
shards stops spending its wall clock in per-file latency.
- 🖧 **Cold model loads are durable jobs** - the per-model advisory lock
no longer spans a multi-GB transfer, which had made a 35.7 GB load look
permanently broken from the operator's seat while staging progressed
normally underneath.
- 🎮 **vllm-cpp covers the cards you own** - CUDA builds went from one or
two architectures to eight on amd64 and five on arm64, picking up A100,
L4, 4090, H100/H200, B200, Jetson Orin and Jetson Thor.

Plus a single shared WebRTC UDP port for Realtime, Metal actually
enabled in the macOS Stable Diffusion and Parakeet builds, backend crash
diagnostics at the default log level, and Portuguese (Brazil) and
Indonesian UI translations.

<p align="center">
<img width="2880" height="1800" alt="ui-models-explore"
src="https://github.com/user-attachments/assets/7dd68a4a-e707-43e6-af1a-dcc3922469eb"
/>
<br><em>Models and backends each get one canonical page, with Explore
and Installed as views rather than separate destinations.</em>
</p>

***

#### 📊 This release in numbers

|                      |                                    |
| -------------------- | ---------------------------------- |
| Pull requests merged | **146**                            |
| Commits              | 148                                |
| Files changed        | 420 (**+27,788** / -4,670)         |
| Development window   | 13 days (2026-08-07 to 2026-08-20) |
| Human contributors   | 10, of whom **3 first-time**       |
| Gallery entries      | 1,622 to **1,707** (+85)           |

Where the work landed:

| Area       | Change                            |
| ---------- | --------------------------------- |
| `core/`    | +17,193 / -4,187 across 279 files |
| `gallery/` | +3,808 / -62                      |
| `backend/` | +3,298 / -177 across 60 files     |
| `pkg/`     | +1,180 / -113 across 31 files     |
| `swagger/` | +1,095 / -5                       |
| `docs/`    | +952 / -98 across 23 files        |

***

#### 📌 TL;DR

| Area | Summary |
| ------------------------------ |
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
|
| 🔐 **Auth by default** | Every method and path requires credentials
unless listed in an explicit public registry. API instructions, Swagger
GETs, the LocalAI well-known document, and the health, login, OAuth,
SPA, asset, branding and node registration-token bootstrap flows stay
public. **Migration:** with database auth or legacy API keys configured,
`/version` and generated audio, image, video and 3D URLs now require
credentials. Embedded deployments can add narrow prefixes through
`ApplicationConfig.PathWithoutAuth`, and the legacy GET exemption flags
remain as explicit compatibility overrides. |
| 🗜️ **Context compression** | Opt-in per-model `compression` config.
Older complete turns are compressed through a configured LocalAI model
before inference, after PII filtering and Assistant/MCP prompt injection
(including later MCP iterations). Leading system/developer prompts, the
newest messages and complete tool-call/result units are preserved; tool
schemas and completion headroom are accounted for with a conservative
offline token bound. Metadata rides non-streaming responses and
streaming usage trailers, with event, ratio and duration metrics
exported. Disabled by default; cloud-proxy passthrough is rejected
(translate mode works). |
| 🖥️ **Unified lifecycle UI** | `/app/models` owns **Explore** and
**Installed** with URL-backed search, state and selection;
`/app/backends` owns **Catalog** and **Installed** while keeping
variants, development builds and target-node scope. Explore offers
capability-aware Open and Manage installation; destructive model
controls stay in Installed. Operate Overview shows host capacity from
its shared summary poller, and the nested Host destination is removed.
`/app/manage` redirects while preserving legacy query state. No API
change. |
| 📥 **Import form rebuild** | The import page moves to `page--medium`
with a work column and the format reference beside it rather than behind
a closed chevron. The source field is the hero and carries its own
Import button, which removes the `aria-hidden` submit that existed only
because the real action sat outside the `<form>`. Simple and Advanced
modes are gone (about 80% the same surface); the real distinction, a
source or YAML, is now two tabs. Also fixes two class bugs: a primary
button with no `className` at all falling through to browser chrome, and
`class="btn btn-primary fas fa-save fa-upload"` setting Font Awesome as
the button's own font while two icons fought over one `::before`. |
| 🎬 **MiniMax-H3 video** | `vllm-cpp` over vllm.cpp ABI v12. A second
engine handle loads the H3 checkpoint set (the DiT is
`parameters.model`, the text encoder and two VAEs are named in
`options:`), and `GenerateVideo` renders video and audio jointly into an
MP4 with a real AAC track. The DiT partition is declared, not detected:
community quantizations strip the release metadata and the FL2VA and
Ref2VA DiTs are byte-structurally identical, so
`checkPartitionConditioning` refuses a reference-conditioned FL2VA
request before the engine runs (it would otherwise render for hours and
return a coloured lattice). ffmpeg comes from the host: libvllm composes
the mux argv and spawns nothing. New gallery entry
`minimax-h3-fl2va-q4`. |
| 🗣️ **Qwen3-TTS on llama-cpp** | TTS through the `llama-cpp` backend on
CUDA, ROCm, SYCL, Vulkan, Metal and L4T, using upstream's GGUF
conversion. Implemented as a slot-based `SERVER_TASK_TYPE_TTS` task,
which is the concurrency-safe integration given that `server_context`
owns the `llama_context` and runs the slot scheduler on its own thread.
Carries the still-draft upstream server hunks as
`patches/0002-add-server-task-type-tts.patch` (delete on merge of
[ggml-org/llama.cpp#26603](https://redirect.github.com/ggml-org/llama.cpp/issues/26603)).
Gallery: `qwen3-tts-llamacpp` and `qwen3-tts-llamacpp-q4`. The existing
`qwen3-tts-cpp` backend is untouched and remains a separate path. |
| 🧭 **KNN routing** | `classifier: knn` routes by similarity-weighted
voting over labelled example prompts, so no classifier model is needed
and label knowledge lives in a corpus you seed and curate. Entries below
`knn.similarity_threshold` cannot vote; when none clears it the router
takes the fallback, and `nearest_similarity` is recorded on decisions
and fallbacks alike. One JSONL file per router under `<data
path>/router-corpus` is the source of truth, with the in-memory index
rebuilt at classifier build time and entries re-embedded when the
embedding model changed. Corpus input is API-only by design: `POST
/api/router/{name}/corpus`, `GET .../corpus/stats` (label counts only,
texts are never returned), `DELETE .../corpus`, admin-gated and exposed
as MCP tools. |
| 📊 **Admission and traces** | Process-wide HTTP admission control,
bounding what was previously only per backend. Backend operations are
represented while in flight, and running backend traces surface in the
UI with immediate log links. |
| 🕵️ **PII pseudonyms** | Opt-in `pii.reverse_in_response`. Masked
request values become unique deterministic pseudonyms within the request
(`EMAIL_001`, `EMAIL_002`) and are restored if the backend returns them,
including SSE tokens split across response writes. Substitution maps are
request-local and never persisted. Irreversible `[REDACTED:...]` remains
the default. |
| 📦 **Parallel HF downloads** | `DownloadFilesWithConcurrency` runs up
to N whole-file transfers through an `errgroup` with `SetLimit`. Single
files are never split, so `.partial` resume and per-file SHA
verification are untouched, and the two non-artifact callers keep
sequential ordering and fail-fast behaviour through a limit-of-1
wrapper. `completedBytes` became an `atomic.Int64` (the race detector
reported three races otherwise) and the caller's status callback stays
serialized. |
| 📞 **Realtime WebRTC port** | `--web-rtc-udp-port` /
`LOCALAI_WEBRTC_UDP_PORT` reuses one Pion ICE UDP mux across Realtime
calls, with bind failures surfaced through signaling and
container/firewall setup documented. The follow-up fix keeps
`LOCALAI_WEBRTC_ICE_INTERFACES` effective when a fixed port is set,
which had been silently ignored: a wildcard mux made pion enumerate
every interface itself, handing browsers unroutable `172.x` candidates
that dropped once ICE consent checks failed. |
| 🖧 **Durable cold loads** | The per-model advisory lock is a dedup
decision measured in milliseconds, not a transfer's lifetime. Cold loads
now run as durable jobs instead of holding it across backend install,
multi-GB staging and checkpoint load, and `WithLockCtx` now defends
against `statement_timeout` as well as `lock_timeout` (both abort the
same blocking `pg_advisory_lock`, only the latter was overridden). |
| 🎮 **vllm-cpp CUDA coverage** | amd64 goes from `120a;121a` to
`80;86;89;90a;100a;103a;120a;121a`, arm64 from `121a` to
`87;90a;100a;110;121a`, split by where the silicon exists. An unlisted
card did not run slower, it died at the first request with `no kernel
image is available for execution on the device`, long after install
reported success. The CUDA 13 guard now covers both branches, and
Triton-AOT stays on. |
| 🌍 **Two new languages** | Portuguese (Brazil), a complete 14-namespace
translation at full key parity with `en/`, and Indonesian for the admin,
media and navigation surfaces. |
| 🧠 **Models** | 85 new gallery entries: Qwen3.8 (9B, 27B, Ridge and
small variants), Gemma 4 agentic and Scotoma 2, DeepSeek V4 Pro 0813,
Ling 3.0 Flash, Nemotron 3.5 Lightning 30B, Tess 4 27B, Ornith 1.0 and
1.5 9B, LFM2.5 230M and VL 1.6B, HunyuanOCR and OvisOCR2, Higgs Audio v3
TTS, MiniMax-H3 Ref2VA, plus vllm.cpp text-generation entries and a
first Carbon genomics family. |

***

#### 🚀 New Features & Major Enhancements

##### 🔐 Authentication now denies by default

The previous classifier gated selected API-style paths by prefix.
Anything whose path was not on that list was public, which meant
unprefixed aliases (`/mcp/chat/completions`, `/moderations`, `/models`,
`/backends`, `/import-model`) could bypass global authentication, and
any newly registered route inherited the same weakness by default.

The middleware is now method-aware and denies by default: a route is
public only if its method and path appear in an explicit public
registry. What stays public is the set required to bootstrap and to be
discoverable: API instructions, Swagger GET routes, the LocalAI
well-known document, and the health, login, OAuth, SPA, asset, branding
and node registration-token flows. Whole-router coverage is asserted in
tests, so a new route cannot become public by omission.

**Migration impact.** When database authentication or legacy API keys
are configured, `/version` and generated audio, image, video and 3D URLs
now require credentials. Embedded deployments can still add narrow
prefixes via `ApplicationConfig.PathWithoutAuth`, and the legacy GET
exemption flags remain available as explicit compatibility overrides.

Thanks to [Naor Yaacov](https://www.linkedin.com/in/naor-yaacov/) for
reporting this class of authentication bypass.

> 🔗 PRs:
[#&#8203;11602](https://redirect.github.com/mudler/LocalAI/issues/11602)

##### 🗜️ End-to-end context compression

A long conversation eventually stops fitting. Compression is opt-in per
model, and when enabled it compresses older complete turns through a
configured LocalAI model before inference rather than truncating them
away.

What it will not touch: leading system and developer safety prompts, the
newest messages, and complete tool-call/result units, which are kept
whole so a compressed history never leaves a call without its result. It
runs after PII filtering and after Assistant/MCP prompt injection,
including on later MCP iterations, so what gets compressed is the prompt
that would actually have been sent. Tool schemas and completion headroom
are accounted for with a conservative offline token bound.

Compression metadata is exposed in non-streaming responses and in
streaming usage trailers, and compression events, ratios and durations
are exported as metrics. Cloud-proxy passthrough configurations reject
compression because LocalAI cannot safely rewrite an opaque provider
payload; translate mode is supported. A late failure in an
already-started stream is returned as an in-band SSE error followed by
`[DONE]`.

> 🔗 PRs:
[#&#8203;11556](https://redirect.github.com/mudler/LocalAI/issues/11556)

##### 🖥️ One canonical page per resource

Models had a gallery and a separate Host management surface. Backends
had a nested Host view for installed binaries. Between them it was not
obvious where a resource lived, and the common lifecycle actions sat one
level deeper than they needed to.

Each resource now has one page. `/app/models` owns **Explore** and
**Installed**, `/app/backends` owns **Catalog** and **Installed**, both
with URL-backed search, state and selection, and backends keep their
variants, development builds and target-node scope. Explore presents
capability-aware **Open** and **Manage installation** actions while
destructive model controls stay in Installed. Operate Overview reads
host capacity from its shared summary poller, and the nested Host
destination is removed. Narrow list/detail views restore focus to the
originating row when you come back from a detail view.

Nothing in the API changed, and existing `/app/manage` bookmarks keep
working through a replace redirect that preserves legacy model and
backend query state.

<p align="center">
<img width="2880" height="1800" alt="ui-backends-catalog"
src="https://github.com/user-attachments/assets/c18e9272-70bd-4875-bda0-319bea616637"
/>
<br><em>Catalog and Installed as views of one page, with target-node
scope intact.</em>
</p>

> 🔗 PRs:
[#&#8203;11548](https://redirect.github.com/mudler/LocalAI/issues/11548)

##### 📥 The import form, rebuilt

The import page had taken the new palette but kept its old layout: a
760px column with the primary action detached from the form it submits.
Two of its problems were outright bugs.

`ImportModel.jsx:808` carried no `className` at all, so the page's
single most important control fell through to the user-agent button,
with system chrome, system font, the wrong radius and no design-system
focus ring. Next to it, `class="btn btn-primary fas fa-save fa-upload"`
set Font Awesome as the button's own font family, which its label text
inherited, while `fa-save` and `fa-upload` fought over one `::before`.

The layout moves to `page--medium` with a work column and the format
reference beside it, since that reference answers the only question a
first-time admin has and used to sit behind a chevron that was closed by
default. Below 1024px it becomes a disclosure instead of disappearing.
The source field is the hero, monospace because it holds something you
paste, and it carries its own Import button, which removes the
`aria-hidden` submit that existed only to compensate for the real action
sitting outside the `<form>`. Simple and Advanced modes are gone: they
were about 80% the same surface, and the overlap cost a mode switch, a
localStorage key and a three-button Keep/Discard/Cancel dialog whose
only job was protecting state the switch would have hidden. What
genuinely differs is the kind of input, which is now two tabs: a source,
or YAML. The size and VRAM estimate reports under the field that
produced it instead of as a banner above the page header.

A follow-up swept the same class of bug across the rest of the UI: eight
header controls on seven pages had two or three elements' classes
collapsed into one string.

<p align="center">
<img width="2880" height="1800" alt="ui-import-model"
src="https://github.com/user-attachments/assets/d8abd9d6-81fa-4ef7-8062-4bcf3836a27a"
/>
<br><em>One form, a collapsible options panel, and the format reference
where you can read it.</em>
</p>

> 🔗 PRs:
[#&#8203;11461](https://redirect.github.com/mudler/LocalAI/issues/11461),
[#&#8203;11462](https://redirect.github.com/mudler/LocalAI/issues/11462),
[#&#8203;11488](https://redirect.github.com/mudler/LocalAI/issues/11488)

##### 🎬 MiniMax-H3: video and audio, jointly

vllm.cpp's stable C ABI grew a video slice at v12, and `vllm-cpp` now
serves two things. Text generation is unchanged. When a model config
declares the H3 checkpoint set, `Load` opens a video engine instead and
`GenerateVideo` renders a clip through LocalAI's existing `/video`
endpoint, with **video and audio generated jointly**, so the MP4 comes
back with a real AAC track rather than silent. Ask for speech in the
prompt and the model lip-syncs it.

Three things shape the integration:

**The video engine is a second handle, not a mode of the first.** H3 is
not a model directory. The DiT, the text encoder and two VAEs are
separate artifacts, and vllm.cpp's two loaders refuse each other's
checkpoints. `parameters.model` is the DiT; the rest of the set is named
in `options:`.

**The DiT partition is declared, not detected.** The FL2VA DiT serves
`t2va` and `fl2va`; `ref2va` is a different checkpoint. Community GGUF
and NVFP4 quantizations strip the release metadata and the two DiTs are
byte-structurally identical, so the engine refuses to generate until it
is told which it has. Handing reference conditioning to an FL2VA DiT
renders for hours and returns a coloured lattice over the frame, so
`checkPartitionConditioning` rejects that combination before the engine
is ever called.

**ffmpeg comes from the host.** libvllm writes the frames and the WAV
and composes the mux argv, then spawns nothing, which is a deliberate
upstream process boundary. The backend substitutes `argv[0]` and execs
it.

Gallery entries `minimax-h3-fl2va-q4` (Q4\_K\_M FL2VA set) and
`minimax-h3-ref2va-q4` ship with it.

> 🔗 PRs:
[#&#8203;11424](https://redirect.github.com/mudler/LocalAI/issues/11424),
[#&#8203;11439](https://redirect.github.com/mudler/LocalAI/issues/11439)

##### 🗣️ Qwen3-TTS through the llama.cpp backend

Qwen3-TTS now runs on the `llama-cpp` backend, which means
text-to-speech on the same accelerator matrix already shipped for text
generation (CUDA, ROCm, SYCL, Vulkan, Metal, L4T) using upstream's own
GGUF conversion.

`grpc-server.cpp` is an adapter over llama.cpp's shared
`server_context`, which owns the `llama_context` and runs the slot
scheduler on its own thread, so a gRPC handler driving the gen-audio
loop itself would race that scheduler. Making TTS a slot-based
`SERVER_TASK_TYPE_TTS` task is the concurrency-safe integration,
following the `0001-add-server-task-type-score.patch` precedent already
in the tree. llama.cpp merged Qwen3-TTS in
[ggml-org/llama.cpp#26254](https://redirect.github.com/ggml-org/llama.cpp/issues/26254);
the server plumbing in
[#&#8203;26603](https://redirect.github.com/mudler/LocalAI/issues/26603)
is still a draft, so it is carried as
`patches/0002-add-server-task-type-tts.patch` and should be deleted once
that merges. `disable-tts-task.sh` keeps turboquant and bonsai
compiling, since they copy `grpc-server.cpp` into forks without our
patches.

Both paths were verified end to end on CPU returning valid 24 kHz mono
16-bit WAV containing real speech, measured rather than eyeballed.
Gallery entries: `qwen3-tts-llamacpp` and `qwen3-tts-llamacpp-q4`.

The existing `qwen3-tts-cpp` backend over qwentts.cpp is untouched. This
is a second, independent path, not a replacement.

> 🔗 PRs:
[#&#8203;11392](https://redirect.github.com/mudler/LocalAI/issues/11392)

##### 🧭 KNN as a first-class router

`classifier: knn` promotes KNN search from a cache for the classifier to
a primary request router. Unlike `score` or `colbert` it needs no
classifier model: label knowledge lives in a corpus of labelled example
prompts that you seed and curate through the admin API, so routing
decisions are deterministic, auditable, and grounded in graded
experience rather than a model's opinion.

There is an explicit epistemic gate. Corpus entries below
`knn.similarity_threshold` cannot vote, and when none clears it the
classifier activates no labels and the router takes the fallback: a
prompt unlike all labelled experience is treated as undecidable, not
guessed. Decisions record `nearest_similarity`, on fallback rows too, so
you can see how far the nearest labelled experience actually was, and
the Routing tab explains out-of-corpus fallbacks and shows per-label
corpus counts.

Persistence is one JSONL file per router under `<data
path>/router-corpus` holding text, labels, vector and embedder
fingerprint. That file is the source of truth; the local-store index is
rebuilt from it at classifier build time and stays a pure in-memory
index, and entries recorded under a different embedding model re-embed
on load. This also corrects the docs' claim that local-store collections
persist: the embedding cache never survived restarts and still does not,
while the corpus does.

Corpus input is API-only by design, since entries may contain example
user content: `POST /api/router/{name}/corpus` seeds (labels validated
against declared policies, embedded server-side, indexed immediately),
`GET .../corpus/stats` inspects and returns label counts only (entry
texts are never returned by any surface), and `DELETE .../corpus` wipes.
All admin-gated like the sibling router endpoints and exposed as MCP
tools.

> 🔗 PRs:
[#&#8203;10652](https://redirect.github.com/mudler/LocalAI/issues/10652)

##### 📊 Global admission control and running backend traces

Admission control existed per backend, which left nothing bounding the
process as a whole, and the traces list could grow without limit. HTTP
admission is now bounded process-wide.

Alongside it, backend operations are represented while they are still in
flight rather than only once they finish, and running backend traces
surface in the UI with immediate links to their logs, so an operation
that is taking too long is something you can look at instead of
something you wait out.

<p align="center">
<img width="2880" height="1800" alt="ui-activity-inflight"
src="https://github.com/user-attachments/assets/fe0acd52-3688-4c20-ae1d-cef9f8b2825e"
/>
<br><em>Operations are represented while they are still running, not
only once they finish.</em>
</p>

> 🔗 PRs:
[#&#8203;11560](https://redirect.github.com/mudler/LocalAI/issues/11560)

##### 🕵️ PII pseudonyms that survive the round trip

The PII middleware masked values irreversibly, which is right for logs
and wrong for a conversation: a model that is handed `[REDACTED:EMAIL]`
twice cannot tell whether it saw one address or two, and anything it
says about them comes back unusable.

Opt-in `pii.reverse_in_response` turns masked request values into unique
deterministic pseudonyms within the request (`EMAIL_001`, `EMAIL_002`)
and restores them if the backend returns them. Restoration handles
normal JSON and SSE tokens split across response writes. Substitution
maps stay request-local and are never persisted. Irreversible
`[REDACTED:...]` remains the default.

> 🔗 PRs:
[#&#8203;11272](https://redirect.github.com/mudler/LocalAI/issues/11272)

##### 📦 Bounded parallel Hugging Face downloads

Snapshot materialization fetched every file through the sequential
executor, so a repository split into many shards spent most of its wall
clock in per-file request latency rather than moving bytes.

`DownloadFilesWithConcurrency` now runs up to N whole-file transfers at
once through an `errgroup` with `SetLimit`. Only whole files run in
parallel: a single file is never split, so the `.partial` resume
machinery and the per-file SHA check are untouched. The two non-artifact
callers (`core/gallery/models.go` and
`core/config/model_config_loader.go`) keep exactly their previous
behaviour through a wrapper passing a limit of 1, so tasks still run in
slice order and the first failure still returns before any later task
starts.

Two consequences of the parallel path are worth knowing:
`completedBytes` is now an `atomic.Int64`, which is not a precaution
(with a plain `int64` the race detector reports three races), and the
caller's status callback is serialized to preserve the guarantee the
sequential path gave it implicitly. `AfterDownload` is deliberately not
serialized, because it does the verify-and-promote work that the
parallelism exists to overlap.

> 🔗 PRs:
[#&#8203;11162](https://redirect.github.com/mudler/LocalAI/issues/11162)

##### 📞 Realtime WebRTC on one UDP port

`--web-rtc-udp-port` / `LOCALAI_WEBRTC_UDP_PORT` reuses a single Pion
ICE UDP mux across Realtime WebRTC calls, so a container or firewall
needs one rule rather than a range. UDP bind failures are surfaced
through signaling instead of failing opaquely, and the container and
firewall setup is documented.

A follow-up closed the gap it opened. `LOCALAI_WEBRTC_ICE_INTERFACES`
was silently ignored whenever a fixed UDP port was set, which is exactly
the combination an operator reaches for: pinning a port to write a
firewall rule and restricting interfaces to keep unreachable `docker0`
and `veth` addresses out of the candidate list usually go together. A
wildcard mux made pion derive candidate addresses by enumerating
interfaces itself with a nil filter, so on a host with docker bridges
the browser received `172.18.0.1`, `172.17.0.1`, `10.10.10.1` and
friends, connected on a good pair, then dropped when ICE consent checks
failed on the others.

> 🔗 PRs:
[#&#8203;11436](https://redirect.github.com/mudler/LocalAI/issues/11436),
[#&#8203;11466](https://redirect.github.com/mudler/LocalAI/issues/11466)

##### 🖧 Cold model loads run as durable jobs

On a two-replica frontend, loading a 35.7 GB GGUF onto a newly added
Jetson Thor worker made the model permanently unloadable from the
operator's seat, while staging was in fact progressing normally
underneath. Replica A held the per-model advisory lock through roughly
twenty minutes of transfer; replica B blocked on `pg_advisory_lock` for
the same model and was killed at 60s by the role's `statement_timeout`,
and every UI retry reproduced it.

Two defects sat behind that one symptom. The lock's lifetime was the
transfer's lifetime, with `Route` wrapping backend install, multi-GB
staging and checkpoint load in `advisorylock.WithLockCtx`, which turns a
millisecond dedup decision into a cluster-wide outage for that model.
And `WithLockCtx` overrode `lock_timeout` but not `statement_timeout`,
though both abort the same blocking call.

Cold loads now run as durable jobs, so the lock is held only for the
decision it exists to make.

> 🔗 PRs:
[#&#8203;11514](https://redirect.github.com/mudler/LocalAI/issues/11514)

##### 🎮 vllm-cpp builds for the cards people own

The `vllm-cpp` CUDA images were built for Blackwell only: `120a;121a` on
amd64 and `121a` alone on arm64, out of the ten architectures vllm.cpp's
own release archive builds.

What makes it worth calling out is the failure mode. An unlisted card is
not slower, it dies at the first request with `no kernel image is
available for execution on the device`, long after `local-ai backends
install vllm-cpp` reported success. This was found on a Jetson Thor node
that had the backend installed and could serve nothing.

|       | before      | after                              |
| ----- | ----------- | ---------------------------------- |
| amd64 | `120a;121a` | `80;86;89;90a;100a;103a;120a;121a` |
| arm64 | `121a`      | `87;90a;100a;110;121a`             |

The split follows where the silicon exists: Jetson (`87` Orin, `110`
Thor) is arm64-only, desktop `120a` is amd64-only, and `90a`/`100a` are
on both because of the SBSA parts. That adds A100, A10/3090, L4/4090/RTX
6000 Ada, H100/H200, B200, B300, Jetson Orin and Jetson Thor. The CUDA
13 guard now covers both branches rather than amd64 alone, and
Triton-AOT stays on.

> 🔗 PRs:
[#&#8203;11512](https://redirect.github.com/mudler/LocalAI/issues/11512)

##### 🌍 Portuguese (Brazil) and Indonesian

A complete `pt-BR` translation of the WebUI: 14 namespaces at full key
parity with `en/` including `modelEditor.json`, with every i18next
interpolation variable and `_one`/`_other` plural key preserved.
Indonesian covers the admin, media and navigation strings. Both keep
brand, model and technical identifiers untranslated, matching existing
locale conventions.

> 🔗 PRs:
[#&#8203;11427](https://redirect.github.com/mudler/LocalAI/issues/11427),
[#&#8203;11493](https://redirect.github.com/mudler/LocalAI/issues/11493)

##### 🧰 Smaller features worth knowing about

- **Metal is actually on in two macOS backends.** `stablediffusion-ggml`
gated its Metal flags on an `OS=Darwin` variable the runner never
defines, so the Darwin workflow's `BUILD_TYPE=metal` produced a build
without `GGML_METAL_EMBED_LIBRARY=ON`, shipping a runtime source path
that failed to expose `kernel_mul_mv_ext_bf16_f32_r1_5`. `parakeet-cpp`
never forwarded `BUILD_TYPE=metal` to `PARAKEET_GGML_METAL` at all: on
an M1 Air the same five-minute sample went from 82.57s to 50.18s with
byte-identical output. A dry-run build-contract test now guards the
Stable Diffusion flags.
- **Backend crashes say why.** An unexpected runtime exit logged its
stderr only at debug level, so at default log level operators saw an
exit code and nothing else. The final non-empty stderr line now rides
the unexpected-exit warning, and a failed gRPC readiness preserves the
process exit code plus the last stderr diagnostic bounded to 4 KiB.
- **MCP servers stay visible when they fail.** Model-level MCP servers
disappeared from Chat whenever connection setup or tool discovery
failed, hiding container DNS, routing and VPN reachability problems.
They now stay listed as disabled error rows with a per-server error,
discovery retries while safely closing partially created sessions, and
the model editor documents the expected `mcp.remote` and `mcp.stdio`
formats.
- **Checksum mismatches retry.** A post-download SHA mismatch is now
treated as a transient transfer failure: LocalAI removes the mismatched
partial and lets the bounded planner retry a stale or corrupted CDN
response, while still refusing unverified bytes. Direct
`URI.DownloadFile` callers still see the mismatch immediately.
- **Audio transform rejects the wrong contract.** The audio-transform
WebSocket accepted `realtime_audio` models and opened the frame-based
`AudioTransformStream` RPC, which failed after the handshake with
`NotImplementedError` for any-to-any models like `liquid-audio`. The use
case is now validated before the backend loads, and any-to-any callers
are pointed at the OpenAI Realtime API.
- **Invalid preload JSON names itself.** `PRELOAD_MODELS` /
`--preload-models` now identifies which input was invalid and rejects
non-array top-level values (including booleans and `null`) with the
expected shape in the error.

> 🔗 PRs:
[#&#8203;11531](https://redirect.github.com/mudler/LocalAI/issues/11531),
[#&#8203;11492](https://redirect.github.com/mudler/LocalAI/issues/11492),
[#&#8203;11532](https://redirect.github.com/mudler/LocalAI/issues/11532),
[#&#8203;11447](https://redirect.github.com/mudler/LocalAI/issues/11447),
[#&#8203;11495](https://redirect.github.com/mudler/LocalAI/issues/11495),
[#&#8203;11536](https://redirect.github.com/mudler/LocalAI/issues/11536),
[#&#8203;11565](https://redirect.github.com/mudler/LocalAI/issues/11565),
[#&#8203;11434](https://redirect.github.com/mudler/LocalAI/issues/11434)

***

#### 🐛 Bug Fixes (recap)

- `fix(auth)`: protect HTTP routes by default -
[#&#8203;11602](https://redirect.github.com/mudler/LocalAI/issues/11602)
- `fix(distributed)`: run cold model loads as durable jobs instead of
holding the advisory lock -
[#&#8203;11514](https://redirect.github.com/mudler/LocalAI/issues/11514)
- `fix(vllm-cpp)`: build every CUDA architecture the platform can host -
[#&#8203;11512](https://redirect.github.com/mudler/LocalAI/issues/11512)
- `fix(realtime)`: keep the ICE interface allow-list working with a
fixed UDP port -
[#&#8203;11466](https://redirect.github.com/mudler/LocalAI/issues/11466)
- `fix(stablediffusion)`: embed Metal library -
[#&#8203;11531](https://redirect.github.com/mudler/LocalAI/issues/11531)
- `fix(parakeet-cpp)`: enable Metal in macOS builds -
[#&#8203;11492](https://redirect.github.com/mudler/LocalAI/issues/11492)
- `fix(model)`: report backend crash diagnostics -
[#&#8203;11532](https://redirect.github.com/mudler/LocalAI/issues/11532)
- `fix(model)`: surface backend startup exits -
[#&#8203;11447](https://redirect.github.com/mudler/LocalAI/issues/11447)
- `fix(downloader)`: retry checksum mismatches -
[#&#8203;11536](https://redirect.github.com/mudler/LocalAI/issues/11536)
- `fix(audio)`: reject incompatible transform streams -
[#&#8203;11565](https://redirect.github.com/mudler/LocalAI/issues/11565)
- `fix(gallery)`: parse harmony output of gpt-oss-\* models correctly -
[#&#8203;11518](https://redirect.github.com/mudler/LocalAI/issues/11518)
- `fix(gallery)`: identify invalid preload JSON -
[#&#8203;11434](https://redirect.github.com/mudler/LocalAI/issues/11434)
- `fix(gallery)`: repair DeepSeek V4 fallback -
[#&#8203;11480](https://redirect.github.com/mudler/LocalAI/issues/11480)
- `fix(gallery)`: correct Higgs Audio v3 checksum -
[#&#8203;11459](https://redirect.github.com/mudler/LocalAI/issues/11459)
- `fix(fish-speech)`: preserve ROCm PyTorch -
[#&#8203;11568](https://redirect.github.com/mudler/LocalAI/issues/11568)
- `fix(vllm)`: align Intel basekit runtime to oneAPI 2025.3.2 -
[#&#8203;11437](https://redirect.github.com/mudler/LocalAI/issues/11437)
- `fix(kokoros)`: add the missing `upscale_image` stub to the Backend
trait impl -
[#&#8203;11414](https://redirect.github.com/mudler/LocalAI/issues/11414)
- `fix`: show MCP connection errors in the UI -
[#&#8203;11495](https://redirect.github.com/mudler/LocalAI/issues/11495)
- `fix(ui)`: unmerge the class strings that left buttons in browser
chrome -
[#&#8203;11462](https://redirect.github.com/mudler/LocalAI/issues/11462)
- `fix(ui)`: keep agent import action visible -
[#&#8203;11488](https://redirect.github.com/mudler/LocalAI/issues/11488)
- `fix`: wrap long TTS request text instead of widening the page -
[#&#8203;11576](https://redirect.github.com/mudler/LocalAI/issues/11576)

***

#### 🧠 Models

85 new gallery entries this cycle, taking the index from 1,622 to 1,707.

Text generation: Qwen3.8 in 9B, 27B, Ridge and small variants, Gemma 4
agentic and Gemma 4 Scotoma 2, DeepSeek V4 Pro 0813, Ling 3.0 Flash,
Nemotron 3.5 Lightning 30B, Tess 4 27B, Ornith 1.0 and 1.5 9B, Muse
Glimmer 30B, Grug 12B, BigBang v1, Genesis Hermes V7, TwIL-LM3, BTL-4
Compact, XYZ Aquila mini, North Mini Code, AREX Turbo, Fara1.5 4B,
MiniCPM5 1B Q8, Shieldstral 1.0 3B and UI-Mate 9B.

Vision and OCR: HunyuanOCR, OvisOCR2, LFM2.5 VL 1.6B, and LFM2.5 230M
alongside it.

Audio and video: Higgs Audio v3 TTS, the Qwen3-TTS llama.cpp entries,
and the MiniMax-H3 FL2VA and Ref2VA video sets.

Also a first Carbon genomics family, and text-generation entries for the
`vllm-cpp` backend.

> 🔗 PRs:
[#&#8203;11622](https://redirect.github.com/mudler/LocalAI/issues/11622),
[#&#8203;11603](https://redirect.github.com/mudler/LocalAI/issues/11603),
[#&#8203;11599](https://redirect.github.com/mudler/LocalAI/issues/11599),
[#&#8203;11598](https://redirect.github.com/mudler/LocalAI/issues/11598),
[#&#8203;11594](https://redirect.github.com/mudler/LocalAI/issues/11594),
[#&#8203;11584](https://redirect.github.com/mudler/LocalAI/issues/11584),
[#&#8203;11573](https://redirect.github.com/mudler/LocalAI/issues/11573),
[#&#8203;11571](https://redirect.github.com/mudler/LocalAI/issues/11571),
[#&#8203;11561](https://redirect.github.com/mudler/LocalAI/issues/11561),
[#&#8203;11559](https://redirect.github.com/mudler/LocalAI/issues/11559),
[#&#8203;11557](https://redirect.github.com/mudler/LocalAI/issues/11557),
[#&#8203;11552](https://redirect.github.com/mudler/LocalAI/issues/11552),
[#&#8203;11551](https://redirect.github.com/mudler/LocalAI/issues/11551),
[#&#8203;11549](https://redirect.github.com/mudler/LocalAI/issues/11549),
[#&#8203;11547](https://redirect.github.com/mudler/LocalAI/issues/11547),
[#&#8203;11540](https://redirect.github.com/mudler/LocalAI/issues/11540),
[#&#8203;11533](https://redirect.github.com/mudler/LocalAI/issues/11533),
[#&#8203;11526](https://redirect.github.com/mudler/LocalAI/issues/11526),
[#&#8203;11519](https://redirect.github.com/mudler/LocalAI/issues/11519),
[#&#8203;11511](https://redirect.github.com/mudler/LocalAI/issues/11511),
[#&#8203;11490](https://redirect.github.com/mudler/LocalAI/issues/11490),
[#&#8203;11479](https://redirect.github.com/mudler/LocalAI/issues/11479),
[#&#8203;11478](https://redirect.github.com/mudler/LocalAI/issues/11478),
[#&#8203;11477](https://redirect.github.com/mudler/LocalAI/issues/11477),
[#&#8203;11458](https://redirect.github.com/mudler/LocalAI/issues/11458),
[#&#8203;11456](https://redirect.github.com/mudler/LocalAI/issues/11456),
[#&#8203;11455](https://redirect.github.com/mudler/LocalAI/issues/11455),
[#&#8203;11449](https://redirect.github.com/mudler/LocalAI/issues/11449),
[#&#8203;11446](https://redirect.github.com/mudler/LocalAI/issues/11446),
[#&#8203;11443](https://redirect.github.com/mudler/LocalAI/issues/11443),
[#&#8203;11441](https://redirect.github.com/mudler/LocalAI/issues/11441),
[#&#8203;11439](https://redirect.github.com/mudler/LocalAI/issues/11439),
[#&#8203;11438](https://redirect.github.com/mudler/LocalAI/issues/11438),
[#&#8203;11435](https://redirect.github.com/mudler/LocalAI/issues/11435)

***

#### 👒 Dependencies

Submodule and pin bumps this cycle:

| Project | Bumps |
|
------------------------------------------------------------------------------------------------------------------
| ------ |
| CrispStrobe/CrispASR | 10 |
| ikawrakow/ik\_llama.cpp | 9 |
| vllm-metal (darwin) | 7 |
| mudler/vllm.cpp | 6 |
| 0xShug0/audio.cpp | 5 |
| ggml-org/llama.cpp | 4 |
| ggml-org/whisper.cpp | 3 |
| leejet/stable-diffusion.cpp | 2 |
| antirez/ds4, mudler/parakeet.cpp, mudler/depth-anything.cpp,
NVIDIA/NeMo-Speech.cpp, vllm-project/vllm cu130 wheel | 1 each |

Plus `golang.org/x/net` to v0.55.0, vllm 0.26.0 and transformers
>=5.15.0 in the Python backends, sentence-transformers 5.7.0, packaging
26.3, dompurify 3.4.13, the Kokoros source pin, inference defaults
refreshed from unsloth, and seven gallery checksum refreshes.

***

#### 📖 Documentation

New pages for context compression and `vllm-cpp`, a substantially
expanded middleware page covering the PII pseudonym and compression
surfaces, and a rewritten authentication page documenting the public and
protected route surfaces after the deny-by-default change.

Video generation gained the MiniMax-H3 setup, text-to-audio the
Qwen3-TTS llama.cpp path, distributed mode the durable cold-load
behaviour, MCP the configuration formats and container networking
implications, and Realtime the shared UDP port and firewall guidance.
Embeddings, model gallery, backends and API discovery all picked up
corrections, and the broken stars counter came out of the site.

> 🔗 PRs:
[#&#8203;11448](https://redirect.github.com/mudler/LocalAI/issues/11448),
[#&#8203;11581](https://redirect.github.com/mudler/LocalAI/issues/11581),
[#&#8203;11582](https://redirect.github.com/mudler/LocalAI/issues/11582),
[#&#8203;11415](https://redirect.github.com/mudler/LocalAI/issues/11415)

***

#### 🙌 New Contributors

- [@&#8203;tom-mi](https://redirect.github.com/tom-mi) made their first
contribution in
[#&#8203;11518](https://redirect.github.com/mudler/LocalAI/issues/11518)
- [@&#8203;kassane](https://redirect.github.com/kassane) made their
first contribution in
[#&#8203;11427](https://redirect.github.com/mudler/LocalAI/issues/11427)
- [@&#8203;fieryWaters](https://redirect.github.com/fieryWaters) made
their first contribution in
[#&#8203;11492](https://redirect.github.com/mudler/LocalAI/issues/11492)

Thanks also to [@&#8203;richiejp](https://redirect.github.com/richiejp),
[@&#8203;jimmykarily](https://redirect.github.com/jimmykarily),
[@&#8203;Dennisadira](https://redirect.github.com/Dennisadira),
[@&#8203;dedyf5](https://redirect.github.com/dedyf5),
[@&#8203;ALameLlama](https://redirect.github.com/ALameLlama) and
[@&#8203;walcz-de](https://redirect.github.com/walcz-de), and to [Naor
Yaacov](https://www.linkedin.com/in/naor-yaacov/) for the authentication
bypass report.

<!-- Release notes generated using configuration in .github/release.yml
at master -->

#### What's Changed

##### Breaking Changes 🛠

- fix(auth): protect HTTP routes by default by
[@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11602](https://redirect.github.com/mudler/LocalAI/pull/11602)

##### Bug fixes :bug:

- fix(kokoros): add missing `upscale_image` stub to Backend trait impl
by [@&#8203;mudler](https://redirect.github.com/mudler) with
[@&#8203;Copilot](https://redirect.github.com/Copilot) in
[#&#8203;11414](https://redirect.github.com/mudler/LocalAI/pull/11414)
- fix(gallery): identify invalid preload JSON by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11434](https://redirect.github.com/mudler/LocalAI/pull/11434)
- fix(vllm): align Intel basekit runtime by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11437](https://redirect.github.com/mudler/LocalAI/pull/11437)
- fix(model): surface backend startup exits by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11447](https://redirect.github.com/mudler/LocalAI/pull/11447)
- fix(ui): unmerge the class strings that left buttons in browser chrome
by [@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11462](https://redirect.github.com/mudler/LocalAI/pull/11462)
- fix(realtime): keep the ICE interface allow-list working with a fixed
UDP port by
[@&#8203;jimmykarily](https://redirect.github.com/jimmykarily) in
[#&#8203;11466](https://redirect.github.com/mudler/LocalAI/pull/11466)
- fix(parakeet-cpp): enable Metal in macOS builds by
[@&#8203;fieryWaters](https://redirect.github.com/fieryWaters) in
[#&#8203;11492](https://redirect.github.com/mudler/LocalAI/pull/11492)
- fix: Show MCP connection errors in the UI by
[@&#8203;richiejp](https://redirect.github.com/richiejp) in
[#&#8203;11495](https://redirect.github.com/mudler/LocalAI/pull/11495)
- fix(vllm-cpp): build every CUDA architecture the platform can host by
[@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11512](https://redirect.github.com/mudler/LocalAI/pull/11512)
- fix (gallery): Parse harmony output of gpt-oss-\* models correctly
([#&#8203;8037](https://redirect.github.com/mudler/LocalAI/issues/8037))
by [@&#8203;tom-mi](https://redirect.github.com/tom-mi) in
[#&#8203;11518](https://redirect.github.com/mudler/LocalAI/pull/11518)
- fix(ui): keep agent import action visible by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11488](https://redirect.github.com/mudler/LocalAI/pull/11488)
- fix(stablediffusion): embed Metal library by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11531](https://redirect.github.com/mudler/LocalAI/pull/11531)
- fix(model): report backend crash diagnostics by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11532](https://redirect.github.com/mudler/LocalAI/pull/11532)
- fix(distributed): run cold model loads as durable jobs instead of
holding the advisory lock by
[@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11514](https://redirect.github.com/mudler/LocalAI/pull/11514)
- fix(downloader): retry checksum mismatches by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11536](https://redirect.github.com/mudler/LocalAI/pull/11536)
- fix(fish-speech): preserve ROCm PyTorch by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11568](https://redirect.github.com/mudler/LocalAI/pull/11568)
- fix(audio): reject incompatible transform streams by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11565](https://redirect.github.com/mudler/LocalAI/pull/11565)
- fix: tts text wrap by
[@&#8203;ALameLlama](https://redirect.github.com/ALameLlama) in
[#&#8203;11576](https://redirect.github.com/mudler/LocalAI/pull/11576)

##### Exciting New Features 🎉

- feat(modelartifacts): support bounded parallel Hugging Face file
downloads by
[@&#8203;Dennisadira](https://redirect.github.com/Dennisadira) in
[#&#8203;11162](https://redirect.github.com/mudler/LocalAI/pull/11162)
- feat(i18n): add pt-BR translation by
[@&#8203;kassane](https://redirect.github.com/kassane) in
[#&#8203;11427](https://redirect.github.com/mudler/LocalAI/pull/11427)
- feat(vllm-cpp): serve MiniMax-H3 video+audio generation by
[@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11424](https://redirect.github.com/mudler/LocalAI/pull/11424)
- feat(pii): restore request-scoped pseudonyms by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11272](https://redirect.github.com/mudler/LocalAI/pull/11272)
- feat(llama-cpp): serve Qwen3-TTS through the llama.cpp backend by
[@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11392](https://redirect.github.com/mudler/LocalAI/pull/11392)
- feat(realtime): add shared WebRTC UDP port by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11436](https://redirect.github.com/mudler/LocalAI/pull/11436)
- feat(ui): rebuild the import form on the restyled design language by
[@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11461](https://redirect.github.com/mudler/LocalAI/pull/11461)
- i18n(id): translate admin, media, and nav UI strings to Indonesian by
[@&#8203;dedyf5](https://redirect.github.com/dedyf5) in
[#&#8203;11493](https://redirect.github.com/mudler/LocalAI/pull/11493)
- feat(ui): unify model and backend lifecycle by
[@&#8203;localai-bot](https://redirect.github.com/localai-bot) in
[#&#8203;11548](https://redirect.github.com/mudler/LocalAI/pull/11548)
- feat: bound global admission and expose running backend traces by
[@&#8203;richiejp](https://redirect.github.com/richiejp) in
[#&#8203;11560](https://redirect.github.com/mudler/LocalAI/pull/11560)
- feat(router): make KNN a first-class classifier with a persisted,
curated corpus by
[@&#8203;richiejp](https://redirect.github.com/richiejp) in
[#&#8203;10652](https://redirect.github.com/mudler/LocalAI/pull/10652)
- feat(chat): add end-to-end context compression by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11556](https://redirect.github.com/mudler/LocalAI/pull/11556)

##### 🧠 Models

- feat(gallery): add Grug 12B variants by
[@&#8203;localai-org-maint-bot](https://redirect.github.com/localai-org-maint-bot)
in
[#&#8203;11438](https://redirect.github.com/mudler/LocalAI/pull/11438)
- 

</details>

---

### Configuration

📅 **Schedule**: (UTC)

- Branch creation
  - At any time (no schedule defined)
- Automerge
  - At any time (no schedule defined)

🚦 **Automerge**: Enabled.

♻ **Rebasing**: Whenever PR becomes conflicted, or you tick the
rebase/retry checkbox.

🔕 **Ignore**: Close this PR and you won't be reminded about this update
again.

---

- [ ] <!-- rebase-check -->If you want to rebase/retry this PR, check
this box

---

This PR has been generated by [Renovate
Bot](https://redirect.github.com/renovatebot/renovate).

<!--renovate-debug:eyJjcmVhdGVkSW5WZXIiOiI0My4xMzAuMSIsInVwZGF0ZWRJblZlciI6IjQzLjEzMC4xIiwidGFyZ2V0QnJhbmNoIjoibWFzdGVyIiwibGFiZWxzIjpbImFwcC9sb2NhbC1haSIsImF1dG9tZXJnZSIsInJlbm92YXRlL2NvbnRhaW5lciIsInR5cGUvbWlub3IiXX0=-->
@Rye426

Rye426 commented Sep 2, 2026

Copy link
Copy Markdown

Requesting support for Qwen3-TTS-12HZ-0.6B. While the code successfully exports the 0.6B model, an error occurs during inference:0.03.232.837 E clip_init: failed to load model 'mmproj-Qwen3-TTS-12Hz-0.6B-Base-F16.gguf': operator(): unable to find tensor a.gen.code.proj_in.weight

@ServeurpersoCom

Copy link
Copy Markdown
Contributor

Hello,

@Rye426, this should be fixed here: #28231. Note that the mmproj GGUF has to be regenerated, the F32 rule lives in the conversion script.

@quine00 the fix touches the same code path, so it is worth retrying once it lands. That said, cloning here is speaker-embedding only: we extract a single timbre vector from your reference clip and condition on that.

The reference implementation also has an in-context mode that feeds the clip's own audio codes plus its transcript, which is what gets you a close likeness, and that one needs the codec encoder, which the mmproj does not currently ship.

@Rye426

Rye426 commented Sep 2, 2026

Copy link
Copy Markdown

Hello,

@Rye426, this should be fixed here: #28231. Note that the mmproj GGUF has to be regenerated, the F32 rule lives in the conversion script.

@quine00 the fix touches the same code path, so it is worth retrying once it lands. That said, cloning here is speaker-embedding only: we extract a single timbre vector from your reference clip and condition on that.

The reference implementation also has an in-context mode that feeds the clip's own audio codes plus its transcript, which is what gets you a close likeness, and that one needs the codec encoder, which the mmproj does not currently ship.

Thank you!
But
There is another issue: when using the Vulkan backend to execute the GET_ROWS operator, the backend does not support the current tensor offset configuration.
llama.cpp/ggml/src/ggml-vulkan/ggml-vulkan.cpp:12111: GGML_ASSERT(dst->op != GGML_OP_GET_ROWS || (a_offset == 0 && b_offset == 0 && d_offset == 0)) failed
[New LWP 761440]
[New LWP 761439]
[New LWP 761438]
[Thread debugging using libthread_db enabled]
Using host libthread_db library "/lib/aarch64-linux-gnu/libthread_db.so.1".
0x0000ffffa16eef68 in ?? () from /lib/aarch64-linux-gnu/libc.so.6
#0 0x0000ffffa16eef68 in ?? () from /lib/aarch64-linux-gnu/libc.so.6
#1 0x0000ffffa16e2618 in ?? () from /lib/aarch64-linux-gnu/libc.so.6
#2 0x0000ffffa16e265c [PAC] in ?? () from /lib/aarch64-linux-gnu/libc.so.6
#3 0x0000ffffa173aa24 [PAC] in wait4 () from /lib/aarch64-linux-gnu/libc.so.6
#4 0x0000aaaabf25216c [PAC] in ggml_print_backtrace ()
#5 0x0000aaaabf2522b0 in ggml_abort ()
#6 0x0000aaaabf07af00 in void init_pushconst_tensor_offsets<vk_op_binary_push_constants>(ggml_backend_vk_context*, vk_op_binary_push_constants&, ggml_tensor const*, ggml_tensor const*, ggml_tensor const*, ggml_tensor const*, ggml_tensor*) ()
#7 0x0000aaaabf20101c in void ggml_vk_op_f32<vk_op_binary_push_constants>(ggml_backend_vk_context*, std::shared_ptr<vk_context_struct>&, ggml_tensor const*, ggml_tensor const*, ggml_tensor const*, ggml_tensor const*, ggml_tensor*, ggml_op, vk_op_binary_push_constants&&) [clone .constprop.0] ()
#8 0x0000aaaabf202998 in ggml_vk_get_rows(ggml_backend_vk_context*, std::shared_ptr<vk_context_struct>&, ggml_tensor const*, ggml_tensor const*, ggml_tensor*) ()
#9 0x0000aaaabf21c230 in ggml_vk_build_graph(ggml_backend_vk_context*, ggml_cgraph*, int, ggml_tensor*, int, bool, bool, bool) [clone .isra.0] ()
#10 0x0000aaaabf21d594 in ggml_backend_vk_graph_compute(ggml_backend*, ggml_cgraph*) ()
#11 0x0000aaaabf273d58 in ggml_backend_sched_compute_splits(ggml_backend_sched*) ()
#12 0x0000aaaabf27490c in ggml_backend_sched_graph_compute ()
#13 0x0000aaaabef21254 in clip_encode(clip_ctx*, clip_encode_params*) ()
#14 0x0000aaaabee8c354 in mtmd_gen_audio_process_impl(mtmd_context*, mtmd_gen_inp const*, mtmd_gen_out*) ()
#15 0x0000aaaabee8c88c in mtmd_gen_audio_process ()
#16 0x0000aaaabef10ac8 in qwen3tts_gen_audio_pipeline::step_gen(int, float const*, float const**, bool*) ()
#17 0x0000aaaabef0e438 in mtmd_helper_gen_audio_step_gen ()
#18 0x0000aaaabea7ea48 in main ()
[Inferior 1 (process 761436) detached]
Aborted

@ServeurpersoCom

ServeurpersoCom commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Hello,
@Rye426, this should be fixed here: #28231. Note that the mmproj GGUF has to be regenerated, the F32 rule lives in the conversion script.

Thank you! But There is another issue: when using the Vulkan backend to execute the GET_ROWS operator, the backend does not support the current tensor offset configuration. llama.cpp/ggml/src/ggml-vulkan/ggml-vulkan.cpp:12111: GGML_ASSERT(dst->op != GGML_OP_GET_ROWS || (a_offset == 0 && b_offset == 0 && d_offset == 0)) failed [New LWP ...

Thanks, that is a separate problem: the gen graph hands GET_ROWS an index tensor that is a view at a small byte offset, and the Vulkan path requires those offsets to be aligned, while CPU and CUDA do not care. It hits the 1.7B the same way.

Please try this fix: #28240

@Rye426

Rye426 commented Sep 7, 2026

Copy link
Copy Markdown

Hello,
@Rye426, this should be fixed here: #28231. Note that the mmproj GGUF has to be regenerated, the F32 rule lives in the conversion script.

Thank you! But There is another issue: when using the Vulkan backend to execute the GET_ROWS operator, the backend does not support the current tensor offset configuration. llama.cpp/ggml/src/ggml-vulkan/ggml-vulkan.cpp:12111: GGML_ASSERT(dst->op != GGML_OP_GET_ROWS || (a_offset == 0 && b_offset == 0 && d_offset == 0)) failed [New LWP ...

Thanks, that is a separate problem: the gen graph hands GET_ROWS an index tensor that is a view at a small byte offset, and the Vulkan path requires those offsets to be aligned, while CPU and CUDA do not care. It hits the 1.7B the same way.

Please try this fix: #28240

Thanks

thecodacus pushed a commit to thecodacus/llama.cpp that referenced this pull request Sep 7, 2026
…gml-org#26254)

* convert text model

* main model load ok

* convert encoder ok

* speaker encoder loading ok

* speaker enc graph

* adapt vocab for backbone (with some tricks)

* add suppress_tokens

* poc new mtmd gen api

* convert code_predictor to gguf

* load gen_code model ok

* add clip_encode

* wire up

* code gen cgraph init version

Co-authored-by: Pascal <admin@serveurperso.com>

* code2wav convert to gguf

* code2wav graph ok

* wire up in/out

* (wip) subgraph

* wire up

* wip, correct code2wav

* demo (to be removed)

* code2wav preserve kv between calls

* demo voice clone

* llama: add llama_model_get_tok_embd

* mtmd_helper_gen_audio API

* fix clamp cold prefix

Co-authored-by: Pascal <admin@serveurperso.com>

* fuse snake op

Co-authored-by: Pascal <admin@serveurperso.com>

* demo: use proper sampling

* update dev docs

* polymorphism helper

* revamp llama-tts binary

* update docs

* fix compile

* fix lint

* nits

* add guide + docs

* more timings info

* clean up code comments

* security fixes

* update docs

* use ggml_build_forward_select, clean up comments

* fix ci

* use ISO 639-1 language code

* rename CODE2WAV --> GEN_WAV, update docs

* clean up

* clean up tts.cpp

* add seq_id

* add step_prompt()

* mtmd_helper_model_can_chat

* clean up comments

---------

Co-authored-by: Pascal <admin@serveurperso.com>
zbrad pushed a commit to zbrad/llama.cpp that referenced this pull request Sep 10, 2026
…gml-org#26254)

* convert text model

* main model load ok

* convert encoder ok

* speaker encoder loading ok

* speaker enc graph

* adapt vocab for backbone (with some tricks)

* add suppress_tokens

* poc new mtmd gen api

* convert code_predictor to gguf

* load gen_code model ok

* add clip_encode

* wire up

* code gen cgraph init version

Co-authored-by: Pascal <admin@serveurperso.com>

* code2wav convert to gguf

* code2wav graph ok

* wire up in/out

* (wip) subgraph

* wire up

* wip, correct code2wav

* demo (to be removed)

* code2wav preserve kv between calls

* demo voice clone

* llama: add llama_model_get_tok_embd

* mtmd_helper_gen_audio API

* fix clamp cold prefix

Co-authored-by: Pascal <admin@serveurperso.com>

* fuse snake op

Co-authored-by: Pascal <admin@serveurperso.com>

* demo: use proper sampling

* update dev docs

* polymorphism helper

* revamp llama-tts binary

* update docs

* fix compile

* fix lint

* nits

* add guide + docs

* more timings info

* clean up code comments

* security fixes

* update docs

* use ggml_build_forward_select, clean up comments

* fix ci

* use ISO 639-1 language code

* rename CODE2WAV --> GEN_WAV, update docs

* clean up

* clean up tts.cpp

* add seq_id

* add step_prompt()

* mtmd_helper_model_can_chat

* clean up comments

---------

Co-authored-by: Pascal <admin@serveurperso.com>
pl752 pushed a commit to pl752/llama.cpp that referenced this pull request Sep 15, 2026
…gml-org#26254)

* convert text model

* main model load ok

* convert encoder ok

* speaker encoder loading ok

* speaker enc graph

* adapt vocab for backbone (with some tricks)

* add suppress_tokens

* poc new mtmd gen api

* convert code_predictor to gguf

* load gen_code model ok

* add clip_encode

* wire up

* code gen cgraph init version

Co-authored-by: Pascal <admin@serveurperso.com>

* code2wav convert to gguf

* code2wav graph ok

* wire up in/out

* (wip) subgraph

* wire up

* wip, correct code2wav

* demo (to be removed)

* code2wav preserve kv between calls

* demo voice clone

* llama: add llama_model_get_tok_embd

* mtmd_helper_gen_audio API

* fix clamp cold prefix

Co-authored-by: Pascal <admin@serveurperso.com>

* fuse snake op

Co-authored-by: Pascal <admin@serveurperso.com>

* demo: use proper sampling

* update dev docs

* polymorphism helper

* revamp llama-tts binary

* update docs

* fix compile

* fix lint

* nits

* add guide + docs

* more timings info

* clean up code comments

* security fixes

* update docs

* use ggml_build_forward_select, clean up comments

* fix ci

* use ISO 639-1 language code

* rename CODE2WAV --> GEN_WAV, update docs

* clean up

* clean up tts.cpp

* add seq_id

* add step_prompt()

* mtmd_helper_model_can_chat

* clean up comments

---------

Co-authored-by: Pascal <admin@serveurperso.com>
frostyautumnleaf pushed a commit to frostyautumnleaf/llama.cpp that referenced this pull request Oct 5, 2026
…gml-org#26254)

* convert text model

* main model load ok

* convert encoder ok

* speaker encoder loading ok

* speaker enc graph

* adapt vocab for backbone (with some tricks)

* add suppress_tokens

* poc new mtmd gen api

* convert code_predictor to gguf

* load gen_code model ok

* add clip_encode

* wire up

* code gen cgraph init version

Co-authored-by: Pascal <admin@serveurperso.com>

* code2wav convert to gguf

* code2wav graph ok

* wire up in/out

* (wip) subgraph

* wire up

* wip, correct code2wav

* demo (to be removed)

* code2wav preserve kv between calls

* demo voice clone

* llama: add llama_model_get_tok_embd

* mtmd_helper_gen_audio API

* fix clamp cold prefix

Co-authored-by: Pascal <admin@serveurperso.com>

* fuse snake op

Co-authored-by: Pascal <admin@serveurperso.com>

* demo: use proper sampling

* update dev docs

* polymorphism helper

* revamp llama-tts binary

* update docs

* fix compile

* fix lint

* nits

* add guide + docs

* more timings info

* clean up code comments

* security fixes

* update docs

* use ggml_build_forward_select, clean up comments

* fix ci

* use ISO 639-1 language code

* rename CODE2WAV --> GEN_WAV, update docs

* clean up

* clean up tts.cpp

* add seq_id

* add step_prompt()

* mtmd_helper_model_can_chat

* clean up comments

---------

Co-authored-by: Pascal <admin@serveurperso.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion documentation Improvements or additions to documentation model Model specific mtmd Related to multimodal functionality (video/image/audio) testing Everything test related

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants