Repository navigation
Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348
Description
Activity
- added a commit that references this issue
on Jun 22, 2026 Update: the currently-vendored version may already contain the fix
While investigating a separate issue (#1632, Ollama tool-calling reliability), verified the current state of the vendored
llama-cpp-sys-2/llama-cpp-2against this issue's acceptance criteria. This does not confirm Gemma 4 12B now passes end-to-end on the native path — only that the specificpeg-gemma4parser-loop defect this issue tracks appears to already be fixed in what's currently pinned, with no version bump needed to get it.Evidence chain
1. Currently resolved version (
Cargo.lock, not just the>=0.1.146constraint inpackages/nlp-engine/Cargo.toml):name = "llama-cpp-sys-2" version = "0.1.146"2. Vendored llama.cpp commit at that exact crate version (GitHub Contents API on the
llama-cpp-sys-2submodule pointer at tag0.1.146):llama.cpp submodule sha: e21cdc11a0461d8b0cbd28cc356d993bf6be7282https://github.com/ggml-org/llama.cpp/tree/e21cdc11a0461d8b0cbd28cc356d993bf6be7282 — commit message
common/gemma4 : handle parsing edge cases (#21760), 2026-04-13.3. Relevant upstream fix PRs and merge commits:
PR Title Merge commit Merged #21661 common : fix ambiguous grammar rule in gemma4ddf03c6d9a347c3b7b7577a643d23aaa8d03a7c82026-04-09 #21697 common : enable reasoning budget sampler for gemma4(this issue's referenced fix)d7ff074c87ecacd57d5760e2f678866ba9fe71492026-04-10 #21704 common : better align to the updated official gemma4 template3fc65063d9c356510b86fc2f15ca8aea711bfc472026-04-10 4. Ancestry confirmation (GitHub Compare API,
{merge_commit}...e21cdc11a0461d8b0cbd28cc356d993bf6be7282):- #21661 →
ahead_by: 49, behind_by: 0 - #21697 →
ahead_by: 39, behind_by: 0 - #21704 →
ahead_by: 30, behind_by: 0
behind_by: 0for all three means the currently-vendored commit contains every commit reachable from each fix's merge — the fixes are present in what's already pinned, not just chronologically prior to it.Note: this issue's body references upstream issue #21375 specifically as the peg-gemma4 loop bug; #21661/#21697/#21704 (found via the llama.cpp PR history around the vendored commit date) look like the actual fix PRs — worth cross-referencing #21375 directly to confirm they're the same fix before treating this as fully closed-out.
What's NOT yet confirmed
- Whether this issue's actual acceptance criterion —
aichat-matrix.ts 12b-gguf,create_schemawith full field definitions, no truncation — now passes. Have not re-run it. - Whether ADR-046's other two Gemma 4 native-path failure modes (Fix Gemma 4 12B tool-call routing: chat template mismatch causes Python-prose output instead of structured tool calls #1346 Python-prose tool calls, Gemma 4 12B runs away (4096-token cap every turn) + OAI parser ffi error -3 through NodeSpace's inference path, despite clean output via raw llama-server #1365 runaway generation / FFI error) are resolved by the same upstream work or remain separate open defects. ADR-046 was written 2026-07-05, three months after these fixes landed upstream (and using this same
0.1.146version), and still recorded Gemma 4 12B failing — so either something about that evaluation didn't exercise this fixed code path, or a different failure mode (Fix Gemma 4 12B tool-call routing: chat template mismatch causes Python-prose output instead of structured tool calls #1346/Gemma 4 12B runs away (4096-token cap every turn) + OAI parser ffi error -3 through NodeSpace's inference path, despite clean output via raw llama-server #1365) is still blocking it independently of the parser-loop bug.
Recommend re-running
aichat-matrix.ts 12b-ggufagainstgemma-4-12b-unsloth-q4kmto get a real answer before updating this issue's status — this comment only establishes that the passive-wait condition described in the issue body may already be satisfied, not that Gemma 4 12B is unblocked.- #21661 →
Re-test attempt: inconclusive on the parser fix, but reconfirms 16GB is not viable for 12B here
Attempted to actually re-run
aichat-matrix.tsagainstgemma-4-12b-unsloth-q4kmper this issue's acceptance criteria, following the reproduce recipe inscripts/aichat-1329-trial.md. Two clean attempts (fresh test daemon each time, no other heavy process competing for GPU at the time) both failed the same way:ggml_metal_synchronize: error: command buffer 0 failed with status 5 error: Insufficient Memory (00000008:kIOGPUCommandBufferCallbackErrorOutOfMemory) ... llama_decode: failed to decode, ret = -3This happened on ordinary conversational turns (a plain "Hi there" greeting), before ever reaching a turn that would exercise the large
create_schematool-call JSON this issue'speg-gemma4parser-loop bug is actually about. So this doesn't tell us anything new about whether the parser fix (see previous comment) actually resolves the loop — the run never got that far.What it does reconfirm: this specific 16GB test machine can't reliably run Gemma 4 12B at all, independent of the parser question. This matches the original #1329 trial's verdict ("12B is not viable on 16GB... the blocker is the f16 KV cache") even though the current catalog entry (
gemma-4-12b-unsloth-q4km) already usestype_k/type_v: Q8_0quantized KV cache specifically to reduce that footprint — still not enough headroom on this machine's actual available memory (confirmed viavm_stat: ~5.4GB recovered immediately after killing the test daemon, i.e. the model's own footprint was consuming that much beyond what the rest of the system needs).Net effect on this issue's re-verification: still open. The earlier evidence (fix PRs present in the vendored
llama-cpp-sys-2 0.1.146, confirmed via commit ancestry) stands as-is — that's about source code, not runtime behavior. But confirming the fix actually works end-to-end needs either:- A machine with more headroom than 16GB (the A/B test Gemma 4 12B (Q4_K_M) vs E4B for the local ai-chat agent #1329 trial's own recommendation — it was written explicitly to be rerun "on more capable hardware"), or
- Running at a reduced context size the way the A/B test Gemma 4 12B (Q4_K_M) vs E4B for the local ai-chat agent #1329 trial's
n_ctx: 8_192probe did, accepting that a reduced-context run doesn't fully validate the full-32K-context production shape.
Not attempting either right now — flagging so the next person who picks this up doesn't repeat the same two OOM cycles I just did.
Final result: 12B is not viable on this 16GB machine, independent of the parser fix
Four total attempts today (this comment supersedes the "inconclusive" framing of the previous one). The last two were run under genuinely clean, contention-free conditions — confirmed via
vm_statand the daemon's ownMiB freelog line at model-load time (11985-12123 MiB free, all 49/49 layers offloaded to GPU, no othernodespacedprocess running) — and both still failed identically:- Turn 1 (a plain greeting) barely completes, right at the 240s timeout ceiling.
- Turn 2 hard OOMs:
ggml_metal_synchronize: error: command buffer 0 failed with status 5/Insufficient Memory (00000008:kIOGPUCommandBufferCallbackErrorOutOfMemory), then retry-loops and fails identically on retry.
This is consistent across every clean attempt, which rules out transient contention as the explanation (earlier same-day attempts were confounded by contention — a different Claude session's own test daemon, and residual memory from an in-place model swap — but the last two controlled for both and got the same result).
Conclusion:
gemma-4-12b-unsloth-q4km— even with thetype_k/type_v: Q8_0quantized KV cache this catalog entry already applies specifically to reduce memory footprint — does not have enough sustained headroom to survive a second inference turn on this 16GB machine. This reconfirms the original #1329 trial's verdict ("12B is not viable on 16GB... the blocker is the KV cache, not the weights"), via a different specific symptom (hard OOM crash-loop rather than the trial's fabrication/silent-failure mode) but the same root cause and the same conclusion.This means the parser-loop fix question this issue tracks remains genuinely unverified — no attempt today ever got far enough into a conversation to reach a
create_schema-sized tool call before OOMing. The source-level evidence (fix present in vendoredllama-cpp-sys-2 0.1.146, previous comment) still stands, but confirming it works end-to-end needs a machine with real headroom beyond 16GB, per the #1329 trial's own recommendation.Useful control: Gemma 4 E4B ran clean on the same engine
For comparison, ran the same
aichat-matrix.tssuite againstgemma-4-e4b-q4kmon this machine, same engine version, no OOM at all — 12 scenarios, fast turns (2-25s, no timeouts), 10/12 passed. The two misses (list/query pickedexecute_queryover the expectedsearch_nodes; update stopped after search without completing theupdate_nodestep) look like skill-prompting/routing gaps, not tool-calling capability or memory issues — worth a follow-up but out of scope for this issue. This at least confirms the current engine version hasn't regressed anything for the model that's actually known to fit this hardware class.Recommend closing the loop on this issue as "source fix confirmed present, runtime verification blocked on hardware — needs re-attempt on a machine with >16GB" rather than leaving it looking like an open investigation on this machine.
Final resolution: peg-gemma4 parser-loop bug is fixed, but 12B remains parked for a different reason
Re-tested end-to-end on a 48GB machine as part of issue #1956 (past both the 16GB machine that OOM'd here previously and the 24GB
min_memory_gbthis issue's other Q8_0-KV-cache tier already declares).This issue's specific acceptance criterion is now confirmed passing:
augment_gemma4_stops(packages/nlp-engine/src/chat/mod.rs) already patches the<turn|>-not-marked-as-EOG gap in code, unconditionally, wheneverchat_format=3(PEG_GEMMA4) is detected — no infinite-loop or non-terminating generation was observed in any run against eithergemma-4-12b-q4kmorgemma-4-12b-unsloth-q4km.However, 12B still fails the
create_schemaacceptance criterion — for a different reason than this issue tracked. Both GGUF sources produce malformed tool-call JSON on the full-field-definitionscreate_schemascenario: escaped underscores in field names, truncated/garbled nested structures, and one run that hit a 2048-token generation cap with empty arguments. This reproduced identically across both sources — investigation found their embedded chat template and tokenizer EOG metadata are byte-for-byte identical, ruling out a template-level explanation.Root cause (via upstream research,
ggml-org/llama.cpp#21839): llama.cpp's grammar constraint for Gemma 4 tool calls only covers the outer envelope, not argument content — so nested-argument JSON correctness depends entirely on the model's own token-level accuracy, with no engine-level safety net. This is a genuine model-capability limitation at Q4_K_M for 12B, not fixable by a further llama-cpp-sys-2 bump.Full writeup: ADR-056's 2026-08-05 Addendum (
../nodespace-docs/decisions/056-gemma-4-e4b-locked-native-model.md).Closing this issue — its specific parser-loop question is answered (fixed). 12B's continued parking is now tracked via ADR-056's re-introduction path, not this issue.
- added a commit that references this issue
on Aug 6, 2026
Overview
Gemma 4 12B tool calls work correctly via the Ollama backend but fail on the GGUF/llama.cpp path. The root cause is that `llama-cpp-sys-2 0.1.146` (released 2026-04-30) predates Gemma 4 12B (released 2026-06-03) and the bundled llama.cpp has known bugs in the `peg-gemma4` parser for 12B-scale tool call payloads.
Root Cause
The `peg-gemma4` streaming parser (llama.cpp PR #21326) enters an infinite repetition loop when the model generates large tool call JSON (issue #21375 upstream). For `create_schema` with full field definitions, 12B generates ~500+ tokens of JSON; the parser re-parses on every token and never exits the tool call block. EOS (`<turn|>`) fires mid-JSON, delivering truncated `{}` args to the tool executor.
Verified in the #1329 A/B trial on M4 Pro / 48GB:
Fix
When a new `llama-cpp-sys-2` release bundles llama.cpp with the upstream peg-gemma4 fixes (PRs #21375, #21697), bump the version constraint in `packages/nlp-engine/Cargo.toml` and re-run the `aichat-matrix.ts` scenario matrix against `gemma-4-12b-unsloth-q4km` to verify.
This is a passive wait — no code change needed on our side until the crate is updated.
Acceptance Criteria
Technical Specifications
Reference Files
Watch For
Related Issues