Skip to content

Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348

Description

@malibio

Overview

Gemma 4 12B tool calls work correctly via the Ollama backend but fail on the GGUF/llama.cpp path. The root cause is that `llama-cpp-sys-2 0.1.146` (released 2026-04-30) predates Gemma 4 12B (released 2026-06-03) and the bundled llama.cpp has known bugs in the `peg-gemma4` parser for 12B-scale tool call payloads.

Root Cause

The `peg-gemma4` streaming parser (llama.cpp PR #21326) enters an infinite repetition loop when the model generates large tool call JSON (issue #21375 upstream). For `create_schema` with full field definitions, 12B generates ~500+ tokens of JSON; the parser re-parses on every token and never exits the tool call block. EOS (`<turn|>`) fires mid-JSON, delivering truncated `{}` args to the tool executor.

Verified in the #1329 A/B trial on M4 Pro / 48GB:

  • E4B GGUF: ✅ works (generates minimal 15-token args, closes cleanly)
  • 12B GGUF (Unsloth): ❌ peg-gemma4 loop, EOS at token 544 with `brace_depth=3`
  • 12B via Ollama: ✅ works (Ollama has its own template handling, uses non-streaming API)

Fix

When a new `llama-cpp-sys-2` release bundles llama.cpp with the upstream peg-gemma4 fixes (PRs #21375, #21697), bump the version constraint in `packages/nlp-engine/Cargo.toml` and re-run the `aichat-matrix.ts` scenario matrix against `gemma-4-12b-unsloth-q4km` to verify.

This is a passive wait — no code change needed on our side until the crate is updated.

Acceptance Criteria

  • `llama-cpp-sys-2` releases a version bundling llama.cpp with peg-gemma4 loop fix
  • Bump version in `packages/nlp-engine/Cargo.toml`
  • Re-run `aichat-matrix.ts 12b-gguf` — `create_schema` executes with full field definitions without truncation
  • E4B GGUF behavior unchanged

Technical Specifications

Reference Files

Watch For

Related Issues

Activity

  1. malibio commented on Jul 11, 2026

    @malibio
    CollaboratorAuthor

    Update: the currently-vendored version may already contain the fix

    While investigating a separate issue (#1632, Ollama tool-calling reliability), verified the current state of the vendored llama-cpp-sys-2/llama-cpp-2 against this issue's acceptance criteria. This does not confirm Gemma 4 12B now passes end-to-end on the native path — only that the specific peg-gemma4 parser-loop defect this issue tracks appears to already be fixed in what's currently pinned, with no version bump needed to get it.

    Evidence chain

    1. Currently resolved version (Cargo.lock, not just the >=0.1.146 constraint in packages/nlp-engine/Cargo.toml):

    name = "llama-cpp-sys-2"
    version = "0.1.146"
    

    2. Vendored llama.cpp commit at that exact crate version (GitHub Contents API on the llama-cpp-sys-2 submodule pointer at tag 0.1.146):

    llama.cpp submodule sha: e21cdc11a0461d8b0cbd28cc356d993bf6be7282
    

    https://github.com/ggml-org/llama.cpp/tree/e21cdc11a0461d8b0cbd28cc356d993bf6be7282 — commit message common/gemma4 : handle parsing edge cases (#21760), 2026-04-13.

    3. Relevant upstream fix PRs and merge commits:

    PR Title Merge commit Merged
    #21661 common : fix ambiguous grammar rule in gemma4 ddf03c6d9a347c3b7b7577a643d23aaa8d03a7c8 2026-04-09
    #21697 common : enable reasoning budget sampler for gemma4 (this issue's referenced fix) d7ff074c87ecacd57d5760e2f678866ba9fe7149 2026-04-10
    #21704 common : better align to the updated official gemma4 template 3fc65063d9c356510b86fc2f15ca8aea711bfc47 2026-04-10

    4. Ancestry confirmation (GitHub Compare API, {merge_commit}...e21cdc11a0461d8b0cbd28cc356d993bf6be7282):

    • #21661 → ahead_by: 49, behind_by: 0
    • #21697 → ahead_by: 39, behind_by: 0
    • #21704 → ahead_by: 30, behind_by: 0

    behind_by: 0 for all three means the currently-vendored commit contains every commit reachable from each fix's merge — the fixes are present in what's already pinned, not just chronologically prior to it.

    Note: this issue's body references upstream issue #21375 specifically as the peg-gemma4 loop bug; #21661/#21697/#21704 (found via the llama.cpp PR history around the vendored commit date) look like the actual fix PRs — worth cross-referencing #21375 directly to confirm they're the same fix before treating this as fully closed-out.

    What's NOT yet confirmed

    Recommend re-running aichat-matrix.ts 12b-gguf against gemma-4-12b-unsloth-q4km to get a real answer before updating this issue's status — this comment only establishes that the passive-wait condition described in the issue body may already be satisfied, not that Gemma 4 12B is unblocked.

  2. malibio commented on Jul 11, 2026

    @malibio
    CollaboratorAuthor

    Re-test attempt: inconclusive on the parser fix, but reconfirms 16GB is not viable for 12B here

    Attempted to actually re-run aichat-matrix.ts against gemma-4-12b-unsloth-q4km per this issue's acceptance criteria, following the reproduce recipe in scripts/aichat-1329-trial.md. Two clean attempts (fresh test daemon each time, no other heavy process competing for GPU at the time) both failed the same way:

    ggml_metal_synchronize: error: command buffer 0 failed with status 5
    error: Insufficient Memory (00000008:kIOGPUCommandBufferCallbackErrorOutOfMemory)
    ...
    llama_decode: failed to decode, ret = -3
    

    This happened on ordinary conversational turns (a plain "Hi there" greeting), before ever reaching a turn that would exercise the large create_schema tool-call JSON this issue's peg-gemma4 parser-loop bug is actually about. So this doesn't tell us anything new about whether the parser fix (see previous comment) actually resolves the loop — the run never got that far.

    What it does reconfirm: this specific 16GB test machine can't reliably run Gemma 4 12B at all, independent of the parser question. This matches the original #1329 trial's verdict ("12B is not viable on 16GB... the blocker is the f16 KV cache") even though the current catalog entry (gemma-4-12b-unsloth-q4km) already uses type_k/type_v: Q8_0 quantized KV cache specifically to reduce that footprint — still not enough headroom on this machine's actual available memory (confirmed via vm_stat: ~5.4GB recovered immediately after killing the test daemon, i.e. the model's own footprint was consuming that much beyond what the rest of the system needs).

    Net effect on this issue's re-verification: still open. The earlier evidence (fix PRs present in the vendored llama-cpp-sys-2 0.1.146, confirmed via commit ancestry) stands as-is — that's about source code, not runtime behavior. But confirming the fix actually works end-to-end needs either:

    Not attempting either right now — flagging so the next person who picks this up doesn't repeat the same two OOM cycles I just did.

  3. malibio commented on Jul 11, 2026

    @malibio
    CollaboratorAuthor

    Final result: 12B is not viable on this 16GB machine, independent of the parser fix

    Four total attempts today (this comment supersedes the "inconclusive" framing of the previous one). The last two were run under genuinely clean, contention-free conditions — confirmed via vm_stat and the daemon's own MiB free log line at model-load time (11985-12123 MiB free, all 49/49 layers offloaded to GPU, no other nodespaced process running) — and both still failed identically:

    • Turn 1 (a plain greeting) barely completes, right at the 240s timeout ceiling.
    • Turn 2 hard OOMs: ggml_metal_synchronize: error: command buffer 0 failed with status 5 / Insufficient Memory (00000008:kIOGPUCommandBufferCallbackErrorOutOfMemory), then retry-loops and fails identically on retry.

    This is consistent across every clean attempt, which rules out transient contention as the explanation (earlier same-day attempts were confounded by contention — a different Claude session's own test daemon, and residual memory from an in-place model swap — but the last two controlled for both and got the same result).

    Conclusion: gemma-4-12b-unsloth-q4km — even with the type_k/type_v: Q8_0 quantized KV cache this catalog entry already applies specifically to reduce memory footprint — does not have enough sustained headroom to survive a second inference turn on this 16GB machine. This reconfirms the original #1329 trial's verdict ("12B is not viable on 16GB... the blocker is the KV cache, not the weights"), via a different specific symptom (hard OOM crash-loop rather than the trial's fabrication/silent-failure mode) but the same root cause and the same conclusion.

    This means the parser-loop fix question this issue tracks remains genuinely unverified — no attempt today ever got far enough into a conversation to reach a create_schema-sized tool call before OOMing. The source-level evidence (fix present in vendored llama-cpp-sys-2 0.1.146, previous comment) still stands, but confirming it works end-to-end needs a machine with real headroom beyond 16GB, per the #1329 trial's own recommendation.

    Useful control: Gemma 4 E4B ran clean on the same engine

    For comparison, ran the same aichat-matrix.ts suite against gemma-4-e4b-q4km on this machine, same engine version, no OOM at all — 12 scenarios, fast turns (2-25s, no timeouts), 10/12 passed. The two misses (list/query picked execute_query over the expected search_nodes; update stopped after search without completing the update_node step) look like skill-prompting/routing gaps, not tool-calling capability or memory issues — worth a follow-up but out of scope for this issue. This at least confirms the current engine version hasn't regressed anything for the model that's actually known to fit this hardware class.

    Recommend closing the loop on this issue as "source fix confirmed present, runtime verification blocked on hardware — needs re-attempt on a machine with >16GB" rather than leaving it looking like an open investigation on this machine.

  4. malibio commented on Aug 5, 2026

    @malibio
    CollaboratorAuthor

    Final resolution: peg-gemma4 parser-loop bug is fixed, but 12B remains parked for a different reason

    Re-tested end-to-end on a 48GB machine as part of issue #1956 (past both the 16GB machine that OOM'd here previously and the 24GB min_memory_gb this issue's other Q8_0-KV-cache tier already declares).

    This issue's specific acceptance criterion is now confirmed passing: augment_gemma4_stops (packages/nlp-engine/src/chat/mod.rs) already patches the <turn|>-not-marked-as-EOG gap in code, unconditionally, whenever chat_format=3 (PEG_GEMMA4) is detected — no infinite-loop or non-terminating generation was observed in any run against either gemma-4-12b-q4km or gemma-4-12b-unsloth-q4km.

    However, 12B still fails the create_schema acceptance criterion — for a different reason than this issue tracked. Both GGUF sources produce malformed tool-call JSON on the full-field-definitions create_schema scenario: escaped underscores in field names, truncated/garbled nested structures, and one run that hit a 2048-token generation cap with empty arguments. This reproduced identically across both sources — investigation found their embedded chat template and tokenizer EOG metadata are byte-for-byte identical, ruling out a template-level explanation.

    Root cause (via upstream research, ggml-org/llama.cpp#21839): llama.cpp's grammar constraint for Gemma 4 tool calls only covers the outer envelope, not argument content — so nested-argument JSON correctness depends entirely on the model's own token-level accuracy, with no engine-level safety net. This is a genuine model-capability limitation at Q4_K_M for 12B, not fixable by a further llama-cpp-sys-2 bump.

    Full writeup: ADR-056's 2026-08-05 Addendum (../nodespace-docs/decisions/056-gemma-4-e4b-locked-native-model.md).

    Closing this issue — its specific parser-loop question is answered (fixed). 12B's continued parking is now tracked via ADR-056's re-introduction path, not this issue.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions