Skip to content

Expose Gemma 4 12B as a selectable native model (once verified on adequate hardware) #1956

Description

@malibio

Context

Gemma 4 12B (gemma-4-12b-unsloth-q4km) is currently withheld from the chat model picker — only gemma-4-e4b-q4km is exposed via EXPOSED_GGUF_MODEL_IDS in packages/desktop-app/src-tauri/src/commands/chat_models.rs:47.

Per ADR-056 and issue #1348, 12B was parked for two reasons:

  1. An upstream peg-gemma4 streaming-parser infinite-loop bug (tool-call JSON generation) — source-level fix confirmed present in the currently vendored llama-cpp-sys-2 0.1.146 (per Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348 comment thread), but never verified end-to-end.
  2. Hardware: prior re-verification attempts on a 16GB machine hit hard GPU OOM (Insufficient Memory kIOGPUCommandBufferCallbackErrorOutOfMemory) on the second inference turn, even with the catalog entry's existing Q8_0 KV-cache quantization. Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348's own conclusion: "12B is not viable on 16GB — the blocker is memory headroom, not the parser."

Why now

Now running on a 48GB Apple M4 Pro — well past the 24GB min_memory_gb the gemma-4-12b-unsloth-q4km catalog entry already declares, and past the 16GB machine that OOM'd in #1348's tests.

Plan

  1. Re-run scripts/aichat-matrix.ts against gemma-4-12b-unsloth-q4km on this machine to get a real, current score — Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348's acceptance criteria (create_schema with full field definitions, no truncation) are still unverified end-to-end on adequate hardware.
  2. If it clears an acceptable bar, add gemma-4-12b-unsloth-q4km to EXPOSED_GGUF_MODEL_IDS in packages/desktop-app/src-tauri/src/commands/chat_models.rs:47.
  3. Update ADR-056 and/or close out Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348 with the verified result.

Status

Not yet started — still evaluating/testing on this machine before committing to exposing it. This issue is tracking intent, not requesting immediate work.

Related

Activity

  1. malibio commented on Aug 4, 2026

    @malibio
    CollaboratorAuthor

    Frontend baseline: 179 test files passed, 4286 tests passed (48.05s). Starting implementation in worktree issue-1956-gemma-12b.

  2. self-assigned this
    on Aug 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

enhancementNew feature or request

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions