You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Expose Gemma 4 12B as a selectable native model (once verified on adequate hardware) #1956
Gemma 4 12B (gemma-4-12b-unsloth-q4km) is currently withheld from the chat model picker — only gemma-4-e4b-q4km is exposed via EXPOSED_GGUF_MODEL_IDS in packages/desktop-app/src-tauri/src/commands/chat_models.rs:47.
Per ADR-056 and issue #1348, 12B was parked for two reasons:
Hardware: prior re-verification attempts on a 16GB machine hit hard GPU OOM (Insufficient Memory kIOGPUCommandBufferCallbackErrorOutOfMemory) on the second inference turn, even with the catalog entry's existing Q8_0 KV-cache quantization. Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348's own conclusion: "12B is not viable on 16GB — the blocker is memory headroom, not the parser."
Why now
Now running on a 48GB Apple M4 Pro — well past the 24GB min_memory_gb the gemma-4-12b-unsloth-q4km catalog entry already declares, and past the 16GB machine that OOM'd in #1348's tests.
If it clears an acceptable bar, add gemma-4-12b-unsloth-q4km to EXPOSED_GGUF_MODEL_IDS in packages/desktop-app/src-tauri/src/commands/chat_models.rs:47.
Not yet started — still evaluating/testing on this machine before committing to exposing it. This issue is tracking intent, not requesting immediate work.
Context
Gemma 4 12B (
gemma-4-12b-unsloth-q4km) is currently withheld from the chat model picker — onlygemma-4-e4b-q4kmis exposed viaEXPOSED_GGUF_MODEL_IDSinpackages/desktop-app/src-tauri/src/commands/chat_models.rs:47.Per ADR-056 and issue #1348, 12B was parked for two reasons:
peg-gemma4streaming-parser infinite-loop bug (tool-call JSON generation) — source-level fix confirmed present in the currently vendoredllama-cpp-sys-2 0.1.146(per Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348 comment thread), but never verified end-to-end.Insufficient Memory kIOGPUCommandBufferCallbackErrorOutOfMemory) on the second inference turn, even with the catalog entry's existing Q8_0 KV-cache quantization. Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348's own conclusion: "12B is not viable on 16GB — the blocker is memory headroom, not the parser."Why now
Now running on a 48GB Apple M4 Pro — well past the 24GB
min_memory_gbthegemma-4-12b-unsloth-q4kmcatalog entry already declares, and past the 16GB machine that OOM'd in #1348's tests.Plan
scripts/aichat-matrix.tsagainstgemma-4-12b-unsloth-q4kmon this machine to get a real, current score — Upgrade llama-cpp-sys-2 when Gemma 4 12B peg-gemma4 parser fixes land upstream #1348's acceptance criteria (create_schemawith full field definitions, no truncation) are still unverified end-to-end on adequate hardware.gemma-4-12b-unsloth-q4kmtoEXPOSED_GGUF_MODEL_IDSinpackages/desktop-app/src-tauri/src/commands/chat_models.rs:47.Status
Not yet started — still evaluating/testing on this machine before committing to exposing it. This issue is tracking intent, not requesting immediate work.
Related