Skip to content

Updating README after running 60B of llama.cpp - #93

Closed
zitterbewegung wants to merge 1 commit into
ggml-org:masterfrom
zitterbewegung:patch-1
Closed

zitterbewegung wants to merge 1 commit into
ggml-org:masterfrom
zitterbewegung:patch-1

Conversation

@zitterbewegung

Copy link
Copy Markdown

No description provided.

@zitterbewegung

Copy link
Copy Markdown
Author

@ggerganov

Copy link
Copy Markdown
Member

More detailed info should be provided in the README for each model. Added TODO

@ggerganov ggerganov closed this Mar 13, 2023
jesusmb1995 pushed a commit to jesusmb1995/llama.cpp that referenced this pull request Mar 20, 2026
InfernalDread pushed a commit to InfernalDread/llama.cpp that referenced this pull request Apr 23, 2026
…se-wht

fix: inverse WHT in test-turbo-quant.c round-trip (#59)
phuongncn pushed a commit to phuongncn/llama.cpp-gx10-dgx-sparks-deepseekv4 that referenced this pull request Apr 28, 2026
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
mndodd added a commit to mndodd/llama.cpp that referenced this pull request Aug 12, 2026
…ead of the caps

Adopts the structure B49's CUDA re-survey identified: ggml_cuda_should_use_mmq
opens with type support, a shared-memory gate, and dp4a/mma availability, and
only THEN reaches its per-(arch,type) ne11 ceilings. Capability first, tuning
second, failing closed.

☠☠☠ WE HAD NO SUCH SEPARATION AND B60 IS THE BILL. q2_K over DEQ+GEMM returns
NaN at EVERY width it serves -- measured at ne[1] 12/20/24/32/64, reordered AND
not. The only thing standing between a user and that NaN was mmq_cap, a
PERFORMANCE constant that happened to keep MMQ in front of it. f310 §7 set that
constant for SPEED reasons and the NaN reappeared at 9..16 -- a correctness
regression produced by a tuning edit, because the two share an integer field.
⇒ A CORRECTNESS CONSTRAINT ENCODED AS A PERFORMANCE CONSTANT IS ONE TUNING EDIT
  FROM BEING VIOLATED, AND NOTHING WILL SAY SO.

☠ SCOPE, STATED IN THE HEADER: this does NOT repair q2_K and does NOT reroute
around it. Above the MMVQ switch width a reordered q2_K is refused by MMQ
(ggml_sycl_mmq_reads_reordered -> false) and DEQ+GEMM is the only remaining
path, so for that cell THERE IS NO CORRECT KERNEL TO CHOOSE. Verified: MMQ_CAP
=4096 does not move it, with or without FORCE_REORDER. What the gate buys is
that the failure stops being SILENT.

Design:
 * default is LEGAL; the table lists only pairs MEASURED wrong. We cannot
   enumerate correctness, only the failures we have found, so an unlisted pair
   is UNTESTED, not blessed. The asymmetry is deliberate.
 * NOT env-gated -- a correctness diagnostic that fires only when someone
   remembered to ask for it is the defect it exists to catch, one level up.
 * warns ONCE per (type,kernel) behind a mutex, so it cannot become
   RIG-HYGIENE ggml-org#93's synchronous-output tax on a hot dispatch path.

Gated both directions:
  must-fire     q2_K -> DEQ+GEMM         fires, names pair + cause
  must-not-fire q8_0/q6_K/q4_K @8,13,20,32   0 lines each
  once-only     q2_K x 4 widths x 3 reps     1 block, not 12
SimonTeixidor pushed a commit to SimonTeixidor/llama.cpp that referenced this pull request Sep 30, 2026
…h-lifetime

qwen4exp: keep detached PLE prefetch inputs alive
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants