Repository navigation
Updating README after running 60B of llama.cpp - #93
Closed
zitterbewegung wants to merge 1 commit into
Closed
zitterbewegung wants to merge 1 commit into
zitterbewegung wants to merge 1 commit into
Conversation
Author
Member
|
More detailed info should be provided in the README for each model. Added TODO |
jesusmb1995
pushed a commit
to jesusmb1995/llama.cpp
that referenced
this pull request
Mar 20, 2026
InfernalDread
pushed a commit
to InfernalDread/llama.cpp
that referenced
this pull request
Apr 23, 2026
…se-wht fix: inverse WHT in test-turbo-quant.c round-trip (#59)
phuongncn
pushed a commit
to phuongncn/llama.cpp-gx10-dgx-sparks-deepseekv4
that referenced
this pull request
Apr 28, 2026
Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
mndodd
added a commit
to mndodd/llama.cpp
that referenced
this pull request
Aug 12, 2026
…ead of the caps Adopts the structure B49's CUDA re-survey identified: ggml_cuda_should_use_mmq opens with type support, a shared-memory gate, and dp4a/mma availability, and only THEN reaches its per-(arch,type) ne11 ceilings. Capability first, tuning second, failing closed. ☠☠☠ WE HAD NO SUCH SEPARATION AND B60 IS THE BILL. q2_K over DEQ+GEMM returns NaN at EVERY width it serves -- measured at ne[1] 12/20/24/32/64, reordered AND not. The only thing standing between a user and that NaN was mmq_cap, a PERFORMANCE constant that happened to keep MMQ in front of it. f310 §7 set that constant for SPEED reasons and the NaN reappeared at 9..16 -- a correctness regression produced by a tuning edit, because the two share an integer field. ⇒ A CORRECTNESS CONSTRAINT ENCODED AS A PERFORMANCE CONSTANT IS ONE TUNING EDIT FROM BEING VIOLATED, AND NOTHING WILL SAY SO. ☠ SCOPE, STATED IN THE HEADER: this does NOT repair q2_K and does NOT reroute around it. Above the MMVQ switch width a reordered q2_K is refused by MMQ (ggml_sycl_mmq_reads_reordered -> false) and DEQ+GEMM is the only remaining path, so for that cell THERE IS NO CORRECT KERNEL TO CHOOSE. Verified: MMQ_CAP =4096 does not move it, with or without FORCE_REORDER. What the gate buys is that the failure stops being SILENT. Design: * default is LEGAL; the table lists only pairs MEASURED wrong. We cannot enumerate correctness, only the failures we have found, so an unlisted pair is UNTESTED, not blessed. The asymmetry is deliberate. * NOT env-gated -- a correctness diagnostic that fires only when someone remembered to ask for it is the defect it exists to catch, one level up. * warns ONCE per (type,kernel) behind a mutex, so it cannot become RIG-HYGIENE ggml-org#93's synchronous-output tax on a hot dispatch path. Gated both directions: must-fire q2_K -> DEQ+GEMM fires, names pair + cause must-not-fire q8_0/q6_K/q4_K @8,13,20,32 0 lines each once-only q2_K x 4 widths x 3 reps 1 block, not 12
SimonTeixidor
pushed a commit
to SimonTeixidor/llama.cpp
that referenced
this pull request
Sep 30, 2026
…h-lifetime qwen4exp: keep detached PLE prefetch inputs alive
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.