Skip to content

Eval bug: llama-server hard crash (cublasSgemm INVALID_VALUE) with --spec-type draft-mtp under KV-cache saturation #26558

Description

@ZisIsNotZis

Name and Version

$ llama-server --version
version: 3645 (6c8dcaa7ae)
built with GNU 13.3.0 for Linux x86_64

Commit tested: 6c8dcaa7ae41fa9f4aa2b3b68ee82cb8b2a03632 (master, sycl: parallelize the non-contiguous concat kernel (#25852))
CUDA build: CMAKE_BUILD_TYPE=Release, GGML_CUDA=ON, CMAKE_CUDA_ARCHITECTURES=89, built with GNU 13.3.0.

Operating systems

  • Linux (Ubuntu 24.04, kernel 7.0.0-28-generic)

GGML backends

  • CUDA

Hardware

NVIDIA GeForce RTX 4090 (24 GiB), driver 595.84, CUDA 13.2.

Models

unsloth/Qwen3.5-0.8B-MTP-GGUF at Q4_K_XL (the Qwen3.5-0.8B-UD-Q4_K_XL.gguf file). This is the Qwen3.5 "unified decoder" (hybrid full-attention / linear-attention / gated delta net) architecture with a built-in MTP/NextN head (qwen35.nextn_predict_layers, one nextn block at blk.24).

Problem description & steps to reproduce

llama-server crashes with a hard CUDA error (cublasSgemm_v2 → CUDA_ERROR_INVALID_VALUE → GGML_ABORT) when run with --spec-type draft-mtp under parallel load with KV-cache saturation. The crash is a GPU-API abort, not a graceful "context size exceeded" return, so it is a bug rather than an expected failure mode of running out of context.

This has been reproducible for a long time ("ever since MTP code merged"); the same crash class is reported in #20049 and #23803, and in Indras-Mirror/llama.cpp-turboq-mtp#17. The maintainers' earlier explanation ("total parallel tokens exceeded total KV slots") describes the trigger (the log is full of context-exceeded/retry events) but not the mechanism: a context-full condition must return decode() == 1 gracefully, never reach cublasSgemm.

Stable, fast reproduction (~20–25 min, no benchmark harness needed). Two independent server configs both crash; both crash on the same op:

Server A (smaller context):

llama-server -m Qwen3.5-0.8B-UD-Q4_K_XL.gguf --spec-type draft-mtp --temp 0 \
  -c 1024 -np 4 --kv-unified --host 127.0.0.1 --port 18081 --no-webui

Server B (larger context):

llama-server -m Qwen3.5-0.8B-UD-Q4_K_XL.gguf --spec-type draft-mtp --temp 0 \
  -c 8192 -np 4 --kv-unified --host 127.0.0.1 --port 18080 --no-webui

Stress client: many parallel /completion requests with mixed prompt lengths / n_predict, designed to keep the KV cache permanently saturated so the server is constantly in the failed to find free space in the KV cache, retrying with smaller batch size / Context size has been exceeded regime (the exact regime shown in the attached log). A minimal Python soak that reproduces it:

import concurrent.futures, json, random, urllib.request
HOST = "http://127.0.0.1:18081"; CONC = 10
WORDS = "the quick brown fox jumps over the lazy dog".split(); TEXT = " ".join(WORDS)
def mk(n): return " ".join([TEXT]*(max(1, int(n*4.2/len(TEXT)))))
def one(_):
    p = {"prompt": mk(random.choice([150,250,350,500,650,850])),
         "n_predict": random.choice([16,32,64,128,256,512]), "temperature": 0}
    req = urllib.request.Request(HOST+"/completion", data=json.dumps(p).encode(),
                                 headers={"Content-Type":"application/json"})
    try: urllib.request.urlopen(req, timeout=600).read()
    except Exception: pass
with concurrent.futures.ThreadPoolExecutor(CONC) as ex:
    futs = [ex.submit(one, i) for i in range(CONC)]
    while True:
        done, _ = concurrent.futures.wait(futs, return_when=concurrent.futures.FIRST_COMPLETED)
        for f in done:
            f.result(); futs.remove(f); futs.append(ex.submit(one, 0))

Both servers crash after ~20–25 min with the identical signature. The original (real-world) report was under hb run -d terminal-bench -a mini-swe-agent against the same server, where it took much longer (many thousands of tasks); the soak above just makes the same KV-saturation regime much denser.

First Bad Commit

Not bisected. Long-standing — the reporter states the crash has occurred "ever since the MTP code was merged", and the same crash class is described in #20049 / #23803 / Indras-Mirror/llama.cpp-turboq-mtp#17.

Relevant log output

Every crash is preceded by the same pattern: one or more Context size has been exceeded events, slots released, then 4 fresh slots launched, then the crash in the first decode:

E srv    send_error: task id = 46435, error: Context size has been exceeded.
I slot      release: id  1 | task 46435 | stop processing: n_tokens = 1293, truncated = 0
E srv    send_error: task id = 46443, error: Context size has been exceeded.
I slot      release: id  2 | task 46443 | stop processing: n_tokens = 2584, truncated = 0
E srv  update_slots: decode() failed: Context size has been exceeded.
I slot get_availabl: id  0 | task -1 | selected slot by LRU, t_last = 536191175449
I slot launch_slot_: id  0 | task 46553 | processing task, is_child = 0
... 4 slots launched ...
ggml/src/ggml-cuda/ggml-cuda.cu:106: CUDA error
E CUDA error: an unsupported value or parameter was passed to the function
E   current device: 0, in function ggml_cuda_mul_mat_cublas_impl at ggml-cuda.cu:1541
E   cublasSgemm_v2(ctx.cublas_handle(), CUBLAS_OP_T, CUBLAS_OP_N, ne01, ne11, ne10, (const float*)alpha, (const float*)src0_ptr, s01, (const float*)src1_ptr, s11, (const float*)beta, (float*)dst_ptr, ne0)

Captured failing GEMM (via a local diagnostic patch, LLAMA_DEBUG_GEMM=1)

I added a one-line diagnostic at the exact crash site (ggml-cuda.cu:1541, F32 cublasSgemm branch) that dumps the failing call's parameters on cublasSgemm_v2 != CUBLAS_STATUS_SUCCESS. Both independent reproductions failed on the identical matmul, with all cuBLAS constraints satisfied:

GEMM-FAIL: status=7 src0='blk.0.ssm_alpha.weight'(0) ne00=1024 ne01=16 nb00=4 nb01=4096
           src1='attn_norm-0'(0) ne10=1024 ne11=16 nb10=4 nb11=4096
           dst='node_43' ne0=16 | s01=1024 s11=1024 ne0=16
           align0=0 align1=0 alignD=0 src0_ptr=0x7f3c3ed64100 src1_ptr=0x7f3c6a04c300 dst_ptr=0x7f3c6a0fc300

Mapping to the call cublasSgemm(handle, CUBLAS_OP_T, CUBLAS_OP_N, m, n, k, …):

param value cuBLAS constraint verdict
m = ne01 16 ≥ 0 ok
n = ne11 16 ≥ 0 ok
k = ne10 1024 ≥ 0 ok
lda = s01 1024 ≥ max(1,k) (transa=T) ok
ldb = s11 1024 ≥ max(1,k) (transb=N) ok
ldc = ne0 16 ≥ max(1,m) ok
src0/src1/dst ptr non-null, 16-byte aligned — ok

status=7 is CUBLAS_STATUS_INVALID_VALUE. The operands are:

  • src0 = blk.0.ssm_alpha.weight — F32 [1024×16] weight of layer-0 linear attention (gated delta net) alpha projection (src/models/qwen35moe.cpp:393),
  • src1 = attn_norm-0 — F32 [1024×16] layer-0 attention norm output,
  • dst = node_43 — [16×16].

Since every printed parameter provably satisfies cuBLAS's own validation rules, the rejection indicates corruption of the cuBLAS handle / its internal state (i.e., memory corruption elsewhere in the MTP path), rather than a bad tensor shape. It is deterministic — the same op fails on independent runs — which is consistent with "first F32 cuBLAS call after the corruption event."

Analysis done

Confirmed:

  • The crash is an F32 (non-batched) cublasSgemm with a valid parameter set → handle/state corruption, not a user-config or context-size issue.
  • Deterministic failing op: layer-0 linear-attention ssm_alpha projection, always right after all slots are cleared via context-exceeded and fresh decodes start.
  • MTP is the only speculative mode with a second full context (ctx_dft, LLAMA_CONTEXT_TYPE_MTP) that re-decodes every target batch and maintains its own KV cache. The draft head graph is LLM_GRAPH_TYPE_DECODER_MTP (src/models/qwen35moe.cpp:553+).

Suspected mechanisms (pending confirmation):

  1. Un-synchronized read of the target's hidden states. llama_context::decode() deliberately does not synchronize at the end (//synchronize(); is commented out at src/llama-context.cpp:2080); nextn embeddings are copied to host with ggml_backend_tensor_get_async (llama-context.cpp:2008), and llama_get_embeddings_nextn() performs no sync (llama-context.cpp:934). The server calls llama_decode(ctx_tgt, …) and immediately common_speculative_process(...) → MTP::process() reads llama_get_embeddings_nextn(ctx_tgt) at common/speculative.cpp:1450 — on a different context/stream. This is a genuine host-memory data race under load.

  2. Batch-retry loop is not MTP-aware. tools/server/server-context.cpp:3684-3692 retries by halving n_batch on target-decode failure; MTP::process() (the second decode pass) is only invoked on target success (server-context.cpp:3697) but its KV/state are coupled to the batches that failed/retried.

  3. Separate CUDA streams + shared device memory between ctx_tgt and ctx_dft (each context owns a backend with its own stream/pool) enabling a use-after-free window on GPU buffers.

Tests in progress:

  • LLAMA_GRAPH_REUSE_DISABLE=1 soak (the logs show very heavy graph reuse: graphs reused = 21941–62952). If disabling graph reuse stops the crash, that pins the vector.
  • A llama_synchronize(ctx_tgt) before MTP::process() test (targets suspicion Merging tensors of larger models #1).

Related PRs/issues

Activity

  1. ZisIsNotZis commented on Aug 4, 2026

    @ZisIsNotZis
    Author

    Update: graph-reuse test result

    LLAMA_GRAPH_REUSE_DISABLE=1 does not prevent the crash. A server run with identical config (the 1024-context variant) but graph reuse disabled crashed on the identical failing op after ~51 min (vs ~21 min baseline). Same GEMM-FAIL signature:

    GEMM-FAIL: status=7 src0="blk.0.ssm_alpha.weight"(0) ne00=1024 ne01=16 nb00=4 nb01=4096
               src1="attn_norm-0"(0) ne10=1024 ne11=16 nb10=4 nb11=4096
               dst="node_43" ne0=16 | s01=1024 s11=1024 ne0=16
               align0=0 align1=0 alignD=0 src0_ptr=0x7f0900d64100 src1_ptr=0x7f0934018300 dst_ptr=0x7f09340c8300
    

    So graph reuse (the graphs reused = 21k-62k counters) is not the corruption vector; the extended time-to-crash matches the server being proportionally slower without graph reuse. The crash is deterministic on the layer-0 linear-attention ssm_alpha projection every time, always right after all slots are cleared via context-exceeded and fresh decodes start, always with a cuBLAS-valid parameter set — pointing to deterministic memory corruption (candidate: cuBLAS handle/workspace state) rather than a timing race.

    Next test being run: forcing llama_synchronize(ctx_tgt) before MTP::process() to test the un-synchronized async h_nextn read.

  2. ZisIsNotZis commented on Aug 4, 2026

    @ZisIsNotZis
    Author

    Update: sync test result — does NOT prevent the crash

    I rebuilt with llama_synchronize(ctx_tgt) inserted in common_speculative_impl_draft_mtp::process() immediately before the llama_get_embeddings_nextn(ctx_tgt) read (common/speculative.cpp:1450), to rule out the un-synchronized async h_nextn copy (llama-context.cpp:2080 has a commented-out synchronize(); the copy is ggml_backend_tensor_get_async at llama-context.cpp:2008).

    Result: still crashes, on the identical failing op, after ~57 min (vs ~21 min baseline):

    GEMM-FAIL: status=7 src0="blk.0.ssm_alpha.weight"(0) ne00=1024 ne01=16 nb00=4 nb01=4096
               src1="attn_norm-0"(0) ne10=1024 ne11=16 nb10=4 nb11=4096
               dst="node_43" ne0=16 | s01=1024 s11=1024 ne0=16
               align0=0 align1=0 alignD=0
    

    Combined test matrix (all 1024-ctx --kv-unified -np 4 config, same soak):

    variant time to crash failing op
    baseline ~21 min ssm_alpha × attn_norm
    LLAMA_GRAPH_REUSE_DISABLE=1 ~51 min ssm_alpha × attn_norm
    llama_synchronize(ctx_tgt) in MTP::process() ~57 min ssm_alpha × attn_norm

    Both changes only delay the crash roughly in proportion to how much they slow the server down; neither removes it. So the crash is robust and is not caused by graph reuse and not caused by the unsynced async h_nextn read. The deterministic same-op failure (always the first F32 cublasSgemm in layer-0 linear attention, always immediately after all slots are cleared via context-exceeded, always with a cuBLAS-valid parameter set) points to a genuine deterministic memory corruption in the MTP dual-context path — e.g. a host-side buffer overflow or a use-after-free between the target and draft contexts` (separate CUDA streams sharing the device memory allocator) — rather than a timing race.

    The crash does not reproduce with --spec-type off (same load), confirming it is specific to the MTP dual-context path.

  3. ZisIsNotZis commented on Aug 5, 2026

    @ZisIsNotZis
    Author

    Update: root mechanism identified — the CUDA stream handed to cuBLAS is corrupted

    Instrumented the crash site with probes that dump, at the exact failing call: cudaPeekAtLastError() (prior async CUDA error), cublasGetStream() (handle alive?), and the stream pointer stored in the handle. All three independent crashes (1024-ctx, 256-ctx, and an ASAN build) show the identical pattern:

    GEMM-FAIL: status=7 src0="blk.0.ssm_alpha.weight" ... src1="attn_norm-0" ...
      | handle=0x624c2c2262b0 stream=0x624c1f0e6fc0 ce_peek=0(no error) cublasGetStream=0 stream_after=0x624c1f0e6fc0
    
    • ce_peek=0 → no prior device kernel error; the stream is not poisoned (rules out a crashed kernel).
    • cublasGetStream returns success → the cuBLAS handle struct itself is intact (rules out handle corruption).
    • BUT the stream pointer stored in the handle (0x624c1f0e6fc0, 0x5b27815f1000, 0x633347fcd1d0) is a host-process-heap address, not a CUDA driver stream handle (valid ones are 0x7f...-range driver pointers on this system). cuBLAS launches the GEMM on this bogus stream → cudaErrorInvalidValue → CUBLAS_STATUS_INVALID_VALUE.

    Since cublasSetStream(handle, ctx.stream()) runs immediately before every GEMM (ggml-cuda.cu:1419), the garbage value almost certainly comes from ctx.stream() — i.e. the streams[device][stream] array inside the ggml_backend_cuda_context struct itself has been overwritten by a host-side write (a heap pointer value landed in the stream slot). That struct also owns the buffer pool and the CUDA-graph cache — consistent with the observed unbounded GPU memory growth (a 0.8B/1024-ctx server using 7.7 GiB, growing over time; reporter observed total GPU footprint creeping 20→23 GiB).

    An ASAN build reproduced the identical crash with no AddressSanitizer report — so the write is either into memory ASAN does not track (the cuBLAS handle's internals) or a same-size pointer write (not a byte-overrun). Pinning the exact writer is the next step; the MTP dual-context setup (separate CUDA streams, two interleaved decode passes) remains the only configuration that triggers it.

  4. ZisIsNotZis commented on Aug 5, 2026

    @ZisIsNotZis
    Author

    Update: CUDA graphs confirmed as the root cause of both the GPU memory leak and the crash

    Running with GGML_CUDA_DISABLE_GRAPHS=1 (which disables the backend-level CUDA graph capture, https://github.com/ggml-org/llama.cpp/blob/6c8dcaa7ae/ggml/src/ggml-cuda/common.cuh#L1258) completely changes the behavior:

    metric GGML_CUDA_DISABLE_GRAPHS=0 (baseline) GGML_CUDA_DISABLE_GRAPHS=1
    GPU memory (after ~10 min soak) 3184 MiB and growing 1430 MiB — flat
    GPU leak rate (post-reservation) ~40-50 MiB/min 0 MiB/min
    Crash (cublas INVALID_VALUE) Always in 21-57 min ALIVE after 46 min, 0 probes, 534k retries, 135k context-exceeded

    The cache is keyed by cgraph->nodes[0] (a ggml tensor pointer). The MTP draft context's graphs are rebuilt constantly (due to sampler-based graph-reuse rejection at https://github.com/ggml-org/llama.cpp/blob/6c8dcaa7ae/src/llama-graph.h#L820-L831), creating a steady stream of new cache entries. Each entry holds a captured cudaGraph with device memory — the eviction sweep (every 5s, evicting idle >=10s, https://github.com/ggml-org/llama.cpp/blob/6c8dcaa7ae/ggml/src/ggml-cuda/common.cuh#L1435-L1441) can't keep up.

    The crash mechanism (the streams array slot inside ggml_backend_cuda_context being overwritten with a host-heap pointer, previously documented in #26558 (comment)) ties to the same ggml_backend_cuda_context struct that owns the cache map, the pool, and the streams — the graph-cache churn corrupts the context's bookkeeping.

    Disabling CUDA graphs (GGML_CUDA_DISABLE_GRAPHS=1) is a clean workaround: no leak, no crash, at the cost of performance (no graph replay).

  5. ZisIsNotZis commented on Aug 5, 2026

    @ZisIsNotZis
    Author

    Final confirmatory experiment: GGML_CUDA_DISABLE_GRAPHS=1 — no crash, no leak

    To test whether the CUDA graph capture mechanism itself is the root cause (rather than the MTP logic per se), I ran the same repro (1024-ctx, --kv-unified -np 4, --spec-type draft-mtp, --temp 0, same soak) with GGML_CUDA_DISABLE_GRAPHS=1, which disables the backend-level CUDA graph capture (common.cuh:1258). The result:

    metric GGML_CUDA_DISABLE_GRAPHS=0 (baseline, all ∼6 runs) GGML_CUDA_DISABLE_GRAPHS=1
    Crash (cublas INVALID_VALUE) Always, in 21–57 min 0 after 1h47m of continuous soak (still running)
    GPU memory (post-reservation) Grows ~40–50 MiB/min Flat at 1430 MiB from startup
    GEMM probes fired ce_peek=0, handle alive, garbage host-heap stream 0 probes
    Retries processed 206k–536k before crash 1.27M (alive)
    Context-exceeded events 52k–92k before crash 325k (alive)

    The server survives 1.7× the longest graphs-ON crash time, with zero memory growth, zero probes, and zero crashes, while processing 1.27M retries and 325k context-exceeded events — far more churn than any baseline run.

    Interpretation: The CUDA graph cache (cuda_graphs unordered_map in ggml_backend_cuda_context, keyed by cgraph->nodes[0] pointer, common.cuh:1428) grows under the MTP draft context because the draft graph is rebuilt constantly (the sampler-based graph-reuse rejection at llama-graph.h:820-831 prevents can_reuse for identical shapes when the token/seq_id composition changes). Each rebuild creates a new cache entry with a captured cudaGraph holding device memory → the leak. The cache churn also corrupts the ggml_backend_cuda_context struct (which owns the cache map, the pool, and the streams[] array) — a heap pointer lands in the stream slot → cublasSgemm launches on a garbage stream → cudaErrorInvalidValue → CUBLAS_STATUS_INVALID_VALUE → GGML_ABORT.

    Workaround: GGML_CUDA_DISABLE_GRAPHS=1 (no crash, no leak — at the cost of CUDA graph replay performance). The proper fix would be to key the graph cache on the graph's structural identity (uid or shape) rather than the transient cgraph->nodes[0] pointer, or to disable the graph cache for the MTP draft context specifically.

    Summary of all findings in this issue:

    1. The crash is a cublasSgemm_v2 returning CUBLAS_STATUS_INVALID_VALUE (7) with a provably valid parameter set — the cuBLAS handle is intact, the pointers are non-null and 16-byte aligned, and all dimension/leading-dimension constraints hold.
    2. The root cause is a garbage CUDA stream pointer stored in the cuBLAS handle — a host-heap address (0x624c... range) instead of a valid driver stream handle, causing cudaLaunchKernel to fail on the bogus stream.
    3. The garbage stream comes from the ggml_backend_cuda_context::streams[] array being corrupted by the CUDA graph cache churn (the cache's cuda_graphs unordered_map, the pool, and the streams array all live in the same struct).
    4. The trigger is always the same: all slots hit context-exceeded, fresh decodes start, and the first F32 cublasSgemm in layer-0's linear-attention ssm_alpha projection fails — because it's the first cuBLAS call after the corruption.
    5. LLAMA_GRAPH_REUSE_DISABLE=1 and llama_synchronize(ctx_tgt) before MTP::process() both delay the crash proportionally to how much they slow the server, but do not prevent it — confirming the corruption is not a timing race.
    6. GGML_CUDA_DISABLE_GRAPHS=1 completely prevents both the leak and the crash — the server survives 1.7× the longest baseline crash time with zero symptoms.
    7. The GPU memory leak (~40-50 MiB/min per server under MTP, observed across multiple independent servers) is also eliminated by GGML_CUDA_DISABLE_GRAPHS=1.
  6. InfinityEngineer commented on Aug 17, 2026

    @InfinityEngineer

    Another data point + independent confirmation, from a different setup (dual RTX 3090, NVLink, Linux, CUDA 12.x, driver 580):

    We hit what looks like exactly this crash class running Qwen3.6-27B-MTP (--spec-type draft-mtp) under sustained batched load, with --split-mode tensor noticeably accelerating time-to-failure vs single-GPU layer split. Same signature: CUDA error: an unsupported value or parameter was passed to the function from cublasGemmEx (CUBLAS_STATUS_INVALID_VALUE), always on device 1, all GEMM params valid on inspection. Reproduced across three builds spanning June–August (4c65955, d69b7e606, 4dee52f).

    Mitigation matrix (build 4dee52f, same load generator, ~55s median time-to-failure baseline):

    config time to failure
    stock (CUDA graphs on) 53–69 s
    LLAMA_GRAPH_REUSE_DISABLE=1 ~7 min (≈7×, but still eventually dies)
    GGML_CUDA_DISABLE_GRAPHS=1 no failure — 15 min / 118 req clean, then 50 min sustained clean

    That matches your finding above almost exactly: graph-reuse-disable only delays the crash roughly in proportion to how much it slows the server down, while full graph-capture disable eliminates it outright. On our side we'd independently landed on the same suspect (the per-context graph cache keyed by cgraph->nodes[0], swept by the timed eviction, colliding/racing with an in-flight cudaGraphExec_t replay) before seeing your root-cause writeup — good to see it nailed down precisely to the ggml_backend_cuda_context::streams array getting a host-heap pointer written over a driver stream handle. Confirms this isn't specific to your 4090/single-GPU/tiny-model repro: same corruption reproduces on multi-GPU tensor-split with a much larger MTP model, and --split-mode tensor makes it worse, not better (more concurrent graph churn across devices presumably shrinks the race window).

    We're running with GGML_CUDA_DISABLE_GRAPHS=1 as a permanent workaround in production. Happy to run further diagnostics on our repro if useful — it fires reliably in under 2 minutes with graphs enabled, and we can pull the same stream/handle probes you used if that helps narrow the fix.

  7. ORippler commented on Aug 21, 2026

    @ORippler
    Collaborator

    I can't seem to reproduce the memory-growth at all. Did you configure the project in a specific way? Running server with ./build-x64-linux-gcc-reldbg/bin/llama-server -m /mnt/share/gguf/unsloth/Qwen3.5-0.8B-MTP-GGUF/Qwen3.5-0.8B-UD-Q4_K_XL.gguf --spec-type draft-mtp --temp 0 -c 1024 -np 4 --kv-unified --host 127.0.0.1 --port 18081 --no-webui and your command above

  8. ORippler commented on Aug 21, 2026

    @ORippler
    Collaborator

    @ZisIsNotZis please run your repro on latest master and see if it re-occurs (#26574 has been merged)

  9. ORippler commented on Aug 21, 2026

    @ORippler
    Collaborator

    CUDA 12.x, driver 580

    For CTK < 12.4, #26574 should fix a cudaGraph-associated memory-leak

  10. ZisIsNotZis commented on Aug 24, 2026

    @ZisIsNotZis
    Author

    @ORippler Interesting, now it seems not crash any more, I'm on Driver Version: 595.84, CUDA Version: 13.2, cuda-toolkit 13.3.1-1, cudnn9-cuda-13 9.25.0.15-1.

  11. ORippler commented on Aug 28, 2026

    @ORippler
    Collaborator

    @ORippler Interesting, now it seems not crash any more, I'm on Driver Version: 595.84, CUDA Version: 13.2, cuda-toolkit 13.3.1-1, cudnn9-cuda-13 9.25.0.15-1.

    Closing as it no longer repros apparently

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions