Skip to content

Eval bug: SYCL multi-GPU crash with Intel Arc Pro B50 + Arc A770 #27888

Description

@nicolataibi

Name and Version

llama.cpp build: b10644-d7a207411

The executable was built locally from the ggml-org/llama.cpp repository.

version: 0.3.0-dev (build 10644, commit d7a2074)
built with IntelLLVM 2026.1.0 for Linux x86_64

Operating systems

Linux

GGML backends

SYCL

Hardware

Operating systems
Fedora Linux x86_64
Kernel: 7.1.10-200.fc44.x86_64 #1 SMP PREEMPT_DYNAMIC Sun Aug 23 16:15:11 UTC 2026 x86_64 GNU/Linux

GGML backends
GGML SYCL
Intel Level Zero

CMake configuration:

GGML_SYCL=ON
GGML_SYCL_TARGET=INTEL
GGML_SYCL_SUPPORT_LEVEL_ZERO_API=ON
GGML_SYCL_HOST_MEM_FALLBACK=ON
GGML_SYCL_DNN=ON
GGML_SYCL_F16=ON
GGML_SYCL_GRAPH=ON

The issue appears to affect the SYCL multi-GPU path. Single-GPU SYCL inference works correctly.
Hardware
CPU:
AMD Ryzen 9 5950X (16 cores / 32 threads)

GPU 0:
Intel Arc Pro B50
16 GB VRAM
SYCL device: SYCL0
Reported memory: 16304 MiB
Reported free memory: approximately 16220 MiB

GPU 1:
Intel Arc A770
16 GB VRAM
SYCL device: SYCL1
Reported memory: 15473 MiB
Reported free memory: approximately 14910 MiB

The two GPUs are from different Intel GPU generations:

Arc Pro B50: Battlemage
Arc A770: Alchemist

Both devices are detected correctly by llama.cpp:

Available devices:
SYCL0: Intel(R) Arc(TM) Pro B50 Graphics (16304 MiB, 16220 MiB free)
SYCL1: Intel(R) Arc(TM) A770 Graphics (15473 MiB, 14910 MiB free)

Models

Model:
gemma-4-12b-it-UD-Q6_K_XL.gguf

Quantization:
Q6_K

Parameters:
approximately 12B

The model is an Unsloth GGUF quantization.

The model loads and runs correctly when using either GPU individually.

Problem description & steps to reproduce

The problem occurs when attempting to use the Intel Arc Pro B50 and Intel Arc A770 simultaneously through the SYCL backend.

Both devices are correctly detected:

Available devices:
SYCL0: Intel(R) Arc(TM) Pro B50 Graphics (16304 MiB, 16220 MiB free)
SYCL1: Intel(R) Arc(TM) A770 Graphics (15473 MiB, 14910 MiB free)

The following environment was used for the multi-GPU tests:

env -u ONEAPI_DEVICE_SELECTOR
-u SYCL_DEVICE_FILTER
-u SYCL_PI_LEVEL_ZERO_DEVICE_SCOPE
-u ZE_ENABLE_VALIDATION_LAYER
-u ZE_ENABLE_TRACING_LAYER
GGML_SYCL_USM_SYSTEM=1
GGML_SYCL_ENABLE_VMM=0
ZES_ENABLE_SYSMAN=1
UR_L0_ENABLE_RELAXED_ALLOCATION_LIMITS=1
~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0,SYCL1
--split-mode layer
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"

This results in:

level_zero backend failed with error: 39 (UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY)

Exception caught at:
ggml/src/ggml-sycl/ggml-sycl.cpp:5513

func: operator()

SYCL error:
CHECK_TRY_ERROR((stream)->memcpy(
data,
(const char *)tensor->data + offset,
size
))

in function:
ggml_backend_sycl_get_tensor_async

Importantly, the model works successfully on each GPU individually.

B50 only:

~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0
--split-mode none
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"

Result: works successfully.

A770 only:

~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0
--split-mode none
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"

Result: works successfully when SYCL0 is mapped to the A770 through the corresponding device-selection environment.

Additional multi-GPU split-mode tests produced different failures:

--split-mode layer
UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY

The failure occurs inside:

ggml_backend_sycl_get_tensor_async()
ggml-sycl.cpp:5513
--split-mode row

Result:

Segmentation fault (core dumped)
--split-mode tensor

Result:

GGML_ASSERT(src_ss[0].axis != GGML_BACKEND_SPLIT_AXIS_0) failed

ggml-backend-meta.cpp:543

The following mitigations were also tested:

GGML_SYCL_ENABLE_VMM=0
GGML_SYCL_USM_SYSTEM=1
ZES_ENABLE_SYSMAN=1
UR_L0_ENABLE_RELAXED_ALLOCATION_LIMITS=1

None resolved the multi-GPU problem.

The following device-selection variables were also explicitly unset during testing:

ONEAPI_DEVICE_SELECTOR
SYCL_DEVICE_FILTER
SYCL_PI_LEVEL_ZERO_DEVICE_SCOPE
ZE_ENABLE_VALIDATION_LAYER
ZE_ENABLE_TRACING_LAYER

This did not resolve the issue.

Therefore, the current evidence suggests that:

both Intel GPUs are correctly detected;
both GPUs can independently execute the same model;
the same model and quantization work independently on each device;
the failure is introduced when llama.cpp attempts multi-device execution;
different multi-GPU split modes fail in different ways;
the layer split reaches a Level Zero out-of-device-memory error during a tensor memory copy;
row results in a segmentation fault;
tensor reaches a GGML backend assertion.

I would appreciate guidance on whether this multi-GPU configuration is currently expected to work with SYCL/Level Zero when combining an Intel Arc Pro B50 (BMG) and an Intel Arc A770 (DG2), or whether there is a known limitation in the current SYCL backend.

First Bad Commit

Unknown.

The issue has not yet been bisected to an earlier commit.

The current reproducible version is:

d7a2074 (HEAD -> master, tag: b10644, origin/master, origin/HEAD)
models : support nanbeige4.2-3B (#27730)

Relevant log output

The most relevant failure is:

level_zero backend failed with error: 39 (UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY)

Exception caught at file:
/home/nick/dev/intel/llama.cpp/ggml/src/ggml-sycl/ggml-sycl.cpp,
line:5513,
func: operator()

SYCL error:
CHECK_TRY_ERROR((stream)->memcpy(
data,
(const char *)tensor->data + offset,
size
))

in function:
ggml_backend_sycl_get_tensor_async

The GDB backtrace subsequently shows the CLI waiting for the server-side completion path:

#0 __syscall_cancel_arch
#1 __internal_syscall_cancel
#2 __syscall_cancel
#3 poll
#4 cli_server::wait_ready(...)
#5 cli_context::init()
#6 llama_cli(int, char**)

The GDB libsycl.so auto-load warning appears after the crash, but I do not believe it is the cause of the failure. It is:

warning: File
"/opt/intel/oneapi/compiler/2026.1/lib/libsycl.so.9.0.0-gdb.py"
auto-loading has been declined

The actual SYCL/Level Zero error occurs before this GDB diagnostic.

The key reproducibility matrix is:

B50 only -> SUCCESS
A770 only -> SUCCESS

B50 + A770
layer -> UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY
row -> segmentation fault
tensor -> GGML_ASSERT in ggml-backend-meta.cpp

This issue can be reproduced with a minimal prompt (-p "Hello") and a very small context size (--ctx-size 1024), so the failure does not appear to be caused by a large context or inference workload.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions