Name and Version
llama.cpp build: b10644-d7a207411
The executable was built locally from the ggml-org/llama.cpp repository.
version: 0.3.0-dev (build 10644, commit d7a2074)
built with IntelLLVM 2026.1.0 for Linux x86_64
Operating systems
Linux
GGML backends
SYCL
Hardware
Operating systems
Fedora Linux x86_64
Kernel: 7.1.10-200.fc44.x86_64 #1 SMP PREEMPT_DYNAMIC Sun Aug 23 16:15:11 UTC 2026 x86_64 GNU/Linux
GGML backends
GGML SYCL
Intel Level Zero
CMake configuration:
GGML_SYCL=ON
GGML_SYCL_TARGET=INTEL
GGML_SYCL_SUPPORT_LEVEL_ZERO_API=ON
GGML_SYCL_HOST_MEM_FALLBACK=ON
GGML_SYCL_DNN=ON
GGML_SYCL_F16=ON
GGML_SYCL_GRAPH=ON
The issue appears to affect the SYCL multi-GPU path. Single-GPU SYCL inference works correctly.
Hardware
CPU:
AMD Ryzen 9 5950X (16 cores / 32 threads)
GPU 0:
Intel Arc Pro B50
16 GB VRAM
SYCL device: SYCL0
Reported memory: 16304 MiB
Reported free memory: approximately 16220 MiB
GPU 1:
Intel Arc A770
16 GB VRAM
SYCL device: SYCL1
Reported memory: 15473 MiB
Reported free memory: approximately 14910 MiB
The two GPUs are from different Intel GPU generations:
Arc Pro B50: Battlemage
Arc A770: Alchemist
Both devices are detected correctly by llama.cpp:
Available devices:
SYCL0: Intel(R) Arc(TM) Pro B50 Graphics (16304 MiB, 16220 MiB free)
SYCL1: Intel(R) Arc(TM) A770 Graphics (15473 MiB, 14910 MiB free)
Models
Model:
gemma-4-12b-it-UD-Q6_K_XL.gguf
Quantization:
Q6_K
Parameters:
approximately 12B
The model is an Unsloth GGUF quantization.
The model loads and runs correctly when using either GPU individually.
Problem description & steps to reproduce
The problem occurs when attempting to use the Intel Arc Pro B50 and Intel Arc A770 simultaneously through the SYCL backend.
Both devices are correctly detected:
Available devices:
SYCL0: Intel(R) Arc(TM) Pro B50 Graphics (16304 MiB, 16220 MiB free)
SYCL1: Intel(R) Arc(TM) A770 Graphics (15473 MiB, 14910 MiB free)
The following environment was used for the multi-GPU tests:
env -u ONEAPI_DEVICE_SELECTOR
-u SYCL_DEVICE_FILTER
-u SYCL_PI_LEVEL_ZERO_DEVICE_SCOPE
-u ZE_ENABLE_VALIDATION_LAYER
-u ZE_ENABLE_TRACING_LAYER
GGML_SYCL_USM_SYSTEM=1
GGML_SYCL_ENABLE_VMM=0
ZES_ENABLE_SYSMAN=1
UR_L0_ENABLE_RELAXED_ALLOCATION_LIMITS=1
~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0,SYCL1
--split-mode layer
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"
This results in:
level_zero backend failed with error: 39 (UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY)
Exception caught at:
ggml/src/ggml-sycl/ggml-sycl.cpp:5513
func: operator()
SYCL error:
CHECK_TRY_ERROR((stream)->memcpy(
data,
(const char *)tensor->data + offset,
size
))
in function:
ggml_backend_sycl_get_tensor_async
Importantly, the model works successfully on each GPU individually.
B50 only:
~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0
--split-mode none
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"
Result: works successfully.
A770 only:
~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0
--split-mode none
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"
Result: works successfully when SYCL0 is mapped to the A770 through the corresponding device-selection environment.
Additional multi-GPU split-mode tests produced different failures:
--split-mode layer
UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY
The failure occurs inside:
ggml_backend_sycl_get_tensor_async()
ggml-sycl.cpp:5513
--split-mode row
Result:
Segmentation fault (core dumped)
--split-mode tensor
Result:
GGML_ASSERT(src_ss[0].axis != GGML_BACKEND_SPLIT_AXIS_0) failed
ggml-backend-meta.cpp:543
The following mitigations were also tested:
GGML_SYCL_ENABLE_VMM=0
GGML_SYCL_USM_SYSTEM=1
ZES_ENABLE_SYSMAN=1
UR_L0_ENABLE_RELAXED_ALLOCATION_LIMITS=1
None resolved the multi-GPU problem.
The following device-selection variables were also explicitly unset during testing:
ONEAPI_DEVICE_SELECTOR
SYCL_DEVICE_FILTER
SYCL_PI_LEVEL_ZERO_DEVICE_SCOPE
ZE_ENABLE_VALIDATION_LAYER
ZE_ENABLE_TRACING_LAYER
This did not resolve the issue.
Therefore, the current evidence suggests that:
both Intel GPUs are correctly detected;
both GPUs can independently execute the same model;
the same model and quantization work independently on each device;
the failure is introduced when llama.cpp attempts multi-device execution;
different multi-GPU split modes fail in different ways;
the layer split reaches a Level Zero out-of-device-memory error during a tensor memory copy;
row results in a segmentation fault;
tensor reaches a GGML backend assertion.
I would appreciate guidance on whether this multi-GPU configuration is currently expected to work with SYCL/Level Zero when combining an Intel Arc Pro B50 (BMG) and an Intel Arc A770 (DG2), or whether there is a known limitation in the current SYCL backend.
First Bad Commit
Unknown.
The issue has not yet been bisected to an earlier commit.
The current reproducible version is:
d7a2074 (HEAD -> master, tag: b10644, origin/master, origin/HEAD)
models : support nanbeige4.2-3B (#27730)
Relevant log output
The most relevant failure is:
level_zero backend failed with error: 39 (UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY)
Exception caught at file:
/home/nick/dev/intel/llama.cpp/ggml/src/ggml-sycl/ggml-sycl.cpp,
line:5513,
func: operator()
SYCL error:
CHECK_TRY_ERROR((stream)->memcpy(
data,
(const char *)tensor->data + offset,
size
))
in function:
ggml_backend_sycl_get_tensor_async
The GDB backtrace subsequently shows the CLI waiting for the server-side completion path:
#0 __syscall_cancel_arch
#1 __internal_syscall_cancel
#2 __syscall_cancel
#3 poll
#4 cli_server::wait_ready(...)
#5 cli_context::init()
#6 llama_cli(int, char**)
The GDB libsycl.so auto-load warning appears after the crash, but I do not believe it is the cause of the failure. It is:
warning: File
"/opt/intel/oneapi/compiler/2026.1/lib/libsycl.so.9.0.0-gdb.py"
auto-loading has been declined
The actual SYCL/Level Zero error occurs before this GDB diagnostic.
The key reproducibility matrix is:
B50 only -> SUCCESS
A770 only -> SUCCESS
B50 + A770
layer -> UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY
row -> segmentation fault
tensor -> GGML_ASSERT in ggml-backend-meta.cpp
This issue can be reproduced with a minimal prompt (-p "Hello") and a very small context size (--ctx-size 1024), so the failure does not appear to be caused by a large context or inference workload.
Name and Version
llama.cpp build: b10644-d7a207411
The executable was built locally from the ggml-org/llama.cpp repository.
version: 0.3.0-dev (build 10644, commit d7a2074)
built with IntelLLVM 2026.1.0 for Linux x86_64
Operating systems
Linux
GGML backends
SYCL
Hardware
Operating systems
Fedora Linux x86_64
Kernel: 7.1.10-200.fc44.x86_64 #1 SMP PREEMPT_DYNAMIC Sun Aug 23 16:15:11 UTC 2026 x86_64 GNU/Linux
GGML backends
GGML SYCL
Intel Level Zero
CMake configuration:
GGML_SYCL=ON
GGML_SYCL_TARGET=INTEL
GGML_SYCL_SUPPORT_LEVEL_ZERO_API=ON
GGML_SYCL_HOST_MEM_FALLBACK=ON
GGML_SYCL_DNN=ON
GGML_SYCL_F16=ON
GGML_SYCL_GRAPH=ON
The issue appears to affect the SYCL multi-GPU path. Single-GPU SYCL inference works correctly.
Hardware
CPU:
AMD Ryzen 9 5950X (16 cores / 32 threads)
GPU 0:
Intel Arc Pro B50
16 GB VRAM
SYCL device: SYCL0
Reported memory: 16304 MiB
Reported free memory: approximately 16220 MiB
GPU 1:
Intel Arc A770
16 GB VRAM
SYCL device: SYCL1
Reported memory: 15473 MiB
Reported free memory: approximately 14910 MiB
The two GPUs are from different Intel GPU generations:
Arc Pro B50: Battlemage
Arc A770: Alchemist
Both devices are detected correctly by llama.cpp:
Available devices:
SYCL0: Intel(R) Arc(TM) Pro B50 Graphics (16304 MiB, 16220 MiB free)
SYCL1: Intel(R) Arc(TM) A770 Graphics (15473 MiB, 14910 MiB free)
Models
Model:
gemma-4-12b-it-UD-Q6_K_XL.gguf
Quantization:
Q6_K
Parameters:
approximately 12B
The model is an Unsloth GGUF quantization.
The model loads and runs correctly when using either GPU individually.
Problem description & steps to reproduce
The problem occurs when attempting to use the Intel Arc Pro B50 and Intel Arc A770 simultaneously through the SYCL backend.
Both devices are correctly detected:
Available devices:
SYCL0: Intel(R) Arc(TM) Pro B50 Graphics (16304 MiB, 16220 MiB free)
SYCL1: Intel(R) Arc(TM) A770 Graphics (15473 MiB, 14910 MiB free)
The following environment was used for the multi-GPU tests:
env -u ONEAPI_DEVICE_SELECTOR
-u SYCL_DEVICE_FILTER
-u SYCL_PI_LEVEL_ZERO_DEVICE_SCOPE
-u ZE_ENABLE_VALIDATION_LAYER
-u ZE_ENABLE_TRACING_LAYER
GGML_SYCL_USM_SYSTEM=1
GGML_SYCL_ENABLE_VMM=0
ZES_ENABLE_SYSMAN=1
UR_L0_ENABLE_RELAXED_ALLOCATION_LIMITS=1
~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0,SYCL1
--split-mode layer
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"
This results in:
level_zero backend failed with error: 39 (UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY)
Exception caught at:
ggml/src/ggml-sycl/ggml-sycl.cpp:5513
func: operator()
SYCL error:
CHECK_TRY_ERROR((stream)->memcpy(
data,
(const char *)tensor->data + offset,
size
))
in function:
ggml_backend_sycl_get_tensor_async
Importantly, the model works successfully on each GPU individually.
B50 only:
~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0
--split-mode none
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"
Result: works successfully.
A770 only:
~/dev/intel/llama.cpp/build/bin/llama-cli
--model "/run/media/nick/KINGSTON/AI/gemma-4-12b-it-UD-Q6_K_XL.gguf"
--device SYCL0
--split-mode none
--ctx-size 1024
--batch-size 32
--threads 4
--threads-batch 4
-p "Hello"
Result: works successfully when SYCL0 is mapped to the A770 through the corresponding device-selection environment.
Additional multi-GPU split-mode tests produced different failures:
--split-mode layer
UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY
The failure occurs inside:
ggml_backend_sycl_get_tensor_async()
ggml-sycl.cpp:5513
--split-mode row
Result:
Segmentation fault (core dumped)
--split-mode tensor
Result:
GGML_ASSERT(src_ss[0].axis != GGML_BACKEND_SPLIT_AXIS_0) failed
ggml-backend-meta.cpp:543
The following mitigations were also tested:
GGML_SYCL_ENABLE_VMM=0
GGML_SYCL_USM_SYSTEM=1
ZES_ENABLE_SYSMAN=1
UR_L0_ENABLE_RELAXED_ALLOCATION_LIMITS=1
None resolved the multi-GPU problem.
The following device-selection variables were also explicitly unset during testing:
ONEAPI_DEVICE_SELECTOR
SYCL_DEVICE_FILTER
SYCL_PI_LEVEL_ZERO_DEVICE_SCOPE
ZE_ENABLE_VALIDATION_LAYER
ZE_ENABLE_TRACING_LAYER
This did not resolve the issue.
Therefore, the current evidence suggests that:
both Intel GPUs are correctly detected;
both GPUs can independently execute the same model;
the same model and quantization work independently on each device;
the failure is introduced when llama.cpp attempts multi-device execution;
different multi-GPU split modes fail in different ways;
the layer split reaches a Level Zero out-of-device-memory error during a tensor memory copy;
row results in a segmentation fault;
tensor reaches a GGML backend assertion.
I would appreciate guidance on whether this multi-GPU configuration is currently expected to work with SYCL/Level Zero when combining an Intel Arc Pro B50 (BMG) and an Intel Arc A770 (DG2), or whether there is a known limitation in the current SYCL backend.
First Bad Commit
Unknown.
The issue has not yet been bisected to an earlier commit.
The current reproducible version is:
d7a2074 (HEAD -> master, tag: b10644, origin/master, origin/HEAD)
models : support nanbeige4.2-3B (#27730)
Relevant log output
The most relevant failure is:
level_zero backend failed with error: 39 (UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY)
Exception caught at file:
/home/nick/dev/intel/llama.cpp/ggml/src/ggml-sycl/ggml-sycl.cpp,
line:5513,
func: operator()
SYCL error:
CHECK_TRY_ERROR((stream)->memcpy(
data,
(const char *)tensor->data + offset,
size
))
in function:
ggml_backend_sycl_get_tensor_async
The GDB backtrace subsequently shows the CLI waiting for the server-side completion path:
#0 __syscall_cancel_arch
#1 __internal_syscall_cancel
#2 __syscall_cancel
#3 poll
#4 cli_server::wait_ready(...)
#5 cli_context::init()
#6 llama_cli(int, char**)
The GDB libsycl.so auto-load warning appears after the crash, but I do not believe it is the cause of the failure. It is:
warning: File
"/opt/intel/oneapi/compiler/2026.1/lib/libsycl.so.9.0.0-gdb.py"
auto-loading has been declined
The actual SYCL/Level Zero error occurs before this GDB diagnostic.
The key reproducibility matrix is:
B50 only -> SUCCESS
A770 only -> SUCCESS
B50 + A770
layer -> UR_RESULT_ERROR_OUT_OF_DEVICE_MEMORY
row -> segmentation fault
tensor -> GGML_ASSERT in ggml-backend-meta.cpp
This issue can be reproduced with a minimal prompt (-p "Hello") and a very small context size (--ctx-size 1024), so the failure does not appear to be caused by a large context or inference workload.