Repository navigation
[GSD-13448] sycl::malloc_device() failed when allocate memory size is more than 19.3GB, through the "Max memory allocation" is 30.3GB on B70 #998
Description
Activity
- addedType: BugGeneral bug report, unexpected behavior or crashGeneral bug report, unexpected behavior or crashOS: LinuxIssue specific to Linux distributions (Ubuntu, Fedora, RHEL, etc.)Issue specific to Linux distributions (Ubuntu, Fedora, RHEL, etc.)
on Sep 14, 2026 - changed the title
[-]sycl::malloc_device() failed when allocate memory size is more than 19.4GB, through the "Max memory allocation" is 30.3GB on B70[/-][+]sycl::malloc_device() failed when allocate memory size is more than 19.3GB, through the "Max memory allocation" is 30.3GB on B70[/+]on Sep 14, 2026 On B60 (same running time version), it's passed:
SYCL device: Intel(R) Arc(TM) Pro B60 Graphics GPU max memory allocation: 22.71 GiB Enter GPU memory to allocate in GiB (0 to quit): 22.71 Free Intel GPU memory BEFORE: 20749.99 MiB Successfully allocated 22.71 GiB at: 0xffffd556b6400000 Free Intel GPU memory AFTER : 329.89 MiB- changed the title
[-]sycl::malloc_device() failed when allocate memory size is more than 19.3GB, through the "Max memory allocation" is 30.3GB on B70[/-][+][GSD-13448] sycl::malloc_device() failed when allocate memory size is more than 19.3GB, through the "Max memory allocation" is 30.3GB on B70[/+]on Sep 15, 2026 Hi @arthw,
Thank you for the submission, and for following the templates when creating it - much appreciated.
Just to confirm: B60 is passing, while B70 returns an error on the same GPU driver version. Is that correct?
Thanks again for your cooperation.
@kgibala
Yes, only B70 has such issue.
Hope this issue be confirmed/reproduced in your lab as soon!Thank you!
@kgibala
Any progress of this issue?Hi @arthw,
We are currently dispatching this issue internally and will provide an update as soon as possible.
Thank you for your patience.
OK!
This issue appears in several user cases.
It's not an isolated case.If you need me test with more case, please add comments!
Thank you!
We tried to reproduce this on our B70, including on the exact driver version you reported, and could
not. Here is what we ran so you can see where our setup differs.Our card: Arc Pro B70 (8086:e223, Sparkle, 0000:04:00.0), Meigao N5A mini PC, Ubuntu 26.04, kernel
7.0.0-30-generic, xe, GuC firmware xe/bmg_guc_70.bin version 70.58.0. Our reporter does what your
description says: one sycl::malloc_device of N GiB from a gpu_selector_v queue, no kernel, no write,
then free. We also built a second variant that memsets the whole buffer, to separate "the allocation
is refused" from "the allocation is refused once the pages are actually populated".On intel-opencl-icd / libze-intel-gpu1 26.31.39395.13 with IGC 2.40.13, Level Zero 1.17.39395+13,
which is the version in your report:global_mem_size 31.89 GiB, max_mem_alloc_size 30.30 GiB (matches your numbers)
19.3 GiB, allocation only OK
19.4 GiB, allocation only OK <- the one that fails for you
20.0 GiB, allocation only OK
30.0 GiB, allocation only OK
19.4 GiB allocated and then fully written OK
kernel log across all five runs zero xe fault, CAT or reset linesFor context, 19.4 and 30.0 GiB also succeed on 26.35.39758.10 with IGC 2.41.5, and a 0.5 GiB step
ladder to 30.0 GiB completed on both 25.48.36300.8 and 26.35.39758.10, so nothing we run on this
card sees a ceiling near 19.3 GiB on any of the three drivers, and the device reports the same
30.30 GiB max_mem_alloc_size you do.So we cannot confirm it from here and it does not look like a limit that applies to every B70. Four
things that might explain the difference, and we will test any of them:- Your B60 passes 22.71 GiB on the same driver, which is above the 19.3 GiB your B70 refuses. That
points at the B70 card or its slot more than at the driver. Does the B70 fail from a fresh boot
with nothing else running, and does it fail on the very first allocation? - If your reporter does anything beyond a plain malloc_device from a gpu_selector_v queue, such as
a context with properties, a different USM kind like malloc_shared, or an allocation made while
another large buffer is already resident, we would like to see the source. Ours is about 20 lines
and we are happy to post it so it can be diffed against yours. - Does yours fail with a nullptr or with a thrown exception? Ours reports those separately and the
answer narrows it down a lot. - Which vendor board is your B70? Ours is Sparkle, and this is the only one we have.
Separately, two real memory path failures on this card that may or may not be related: 25.48.36300.8
faults the blit (bcs) engine within seconds of our own allocate and memset workload, and 26.35 fails
intermittently under sustained serving load with ccs and bcs queue resets. Both are written up with
logs and coredumps in #948. Different symptom from
the allocation size question, but both are on the same part.Happy to run a specific allocation sequence if that would help pin it down.
- Your B60 passes 22.71 GiB on the same driver, which is above the 19.3 GiB your B70 refuses. That
I compare the software/driver/firmware version.
xe/bmg_guc_70.bin version 70.72.1 (mine) 70.58.0(your)- download it to 70.58.0 - not fix
- upgrade driver to 26.35.39758.10 - not fix.
Here are my answers for your questions:
-
Your B60 passes 22.71 GiB on the same driver, which is above the 19.3 GiB your B70 refuses. That
points at the B70 card or its slot more than at the driver. Does the B70 fail from a fresh boot
with nothing else running, and does it fail on the very first allocation?
A:
My B60 is installed in another PC. B70 is installed in another PC, there is only a B70 dGPU and iGPU in 13400.
B70 is failed on the very first allocation after reboot. -
If your reporter does anything beyond a plain malloc_device from a gpu_selector_v queue, such as
a context with properties, a different USM kind like malloc_shared, or an allocation made while
another large buffer is already resident, we would like to see the source. Ours is about 20 lines
and we are happy to post it so it can be diffed against yours.
A:
Yes, I'd like to see your test code. If it works well on my B70, I will update the llama.cpp SYCL backend code as your code. -
Does yours fail with a nullptr or with a thrown exception? Ours reports those separately and theanswer narrows it down a lot.
You can refer to my reproduce code for this question. -
Which vendor board is your B70? Ours is Sparkle, and this is the only one we have.
My B70 vendor isGunnir.
You can see it in above log:
lspci -vvv -k -s 0000:03:00.0 03:00.0 VGA compatible controller: Intel Corporation Battlemage G31 [Arc Pro B70] (prog-if 00 [VGA controller]) Subsystem: Shenzhen GunnirDo you need I test the code with more debug optional?
Thank you!
@arthw Here is the code. It is a test program, not a fix, so I would not expect it to change
anything on your card, but it lets you diff like for like. The core is three lines: a
gpu_selector_v queue, sycl::malloc_device(bytes, q), sycl::free. I kept two extra modes in the same
file because your report and our own workload differ in whether the buffer is ever written, and on
the other fault path we chase that distinction turned out to matter.Build and run:
source /opt/intel/oneapi/setvars.sh icpx -fsycl -O2 -o b70_alloc_single b70_alloc_single.cpp ONEAPI_DEVICE_SELECTOR=level_zero:0 ./b70_alloc_single 19.4 alloc// b70_alloc_single.cpp - faithful replica of the reproducer in intel/compute-runtime#998. // // arthw's program allocates ONE sycl::malloc_device buffer of a given size and never touches it: // 19.3 GiB succeeds, 19.4 GiB returns nullptr, on driver 26.31.39395.13, on a B70. // // Our earlier ladder probe wrote 4KB into every allocation, so this exists to reproduce his // request exactly (mode=alloc, no touch at all) and then to separate "the allocation is refused" // from "the allocation is refused once the pages are actually populated" (mode=touch / fill). // // Usage: ./b70_alloc_single <gb> [alloc|touch|fill] // alloc malloc_device only, no kernel, no write (his test) // touch malloc_device plus a 4KB memset (our ladder behaviour) // fill malloc_device plus a full memset of the buffer // // Exit codes: 0 success, 2 malloc returned nullptr, 3 a SYCL exception (e.g. DEVICE_LOST). #include <sycl/sycl.hpp> #include <cstdio> #include <cstdlib> #include <cstring> int main(int argc, char **argv) { const double gb = (argc > 1) ? atof(argv[1]) : 19.4; const char *mode = (argc > 2) ? argv[2] : "alloc"; try { sycl::queue q(sycl::gpu_selector_v); const sycl::device dev = q.get_device(); const double global = dev.get_info<sycl::info::device::global_mem_size>() / 1073741824.0; const double maxalloc = dev.get_info<sycl::info::device::max_mem_alloc_size>() / 1073741824.0; printf("device : %s\n", dev.get_info<sycl::info::device::name>().c_str()); printf("driver : %s\n", dev.get_info<sycl::info::device::driver_version>().c_str()); printf("global_mem_size : %.2f GiB\n", global); printf("max_mem_alloc_size : %.2f GiB\n", maxalloc); const size_t bytes = (size_t)(gb * 1073741824.0); printf("requested : %.2f GiB (%zu bytes)\n", gb, bytes); printf("mode : %s\n", mode); fflush(stdout); void *p = sycl::malloc_device(bytes, q); if (p == nullptr) { printf("RESULT: malloc_device returned nullptr\n"); return 2; } printf("allocated at : %p\n", p); if (strcmp(mode, "touch") == 0) { q.memset(p, 0x5a, 4096).wait(); printf("4KB memset : OK\n"); } else if (strcmp(mode, "fill") == 0) { q.memset(p, 0x5a, bytes).wait(); printf("full memset : OK\n"); } sycl::free(p, q); printf("RESULT: OK (allocated and freed)\n"); return 0; } catch (const sycl::exception &e) { printf("RESULT: SYCL exception: %s\n", e.what()); return 3; } catch (const std::exception &e) { printf("RESULT: std::exception: %s\n", e.what()); return 3; } }
For comparison, this is our card on 26.31.39395.13 with IGC 2.40.13 and Level Zero 1.17.39395+13:
device reports global_mem_size 31.89 GiB, max_mem_alloc_size 30.30 GiB (same as yours) 19.3 GiB, alloc only OK 19.4 GiB, alloc only OK <- returns nullptr for you 20.0 GiB, alloc only OK 30.0 GiB, alloc only OK 19.4 GiB then a full memset of the buffer OK kernel log across all five attempts no xe fault, CAT or reset lines19.4 and 30.0 GiB also pass on 26.35.39758.10 with IGC 2.41.5, and a 0.5 GiB step ladder up to 30.0
GiB completed on 25.48.36300.8 as well. Our card reports the same 32GB BAR as yours and the same
advertised maximum allocation, so neither of those is the difference.Two things in your output that I keep coming back to:
- The 19.4 GiB attempt is refused while 32586 MiB is free and the device advertises 30.30 GiB. That
reads as a cap, not as exhaustion, and 19.3 GiB out of 30.3 GiB is close to 60 percent. I want to be
clear that this is a guess and not a finding. - Your B60 allocating 22.71 GiB on the same driver is about 94 percent of its 24 GB, which rules out
a simple percentage rule applied across the runtime. So whatever refuses your 19.4 GiB looks specific
to that B70.
Yes please to more debug. The three things that would help most:
- If your reproducer does anything beyond malloc_device and free, posting it lets us diff properly.
Queue and context properties alone could matter. - Allocation logging on the failing attempt. If the runtime is the layer refusing it, the log should
say so. marfrit's report in [GSD-13010] Sporadic permanent GPU wedge on dual Arc Pro B70 (BMG G31): ccs/bcs engine reset + "Fault response: Unsuccessful -ENOENT/-EINVAL" under sustained Level-Zero inference load #948 shows the allocation path is visible with the runtime debug keys
enabled. Our 19.4 GiB attempt produces no such line because it succeeds, so yours is the useful one. - You already ruled out the GuC downgrade and your BAR matches ours, so the visible differences left
are the board vendor (Gunnir there, Sparkle here) and the slot. Any other B70 in that machine, or that
card in another machine, would split it in one run.
Happy to run whatever specific sequence you want if you post it.
- The 19.4 GiB attempt is refused while 32586 MiB is free and the device advertises 30.30 GiB. That
Hi @arthw,
Thank you for the additional tests and for confirming that the B60 and B70 are installed in different PCs. We also noted that the B70 fails on the first allocation after reboot, and that neither the GuC firmware downgrade nor the update to driver 26.35.39758.10 resolved it.
We tested your reproducer on our B70, including with driver 26.31.39395.13, and both 19.3 GiB and 19.4 GiB allocations succeeded in that setup.
However, further testing gave us a useful lead: reducing system RAM on a B70 test machine, while keeping swap unchanged, lowered the allocation size at which the test started failing. In one configuration, 29.1 GiB succeeded and 29.2 GiB failed. With less RAM, 21.3 GiB succeeded and 21.4 GiB failed. This suggests a relationship with system RAM and swap capacity, but we have not yet confirmed whether it explains the 19.4 GiB failure on your machine or whether this behavior is expected.
Since your B60 and B70 are in different PCs, could you share the following output from both systems, identifying which is which?
free -h cat /proc/meminfo swapon --show
Debug output from the failing B70 run would also help. Please use your original
alloc_20gbreproducer in the environment where the failure occurs and run:( set -x NEOReadDebugKeys=1 PrintDebugSettings=1 PrintDebugMessages=1 LogAllocationSummaryReport=1 \ LogAllocationType=1 LogAllocationStdout=1 LogAllocationMemoryPool=1 \ PrintBOCreateDestroyResult=1 PrintBOBindingResult=1 \ PrintIoctlEntries=1 PrintXeLogs=1 \ ZEL_ENABLE_LOADER_LOGGING=1 ZEL_LOADER_LOGGING_LEVEL=trace ZEL_LOADER_LOG_CONSOLE=1 \ ZE_ENABLE_VALIDATION_LAYER=1 ZEL_LOADER_LOGGING_ENABLE_SUCCESS_PRINT=1 \ ./alloc_20gb 19.4 ) > b70-alloc-19.4.log 2>&1
Please attach
b70-alloc-19.4.log. If possible, repeat with19.3and save it asb70-alloc-19.3.logso we can compare the successful and failing attempts. Please also includesycl-lsandicpx --versionoutput from the same environment and indicate which driver version you used for these runs.If you have time for some additional testing, could you compare a few RAM and swap configurations on the B70 system and check whether the allocation failure threshold changes? The following scenarios would help us investigate the relationship observed in our lab:
- Your current RAM and swap configuration as a baseline.
- A few lower RAM limits, keeping swap unchanged. For example,
mem=16Gandmem=8G, if both are below your normal RAM capacity and leave enough memory to run the system and reproducer. - Optionally, a larger swap configuration with the original RAM configuration restored, to check whether the threshold increases.
Please change one variable at a time and keep the driver and reproducer unchanged. The RAM limits above are examples for comparison. Reducing RAM may lower the maximum successful allocation size.
To test a RAM limit, edit
/etc/default/grubwith administrator privileges and add amem=option inside the quotes of the existingGRUB_CMDLINE_LINUX_DEFAULTvalue. For example, for the 8 GiB scenario:GRUB_CMDLINE_LINUX_DEFAULT="... mem=8G ..."Here,
...represents your existing boot parameters. Keep those parameters and do not enter the dots literally. Replacemem=8Gwith the value for each scenario, keeping only onemem=option.Save the file, then apply the change and reboot after saving any open work:
sudo update-grub sudo reboot
For each scenario, capture
cat /proc/cmdline,free -h, andswapon --show, and repeat the 19.3 and 19.4 GiB allocation tests. Try smaller or larger allocations as needed to identify the largest successful size and the next failing size. Please share those results together with the RAM and swap configuration for each run. The usable RAM shown by Linux may be below the selected limit because of reserved memory.After the RAM tests, remove the added
mem=option from/etc/default/grub, runsudo update-grub, and reboot again to restore normal RAM availability. If a limited-memory boot prevents normal startup, remove that option from the kernel command line using the GRUB menu's boot-entry editor for that boot, then undo the persistent change as described above. Restore your original swap configuration after any swap experiments.These results should help us determine whether the RAM/swap-related behavior observed in our tests also applies to your system.
Thank you for helping us investigate this.
- addedStatus: Needs FeedbackWaiting for additional information from reporterWaiting for additional information from reporter
on Sep 21, 2026 @NateHag
Here is the result of your code on my B70:
Failed../build_b70_alloc_single.sh device : Intel(R) Arc(TM) Pro B70 Graphics driver : 1.17.39758+10 global_mem_size : 31.89 GiB max_mem_alloc_size : 30.30 GiB requested : 19.40 GiB (20830591385 bytes) mode : alloc RESULT: malloc_device returned nullptr@kgibala
B70:Log file:
sycl-ls [level_zero:gpu][level_zero:0] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) Arc(TM) Pro B70 Graphics 20.2.0 [1.17.39758+10] [level_zero:gpu][level_zero:1] Intel(R) oneAPI Unified Runtime over Level-Zero V2, Intel(R) UHD Graphics 730 12.2.0 [1.17.39758+10] [opencl:cpu][opencl:0] Intel(R) OpenCL, 13th Gen Intel(R) Core(TM) i5-13400 OpenCL 3.0 (Build 0) [2026.21.3.0.31_160000] [opencl:gpu][opencl:1] Intel(R) OpenCL Graphics (discrete), Intel(R) Arc(TM) Pro B70 Graphics OpenCL 3.0 NEO [26.35.39758.10] [opencl:gpu][opencl:2] Intel(R) OpenCL Graphics (integrated), Intel(R) UHD Graphics 730 OpenCL 3.0 NEO [26.35.39758.10] icpx --version Intel(R) oneAPI DPC++/C++ Compiler 2026.0.0 (2026.0.0.20260331) Target: x86_64-unknown-linux-gnu Thread model: posix InstalledDir: /opt/intel/oneapi/compiler/2026.0/bin/compiler Configuration file: /opt/intel/oneapi/compiler/2026.0/bin/compiler/../icpx.cfg free -h cat /proc/meminfo swapon --show total used free shared buff/cache available Mem: 15Gi 2.5Gi 10Gi 27Mi 2.8Gi 12Gi Swap: 4.0Gi 0B 4.0Gi MemTotal: 16126248 kB MemFree: 10948452 kB MemAvailable: 13545420 kB Buffers: 108212 kB Cached: 2729596 kB SwapCached: 0 kB Active: 2159824 kB Inactive: 1579360 kB Active(anon): 928872 kB Inactive(anon): 0 kB Active(file): 1230952 kB Inactive(file): 1579360 kB Unevictable: 64 kB Mlocked: 64 kB SwapTotal: 4194300 kB SwapFree: 4194300 kB Zswap: 0 kB Zswapped: 0 kB Dirty: 64664 kB Writeback: 0 kB AnonPages: 901672 kB Mapped: 629804 kB Shmem: 27772 kB KReclaimable: 91796 kB Slab: 335192 kB SReclaimable: 91796 kB SUnreclaim: 243396 kB KernelStack: 15216 kB PageTables: 27572 kB SecPageTables: 3576 kB NFS_Unstable: 0 kB Bounce: 0 kB WritebackTmp: 0 kB CommitLimit: 12257424 kB Committed_AS: 7143448 kB VmallocTotal: 34359738367 kB VmallocUsed: 71412 kB VmallocChunk: 0 kB Percpu: 17664 kB HardwareCorrupted: 0 kB AnonHugePages: 0 kB ShmemHugePages: 0 kB ShmemPmdMapped: 0 kB FileHugePages: 583680 kB FilePmdMapped: 53248 kB CmaTotal: 0 kB CmaFree: 1824768 kB Unaccepted: 0 kB Balloon: 0 kB HugePages_Total: 0 HugePages_Free: 0 HugePages_Rsvd: 0 HugePages_Surp: 0 Hugepagesize: 2048 kB Hugetlb: 0 kB DirectMap4k: 277012 kB DirectMap2M: 5771264 kB DirectMap1G: 11534336 kB NAME TYPE SIZE USED PRIO /swap.img file 4G 0B -1B70- after set host memory to 8G (base is 16G)
The max mem of succssful is changed from 19.3 to 15.7
./alloc_20gb 15.7 SYCL device: Intel(R) Arc(TM) Pro B70 Graphics GPU max memory allocation: 30.30 GiB Starting allocations... === Allocation 15.70(GiB) === Free Intel GPU memory BEFORE: 32155.87 MiB Successfully allocated 15.70 GiB at: 0xffffd556b6400000 Free Intel GPU memory AFTER : 16078.98 MiB Memory will be released when the program ends. Freeing all device memory... All device memory freed. ./alloc_20gb 15.8 SYCL device: Intel(R) Arc(TM) Pro B70 Graphics GPU max memory allocation: 30.30 GiB Starting allocations... === Allocation 15.80(GiB) === Free Intel GPU memory BEFORE: 32155.87 MiB Failed to allocate 15.8 GiB on device! Memory will be released when the program ends. Freeing all device memory... All device memory freed.free -h cat /proc/meminfo swapon --show total used free shared buff/cache available Mem: 5.1Gi 2.2Gi 1.7Gi 27Mi 1.4Gi 2.9Gi Swap: 4.0Gi 0B 4.0Gi MemTotal: 5304628 kB MemFree: 1797500 kB MemAvailable: 3037152 kB Buffers: 88472 kB Cached: 1359600 kB SwapCached: 0 kB Active: 1966140 kB Inactive: 344648 kB Active(anon): 889992 kB Inactive(anon): 0 kB Active(file): 1076148 kB Inactive(file): 344648 kB Unevictable: 64 kB Mlocked: 64 kB SwapTotal: 4194300 kB SwapFree: 4194300 kB Zswap: 0 kB Zswapped: 0 kB Dirty: 0 kB Writeback: 0 kB AnonPages: 862856 kB Mapped: 596524 kB Shmem: 27664 kB KReclaimable: 71520 kB Slab: 312376 kB SReclaimable: 71520 kB SUnreclaim: 240856 kB KernelStack: 17600 kB PageTables: 30568 kB SecPageTables: 3016 kB NFS_Unstable: 0 kB Bounce: 0 kB WritebackTmp: 0 kB CommitLimit: 6846612 kB Committed_AS: 7385452 kB VmallocTotal: 34359738367 kB VmallocUsed: 73820 kB VmallocChunk: 0 kB Percpu: 17600 kB HardwareCorrupted: 0 kB AnonHugePages: 0 kB ShmemHugePages: 0 kB ShmemPmdMapped: 0 kB FileHugePages: 0 kB FilePmdMapped: 0 kB CmaTotal: 0 kB CmaFree: 696284 kB Unaccepted: 0 kB Balloon: 0 kB HugePages_Total: 0 HugePages_Free: 0 HugePages_Rsvd: 0 HugePages_Surp: 0 Hugepagesize: 2048 kB Hugetlb: 0 kB DirectMap4k: 295444 kB DirectMap2M: 4184064 kB DirectMap1G: 3145728 kB NAME TYPE SIZE USED PRIO /swap.img file 4G 0B -1@kgibala
Yes, as you said, it's impacted by host memory.
But the DDR5 is too expensive now, so I won't buy more recently. :<Thank you!
B60
free -h cat /proc/meminfo swapon --show total used free shared buff/cache available Mem: 62Gi 12Gi 14Gi 87Mi 36Gi 50Gi Swap: 8.0Gi 140Mi 7.9Gi MemTotal: 65296752 kB MemFree: 15381884 kB MemAvailable: 52470100 kB Buffers: 432080 kB Cached: 36080976 kB SwapCached: 18080 kB Active: 11939896 kB Inactive: 35096896 kB Active(anon): 9150208 kB Inactive(anon): 1462568 kB Active(file): 2789688 kB Inactive(file): 33634328 kB Unevictable: 88708 kB Mlocked: 0 kB SwapTotal: 8388604 kB SwapFree: 8244388 kB Zswap: 0 kB Zswapped: 0 kB Dirty: 65572 kB Writeback: 0 kB AnonPages: 10607124 kB Mapped: 1257800 kB Shmem: 89192 kB KReclaimable: 1374332 kB Slab: 1831012 kB SReclaimable: 1374332 kB SUnreclaim: 456680 kB KernelStack: 17104 kB PageTables: 233696 kB SecPageTables: 8400 kB NFS_Unstable: 0 kB Bounce: 0 kB WritebackTmp: 0 kB CommitLimit: 41036980 kB Committed_AS: 17046832 kB VmallocTotal: 34359738367 kB VmallocUsed: 69040 kB VmallocChunk: 0 kB Percpu: 28400 kB HardwareCorrupted: 0 kB AnonHugePages: 0 kB ShmemHugePages: 16384 kB ShmemPmdMapped: 0 kB FileHugePages: 7421952 kB FilePmdMapped: 2048 kB CmaTotal: 0 kB CmaFree: 1073980 kB Unaccepted: 0 kB Balloon: 0 kB HugePages_Total: 0 HugePages_Free: 0 HugePages_Rsvd: 0 HugePages_Surp: 0 Hugepagesize: 2048 kB Hugetlb: 0 kB DirectMap4k: 363784 kB DirectMap2M: 7477248 kB DirectMap1G: 58720256 kB NAME TYPE SIZE USED PRIO /swap.img file 8G 140.8M -1@kgibala
Any update or progress?
I guess you have known some cause about this issue.
It's not common sense that the malloc on GPU is impacted by the host memory size.Thank you!
Title: [B65] Large single device allocation threshold appears to track host RAM + swap
Setup: Arc Pro B65 (BMG-G31, PCI <8086:e222>, ), Ubuntu 26.04, kernel 7.0.0-34-generic, xe driver,
llama.cpp SYCL in official image ghcr.io/ggml-org/llama.cpp:server-intel (build 11312, 0c1e57098),
Host: MemTotal 15219 MiB (14.86 GiB), swap initially 4 GiB.Observations (single B65, VRAM free 32585 MiB in all runs unless noted):
- Swap 4 GiB (RAM+swap = 18.86 GiB): one weights buffer of 18.74 GiB allocates, 19.29 GiB
fails with "can't allocate" (older build 9765, no workaround). - Build 11312 with the 60% cap from llama.cpp#28953: first chunk 19.08 GiB (device
max_mem_alloc_size ~ GiB) fails <free VRAM: ...>; with factor 0.55 (17.5 GiB chunks)
the full 25.4 GiB model loads. - Swap raised to 16 GiB (RAM+swap ~30.9 GiB), same official image 11312, no modification:
model loads and the server reaches "listening". <A/B result: swap off -> fail, on -> pass>
The threshold coincides with MemTotal+SwapTotal here, not with CommitLimit (~11.4 GiB on this
machine: swap + 50% RAM). Caveats: one machine, load only (no inference), no debug logs yet.
Happy to run the NEO debug-key run requested above on this card.- Swap 4 GiB (RAM+swap = 18.86 GiB): one weights buffer of 18.74 GiB allocates, 19.29 GiB
-
Is there any progress of this issue?
-
Is there any workaround for it, except to add new host memory?
Thank you!
-
* Is there any progress of this issue? * Is there any workaround for it, except to add new host memory?Thank you!
Since it works fine with the latest intel docker image, I didn't change anything in my system.
Raising the swap from 4GiB to 16GiB solved the problem.
It seemed to me that the Intel SYCL tried to allocate - or at least to check - whether there was enough - system and swap - ram (60% of AI gguf size).
Nvidia and AMD handle VRAM allocation for AI workloads differently.Best regards
@szabados-marton
Got it! It's a workaround.But this issue is common issue of SYCL code, instead of llama.cpp only.
I guess some high level feature lead to such abnormal behavior.
Allocate memory on GPU shouldn't be impacted by host memory size in principle.Thank you for your sharing!
Pre-submission Checklist
GPU Hardware
B70
DRI Devices Information
GPU Detailed Information (lspci output)
Driver Version
26.31.39395.13
Installed GPU Driver Packages
Driver Installation Details
Download deb files in https://github.com/intel/compute-runtime/releases/tag/26.31.39395.13.
sudo dpkg -i *.deb.
Linux Distribution
Other (please specify below)
Other Linux Distribution
Ubuntu 26.04
Kernel Version & Boot Parameters
Actual Behavior
sycl::malloc_device() failed when allocate memory size is more than 19.4GB, through the "Max memory allocation" is 30.3GB (clinfo) on B70.
When allocate <=19.3GB, it's successful.
Expected Behavior
sycl::malloc_device() should be successful when allocate memory size<= 30.3GB (Max memory allocation) on B70.
When allocate <=19.3GB, it's successful.
Reproduction Rate
Always reproduces - 100%
Steps to Reproduce
icpx -fsycl -o alloc_20gb alloc_20gb.cpp
./alloc_20gb 19.3
./alloc_20gb 19.4
alloc_20gb.cpp
Is this a regression?
Last Known Working Driver Version
No response
First Known Failing Driver Version
No response
API Call Logs
19.3 GB - passed
19.4GB - failed
strace Logs
No response
System Logs / dmesg Output
No response
Backtrace (if crash or hang occurred)
No response
Source Code / Reproducer
No response
Command Line / Application Details
No response
oneAPI Version (if applicable)
No response
Screenshots / Video
No response
Additional Notes
It's impact the llama.cpp SYCL backend. Refer to issue: ggml-org/llama.cpp#28778