Getting Started · Platform Support · Documentation · Citation · Artifacts
libumsh is a C++ library for sharing memory between CPUs and GPUs on unified memory architectures. It provides a common interface for memory binding, data-access policies, and control operations across Intel integrated GPUs and ARM64 NVIDIA systems.
Sharing physical memory does not by itself make CPU and GPU accesses coherent. libumsh separates how memory is allocated and mapped from how each endpoint reads, writes, and synchronizes. Applications can inspect the effective binding and available operations, then choose a protocol for their hardware.
The library serves two kinds of communication:
- Data-plane transmission: share bulk data such as model weights, images, and inference outputs, with binding and access policies selected for the transfer direction and kernel boundary.
- Control-plane coordination: exchange flags and ownership state to coordinate CPU/GPU work, using signaling for producer/consumer dependencies and locks for shared resources.
- Explicit memory binding: request page size, CPU/GPU cache policy, and coherency, and inspect the effective attributes and backing verification.
- Data-access policies: describe reads and writes in terms of cache invalidation, flushing, and bypass, with backend-specific implementations.
- Control operations: use
acquire/releasefor ownership,lock/unlockfor mutual exclusion, andwait/setfor signaling on supported platforms. - Capability-aware selection: query available implementations and use platform defaults with strict checks for unsupported combinations.
- Tools and extensibility: use correctness tests and standalone benchmarks to validate a configuration, or add a backend or device profile.
libumsh builds natively on Linux. Choose the deployment guide for your platform:
| Platform | Build target | Backend | Deployment guide |
|---|---|---|---|
| Intel Raptor Lake | RAPTOR |
OpenCL USM | Intel |
| Intel Arrow Lake | ARROW |
OpenCL USM | Intel |
| Intel Lunar Lake | LUNAR |
OpenCL USM | Intel |
| Intel Meteor Lake | METEOR |
OpenCL USM | Intel |
| NVIDIA Jetson AGX Orin | ORIN |
CUDA (sm_87) |
NVIDIA |
| NVIDIA DGX Spark / GB10 | SPARK |
CUDA (sm_121) |
NVIDIA |
The CUDA backend targets the ARM64 UMA systems above. See implementation status for available binding configurations, read/write policies, default plans, and control operations. For another platform, follow the extension guide.
Build on the target Linux machine with CMake 3.24+, a C++23 compiler, Python 3,
and the backend's development headers and runtime. Intel requires OpenCL with
cl_intel_unified_shared_memory; NVIDIA requires a CUDA toolkit that can
generate native code for the target GPU. The platform guides provide the
installation details.
From the repository root:
cmake -S . -B build-release \
-DCMAKE_BUILD_TYPE=Release \
-DUMSH_TARGET_PLATFORM=AUTO \
-DUMSH_BUILD_EXAMPLES=ON \
-DUMSH_BUILD_KERNEL_MODULE=OFF
cmake --build build-release --parallel
ctest --test-dir build-release --output-on-failureAUTO detects the native platform. An explicit build target from the table
selects its identity; it does not enable cross-compilation. Examples fetch a
pinned Abseil dependency during configuration. Set UMSH_BUILD_EXAMPLES=OFF
to build only the library and enabled host tests without fetching Abseil;
BUILD_TESTING=OFF disables those tests.
The default build uses the stock GPU driver. Optional kernel modules enable specific backing or cache-maintenance extensions; their build and loading instructions are in the platform guides.
The Core interface creates a binding from a size and explicit options:
#include <umsh/core/core.hpp>
#include <umsh/utils/error.hpp>
int main() {
auto core = UMSH_CHECK(umsh::core::create_core());
UMSH_CHECK(core->initialize(0));
umsh::binding::BindOptions options{
.page_size = umsh::binding::PageSize::Default,
.cpu_cache = umsh::binding::CpuCachePolicy::WriteBack,
.gpu_cache = umsh::binding::GpuCachePolicy::RuntimeDefault,
.coherency = umsh::binding::Coherency::RuntimeDefault,
};
auto memory = UMSH_CHECK(core->bind(4096, options));
// Use memory.host_ptr on the CPU and the backend's GPU access interface.
// Finish all CPU/GPU users before releasing the allocation.
UMSH_CHECK(core->unbind(memory));
}The APIs return umsh::Result<T>; UMSH_CHECK is a convenience for examples.
Applications can inspect the returned error and recover instead.
core->binding_info(memory) describes the requested and effective binding.
Check same_backing and its verification method when shared physical backing
is required. Successful allocation alone does not establish a concurrent
CPU/GPU data-transfer protocol.
Install the library and headers:
cmake --install build-release --prefix /path/to/umsh-installSave the allocation example as main.cpp in a separate application directory
and link the exported CMake target:
cmake_minimum_required(VERSION 3.24)
project(my_app LANGUAGES CXX)
find_package(umsh CONFIG REQUIRED)
add_executable(my_app main.cpp)
target_link_libraries(my_app PRIVATE umsh::umsh)Configure and build from that application directory:
cmake -S . -B build -DCMAKE_PREFIX_PATH=/path/to/umsh-install
cmake --build build --parallelThe exported target carries the backend/platform definitions and C++23 requirement. Applications with CUDA device code additionally need a CUDA-enabled CMake project.
Start with the binding and capability probe:
build-release/examples/custom/policy_probe/umsh_policy_probeIt reports the effective binding and available policies. Inspect the
same_backing and verification fields separately from the probe's PASS.
For CPU/GPU communication, run the hardware correctness tests described below.
- Control operations: CPU/GPU signaling and ownership examples, with tests for publication, bounded waiting, and mutual exclusion.
- Policy-transfer tests: exercise reads and writes while a GPU kernel is running. The matrix includes weak policies and unsupported combinations, so not every row is expected to pass.
- File loading: load a file into a caller-owned shared allocation using buffered, mmap, or direct I/O.
- Standalone benchmarks: measure binding costs, CPU/GPU reads and writes, GPU contention, and SSD loading.
ctest covers host-side default selection and, with examples enabled,
storage-report handling. GPU correctness tests run explicitly on the target
machine; follow the Intel or
NVIDIA commands before measuring performance.
Applications bind memory through Core, inspect its capabilities, and select
the data and control operations needed by their communication protocol. Native
backends supply the allocation, cache-maintenance, and device-access mechanisms.
| Component | Responsibility |
|---|---|
| Core and backends | Allocate/import shared memory, report effective attributes, and expose backend accessors. |
| Policy contracts | Describe reads, writes, scopes, support levels, and implementation identities. |
| Platform defaults | Specify compile-time transfer plans and reject unavailable selections. |
| Control layer and implementations | Provide ownership primitives and signaling/locking helpers using dedicated control memory. |
| Optional kernel extensions | Supply additional memory attributes and cache-maintenance mechanisms. |
Read policies are Alone, AfterInvalidate, or Bypass; write policies are
Alone, ThenFlush, or Bypass. Query support before choosing an
implementation. CUDA cache operators and fences are separate raw
characterization primitives, with their own semantics.
A default plan describes a platform's intended binding and endpoint operations. If a required binding or implementation is unavailable, strict selection rejects that plan. New platforms need their own validated defaults.
Control words have a separate lifetime from payload memory. Ownership and signaling order accesses, while the payload's policy supplies any required cache maintenance. Complete all CPU/GPU users before releasing either allocation; see the control-plane guide for examples and ordering rules.
As described in Section 5 of the paper, wait/set build on backend read/write
operations, while lock/unlock build on acquire/release. The higher-level
helpers retain their backend's scope and ordering requirements.
- Intel: OpenCL setup, correctness checks, and optional x86 modules.
- NVIDIA: ARM64 CUDA setup, correctness checks, and optional driver extensions.
- Implementation status: binding configurations, read/write policies, default plans, and control-operation availability.
- Control plane: API contracts, backend support, and synchronization examples.
- Benchmarks: performance tools, page-system prerequisites, and result interpretation.
- Platform extension: adding a backend or device profile and validating its capabilities.
Contributions are welcome: add support for a UMA platform, implement a data or control operation, improve tests and examples, or report and fix bugs.
For hardware support, start with the platform extension guide. Describe the binding and operation contracts, keep unsupported combinations explicit, and include correctness results from the target hardware. Bug reports should include the platform, driver/toolkit versions, build options, and a minimal reproducer.
libumsh is licensed under Apache-2.0, with the following exceptions:
the three helper kernel modules and shared ioctl header are dual-licensed under
Apache-2.0 OR GPL-2.0-only, and the NVIDIA UVM overlay retains its MIT license
and upstream copyright notices. The Figure 3 image is distributed under
CC BY 4.0.
See Licensing for the exact file scope and third-party notices.
If you use libumsh in your research, please cite:
% TODOThe artifacts of libumsh are published on GitHub and Zenodo.
