Skip to content
XpuOSPublic

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

libumsh: Unified Memory Sharing for CPU–GPU Systems

License

Getting Started · Platform Support · Documentation · Citation · Artifacts

About

libumsh is a C++ library for sharing memory between CPUs and GPUs on unified memory architectures. It provides a common interface for memory binding, data-access policies, and control operations across Intel integrated GPUs and ARM64 NVIDIA systems.

Sharing physical memory does not by itself make CPU and GPU accesses coherent. libumsh separates how memory is allocated and mapped from how each endpoint reads, writes, and synchronizes. Applications can inspect the effective binding and available operations, then choose a protocol for their hardware.

Data copying uses two memory regions and an extra copy; direct access lets the sender and receiver use one shared region.

The library serves two kinds of communication:

  • Data-plane transmission: share bulk data such as model weights, images, and inference outputs, with binding and access policies selected for the transfer direction and kernel boundary.
  • Control-plane coordination: exchange flags and ownership state to coordinate CPU/GPU work, using signaling for producer/consumer dependencies and locks for shared resources.

Features

  • Explicit memory binding: request page size, CPU/GPU cache policy, and coherency, and inspect the effective attributes and backing verification.
  • Data-access policies: describe reads and writes in terms of cache invalidation, flushing, and bypass, with backend-specific implementations.
  • Control operations: use acquire/release for ownership, lock/unlock for mutual exclusion, and wait/set for signaling on supported platforms.
  • Capability-aware selection: query available implementations and use platform defaults with strict checks for unsupported combinations.
  • Tools and extensibility: use correctness tests and standalone benchmarks to validate a configuration, or add a backend or device profile.

Platform Support

libumsh builds natively on Linux. Choose the deployment guide for your platform:

Platform Build target Backend Deployment guide
Intel Raptor Lake RAPTOR OpenCL USM Intel
Intel Arrow Lake ARROW OpenCL USM Intel
Intel Lunar Lake LUNAR OpenCL USM Intel
Intel Meteor Lake METEOR OpenCL USM Intel
NVIDIA Jetson AGX Orin ORIN CUDA (sm_87) NVIDIA
NVIDIA DGX Spark / GB10 SPARK CUDA (sm_121) NVIDIA

The CUDA backend targets the ARM64 UMA systems above. See implementation status for available binding configurations, read/write policies, default plans, and control operations. For another platform, follow the extension guide.

Getting Started

Build

Build on the target Linux machine with CMake 3.24+, a C++23 compiler, Python 3, and the backend's development headers and runtime. Intel requires OpenCL with cl_intel_unified_shared_memory; NVIDIA requires a CUDA toolkit that can generate native code for the target GPU. The platform guides provide the installation details.

From the repository root:

cmake -S . -B build-release \
  -DCMAKE_BUILD_TYPE=Release \
  -DUMSH_TARGET_PLATFORM=AUTO \
  -DUMSH_BUILD_EXAMPLES=ON \
  -DUMSH_BUILD_KERNEL_MODULE=OFF
cmake --build build-release --parallel
ctest --test-dir build-release --output-on-failure

AUTO detects the native platform. An explicit build target from the table selects its identity; it does not enable cross-compilation. Examples fetch a pinned Abseil dependency during configuration. Set UMSH_BUILD_EXAMPLES=OFF to build only the library and enabled host tests without fetching Abseil; BUILD_TESTING=OFF disables those tests.

The default build uses the stock GPU driver. Optional kernel modules enable specific backing or cache-maintenance extensions; their build and loading instructions are in the platform guides.

Allocate shared memory

The Core interface creates a binding from a size and explicit options:

#include <umsh/core/core.hpp>
#include <umsh/utils/error.hpp>

int main() {
    auto core = UMSH_CHECK(umsh::core::create_core());
    UMSH_CHECK(core->initialize(0));

    umsh::binding::BindOptions options{
        .page_size = umsh::binding::PageSize::Default,
        .cpu_cache = umsh::binding::CpuCachePolicy::WriteBack,
        .gpu_cache = umsh::binding::GpuCachePolicy::RuntimeDefault,
        .coherency = umsh::binding::Coherency::RuntimeDefault,
    };
    auto memory = UMSH_CHECK(core->bind(4096, options));

    // Use memory.host_ptr on the CPU and the backend's GPU access interface.
    // Finish all CPU/GPU users before releasing the allocation.
    UMSH_CHECK(core->unbind(memory));
}

The APIs return umsh::Result<T>; UMSH_CHECK is a convenience for examples. Applications can inspect the returned error and recover instead.

core->binding_info(memory) describes the requested and effective binding. Check same_backing and its verification method when shared physical backing is required. Successful allocation alone does not establish a concurrent CPU/GPU data-transfer protocol.

Link with libumsh

Install the library and headers:

cmake --install build-release --prefix /path/to/umsh-install

Save the allocation example as main.cpp in a separate application directory and link the exported CMake target:

cmake_minimum_required(VERSION 3.24)
project(my_app LANGUAGES CXX)
find_package(umsh CONFIG REQUIRED)
add_executable(my_app main.cpp)
target_link_libraries(my_app PRIVATE umsh::umsh)

Configure and build from that application directory:

cmake -S . -B build -DCMAKE_PREFIX_PATH=/path/to/umsh-install
cmake --build build --parallel

The exported target carries the backend/platform definitions and C++23 requirement. Applications with CUDA device code additionally need a CUDA-enabled CMake project.

Examples, Tests, and Benchmarks

Start with the binding and capability probe:

build-release/examples/custom/policy_probe/umsh_policy_probe

It reports the effective binding and available policies. Inspect the same_backing and verification fields separately from the probe's PASS. For CPU/GPU communication, run the hardware correctness tests described below.

  • Control operations: CPU/GPU signaling and ownership examples, with tests for publication, bounded waiting, and mutual exclusion.
  • Policy-transfer tests: exercise reads and writes while a GPU kernel is running. The matrix includes weak policies and unsupported combinations, so not every row is expected to pass.
  • File loading: load a file into a caller-owned shared allocation using buffered, mmap, or direct I/O.
  • Standalone benchmarks: measure binding costs, CPU/GPU reads and writes, GPU contention, and SSD loading.

ctest covers host-side default selection and, with examples enabled, storage-report handling. GPU correctness tests run explicitly on the target machine; follow the Intel or NVIDIA commands before measuring performance.

Architecture and Workflow

Applications bind memory through Core, inspect its capabilities, and select the data and control operations needed by their communication protocol. Native backends supply the allocation, cache-maintenance, and device-access mechanisms.

Component Responsibility
Core and backends Allocate/import shared memory, report effective attributes, and expose backend accessors.
Policy contracts Describe reads, writes, scopes, support levels, and implementation identities.
Platform defaults Specify compile-time transfer plans and reject unavailable selections.
Control layer and implementations Provide ownership primitives and signaling/locking helpers using dedicated control memory.
Optional kernel extensions Supply additional memory attributes and cache-maintenance mechanisms.

Read policies are Alone, AfterInvalidate, or Bypass; write policies are Alone, ThenFlush, or Bypass. Query support before choosing an implementation. CUDA cache operators and fences are separate raw characterization primitives, with their own semantics.

A default plan describes a platform's intended binding and endpoint operations. If a required binding or implementation is unavailable, strict selection rejects that plan. New platforms need their own validated defaults.

Control words have a separate lifetime from payload memory. Ownership and signaling order accesses, while the payload's policy supplies any required cache maintenance. Complete all CPU/GPU users before releasing either allocation; see the control-plane guide for examples and ordering rules.

As described in Section 5 of the paper, wait/set build on backend read/write operations, while lock/unlock build on acquire/release. The higher-level helpers retain their backend's scope and ordering requirements.

Documentation

  • Intel: OpenCL setup, correctness checks, and optional x86 modules.
  • NVIDIA: ARM64 CUDA setup, correctness checks, and optional driver extensions.
  • Implementation status: binding configurations, read/write policies, default plans, and control-operation availability.
  • Control plane: API contracts, backend support, and synchronization examples.
  • Benchmarks: performance tools, page-system prerequisites, and result interpretation.
  • Platform extension: adding a backend or device profile and validating its capabilities.

Contributing

Contributions are welcome: add support for a UMA platform, implement a data or control operation, improve tests and examples, or report and fix bugs.

For hardware support, start with the platform extension guide. Describe the binding and operation contracts, keep unsupported combinations explicit, and include correctness results from the target hardware. Bug reports should include the platform, driver/toolkit versions, build options, and a minimal reproducer.

License

libumsh is licensed under Apache-2.0, with the following exceptions: the three helper kernel modules and shared ioctl header are dual-licensed under Apache-2.0 OR GPL-2.0-only, and the NVIDIA UVM overlay retains its MIT license and upstream copyright notices. The Figure 3 image is distributed under CC BY 4.0. See Licensing for the exact file scope and third-party notices.

Citation

If you use libumsh in your research, please cite:

% TODO

The artifacts of libumsh are published on GitHub and Zenodo.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages