Summary
I'd like to propose an application-managed external dma-buf registration API, so GPU runtimes that are not natively integrated into uGDS (Intel Level Zero first among them) can participate in NVMe P2P DMA by exporting their buffers as dma-buf file descriptors and handing those fds to uGDS.
Unlike the existing import path (#1 for AMD HIP) and the fd export direction (#16), the library would not link or dlopen any Intel oneAPI component. The application performs the export itself and passes the fd plus offset to a new registration entry point.
Background
uGDS currently supports NVIDIA CUDA and AMD HIP/ROCm through dedicated backends. Intel discrete GPUs (Arc Pro B60/B50 on the Xe kernel driver) expose their device memory through the standard Linux DMA-buf framework, but there is no path to register such a buffer with uGDS; applications fall back to copies through system memory.
Two properties make a generic external-import path worth adding:
- On the Xe driver,
dma_buf_map_attachment against the NVMe device returns either peer BAR bus addresses or system-memory addresses, so the kernel can verify with authority whether a mapping is truly P2P before any DMA address reaches userspace. Heuristic fd-based checks (fdinfo) are not reliable for this.
- Untrusted fds could otherwise pin unbounded GPU memory, so the kernel side needs per-file and global pin accounting with administrable limits.
The driver currently has no explicit vendor-neutral dma-buf-only build mode. BUILD_DMABUF=1 adds a standalone importer build without NVIDIA driver sources and without selecting the HIP backend.
Proposed API
uGDSDmabufRegParams_t params = {
.struct_size = sizeof(params),
.version = UGDS_DMABUF_REG_PARAMS_VERSION_1,
.dmabuf_fd = export_fd,
.dmabuf_offset = (uintptr_t)ptr - (uintptr_t)base,
.flags = UGDS_DMABUF_REQUIRE_P2P, /* optional strict mode */
};
uGDSBufRegisterDmabuf(ptr, size, ¶ms);
uGDSDmabufCaps_t caps;
uGDSQueryDmabufSupport(&caps); /* runtime feature detection */
Kernel side additions: a versioned V2 mapping ioctl with mapping-class output (PEER_BAR / SYSTEM / UNKNOWN plus failure diagnostics), a capability ioctl, kernel-side P2P BAR classification of every mapped SG address, and read-only pin-limit module parameters.
Status
I have a working implementation stacked on #30 (SGL support) on my fork. Builds verified on CUDA-only, HIP-only, and dual-backend configurations. The kernel module also builds in a vendor-neutral dma-buf-only configuration with BUILD_DMABUF=1, and a CUDA/ROCm-free library-only compile check passes. CMake already accepts the library-only configuration on the base branch; this PR removes the unconditional cuda_runtime.h dependency that prevented it from compiling. Runtime validation on Intel Arc hardware is still pending. I'd be happy to restructure or adjust the approach based on feedback.
Summary
I'd like to propose an application-managed external dma-buf registration API, so GPU runtimes that are not natively integrated into uGDS (Intel Level Zero first among them) can participate in NVMe P2P DMA by exporting their buffers as dma-buf file descriptors and handing those fds to uGDS.
Unlike the existing import path (#1 for AMD HIP) and the fd export direction (#16), the library would not link or dlopen any Intel oneAPI component. The application performs the export itself and passes the fd plus offset to a new registration entry point.
Background
uGDS currently supports NVIDIA CUDA and AMD HIP/ROCm through dedicated backends. Intel discrete GPUs (Arc Pro B60/B50 on the Xe kernel driver) expose their device memory through the standard Linux DMA-buf framework, but there is no path to register such a buffer with uGDS; applications fall back to copies through system memory.
Two properties make a generic external-import path worth adding:
dma_buf_map_attachmentagainst the NVMe device returns either peer BAR bus addresses or system-memory addresses, so the kernel can verify with authority whether a mapping is truly P2P before any DMA address reaches userspace. Heuristic fd-based checks (fdinfo) are not reliable for this.The driver currently has no explicit vendor-neutral dma-buf-only build mode.
BUILD_DMABUF=1adds a standalone importer build without NVIDIA driver sources and without selecting the HIP backend.Proposed API
Kernel side additions: a versioned V2 mapping ioctl with mapping-class output (PEER_BAR / SYSTEM / UNKNOWN plus failure diagnostics), a capability ioctl, kernel-side P2P BAR classification of every mapped SG address, and read-only pin-limit module parameters.
Status
I have a working implementation stacked on #30 (SGL support) on my fork. Builds verified on CUDA-only, HIP-only, and dual-backend configurations. The kernel module also builds in a vendor-neutral dma-buf-only configuration with
BUILD_DMABUF=1, and a CUDA/ROCm-free library-only compile check passes. CMake already accepts the library-only configuration on the base branch; this PR removes the unconditionalcuda_runtime.hdependency that prevented it from compiling. Runtime validation on Intel Arc hardware is still pending. I'd be happy to restructure or adjust the approach based on feedback.