Summary
I'd like to propose scatter-gather (vectored) I/O support, so a single call can stream data across multiple registered GPU memory regions instead of issuing one call per region. The segment list is split into MDTS-bounded command windows using NVMe PRP lists, and the same windowed engine serves the synchronous, batch, and asynchronous API families.
Background
Today the public API takes one registered buffer per call (or per batch element / async request). Workloads that keep several GPU regions hot (KV-cache style paging, weight streaming, log-structured writes) either issue many small commands or re-register larger spans, and neither option composes well with the batch and async infrastructure that already exists.
A vectored entry point on top of the existing registration model fits more naturally: segments reference already-registered bases with offsets, so no new memory management is introduced, and the controller-side constraints (MDTS, MPS alignment, PRP contiguity) are handled inside the library.
Proposed API
typedef struct uGDSIoSegment {
void* base; /* registered buffer base (exact registry key) */
off_t offset; /* byte offset inside base; must be MPS-aligned */
size_t size; /* transfer size in bytes; must be block-multiple */
} uGDSIoSegment_t;
ssize_t uGDSReadv(uGDSHandle_t fh, const uGDSIoSegment_t* segs,
unsigned nr_segs, off_t file_offset);
ssize_t uGDSWritev(uGDSHandle_t fh, const uGDSIoSegment_t* segs,
unsigned nr_segs, off_t file_offset);
uGDSError_t uGDSBatchIOSubmitv(uGDSBatchHandle_t batch, unsigned nr,
uGDSIOSegParams_t* iocb, unsigned flags);
uGDSError_t uGDSReadvAsync(uGDSHandle_t fh, uGDSIoSegment_t* segs,
unsigned nr_segs, off_t* file_offset_p,
ssize_t* bytes_read_p, void* stream);
uGDSError_t uGDSWritevAsync(uGDSHandle_t fh, uGDSIoSegment_t* segs,
unsigned nr_segs, off_t* file_offset_p,
ssize_t* bytes_written_p, void* stream);
Limits: UGDS_IOV_MAX (1024) segments per sync/async call, UGDS_BATCH_IOV_MAX (128) per batch element. Plain and vectored elements may be mixed on the same batch handle.
The implementation uses NVMe PRP lists rather than native SGL descriptors, since the underlying libnvm layer does not support SGL command building and PRP already expresses non-contiguous pages.
To keep review manageable, the work is split into a stacked series covering (1) buffer registration hardening, (2) the shared streaming engine and sync vectored API, (3) the batch lifecycle and vectored batch submit, and (4) vectored async.
Status
The full implementation is complete and builds clean on CUDA-only, HIP-only, and dual-backend configurations. I will publish the stacked series against this issue.
Summary
I'd like to propose scatter-gather (vectored) I/O support, so a single call can stream data across multiple registered GPU memory regions instead of issuing one call per region. The segment list is split into MDTS-bounded command windows using NVMe PRP lists, and the same windowed engine serves the synchronous, batch, and asynchronous API families.
Background
Today the public API takes one registered buffer per call (or per batch element / async request). Workloads that keep several GPU regions hot (KV-cache style paging, weight streaming, log-structured writes) either issue many small commands or re-register larger spans, and neither option composes well with the batch and async infrastructure that already exists.
A vectored entry point on top of the existing registration model fits more naturally: segments reference already-registered bases with offsets, so no new memory management is introduced, and the controller-side constraints (MDTS, MPS alignment, PRP contiguity) are handled inside the library.
Proposed API
Limits:
UGDS_IOV_MAX(1024) segments per sync/async call,UGDS_BATCH_IOV_MAX(128) per batch element. Plain and vectored elements may be mixed on the same batch handle.The implementation uses NVMe PRP lists rather than native SGL descriptors, since the underlying libnvm layer does not support SGL command building and PRP already expresses non-contiguous pages.
To keep review manageable, the work is split into a stacked series covering (1) buffer registration hardening, (2) the shared streaming engine and sync vectored API, (3) the batch lifecycle and vectored batch submit, and (4) vectored async.
Status
The full implementation is complete and builds clean on CUDA-only, HIP-only, and dual-backend configurations. I will publish the stacked series against this issue.