Skip to content

[Feature] Add scatter-gather (vectored) I/O support #48

Description

@potatogim

Summary

I'd like to propose scatter-gather (vectored) I/O support, so a single call can stream data across multiple registered GPU memory regions instead of issuing one call per region. The segment list is split into MDTS-bounded command windows using NVMe PRP lists, and the same windowed engine serves the synchronous, batch, and asynchronous API families.

Background

Today the public API takes one registered buffer per call (or per batch element / async request). Workloads that keep several GPU regions hot (KV-cache style paging, weight streaming, log-structured writes) either issue many small commands or re-register larger spans, and neither option composes well with the batch and async infrastructure that already exists.

A vectored entry point on top of the existing registration model fits more naturally: segments reference already-registered bases with offsets, so no new memory management is introduced, and the controller-side constraints (MDTS, MPS alignment, PRP contiguity) are handled inside the library.

Proposed API

typedef struct uGDSIoSegment {
    void*   base;     /* registered buffer base (exact registry key) */
    off_t   offset;   /* byte offset inside base; must be MPS-aligned */
    size_t  size;     /* transfer size in bytes; must be block-multiple */
} uGDSIoSegment_t;

ssize_t     uGDSReadv(uGDSHandle_t fh, const uGDSIoSegment_t* segs,
                      unsigned nr_segs, off_t file_offset);
ssize_t     uGDSWritev(uGDSHandle_t fh, const uGDSIoSegment_t* segs,
                       unsigned nr_segs, off_t file_offset);

uGDSError_t uGDSBatchIOSubmitv(uGDSBatchHandle_t batch, unsigned nr,
                               uGDSIOSegParams_t* iocb, unsigned flags);

uGDSError_t uGDSReadvAsync(uGDSHandle_t fh, uGDSIoSegment_t* segs,
                           unsigned nr_segs, off_t* file_offset_p,
                           ssize_t* bytes_read_p, void* stream);
uGDSError_t uGDSWritevAsync(uGDSHandle_t fh, uGDSIoSegment_t* segs,
                            unsigned nr_segs, off_t* file_offset_p,
                            ssize_t* bytes_written_p, void* stream);

Limits: UGDS_IOV_MAX (1024) segments per sync/async call, UGDS_BATCH_IOV_MAX (128) per batch element. Plain and vectored elements may be mixed on the same batch handle.

The implementation uses NVMe PRP lists rather than native SGL descriptors, since the underlying libnvm layer does not support SGL command building and PRP already expresses non-contiguous pages.

To keep review manageable, the work is split into a stacked series covering (1) buffer registration hardening, (2) the shared streaming engine and sync vectored API, (3) the batch lifecycle and vectored batch submit, and (4) vectored async.

Status

The full implementation is complete and builds clean on CUDA-only, HIP-only, and dual-backend configurations. I will publish the stacked series against this issue.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions