Skip to content

Interrupt mode: known limitations and follow-up work #26

Description

@chengy-sysu

PR #25 landed a working MSI-X + eventfd interrupt mode for synchronous IO. This issue tracks the known limitations and potential follow-up work, roughly ordered by impact.

Functional scope

1. Sync-only interrupt support

Interrupt vectors are only assigned to sync QPs. The batch CQ (ugds_handle.cpp:289) is always polling, and async IO inherits the sync path without its own eventfd integration. UGDS_INTERRUPT_MODE=1 means "sync IO uses interrupts", not "all of uGDS uses interrupts".

2. One eventfd per vector — no sharing or multi-process

irq.c returns -EBUSY if a vector is already armed. Since every handle starts requesting from vector 0, the first handle typically claims all sync vectors; subsequent handles silently fall back to polling. This effectively limits interrupt mode to one handle per controller.

3. Mode is decided at handle creation time

The environment variable is read once during uGDSHandleRegister. There is no way to:

  • Dynamically switch between polling and interrupt
  • Select mode per IO size (like GDS poll_mode_max_size_kb)
  • Adapt based on load
  • Query or configure mode via a public API

4. Large-IO chunking: one interrupt per MDTS chunk

When an IO exceeds max_transfer_size, do_io_internal splits it into MDTS-sized chunks and waits for each chunk sequentially. In interrupt mode every chunk triggers a separate IRQ → eventfd → ppoll cycle. There is no batched-submit or coalesced-wait path for multi-chunk IOs.

Performance

5. Sync QP is QD=1

Each QP does submit-one / wait-one, so there is no opportunity to accumulate multiple completions per interrupt. Interrupt mode saves CPU but does not improve throughput.

6. Small-IO throughput/latency penalty

Every IO adds IRQ delivery, eventfd signal, scheduler wakeup, ppoll, and read syscalls. At 4KB this overhead dominates — interrupt mode is measurably slower than busy-poll. There is no hybrid strategy (e.g., spin for N µs then fall back to interrupt).

7. No interrupt coalescing or CPU affinity

NVMe interrupt coalescing registers are not programmed. IRQ CPU affinity is left to the kernel / irqbalance with no NUMA or PCIe-topology-aware pinning.

Fallback and observability

8. Silent per-QP fallback

If eventfd creation or NVM_REGISTER_INTERRUPT fails for a QP, that QP silently continues in polling mode. The handle can end up in a mixed state (some QPs interrupt, some polling) with no way for the application to detect this.

9. Benchmark completion label may be inaccurate

The benchmark prints interrupt (MSI-X) based on the environment variable, not the actual registration result. It can claim interrupt mode when:

  • The device has no MSI-X vectors (runtime fell back to polling)
  • Vectors were already claimed by another handle
  • The IO path is batch or async (always polling)

10. End-to-end test cannot prove IRQs actually fired

test_interrupt_mode.cu verifies data correctness. If the runtime silently fell back to polling, the test still passes. Confirming real MSI-X delivery currently requires inspecting /proc/interrupts manually.

Hardware and resources

11. MSI-X only, max 64 vectors

No MSI or legacy INTx fallback. The driver hard-caps at 64 vectors. If num_vectors < num_sync_qps, the excess QPs use polling.

12. MSI-X allocated eagerly at probe

ugds_irq_ctrl_init runs during PCI probe regardless of whether any process enables interrupt mode. The vectors and kernel IRQ descriptors are allocated even in poll-only deployments.

13. Limited hardware validation

Tested on NVIDIA A100 + Samsung 990 PRO (CUDA, Linux 6.8). The HIP/ROCm backend compiles with interrupt support but has not been validated on AMD GPU hardware.

Potential follow-up

  • Batch / async interrupt support
  • Multi-CQ vector sharing (multiple CQs per MSI-X vector)
  • NVMe interrupt coalescing configuration
  • Lazy MSI-X allocation (allocate on first REGISTER_INTERRUPT, not at probe)
  • IRQ CPU affinity control (NUMA-aware vector assignment)
  • Multi-process / per-QP dynamic vector registration
  • Hybrid poll-then-interrupt wait strategy
  • Coalesced wait for multi-chunk large IOs
  • Per-QP mode query API
  • Automated IRQ-fired verification in tests

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions