Skip to content

perf(rocm): add host-gap accounting per port unit to the decode profile #2148

Description

@inureyes

Part of #1801

Problem / Background

scripts/rocm_decode_profile.py (#2061, PR #2086) ranks the #1814 ports by their share of decode GPU kernel time. That share implies a speedup ceiling of 1/(1 - share), but it ignores the host time between dispatches. For ports that replace many small dispatches the ceiling is wrong: the SSM port (#2067, PR #2099) measured 1.46x decode on granite-4.0-h-tiny and 1.45x on Nemotron-3-Nano, above the 1/(1 - 0.298) = 1.42x and 1/(1 - 0.200) = 1.25x the profile's 29.82% and 20.01% shares allow (docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md:114,117). The same page reports 3197 and 2008 dispatches per token and 4.8 and 5.1 ms per token of host gap for the two hybrids (line 73), but only as a whole-run number. The next ranking will underrate dispatch-heavy ports the same way.

Current Behavior

  • summarize() (scripts/rocm_decode_profile.py:400-521) computes host_gap_ms_per_token, host_gap_pct_of_wall and plain_host_gap_ms_per_token_est for the whole decode window only (lines 489-499).
  • Per port unit (UNIT_ROLES, line 350) it records fallback_share_pct, reached_default_share_pct, reached_optin_share_pct and dispatches per token (lines 453-466), all relative to gpu_sum (the sum of kernel durations, share() at line 437).
  • report() (lines 544-567) prints the unit table as reached (fallback) GPU-time shares.

Proposed Solution

Attribute idle time to the dispatch that ends it, then derive a wall-time share and ceiling per role and per port unit. Keep the existing GPU-time fields unchanged.

  • New pure function attribute_gaps(decode: list[Dispatch], roles: list[str | None], lo: int, hi: int) -> dict[str, int]: walk dispatches in start order, keep frontier = max(lo, max end so far); for each dispatch gap = max(0, start - frontier), added to that dispatch's role ("unattributed" for None). The idle tail from the last end to hi goes to "tail". The total equals wall - gpu_busy from busy_ns().
  • Summary additions: top-level role_host_gap_ms_per_token; per unit fallback_host_gap_ms_per_token, fallback_host_gap_us_per_dispatch, fallback_wall_share_pct = 100 * (unit kernel ns + unit gap ns) / wall, ceiling_gpu_share = 1 / (1 - fallback_share_pct / 100) and ceiling_wall = 1 / (1 - fallback_wall_share_pct / 100), both rounded to 2 decimals. When a plain run exists, plain_ceiling_wall_est scales the unit's gap by plain_host_gap_ms_per_token_est / host_gap_ms_per_token (the tracer inflates host gaps, not kernel time) and divides by the plain wall per token. Same fields for the reached-default roles.
  • report(): a second unit table with ceiling_wall (plain_ceiling_wall_est when present) and ceiling_gpu_share per run and unit.
  • Edge cases: zero tokens leave the per-token fields None as today; a unit with no dispatches reports 0 gap and a ceiling of 1.0; a share of 100% must not divide by zero (report None).
  • Docs: in docs/benchmarks.md (section at line 220) say the GPU share ignores launch gaps and that ceiling_wall is the bound to compare a port's measured speedup against. Add a dated note to the 2026-09-30 profile page that the perf(rocm): port ssm_update_kernel to HIP and fix ssm_kernel_available #2067 shares understated the gain, with the new ceilings.

Acceptance Criteria

  • Unit tests in tests/test_rocm_decode_profile.py (WindowTests): attribute_gaps on a synthetic trace with overlapping dispatches, a leading gap and an idle tail gives the expected per-role totals, and their sum equals wall - busy_ns(...); summarize() on the existing synthetic trace emits the new keys.
  • On gfx1151, scripts/rocm_decode_profile.sh for granite-4.0-h-tiny-4bit and NVIDIA-Nemotron-3-Nano-30B-A3B-4bit with MLXCEL_SSM_KERNEL=0 (the graph path) gives a perf(rocm): port ssm_update_kernel to HIP and fix ssm_kernel_available #2067 plain_ceiling_wall_est at or above the measured 1.46x and 1.45x. If it does not, the PR explains the remaining difference before merging.
  • The report output and both docs changes are in the PR.

Verification

python3 -m unittest tests/test_rocm_decode_profile.py -v
cargo build --release --features rocm --bin mlxcel-bench-decode
MLXCEL_SSM_KERNEL=0 scripts/rocm_decode_profile.sh models/mlx/granite-4.0-h-tiny-4bit models/mlx/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit
python3 scripts/rocm_decode_profile.py report benchmarks/rocm_profiles/gfx1151_<commit>

Related: #2061, PR #2086, #2067, PR #2099, #1814.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:benchmarkBenchmark harness and performance measurement (bench_*.sh, /update-benchmarks)platform:linuxLinux (CUDA / packaging) specificpriority:lowLow prioritystatus:readyReady to be worked ontype:enhancementNew features, capabilities, or significant additions

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions