You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
perf(rocm): add host-gap accounting per port unit to the decode profile #2148
scripts/rocm_decode_profile.py (#2061, PR #2086) ranks the #1814 ports by their share of decode GPU kernel time. That share implies a speedup ceiling of 1/(1 - share), but it ignores the host time between dispatches. For ports that replace many small dispatches the ceiling is wrong: the SSM port (#2067, PR #2099) measured 1.46x decode on granite-4.0-h-tiny and 1.45x on Nemotron-3-Nano, above the 1/(1 - 0.298) = 1.42x and 1/(1 - 0.200) = 1.25x the profile's 29.82% and 20.01% shares allow (docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md:114,117). The same page reports 3197 and 2008 dispatches per token and 4.8 and 5.1 ms per token of host gap for the two hybrids (line 73), but only as a whole-run number. The next ranking will underrate dispatch-heavy ports the same way.
Current Behavior
summarize() (scripts/rocm_decode_profile.py:400-521) computes host_gap_ms_per_token, host_gap_pct_of_wall and plain_host_gap_ms_per_token_est for the whole decode window only (lines 489-499).
Per port unit (UNIT_ROLES, line 350) it records fallback_share_pct, reached_default_share_pct, reached_optin_share_pct and dispatches per token (lines 453-466), all relative to gpu_sum (the sum of kernel durations, share() at line 437).
report() (lines 544-567) prints the unit table as reached (fallback) GPU-time shares.
Proposed Solution
Attribute idle time to the dispatch that ends it, then derive a wall-time share and ceiling per role and per port unit. Keep the existing GPU-time fields unchanged.
New pure function attribute_gaps(decode: list[Dispatch], roles: list[str | None], lo: int, hi: int) -> dict[str, int]: walk dispatches in start order, keep frontier = max(lo, max end so far); for each dispatch gap = max(0, start - frontier), added to that dispatch's role ("unattributed" for None). The idle tail from the last end to hi goes to "tail". The total equals wall - gpu_busy from busy_ns().
Summary additions: top-level role_host_gap_ms_per_token; per unit fallback_host_gap_ms_per_token, fallback_host_gap_us_per_dispatch, fallback_wall_share_pct = 100 * (unit kernel ns + unit gap ns) / wall, ceiling_gpu_share = 1 / (1 - fallback_share_pct / 100) and ceiling_wall = 1 / (1 - fallback_wall_share_pct / 100), both rounded to 2 decimals. When a plain run exists, plain_ceiling_wall_est scales the unit's gap by plain_host_gap_ms_per_token_est / host_gap_ms_per_token (the tracer inflates host gaps, not kernel time) and divides by the plain wall per token. Same fields for the reached-default roles.
report(): a second unit table with ceiling_wall (plain_ceiling_wall_est when present) and ceiling_gpu_share per run and unit.
Edge cases: zero tokens leave the per-token fields None as today; a unit with no dispatches reports 0 gap and a ceiling of 1.0; a share of 100% must not divide by zero (report None).
Docs: in docs/benchmarks.md (section at line 220) say the GPU share ignores launch gaps and that ceiling_wall is the bound to compare a port's measured speedup against. Add a dated note to the 2026-09-30 profile page that the perf(rocm): port ssm_update_kernel to HIP and fix ssm_kernel_available #2067 shares understated the gain, with the new ceilings.
Acceptance Criteria
Unit tests in tests/test_rocm_decode_profile.py (WindowTests): attribute_gaps on a synthetic trace with overlapping dispatches, a leading gap and an idle tail gives the expected per-role totals, and their sum equals wall - busy_ns(...); summarize() on the existing synthetic trace emits the new keys.
On gfx1151, scripts/rocm_decode_profile.sh for granite-4.0-h-tiny-4bit and NVIDIA-Nemotron-3-Nano-30B-A3B-4bit with MLXCEL_SSM_KERNEL=0 (the graph path) gives a perf(rocm): port ssm_update_kernel to HIP and fix ssm_kernel_available #2067plain_ceiling_wall_est at or above the measured 1.46x and 1.45x. If it does not, the PR explains the remaining difference before merging.
The report output and both docs changes are in the PR.
Part of #1801
Problem / Background
scripts/rocm_decode_profile.py(#2061, PR #2086) ranks the #1814 ports by their share of decode GPU kernel time. That share implies a speedup ceiling of 1/(1 - share), but it ignores the host time between dispatches. For ports that replace many small dispatches the ceiling is wrong: the SSM port (#2067, PR #2099) measured 1.46x decode on granite-4.0-h-tiny and 1.45x on Nemotron-3-Nano, above the 1/(1 - 0.298) = 1.42x and 1/(1 - 0.200) = 1.25x the profile's 29.82% and 20.01% shares allow (docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md:114,117). The same page reports 3197 and 2008 dispatches per token and 4.8 and 5.1 ms per token of host gap for the two hybrids (line 73), but only as a whole-run number. The next ranking will underrate dispatch-heavy ports the same way.Current Behavior
summarize()(scripts/rocm_decode_profile.py:400-521) computeshost_gap_ms_per_token,host_gap_pct_of_wallandplain_host_gap_ms_per_token_estfor the whole decode window only (lines 489-499).UNIT_ROLES, line 350) it recordsfallback_share_pct,reached_default_share_pct,reached_optin_share_pctand dispatches per token (lines 453-466), all relative togpu_sum(the sum of kernel durations,share()at line 437).report()(lines 544-567) prints the unit table asreached (fallback)GPU-time shares.Proposed Solution
Attribute idle time to the dispatch that ends it, then derive a wall-time share and ceiling per role and per port unit. Keep the existing GPU-time fields unchanged.
attribute_gaps(decode: list[Dispatch], roles: list[str | None], lo: int, hi: int) -> dict[str, int]: walk dispatches in start order, keepfrontier = max(lo, max end so far); for each dispatchgap = max(0, start - frontier), added to that dispatch's role ("unattributed"forNone). The idle tail from the last end tohigoes to"tail". The total equalswall - gpu_busyfrombusy_ns().role_host_gap_ms_per_token; per unitfallback_host_gap_ms_per_token,fallback_host_gap_us_per_dispatch,fallback_wall_share_pct = 100 * (unit kernel ns + unit gap ns) / wall,ceiling_gpu_share = 1 / (1 - fallback_share_pct / 100)andceiling_wall = 1 / (1 - fallback_wall_share_pct / 100), both rounded to 2 decimals. When a plain run exists,plain_ceiling_wall_estscales the unit's gap byplain_host_gap_ms_per_token_est / host_gap_ms_per_token(the tracer inflates host gaps, not kernel time) and divides by the plain wall per token. Same fields for the reached-default roles.report(): a second unit table withceiling_wall(plain_ceiling_wall_estwhen present) andceiling_gpu_shareper run and unit.Noneas today; a unit with no dispatches reports 0 gap and a ceiling of 1.0; a share of 100% must not divide by zero (reportNone).docs/benchmarks.md(section at line 220) say the GPU share ignores launch gaps and thatceiling_wallis the bound to compare a port's measured speedup against. Add a dated note to the 2026-09-30 profile page that the perf(rocm): port ssm_update_kernel to HIP and fix ssm_kernel_available #2067 shares understated the gain, with the new ceilings.Acceptance Criteria
tests/test_rocm_decode_profile.py(WindowTests):attribute_gapson a synthetic trace with overlapping dispatches, a leading gap and an idle tail gives the expected per-role totals, and their sum equalswall - busy_ns(...);summarize()on the existing synthetic trace emits the new keys.scripts/rocm_decode_profile.shfor granite-4.0-h-tiny-4bit and NVIDIA-Nemotron-3-Nano-30B-A3B-4bit withMLXCEL_SSM_KERNEL=0(the graph path) gives a perf(rocm): port ssm_update_kernel to HIP and fix ssm_kernel_available #2067plain_ceiling_wall_estat or above the measured 1.46x and 1.45x. If it does not, the PR explains the remaining difference before merging.Verification
Related: #2061, PR #2086, #2067, PR #2099, #1814.