From 0908c8883663b3489ff5fbc5a9b3bf00517c4474 Mon Sep 17 00:00:00 2001 From: Jeongkyu Shin Date: Wed, 30 Sep 2026 22:59:05 +0900 Subject: [PATCH 1/6] test(rocm): run MLX gather_mm numeric tests on every GPU backend The three grouped_gemm_numeric_tests gated on Metal or CUDA, so on ROCm they returned before touching the GPU and MLX's ROCm GatherMM was never checked. Run by exact name on gfx1151 with the gate widened, all three pass against the f64 host reference, and a rocprofv3 kernel trace shows the overlay's gather_batched_gemm_kernel (f32, bf16, f16) and a hipBLASLt GEMM for the sorted single-row case, so the pass is not vacuous. With the reference pointed at the wrong expert, all three fail at the value assertion. The gates now read gpu_backend_available(), the file leaves BACKEND_ENUMERATION_TODO in check_kernel_port_dispatch.py, and the checker reports 0 awaiting a predicate. Metal and CUDA still run the tests (not run here). Refs #2061 --- scripts/ci/check_kernel_port_dispatch.py | 21 +++++++------------ .../src/grouped_gemm_numeric_tests.rs | 12 +++++++---- 2 files changed, 16 insertions(+), 17 deletions(-) diff --git a/scripts/ci/check_kernel_port_dispatch.py b/scripts/ci/check_kernel_port_dispatch.py index 8f7d73bbf..37f2246a0 100755 --- a/scripts/ci/check_kernel_port_dispatch.py +++ b/scripts/ci/check_kernel_port_dispatch.py @@ -58,10 +58,10 @@ ``*_available()`` predicate for "does this backend have this port", or ``gpu_backend_available()`` for "is there a GPU at all". -Rust files listed in ``BACKEND_ENUMERATION_TODO`` are exempt from rule 4, each -with the predicate it is waiting on. Unlike ``UNCONVERTED`` these are not a -convention that was skipped: they need a support predicate to be exported first -(#1814). +Rust files listed in ``BACKEND_ENUMERATION_TODO`` (empty since #2061) are +exempt from rule 4, each with the predicate it is waiting on. Unlike +``UNCONVERTED`` these are not a convention that was skipped: they need a support +predicate to be exported first (#1814). """ from __future__ import annotations @@ -88,15 +88,10 @@ # # Not a general exemption list: every entry names a predicate that does not exist # yet, so the entry disappears when that predicate lands rather than when someone -# remembers to look. -BACKEND_ENUMERATION_TODO = { - # Gates MLX's own `gather_mm`, not an mlxcel port, so no port table applies. - # `gpu_backend_available()` would be the honest predicate, but whether MLX's - # ROCm backend implements the grouped-GEMM path is unverified here and - # widening the gate on an assumption is how the abort in #1806 was reached - # (#1814). - "src/lib/mlxcel-core/src/grouped_gemm_numeric_tests.rs", -} +# remembers to look. Empty since #2061, which ran the last entry's tests +# (`grouped_gemm_numeric_tests.rs`, MLX's own `gather_mm`) on gfx1151, saw them +# pass, and moved them to `gpu_backend_available()`. +BACKEND_ENUMERATION_TODO: set[str] = set() # The helper's own translation units, which are allowed to name the backend. HELPERS = { diff --git a/src/lib/mlxcel-core/src/grouped_gemm_numeric_tests.rs b/src/lib/mlxcel-core/src/grouped_gemm_numeric_tests.rs index 6084527dc..225ad6241 100644 --- a/src/lib/mlxcel-core/src/grouped_gemm_numeric_tests.rs +++ b/src/lib/mlxcel-core/src/grouped_gemm_numeric_tests.rs @@ -38,7 +38,11 @@ //! `cutlass_grouped_gemm_unaligned` for the sorted single-row case that //! non-quantized MoE decode takes. //! -//! GPU-only (Metal or CUDA); these skip on a CPU-only build. +//! GPU-only; these skip on a CPU-only build. They run on every GPU backend +//! (`gpu_backend_available()`), ROCm included: MLX's ROCm overlay implements +//! `GatherMM` itself (`gather_batched_gemm_kernel`, and hipBLASLt for the +//! sorted single-row segments), and all three passed on gfx1151 with the +//! kernels confirmed in a `rocprofv3` trace (#2061). use super::*; @@ -184,7 +188,7 @@ impl Case { /// `n = 703` the unaligned arm, which are separate template instantiations. #[test] fn gather_mm_matches_dense_per_expert_reference() { - if !crate::metal_is_available() && !crate::cuda_is_available() { + if !crate::gpu_backend_available() { return; } @@ -283,7 +287,7 @@ fn gather_mm_matches_dense_per_expert_reference() { /// every element of a slab by a whole multiple rather than by a rounding error. #[test] fn gather_mm_selects_the_indexed_expert() { - if !crate::metal_is_available() && !crate::cuda_is_available() { + if !crate::gpu_backend_available() { return; } @@ -350,7 +354,7 @@ fn gather_mm_selects_the_indexed_expert() { /// ever reached a tensor-core configuration. #[test] fn gather_mm_half_precision_matches_reference() { - if !crate::metal_is_available() && !crate::cuda_is_available() { + if !crate::gpu_backend_available() { return; } From 10bee8567d1c12e9f40f4f0dbc89c0db19e29239 Mon Sep 17 00:00:00 2001 From: Jeongkyu Shin Date: Wed, 30 Sep 2026 22:59:16 +0900 Subject: [PATCH 2/6] feat(bench): add a ROCm per-kernel decode profile harness ROCm had end-to-end tok/s only, so the #1814 kernel ports had no measured order. This adds what a per-kernel decode profile needs, as reusable tooling rather than a one-off: - scripts/rocm_gpu_guard.sh: the idle-GPU guard the #2056 baseline described, as a script (90 s with /sys/class/kfd/kfd/proc empty and no compiler, a 1 Hz monitor that ignores the command's own GPU processes, rerun on contention, every sample logged). - mlxcel-bench-decode: --temperature and --top-p (default greedy, unchanged), and MLXCEL_BENCH_PHASE_MARKS=1, which prints the warmup, measured, decode-start and end times on CLOCK_MONOTONIC and CLOCK_BOOTTIME so a trace can be cut to the measured decode by timestamp. - scripts/rocm_decode_profile.sh runs a plain and a rocprofv3 --kernel-trace --hip-graph-trace --stats run per model under the guard; scripts/rocm_decode_profile.py cuts the decode window, reports GPU time and host gap per token, profiler cost, top kernels, and attributes dispatches to mlxcel ops and #1814 port units by dispatch order, with which ports mlxcel actually reaches per model. - tests/test_rocm_decode_profile.py covers the guard against a fake KFD directory and the cut and attribution rules on synthetic steps; docs/benchmarks.md documents the harness. Refs #2061 --- docs/benchmarks.md | 38 ++ scripts/rocm_decode_profile.py | 595 ++++++++++++++++++++++++++++ scripts/rocm_decode_profile.sh | 133 +++++++ scripts/rocm_gpu_guard.sh | 175 ++++++++ src/bin/bench_decode.rs | 59 ++- src/bin/bench_decode/phase_marks.rs | 108 +++++ tests/test_rocm_decode_profile.py | 225 +++++++++++ 7 files changed, 1326 insertions(+), 7 deletions(-) create mode 100644 scripts/rocm_decode_profile.py create mode 100755 scripts/rocm_decode_profile.sh create mode 100755 scripts/rocm_gpu_guard.sh create mode 100644 src/bin/bench_decode/phase_marks.rs create mode 100644 tests/test_rocm_decode_profile.py diff --git a/docs/benchmarks.md b/docs/benchmarks.md index 0347e0c51..5c82ff175 100644 --- a/docs/benchmarks.md +++ b/docs/benchmarks.md @@ -201,6 +201,44 @@ MLXLM_PYTHON=/bin/python LD_LIBRARY_PATH=/opt/rocm/lib \ `bench_mlxlm.py`'s `--big-cooldown` replaces the normal cooldown after a big model where `bench_decode.sh` adds it, so 60 there matches 30 + 30 here. +`scripts/rocm_gpu_guard.sh -- ` is that check as a script: it waits +for 90 s with `/sys/class/kfd/kfd/proc` empty and no compiler process, runs the +command under a 1 Hz monitor, reruns it if any sample shows another GPU process +or a compiler, and appends every sample to `--log`, which is the evidence a +published page cites. + +### ROCm per-kernel decode profile (issue #2061) + +`scripts/rocm_decode_profile.sh` answers where ROCm decode time goes, per +kernel. For each model it takes one plain `mlxcel-bench-decode` run and one +under `rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv`, both at the +default pp512/tg128 shape and both through `rocm_gpu_guard.sh`. The bench runs +with `MLXCEL_BENCH_PHASE_MARKS=1`, which prints the host-clock time of its +warmup, measured prefill, measured decode and end +(`src/bin/bench_decode/phase_marks.rs`), and +`scripts/rocm_decode_profile.py` keeps the dispatches that start inside the +measured decode. From those it writes a per-kernel table +(`_decode_kernels.csv`: calls, GPU time, share of decode GPU time, kernel +class, #1814 port unit) and a summary (`_summary.json`: decode GPU time and +host gap per token, profiler overhead, share per port unit), next to +rocprofv3's own whole-process `_kernel_stats.csv`. + +```bash +cargo build --release --features rocm --bin mlxcel-bench-decode +scripts/rocm_decode_profile.sh models/mlx/Meta-Llama-3.1-8B-Instruct-4bit models/mlx/Qwen3-30B-A3B-4bit +scripts/rocm_decode_profile.sh --temperature 0.7 --top-p 0.95 --no-plain models/mlx/Meta-Llama-3.1-8B-Instruct-4bit +python3 scripts/rocm_decode_profile.py report benchmarks/rocm_profiles/gfx1151_ +``` + +Results go to `benchmarks/rocm_profiles/gfx1151_/`; the full traces go +to a temporary directory (`--trace-dir` to keep them) because they run to +hundreds of MB. Read shares rather than absolute times from the profiled run: +the tracer slows decode, and the summary records by how much against the plain +run. `--temperature` and `--top-p` exist so a profile can see the sampler; the +greedy default never dispatches it. The published profile and how each kernel +name was attributed to a port unit are in +[rocm-decode-profile-gfx1151-2026-09-30.md](benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md). + ### An op-level number is not a decode number (issue #901) `examples/rejection_sampling_microbench.rs` reports two speedups per row, and diff --git a/scripts/rocm_decode_profile.py b/scripts/rocm_decode_profile.py new file mode 100644 index 000000000..3b82d3b1a --- /dev/null +++ b/scripts/rocm_decode_profile.py @@ -0,0 +1,595 @@ +#!/usr/bin/env python3 +# Copyright 2025-2026 Lablup Inc. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Cut the measured decode out of a rocprofv3 kernel trace and attribute it. + +Issue #2061. `scripts/rocm_decode_profile.sh` runs `mlxcel-bench-decode` under +`rocprofv3 --kernel-trace` with `MLXCEL_BENCH_PHASE_MARKS=1`; this script reads +the trace and the bench's log and writes, per run: + +* ``_decode_kernels.csv``: one row per kernel name dispatched during the + measured decode, with calls, total GPU time, share of decode GPU time, its + kernel class and the #1814 port unit it is attributed to; +* ``_summary.json``: decode GPU time and host gap per token, tok/s with + and without the profiler, share per port unit, and the checks below. + +The measured decode is the dispatches that start between the bench's +``decode_start`` and ``measured_end`` marks (see +``src/bin/bench_decode/phase_marks.rs``). The prefill before it ends in a +blocking ``eval`` of the first token, so no kernel should straddle the cut; +the summary records how many do (expected 0) and the idle gap before the +first decode dispatch. + +``report`` renders the summaries of a directory as the Markdown tables of the +published profile. + +Run the unit tests with ``python3 -m unittest tests/test_rocm_decode_profile.py``. +""" + +from __future__ import annotations + +import argparse +import csv +import json +import pathlib +import re +import sys +from collections import defaultdict +from dataclasses import dataclass, field + +# -------------------------------------------------------------------------- +# Kernel classes. First match wins. Patterns run on the demangled kernel name +# as rocprofv3 writes it; see the profile doc for how each was checked against +# the MLX ROCm overlay source. +# -------------------------------------------------------------------------- +KERNEL_CLASSES: list[tuple[str, re.Pattern[str]]] = [ + ("gather_qmv", re.compile(r"gather.*q(mv|mm)|q(mv|mm).*gather", re.I)), + ("gather_mm", re.compile(r"gather_mm|segmented|sorted_rhs", re.I)), + ("qmv", re.compile(r"\bqmv|qmv_|_qmv|qvm", re.I)), + ("qmm", re.compile(r"\bqmm|qmm_|_qmm", re.I)), + ("gemm", re.compile(r"Cijk_|gemm|gemv|rocblas|hipblaslt|wmma_matmul", re.I)), + ("rms_norm", re.compile(r"rms_?norm", re.I)), + ("layer_norm", re.compile(r"layer_?norm", re.I)), + ("rope", re.compile(r"rope", re.I)), + ("sdpa", re.compile(r"sdpa|attention|flash", re.I)), + ("softmax", re.compile(r"softmax", re.I)), + ("logsumexp", re.compile(r"logsumexp", re.I)), + ("arg_reduce", re.compile(r"arg_?reduce|argmax|argmin", re.I)), + ("scan", re.compile(r"scan|cumsum", re.I)), + ("sort", re.compile(r"sort|partition", re.I)), + ("random", re.compile(r"rbits|random|philox|threefry", re.I)), + ("conv", re.compile(r"conv", re.I)), + ("reduce", re.compile(r"reduce|all_reduce|row_reduce|col_reduce", re.I)), + ("gather_scatter", re.compile(r"gather|scatter|take|slice_update|masked", re.I)), + ("copy", re.compile(r"copy|fill|arange", re.I)), + ("binary", re.compile(r"binary", re.I)), + ("ternary", re.compile(r"ternary|select|where", re.I)), + ("unary", re.compile(r"unary", re.I)), + ("dequantize", re.compile(r"dequantize", re.I)), + # MLX's JIT-compiled elementwise kernels (`mlx::core::compile`) are named + # after their encoded tape plus `_contiguous` / `_strided`, e.g. + # `CV2ISigmoid...Multiply..._contiguous` for mlxcel's compiled SwiGLU and + # `BV2ISigmoid...` for SiLU; `compiled.cpp` in the overlay builds the name. + ("compiled", re.compile(r"rocm::[A-Z][A-Za-z0-9]+_[A-Za-z0-9_]*_(contiguous|strided)<")), +] + +# Port units split from #1814, in the order #1814 listed them. +PORT_UNITS = { + "2063": "fused_add_rms_norm + fused_rope_qk_append", + "2064": "samplers (gumbel_max_sample, rejection_sample)", + "2065": "fused MoE decode (moe_gateup, moe_down)", + "2067": "SSM update (ssm_update_kernel)", + "2068": "paged attention (v1, v2, merge)", +} + + +def classify(name: str) -> str: + for cls, pat in KERNEL_CLASSES: + if pat.search(name): + return cls + return "other" + + +# -------------------------------------------------------------------------- +# Inputs +# -------------------------------------------------------------------------- +@dataclass +class Dispatch: + name: str + start: int + end: int + grid: int = 0 + + @property + def dur(self) -> int: + return self.end - self.start + + +PHASE_RE = re.compile(r"^\[phase\] (\w+) monotonic_ns=(\S+) boottime_ns=(\S+)") +DECODE_RE = re.compile(r"^\s*Decode:\s+([\d.]+) ms \(([\d.]+) tok/s\)") +PREFILL_RE = re.compile(r"^\s*Prefill:\s+([\d.]+) ms \(([\d.]+) tok/s\)") +GEN_RE = re.compile(r"^\s*Generated tokens:\s+(\d+)") +PROMPT_RE = re.compile(r"^\s*Prompt tokens:\s+(\d+)") + + +@dataclass +class BenchLog: + marks: dict[str, dict[str, int]] = field(default_factory=dict) + decode_ms: float | None = None + decode_tok_s: float | None = None + prefill_ms: float | None = None + prefill_tok_s: float | None = None + generated: int | None = None + prompt: int | None = None + + +def parse_log(text: str) -> BenchLog: + log = BenchLog() + for line in text.splitlines(): + if m := PHASE_RE.match(line): + clocks = {} + for clock, value in (("monotonic", m.group(2)), ("boottime", m.group(3))): + if value != "na": + clocks[clock] = int(value) + log.marks[m.group(1)] = clocks + elif m := DECODE_RE.match(line): + log.decode_ms, log.decode_tok_s = float(m.group(1)), float(m.group(2)) + elif m := PREFILL_RE.match(line): + log.prefill_ms, log.prefill_tok_s = float(m.group(1)), float(m.group(2)) + elif m := GEN_RE.match(line): + log.generated = int(m.group(1)) + elif m := PROMPT_RE.match(line): + log.prompt = int(m.group(1)) + return log + + +def _col(header: list[str], *names: str) -> int: + for n in names: + if n in header: + return header.index(n) + raise ValueError(f"kernel trace has none of the columns {names}; header: {header}") + + +def read_trace(path: pathlib.Path) -> list[Dispatch]: + with path.open(newline="") as fh: + rows = csv.reader(fh) + header = next(rows) + i_name = _col(header, "Kernel_Name", "KernelName") + i_start = _col(header, "Start_Timestamp", "BeginNs") + i_end = _col(header, "End_Timestamp", "EndNs") + grid_cols = [header.index(c) for c in ("Grid_Size_X", "Grid_Size_Y", "Grid_Size_Z") + if c in header] + out = [] + for r in rows: + grid = 1 + for c in grid_cols: + grid *= max(int(r[c] or 1), 1) + out.append(Dispatch(r[i_name], int(r[i_start]), int(r[i_end]), + grid if grid_cols else 0)) + out.sort(key=lambda d: d.start) + return out + + +# -------------------------------------------------------------------------- +# Decode window +# -------------------------------------------------------------------------- +def pick_clock(log: BenchLog, dispatches: list[Dispatch]) -> str: + """The host clock whose marks actually bracket the traced dispatches. + + rocprofv3 stamps dispatches on one host clock; the bench prints two. The + right one is the one under which the bench's own passes contain the + dispatches, so count dispatches inside [warmup_start, measured_end] for + each and keep the larger (they tie when the host has never suspended). + """ + best, best_n = None, -1 + for clock in ("boottime", "monotonic"): + try: + lo = log.marks["warmup_start"][clock] + hi = log.marks["measured_end"][clock] + except KeyError: + continue + n = sum(1 for d in dispatches if lo <= d.start <= hi) + if n > best_n: + best, best_n = clock, n + if best is None or best_n <= 0: + raise ValueError("no dispatch lies between the warmup_start and measured_end marks " + "on either clock; was MLXCEL_BENCH_PHASE_MARKS=1 set?") + return best + + +def busy_ns(ds: list[Dispatch], lo: int, hi: int) -> int: + """Union of dispatch intervals clipped to [lo, hi].""" + total, cur_s, cur_e = 0, None, None + for d in ds: + s, e = max(d.start, lo), min(d.end, hi) + if e <= s: + continue + if cur_e is None or s > cur_e: + if cur_e is not None: + total += cur_e - cur_s + cur_s, cur_e = s, e + else: + cur_e = max(cur_e, e) + if cur_e is not None: + total += cur_e - cur_s + return total + + +# -------------------------------------------------------------------------- +# Roles: what mlxcel op each decode dispatch belongs to +# -------------------------------------------------------------------------- +# Generic kernels (binary, copy, reduce) are emitted by many ops, so a name alone +# cannot say which op dispatched them. MLX evaluates a decode step's graph in a +# fixed order, so each op leaves a recognisable run of dispatches; the rules +# below read those runs. Each was checked against the model code and a printed +# decode step (see the profile doc, "Attribution"). +ROLES = ( + "sampler_tail", # after the lm_head GEMV: logit bias, sampling + "add_rms_join_post_attn", # residual add + RMSNorm after attention's o_proj + "add_rms_join", # the other residual add + RMSNorm pairs + "rope_append", # q/k RoPE and the K/V cache writes between them + "moe_expert_gemv", # gather_qmm expert GEMVs + "moe_activation", # SwitchGLU activation between the expert GEMVs + "moe_weighted_sum", # score-weighted sum after the down GEMV + "moe_gather_indices", # the arange each gather_qmm builds + "ssm_step", # the Mamba2 SSD step graph (ssm_step) + "ssm_conv", # depthwise conv1d and its state copies + "ssm_silu", # SiLU on the conv output and on the gate + "ssm_gated_norm", # gated RMSNorm closing the mixer + "paged_attention", # paged-attention graph fallback +) + + +def assign_roles(decode: list[Dispatch]) -> list[str | None]: + n = len(decode) + cls = [classify(d.name) for d in decode] + role: list[str | None] = [None] * n + + def is_add(i: int) -> bool: + return cls[i] == "binary" and "::Add," in decode[i].name + + # Sampler tail: the lm_head is the widest qmv of the step (vocab rows). + qmv_grids = [d.grid for d, c in zip(decode, cls) if c == "qmv"] + if qmv_grids: + head_grid = max(qmv_grids) + stop = {"qmv", "rms_norm", "dequantize", "gather_qmv", "sdpa", "gemm"} + for i in range(n): + if cls[i] == "qmv" and decode[i].grid == head_grid: + j = i + 1 + while j < n and cls[j] not in stop and "gather_rows" not in decode[j].name: + role[j] = "sampler_tail" + j += 1 + + # Mamba2 mixer: the dispatches between the qmv before a conv1d (in_proj) + # and the qmv after it (out_proj). + for i in range(n): + if cls[i] != "conv" or role[i] is not None: + continue + lo = i + while lo > 0 and cls[lo - 1] not in ("qmv", "gather_qmv"): + lo -= 1 + hi = i + while hi + 1 < n and cls[hi + 1] not in ("qmv", "gather_qmv"): + hi += 1 + span = range(lo, hi + 1) + norm_at = max((k for k in span if cls[k] == "rms_norm"), default=None) + for k in span: + if norm_at is not None and k >= norm_at: + role[k] = "ssm_gated_norm" + elif cls[k] == "conv" or "copy_gg" in decode[k].name: + role[k] = "ssm_conv" + elif cls[k] == "compiled" and "Sigmoid" in decode[k].name: + role[k] = "ssm_silu" + else: + role[k] = "ssm_step" + + # MoE: clusters of gather_qmv dispatches. + g = [i for i in range(n) if cls[i] == "gather_qmv"] + clusters: list[list[int]] = [] + for i in g: + if clusters and i - clusters[-1][-1] <= 12: + clusters[-1].append(i) + else: + clusters.append([i]) + for c in clusters: + for i in c: + role[i] = "moe_expert_gemv" + for k in range(c[0] + 1, c[-1]): + if role[k] is None and cls[k] in ("compiled", "binary", "copy", "unary") \ + and "::Divide," not in decode[k].name: + role[k] = "moe_activation" + k = c[-1] + 1 + while k < n and role[k] is None and not is_add(k) \ + and cls[k] in ("binary", "copy", "reduce", "unary"): + role[k] = "moe_weighted_sum" + k += 1 + if g: + for i in range(n): + if role[i] is None and "arange_kernel" in decode[i].name: + role[i] = "moe_gather_indices" + + # RoPE and the cache writes it brackets. + for i in range(n): + if role[i] is None and cls[i] == "rope": + role[i] = "rope_append" + k = i + 1 + while k < n and role[k] is None and "copy_gg" in decode[k].name: + role[k] = "rope_append" + k += 1 + + # Residual add + RMSNorm pairs; post-attention when the add follows o_proj + # (a qmv) that follows the attention kernel. + for i in range(n - 1): + if role[i] is None and role[i + 1] is None and is_add(i) and cls[i + 1] == "rms_norm": + post = i >= 2 and cls[i - 1] == "qmv" and cls[i - 2] == "sdpa" + r = "add_rms_join_post_attn" if post else "add_rms_join" + role[i] = role[i + 1] = r + + for i in range(n): + if role[i] is None and "paged" in decode[i].name.lower(): + role[i] = "paged_attention" + return role + + +# -------------------------------------------------------------------------- +# Port units: which roles each would replace, and which the shipped code reaches +# -------------------------------------------------------------------------- +UNIT_ROLES = { + "2063": ("add_rms_join_post_attn", "rope_append"), + "2064": ("sampler_tail",), + "2065": ("moe_expert_gemv", "moe_activation", "moe_weighted_sum", "moe_gather_indices"), + "2067": ("ssm_step",), + "2068": ("paged_attention",), +} + + +def reach(unit: str, config: dict, sampled: bool) -> tuple[tuple[str, ...], tuple[str, ...], str]: + """(roles reached with shipped defaults, roles reached with opt-ins, note). + + Read from mlxcel at the profiled commit; see the doc for the line numbers. + """ + mt = config.get("model_type", "") + rope_scaling = config.get("rope_scaling") or {} + rope_table = (rope_scaling.get("rope_type") or rope_scaling.get("type")) in ("llama3", "yarn") + if unit == "2063": + if mt != "llama": + return (), (), f"{mt} never calls fused_add_rms_norm or forward_fused_rope_append" + optin = ("add_rms_join_post_attn",) + (() if rope_table else ("rope_append",)) + note = ("both paths ship off (FUSED_ADD_RMSNORM_DEFAULT / FUSED_ROPE_APPEND_DEFAULT " + "false); only the post-attention join calls the fused norm") + if rope_table: + note += "; rope_scaling builds a frequency table, which the RoPE kernel cannot take" + return (), optin, note + if unit == "2064": + if sampled: + return ("sampler_tail",), ("sampler_tail",), "sampled run: the draw is the fallback chain" + return (), (), "greedy argmax dispatches neither sampler kernel" + if unit == "2065": + if mt in ("qwen3_moe", "qwen3_next", "mixtral", "dbrx", "cohere2_moe", "klear", + "afmoe", "lfm2", "bailing_moe"): + return UNIT_ROLES["2065"], UNIT_ROLES["2065"], "forward_fused_kernel caller" + if mt == "nemotron_h": + return (), (), ("fused_moe_forward's default branch is gather_qmm; the kernel path " + "needs MLXCEL_FUSED_MOE_RELU2 and moe_fc1_relu2 (#2069)") + return (), (), f"{mt} does not call forward_fused_kernel" + if unit == "2067": + if mt in ("granitemoehybrid", "nemotron_h", "falcon_h1", "plamo2"): + return ("ssm_step",), ("ssm_step",), "ssm_kernel_available() gates the decode step" + return (), (), "no Mamba2 layer" + if unit == "2068": + return (), (), "the bench decodes into a dense KVCache; the paged path is not taken" + raise ValueError(unit) + + +# -------------------------------------------------------------------------- +# summarize +# -------------------------------------------------------------------------- +def summarize(trace: pathlib.Path, log_path: pathlib.Path, plain_log: pathlib.Path | None, + model_dir: pathlib.Path | None, name: str, out_dir: pathlib.Path, + temperature: float = 0.0) -> dict: + log = parse_log(log_path.read_text(errors="replace")) + dispatches = read_trace(trace) + clock = pick_clock(log, dispatches) + lo = log.marks["decode_start"][clock] + hi = log.marks["measured_end"][clock] + decode = [d for d in dispatches if lo <= d.start <= hi] + if not decode: + raise ValueError("no dispatch inside the decode window") + before = [d for d in dispatches if d.end <= lo] + straddling = sum(1 for d in dispatches if d.start < lo < d.end) + tokens = log.generated or 0 + gpu_sum = sum(d.dur for d in decode) + gpu_busy = busy_ns(decode, lo, hi) + wall = hi - lo + + config = {} + if model_dir is not None and (model_dir / "config.json").exists(): + config = json.loads((model_dir / "config.json").read_text()) + + roles = assign_roles(decode) + per_kernel: dict[str, list] = {} + for d, r in zip(decode, roles): + row = per_kernel.setdefault(d.name, [0, 0, defaultdict(int)]) + row[0] += 1 + row[1] += d.dur + row[2][r or ""] += d.dur + role_ns: dict[str, int] = defaultdict(int) + role_calls: dict[str, int] = defaultdict(int) + class_ns: dict[str, int] = defaultdict(int) + for d, r in zip(decode, roles): + role_ns[r or "unattributed"] += d.dur + role_calls[r or "unattributed"] += 1 + class_ns[classify(d.name)] += d.dur + + def share(ns: int) -> float: + return round(100.0 * ns / gpu_sum, 2) + + unit_of_role = {r: u for u, rs in UNIT_ROLES.items() for r in rs} + out_dir.mkdir(parents=True, exist_ok=True) + with (out_dir / f"{name}_decode_kernels.csv").open("w", newline="") as fh: + w = csv.writer(fh) + w.writerow(["kernel", "class", "roles", "port_units", "calls", "calls_per_token", + "total_ns", "avg_ns", "share_pct"]) + for kname, (calls, ns, by_role) in sorted(per_kernel.items(), key=lambda kv: -kv[1][1]): + rs = sorted(r for r in by_role if r) + units = sorted({unit_of_role[r] for r in rs if r in unit_of_role}) + w.writerow([kname, classify(kname), ";".join(rs), ";".join(units), calls, + round(calls / tokens, 2) if tokens else "", ns, round(ns / calls), + f"{100.0 * ns / gpu_sum:.3f}"]) + + units = {} + for unit, rs in UNIT_ROLES.items(): + default_roles, optin_roles, note = reach(unit, config, temperature > 0) + units[unit] = { + "port": PORT_UNITS[unit], + "fallback_share_pct": share(sum(role_ns[r] for r in rs)), + "fallback_dispatches_per_token": round(sum(role_calls[r] for r in rs) / tokens, 1) + if tokens else None, + "reached_default_share_pct": share(sum(role_ns[r] for r in default_roles)), + "reached_default_dispatches_per_token": + round(sum(role_calls[r] for r in default_roles) / tokens, 1) if tokens else None, + "reached_optin_share_pct": share(sum(role_ns[r] for r in optin_roles)), + "note": note, + } + + plain = parse_log(plain_log.read_text(errors="replace")) if plain_log else None + top = sorted(per_kernel.items(), key=lambda kv: -kv[1][1])[:10] + summary = { + "name": name, + "model_type": config.get("model_type"), + "temperature": temperature, + "clock": clock, + "prompt_tokens": log.prompt, + "generated_tokens": tokens, + "profiled_decode_tok_s": log.decode_tok_s, + "profiled_prefill_tok_s": log.prefill_tok_s, + "plain_decode_tok_s": plain.decode_tok_s if plain else None, + "plain_prefill_tok_s": plain.prefill_tok_s if plain else None, + "profiler_decode_slowdown_pct": round(100.0 * (plain.decode_tok_s / log.decode_tok_s - 1), 2) + if plain and plain.decode_tok_s and log.decode_tok_s else None, + "decode_dispatches": len(decode), + "dispatches_per_token": round(len(decode) / tokens, 1) if tokens else None, + "decode_wall_ms": round(wall / 1e6, 3), + "decode_gpu_sum_ms": round(gpu_sum / 1e6, 3), + "decode_gpu_busy_ms": round(gpu_busy / 1e6, 3), + "gpu_ms_per_token": round(gpu_busy / 1e6 / tokens, 4) if tokens else None, + "host_gap_ms_per_token": round((wall - gpu_busy) / 1e6 / tokens, 4) if tokens else None, + "host_gap_pct_of_wall": round(100.0 * (wall - gpu_busy) / wall, 2), + # The tracer adds host time per dispatch but barely changes kernel + # durations, so the unprofiled run's wall time per token minus the + # traced GPU time is the better host-gap estimate for dispatch-heavy + # models. Same denominator as tok/s (generated tokens). + "plain_wall_ms_per_token": round(1000.0 / plain.decode_tok_s, 4) + if plain and plain.decode_tok_s else None, + "plain_host_gap_ms_per_token_est": + round(1000.0 / plain.decode_tok_s - gpu_busy / 1e6 / tokens, 4) + if plain and plain.decode_tok_s and tokens else None, + "checks": { + "kernels_straddling_decode_start": straddling, + "idle_gap_before_first_decode_dispatch_us": + round((decode[0].start - before[-1].end) / 1e3, 1) if before else None, + "last_dispatch_before_decode": before[-1].name[:120] if before else None, + }, + "top_kernels": [ + {"kernel": k, "class": classify(k), "calls_per_token": round(c / tokens, 2) if tokens else None, + "share_pct": share(ns)} + for k, (c, ns, _) in top + ], + "class_share_pct": {c: share(ns) for c, ns in sorted(class_ns.items(), key=lambda kv: -kv[1])}, + "role_share_pct": {r: share(ns) for r, ns in sorted(role_ns.items(), key=lambda kv: -kv[1])}, + "role_dispatches_per_token": {r: round(c / tokens, 1) for r, c in role_calls.items()} + if tokens else {}, + "port_units": units, + } + (out_dir / f"{name}_summary.json").write_text(json.dumps(summary, indent=2) + "\n") + return summary + + +# -------------------------------------------------------------------------- +# report +# -------------------------------------------------------------------------- +def short_kernel(name: str) -> str: + """`binary_vv` style: kernel, op functor and first dtype.""" + if name.startswith("Cijk_"): + return "Cijk_* (Tensile GEMM)" + m = re.search(r"::([A-Za-z0-9_]+)(<[^(]*>)?\(", name) + base = m.group(1) if m else name.split("(")[0] + if base[:1].isupper() and ("_contiguous" in base or "_strided" in base): + return "compiled " + base[:40] + "..." + parts = [] + if m and m.group(2): + args = m.group(2)[1:-1] + if op := re.search(r"rocm::(\w+)(<[^>]*>)?,", args): + parts.append(op.group(1)) + args = args[op.end():] + if dt := re.search(r"(hip_bfloat16|__half|float|int|unsigned int|bool)", args): + parts.append({"hip_bfloat16": "bf16", "__half": "f16", "float": "f32", + "unsigned int": "u32"}.get(dt.group(1), dt.group(1))) + return f"{base}<{', '.join(parts)}>" if parts else base + + +def report(out_dir: pathlib.Path) -> str: + rows = [json.loads(p.read_text()) for p in sorted(out_dir.glob("*_summary.json"))] + lines = ["| Run | plain tok/s | profiled tok/s | profiler cost | dispatches/token | " + "GPU ms/token | host gap ms/token | host gap % |", + "|---|---:|---:|---:|---:|---:|---:|---:|"] + for s in rows: + lines.append( + f"| `{s['name']}` | {s['plain_decode_tok_s'] or '-'} | {s['profiled_decode_tok_s']} | " + f"{'-' if s['profiler_decode_slowdown_pct'] is None else str(s['profiler_decode_slowdown_pct']) + '%'} | " + f"{s['dispatches_per_token']} | {s['gpu_ms_per_token']} | {s['host_gap_ms_per_token']} | " + f"{s['host_gap_pct_of_wall']} |") + lines += ["", "| Run | " + " | ".join(f"#{u}" for u in UNIT_ROLES) + " |", + "|---|" + "---:|" * len(UNIT_ROLES)] + for s in rows: + cells = [] + for u in UNIT_ROLES: + p = s["port_units"][u] + cells.append(f"{p['reached_default_share_pct']} ({p['fallback_share_pct']})") + lines.append(f"| `{s['name']}` | " + " | ".join(cells) + " |") + for s in rows: + lines += ["", f"Top kernels, `{s['name']}` (share of decode GPU time, calls per token):", ""] + for k in s["top_kernels"]: + lines.append(f"- `{short_kernel(k['kernel'])}` {k['share_pct']}%, {k['calls_per_token']}") + return "\n".join(lines) + "\n" + + +def main(argv: list[str] | None = None) -> int: + ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) + sub = ap.add_subparsers(dest="cmd", required=True) + s = sub.add_parser("summarize", help="cut and attribute one profiled run") + s.add_argument("--trace", type=pathlib.Path, required=True) + s.add_argument("--log", type=pathlib.Path, required=True) + s.add_argument("--plain-log", type=pathlib.Path) + s.add_argument("--model-dir", type=pathlib.Path) + s.add_argument("--name", required=True) + s.add_argument("--out-dir", type=pathlib.Path, required=True) + s.add_argument("--temperature", type=float, default=0.0, + help="the run's sampling temperature (0 = greedy)") + r = sub.add_parser("report", help="render the summaries in a directory as Markdown") + r.add_argument("out_dir", type=pathlib.Path) + args = ap.parse_args(argv) + if args.cmd == "summarize": + s = summarize(args.trace, args.log, args.plain_log, args.model_dir, args.name, + args.out_dir, args.temperature) + print(f"{s['name']}: {s['decode_dispatches']} decode dispatches, " + f"{s['gpu_ms_per_token']} GPU ms/token, {s['host_gap_ms_per_token']} host gap ms/token") + else: + sys.stdout.write(report(args.out_dir)) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/rocm_decode_profile.sh b/scripts/rocm_decode_profile.sh new file mode 100755 index 000000000..1b4595854 --- /dev/null +++ b/scripts/rocm_decode_profile.sh @@ -0,0 +1,133 @@ +#!/usr/bin/env bash +# Per-kernel ROCm decode profile of mlxcel-bench-decode (issue #2061). +# +# Usage: +# scripts/rocm_decode_profile.sh [options] MODEL_DIR [MODEL_DIR...] +# +# For each model this runs, each under scripts/rocm_gpu_guard.sh (90 s of idle +# GPU and no compiler before, a 1 Hz monitor during, rerun on contention): +# +# 1. a plain `mlxcel-bench-decode` run, for tok/s without the profiler; +# 2. the same run under `rocprofv3 --kernel-trace --hip-graph-trace --stats +# -f csv` with MLXCEL_BENCH_PHASE_MARKS=1; +# +# then scripts/rocm_decode_profile.py cuts the measured decode out of the +# kernel trace by the bench's phase marks and writes the per-kernel decode +# table and a summary. The shape is scripts/bench_decode.sh's default, pp512 / +# tg128 with a 20-token same-process warmup and --ignore-eos, greedy unless +# --temperature is given. +# +# Options: +# --out DIR output directory (default benchmarks/rocm_profiles/gfx1151_) +# --trace-dir DIR where full kernel traces go (default: a temp dir; they are +# tens to hundreds of MB and are not meant to be committed) +# --temperature T sampling temperature (default 0, greedy) +# --top-p P top-p (default 1.0, off) +# --tag NAME run name suffix (default greedy, or t[-p

]) +# --no-plain skip the unprofiled run (no overhead figure) +# --idle-secs N guard idle window (default 90) +# --bench PATH bench binary (default target/release/mlxcel-bench-decode) +# +# Build first: cargo build --release --features rocm --bin mlxcel-bench-decode +# Render the tables: python3 scripts/rocm_decode_profile.py report + +set -euo pipefail + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +BENCH="$ROOT/target/release/mlxcel-bench-decode" +ROCPROF="${ROCPROFV3:-$(command -v rocprofv3 || echo /opt/rocm/bin/rocprofv3)}" +OUT="" +TRACE_DIR="" +TEMPERATURE=0 +TOP_P=1.0 +TAG="" +PLAIN=1 +IDLE_SECS=90 +PROMPT_TOKENS=512 +MAX_TOKENS=128 +WARMUP_TOKENS=20 + +usage() { sed -n '2,33p' "$0" | sed 's/^# \{0,1\}//'; } + +MODELS=() +while [[ $# -gt 0 ]]; do + case "$1" in + --out) OUT="$2"; shift 2 ;; + --trace-dir) TRACE_DIR="$2"; shift 2 ;; + --temperature) TEMPERATURE="$2"; shift 2 ;; + --top-p) TOP_P="$2"; shift 2 ;; + --tag) TAG="$2"; shift 2 ;; + --no-plain) PLAIN=0; shift ;; + --idle-secs) IDLE_SECS="$2"; shift 2 ;; + --bench) BENCH="$2"; shift 2 ;; + -h|--help) usage; exit 0 ;; + -*) echo "unknown option $1" >&2; usage >&2; exit 2 ;; + *) MODELS+=("$1"); shift ;; + esac +done +[[ ${#MODELS[@]} -gt 0 ]] || { usage >&2; exit 2; } +[[ -x "$BENCH" ]] || { echo "bench binary not found: $BENCH (build it first)" >&2; exit 2; } +[[ -x "$ROCPROF" ]] || { echo "rocprofv3 not found (set ROCPROFV3)" >&2; exit 2; } + +COMMIT=$(git -C "$ROOT" rev-parse --short=8 HEAD) +if [[ -n "$(git -C "$ROOT" status --porcelain --untracked-files=no -- src Cargo.toml Cargo.lock)" ]]; then + echo "note: the source tree differs from $COMMIT; record the diff with the results" >&2 +fi +OUT="${OUT:-$ROOT/benchmarks/rocm_profiles/gfx1151_${COMMIT}}" +mkdir -p "$OUT" +if [[ -z "$TRACE_DIR" ]]; then + TRACE_DIR=$(mktemp -d -t rocm-decode-trace.XXXXXX) +fi +mkdir -p "$TRACE_DIR" +if [[ -z "$TAG" ]]; then + if [[ "$TEMPERATURE" == 0 || "$TEMPERATURE" == 0.0 ]]; then + TAG=greedy + else + TAG="t${TEMPERATURE}" + [[ "$TOP_P" == 1 || "$TOP_P" == 1.0 ]] || TAG="${TAG}-p${TOP_P}" + fi +fi +GUARD_LOG="$OUT/guard.log" + +bench_args() { + printf '%s\n' -m "$1" -p "profile" -n "$MAX_TOKENS" --warmup-tokens "$WARMUP_TOKENS" \ + --ignore-eos --prompt-tokens "$PROMPT_TOKENS" --temperature "$TEMPERATURE" --top-p "$TOP_P" +} + +for model in "${MODELS[@]}"; do + name="$(basename "$model")_${TAG}" + mapfile -t args < <(bench_args "$model") + echo ">>> $name" >&2 + + run_dir="$TRACE_DIR/$name" + mkdir -p "$run_dir" + + if [[ "$PLAIN" == 1 ]]; then + echo "== $name plain" >> "$GUARD_LOG" + "$ROOT/scripts/rocm_gpu_guard.sh" --idle-secs "$IDLE_SECS" --log "$GUARD_LOG" -- \ + "$BENCH" "${args[@]}" > "$run_dir/plain_bench.log" 2>&1 + # The guard's own lines are in guard.log already. + grep -v '^rocm_gpu_guard: ' "$run_dir/plain_bench.log" > "$OUT/${name}_plain_bench.log" || true + fi + + echo "== $name profiled" >> "$GUARD_LOG" + MLXCEL_BENCH_PHASE_MARKS=1 "$ROOT/scripts/rocm_gpu_guard.sh" --idle-secs "$IDLE_SECS" \ + --log "$GUARD_LOG" -- \ + "$ROCPROF" --kernel-trace --hip-graph-trace --stats -f csv -d "$run_dir" -o "$name" -- \ + "$BENCH" "${args[@]}" > "$run_dir/bench.log" 2>&1 + # Drop rocprofv3's own glog lines (they name local temp paths); keep the bench's. + grep -Ev '^[WEI][0-9]{4} |^rocm_gpu_guard: ' "$run_dir/bench.log" > "$OUT/${name}_bench.log" || true + cp "$run_dir/${name}_kernel_stats.csv" "$OUT/${name}_kernel_stats.csv" + + plain_args=() + [[ "$PLAIN" == 1 ]] && plain_args=(--plain-log "$OUT/${name}_plain_bench.log") + python3 "$ROOT/scripts/rocm_decode_profile.py" summarize \ + --trace "$run_dir/${name}_kernel_trace.csv" --log "$OUT/${name}_bench.log" \ + ${plain_args[@]+"${plain_args[@]}"} --model-dir "$model" --name "$name" --out-dir "$OUT" \ + --temperature "$TEMPERATURE" \ + || echo "summarize failed for $name; rerun it on $run_dir/${name}_kernel_trace.csv" >&2 +done +# guard.log is published evidence; keep local paths out of it. +sed -i "s#${ROOT}/##g; s#${TRACE_DIR}##g" "$GUARD_LOG" +echo "traces: $TRACE_DIR" >&2 +echo "results: $OUT" >&2 diff --git a/scripts/rocm_gpu_guard.sh b/scripts/rocm_gpu_guard.sh new file mode 100755 index 000000000..bf8b9dc27 --- /dev/null +++ b/scripts/rocm_gpu_guard.sh @@ -0,0 +1,175 @@ +#!/usr/bin/env bash +# Run a command only while the ROCm GPU is otherwise idle, and prove it was. +# +# Usage: +# scripts/rocm_gpu_guard.sh [--idle-secs N] [--max-attempts N] [--max-wait SECS] +# [--log FILE] -- COMMAND [ARGS...] +# +# A benchmark that shares the GPU with another process, or the UMA memory bus +# with a compiler, is not a measurement. This is the guard the gfx1151 baseline +# (docs/benchmark_results/rocm-baseline-gfx1151-2026-09-30.md, #2056) describes, +# as a script so every ROCm measurement can use the same one (#2061): +# +# 1. wait until the host has gone --idle-secs (default 90) consecutive seconds +# with /sys/class/kfd/kfd/proc empty (the process list `rocm-smi --showpids` +# reads) and no compiler process running; +# 2. run COMMAND while a monitor samples both once per second; +# 3. reject the attempt if any sample shows a GPU process that is not COMMAND +# or one of its descendants, or a compiler, and go back to step 1. +# +# Every sample is appended to --log (default: stderr only), so the published +# evidence is the log itself. Exit status: COMMAND's status from the first clean +# attempt; 75 when every attempt was contended or --max-wait ran out, in which +# case nothing COMMAND printed should be used. COMMAND's stdout and stderr pass +# through unchanged, so callers capture them as usual. +# +# The sampling interval is one second: a GPU job shorter than that can in +# principle be missed. The compiler list matches /proc//comm exactly. + +set -euo pipefail + +IDLE_SECS=90 +MAX_ATTEMPTS=5 +MAX_WAIT=0 +LOG="" +KFD_PROC_DIR="${ROCM_GPU_GUARD_KFD_DIR:-/sys/class/kfd/kfd/proc}" +# Build tools whose memory traffic or CPU load would distort a UMA measurement. +# Driver processes (make, cmake, ninja, build scripts) are left out: they are +# idle while their compiler children, which are listed, do the work. +# ROCM_GPU_GUARD_COMPILER_RE overrides the list (the unit tests use it). +COMPILER_RE="${ROCM_GPU_GUARD_COMPILER_RE:-^(cargo|rustc|clang|clang\+\+|clang-[0-9]+|hipcc|nvcc|cc1|cc1plus|ld|ld\.lld|ld\.gold|ld\.bfd|lld|collect2)$}" + +usage() { + sed -n '2,27p' "$0" | sed 's/^# \{0,1\}//' +} + +while [[ $# -gt 0 ]]; do + case "$1" in + --idle-secs) IDLE_SECS="$2"; shift 2 ;; + --max-attempts) MAX_ATTEMPTS="$2"; shift 2 ;; + --max-wait) MAX_WAIT="$2"; shift 2 ;; + --log) LOG="$2"; shift 2 ;; + -h|--help) usage; exit 0 ;; + --) shift; break ;; + *) echo "rocm_gpu_guard: unknown option $1" >&2; usage >&2; exit 2 ;; + esac +done +[[ $# -gt 0 ]] || { echo "rocm_gpu_guard: no command given" >&2; usage >&2; exit 2; } +[[ -d "$KFD_PROC_DIR" ]] || { echo "rocm_gpu_guard: $KFD_PROC_DIR not found (no ROCm KFD driver?)" >&2; exit 2; } + +note() { + local line + line="$(date '+%Y-%m-%dT%H:%M:%S%z') $*" + [[ -n "$LOG" ]] && printf '%s\n' "$line" >> "$LOG" + printf 'rocm_gpu_guard: %s\n' "$line" >&2 +} + +# pid:comm of every process holding the GPU, comma separated. +kfd_holders() { + local d pid comm out=() + for d in "$KFD_PROC_DIR"/*; do + [[ -e "$d" ]] || continue + pid="${d##*/}" + comm=$(cat "/proc/$pid/comm" 2>/dev/null || echo '?') + out+=("$pid:$comm") + done + local IFS=, + echo "${out[*]}" +} + +# pid:comm of every running compiler, comma separated. +compilers() { + local f pid comm out=() + for f in /proc/[0-9]*/comm; do + comm=$(cat "$f" 2>/dev/null) || continue + if [[ "$comm" =~ $COMPILER_RE ]]; then + pid="${f#/proc/}"; pid="${pid%/comm}" + out+=("$pid:$comm") + fi + done + local IFS=, + echo "${out[*]}" +} + +# True when $1 is $2 or a descendant of it. +descends_from() { + local pid="$1" root="$2" ppid + while [[ -n "$pid" && "$pid" != 0 && "$pid" != 1 ]]; do + [[ "$pid" == "$root" ]] && return 0 + ppid=$(awk '{print $4}' "/proc/$pid/stat" 2>/dev/null) || return 1 + pid="$ppid" + done + return 1 +} + +# Holders in $2 (a kfd_holders reading) that are not $1 or its descendants. +foreign_holders() { + local root="$1" entry pid out=() + local IFS=, + for entry in $2; do + pid="${entry%%:*}" + descends_from "$pid" "$root" || out+=("$entry") + done + echo "${out[*]}" +} + +wait_idle() { + local quiet=0 waited=0 holders comps + note "waiting for ${IDLE_SECS}s of idle GPU and no compiler" + while (( quiet < IDLE_SECS )); do + holders=$(kfd_holders); comps=$(compilers) + if [[ -z "$holders" && -z "$comps" ]]; then + quiet=$((quiet + 1)) + else + (( quiet > 0 )) && note "idle streak reset after ${quiet}s: gpu=[${holders}] compilers=[${comps}]" + quiet=0 + fi + if (( MAX_WAIT > 0 && waited >= MAX_WAIT )); then + note "gave up after ${waited}s without ${IDLE_SECS}s of idle" + return 1 + fi + sleep 1 + waited=$((waited + 1)) + done + note "idle for ${quiet}s: kfd proc empty and no compiler in every 1 Hz sample" +} + +attempt=0 +while (( attempt < MAX_ATTEMPTS )); do + attempt=$((attempt + 1)) + wait_idle || exit 75 + note "attempt ${attempt}/${MAX_ATTEMPTS}: start: $*" + "$@" & + cmd_pid=$! + # The monitor exits 1 if any of its samples was contended. + ( + n=0 contended=0 + while kill -0 "$cmd_pid" 2>/dev/null; do + holders=$(kfd_holders); comps=$(compilers) + foreign=$(foreign_holders "$cmd_pid" "$holders") + n=$((n + 1)) + if [[ -n "$foreign" || -n "$comps" ]]; then + note "sample ${n}: CONTENDED foreign_gpu=[${foreign}] compilers=[${comps}]" + contended=1 + elif [[ -n "$LOG" ]]; then + printf '%s sample %d: clean gpu=[%s]\n' "$(date '+%Y-%m-%dT%H:%M:%S%z')" "$n" "$holders" >> "$LOG" + fi + sleep 1 + done + note "monitor: ${n} samples" + exit "$contended" + ) & + mon_pid=$! + rc=0 + wait "$cmd_pid" || rc=$? + mon_rc=0 + wait "$mon_pid" || mon_rc=$? + if (( mon_rc != 0 )); then + note "attempt ${attempt}: REJECTED (contended), exit ${rc}; rerunning" + continue + fi + note "attempt ${attempt}: CLEAN, exit ${rc}" + exit "$rc" +done +note "every attempt (${MAX_ATTEMPTS}) was contended" +exit 75 diff --git a/src/bin/bench_decode.rs b/src/bin/bench_decode.rs index b9a2848c0..e5c837c17 100644 --- a/src/bin/bench_decode.rs +++ b/src/bin/bench_decode.rs @@ -40,6 +40,9 @@ use mlxcel::vision::merge::InputEmbeddings; use mlxcel::{CxxGenerator, LanguageModel, LoadedModel, SamplingConfig}; use mlxcel_core::cache::KVCacheMode; +#[path = "bench_decode/phase_marks.rs"] +mod phase_marks; + /// Same-process benchmark for `scripts/bench_decode.sh`. #[derive(Parser, Debug)] #[command(name = "mlxcel-bench-decode")] @@ -70,6 +73,18 @@ struct Args { #[arg(long)] ignore_eos: bool, + /// Sampling temperature for both passes. The default 0 is greedy argmax, + /// which is what the published sweeps measure; a positive value routes + /// the draw through the categorical sampler (and its fused kernel where + /// the backend has one), so a profile can see what sampling costs (#2061). + #[arg(long, default_value_t = 0.0)] + temperature: f32, + + /// Nucleus (top-p) threshold for both passes. 1.0 disables it. Only has an + /// effect with a positive `--temperature`. + #[arg(long, default_value_t = 1.0)] + top_p: f32, + /// Generated tokens in the warmup pass. #[arg(long, default_value_t = 20)] warmup_tokens: usize, @@ -302,11 +317,17 @@ fn prepare_prompt( /// both the model's built-in ids and the ones read from the checkpoint config /// are covered; biasing only one of the two leaves the other able to end the /// run early. -fn sampling_config(model_path: &Path, model: &LoadedModel, ignore_eos: bool) -> SamplingConfig { +fn sampling_config( + model_path: &Path, + model: &LoadedModel, + ignore_eos: bool, + temperature: f32, + top_p: f32, +) -> SamplingConfig { let mut config = build_sampling_config(ResolvedSamplingParams { - temperature: 0.0, + temperature, top_k: 0, - top_p: 1.0, + top_p, min_p: 0.0, seed: None, repetition_penalty: 1.0, @@ -387,7 +408,7 @@ fn measured( max_tokens: usize, sampling: &SamplingConfig, kv_cache_mode: KVCacheMode, -) -> Result { +) -> Result<(mlxcel::GenerationStats, phase_marks::Stamp)> { let mut generator = CxxGenerator::new_with_kv_mode(model.num_layers(), kv_cache_mode); let (_tokens, stats) = if let Some(embeddings) = prepared.embeddings.as_ref() { let (input_embeds, mask) = mlxcel::vlm_runtime::prepared_embedding_refs(embeddings)?; @@ -402,8 +423,11 @@ fn measured( } else { generator.generate_with_stats(model, &prepared.tokens, max_tokens, sampling) }; + // Read the clocks before the trailing synchronize: the decode loop has + // already waited on its last token, so this is where `decode_time_ms` ends. + let end = phase_marks::Stamp::now(); mlxcel_core::synchronize_default(); - Ok(stats) + Ok((stats, end)) } /// Print the MLX allocator counters at a phase boundary (issue #2062). @@ -510,7 +534,13 @@ fn main() -> Result<()> { let mut run_peak = mlxcel_core::get_peak_memory(); print_memory_phase("after load")?; let tokenizer = load_tokenizer(&args.model).unwrap_or(loaded_tokenizer); - let sampling = sampling_config(&args.model, &model, args.ignore_eos); + let sampling = sampling_config( + &args.model, + &model, + args.ignore_eos, + args.temperature, + args.top_p, + ); // `--prompt-tokens N` synthesizes a deterministic long prompt for prefill // benchmarking; otherwise the short-prompt `--prompt` path runs unchanged. @@ -540,11 +570,15 @@ fn main() -> Result<()> { } }; + let phase_marks = phase_marks::enabled(); let prepared = make_prepared()?; // Per-phase peaks (issue #2062): reset the high-water mark at each phase // boundary and fold the phase peaks into `run_peak`, so the whole-run // figure printed at the end is unchanged while each phase is attributable. mlxcel_core::reset_peak_memory(); + if phase_marks { + phase_marks::Stamp::now().print("warmup_start"); + } warmup( &model, &prepared, @@ -564,7 +598,18 @@ fn main() -> Result<()> { // VLM it re-runs the vision encoder against now-warm MLX/Metal state. let prepared = make_prepared()?; mlxcel_core::reset_peak_memory(); - let stats = measured(&model, &prepared, args.max_tokens, &sampling, kv_cache_mode)?; + if phase_marks { + phase_marks::Stamp::now().print("measured_start"); + } + let (stats, measured_end) = + measured(&model, &prepared, args.max_tokens, &sampling, kv_cache_mode)?; + if phase_marks { + // The generator times its decode loop itself; its start is that long + // before the loop returned. + let decode_ns = (stats.decode_time_ms * 1e6) as u128; + measured_end.earlier_by(decode_ns).print("decode_start"); + measured_end.print("measured_end"); + } run_peak = run_peak.max(mlxcel_core::get_peak_memory()); print_memory_phase("after measured pass")?; diff --git a/src/bin/bench_decode/phase_marks.rs b/src/bin/bench_decode/phase_marks.rs new file mode 100644 index 000000000..3ca6014ed --- /dev/null +++ b/src/bin/bench_decode/phase_marks.rs @@ -0,0 +1,108 @@ +// Copyright 2025-2026 Lablup Inc. +// +// Licensed under the Apache License, Version 2.0 (the "License"); +// you may not use this file except in compliance with the License. +// You may obtain a copy of the License at +// +// http://www.apache.org/licenses/LICENSE-2.0 +// +// Unless required by applicable law or agreed to in writing, software +// distributed under the License is distributed on an "AS IS" BASIS, +// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +// See the License for the specific language governing permissions and +// limitations under the License. + +//! Host-clock phase marks for per-kernel profiling (issue #2061). +//! +//! A kernel trace (`rocprofv3 --kernel-trace`, or any tracer that stamps +//! dispatches with a host clock) holds the warmup pass and the measured prefill +//! as well as the measured decode. `MLXCEL_BENCH_PHASE_MARKS=1` makes the bench +//! print where those phases sit on the host clocks, so a post-processor can cut +//! the measured decode out of the trace by timestamp instead of guessing it +//! from kernel names. `scripts/rocm_decode_profile.py` reads these lines. +//! +//! Each mark is one stderr line: +//! +//! ```text +//! [phase] decode_start monotonic_ns=123 boottime_ns=456 +//! ``` +//! +//! Both clocks are printed because tracers differ in which one they stamp +//! with; they agree on durations and differ only by time spent suspended. +//! `decode_start` is derived as `measured_end` minus the generator's own +//! `decode_time_ms`, because the decode loop's start is inside the generator. +//! The prefill that precedes it ends in a blocking `eval` of the first token, +//! so the device is idle at that instant and the cut does not split a kernel. + +const ENV: &str = "MLXCEL_BENCH_PHASE_MARKS"; + +/// Whether phase marks were asked for. Off unless the variable is `1`, so the +/// bench's output is unchanged for `scripts/bench_decode.sh`. +pub(crate) fn enabled() -> bool { + std::env::var(ENV).is_ok_and(|v| v.trim() == "1") +} + +/// A reading of both host clocks, in nanoseconds. +#[derive(Debug, Clone, Copy)] +pub(crate) struct Stamp { + monotonic_ns: Option, + boottime_ns: Option, +} + +impl Stamp { + pub(crate) fn now() -> Self { + Self { + monotonic_ns: monotonic_ns(), + boottime_ns: boottime_ns(), + } + } + + /// This stamp moved `ns` earlier on both clocks. + pub(crate) fn earlier_by(self, ns: u128) -> Self { + Self { + monotonic_ns: self.monotonic_ns.map(|t| t.saturating_sub(ns)), + boottime_ns: self.boottime_ns.map(|t| t.saturating_sub(ns)), + } + } + + pub(crate) fn print(self, label: &str) { + let show = |v: Option| v.map_or_else(|| "na".to_string(), |n| n.to_string()); + eprintln!( + "[phase] {label} monotonic_ns={} boottime_ns={}", + show(self.monotonic_ns), + show(self.boottime_ns) + ); + } +} + +#[cfg(unix)] +fn clock_ns(id: libc::clockid_t) -> Option { + let mut ts = libc::timespec { + tv_sec: 0, + tv_nsec: 0, + }; + // SAFETY: `ts` is a live, writable `timespec` for the whole call, which is + // all `clock_gettime` requires of its out-pointer. + let rc = unsafe { libc::clock_gettime(id, &mut ts) }; + (rc == 0).then(|| ts.tv_sec as u128 * 1_000_000_000 + ts.tv_nsec as u128) +} + +#[cfg(unix)] +fn monotonic_ns() -> Option { + clock_ns(libc::CLOCK_MONOTONIC) +} + +#[cfg(not(unix))] +fn monotonic_ns() -> Option { + None +} + +#[cfg(any(target_os = "linux", target_os = "android"))] +fn boottime_ns() -> Option { + clock_ns(libc::CLOCK_BOOTTIME) +} + +#[cfg(not(any(target_os = "linux", target_os = "android")))] +fn boottime_ns() -> Option { + None +} diff --git a/tests/test_rocm_decode_profile.py b/tests/test_rocm_decode_profile.py new file mode 100644 index 000000000..1d12f6369 --- /dev/null +++ b/tests/test_rocm_decode_profile.py @@ -0,0 +1,225 @@ +# Copyright 2025-2026 Lablup Inc. +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +"""Tests for the ROCm per-kernel decode profile tooling (issue #2061). + +Covers scripts/rocm_decode_profile.py (the decode cut, the busy-time union, +the port-unit attribution) on synthetic traces, and scripts/rocm_gpu_guard.sh +against a fake KFD process directory. No GPU is needed. + +Run with: + python3 -m unittest tests/test_rocm_decode_profile.py +""" + +import os +import pathlib +import subprocess +import sys +import tempfile +import unittest + +ROOT = pathlib.Path(__file__).resolve().parents[1] +GUARD = ROOT / "scripts" / "rocm_gpu_guard.sh" +sys.path.insert(0, str(ROOT / "scripts")) + +import rocm_decode_profile as rdp # noqa: E402 + +# A compiler pattern nothing on the host matches, so the guard's compiler check +# cannot make these tests depend on what else is running. +NO_COMPILERS = "^no-such-compiler-for-tests$" + + +def run_guard(kfd_dir, *args, env_extra=None): + env = dict(os.environ, ROCM_GPU_GUARD_KFD_DIR=str(kfd_dir), + ROCM_GPU_GUARD_COMPILER_RE=NO_COMPILERS) + env.update(env_extra or {}) + return subprocess.run(["bash", str(GUARD), *args], env=env, capture_output=True, + text=True, timeout=60) + + +class GuardTests(unittest.TestCase): + def test_idle_gpu_runs_the_command_and_passes_its_status(self): + with tempfile.TemporaryDirectory() as kfd, tempfile.TemporaryDirectory() as out: + log = pathlib.Path(out) / "guard.log" + r = run_guard(kfd, "--idle-secs", "1", "--log", str(log), "--", + "bash", "-c", "sleep 1.5; exit 3") + self.assertEqual(r.returncode, 3, r.stderr) + text = log.read_text() + self.assertIn("CLEAN", text) + self.assertIn("sample 1: clean", text) + + def test_a_foreign_gpu_process_during_the_run_rejects_every_attempt(self): + with tempfile.TemporaryDirectory() as kfd: + # The command itself registers a foreign GPU holder (pid 1) once it + # is running, so the idle wait passes and the monitor must catch it. + cmd = f"mkdir -p {kfd}/1; sleep 2.5; rmdir {kfd}/1" + r = run_guard(kfd, "--idle-secs", "1", "--max-attempts", "2", "--", + "bash", "-c", cmd) + self.assertEqual(r.returncode, 75, r.stderr) + self.assertIn("CONTENDED foreign_gpu=[1:", r.stderr) + self.assertIn("every attempt (2) was contended", r.stderr) + + def test_the_commands_own_gpu_process_is_not_contention(self): + with tempfile.TemporaryDirectory() as kfd: + # $$ of the inner bash is a descendant of the guarded command. + cmd = f"mkdir -p {kfd}/$$; sleep 2.5; rmdir {kfd}/$$" + r = run_guard(kfd, "--idle-secs", "1", "--", "bash", "-c", cmd) + self.assertEqual(r.returncode, 0, r.stderr) + self.assertNotIn("CONTENDED", r.stderr) + + def test_a_busy_gpu_before_the_run_times_out_without_running(self): + with tempfile.TemporaryDirectory() as kfd, tempfile.TemporaryDirectory() as out: + os.mkdir(pathlib.Path(kfd) / "1") + marker = pathlib.Path(out) / "ran" + r = run_guard(kfd, "--idle-secs", "1", "--max-wait", "2", "--", + "touch", str(marker)) + self.assertEqual(r.returncode, 75, r.stderr) + self.assertFalse(marker.exists()) + + def test_a_compiler_counts_as_contention(self): + with tempfile.TemporaryDirectory() as kfd: + r = run_guard(kfd, "--idle-secs", "1", "--max-wait", "2", "--", "true", + env_extra={"ROCM_GPU_GUARD_COMPILER_RE": "^(bash)$"}) + self.assertEqual(r.returncode, 75, r.stderr) + self.assertIn("gave up after", r.stderr) + + +def k(op: str) -> str: + """A kernel name spelled the way rocprofv3 writes the overlay's kernels.""" + names = { + "qmv": "void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*)", + "gqmv": "void mlx::core::rocm::gather_qmv_wide_kernel(x)", + "rms": "void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*)", + "add": "void mlx::core::rocm::binary_vv(x)", + "mul": "void mlx::core::rocm::binary_g(x)", + "div": "void mlx::core::rocm::binary_vs(x)", + "rope": "void mlx::core::rocm::rope_single_1d<__half, false, true>(x)", + "kv": "void mlx::core::rocm::copy_gg_byval<__half, __half, long>(x)", + "cast": "void mlx::core::rocm::copy_v<__half, float, unsigned int, 4>(x)", + "sdpa": "void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(x)", + "swiglu": "void mlx::core::rocm::CV2ISigmoidADV2IMultiplyGH_VV_V2V2_614_contiguous(x)", + "silu": "void mlx::core::rocm::BV2ISigmoidACV2OMultiplyCD_V_V2_614_contiguous(x)", + "argmax": "void mlx::core::rocm::arg_reduce_final<__half, mlx::core::rocm::ArgMax<__half>, 256>(x)", + "embed": "void mlx::core::rocm::gather_rows_kernel(x)", + "sort": "void mlx::core::rocm::block_sort_kernel(x)", + "arange": "void mlx::core::rocm::arange_kernel(unsigned int*)", + "colsum": "void mlx::core::rocm::col_reduce_small(x)", + "conv": "void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(x)", + "exp": "void mlx::core::rocm::unary_v(x)", + "scan": "void mlx::core::rocm::contiguous_scan(x)", + } + return names[op] + + +def seq(*ops, grid=None): + grid = grid or {} + out, t = [], 1000 + for op in ops: + out.append(rdp.Dispatch(k(op.split(":")[0]), t, t + 10, int(op.split(":")[1]) if ":" in op else 256)) + t += 20 + return out + + +class RoleTests(unittest.TestCase): + def test_dense_layer_and_sampler_tail(self): + d = seq("embed", "rms", "qmv:1000", "rope", "kv", "kv", "rope", "sdpa", "qmv:1000", + "add", "rms", "qmv:3000", "qmv:3000", "swiglu", "qmv:1000", "add", "rms", + "qmv:90000", "cast", "argmax", "embed") + roles = rdp.assign_roles(d) + self.assertEqual(roles[3:7], ["rope_append"] * 4) + self.assertEqual(roles[9:11], ["add_rms_join_post_attn"] * 2) + self.assertEqual(roles[15:17], ["add_rms_join"] * 2) + self.assertEqual(roles[18:20], ["sampler_tail"] * 2) + self.assertIsNone(roles[17]) # the lm_head GEMV itself + self.assertIsNone(roles[20]) # the next step's embedding + + def test_moe_block(self): + d = seq("add", "rms", "qmv:64", "sort", "arange", "arange", "div", "gqmv", "gqmv", + "swiglu", "gqmv", "mul", "cast", "colsum", "add", "rms", "qmv:90000", "argmax") + roles = rdp.assign_roles(d) + self.assertEqual([roles[i] for i in (7, 8, 10)], ["moe_expert_gemv"] * 3) + self.assertEqual(roles[9], "moe_activation") + self.assertEqual(roles[11:14], ["moe_weighted_sum"] * 3) + self.assertEqual(roles[4:6], ["moe_gather_indices"] * 2) + self.assertIsNone(roles[3]) # router top-k stays outside the fused kernel + self.assertIsNone(roles[6]) # score normalisation too + self.assertEqual(roles[14:16], ["add_rms_join"] * 2) + + def test_mamba_mixer(self): + d = seq("rms", "qmv:200000", "cast", "kv", "conv", "exp", "mul", "silu", "scan", + "mul", "silu", "rms", "mul", "qmv:50000", "qmv:900000", "argmax") + roles = rdp.assign_roles(d) + self.assertEqual(roles[2], "ssm_step") + self.assertEqual(roles[3:5], ["ssm_conv"] * 2) + self.assertEqual([roles[i] for i in (5, 6, 8, 9)], ["ssm_step"] * 4) + self.assertEqual([roles[i] for i in (7, 10)], ["ssm_silu"] * 2) + self.assertEqual(roles[11:13], ["ssm_gated_norm"] * 2) + self.assertIsNone(roles[1]) + self.assertIsNone(roles[13]) + + +class ReachTests(unittest.TestCase): + def test_fused_norm_and_rope_ship_off_and_rope_tables_never_reach(self): + default, optin, _ = rdp.reach("2063", {"model_type": "llama", + "rope_scaling": {"rope_type": "llama3"}}, False) + self.assertEqual(default, ()) + self.assertEqual(optin, ("add_rms_join_post_attn",)) + _, optin, _ = rdp.reach("2063", {"model_type": "llama"}, False) + self.assertEqual(optin, ("add_rms_join_post_attn", "rope_append")) + + def test_moe_reach_follows_the_caller(self): + self.assertTrue(rdp.reach("2065", {"model_type": "qwen3_moe"}, False)[0]) + self.assertFalse(rdp.reach("2065", {"model_type": "granitemoehybrid"}, False)[0]) + self.assertFalse(rdp.reach("2065", {"model_type": "nemotron_h"}, False)[0]) + + def test_samplers_only_in_sampled_runs(self): + self.assertFalse(rdp.reach("2064", {}, False)[0]) + self.assertTrue(rdp.reach("2064", {}, True)[0]) + + +class WindowTests(unittest.TestCase): + def test_busy_is_the_union_of_overlapping_dispatches(self): + d = [rdp.Dispatch("a", 0, 10), rdp.Dispatch("b", 5, 15), rdp.Dispatch("c", 20, 30)] + self.assertEqual(rdp.busy_ns(d, 0, 100), 25) + self.assertEqual(rdp.busy_ns(d, 8, 25), 12) + + def test_summarize_cuts_the_decode_window_by_the_phase_marks(self): + with tempfile.TemporaryDirectory() as tmp: + tmp = pathlib.Path(tmp) + trace = tmp / "t_kernel_trace.csv" + rows = [("warm", 100_000, 110_000), ("prefill", 500_000, 590_000), + ("decode", 700_000, 710_000), ("decode", 720_000, 760_000), + ("after", 5_000_000, 5_010_000)] + trace.write_text('"Kind","Kernel_Name","Start_Timestamp","End_Timestamp"\n' + "".join( + f'"KERNEL_DISPATCH","void mlx::core::rocm::{n}_kernel(x)",{s},{e}\n' + for n, s, e in rows)) + log = tmp / "bench.log" + log.write_text( + "[phase] warmup_start monotonic_ns=1 boottime_ns=90000\n" + "[phase] measured_start monotonic_ns=2 boottime_ns=400000\n" + "[phase] decode_start monotonic_ns=3 boottime_ns=600000\n" + "[phase] measured_end monotonic_ns=4 boottime_ns=800000\n" + " Prompt tokens: 512\n Generated tokens: 2\n" + " Prefill: 1.00 ms (512.00 tok/s)\n Decode: 0.20 ms (10000.00 tok/s)\n") + s = rdp.summarize(trace, log, None, None, "t", tmp) + self.assertEqual(s["clock"], "boottime") + self.assertEqual(s["decode_dispatches"], 2) + self.assertEqual(s["decode_gpu_busy_ms"], 0.05) + self.assertEqual(s["checks"]["kernels_straddling_decode_start"], 0) + self.assertEqual(s["checks"]["idle_gap_before_first_decode_dispatch_us"], 110.0) + self.assertTrue((tmp / "t_decode_kernels.csv").exists()) + + +if __name__ == "__main__": + unittest.main() From c098e02d606a60fcc518d282e70a51ac5b08aaf7 Mon Sep 17 00:00:00 2001 From: Jeongkyu Shin Date: Wed, 30 Sep 2026 23:09:27 +0900 Subject: [PATCH 3/6] docs(rocm): publish the gfx1151 decode profile and #1814 port order Profiles greedy and sampled decode (pp512/tg128) of Llama-3.1-8B, Qwen3-30B-A3B, granite-4.0-h-tiny and Nemotron-3-Nano-30B-A3B on gfx1151 with rocprofv3, every run under the idle-GPU guard (16 accepted runs; one rejected by the guard, one discarded and rerun by hand), and attributes decode GPU time to the port units split from #1814. Shipped-settings share of decode GPU time: #2067 SSM update 29.8% (granite) and 20.0% (Nemotron) plus more than half their dispatches; #2065 fused MoE 46.8% of Qwen3, mostly GEMVs already at about 181 GB/s; #2064 samplers 0 greedy, 0.4 to 3.8% sampled; #2063 zero (both fusions ship off) and #2068 zero (no paged path in single-stream decode). Implied order: 7, 5, 4, 8, 3. The rocprofv3 stats CSVs, per-kernel decode tables, summaries, bench logs and the guard log are under benchmarks/rocm_profiles/gfx1151_929c80ab/. The report also records the gather_mm test outcome. Refs #2061 --- ...lama-3.1-8B-Instruct-4bit_greedy_bench.log | 16 + ...8B-Instruct-4bit_greedy_decode_kernels.csv | 20 + ...1-8B-Instruct-4bit_greedy_kernel_stats.csv | 27 ++ ....1-8B-Instruct-4bit_greedy_plain_bench.log | 12 + ...a-3.1-8B-Instruct-4bit_greedy_summary.json | 175 ++++++++ ...-3.1-8B-Instruct-4bit_t0.7-p0.95_bench.log | 16 + ...nstruct-4bit_t0.7-p0.95_decode_kernels.csv | 53 +++ ...-Instruct-4bit_t0.7-p0.95_kernel_stats.csv | 59 +++ ...1-8B-Instruct-4bit_t0.7-p0.95_summary.json | 182 ++++++++ ...-Llama-3.1-8B-Instruct-4bit_t0.7_bench.log | 16 + ...1-8B-Instruct-4bit_t0.7_decode_kernels.csv | 38 ++ ...3.1-8B-Instruct-4bit_t0.7_kernel_stats.csv | 45 ++ ...ama-3.1-8B-Instruct-4bit_t0.7_summary.json | 180 ++++++++ ...otron-3-Nano-30B-A3B-4bit_greedy_bench.log | 21 + ...ano-30B-A3B-4bit_greedy_decode_kernels.csv | 51 +++ ...-Nano-30B-A3B-4bit_greedy_kernel_stats.csv | 63 +++ ...3-Nano-30B-A3B-4bit_greedy_plain_bench.log | 17 + ...on-3-Nano-30B-A3B-4bit_greedy_summary.json | 188 +++++++++ ...n-3-Nano-30B-A3B-4bit_t0.7-p0.95_bench.log | 21 + ...30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv | 82 ++++ ...o-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv | 93 ++++ ...-Nano-30B-A3B-4bit_t0.7-p0.95_summary.json | 190 +++++++++ ...emotron-3-Nano-30B-A3B-4bit_t0.7_bench.log | 21 + ...-Nano-30B-A3B-4bit_t0.7_decode_kernels.csv | 67 +++ ...-3-Nano-30B-A3B-4bit_t0.7_kernel_stats.csv | 79 ++++ ...tron-3-Nano-30B-A3B-4bit_t0.7_summary.json | 188 +++++++++ .../Qwen3-30B-A3B-4bit_greedy_bench.log | 15 + ...en3-30B-A3B-4bit_greedy_decode_kernels.csv | 28 ++ ...Qwen3-30B-A3B-4bit_greedy_kernel_stats.csv | 43 ++ .../Qwen3-30B-A3B-4bit_greedy_plain_bench.log | 11 + .../Qwen3-30B-A3B-4bit_greedy_summary.json | 179 ++++++++ .../Qwen3-30B-A3B-4bit_t0.7-p0.95_bench.log | 15 + ...30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv | 58 +++ ...3-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv | 72 ++++ ...Qwen3-30B-A3B-4bit_t0.7-p0.95_summary.json | 183 ++++++++ .../Qwen3-30B-A3B-4bit_t0.7_bench.log | 15 + ...Qwen3-30B-A3B-4bit_t0.7_decode_kernels.csv | 44 ++ .../Qwen3-30B-A3B-4bit_t0.7_kernel_stats.csv | 58 +++ .../Qwen3-30B-A3B-4bit_t0.7_summary.json | 182 ++++++++ .../granite-4.0-h-tiny-4bit_greedy_bench.log | 32 ++ ...-4.0-h-tiny-4bit_greedy_decode_kernels.csv | 54 +++ ...te-4.0-h-tiny-4bit_greedy_kernel_stats.csv | 77 ++++ ...ite-4.0-h-tiny-4bit_greedy_plain_bench.log | 12 + ...ranite-4.0-h-tiny-4bit_greedy_summary.json | 189 +++++++++ ...anite-4.0-h-tiny-4bit_t0.7-p0.95_bench.log | 16 + ...-h-tiny-4bit_t0.7-p0.95_decode_kernels.csv | 82 ++++ ....0-h-tiny-4bit_t0.7-p0.95_kernel_stats.csv | 103 +++++ ...te-4.0-h-tiny-4bit_t0.7-p0.95_summary.json | 190 +++++++++ .../granite-4.0-h-tiny-4bit_t0.7_bench.log | 16 + ...te-4.0-h-tiny-4bit_t0.7_decode_kernels.csv | 68 +++ ...nite-4.0-h-tiny-4bit_t0.7_kernel_stats.csv | 90 ++++ .../granite-4.0-h-tiny-4bit_t0.7_summary.json | 189 +++++++++ .../rocm_profiles/gfx1151_929c80ab/guard.log | 396 ++++++++++++++++++ .../rocm-decode-profile-gfx1151-2026-09-30.md | 162 +++++++ docs/installation.md | 3 + 55 files changed, 4502 insertions(+) create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_plain_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_plain_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_plain_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_plain_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_bench.log create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_decode_kernels.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_kernel_stats.csv create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_summary.json create mode 100644 benchmarks/rocm_profiles/gfx1151_929c80ab/guard.log create mode 100644 docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_bench.log new file mode 100644 index 000000000..815efb3dd --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_bench.log @@ -0,0 +1,16 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.165 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] warmup_start monotonic_ns=169763750669641 boottime_ns=169763750669691 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] measured_start monotonic_ns=169765813493615 boottime_ns=169765813493645 +[phase] decode_start monotonic_ns=169766321703540 boottime_ns=169766321703560 +[phase] measured_end monotonic_ns=169769539602196 boottime_ns=169769539602216 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 478.76 ms (1069.42 tok/s) + Decode: 3217.90 ms (39.78 tok/s) + MLX peak memory: 20.60 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_decode_kernels.csv new file mode 100644 index 000000000..729af17e6 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_decode_kernels.csv @@ -0,0 +1,20 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)",qmv,,,20447,159.74,2832888679,138548,94.812 +"void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)",sdpa,,,4064,31.75,63928915,15731,2.140 +"void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,8255,64.49,34238309,4148,1.146 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn,2063,8128,63.5,16540628,2035,0.554 +"void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,rope_append,2063,8256,64.5,15280116,1851,0.511 +"void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)",rope,rope_append,2063,8128,63.5,12891275,1586,0.431 +"void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)",compiled,,,4064,31.75,8863779,2181,0.297 +"void mlx::core::rocm::arg_reduce_partial<__half, mlx::core::rocm::ArgMax<__half>, 256>(__half const*, __half*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",arg_reduce,sampler_tail,2064,127,0.99,713374,5617,0.024 +"void mlx::core::rocm::copy_v<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",copy,sampler_tail,2064,127,0.99,388288,3057,0.013 +"void mlx::core::rocm::gather_rows_kernel<__half, int>(__half const*, int const*, __half*, long, unsigned int, int)",gather_scatter,,,254,1.98,370163,1457,0.012 +"void mlx::core::rocm::gather_axis_kernel<__half, int, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,127,0.99,276394,2176,0.009 +"void mlx::core::rocm::arg_reduce_final<__half, mlx::core::rocm::ArgMax<__half>, 256>(__half const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",arg_reduce,sampler_tail,2064,127,0.99,245578,1934,0.008 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,243372,1916,0.008 +"void mlx::core::rocm::scatter_axis_kernel<__half, int, false, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,237324,1869,0.008 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,199572,1571,0.007 +"void mlx::core::rocm::copy_v<__half, float, unsigned int, 4>(__half const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,187838,1479,0.006 +"void mlx::core::rocm::affine_dequantize_packed_kernel<__half, 4>(unsigned char const*, __half const*, __half const*, __half*, unsigned long, int)",dequantize,,,127,0.99,144035,1134,0.005 +"void mlx::core::rocm::copy_s<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",copy,,,64,0.5,132975,2078,0.004 +"void mlx::core::rocm::copy_v(float const*, __half*, unsigned int)",copy,sampler_tail,2064,127,0.99,124713,982,0.004 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_kernel_stats.csv new file mode 100644 index 000000000..758093603 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_kernel_stats.csv @@ -0,0 +1,27 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)",23508,3251875430,138330.586609,73.19,57067,2458585,103690.064508 +"Cijk_Alik_Bljk_HHS_BH_Bias_HA_S_SAV_UserArgs_MT80x128x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT5_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR0_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG16_8_1_WGMXCC1",192,393932333,2051730.901042,8.87,1199246,2792190,490966.644453 +"Cijk_Alik_Bljk_HHS_BH_Bias_HA_S_SAV_UserArgs_MT64x128x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT4_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR0_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG16_8_1_WGMXCC1",128,236053820,1844170.468750,5.31,895597,3243274,837111.458130 +"void mlx::core::rocm::affine_dequantize_packed_kernel<__half, 4>(unsigned char const*, __half const*, __half const*, __half*, unsigned long, int)",468,235582978,503382.431624,5.30,961,1783020,483086.513258 +"void mlx::core::rocm::kernel_sdpa_flash_wmma<__half, true, 128, 64, 64>(__half const*, __half const*, __half const*, __half*, float*, mlx::core::rocm::FAWmmaParams)",64,95527733,1492620.828125,2.15,1450638,2002912,67198.957956 +"void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)",4672,72747549,15570.965111,1.64,8255,32421,1698.779358 +"void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)",9620,46824909,4867.454158,1.05,3647,139662,8129.922147 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",9472,26370407,2784.037901,0.5935,1562,306414,7041.504417 +"void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",9792,23867678,2437.467116,0.5372,1001,176971,7721.667921 +"void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)",4736,19325787,4080.613809,0.4350,1683,241412,16195.769640 +"void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)",9344,15007453,1606.105843,0.3378,1082,17553,1017.044898 +"void mlx::core::rocm::copy_gg_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,11075268,115367.375000,0.2493,27051,573434,107719.186459 +"void mlx::core::rocm::rope_freqs<__half, false, true>(__half const*, __half*, int const*, float const*, float, float, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, long, int, HIP_vector_type, long)",128,6346774,49584.171875,0.1428,9739,99587,38078.550437 +"void mlx::core::rocm::copy_g_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",64,4396070,68688.593750,0.0989,63399,310862,30857.517482 +"void mlx::core::rocm::arg_reduce_partial<__half, mlx::core::rocm::ArgMax<__half>, 256>(__half const*, __half*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",148,848108,5730.459459,0.0191,5130,10700,795.267314 +"void mlx::core::rocm::gather_rows_kernel<__half, int>(__half const*, int const*, __half*, long, unsigned int, int)",296,480889,1624.625000,0.0108,962,27973,1923.671979 +"void mlx::core::rocm::copy_v<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",148,456336,3083.351351,0.0103,2364,19516,1697.485839 +"void mlx::core::rocm::gather_axis_kernel<__half, int, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",148,336504,2273.675676,7.574e-03,1522,6853,1183.689017 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,328052,2216.567568,7.384e-03,1282,25288,2340.444749 +"void mlx::core::rocm::scatter_axis_kernel<__half, int, false, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,284974,1925.500000,6.414e-03,1322,10019,1369.125897 +"void mlx::core::rocm::arg_reduce_final<__half, mlx::core::rocm::ArgMax<__half>, 256>(__half const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",148,283365,1914.628378,6.378e-03,1643,10339,898.730286 +"void mlx::core::rocm::copy_s<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",128,251754,1966.828125,5.666e-03,1563,5330,556.919227 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",148,248064,1676.108108,5.583e-03,881,6452,1437.128032 +"void mlx::core::rocm::copy_v<__half, float, unsigned int, 4>(__half const*, float*, unsigned int)",148,222987,1506.668919,5.019e-03,1042,8296,991.654252 +"__amd_rocclr_fillBufferUnAligned",1,167714,167714.000000,3.775e-03,167714,167714,0.00000000e+00 +"void mlx::core::rocm::copy_v(float const*, __half*, unsigned int)",148,158217,1069.033784,3.561e-03,801,6372,576.180957 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_plain_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_plain_bench.log new file mode 100644 index 000000000..e764fd20d --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_plain_bench.log @@ -0,0 +1,12 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.200 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 475.90 ms (1075.86 tok/s) + Decode: 3452.34 ms (37.08 tok/s) + MLX peak memory: 20.60 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_summary.json new file mode 100644 index 000000000..bc0d89b5e --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_greedy_summary.json @@ -0,0 +1,175 @@ +{ + "name": "Meta-Llama-3.1-8B-Instruct-4bit_greedy", + "model_type": "llama", + "temperature": 0.0, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 39.78, + "profiled_prefill_tok_s": 1069.42, + "plain_decode_tok_s": 37.08, + "plain_prefill_tok_s": 1075.86, + "profiler_decode_slowdown_pct": -6.79, + "decode_dispatches": 62930, + "dispatches_per_token": 491.6, + "decode_wall_ms": 3217.899, + "decode_gpu_sum_ms": 2987.895, + "decode_gpu_busy_ms": 2987.895, + "gpu_ms_per_token": 23.3429, + "host_gap_ms_per_token": 1.7969, + "host_gap_pct_of_wall": 7.15, + "plain_wall_ms_per_token": 26.9687, + "plain_host_gap_ms_per_token_est": 3.6258, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 30510.9, + "last_dispatch_before_decode": "void mlx::core::rocm::arg_reduce_final<__half, mlx::core::rocm::ArgMax<__half>, 256>(__half const*, unsigned int const*," + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 159.74, + "share_pct": 94.81 + }, + { + "kernel": "void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)", + "class": "sdpa", + "calls_per_token": 31.75, + "share_pct": 2.14 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 64.49, + "share_pct": 1.15 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)", + "class": "binary", + "calls_per_token": 63.5, + "share_pct": 0.55 + }, + { + "kernel": "void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "copy", + "calls_per_token": 64.5, + "share_pct": 0.51 + }, + { + "kernel": "void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)", + "class": "rope", + "calls_per_token": 63.5, + "share_pct": 0.43 + }, + { + "kernel": "void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)", + "class": "compiled", + "calls_per_token": 31.75, + "share_pct": 0.3 + }, + { + "kernel": "void mlx::core::rocm::arg_reduce_partial<__half, mlx::core::rocm::ArgMax<__half>, 256>(__half const*, __half*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)", + "class": "arg_reduce", + "calls_per_token": 0.99, + "share_pct": 0.02 + }, + { + "kernel": "void mlx::core::rocm::copy_v<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)", + "class": "copy", + "calls_per_token": 0.99, + "share_pct": 0.01 + }, + { + "kernel": "void mlx::core::rocm::gather_rows_kernel<__half, int>(__half const*, int const*, __half*, long, unsigned int, int)", + "class": "gather_scatter", + "calls_per_token": 1.98, + "share_pct": 0.01 + } + ], + "class_share_pct": { + "qmv": 94.81, + "sdpa": 2.14, + "rms_norm": 1.15, + "binary": 0.56, + "copy": 0.54, + "rope": 0.43, + "compiled": 0.3, + "gather_scatter": 0.04, + "arg_reduce": 0.03, + "dequantize": 0.0 + }, + "role_share_pct": { + "unattributed": 97.41, + "add_rms_join": 0.85, + "add_rms_join_post_attn": 0.83, + "rope_append": 0.83, + "sampler_tail": 0.08, + "moe_expert_gemv": 0.0, + "moe_activation": 0.0, + "moe_weighted_sum": 0.0, + "moe_gather_indices": 0.0, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 229.2, + "rope_append": 127.5, + "add_rms_join_post_attn": 63.5, + "add_rms_join": 63.5, + "sampler_tail": 7.9, + "moe_expert_gemv": 0.0, + "moe_activation": 0.0, + "moe_weighted_sum": 0.0, + "moe_gather_indices": 0.0, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 1.66, + "fallback_dispatches_per_token": 191.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.83, + "note": "both paths ship off (FUSED_ADD_RMSNORM_DEFAULT / FUSED_ROPE_APPEND_DEFAULT false); only the post-attention join calls the fused norm; rope_scaling builds a frequency table, which the RoPE kernel cannot take" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.08, + "fallback_dispatches_per_token": 7.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "greedy argmax dispatches neither sampler kernel" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "llama does not call forward_fused_kernel" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "no Mamba2 layer" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_bench.log new file mode 100644 index 000000000..b8268d4cd --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_bench.log @@ -0,0 +1,16 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.175 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] warmup_start monotonic_ns=178268956285412 boottime_ns=178268956285462 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] measured_start monotonic_ns=178274491402960 boottime_ns=178274491403000 +[phase] decode_start monotonic_ns=178274995511561 boottime_ns=178274995511621 +[phase] measured_end monotonic_ns=178278446373119 boottime_ns=178278446373179 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 477.15 ms (1073.04 tok/s) + Decode: 3450.86 ms (37.09 tok/s) + MLX peak memory: 20.60 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_decode_kernels.csv new file mode 100644 index 000000000..8255ba2eb --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_decode_kernels.csv @@ -0,0 +1,53 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)",qmv,,,20447,159.74,2855368631,139647,93.556 +"void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)",sdpa,,,4064,31.75,59859806,14729,1.961 +"void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,8255,64.49,37444045,4536,1.227 +"void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,rope_append,2063,8256,64.5,19348408,2344,0.634 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn,2063,8128,63.5,17761039,2185,0.582 +"void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)",rope,rope_append,2063,8128,63.5,15526685,1910,0.509 +"void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)",compiled,,,4064,31.75,9702254,2387,0.318 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",sort,sampler_tail,2064,889,6.95,7194579,8093,0.236 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,3456062,27213,0.113 +"void mlx::core::rocm::contiguous_scan<__half, __half, mlx::core::rocm::Sum, 4, true, false>(__half const*, __half*, int)",scan,sampler_tail,2064,127,0.99,3164368,24916,0.104 +"void mlx::core::rocm::softmax_kernel<__half, __half, 1024, 4>(__half const*, __half*, int)",softmax,sampler_tail,2064,127,0.99,2848690,22431,0.093 +"void mlx::core::rocm::gather_axis_kernel<__half, int, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,508,3.97,2634665,5186,0.086 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(__half*, std::iterator_traits<__half*>::value_type*, __half*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()<__half*, __half*, unsigned int*, unsigned int*>(__half*, __half*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(__half*, std::iterator_traits<__half*>::value_type*, __half*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()<__half*, __half*, unsigned int*, unsigned int*>(__half*, __half*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",sort,sampler_tail,2064,254,1.98,1764580,6947,0.058 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,1759696,13856,0.058 +__amd_rocclr_copyBuffer,copy,sampler_tail,2064,508,3.97,1422428,2800,0.047 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail,2064,127,0.99,828107,6521,0.027 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",binary,sampler_tail,2064,127,0.99,712171,5608,0.023 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,615596,4847,0.020 +"void mlx::core::rocm::binary_vs(__half const*, __half const*, __half*, unsigned int)",binary,sampler_tail,2064,127,0.99,602728,4746,0.020 +"void mlx::core::rocm::copy_v<__half, float, unsigned int, 4>(__half const*, float*, unsigned int)",copy,sampler_tail,2064,254,1.98,590096,2323,0.019 +"void mlx::core::rocm::ternary_g(bool const*, __half const*, __half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,551554,4343,0.018 +"void mlx::core::rocm::gather_rows_kernel<__half, int>(__half const*, int const*, __half*, long, unsigned int, int)",gather_scatter,,,254,1.98,550476,2167,0.018 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,481584,3792,0.016 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",other,sampler_tail,2064,254,1.98,475782,1873,0.016 +__amd_rocclr_fillBufferUnAligned,copy,sampler_tail,2064,381,2.98,440902,1157,0.014 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,435848,1716,0.014 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,431202,3395,0.014 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,429129,1689,0.014 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::default_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1})",ternary,sampler_tail,2064,254,1.98,422401,1663,0.014 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,411123,1619,0.013 +"void mlx::core::rocm::copy_v<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",copy,sampler_tail,2064,127,0.99,379679,2990,0.012 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,374909,2952,0.012 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,331871,1307,0.011 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,308854,2432,0.010 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,278841,1098,0.009 +"void mlx::core::rocm::binary_vs(__half const*, __half const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,259366,2042,0.008 +"void mlx::core::rocm::unary_v(__half const*, __half*, unsigned int)",unary,sampler_tail,2064,127,0.99,252994,1992,0.008 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,243659,1919,0.008 +"void mlx::core::rocm::scatter_axis_kernel<__half, int, false, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,242172,1907,0.008 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,207390,1633,0.007 +"void mlx::core::rocm::affine_dequantize_packed_kernel<__half, 4>(unsigned char const*, __half const*, __half const*, __half*, unsigned long, int)",dequantize,,,127,0.99,206303,1624,0.007 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,189955,1496,0.006 +"void mlx::core::rocm::copy_v(float const*, __half*, unsigned int)",copy,sampler_tail,2064,127,0.99,171960,1354,0.006 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,164910,1299,0.005 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",sort,sampler_tail,2064,127,0.99,163219,1285,0.005 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,155605,1225,0.005 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,154692,1218,0.005 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,150164,1182,0.005 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,148882,1172,0.005 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,140544,1107,0.005 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,139788,1101,0.005 +"void mlx::core::rocm::copy_s<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",copy,,,64,0.5,137536,2149,0.005 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_kernel_stats.csv new file mode 100644 index 000000000..9c21569fc --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_kernel_stats.csv @@ -0,0 +1,59 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)",23508,3278204582,139450.594776,72.08,58028,1781493,103015.275305 +"Cijk_Alik_Bljk_HHS_BH_Bias_HA_S_SAV_UserArgs_MT80x128x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT5_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR0_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG16_8_1_WGMXCC1",192,386816489,2014669.213542,8.51,1163857,2558666,483540.151917 +"void mlx::core::rocm::affine_dequantize_packed_kernel<__half, 4>(unsigned char const*, __half const*, __half const*, __half*, unsigned long, int)",468,261363907,558469.886752,5.75,1242,2244589,529620.414317 +"Cijk_Alik_Bljk_HHS_BH_Bias_HA_S_SAV_UserArgs_MT64x128x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT4_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR0_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG16_8_1_WGMXCC1",128,229354058,1791828.578125,5.04,764430,2997327,832824.209656 +"void mlx::core::rocm::kernel_sdpa_flash_wmma<__half, true, 128, 64, 64>(__half const*, __half const*, __half const*, __half*, float*, mlx::core::rocm::FAWmmaParams)",64,94732093,1480188.953125,2.08,1442939,1517920,19705.453690 +"void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)",4672,67998918,14554.562928,1.50,8336,46166,1719.318326 +"void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)",9620,51710380,5375.299376,1.14,3687,164067,9551.685322 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",9472,38296866,4043.165752,0.8421,1562,345285,20663.201607 +"void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",9792,31509174,3217.848652,0.6928,1042,172844,10339.578540 +"void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)",4736,23323274,4924.677787,0.5128,1683,292266,23005.578195 +"void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)",9344,17755784,1900.233733,0.3904,1082,10580,1434.020259 +"void mlx::core::rocm::copy_g_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",64,9372693,146448.328125,0.2061,63518,315310,117841.318223 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",1036,8405324,8113.247104,0.1848,6813,16911,1077.404152 +"void mlx::core::rocm::copy_gg_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,7653241,79721.260417,0.1683,12503,415337,73294.097240 +"void mlx::core::rocm::rope_freqs<__half, false, true>(__half const*, __half*, int const*, float const*, float, float, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, long, int, HIP_vector_type, long)",128,6342221,49548.601562,0.1394,8936,104596,37986.979572 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,4027253,27211.168919,0.0885,26570,31619,501.447638 +"void mlx::core::rocm::contiguous_scan<__half, __half, mlx::core::rocm::Sum, 4, true, false>(__half const*, __half*, int)",148,3787332,25590.081081,0.0833,24286,125595,8281.492070 +"void mlx::core::rocm::softmax_kernel<__half, __half, 1024, 4>(__half const*, __half*, int)",148,3396954,22952.391892,0.0747,18756,99226,6715.545261 +"void mlx::core::rocm::gather_axis_kernel<__half, int, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",592,3101420,5238.885135,0.0682,1803,27091,2137.614339 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(__half*, std::iterator_traits<__half*>::value_type*, __half*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()<__half*, __half*, unsigned int*, unsigned int*>(__half*, __half*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(__half*, std::iterator_traits<__half*>::value_type*, __half*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()<__half*, __half*, unsigned int*, unsigned int*>(__half*, __half*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",296,2102932,7104.500000,0.0462,5089,33302,2505.524165 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,2055530,13888.716216,0.0452,13505,20038,546.979685 +"__amd_rocclr_copyBuffer",592,1673576,2826.986486,0.0368,1322,12824,1934.925211 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",148,961595,6497.263514,0.0211,6171,7654,237.041096 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",148,853713,5768.331081,0.0188,4969,27812,1833.899212 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,718346,4853.689189,0.0158,4568,5490,138.132582 +"void mlx::core::rocm::binary_vs(__half const*, __half const*, __half*, unsigned int)",148,704436,4759.702703,0.0155,4408,5931,201.502737 +"void mlx::core::rocm::copy_v<__half, float, unsigned int, 4>(__half const*, float*, unsigned int)",296,688601,2326.354730,0.0151,1243,8175,893.774346 +"void mlx::core::rocm::gather_rows_kernel<__half, int>(__half const*, int const*, __half*, long, unsigned int, int)",296,668930,2259.898649,0.0147,1122,20999,1332.338081 +"void mlx::core::rocm::ternary_g(bool const*, __half const*, __half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,643045,4344.898649,0.0141,3447,4769,207.121271 +"__amd_rocclr_fillBufferUnAligned",445,621439,1396.492135,0.0137,761,88486,4268.257512 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",296,586350,1980.912162,0.0129,1362,23524,1340.113835 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,578164,3906.513514,0.0127,3606,19997,1369.963853 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,527382,1781.695946,0.0116,1122,12944,862.444194 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,505183,3413.398649,0.0111,3046,4970,175.672226 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,498421,1683.854730,0.0110,1082,5170,339.920513 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::default_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1})",296,492292,1663.148649,0.0108,1403,2164,130.088698 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,484298,1636.141892,0.0106,1242,3126,323.010553 +"void mlx::core::rocm::copy_v<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",148,451974,3053.878378,9.938e-03,2605,12463,841.552880 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,434701,2937.168919,9.558e-03,2685,3646,153.727535 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,385648,1302.864865,8.479e-03,842,2565,264.811456 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,384915,2600.777027,8.463e-03,1683,25849,2338.933683 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,326607,1103.402027,7.181e-03,841,2324,128.053553 +"void mlx::core::rocm::unary_v(__half const*, __half*, unsigned int)",148,306652,2071.972973,6.743e-03,1883,11822,856.648587 +"void mlx::core::rocm::binary_vs(__half const*, __half const*, bool*, unsigned int)",148,305293,2062.790541,6.713e-03,1523,2965,190.813876 +"void mlx::core::rocm::scatter_axis_kernel<__half, int, false, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,288417,1948.763514,6.342e-03,1403,7174,952.287260 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,285135,1926.587838,6.269e-03,1482,2324,143.032984 +"void mlx::core::rocm::copy_s<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",128,258647,2020.679688,5.687e-03,1603,5451,571.685255 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,242497,1638.493243,5.332e-03,1242,2725,191.085424 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",148,231594,1564.824324,5.092e-03,1162,6211,747.556507 +"void mlx::core::rocm::copy_v(float const*, __half*, unsigned int)",148,200972,1357.918919,4.419e-03,962,5330,1005.209624 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(__half*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",148,196599,1328.371622,4.323e-03,1122,5530,490.336595 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,189757,1282.141892,4.172e-03,1042,6011,438.501297 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",148,180530,1219.797297,3.969e-03,1002,4648,323.834067 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,178936,1209.027027,3.934e-03,961,2285,166.933729 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,175373,1184.952703,3.856e-03,962,2284,186.722404 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,172808,1167.621622,3.800e-03,1041,2325,168.338920 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,165632,1119.135135,3.642e-03,1001,1603,81.134947 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,162389,1097.222973,3.571e-03,921,1523,119.330720 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_summary.json new file mode 100644 index 000000000..099ed8541 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95_summary.json @@ -0,0 +1,182 @@ +{ + "name": "Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95", + "model_type": "llama", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 37.09, + "profiled_prefill_tok_s": 1073.04, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 70042, + "dispatches_per_token": 547.2, + "decode_wall_ms": 3450.862, + "decode_gpu_sum_ms": 3052.038, + "decode_gpu_busy_ms": 3052.038, + "gpu_ms_per_token": 23.844, + "host_gap_ms_per_token": 3.1158, + "host_gap_pct_of_wall": 11.56, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 27567.2, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 159.74, + "share_pct": 93.56 + }, + { + "kernel": "void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)", + "class": "sdpa", + "calls_per_token": 31.75, + "share_pct": 1.96 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 64.49, + "share_pct": 1.23 + }, + { + "kernel": "void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "copy", + "calls_per_token": 64.5, + "share_pct": 0.63 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)", + "class": "binary", + "calls_per_token": 63.5, + "share_pct": 0.58 + }, + { + "kernel": "void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)", + "class": "rope", + "calls_per_token": 63.5, + "share_pct": 0.51 + }, + { + "kernel": "void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)", + "class": "compiled", + "calls_per_token": 31.75, + "share_pct": 0.32 + }, + { + "kernel": "void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})", + "class": "sort", + "calls_per_token": 6.95, + "share_pct": 0.24 + }, + { + "kernel": "void mlx::core::rocm::contiguous_scan(float const*, float*, int)", + "class": "scan", + "calls_per_token": 0.99, + "share_pct": 0.11 + }, + { + "kernel": "void mlx::core::rocm::contiguous_scan<__half, __half, mlx::core::rocm::Sum, 4, true, false>(__half const*, __half*, int)", + "class": "scan", + "calls_per_token": 0.99, + "share_pct": 0.1 + } + ], + "class_share_pct": { + "qmv": 93.56, + "sdpa": 1.96, + "rms_norm": 1.23, + "copy": 0.75, + "binary": 0.71, + "rope": 0.51, + "sort": 0.39, + "compiled": 0.32, + "scan": 0.22, + "gather_scatter": 0.12, + "softmax": 0.09, + "ternary": 0.05, + "unary": 0.04, + "reduce": 0.03, + "other": 0.02, + "random": 0.01, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 96.0, + "sampler_tail": 1.17, + "rope_append": 1.03, + "add_rms_join": 0.9, + "add_rms_join_post_attn": 0.89, + "moe_expert_gemv": 0.0, + "moe_activation": 0.0, + "moe_weighted_sum": 0.0, + "moe_gather_indices": 0.0, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 229.2, + "rope_append": 127.5, + "add_rms_join_post_attn": 63.5, + "add_rms_join": 63.5, + "sampler_tail": 63.5, + "moe_expert_gemv": 0.0, + "moe_activation": 0.0, + "moe_weighted_sum": 0.0, + "moe_gather_indices": 0.0, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 1.92, + "fallback_dispatches_per_token": 191.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.89, + "note": "both paths ship off (FUSED_ADD_RMSNORM_DEFAULT / FUSED_ROPE_APPEND_DEFAULT false); only the post-attention join calls the fused norm; rope_scaling builds a frequency table, which the RoPE kernel cannot take" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 1.17, + "fallback_dispatches_per_token": 63.5, + "reached_default_share_pct": 1.17, + "reached_default_dispatches_per_token": 63.5, + "reached_optin_share_pct": 1.17, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "llama does not call forward_fused_kernel" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "no Mamba2 layer" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_bench.log new file mode 100644 index 000000000..d2ec33480 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_bench.log @@ -0,0 +1,16 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.157 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] warmup_start monotonic_ns=176852470403700 boottime_ns=176852470403740 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] measured_start monotonic_ns=176854979515399 boottime_ns=176854979515419 +[phase] decode_start monotonic_ns=176855486353560 boottime_ns=176855486353580 +[phase] measured_end monotonic_ns=176858719004524 boottime_ns=176858719004544 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 478.31 ms (1070.44 tok/s) + Decode: 3232.65 ms (39.60 tok/s) + MLX peak memory: 20.60 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_decode_kernels.csv new file mode 100644 index 000000000..44f0f876e --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_decode_kernels.csv @@ -0,0 +1,38 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)",qmv,,,20447,159.74,2838094833,138803,94.696 +"void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)",sdpa,,,4064,31.75,59003854,14519,1.969 +"void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,8255,64.49,34316389,4157,1.145 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn,2063,8128,63.5,16538053,2035,0.552 +"void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,rope_append,2063,8256,64.5,15124796,1832,0.505 +"void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)",rope,rope_append,2063,8128,63.5,13005851,1600,0.434 +"void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)",compiled,,,4064,31.75,9066642,2231,0.303 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,3404982,26811,0.114 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail,2064,127,0.99,816052,6426,0.027 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,602525,4744,0.020 +"void mlx::core::rocm::copy_v<__half, float, unsigned int, 4>(__half const*, float*, unsigned int)",copy,sampler_tail,2064,254,1.98,497509,1959,0.017 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,485315,1911,0.016 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,426444,3358,0.014 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,423072,1666,0.014 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,413842,1629,0.014 +"void mlx::core::rocm::gather_rows_kernel<__half, int>(__half const*, int const*, __half*, long, unsigned int, int)",gather_scatter,,,254,1.98,393617,1550,0.013 +"void mlx::core::rocm::copy_v<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",copy,sampler_tail,2064,127,0.99,385365,3034,0.013 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,336662,2651,0.011 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,311876,1228,0.010 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,276992,1091,0.009 +"void mlx::core::rocm::gather_axis_kernel<__half, int, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,127,0.99,267903,2109,0.009 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,265699,2092,0.009 +"void mlx::core::rocm::binary_vs(__half const*, __half const*, __half*, unsigned int)",binary,sampler_tail,2064,127,0.99,254672,2005,0.008 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,230030,1811,0.008 +"void mlx::core::rocm::scatter_axis_kernel<__half, int, false, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,227550,1792,0.008 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,208510,1642,0.007 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,185511,1461,0.006 +"void mlx::core::rocm::affine_dequantize_packed_kernel<__half, 4>(unsigned char const*, __half const*, __half const*, __half*, unsigned long, int)",dequantize,,,127,0.99,171156,1348,0.006 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,169674,1336,0.006 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,163906,1291,0.005 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,153249,1207,0.005 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,151853,1196,0.005 +"void mlx::core::rocm::copy_s<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",copy,,,64,0.5,140506,2195,0.005 +"void mlx::core::rocm::copy_v(float const*, __half*, unsigned int)",copy,sampler_tail,2064,127,0.99,140419,1106,0.005 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,136540,1075,0.005 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,131368,1034,0.004 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,130405,1027,0.004 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_kernel_stats.csv new file mode 100644 index 000000000..0d3ff3662 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_kernel_stats.csv @@ -0,0 +1,45 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)",23508,3258511471,138612.875234,73.13,58029,2252252,103423.935873 +"Cijk_Alik_Bljk_HHS_BH_Bias_HA_S_SAV_UserArgs_MT80x128x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT5_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR0_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG16_8_1_WGMXCC1",192,392570367,2044637.328125,8.81,1240886,2704791,486333.001609 +"Cijk_Alik_Bljk_HHS_BH_Bias_HA_S_SAV_UserArgs_MT64x128x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT4_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR0_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG16_8_1_WGMXCC1",128,237320750,1854068.359375,5.33,849072,3328921,855638.656923 +"void mlx::core::rocm::affine_dequantize_packed_kernel<__half, 4>(unsigned char const*, __half const*, __half const*, __half*, unsigned long, int)",468,236550377,505449.523504,5.31,1162,1962119,485793.245738 +"void mlx::core::rocm::kernel_sdpa_flash_wmma<__half, true, 128, 64, 64>(__half const*, __half const*, __half const*, __half*, float*, mlx::core::rocm::FAWmmaParams)",64,95079814,1485622.093750,2.13,1448356,1533034,18629.596051 +"void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)",4672,67199330,14383.418236,1.51,8175,31500,1447.241482 +"void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)",9620,46739725,4858.599272,1.05,3647,138580,8009.864878 +"void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)",9472,26354636,2782.372889,0.5915,1523,301525,6544.860245 +"void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",9792,23482207,2398.101205,0.5270,1001,176090,7785.133664 +"void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)",4736,19419893,4100.484164,0.4358,1683,244859,15806.780204 +"void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)",9344,15231463,1630.079516,0.3418,1082,10660,934.715642 +"void mlx::core::rocm::copy_gg_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,11960102,124584.395833,0.2684,26971,645810,139085.822829 +"void mlx::core::rocm::rope_freqs<__half, false, true>(__half const*, __half*, int const*, float const*, float, float, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, long, int, HIP_vector_type, long)",128,6480668,50630.218750,0.1454,9698,127760,38833.040938 +"void mlx::core::rocm::copy_g_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",64,4404698,68823.406250,0.0989,63439,310743,30826.417261 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,4091345,27644.222973,0.0918,26129,135414,8988.056053 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",148,986890,6668.175676,0.0221,6091,33302,2323.682858 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,708281,4785.682432,0.0159,4368,9538,606.327632 +"void mlx::core::rocm::copy_v<__half, float, unsigned int, 4>(__half const*, float*, unsigned int)",296,605267,2044.820946,0.0136,1042,8416,983.610781 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,563666,1904.277027,0.0127,1242,10219,1208.508134 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,523224,3535.297297,0.0117,2925,26530,2070.638642 +"void mlx::core::rocm::gather_rows_kernel<__half, int>(__half const*, int const*, __half*, long, unsigned int, int)",296,514444,1737.986486,0.0115,1041,21601,1746.541918 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,508073,1716.462838,0.0114,1122,9298,1096.541354 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,483813,1634.503378,0.0109,1082,6853,754.911520 +"void mlx::core::rocm::copy_v<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",148,457663,3092.317568,0.0103,2484,12744,1146.736772 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,397577,2686.331081,8.923e-03,2404,7614,721.865133 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,366540,1238.310811,8.226e-03,761,9177,1015.201173 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,327043,1104.875000,7.340e-03,721,9778,824.011622 +"void mlx::core::rocm::gather_axis_kernel<__half, int, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",148,318154,2149.689189,7.140e-03,1603,6652,1107.816697 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,310581,2098.520270,6.971e-03,1322,7374,1131.522576 +"void mlx::core::rocm::binary_vs(__half const*, __half const*, __half*, unsigned int)",148,298073,2014.006757,6.690e-03,1643,6573,849.985662 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,289865,1958.547297,6.506e-03,1282,19036,1650.360865 +"void mlx::core::rocm::scatter_axis_kernel<__half, int, false, 1, true, true>(__half const*, int const*, __half*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,281169,1899.790541,6.310e-03,1323,7213,1164.470662 +"void mlx::core::rocm::copy_s<__half, __half, unsigned int, 4>(__half const*, __half*, unsigned int)",128,267783,2092.054688,6.010e-03,1603,9979,1034.534300 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",148,254358,1718.635135,5.709e-03,1042,6212,1151.569548 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,224426,1516.391892,5.037e-03,1162,7453,1045.963727 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,203983,1378.263514,4.578e-03,722,8656,1130.154197 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,198731,1342.777027,4.460e-03,962,8897,1007.844780 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,185675,1254.560811,4.167e-03,961,5931,566.300661 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,185150,1251.013514,4.155e-03,761,7174,942.574337 +"void mlx::core::rocm::copy_v(float const*, __half*, unsigned int)",148,172883,1168.128378,3.880e-03,801,6091,786.298910 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",148,168399,1137.831081,3.779e-03,721,7294,755.983824 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,162665,1099.087838,3.651e-03,721,6693,624.467054 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,155174,1048.472973,3.483e-03,761,1684,175.894273 +"__amd_rocclr_fillBufferUnAligned",1,127158,127158.000000,2.854e-03,127158,127158,0.00000000e+00 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_summary.json new file mode 100644 index 000000000..2479a919f --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Meta-Llama-3.1-8B-Instruct-4bit_t0.7_summary.json @@ -0,0 +1,180 @@ +{ + "name": "Meta-Llama-3.1-8B-Instruct-4bit_t0.7", + "model_type": "llama", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 39.6, + "profiled_prefill_tok_s": 1070.44, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 65978, + "dispatches_per_token": 515.5, + "decode_wall_ms": 3232.651, + "decode_gpu_sum_ms": 2997.054, + "decode_gpu_busy_ms": 2997.054, + "gpu_ms_per_token": 23.4145, + "host_gap_ms_per_token": 1.8406, + "host_gap_pct_of_wall": 7.29, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 29615.3, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel<__half, __half, 64, 4>(__half const*, unsigned char const*, __half const*, __half const*, __half*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 159.74, + "share_pct": 94.7 + }, + { + "kernel": "void mlx::core::rocm::kernel_sdpav_1pass<__half, false, 128>(__half const*, __half const*, __half const*, __half*, __half const*, mlx::core::rocm::AttnParams)", + "class": "sdpa", + "calls_per_token": 31.75, + "share_pct": 1.97 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel<__half, 256, 4>(__half const*, __half const*, __half*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 64.49, + "share_pct": 1.15 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(__half const*, __half const*, __half*, unsigned int)", + "class": "binary", + "calls_per_token": 63.5, + "share_pct": 0.55 + }, + { + "kernel": "void mlx::core::rocm::copy_gg_byval<__half, __half, long>(__half const*, __half*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "copy", + "calls_per_token": 64.5, + "share_pct": 0.5 + }, + { + "kernel": "void mlx::core::rocm::rope_single_freqs_1d<__half, false, true>(__half const*, __half*, int const*, float const*, float, long, unsigned int, unsigned int, long)", + "class": "rope", + "calls_per_token": 63.5, + "share_pct": 0.43 + }, + { + "kernel": "void mlx::core::rocm::Cf2ISigmoidADf2IBroadcastACEf2IBroadcastCAFf2IMultiplyDEGf2IBroadcastFBHf2IBroadcastBFIf2OMultiplyGH_VV_f2f2_6142509188972423790_contiguous(__half const*, __half const*, __half*, unsigned int)", + "class": "compiled", + "calls_per_token": 31.75, + "share_pct": 0.3 + }, + { + "kernel": "void mlx::core::rocm::contiguous_scan(float const*, float*, int)", + "class": "scan", + "calls_per_token": 0.99, + "share_pct": 0.11 + }, + { + "kernel": "void mlx::core::rocm::unary_v(float const*, float*, unsigned int)", + "class": "unary", + "calls_per_token": 0.99, + "share_pct": 0.03 + }, + { + "kernel": "void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "ternary", + "calls_per_token": 0.99, + "share_pct": 0.02 + } + ], + "class_share_pct": { + "qmv": 94.7, + "sdpa": 1.97, + "rms_norm": 1.15, + "binary": 0.64, + "copy": 0.55, + "rope": 0.43, + "compiled": 0.3, + "scan": 0.11, + "gather_scatter": 0.04, + "reduce": 0.03, + "unary": 0.03, + "ternary": 0.02, + "random": 0.02, + "sort": 0.01, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 97.13, + "add_rms_join": 0.85, + "add_rms_join_post_attn": 0.83, + "rope_append": 0.82, + "sampler_tail": 0.37, + "moe_expert_gemv": 0.0, + "moe_activation": 0.0, + "moe_weighted_sum": 0.0, + "moe_gather_indices": 0.0, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 229.2, + "rope_append": 127.5, + "add_rms_join_post_attn": 63.5, + "add_rms_join": 63.5, + "sampler_tail": 31.8, + "moe_expert_gemv": 0.0, + "moe_activation": 0.0, + "moe_weighted_sum": 0.0, + "moe_gather_indices": 0.0, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 1.65, + "fallback_dispatches_per_token": 191.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.83, + "note": "both paths ship off (FUSED_ADD_RMSNORM_DEFAULT / FUSED_ROPE_APPEND_DEFAULT false); only the post-attention join calls the fused norm; rope_scaling builds a frequency table, which the RoPE kernel cannot take" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.37, + "fallback_dispatches_per_token": 31.8, + "reached_default_share_pct": 0.37, + "reached_default_dispatches_per_token": 31.8, + "reached_optin_share_pct": 0.37, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "llama does not call forward_fused_kernel" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "no Mamba2 layer" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_bench.log new file mode 100644 index 000000000..7fb59088f --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_bench.log @@ -0,0 +1,21 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[NemotronH] Loading config... +[NemotronH] Config loaded: 52 layers (23 mamba, 6 attention, 0 mlp, 23 moe) +[NemotronH] Loading weights from safetensors... +[NemotronH] Building model... +[NemotronH] Model loaded successfully +[Load] wall: 0.285 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[phase] warmup_start monotonic_ns=175856462625683 boottime_ns=175856462625713 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[phase] measured_start monotonic_ns=175867172119766 boottime_ns=175867172119806 +[phase] decode_start monotonic_ns=175869076904088 boottime_ns=175869076904108 +[phase] measured_end monotonic_ns=175872047916422 boottime_ns=175872047916442 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1902.81 ms (269.08 tok/s) + Decode: 2971.01 ms (43.08 tok/s) + MLX peak memory: 18.48 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_decode_kernels.csv new file mode 100644 index 000000000..12bf85a6c --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_decode_kernels.csv @@ -0,0 +1,51 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,14859,116.09,705875073,47505,38.627 +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,5842,45.64,509775289,87260,27.896 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,sampler_tail;ssm_step,2064;2067,11811,92.27,74872146,6339,4.097 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,ssm_step,2067,14605,114.1,61945074,4241,3.390 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",gemm,ssm_step,2067,5842,45.64,39744904,6803,2.175 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail;ssm_gated_norm;ssm_step,2064;2065;2067,29337,229.2,38353113,1307,2.099 +"void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)",sort,,,2921,22.82,38065091,13032,2.083 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,6731,52.59,30069178,4467,1.645 +"void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)",gemm,,,2921,22.82,27566013,9437,1.508 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn;ssm_step,2063;2067,12446,97.23,23473520,1886,1.285 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_gated_norm;ssm_step,2067,11684,91.28,20702580,1772,1.133 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_activation;sampler_tail;ssm_gated_norm;ssm_step,2064;2065;2067,11811,92.27,16525718,1399,0.904 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",copy,ssm_step,2067,11684,91.28,15929240,1363,0.872 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,8763,68.46,15259237,1741,0.835 +Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,2921,22.82,14347885,4912,0.785 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,ssm_step,2067,8763,68.46,14105062,1610,0.772 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_conv,,7390,57.73,13267553,1795,0.726 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,ssm_step,2067,5842,45.64,12922369,2212,0.707 +"void mlx::core::rocm::CV2IAsTypeBDV2IBroadcastACEV2IBroadcastCAFV2IMaximumDEGV2OSquareF_VC_V2_2297668033614959926_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,5842,45.64,12914882,2211,0.707 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",copy,ssm_step,2067,11684,91.28,12646742,1082,0.692 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,ssm_silu,,2921,22.82,10401063,3561,0.569 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,762,5.95,9195034,12067,0.503 +Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,2921,22.82,8790725,3009,0.481 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",binary,ssm_step,2067,5842,45.64,8001995,1370,0.438 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,,,2921,22.82,7106430,2433,0.389 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",conv,ssm_conv,,2921,22.82,6918827,2369,0.379 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,6893748,2360,0.377 +"void mlx::core::rocm::rms_norm_kernel(float const*, float const*, float*, float, unsigned int, long)",rms_norm,ssm_gated_norm,,2921,22.82,5945517,2035,0.325 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,ssm_step,2067,2921,22.82,5736256,1964,0.314 +"void mlx::core::rocm::Bf4ISigmoidACf4IBroadcastABDf4IBroadcastBAEf4OMultiplyCD_V_f4_6142509188972423790_contiguous(float const*, float*, unsigned int)",compiled,ssm_silu,,2921,22.82,5713785,1956,0.313 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,ssm_step,2067,2921,22.82,5685716,1946,0.311 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,,,2921,22.82,5368756,1838,0.294 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,5842,45.64,5205451,891,0.285 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,5168207,1769,0.283 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",reduce,,,2921,22.82,5145407,1762,0.282 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",scan,ssm_step,2067,2921,22.82,5030006,1722,0.275 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,2921,22.82,5004572,1713,0.274 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<1, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",compiled,,,2921,22.82,4962481,1699,0.272 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,4947631,1694,0.271 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<1, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",compiled,,,2921,22.82,4315396,1477,0.236 +"void mlx::core::rocm::arg_reduce_partial, 256>(hip_bfloat16 const*, hip_bfloat16*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",arg_reduce,sampler_tail,2064,127,0.99,1250258,9845,0.068 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,426357,1679,0.023 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,400754,3156,0.022 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,127,0.99,388570,3060,0.021 +"void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",arg_reduce,sampler_tail,2064,127,0.99,304208,2395,0.017 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,246818,1943,0.014 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,243900,1920,0.013 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,243050,1914,0.013 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,,,12,0.09,15667,1306,0.001 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,2,0.02,4248,2124,0.000 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_kernel_stats.csv new file mode 100644 index 000000000..f51ec38af --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_kernel_stats.csv @@ -0,0 +1,63 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",6808,3502604822,514483.669506,59.37,75581,116550103,3849969.349749 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",17084,814224030,47660.034535,13.80,3486,1483680,78278.064441 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",232,408753437,1761868.262931,6.93,174367,4364979,942964.209609 +"Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",3404,114886039,33750.305229,1.95,2124,2556511,262989.170446 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",17020,86859762,5103.393772,1.47,1002,261971,10177.431337 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",13580,86226461,6349.518483,1.46,801,94457,8350.827997 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",13524,72523436,5362.572907,1.23,841,1043275,57375.012453 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",6808,63156313,9276.779230,1.07,1242,808916,66680.328458 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",10166,63146482,6211.536691,1.07,1202,522820,46749.761642 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",33958,51719626,1523.046881,0.8767,682,109726,3988.539583 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",6716,45756941,6813.124032,0.7756,1323,21360,5082.820204 +"void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)",3404,44570506,13093.568155,0.7555,10219,68689,1663.953522 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",3404,44019984,12931.840188,0.7461,1242,1001516,95885.056525 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",10166,43969353,4325.138009,0.7453,881,701755,40106.841798 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",7844,38155609,4864.305074,0.6467,3887,137256,4902.848376 +"void mlx::core::rocm::kernel_sdpa_flash_wmma(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float*, mlx::core::rocm::FAWmmaParams)",12,33100063,2758338.583333,0.5610,1539725,15672246,4066888.563023 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",14458,32793557,2268.194564,0.5559,1122,47208,3743.263493 +"void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)",3358,31804013,9471.117630,0.5391,8336,27572,837.111241 +"Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",3404,27624808,8115.396005,0.4682,4368,270947,27295.038466 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",13764,23911372,1737.240046,0.4053,841,105037,4544.591044 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",3404,20662416,6070.039953,0.3502,2003,292187,30955.771773 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",8608,19611823,2278.325163,0.3324,961,377187,6943.214856 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",13616,18722336,1375.024677,0.3173,721,10379,1055.830269 +"void mlx::core::rocm::CV2IAsTypeBDV2IBroadcastACEV2IBroadcastCAFV2IMaximumDEGV2OSquareF_VC_V2_2297668033614959926_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",6808,18569974,2727.669506,0.3148,1763,63559,4691.548683 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",3404,16501047,4847.546122,0.2797,1483,244739,26740.654597 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",13616,14610997,1073.075573,0.2477,641,10380,1148.594504 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",3404,14490924,4257.028202,0.2456,1683,316994,16260.965858 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",3404,13586906,3991.452996,0.2303,2765,149800,4134.145505 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",876,10470905,11953.087900,0.1775,9377,45365,2007.634763 +"Cijk_Alik_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",46,9416842,204713.956522,0.1596,196167,483185,42006.761696 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",6716,9246394,1376.770995,0.1567,761,10700,1421.198724 +"void mlx::core::rocm::rms_norm_kernel(float const*, float const*, float*, float, unsigned int, long)",3404,8833620,2595.070505,0.1497,1523,192360,5447.592959 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",3404,8158610,2396.771445,0.1383,1923,10940,1110.417206 +"void mlx::core::rocm::Bf4ISigmoidACf4IBroadcastABDf4IBroadcastBAEf4OMultiplyCD_V_f4_6142509188972423790_contiguous(float const*, float*, unsigned int)",3404,7932649,2330.390423,0.1345,1322,109165,3461.257818 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",230,7142755,31055.456522,0.1211,1803,237645,28761.805021 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",3404,6715053,1972.694771,0.1138,1322,11020,1663.968804 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",3358,6686903,1991.335021,0.1133,1242,11221,1381.393631 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",3358,6232894,1856.132817,0.1056,1202,7574,396.851074 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",6808,6141092,902.040541,0.1041,601,6412,252.164259 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",3404,6091878,1789.623384,0.1033,1442,12744,404.720004 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",3404,6087635,1788.376910,0.1032,1202,15549,1030.495839 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",3404,5770768,1695.290247,0.0978,1042,10740,1603.713343 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<1, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",3358,5764069,1716.518463,0.0977,1402,12784,332.570414 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",92,5186398,56373.891304,0.0879,1803,116418,53839.465105 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<1, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",3358,4986446,1484.945205,0.0845,801,6251,365.628634 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",46,4041819,87865.630435,0.0685,83437,247344,24090.806915 +"void mlx::core::rocm::arg_reduce_partial, 256>(hip_bfloat16 const*, hip_bfloat16*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",148,1443499,9753.371622,0.0245,8656,14507,1665.801914 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",58,1374216,23693.379310,0.0233,1242,312746,60071.456808 +"Cijk_Alik_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x64x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT2_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR1_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG32_4_1_WGMXCC1",46,1223576,26599.478261,0.0207,22282,85881,9127.465881 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",92,794333,8634.054348,0.0135,7734,16230,1666.464108 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",296,538926,1820.695946,9.135e-03,1282,26530,1604.167414 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,467559,3159.182432,7.925e-03,2484,7974,1158.139976 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",148,451209,3048.709459,7.648e-03,2043,8095,1707.491936 +"void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",148,359791,2431.020270,6.098e-03,1843,10059,1413.167083 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,330693,2234.412162,5.605e-03,1763,30537,2540.172786 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,314152,2122.648649,5.325e-03,1683,24206,1884.688541 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,289975,1959.290541,4.915e-03,1482,6532,1140.268170 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<2, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",46,188911,4106.760870,3.202e-03,2965,15308,1725.419024 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",46,147156,3199.043478,2.494e-03,1923,14027,2467.995254 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",70,119183,1702.614286,2.020e-03,961,6332,976.514673 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<2, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",46,110163,2394.847826,1.867e-03,1282,10099,1738.833338 +"__amd_rocclr_fillBufferUnAligned",1,52418,52418.000000,8.885e-04,52418,52418,0.00000000e+00 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_plain_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_plain_bench.log new file mode 100644 index 000000000..075335cec --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_plain_bench.log @@ -0,0 +1,17 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[NemotronH] Loading config... +[NemotronH] Config loaded: 52 layers (23 mamba, 6 attention, 0 mlp, 23 moe) +[NemotronH] Loading weights from safetensors... +[NemotronH] Building model... +[NemotronH] Model loaded successfully +[Load] wall: 0.325 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1840.13 ms (278.24 tok/s) + Decode: 2473.73 ms (51.74 tok/s) + MLX peak memory: 18.48 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_summary.json new file mode 100644 index 000000000..81aa09b23 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy_summary.json @@ -0,0 +1,188 @@ +{ + "name": "NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy", + "model_type": "nemotron_h", + "temperature": 0.0, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 43.08, + "profiled_prefill_tok_s": 269.08, + "plain_decode_tok_s": 51.74, + "plain_prefill_tok_s": 278.24, + "profiler_decode_slowdown_pct": 20.1, + "decode_dispatches": 256959, + "dispatches_per_token": 2007.5, + "decode_wall_ms": 2971.012, + "decode_gpu_sum_ms": 1827.422, + "decode_gpu_busy_ms": 1827.384, + "gpu_ms_per_token": 14.2764, + "host_gap_ms_per_token": 8.9346, + "host_gap_pct_of_wall": 38.49, + "plain_wall_ms_per_token": 19.3274, + "plain_host_gap_ms_per_token_est": 5.051, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 2095.1, + "last_dispatch_before_decode": "void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, un" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 116.09, + "share_pct": 38.63 + }, + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 45.64, + "share_pct": 27.9 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)", + "class": "binary", + "calls_per_token": 92.27, + "share_pct": 4.1 + }, + { + "kernel": "void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "binary", + "calls_per_token": 114.1, + "share_pct": 3.39 + }, + { + "kernel": "void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)", + "class": "gemm", + "calls_per_token": 45.64, + "share_pct": 2.17 + }, + { + "kernel": "void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)", + "class": "copy", + "calls_per_token": 229.2, + "share_pct": 2.1 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 22.82, + "share_pct": 2.08 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 52.59, + "share_pct": 1.65 + }, + { + "kernel": "void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)", + "class": "gemm", + "calls_per_token": 22.82, + "share_pct": 1.51 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 97.23, + "share_pct": 1.28 + } + ], + "class_share_pct": { + "qmv": 38.63, + "gather_qmv": 27.9, + "binary": 11.55, + "copy": 6.44, + "gemm": 4.95, + "compiled": 2.1, + "sort": 2.08, + "rms_norm": 1.97, + "unary": 1.09, + "ternary": 0.71, + "reduce": 0.67, + "scan": 0.59, + "sdpa": 0.5, + "conv": 0.38, + "gather_scatter": 0.37, + "arg_reduce": 0.09, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 45.26, + "moe_expert_gemv": 27.9, + "ssm_step": 20.01, + "add_rms_join": 2.07, + "ssm_gated_norm": 1.14, + "ssm_conv": 0.98, + "ssm_silu": 0.88, + "moe_activation": 0.61, + "moe_weighted_sum": 0.43, + "moe_gather_indices": 0.28, + "add_rms_join_post_attn": 0.26, + "sampler_tail": 0.17, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 367.4, + "ssm_step": 1141.0, + "ssm_conv": 68.5, + "ssm_silu": 45.6, + "ssm_gated_norm": 91.2, + "add_rms_join": 91.3, + "moe_gather_indices": 45.6, + "moe_expert_gemv": 45.6, + "moe_activation": 45.6, + "moe_weighted_sum": 45.6, + "add_rms_join_post_attn": 11.9, + "sampler_tail": 7.9, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 0.26, + "fallback_dispatches_per_token": 11.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "nemotron_h never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.17, + "fallback_dispatches_per_token": 7.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "greedy argmax dispatches neither sampler kernel" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 29.22, + "fallback_dispatches_per_token": 182.5, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "fused_moe_forward's default branch is gather_qmm; the kernel path needs MLXCEL_FUSED_MOE_RELU2 and moe_fc1_relu2 (#2069)" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 20.01, + "fallback_dispatches_per_token": 1141.0, + "reached_default_share_pct": 20.01, + "reached_default_dispatches_per_token": 1141.0, + "reached_optin_share_pct": 20.01, + "note": "ssm_kernel_available() gates the decode step" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_bench.log new file mode 100644 index 000000000..d530fa798 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_bench.log @@ -0,0 +1,21 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[NemotronH] Loading config... +[NemotronH] Config loaded: 52 layers (23 mamba, 6 attention, 0 mlp, 23 moe) +[NemotronH] Loading weights from safetensors... +[NemotronH] Building model... +[NemotronH] Model loaded successfully +[Load] wall: 0.279 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[phase] warmup_start monotonic_ns=178921979565537 boottime_ns=178921979565577 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[phase] measured_start monotonic_ns=178932846034231 boottime_ns=178932846034251 +[phase] decode_start monotonic_ns=178934893741096 boottime_ns=178934893741166 +[phase] measured_end monotonic_ns=178938032938240 boottime_ns=178938032938310 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 2044.81 ms (250.39 tok/s) + Decode: 3139.20 ms (40.77 tok/s) + MLX peak memory: 18.48 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv new file mode 100644 index 000000000..d779799f1 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv @@ -0,0 +1,82 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,14859,116.09,707890371,47641,37.136 +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,5842,45.64,471160504,80651,24.717 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,sampler_tail;ssm_step,2064;2067,11811,92.27,75297869,6375,3.950 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,ssm_step,2067,14605,114.1,62009027,4246,3.253 +Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,2921,22.82,51377795,17589,2.695 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",gemm,ssm_step,2067,5842,45.64,39754254,6805,2.086 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail;ssm_gated_norm;ssm_step,2064;2065;2067,29464,230.19,39584088,1343,2.077 +"void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)",sort,,,2921,22.82,38262523,13099,2.007 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",sort,sampler_tail,2064,889,6.95,36856786,41459,1.934 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,6731,52.59,30283999,4499,1.589 +"void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)",gemm,,,2921,22.82,27878605,9544,1.463 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn;ssm_step,2063;2067,12446,97.23,23618263,1898,1.239 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_gated_norm;ssm_step,2067,11684,91.28,20577724,1761,1.080 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_activation;sampler_tail;ssm_gated_norm;ssm_step,2064;2065;2067,11811,92.27,16676159,1412,0.875 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,8763,68.46,15797982,1803,0.829 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",copy,ssm_step,2067,11684,91.28,15471437,1324,0.812 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail;ssm_step,2064;2067,8890,69.45,14677178,1651,0.770 +Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,2921,22.82,14447594,4946,0.758 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_conv,,7390,57.73,13551110,1834,0.711 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail;ssm_step,2064;2067,5969,46.63,13464109,2256,0.706 +"void mlx::core::rocm::CV2IAsTypeBDV2IBroadcastACEV2IBroadcastCAFV2IMaximumDEGV2OSquareF_VC_V2_2297668033614959926_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,5842,45.64,13053385,2234,0.685 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",copy,ssm_step,2067,11684,91.28,12403191,1062,0.651 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,ssm_silu,,2921,22.82,10660035,3649,0.559 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,762,5.95,9157718,12018,0.480 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",binary,ssm_step,2067,5842,45.64,8253920,1413,0.433 +"void mlx::core::rocm::contiguous_scan(hip_bfloat16 const*, hip_bfloat16*, int)",scan,sampler_tail,2064,127,0.99,7761022,61110,0.407 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,,,2921,22.82,7358989,2519,0.386 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,6951042,2380,0.365 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",conv,ssm_conv,,2921,22.82,6927475,2372,0.363 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,ssm_step,2067,2921,22.82,5967604,2043,0.313 +"void mlx::core::rocm::rms_norm_kernel(float const*, float const*, float*, float, unsigned int, long)",rms_norm,ssm_gated_norm,,2921,22.82,5959450,2040,0.313 +"void mlx::core::rocm::Bf4ISigmoidACf4IBroadcastABDf4IBroadcastBAEf4OMultiplyCD_V_f4_6142509188972423790_contiguous(float const*, float*, unsigned int)",compiled,ssm_silu,,2921,22.82,5851919,2003,0.307 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,ssm_step,2067,2921,22.82,5580198,1910,0.293 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",softmax,sampler_tail,2064,127,0.99,5333159,41993,0.280 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,,,2921,22.82,5283140,1809,0.277 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",reduce,,,2921,22.82,5282756,1809,0.277 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<1, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",compiled,,,2921,22.82,5263397,1802,0.276 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,5842,45.64,5213460,892,0.273 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,2921,22.82,5092327,1743,0.267 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,4976710,1704,0.261 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",scan,ssm_step,2067,2921,22.82,4964946,1700,0.260 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,4622515,1583,0.242 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<1, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",compiled,,,2921,22.82,4050478,1387,0.212 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,3453229,27191,0.181 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,508,3.97,2767633,5448,0.145 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,2004673,15785,0.105 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",sort,sampler_tail,2064,254,1.98,1831704,7211,0.096 +__amd_rocclr_copyBuffer,copy,sampler_tail,2064,508,3.97,1136326,2237,0.060 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,127,0.99,705121,5552,0.037 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,127,0.99,649054,5111,0.034 +__amd_rocclr_fillBufferUnAligned,copy,sampler_tail,2064,381,2.98,608669,1598,0.032 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,577680,2274,0.030 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",other,sampler_tail,2064,254,1.98,510717,2011,0.027 +"void mlx::core::rocm::ternary_g(bool const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,494525,3894,0.026 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::default_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1})",ternary,sampler_tail,2064,254,1.98,488541,1923,0.026 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,450810,1775,0.024 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,431170,3395,0.023 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,422353,3326,0.022 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,398946,1571,0.021 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,391375,1541,0.021 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,368809,2904,0.019 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,368601,2902,0.019 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,sampler_tail,2064,127,0.99,337235,2655,0.018 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,320810,1263,0.017 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,282882,1114,0.015 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,263335,2074,0.014 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,244816,1928,0.013 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,237604,1871,0.012 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,225746,1778,0.012 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,225667,1777,0.012 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",sort,sampler_tail,2064,127,0.99,198488,1563,0.010 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,177181,1395,0.009 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,153008,1205,0.008 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,151401,1192,0.008 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,146234,1151,0.008 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,144950,1141,0.008 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,142104,1119,0.007 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,139218,1096,0.007 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,136457,1074,0.007 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,,,12,0.09,16150,1346,0.001 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,2,0.02,4289,2144,0.000 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv new file mode 100644 index 000000000..4738c80e1 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv @@ -0,0 +1,93 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",6808,3382128431,496787.372356,57.39,75621,44428902,3570349.536722 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",17084,816805965,47811.166296,13.86,3526,1969589,78902.101175 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",232,410033363,1767385.185345,6.96,174327,4705763,963148.797511 +"Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",3404,159527837,46864.816980,2.71,1643,43098510,785514.348424 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",17020,87132433,5119.414395,1.48,1041,261569,10205.242897 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",13580,86655122,6381.084094,1.47,801,95178,8338.812819 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",13524,72075415,5329.445061,1.22,801,1030891,57330.930029 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",6956,64090779,9213.740512,1.09,1202,858408,66218.054298 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",10166,63694823,6265.475408,1.08,1202,524944,46711.206492 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",34106,52519698,1539.896147,0.8912,682,108443,3949.546207 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",6716,45733519,6809.636540,0.7761,1323,30497,5048.699896 +"void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)",3404,44736455,13142.319330,0.7591,10740,68528,1423.295607 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",10314,44593146,4323.554974,0.7567,881,614431,39765.437285 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",3404,43893383,12894.648355,0.7448,1242,1033175,95706.310827 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",7844,38429549,4899.228582,0.6521,3927,134451,5066.998246 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",1036,38164235,36838.064672,0.6476,6973,28997190,900624.480008 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",14458,32889019,2274.797275,0.5581,1162,46687,3750.483052 +"void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)",3358,32070042,9550.340083,0.5442,8295,27932,728.329404 +"Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",3404,27533695,8088.629553,0.4672,4288,264896,26915.480816 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",13764,24051868,1747.447544,0.4081,841,113613,4577.988416 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",3404,20956807,6156.523796,0.3556,2004,290023,31052.204101 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",8608,19746763,2294.001278,0.3351,961,377146,6940.179214 +"void mlx::core::rocm::kernel_sdpa_flash_wmma(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float*, mlx::core::rocm::FAWmmaParams)",12,19024671,1585389.250000,0.3228,1543771,1623040,24494.969677 +"void mlx::core::rocm::CV2IAsTypeBDV2IBroadcastACEV2IBroadcastCAFV2IMaximumDEGV2OSquareF_VC_V2_2297668033614959926_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",6808,18796549,2760.950206,0.3190,1763,171161,5088.384800 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",13616,18194368,1336.249119,0.3087,721,10340,957.253982 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",3404,16687454,4902.307286,0.2832,1522,252792,26990.478333 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",3404,14767288,4338.216216,0.2506,1683,646170,18865.557132 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",13616,14547509,1068.412823,0.2469,601,10339,1103.398721 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",3404,13787520,4050.387779,0.2340,2805,153207,4188.662286 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",876,10439918,11917.714612,0.1772,9217,45125,1296.359851 +"Cijk_Alik_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",46,9842553,213968.543478,0.1670,195526,913792,105500.808309 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",6716,9534320,1419.642644,0.1618,801,10780,1530.951555 +"void mlx::core::rocm::contiguous_scan(hip_bfloat16 const*, hip_bfloat16*, int)",148,9279308,62698.027027,0.1575,60474,295794,19294.196985 +"void mlx::core::rocm::rms_norm_kernel(float const*, float const*, float*, float, unsigned int, long)",3404,8667808,2546.359577,0.1471,1523,42680,4388.095753 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",3404,8213297,2412.836957,0.1394,1923,11261,1180.650958 +"void mlx::core::rocm::Bf4ISigmoidACf4IBroadcastABDf4IBroadcastBAEf4OMultiplyCD_V_f4_6142509188972423790_contiguous(float const*, float*, unsigned int)",3404,7969429,2341.195358,0.1352,1362,29254,2957.125472 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",230,7105303,30892.621739,0.1206,1803,130643,26182.975006 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",3358,6849383,2039.720965,0.1162,1242,10580,1287.591447 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",3404,6512342,1913.143948,0.1105,1122,10980,1476.941666 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",148,6220745,42032.060811,0.1056,41037,46728,1247.926003 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",3404,6182961,1816.381022,0.1049,1442,9859,320.117550 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",6808,6138880,901.715629,0.1042,601,13985,269.837074 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",3358,6075113,1809.146218,0.1031,1203,7334,271.418225 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<1, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",3358,6038737,1798.313580,0.1025,1202,8015,304.837099 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",3404,5937703,1744.331081,0.1008,1202,15068,889.957665 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",3404,5400728,1586.582844,0.0916,1042,10540,1426.192335 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",92,5246706,57029.413043,0.0890,1803,130564,54549.686355 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<1, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",3358,4678454,1393.226325,0.0794,801,5010,234.566440 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,4045298,27333.094595,0.0686,26730,48090,1732.272085 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",46,3877429,84291.934783,0.0658,83316,88526,1636.177108 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",592,3267452,5519.344595,0.0554,2043,29055,2239.059097 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,2338017,15797.412162,0.0397,15429,17593,218.992589 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",296,2175148,7348.472973,0.0369,5290,32181,2542.632342 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",58,1392042,24000.724138,0.0236,1322,313306,60061.162927 +"__amd_rocclr_copyBuffer",592,1341911,2266.741554,0.0228,1322,11180,1571.818803 +"Cijk_Alik_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x64x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT2_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR1_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG32_4_1_WGMXCC1",46,1224451,26618.500000,0.0208,22321,85199,8971.361757 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,839410,5671.689189,0.0142,5129,24205,1543.143505 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",92,826307,8981.597826,0.0140,7774,40837,3696.862341 +"__amd_rocclr_fillBufferUnAligned",445,803510,1805.640449,0.0136,721,89126,4462.788321 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,755853,5107.114865,0.0128,4728,5571,175.123767 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",296,708763,2394.469595,0.0120,1283,26409,1583.686150 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",296,598847,2023.131757,0.0102,1402,10861,918.394690 +"void mlx::core::rocm::ternary_g(bool const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,576640,3896.216216,9.785e-03,3446,5130,290.941599 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::default_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1})",296,569252,1923.148649,9.660e-03,1442,7655,1033.590680 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,540858,1827.222973,9.178e-03,1082,13184,901.644820 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,505995,3418.885135,8.586e-03,2565,8015,1441.394712 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,492843,3330.020270,8.363e-03,3125,4128,136.206300 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,476089,1608.408784,8.079e-03,1042,9097,646.322094 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,466435,1575.793919,7.915e-03,1122,9338,581.965671 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,443461,2996.358108,7.525e-03,2524,11301,1271.305489 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,441104,2980.432432,7.485e-03,2645,14187,955.367345 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,382561,2584.871622,6.492e-03,1883,7414,1318.127115 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,379240,1281.216216,6.435e-03,841,5130,419.528146 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,335548,2267.216216,5.694e-03,1723,27010,2072.264437 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,334100,1128.716216,5.669e-03,762,4208,282.319342 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,330253,2231.439189,5.604e-03,1603,29775,2496.796517 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,281566,1902.472973,4.778e-03,1482,10620,1205.942945 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,269787,1822.885135,4.578e-03,1603,6973,465.921346 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, bool*, unsigned int)",148,267624,1808.270270,4.541e-03,1282,7534,625.544460 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",148,231751,1565.885135,3.933e-03,1122,6171,1196.853684 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,209562,1415.959459,3.556e-03,1202,5490,413.066727 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<2, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",46,190594,4143.347826,3.234e-03,3326,15629,1758.579063 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,179534,1213.067568,3.047e-03,921,4769,427.946762 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,178896,1208.756757,3.036e-03,962,2765,345.116004 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,172882,1168.121622,2.934e-03,921,4208,316.377023 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,172240,1163.783784,2.923e-03,961,4208,292.780258 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",148,168471,1138.317568,2.859e-03,961,3928,253.204900 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,166307,1123.695946,2.822e-03,922,3927,281.676484 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,163749,1106.412162,2.779e-03,882,5651,400.924644 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",46,129605,2817.500000,2.199e-03,2084,7373,1281.289867 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",70,124957,1785.100000,2.120e-03,1002,6332,1116.075242 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<2, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",46,98461,2140.456522,1.671e-03,1282,8215,1285.901339 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_summary.json new file mode 100644 index 000000000..b7fc40ab3 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95_summary.json @@ -0,0 +1,190 @@ +{ + "name": "NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95", + "model_type": "nemotron_h", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 40.77, + "profiled_prefill_tok_s": 250.39, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 264071, + "dispatches_per_token": 2063.1, + "decode_wall_ms": 3139.197, + "decode_gpu_sum_ms": 1906.214, + "decode_gpu_busy_ms": 1906.214, + "gpu_ms_per_token": 14.8923, + "host_gap_ms_per_token": 9.6327, + "host_gap_pct_of_wall": 39.28, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 3009.4, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 116.09, + "share_pct": 37.14 + }, + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 45.64, + "share_pct": 24.72 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)", + "class": "binary", + "calls_per_token": 92.27, + "share_pct": 3.95 + }, + { + "kernel": "void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "binary", + "calls_per_token": 114.1, + "share_pct": 3.25 + }, + { + "kernel": "Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8", + "class": "gemm", + "calls_per_token": 22.82, + "share_pct": 2.7 + }, + { + "kernel": "void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)", + "class": "gemm", + "calls_per_token": 45.64, + "share_pct": 2.09 + }, + { + "kernel": "void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)", + "class": "copy", + "calls_per_token": 230.19, + "share_pct": 2.08 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 22.82, + "share_pct": 2.01 + }, + { + "kernel": "void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})", + "class": "sort", + "calls_per_token": 6.95, + "share_pct": 1.93 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 52.59, + "share_pct": 1.59 + } + ], + "class_share_pct": { + "qmv": 37.14, + "gather_qmv": 24.72, + "binary": 11.28, + "gemm": 7.0, + "copy": 6.36, + "sort": 4.19, + "compiled": 2.04, + "rms_norm": 1.9, + "scan": 1.16, + "unary": 1.08, + "ternary": 0.76, + "reduce": 0.71, + "sdpa": 0.48, + "gather_scatter": 0.48, + "conv": 0.36, + "softmax": 0.28, + "other": 0.03, + "random": 0.02, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 43.56, + "moe_expert_gemv": 24.72, + "ssm_step": 21.5, + "sampler_tail": 3.81, + "add_rms_join": 2.0, + "ssm_gated_norm": 1.07, + "ssm_conv": 0.95, + "ssm_silu": 0.87, + "moe_activation": 0.59, + "moe_weighted_sum": 0.41, + "moe_gather_indices": 0.27, + "add_rms_join_post_attn": 0.25, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 367.4, + "ssm_step": 1141.0, + "ssm_conv": 68.5, + "ssm_silu": 45.6, + "ssm_gated_norm": 91.3, + "add_rms_join": 91.3, + "moe_gather_indices": 45.6, + "moe_expert_gemv": 45.6, + "moe_activation": 45.6, + "moe_weighted_sum": 45.6, + "add_rms_join_post_attn": 11.9, + "sampler_tail": 63.5, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 0.25, + "fallback_dispatches_per_token": 11.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "nemotron_h never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 3.81, + "fallback_dispatches_per_token": 63.5, + "reached_default_share_pct": 3.81, + "reached_default_dispatches_per_token": 63.5, + "reached_optin_share_pct": 3.81, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 25.99, + "fallback_dispatches_per_token": 182.6, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "fused_moe_forward's default branch is gather_qmm; the kernel path needs MLXCEL_FUSED_MOE_RELU2 and moe_fc1_relu2 (#2069)" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 21.5, + "fallback_dispatches_per_token": 1141.0, + "reached_default_share_pct": 21.5, + "reached_default_dispatches_per_token": 1141.0, + "reached_optin_share_pct": 21.5, + "note": "ssm_kernel_available() gates the decode step" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_bench.log new file mode 100644 index 000000000..a8063069c --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_bench.log @@ -0,0 +1,21 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[NemotronH] Loading config... +[NemotronH] Config loaded: 52 layers (23 mamba, 6 attention, 0 mlp, 23 moe) +[NemotronH] Loading weights from safetensors... +[NemotronH] Building model... +[NemotronH] Model loaded successfully +[Load] wall: 0.285 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[phase] warmup_start monotonic_ns=177880118775030 boottime_ns=177880118775071 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=None, reserved 128 for generation) +[phase] measured_start monotonic_ns=177891602566085 boottime_ns=177891602566105 +[phase] decode_start monotonic_ns=177893431245234 boottime_ns=177893431245264 +[phase] measured_end monotonic_ns=177896364621427 boottime_ns=177896364621457 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1825.65 ms (280.45 tok/s) + Decode: 2933.38 ms (43.64 tok/s) + MLX peak memory: 18.48 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_decode_kernels.csv new file mode 100644 index 000000000..8bb8cacf2 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_decode_kernels.csv @@ -0,0 +1,67 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,14859,116.09,706288374,47533,39.404 +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,5842,45.64,468534653,80201,26.139 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,sampler_tail;ssm_step,2064;2067,11811,92.27,74876068,6340,4.177 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,ssm_step,2067,14605,114.1,61873542,4236,3.452 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",gemm,ssm_step,2067,5842,45.64,39556630,6771,2.207 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail;ssm_gated_norm;ssm_step,2064;2065;2067,29464,230.19,38317736,1300,2.138 +"void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)",sort,,,2921,22.82,38170513,13068,2.130 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,6731,52.59,30059306,4466,1.677 +"void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)",gemm,,,2921,22.82,27667128,9472,1.544 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn;ssm_step,2063;2067,12446,97.23,23643215,1900,1.319 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_gated_norm;ssm_step,2067,11684,91.28,20614076,1764,1.150 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_activation;sampler_tail;ssm_gated_norm;ssm_step,2064;2065;2067,11811,92.27,16667997,1411,0.930 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,8763,68.46,15767118,1799,0.880 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",copy,ssm_step,2067,11684,91.28,15617110,1337,0.871 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail;ssm_step,2064;2067,8890,69.45,14823460,1667,0.827 +Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,2921,22.82,14299837,4896,0.798 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail;ssm_step,2064;2067,5969,46.63,13502162,2262,0.753 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_conv,,7390,57.73,12960725,1754,0.723 +"void mlx::core::rocm::CV2IAsTypeBDV2IBroadcastACEV2IBroadcastCAFV2IMaximumDEGV2OSquareF_VC_V2_2297668033614959926_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,5842,45.64,12942628,2215,0.722 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",copy,ssm_step,2067,11684,91.28,12389290,1060,0.691 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,ssm_silu,,2921,22.82,10408660,3563,0.581 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,762,5.95,9126043,11976,0.509 +Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,2921,22.82,7801133,2671,0.435 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",binary,ssm_step,2067,5842,45.64,7478353,1280,0.417 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,,,2921,22.82,7076124,2423,0.395 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,6864025,2350,0.383 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",conv,ssm_conv,,2921,22.82,6734212,2305,0.376 +"void mlx::core::rocm::rms_norm_kernel(float const*, float const*, float*, float, unsigned int, long)",rms_norm,ssm_gated_norm,,2921,22.82,5865750,2008,0.327 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,ssm_step,2067,2921,22.82,5656841,1937,0.316 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,ssm_step,2067,2921,22.82,5624914,1926,0.314 +"void mlx::core::rocm::Bf4ISigmoidACf4IBroadcastABDf4IBroadcastBAEf4OMultiplyCD_V_f4_6142509188972423790_contiguous(float const*, float*, unsigned int)",compiled,ssm_silu,,2921,22.82,5526126,1892,0.308 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,,,2921,22.82,5338718,1828,0.298 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,5240400,1794,0.292 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,5842,45.64,5165645,884,0.288 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",reduce,,,2921,22.82,5144760,1761,0.287 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,2921,22.82,5037535,1725,0.281 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,2921,22.82,5018697,1718,0.280 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",scan,ssm_step,2067,2921,22.82,4973741,1703,0.277 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<1, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",compiled,,,2921,22.82,4973435,1703,0.277 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<1, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",compiled,,,2921,22.82,3969179,1359,0.221 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,3517003,27693,0.196 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,664763,2617,0.037 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,652175,5135,0.036 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,127,0.99,576282,4538,0.032 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,551598,2172,0.031 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,411859,1621,0.023 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,407837,1606,0.023 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,384874,1515,0.021 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,372022,2929,0.021 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,358226,2821,0.020 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,127,0.99,290472,2287,0.016 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,279244,1099,0.016 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,274750,2163,0.015 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,242133,1907,0.014 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,236565,1863,0.013 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,236443,1862,0.013 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,219006,1724,0.012 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,189629,1493,0.011 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,185503,1461,0.010 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,171525,1351,0.010 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,162469,1279,0.009 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,161864,1275,0.009 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,141592,1115,0.008 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,136769,1077,0.008 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,,,12,0.09,16271,1356,0.001 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,2,0.02,4128,2064,0.000 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_kernel_stats.csv new file mode 100644 index 000000000..02970b385 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_kernel_stats.csv @@ -0,0 +1,79 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",6808,3426421258,503293.369271,58.88,75422,64161396,3663359.985598 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",17084,814305875,47664.825275,13.99,3486,1478851,78359.664750 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",232,411277846,1772749.336207,7.07,174848,4533175,950319.057386 +"Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",3404,114637291,33677.230024,1.97,1684,2727993,265331.122928 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",17020,86774534,5098.386251,1.49,1002,147636,10063.544402 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",13580,86305397,6355.331149,1.48,801,92894,8342.219112 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",13524,72797305,5382.823499,1.25,801,1023399,57686.972235 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",6956,63986368,9198.730305,1.10,1242,793607,66087.485245 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",10166,63690846,6265.084202,1.09,1162,524743,46782.003722 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",34106,51122348,1498.925350,0.8785,681,109485,3950.514807 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",6716,45479843,6771.864652,0.7815,1362,31299,5051.960443 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",10314,44689239,4332.871728,0.7679,881,613630,39742.306502 +"void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)",3404,44683039,13126.627203,0.7678,10860,68328,1556.463423 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",3404,43901995,12897.178320,0.7544,1242,970659,95596.805640 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",7844,38216474,4872.064508,0.6567,3927,184426,5309.165706 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",14458,32845499,2271.787177,0.5644,1122,46888,3749.688222 +"void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)",3358,31861786,9488.322216,0.5475,8295,27171,762.777451 +"Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",3404,27495215,8077.325206,0.4725,4207,266019,27141.454012 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",13764,23981401,1742.327884,0.4121,841,109085,4534.744851 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",3404,20631691,6061.013807,0.3545,2003,288661,30969.348544 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",8608,19018163,2209.359085,0.3268,961,375985,6918.744789 +"void mlx::core::rocm::kernel_sdpa_flash_wmma(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float*, mlx::core::rocm::FAWmmaParams)",12,18997767,1583147.250000,0.3264,1537922,1640515,29252.031112 +"void mlx::core::rocm::CV2IAsTypeBDV2IBroadcastACEV2IBroadcastCAFV2IMaximumDEGV2OSquareF_VC_V2_2297668033614959926_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",6808,18669077,2742.226351,0.3208,1763,160300,5054.914865 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",13616,18348066,1347.537162,0.3153,721,10299,966.403175 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",3404,16513766,4851.282609,0.2838,1522,246542,26732.833406 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",3404,14492780,4257.573443,0.2490,1683,571631,18165.485753 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",13616,14487988,1064.041422,0.2490,642,10179,1124.026110 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",3404,13478110,3959.491774,0.2316,2805,153929,4156.373402 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",876,10402280,11874.748858,0.1787,9338,16110,674.248452 +"Cijk_Alik_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",46,9606362,208833.956522,0.1651,196048,689613,72479.228925 +"void mlx::core::rocm::rms_norm_kernel(float const*, float const*, float*, float, unsigned int, long)",3404,8736680,2566.592244,0.1501,1522,190637,5427.917772 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",6716,8586269,1278.479601,0.1475,762,10700,1173.613247 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",3404,8076971,2372.788190,0.1388,1844,11421,1097.730536 +"void mlx::core::rocm::Bf4ISigmoidACf4IBroadcastABDf4IBroadcastBAEf4OMultiplyCD_V_f4_6142509188972423790_contiguous(float const*, float*, unsigned int)",3404,7717546,2267.199177,0.1326,1323,93174,3354.759478 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",230,6957475,30249.891304,0.1196,1803,86803,25412.650725 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",3404,6583493,1934.046122,0.1131,1322,11141,1608.374167 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",3358,6528054,1944.030375,0.1122,1242,10059,1336.835613 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",3358,6118262,1821.995831,0.1051,1202,8697,367.590312 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",6808,6100151,896.026880,0.1048,601,14025,285.035982 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",3404,6078316,1785.639248,0.1044,1042,10740,1729.135972 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",3404,6036146,1773.250881,0.1037,1442,9859,323.543549 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",3404,5968004,1753.232667,0.1026,1162,14628,941.620384 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<1, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",3358,5732254,1707.044074,0.0985,1363,7213,284.318705 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",92,5239143,56947.206522,0.0900,1763,125956,54413.635580 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<1, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",3358,4586818,1365.937463,0.0788,842,6332,252.573777 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,4194835,28343.479730,0.0721,26650,130364,8552.828390 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",46,3974490,86401.956522,0.0683,83397,164749,12077.196707 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",58,1869191,32227.431034,0.0321,1283,314109,80303.652510 +"Cijk_Alik_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x64x64_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA128_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB8_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT2_2_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR1_PKA0_SGROB0_SIA1_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG32_4_1_WGMXCC1",46,1211266,26331.869565,0.0208,22362,85841,9095.824374 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",92,798567,8680.076087,0.0137,7775,15789,1752.230053 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,746678,2522.560811,0.0128,1122,7734,1801.795879 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,734651,4963.858108,0.0126,3126,14748,2491.669762 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,626618,2116.952703,0.0108,1122,11101,2137.825846 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",148,623291,4211.425676,0.0107,2003,7694,2368.989377 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",296,524152,1770.783784,9.007e-03,1282,25288,1519.130648 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,498646,1684.614865,8.568e-03,1082,10179,1111.550492 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,480455,1623.158784,8.256e-03,841,10500,1799.509522 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,460831,3113.722973,7.919e-03,2524,7775,957.371237 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,453485,3064.087838,7.792e-03,2685,13345,1291.123868 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,347619,2348.777027,5.973e-03,1923,7454,913.620166 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,333145,1125.489865,5.725e-03,761,8977,878.000749 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,320154,2163.202703,5.501e-03,1683,7094,448.060734 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,318038,2148.905405,5.465e-03,1683,25448,2052.852438 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,315351,2130.750000,5.419e-03,1723,29776,2451.308574 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,278555,1882.128378,4.787e-03,1443,6773,986.250884 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,270267,1826.128378,4.644e-03,922,7334,1782.250963 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,220526,1490.040541,3.789e-03,961,10419,1615.576731 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,219137,1480.655405,3.766e-03,1242,8857,878.643471 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",148,211752,1430.756757,3.639e-03,961,9859,1548.547002 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,191604,1294.621622,3.292e-03,921,8737,1312.134307 +"void mlx::core::rocm::Cf4IAsTypeADf4OSigmoidCEf4IBroadcastDBFf4IBroadcastBDGf4IAddEFHf4ONegativeG_VV_V2f4_6142509188972423790_strided<2, unsigned int, 8>(hip_bfloat16 const*, hip::std::array, float const*, hip::std::array, float*, float*, hip::std::array, unsigned int)",46,190155,4133.804348,3.268e-03,3246,15389,1719.148958 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,187592,1267.513514,3.223e-03,921,6092,979.760852 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,180891,1222.236486,3.108e-03,961,7975,868.076130 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,169766,1147.067568,2.917e-03,961,5010,348.388470 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",70,130123,1858.900000,2.236e-03,962,6331,1187.304463 +"void mlx::core::rocm::gather_axis_kernel(float const*, int const*, float*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",46,127840,2779.130435,2.197e-03,1923,7414,1312.869912 +"void mlx::core::rocm::Ef4IBroadcastBCFf4IBroadcastCBGf4IAddEFHf4IBroadcastAGIf4IBroadcastGAJf4IDivideHIKf4IBroadcastJDLf4IBroadcastDJMf4OMultiplyKL_VVCS_f4f4f4_9971424201298772266_strided<2, unsigned int, 1>(float const*, hip::std::array, float const*, hip::std::array, float const*, float*, hip::std::array, unsigned int)",46,94656,2057.739130,1.627e-03,1242,7894,1080.356822 +"__amd_rocclr_fillBufferUnAligned",1,56787,56787.000000,9.758e-04,56787,56787,0.00000000e+00 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_summary.json new file mode 100644 index 000000000..1348c4d2d --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7_summary.json @@ -0,0 +1,188 @@ +{ + "name": "NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7", + "model_type": "nemotron_h", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 43.64, + "profiled_prefill_tok_s": 280.45, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 260007, + "dispatches_per_token": 2031.3, + "decode_wall_ms": 2933.376, + "decode_gpu_sum_ms": 1792.441, + "decode_gpu_busy_ms": 1792.441, + "gpu_ms_per_token": 14.0034, + "host_gap_ms_per_token": 8.9136, + "host_gap_pct_of_wall": 38.89, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 3194.5, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 116.09, + "share_pct": 39.4 + }, + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 45.64, + "share_pct": 26.14 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)", + "class": "binary", + "calls_per_token": 92.27, + "share_pct": 4.18 + }, + { + "kernel": "void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "binary", + "calls_per_token": 114.1, + "share_pct": 3.45 + }, + { + "kernel": "void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)", + "class": "gemm", + "calls_per_token": 45.64, + "share_pct": 2.21 + }, + { + "kernel": "void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)", + "class": "copy", + "calls_per_token": 230.19, + "share_pct": 2.14 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(float const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 22.82, + "share_pct": 2.13 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 52.59, + "share_pct": 1.68 + }, + { + "kernel": "void mlx::core::rocm::gemv_single(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int)", + "class": "gemm", + "calls_per_token": 22.82, + "share_pct": 1.54 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 97.23, + "share_pct": 1.32 + } + ], + "class_share_pct": { + "qmv": 39.4, + "gather_qmv": 26.14, + "binary": 11.89, + "copy": 6.57, + "gemm": 4.98, + "sort": 2.17, + "compiled": 2.11, + "rms_norm": 2.0, + "unary": 1.14, + "scan": 0.79, + "ternary": 0.75, + "reduce": 0.74, + "sdpa": 0.51, + "gather_scatter": 0.38, + "conv": 0.38, + "random": 0.03, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 46.16, + "moe_expert_gemv": 26.14, + "ssm_step": 20.28, + "add_rms_join": 2.12, + "ssm_gated_norm": 1.15, + "ssm_conv": 0.97, + "ssm_silu": 0.89, + "sampler_tail": 0.68, + "moe_activation": 0.62, + "moe_weighted_sum": 0.44, + "moe_gather_indices": 0.29, + "add_rms_join_post_attn": 0.26, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 367.4, + "ssm_step": 1141.0, + "ssm_conv": 68.5, + "ssm_silu": 45.6, + "ssm_gated_norm": 91.3, + "add_rms_join": 91.3, + "moe_gather_indices": 45.6, + "moe_expert_gemv": 45.6, + "moe_activation": 45.6, + "moe_weighted_sum": 45.6, + "add_rms_join_post_attn": 11.9, + "sampler_tail": 31.8, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 0.26, + "fallback_dispatches_per_token": 11.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "nemotron_h never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.68, + "fallback_dispatches_per_token": 31.8, + "reached_default_share_pct": 0.68, + "reached_default_dispatches_per_token": 31.8, + "reached_optin_share_pct": 0.68, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 27.48, + "fallback_dispatches_per_token": 182.6, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "fused_moe_forward's default branch is gather_qmm; the kernel path needs MLXCEL_FUSED_MOE_RELU2 and moe_fc1_relu2 (#2069)" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 20.28, + "fallback_dispatches_per_token": 1141.0, + "reached_default_share_pct": 20.28, + "reached_default_dispatches_per_token": 1141.0, + "reached_optin_share_pct": 20.28, + "note": "ssm_kernel_available() gates the decode step" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_bench.log new file mode 100644 index 000000000..f75286b9e --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_bench.log @@ -0,0 +1,15 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.190 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[phase] warmup_start monotonic_ns=170052628986246 boottime_ns=170052628986286 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[phase] measured_start monotonic_ns=170061773621421 boottime_ns=170061773621441 +[phase] decode_start monotonic_ns=170063436650806 boottime_ns=170063436650826 +[phase] measured_end monotonic_ns=170065611275141 boottime_ns=170065611275161 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1616.36 ms (316.76 tok/s) + Decode: 2174.62 ms (58.86 tok/s) + MLX peak memory: 23.56 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_decode_kernels.csv new file mode 100644 index 000000000..6df54c7e9 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_decode_kernels.csv @@ -0,0 +1,28 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,18288,142.88,720179835,39380,42.119 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,18415,143.87,565068463,30685,33.048 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,6096,47.62,77874564,12775,4.554 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",sort,,,6096,47.62,73834782,12112,4.318 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,24511,191.49,66594921,2717,3.895 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,6096,47.62,27849302,4568,1.629 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn,2063,12192,95.25,22669258,1859,1.326 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",softmax,,,6096,47.62,22513437,3693,1.317 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,rope_append,2063,12384,96.75,17827064,1440,1.043 +"void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)",rope,rope_append,2063,12192,95.25,16838820,1381,0.985 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,18288,142.88,15623995,854,0.914 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_weighted_sum;sampler_tail,2064;2065,12319,96.24,15320985,1244,0.896 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail,2064;2065,12319,96.24,13592392,1103,0.795 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,moe_weighted_sum,2065,6096,47.62,12765542,2094,0.747 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,6223,48.62,9708294,1560,0.568 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,6096,47.62,9577096,1571,0.560 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",reduce,,,6096,47.62,9561554,1568,0.559 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,,,6096,47.62,9180366,1506,0.537 +"void mlx::core::rocm::arg_reduce_partial, 256>(hip_bfloat16 const*, hip_bfloat16*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",arg_reduce,sampler_tail,2064,127,0.99,1210867,9534,0.071 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,435809,3432,0.025 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,392090,1544,0.023 +"void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",arg_reduce,sampler_tail,2064,127,0.99,267539,2107,0.016 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,235123,1851,0.014 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,207993,1638,0.012 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,192764,1518,0.011 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,176288,1388,0.010 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,,,96,0.75,157453,1640,0.009 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_kernel_stats.csv new file mode 100644 index 000000000..6ab3013cc --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_kernel_stats.csv @@ -0,0 +1,43 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",21312,3510144491,164702.725741,66.88,33542,13400806,1099849.631022 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",21172,651342229,30764.322171,12.41,2524,1444586,63546.096151 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",288,267972292,930459.347222,5.11,134411,2470225,579692.964348 +"void mlx::core::rocm::kernel_sdpa_flash_wmma(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float*, mlx::core::rocm::FAWmmaParams)",96,164436860,1712883.958333,3.13,1648006,1916509,37350.938949 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",7008,88739819,12662.645405,1.69,8736,19556,916.947204 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",7104,86620775,12193.239724,1.65,8256,70051,2473.696042 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",28372,80691040,2844.037784,1.54,1923,155601,2009.888702 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",7104,48684830,6853.157376,0.9276,1763,1222289,49120.142027 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",7104,43530409,6127.591357,0.8294,3326,249828,15873.421632 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",7104,34328709,4832.307010,0.6541,1322,253515,27688.058325 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",14208,32409769,2281.092976,0.6175,1403,174045,4320.536840 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",14356,30795169,2145.107899,0.5868,762,521657,14769.745821 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",7104,27276956,3839.661599,0.5197,3166,67767,1591.819973 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",14688,23993368,1633.535403,0.4572,961,92573,2507.171156 +"void mlx::core::rocm::rms_norm_strided_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long, int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array)",192,21918970,114161.302083,0.4176,25648,239088,86930.216246 +"void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)",14016,19593551,1397.941709,0.3733,962,10179,541.515525 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",14356,19321782,1345.902898,0.3681,801,44283,1397.119186 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",21312,18834387,883.745636,0.3589,601,16511,741.627711 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",488,18312446,37525.504098,0.3489,1002,461665,62476.634606 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",7104,11298754,1590.477759,0.2153,1242,10901,450.688213 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",7156,11237486,1570.358580,0.2141,1202,7854,494.857934 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",7008,10484596,1496.089612,0.1998,1002,8937,662.175844 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,7472307,77836.531250,0.1424,63439,319719,56740.811584 +"void mlx::core::rocm::block_sort_kernel(unsigned int const*, unsigned int*, int, long, long, long, long)",192,4974473,25908.713542,0.0948,13305,86963,22619.961799 +"void mlx::core::rocm::copy_gg_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,4850764,33685.861111,0.0924,3967,230992,35842.924242 +"void mlx::core::rocm::rope(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, long, int, HIP_vector_type)",192,3501632,18237.666667,0.0667,4167,91812,14451.117011 +"void mlx::core::rocm::arg_reduce_partial, 256>(hip_bfloat16 const*, hip_bfloat16*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",148,1411644,9538.135135,0.0269,8977,14147,485.434088 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,509108,3439.918919,9.700e-03,2725,8255,1195.113774 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,389734,2633.337838,7.426e-03,1283,110086,9586.189975 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",96,343681,3580.010417,6.548e-03,2044,8977,2004.489852 +"void mlx::core::rocm::gather_general_kernel(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, mlx::core::rocm::hip_array, unsigned int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,331337,3451.427083,6.313e-03,2004,16110,2362.113633 +"void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",148,316031,2135.344595,6.022e-03,1843,6812,675.397517 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",192,304326,1585.031250,5.799e-03,1162,6211,657.164382 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,280526,1895.445946,5.345e-03,1483,5290,416.266158 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",96,264737,2757.677083,5.044e-03,1963,9658,1657.220085 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,258367,2691.322917,4.923e-03,1603,8135,1702.656800 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,250190,1690.472973,4.767e-03,1162,16030,1575.310337 +"void mlx::core::rocm::copy_g_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,206789,2154.052083,3.940e-03,1282,12663,1793.981660 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,202173,1366.033784,3.852e-03,1001,5370,732.688363 +"void mlx::core::rocm::copy_v(int const*, float*, unsigned int)",96,192638,2006.645833,3.670e-03,1082,11020,1835.307586 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",96,161589,1683.218750,3.079e-03,1001,6532,1326.760211 +"void mlx::core::rocm::copy_v(float const*, int*, unsigned int)",96,160269,1669.468750,3.054e-03,1002,7734,1549.375018 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_plain_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_plain_bench.log new file mode 100644 index 000000000..e02666579 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_plain_bench.log @@ -0,0 +1,11 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.229 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 2373.42 ms (215.72 tok/s) + Decode: 2064.86 ms (61.99 tok/s) + MLX peak memory: 23.56 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_summary.json new file mode 100644 index 000000000..daba4d466 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_greedy_summary.json @@ -0,0 +1,179 @@ +{ + "name": "Qwen3-30B-A3B-4bit_greedy", + "model_type": "qwen3_moe", + "temperature": 0.0, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 58.86, + "profiled_prefill_tok_s": 316.76, + "plain_decode_tok_s": 61.99, + "plain_prefill_tok_s": 215.72, + "profiler_decode_slowdown_pct": 5.32, + "decode_dispatches": 197138, + "dispatches_per_token": 1540.1, + "decode_wall_ms": 2174.624, + "decode_gpu_sum_ms": 1709.857, + "decode_gpu_busy_ms": 1709.857, + "gpu_ms_per_token": 13.3583, + "host_gap_ms_per_token": 3.631, + "host_gap_pct_of_wall": 21.37, + "plain_wall_ms_per_token": 16.1316, + "plain_host_gap_ms_per_token_est": 2.7734, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 49311.5, + "last_dispatch_before_decode": "void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, un" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 142.88, + "share_pct": 42.12 + }, + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 143.87, + "share_pct": 33.05 + }, + { + "kernel": "void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)", + "class": "sdpa", + "calls_per_token": 47.62, + "share_pct": 4.55 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 47.62, + "share_pct": 4.32 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 191.49, + "share_pct": 3.89 + }, + { + "kernel": "void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 47.62, + "share_pct": 1.63 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 95.25, + "share_pct": 1.33 + }, + { + "kernel": "void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)", + "class": "softmax", + "calls_per_token": 47.62, + "share_pct": 1.32 + }, + { + "kernel": "void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "copy", + "calls_per_token": 96.75, + "share_pct": 1.04 + }, + { + "kernel": "void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)", + "class": "rope", + "calls_per_token": 95.25, + "share_pct": 0.98 + } + ], + "class_share_pct": { + "gather_qmv": 42.12, + "qmv": 33.05, + "sdpa": 4.55, + "sort": 4.32, + "rms_norm": 3.89, + "copy": 3.68, + "binary": 2.43, + "compiled": 1.63, + "softmax": 1.32, + "reduce": 1.31, + "rope": 0.98, + "gather_scatter": 0.62, + "arg_reduce": 0.09, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 48.4, + "moe_expert_gemv": 42.12, + "moe_weighted_sum": 2.13, + "add_rms_join": 1.82, + "add_rms_join_post_attn": 1.8, + "moe_activation": 1.63, + "rope_append": 1.02, + "moe_gather_indices": 0.91, + "sampler_tail": 0.17, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 721.8, + "rope_append": 96.0, + "add_rms_join_post_attn": 95.2, + "moe_gather_indices": 142.9, + "moe_expert_gemv": 142.9, + "moe_activation": 47.6, + "moe_weighted_sum": 190.5, + "add_rms_join": 95.2, + "sampler_tail": 7.9, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 2.82, + "fallback_dispatches_per_token": 191.2, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "qwen3_moe never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.17, + "fallback_dispatches_per_token": 7.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "greedy argmax dispatches neither sampler kernel" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 46.79, + "fallback_dispatches_per_token": 523.9, + "reached_default_share_pct": 46.79, + "reached_default_dispatches_per_token": 523.9, + "reached_optin_share_pct": 46.79, + "note": "forward_fused_kernel caller" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "no Mamba2 layer" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_bench.log new file mode 100644 index 000000000..83388a775 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_bench.log @@ -0,0 +1,15 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.176 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[phase] warmup_start monotonic_ns=178414456913588 boottime_ns=178414456913639 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[phase] measured_start monotonic_ns=178418268685543 boottime_ns=178418268685563 +[phase] decode_start monotonic_ns=178419917177432 boottime_ns=178419917177452 +[phase] measured_end monotonic_ns=178422375676783 boottime_ns=178422375676803 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1604.46 ms (319.11 tok/s) + Decode: 2458.50 ms (52.06 tok/s) + MLX peak memory: 23.56 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv new file mode 100644 index 000000000..5a40ade09 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_decode_kernels.csv @@ -0,0 +1,58 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,18288,142.88,731982803,40025,40.486 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,18415,143.87,576138799,31286,31.866 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,6096,47.62,87869590,14414,4.860 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",sort,,,6096,47.62,73434474,12046,4.062 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,24511,191.49,65644572,2678,3.631 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,6096,47.62,28386533,4657,1.570 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,moe_weighted_sum,2065,6096,47.62,25796363,4232,1.427 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,18288,142.88,23796479,1301,1.316 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn,2063,12192,95.25,23031340,1889,1.274 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",softmax,,,6096,47.62,22328471,3663,1.235 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,rope_append,2063,12384,96.75,18288958,1477,1.012 +"void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)",rope,rope_append,2063,12192,95.25,16372937,1343,0.906 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_weighted_sum;sampler_tail,2064;2065,12319,96.24,13921162,1130,0.770 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail,2064;2065,12446,97.23,12874458,1034,0.712 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,6604,51.59,12384547,1875,0.685 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",sort,sampler_tail,2064,1016,7.94,10374857,10211,0.574 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",reduce,,,6096,47.62,9622810,1579,0.532 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,6096,47.62,9542226,1565,0.528 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,6223,48.62,9451474,1519,0.523 +"void mlx::core::rocm::contiguous_scan(hip_bfloat16 const*, hip_bfloat16*, int)",scan,sampler_tail,2064,127,0.99,9146557,72020,0.506 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",softmax,sampler_tail,2064,127,0.99,6345075,49961,0.351 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,4041882,31826,0.224 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,2264477,17831,0.125 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",sort,sampler_tail,2064,254,1.98,1786952,7035,0.099 +__amd_rocclr_copyBuffer,copy,sampler_tail,2064,508,3.97,941180,1853,0.052 +__amd_rocclr_fillBufferUnAligned,copy,sampler_tail,2064,381,2.98,932249,2447,0.052 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail,2064,127,0.99,801507,6311,0.044 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,127,0.99,766599,6036,0.042 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,725355,5711,0.040 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,693662,2731,0.038 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,656027,5166,0.036 +"void mlx::core::rocm::ternary_g(bool const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,551183,4340,0.030 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",other,sampler_tail,2064,254,1.98,486427,1915,0.027 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,461140,1816,0.026 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,435574,3430,0.024 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,434294,3420,0.024 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,419785,1653,0.023 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,418548,1648,0.023 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,406640,3202,0.022 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,365533,2878,0.020 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,341891,1346,0.019 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,333947,1315,0.018 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,324406,1277,0.018 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,276314,2176,0.015 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,sampler_tail,2064,127,0.99,275155,2167,0.015 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,256042,2016,0.014 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,249140,1962,0.014 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,201536,1587,0.011 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,190065,1497,0.011 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,,,96,0.75,165548,1724,0.009 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,158496,1248,0.009 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,157137,1237,0.009 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,154491,1216,0.009 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,150566,1186,0.008 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,144590,1139,0.008 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,140097,1103,0.008 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",sort,sampler_tail,2064,127,0.99,137217,1080,0.008 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv new file mode 100644 index 000000000..d2dca25a1 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_kernel_stats.csv @@ -0,0 +1,72 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",21312,3489763857,163746.427224,65.54,34184,13262701,1087550.162298 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",21172,663746515,31350.203807,12.47,2524,1016302,62780.872776 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",288,264165125,917240.017361,4.96,133971,2352923,558988.823041 +"void mlx::core::rocm::kernel_sdpa_flash_wmma(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float*, mlx::core::rocm::FAWmmaParams)",96,164198415,1710400.156250,3.08,1644638,2068993,48734.240258 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",7008,99950566,14262.352454,1.88,8776,49293,2525.543926 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",7104,86302994,12148.507038,1.62,8095,69089,2539.908785 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",28372,79134915,2789.190575,1.49,1923,92413,1572.990469 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",7104,67708604,9531.053491,1.27,1683,1240802,56855.541307 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",7104,41186715,5797.679476,0.7736,3366,250148,11866.844023 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",7104,34174081,4810.540681,0.6418,1282,266739,27649.738466 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",14208,32458607,2284.530335,0.6096,1443,169156,3647.963031 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",14504,31748505,2188.948221,0.5963,721,541013,17206.102474 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",21312,28028280,1315.140766,0.5264,601,16992,1516.529300 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",7104,27039526,3806.239583,0.5078,3126,24806,1446.232114 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",14688,24284990,1653.389842,0.4561,962,44363,2181.773158 +"void mlx::core::rocm::rms_norm_strided_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long, int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array)",192,21692472,112981.625000,0.4074,25367,232715,85865.729782 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",488,19588794,40140.971311,0.3679,1122,435213,67546.531180 +"void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)",14016,18887389,1347.559147,0.3547,961,9738,512.093761 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",14356,17663795,1230.412023,0.3318,801,45445,1316.089617 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",7600,14343644,1887.321579,0.2694,1202,33944,1377.149627 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",1184,12352956,10433.239865,0.2320,7615,51777,3636.428613 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",7104,11336000,1595.720721,0.2129,1242,11141,515.766816 +"void mlx::core::rocm::contiguous_scan(hip_bfloat16 const*, hip_bfloat16*, int)",148,10946964,73965.972973,0.2056,71453,357669,23483.080335 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",7156,10893157,1522.241056,0.2046,1002,17753,784.916367 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",148,7395202,49967.581081,0.1389,49052,53500,458.044869 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,6496996,67677.041667,0.1220,63518,315391,25721.045686 +"void mlx::core::rocm::copy_gg_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,5555681,38581.118056,0.1043,4208,287658,42146.757018 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,4843662,32727.445946,0.0910,31378,165269,10971.949132 +"void mlx::core::rocm::block_sort_kernel(unsigned int const*, unsigned int*, int, long, long, long, long)",192,3603354,18767.468750,0.0677,13305,83837,12857.742498 +"void mlx::core::rocm::rope(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, long, int, HIP_vector_type)",192,3439274,17912.885417,0.0646,4368,88125,13784.931721 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,2707429,18293.439189,0.0508,17432,86001,5605.392289 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",296,2141574,7235.047297,0.0402,5490,31660,2361.078650 +"__amd_rocclr_copyBuffer",592,1122880,1896.756757,0.0211,1402,9137,927.063348 +"__amd_rocclr_fillBufferUnAligned",444,1048547,2361.592342,0.0197,761,10059,2499.189936 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",148,935358,6319.986486,0.0176,6051,6853,152.391368 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,913714,6173.743243,0.0172,5650,26489,1691.174358 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,864412,5840.621622,0.0162,5450,28013,1841.814033 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,757820,5120.405405,0.0142,1684,9458,2193.198654 +"void mlx::core::rocm::ternary_g(bool const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,642193,4339.141892,0.0121,4127,5370,177.315891 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",296,583011,1969.631757,0.0109,1643,11421,760.186522 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,538567,1819.483108,0.0101,1122,3687,641.456441 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,520295,3515.506757,9.772e-03,3046,17193,1194.327389 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,506829,3424.520270,9.519e-03,3126,3847,119.101424 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,501259,1693.442568,9.414e-03,1082,9378,679.936859 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,482951,1631.591216,9.071e-03,1202,2605,326.169061 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,469119,3169.722973,8.811e-03,2765,8255,576.185271 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,432059,2919.317568,8.115e-03,2365,10620,652.105298 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,396555,1339.712838,7.448e-03,882,2645,339.097169 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,386366,1305.290541,7.257e-03,841,5370,611.694864 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,371053,1253.557432,6.969e-03,801,3807,473.579908 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,334104,2257.459459,6.275e-03,1964,12183,973.195413 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,326777,2207.952703,6.137e-03,1643,22763,1829.217316 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,325447,2198.966216,6.112e-03,1523,6733,401.618112 +"void mlx::core::rocm::gather_general_kernel(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, mlx::core::rocm::hip_array, unsigned int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,322845,3362.968750,6.064e-03,2043,10620,2117.090976 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",244,322603,1322.143443,6.059e-03,922,6452,822.396208 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",96,318595,3318.697917,5.984e-03,2124,10620,1960.411386 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",192,309174,1610.281250,5.807e-03,1202,6212,724.080601 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,304041,2054.331081,5.710e-03,1563,15669,1143.331612 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",96,246860,2571.458333,4.636e-03,1963,9978,1438.766196 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,242410,2525.104167,4.553e-03,1563,7734,1573.670343 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, bool*, unsigned int)",148,242333,1637.385135,4.551e-03,1323,8256,778.018716 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,219959,1486.209459,4.131e-03,1322,2565,234.354033 +"void mlx::core::rocm::copy_g_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,193755,2018.281250,3.639e-03,1282,9779,1655.572038 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,190156,1284.837838,3.571e-03,921,9458,1220.735872 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,183226,1238.013514,3.441e-03,962,4208,385.565421 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,182626,1233.959459,3.430e-03,961,2285,317.888974 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,176694,1193.878378,3.319e-03,921,3567,253.695326 +"void mlx::core::rocm::copy_v(int const*, float*, unsigned int)",96,168235,1752.447917,3.160e-03,1042,7454,1450.651432 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,168033,1135.358108,3.156e-03,961,4368,290.317865 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",148,166872,1127.513514,3.134e-03,961,9698,820.195451 +"void mlx::core::rocm::copy_v(float const*, int*, unsigned int)",96,161697,1684.343750,3.037e-03,962,9497,1662.734424 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_summary.json new file mode 100644 index 000000000..1675504e8 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7-p0.95_summary.json @@ -0,0 +1,183 @@ +{ + "name": "Qwen3-30B-A3B-4bit_t0.7-p0.95", + "model_type": "qwen3_moe", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 52.06, + "profiled_prefill_tok_s": 319.11, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 204123, + "dispatches_per_token": 1594.7, + "decode_wall_ms": 2458.499, + "decode_gpu_sum_ms": 1807.98, + "decode_gpu_busy_ms": 1807.98, + "gpu_ms_per_token": 14.1248, + "host_gap_ms_per_token": 5.0822, + "host_gap_pct_of_wall": 26.46, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 46309.8, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 142.88, + "share_pct": 40.49 + }, + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 143.87, + "share_pct": 31.87 + }, + { + "kernel": "void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)", + "class": "sdpa", + "calls_per_token": 47.62, + "share_pct": 4.86 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 47.62, + "share_pct": 4.06 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 191.49, + "share_pct": 3.63 + }, + { + "kernel": "void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 47.62, + "share_pct": 1.57 + }, + { + "kernel": "void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)", + "class": "reduce", + "calls_per_token": 47.62, + "share_pct": 1.43 + }, + { + "kernel": "void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)", + "class": "copy", + "calls_per_token": 142.88, + "share_pct": 1.32 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 95.25, + "share_pct": 1.27 + }, + { + "kernel": "void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)", + "class": "softmax", + "calls_per_token": 47.62, + "share_pct": 1.23 + } + ], + "class_share_pct": { + "gather_qmv": 40.49, + "qmv": 31.87, + "sort": 4.91, + "sdpa": 4.86, + "copy": 3.97, + "rms_norm": 3.63, + "binary": 2.51, + "reduce": 2.01, + "softmax": 1.59, + "compiled": 1.57, + "rope": 0.91, + "gather_scatter": 0.77, + "scan": 0.73, + "ternary": 0.07, + "unary": 0.06, + "other": 0.03, + "random": 0.02, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 46.86, + "moe_expert_gemv": 40.49, + "sampler_tail": 2.75, + "moe_weighted_sum": 2.66, + "add_rms_join_post_attn": 1.73, + "add_rms_join": 1.68, + "moe_activation": 1.57, + "moe_gather_indices": 1.32, + "rope_append": 0.94, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 721.8, + "rope_append": 96.0, + "add_rms_join_post_attn": 95.2, + "moe_gather_indices": 142.9, + "moe_expert_gemv": 142.9, + "moe_activation": 47.6, + "moe_weighted_sum": 190.5, + "add_rms_join": 95.2, + "sampler_tail": 62.5, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 2.67, + "fallback_dispatches_per_token": 191.2, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "qwen3_moe never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 2.75, + "fallback_dispatches_per_token": 62.5, + "reached_default_share_pct": 2.75, + "reached_default_dispatches_per_token": 62.5, + "reached_optin_share_pct": 2.75, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 46.03, + "fallback_dispatches_per_token": 523.9, + "reached_default_share_pct": 46.03, + "reached_default_dispatches_per_token": 523.9, + "reached_optin_share_pct": 46.03, + "note": "forward_fused_kernel caller" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "no Mamba2 layer" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_bench.log new file mode 100644 index 000000000..a6092c0b2 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_bench.log @@ -0,0 +1,15 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.181 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[phase] warmup_start monotonic_ns=179101868751715 boottime_ns=179101868751755 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(40960), reserved 128 for generation) +[phase] measured_start monotonic_ns=179107438482593 boottime_ns=179107438482613 +[phase] decode_start monotonic_ns=179109266733283 boottime_ns=179109266733303 +[phase] measured_end monotonic_ns=179111460625366 boottime_ns=179111460625386 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1782.21 ms (287.28 tok/s) + Decode: 2193.89 ms (58.34 tok/s) + MLX peak memory: 23.56 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_decode_kernels.csv new file mode 100644 index 000000000..5b94cbaa8 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_decode_kernels.csv @@ -0,0 +1,44 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,18288,142.88,719070079,39319,41.855 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,18415,143.87,566122742,30742,32.953 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,6096,47.62,77889567,12777,4.534 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",sort,,,6096,47.62,73355261,12033,4.270 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;add_rms_join_post_attn,2063,24511,191.49,66659073,2720,3.880 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,6096,47.62,27918668,4580,1.625 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;add_rms_join_post_attn,2063,12192,95.25,22739003,1865,1.324 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",softmax,,,6096,47.62,22451409,3683,1.307 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,rope_append,2063,12384,96.75,17867408,1443,1.040 +"void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)",rope,rope_append,2063,12192,95.25,16276293,1335,0.947 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,18288,142.88,15910435,870,0.926 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_weighted_sum;sampler_tail,2064;2065,12319,96.24,15138401,1229,0.881 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail,2064;2065,12446,97.23,13578811,1091,0.790 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,moe_weighted_sum,2065,6096,47.62,12942786,2123,0.753 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,6223,48.62,9619001,1546,0.560 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,6096,47.62,9617652,1578,0.560 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,6223,48.62,9529022,1531,0.555 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",reduce,,,6096,47.62,9501120,1559,0.553 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,4011274,31585,0.233 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail,2064,127,0.99,811919,6393,0.047 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,721956,5685,0.042 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,507860,1999,0.030 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,465399,1832,0.027 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,425960,3354,0.025 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,421740,3321,0.025 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,420384,3310,0.024 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,396709,1562,0.023 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,367682,1448,0.021 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,359267,1414,0.021 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,333984,1315,0.019 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,289183,1139,0.017 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,227746,1793,0.013 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,220135,1733,0.013 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,215409,1696,0.013 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,210035,1654,0.012 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,190954,1504,0.011 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,190846,1503,0.011 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,181466,1429,0.011 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,178052,1402,0.010 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,176045,1386,0.010 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,172202,1356,0.010 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,,,96,0.75,155331,1618,0.009 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,151877,1196,0.009 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_kernel_stats.csv new file mode 100644 index 000000000..e820fbff9 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_kernel_stats.csv @@ -0,0 +1,58 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",21312,3716684642,174393.986580,67.94,32902,109675112,1372548.701951 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",21172,653372274,30860.205649,11.94,2524,1443183,63597.127465 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",288,268099202,930900.006944,4.90,134211,2871557,574429.414533 +"void mlx::core::rocm::kernel_sdpa_flash_wmma(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float*, mlx::core::rocm::FAWmmaParams)",96,166830622,1737818.979167,3.05,1652736,2215711,65427.380652 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",7008,88821146,12674.250285,1.62,8696,29816,896.992908 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",7104,86125509,12123.523226,1.57,8215,69250,2439.448526 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",28372,80580115,2840.128119,1.47,1923,89317,1713.201836 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",7104,48870833,6879.340231,0.8933,1763,1212672,49140.428859 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",7104,43933243,6184.296593,0.8031,3326,250670,16105.519012 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",7104,35237156,4960.185248,0.6441,1322,263934,28562.532701 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",14208,32275591,2271.649141,0.5900,1443,165148,3681.239258 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",14504,30924052,2132.105074,0.5653,762,529832,14973.742719 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",7104,27187996,3827.139077,0.4970,3166,25247,1436.679494 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",14688,23882176,1625.965142,0.4366,961,44162,2259.475860 +"void mlx::core::rocm::rms_norm_strided_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long, int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array)",192,22288235,116084.557292,0.4074,25568,249066,88376.121481 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",21312,19168668,899.430743,0.3504,641,15068,669.169272 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",488,19102440,39144.344262,0.3492,1041,446196,66506.741034 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",14356,19050147,1326.981541,0.3482,801,45204,1371.003310 +"void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)",14016,18825618,1343.151969,0.3441,962,6773,325.769800 +"void mlx::core::rocm::row_reduce_simple_kernel(float const*, float*, unsigned long, int)",7104,11259655,1584.973958,0.2058,1242,10901,444.131832 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",7156,11166401,1560.424958,0.2041,1162,11020,507.180257 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",7156,11106970,1552.119899,0.2030,1002,9137,782.856718 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,6495968,67666.333333,0.1187,63559,315111,25594.749890 +"void mlx::core::rocm::copy_gg_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,4989864,34651.833333,0.0912,4167,289863,39261.920121 +"void mlx::core::rocm::block_sort_kernel(unsigned int const*, unsigned int*, int, long, long, long, long)",192,4939849,25728.380208,0.0903,13345,84959,21818.738699 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,4818426,32556.932432,0.0881,31098,161983,10755.525706 +"void mlx::core::rocm::rope(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, long, int, HIP_vector_type)",192,3580744,18649.708333,0.0655,4408,95138,14997.022024 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",148,970133,6554.952703,0.0177,6051,31900,2146.022901 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,842421,5692.033784,0.0154,5330,12303,797.683273 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,606444,2048.797297,0.0111,1122,14267,1187.830750 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,545507,1842.929054,9.972e-03,1082,10220,1292.166707 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,515446,3482.743243,9.422e-03,2926,14868,1331.041713 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,504101,3406.087838,9.215e-03,2926,17673,1197.657521 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,491955,3324.020270,8.993e-03,2805,8536,857.834939 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,481826,1627.790541,8.808e-03,1162,8857,790.709128 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,426672,1441.459459,7.799e-03,841,7814,1124.855753 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,400505,1353.057432,7.321e-03,881,6011,788.362663 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",244,383428,1571.426230,7.009e-03,961,10420,1375.752068 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,346932,1172.067568,6.342e-03,802,8456,579.262853 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",96,333351,3472.406250,6.094e-03,2204,8095,1886.682394 +"void mlx::core::rocm::gather_general_kernel(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, mlx::core::rocm::hip_array, unsigned int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,305329,3180.510417,5.581e-03,2043,9698,1922.874526 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,299040,2020.540541,5.466e-03,1242,23524,2156.122847 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",192,293633,1529.338542,5.368e-03,1162,6172,522.464283 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,286492,1935.756757,5.237e-03,1322,43722,3883.227807 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,273312,2847.000000,4.996e-03,1643,7334,1784.797320 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,264982,1790.418919,4.844e-03,1322,6572,757.107061 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,262892,1776.297297,4.806e-03,1442,9538,1001.929835 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",96,257002,2677.104167,4.698e-03,1963,9818,1629.911237 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,235080,1588.378378,4.297e-03,1002,7855,1289.964475 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,228237,1542.141892,4.172e-03,962,5931,1165.597917 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,221656,1497.675676,4.052e-03,962,6172,1086.286235 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,220421,1489.331081,4.029e-03,1322,6372,537.412399 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,201216,1359.567568,3.678e-03,1002,6091,781.541066 +"void mlx::core::rocm::copy_g_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,194807,2029.239583,3.561e-03,1282,12863,1673.521058 +"void mlx::core::rocm::copy_v(int const*, float*, unsigned int)",96,185783,1935.239583,3.396e-03,1122,7574,1601.492894 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,181853,1228.736486,3.324e-03,962,5851,401.274778 +"void mlx::core::rocm::copy_v(float const*, int*, unsigned int)",96,139342,1451.479167,2.547e-03,1002,6292,1178.374447 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_summary.json new file mode 100644 index 000000000..c1fdf04bf --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/Qwen3-30B-A3B-4bit_t0.7_summary.json @@ -0,0 +1,182 @@ +{ + "name": "Qwen3-30B-A3B-4bit_t0.7", + "model_type": "qwen3_moe", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 58.34, + "profiled_prefill_tok_s": 287.28, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 200186, + "dispatches_per_token": 1564.0, + "decode_wall_ms": 2193.892, + "decode_gpu_sum_ms": 1717.99, + "decode_gpu_busy_ms": 1717.904, + "gpu_ms_per_token": 13.4211, + "host_gap_ms_per_token": 3.7187, + "host_gap_pct_of_wall": 21.7, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 50307.0, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 142.88, + "share_pct": 41.86 + }, + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 143.87, + "share_pct": 32.95 + }, + { + "kernel": "void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)", + "class": "sdpa", + "calls_per_token": 47.62, + "share_pct": 4.53 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 47.62, + "share_pct": 4.27 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 191.49, + "share_pct": 3.88 + }, + { + "kernel": "void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 47.62, + "share_pct": 1.63 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 95.25, + "share_pct": 1.32 + }, + { + "kernel": "void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)", + "class": "softmax", + "calls_per_token": 47.62, + "share_pct": 1.31 + }, + { + "kernel": "void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "copy", + "calls_per_token": 96.75, + "share_pct": 1.04 + }, + { + "kernel": "void mlx::core::rocm::rope_single_1d(hip_bfloat16 const*, hip_bfloat16*, int const*, float, float, long, unsigned int, unsigned int)", + "class": "rope", + "calls_per_token": 95.25, + "share_pct": 0.95 + } + ], + "class_share_pct": { + "gather_qmv": 41.86, + "qmv": 32.95, + "sdpa": 4.53, + "sort": 4.29, + "rms_norm": 3.88, + "copy": 3.69, + "binary": 2.59, + "compiled": 1.63, + "reduce": 1.36, + "softmax": 1.31, + "rope": 0.95, + "gather_scatter": 0.6, + "scan": 0.23, + "unary": 0.05, + "ternary": 0.04, + "random": 0.02, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 48.21, + "moe_expert_gemv": 41.86, + "moe_weighted_sum": 2.12, + "add_rms_join": 1.81, + "add_rms_join_post_attn": 1.78, + "moe_activation": 1.63, + "rope_append": 0.99, + "moe_gather_indices": 0.93, + "sampler_tail": 0.69, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 721.8, + "rope_append": 96.0, + "add_rms_join_post_attn": 95.2, + "moe_gather_indices": 142.9, + "moe_expert_gemv": 142.9, + "moe_activation": 47.6, + "moe_weighted_sum": 190.5, + "add_rms_join": 95.2, + "sampler_tail": 31.8, + "ssm_step": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 2.77, + "fallback_dispatches_per_token": 191.2, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "qwen3_moe never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.69, + "fallback_dispatches_per_token": 31.8, + "reached_default_share_pct": 0.69, + "reached_default_dispatches_per_token": 31.8, + "reached_optin_share_pct": 0.69, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 46.52, + "fallback_dispatches_per_token": 523.9, + "reached_default_share_pct": 46.52, + "reached_default_dispatches_per_token": 523.9, + "reached_optin_share_pct": 46.52, + "note": "forward_fused_kernel caller" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "no Mamba2 layer" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_bench.log new file mode 100644 index 000000000..2cfd310b9 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_bench.log @@ -0,0 +1,32 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.131 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] warmup_start monotonic_ns=174829528632894 boottime_ns=174829528632934 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] measured_start monotonic_ns=174832296156829 boottime_ns=174832296156849 +[phase] decode_start monotonic_ns=174833240264116 boottime_ns=174833240264147 +[phase] measured_end monotonic_ns=174836126534702 boottime_ns=174836126534733 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 898.51 ms (569.83 tok/s) + Decode: 2886.27 ms (44.35 tok/s) + MLX peak memory: 13.92 GB +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.123 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] warmup_start monotonic_ns=175562947751488 boottime_ns=175562947751528 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] measured_start monotonic_ns=175567677762974 boottime_ns=175567677763004 +[phase] decode_start monotonic_ns=175568626380535 boottime_ns=175568626380745 +[phase] measured_end monotonic_ns=175571193599854 boottime_ns=175571193600064 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 902.58 ms (567.26 tok/s) + Decode: 2567.22 ms (49.86 tok/s) + MLX peak memory: 13.84 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_decode_kernels.csv new file mode 100644 index 000000000..065ae720c --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_decode_kernels.csv @@ -0,0 +1,54 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,21463,167.68,355922725,16583,24.249 +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,15240,119.06,321360442,21087,21.894 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,18252,142.59,96410275,5282,6.568 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,ssm_step,2067,22824,178.31,81648807,3577,5.563 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",sort,,,5080,39.69,52475632,10330,3.575 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;ssm_gated_norm,,14859,116.09,50252375,3382,3.424 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail;ssm_step,2064;2065;2067,42291,330.4,41568501,983,2.832 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,10160,79.38,40802335,4016,2.780 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;ssm_step,2067,19812,154.78,36085950,1821,2.458 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,ssm_silu,,9144,71.44,26765267,2927,1.823 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",gemm,ssm_step,2067,4572,35.72,23331619,5103,1.590 +Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,4536,35.44,19377568,4272,1.320 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,,,10287,80.37,19376577,1884,1.320 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,ssm_step,2067,9144,71.44,18230628,1994,1.242 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_activation;sampler_tail;ssm_step,2064;2065;2067,14859,116.09,17538267,1180,1.195 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,ssm_gated_norm;ssm_step,2067,9144,71.44,17270096,1889,1.177 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,ssm_gated_norm;ssm_step,2067,14994,117.14,15779959,1052,1.075 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,5080,39.69,14922744,2938,1.017 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",copy,ssm_step,2067,18288,142.88,14772760,808,1.006 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,ssm_step,2067,9144,71.44,14520567,1588,0.989 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",copy,ssm_step,2067,13716,107.16,14264281,1040,0.972 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_conv,,10176,79.5,12860747,1264,0.876 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,15240,119.06,12390726,813,0.844 +"void mlx::core::rocm::softmax_kernel(float const*, float*, int)",softmax,,,5080,39.69,11319548,2228,0.771 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",binary,ssm_step,2067,9144,71.44,10452622,1143,0.712 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,9931603,2172,0.677 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,,,5080,39.69,9218058,1815,0.628 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",conv,ssm_conv,,4572,35.72,8180209,1789,0.557 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,5207,40.68,7914883,1520,0.539 +Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,4572,35.72,7302805,1597,0.498 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7218215,1579,0.492 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,ssm_step,2067,4572,35.72,7206384,1576,0.491 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,5080,39.69,7195557,1416,0.490 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7140876,1562,0.486 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7109022,1555,0.484 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,ssm_step,2067,4572,35.72,6935755,1517,0.473 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",scan,ssm_step,2067,4572,35.72,6520583,1426,0.444 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,ssm_step,2067,4572,35.72,6284750,1375,0.428 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,508,3.97,6071830,11952,0.414 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,4572,35.72,5748709,1257,0.392 +"void mlx::core::rocm::binary_sv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,5426278,1187,0.370 +"void mlx::core::rocm::gemv_single(float const*, float const*, float*, int, int)",gemm,ssm_step,2067,4572,35.72,5144757,1125,0.351 +Cijk_Ailk_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,36,0.28,4235138,117643,0.289 +"void mlx::core::rocm::arg_reduce_partial, 256>(hip_bfloat16 const*, hip_bfloat16*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",arg_reduce,sampler_tail,2064,127,0.99,778296,6128,0.053 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,127,0.99,669502,5272,0.046 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,431329,1698,0.029 +"void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",arg_reduce,sampler_tail,2064,127,0.99,268264,2112,0.018 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,245225,1931,0.017 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,238409,1877,0.016 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,217488,1713,0.015 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,201461,1586,0.014 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,196776,1549,0.013 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,36,0.28,73542,2043,0.005 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_kernel_stats.csv new file mode 100644 index 000000000..41243cd2d --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_kernel_stats.csv @@ -0,0 +1,77 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",17760,1500062709,84462.990372,42.46,14667,51761870,692050.541243 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",24674,408540107,16557.514266,11.56,3487,581169,30499.172326 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",338,254481436,752903.656805,7.20,112090,17563430,1375512.718359 +"Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",5328,167609148,31458.173423,4.74,1402,2528618,255669.606792 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",26496,144623065,5458.298045,4.09,1041,545222,28664.525021 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",20952,110337816,5266.218786,3.12,1242,76263,6171.672972 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",10658,76646053,7191.410490,2.17,1362,605735,48913.303992 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",17316,61950563,3577.648591,1.75,2925,69290,2019.996624 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",5920,61688117,10420.290034,1.75,7173,69571,1882.387426 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",5328,55361821,10390.732170,1.57,1282,795050,76611.434720 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",11760,52527518,4466.625680,1.49,2965,117420,6545.677240 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",48996,51420620,1049.486080,1.46,681,41919,1525.255256 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",10584,48002904,4535.421769,1.36,1322,473928,35595.219601 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",23016,47705963,2072.730405,1.35,1362,125114,2830.216300 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",10584,32187579,3041.154478,0.9111,2484,91010,1612.817283 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",5256,26816299,5102.035578,0.7590,4488,24566,558.985139 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",11988,26123309,2179.121538,0.7394,1362,140502,2884.619129 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",10584,23056004,2178.382842,0.6526,1562,49292,3540.101192 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",17324,22476206,1297.402794,0.6362,801,18395,1301.277601 +"Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",5184,22137931,4270.434221,0.6266,3607,13064,474.958824 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",5992,22065155,3682.435748,0.6245,1162,140142,15770.579827 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",5920,18801479,3175.925507,0.5322,1483,202620,11773.730387 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",17552,18633119,1061.595203,0.5274,761,10379,408.566812 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",11856,18358315,1548.440874,0.5196,881,206827,3843.611229 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",21316,17378422,815.275943,0.4919,601,9939,455.669977 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",5840,17157456,2937.920548,0.4856,2444,11462,345.083699 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",15984,16717897,1045.914477,0.4732,761,10059,427.780120 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",5328,14870573,2791.023461,0.4209,1603,81152,8564.631428 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",17760,14863248,836.894595,0.4207,641,14307,506.641581 +"void mlx::core::rocm::softmax_kernel(float const*, float*, int)",5920,13635206,2303.244257,0.3859,1964,12183,803.954144 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",10512,12006903,1142.209189,0.3398,801,9859,420.803513 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",5328,11691258,2194.305180,0.3309,1883,10940,431.057106 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",5988,9112122,1521.730461,0.2579,1202,10861,521.289350 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",5328,8514573,1598.080518,0.2410,1243,10380,482.715727 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",5328,8440046,1584.092718,0.2389,1242,10580,475.061365 +"Cijk_Ailk_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",72,8435329,117157.347222,0.2388,112010,127399,2900.717692 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",5328,8412404,1578.904655,0.2381,1162,15549,475.812223 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",80,8273620,103420.250000,0.2342,101790,107762,1086.772687 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",5256,8218629,1563.666096,0.2326,1242,8536,377.959227 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",5256,7970747,1516.504376,0.2256,1322,9578,430.137959 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",5328,7408035,1390.396959,0.2097,1202,10540,455.310684 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",584,6914891,11840.566781,0.1957,7213,18234,886.428155 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",456,6907746,15148.565789,0.1955,1002,136055,19867.863627 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",5256,6611953,1257.981925,0.1871,1082,10059,456.848178 +"void mlx::core::rocm::binary_sv(float const*, float const*, float*, unsigned int)",5256,6215720,1182.595129,0.1759,842,9458,380.800406 +"void mlx::core::rocm::gemv_single(float const*, float const*, float*, int, int)",5256,5941964,1130.510654,0.1682,881,9658,409.110304 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,5259621,36525.145833,0.1489,1523,78787,34275.935934 +"void mlx::core::rocm::block_sort_kernel(unsigned int const*, unsigned int*, int, long, long, long, long)",160,4628595,28928.718750,0.1310,13345,78988,24487.536199 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,4384053,54800.662500,0.1241,45966,398225,40339.988318 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",288,4232929,14697.670139,0.1198,1683,173806,21571.850409 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,3399375,22968.750000,0.0962,4729,1315986,151641.754946 +"Cijk_Alik_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",72,2045230,28405.972222,0.0579,26209,130124,12232.148325 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_strided<2, unsigned int, 16>(hip_bfloat16 const*, hip::std::array, hip_bfloat16*, hip::std::array, unsigned int)",72,1942880,26984.444444,0.0550,24365,32661,1828.607620 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,1243011,8632.020833,0.0352,7774,16270,1532.009036 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_strided<2, unsigned int, 16>(hip_bfloat16 const*, hip::std::array, hip_bfloat16 const*, hip::std::array, hip_bfloat16*, hip::std::array, unsigned int)",80,1223213,15290.162500,0.0346,13545,36228,3220.715969 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,1084192,11293.666667,0.0307,1442,142667,21777.296723 +"void mlx::core::rocm::arg_reduce_partial, 256>(hip_bfloat16 const*, hip_bfloat16*, unsigned int*, unsigned long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, long, int, int)",148,904291,6110.074324,0.0256,5210,11982,1362.627869 +"Cijk_Ailk_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x112x32_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR1_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB4_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA2048_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA16_LPB16_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT1_7_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR1_PKA0_SGROB0_SIA3_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG64_2_1_WGMXCC1",8,718465,89808.125000,0.0203,62076,277881,75995.548746 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",8,713538,89192.250000,0.0202,76664,165029,30652.678926 +"Cijk_Alik_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x32x32_MI16x16x1_SN_LDSB1_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS1_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU2_K1_LDSTI0_LBSPPA256_LBSPPB0_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB0_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT4_1_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR1_PLR1_PKA0_SGROB0_SIA2_SS1_SPO0_SRVW0_SSO0_SVW4_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSMn1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA4_VWB1_WSGRA0_WSGRB0_WS32_WG16_4_1_WGMXCC1",8,615473,76934.125000,0.0174,53060,235642,64133.009012 +"void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, unsigned int const*, unsigned int*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, unsigned long, int)",148,315913,2134.547297,8.942e-03,1843,7174,811.822526 +"void mlx::core::rocm::gather_general_kernel(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, mlx::core::rocm::hip_array, unsigned int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,299078,3738.475000,8.465e-03,2044,9939,2363.870212 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",80,297442,3718.025000,8.419e-03,2445,8456,1713.499491 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,289823,1958.263514,8.203e-03,1362,34264,2759.297701 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,281053,1899.006757,7.955e-03,1563,6532,931.307073 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,274515,1854.831081,7.770e-03,1362,6813,729.568434 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,244461,1651.763514,6.919e-03,1162,11902,932.595423 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",80,231835,2897.937500,6.562e-03,1924,11822,2074.492308 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,222824,1505.567568,6.307e-03,1082,6212,1055.404105 +"void mlx::core::rocm::binary_sv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",8,179858,22482.250000,5.091e-03,20479,25969,2102.269097 +"void mlx::core::rocm::copy_g_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,160748,2009.350000,4.550e-03,1282,8094,1475.616757 +"void mlx::core::rocm::copy_v(int const*, float*, unsigned int)",80,146113,1826.412500,4.136e-03,1042,6612,1476.824722 +"void mlx::core::rocm::copy_v(float const*, int*, unsigned int)",80,139705,1746.312500,3.954e-03,1002,6532,1623.011949 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",80,112935,1411.687500,3.197e-03,1002,6051,715.358096 +"__amd_rocclr_fillBufferUnAligned",1,86202,86202.000000,2.440e-03,86202,86202,0.00000000e+00 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",2,27251,13625.500000,7.713e-04,13505,13746,170.412734 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_plain_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_plain_bench.log new file mode 100644 index 000000000..f55db7ef4 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_plain_bench.log @@ -0,0 +1,12 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.132 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1352.12 ms (378.66 tok/s) + Decode: 2086.50 ms (61.35 tok/s) + MLX peak memory: 13.92 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_summary.json new file mode 100644 index 000000000..425d907a9 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_greedy_summary.json @@ -0,0 +1,189 @@ +{ + "name": "granite-4.0-h-tiny-4bit_greedy", + "model_type": "granitemoehybrid", + "temperature": 0.0, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 49.86, + "profiled_prefill_tok_s": 567.26, + "plain_decode_tok_s": 61.35, + "plain_prefill_tok_s": 378.66, + "profiler_decode_slowdown_pct": 23.04, + "decode_dispatches": 409182, + "dispatches_per_token": 3196.7, + "decode_wall_ms": 2567.219, + "decode_gpu_sum_ms": 1467.807, + "decode_gpu_busy_ms": 1467.807, + "gpu_ms_per_token": 11.4672, + "host_gap_ms_per_token": 8.5892, + "host_gap_pct_of_wall": 42.83, + "plain_wall_ms_per_token": 16.2999, + "plain_host_gap_ms_per_token_est": 4.8327, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 50359.2, + "last_dispatch_before_decode": "void mlx::core::rocm::arg_reduce_final, 256>(hip_bfloat16 const*, un" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 167.68, + "share_pct": 24.25 + }, + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 119.06, + "share_pct": 21.89 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)", + "class": "binary", + "calls_per_token": 142.59, + "share_pct": 6.57 + }, + { + "kernel": "void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "binary", + "calls_per_token": 178.31, + "share_pct": 5.56 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 39.69, + "share_pct": 3.58 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 116.09, + "share_pct": 3.42 + }, + { + "kernel": "void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)", + "class": "copy", + "calls_per_token": 330.4, + "share_pct": 2.83 + }, + { + "kernel": "void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 79.38, + "share_pct": 2.78 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 154.78, + "share_pct": 2.46 + }, + { + "kernel": "void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 71.44, + "share_pct": 1.82 + } + ], + "class_share_pct": { + "qmv": 25.27, + "gather_qmv": 21.89, + "binary": 20.86, + "copy": 9.21, + "compiled": 4.6, + "gemm": 4.05, + "sort": 3.58, + "rms_norm": 3.42, + "unary": 1.91, + "ternary": 1.24, + "scan": 0.92, + "softmax": 0.77, + "reduce": 0.63, + "gather_scatter": 0.6, + "conv": 0.56, + "sdpa": 0.41, + "arg_reduce": 0.07, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 35.89, + "ssm_step": 29.82, + "moe_expert_gemv": 21.89, + "add_rms_join": 3.42, + "ssm_gated_norm": 2.08, + "moe_activation": 1.88, + "ssm_silu": 1.82, + "ssm_conv": 1.33, + "moe_gather_indices": 0.84, + "moe_weighted_sum": 0.8, + "sampler_tail": 0.2, + "add_rms_join_post_attn": 0.0, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 669.9, + "ssm_step": 1678.5, + "ssm_conv": 107.2, + "ssm_silu": 71.4, + "ssm_gated_norm": 107.2, + "add_rms_join": 156.8, + "moe_gather_indices": 119.1, + "moe_expert_gemv": 119.1, + "moe_activation": 79.4, + "moe_weighted_sum": 79.4, + "sampler_tail": 8.9, + "add_rms_join_post_attn": 0.0, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "granitemoehybrid never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.2, + "fallback_dispatches_per_token": 8.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "greedy argmax dispatches neither sampler kernel" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 25.42, + "fallback_dispatches_per_token": 396.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "granitemoehybrid does not call forward_fused_kernel" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 29.82, + "fallback_dispatches_per_token": 1678.5, + "reached_default_share_pct": 29.82, + "reached_default_dispatches_per_token": 1678.5, + "reached_optin_share_pct": 29.82, + "note": "ssm_kernel_available() gates the decode step" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_bench.log new file mode 100644 index 000000000..2fa5a616f --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_bench.log @@ -0,0 +1,16 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.123 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] warmup_start monotonic_ns=178565955789417 boottime_ns=178565955789447 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] measured_start monotonic_ns=178571197140456 boottime_ns=178571197140477 +[phase] decode_start monotonic_ns=178573364115358 boottime_ns=178573364115428 +[phase] measured_end monotonic_ns=178576347348162 boottime_ns=178576347348232 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 2123.15 ms (241.15 tok/s) + Decode: 2983.23 ms (42.91 tok/s) + MLX peak memory: 13.84 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_decode_kernels.csv new file mode 100644 index 000000000..3c4aed988 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_decode_kernels.csv @@ -0,0 +1,82 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,21463,167.68,364501216,16983,23.861 +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,15240,119.06,275912607,18105,18.062 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,18252,142.59,105434486,5777,6.902 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,ssm_step,2067,22824,178.31,88816040,3891,5.814 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",sort,,,5080,39.69,54654651,10759,3.578 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;ssm_gated_norm,,14859,116.09,53328914,3589,3.491 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail;ssm_step,2064;2065;2067,42418,331.39,46095621,1087,3.017 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,10160,79.38,42439016,4177,2.778 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;ssm_step,2067,19812,154.78,37944486,1915,2.484 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,ssm_silu,,9144,71.44,29192298,3193,1.911 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",gemm,ssm_step,2067,4572,35.72,25497242,5577,1.669 +Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,4536,35.44,21035089,4637,1.377 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,,,10287,80.37,20456822,1989,1.339 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail;ssm_step,2064;2067,9271,72.43,19937773,2151,1.305 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_activation;sampler_tail;ssm_step,2064;2065;2067,14859,116.09,18739555,1261,1.227 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,ssm_gated_norm;ssm_step,2067,9144,71.44,18647315,2039,1.221 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",copy,ssm_step,2067,18288,142.88,16880431,923,1.105 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail;ssm_step,2064;2067,9271,72.43,16608705,1791,1.087 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,ssm_gated_norm;ssm_step,2067,14994,117.14,16289177,1086,1.066 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,5080,39.69,15475363,3046,1.013 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",copy,ssm_step,2067,13716,107.16,15280867,1114,1.000 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_conv,,10176,79.5,14499563,1425,0.949 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,15240,119.06,12932583,849,0.847 +"void mlx::core::rocm::softmax_kernel(float const*, float*, int)",softmax,,,5080,39.69,11739495,2311,0.768 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",binary,ssm_step,2067,9144,71.44,11080091,1212,0.725 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,5588,43.66,10887544,1948,0.713 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,10845758,2372,0.710 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,,,5080,39.69,9768939,1923,0.639 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",conv,ssm_conv,,4572,35.72,9007101,1970,0.590 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,sampler_tail;ssm_step,2064;2067,4699,36.71,8066774,1717,0.528 +Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,4572,35.72,8027363,1756,0.525 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7855018,1718,0.514 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7797043,1705,0.510 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,ssm_step,2067,4572,35.72,7743118,1694,0.507 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7679542,1680,0.503 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,5080,39.69,7669271,1510,0.502 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",scan,ssm_step,2067,4572,35.72,7191170,1573,0.471 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,ssm_step,2067,4572,35.72,6886422,1506,0.451 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,4572,35.72,6535808,1430,0.428 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,508,3.97,6133324,12073,0.401 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",sort,sampler_tail,2064,889,6.95,6119552,6884,0.401 +"void mlx::core::rocm::contiguous_scan(hip_bfloat16 const*, hip_bfloat16*, int)",scan,sampler_tail,2064,127,0.99,6101665,48045,0.399 +"void mlx::core::rocm::binary_sv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,5788451,1266,0.379 +"void mlx::core::rocm::gemv_single(float const*, float const*, float*, int, int)",gemm,ssm_step,2067,4572,35.72,5759834,1260,0.377 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",softmax,sampler_tail,2064,127,0.99,4285745,33746,0.281 +Cijk_Ailk_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,36,0.28,4221459,117263,0.276 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,2907204,22891,0.190 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",sort,sampler_tail,2064,254,1.98,1878914,7397,0.123 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,1535036,12087,0.100 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,1412892,5563,0.092 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,254,1.98,1331529,5242,0.087 +__amd_rocclr_copyBuffer,copy,sampler_tail,2064,508,3.97,989970,1949,0.065 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,650300,5120,0.043 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,127,0.99,647378,5097,0.042 +__amd_rocclr_fillBufferUnAligned,copy,sampler_tail,2064,381,2.98,585628,1537,0.038 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,579879,4566,0.038 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",other,sampler_tail,2064,254,1.98,507665,1999,0.033 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,453891,1787,0.030 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,452450,3563,0.030 +"void mlx::core::rocm::ternary_g(bool const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail,2064,127,0.99,447964,3527,0.029 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::default_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1})",ternary,sampler_tail,2064,254,1.98,441023,1736,0.029 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,435696,1715,0.029 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,418703,1648,0.027 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",sort,sampler_tail,2064,127,0.99,391131,3080,0.026 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,359389,1415,0.024 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,347257,1367,0.023 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,339667,2675,0.022 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,332825,2621,0.022 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,298599,1176,0.020 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,263296,2073,0.017 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,246263,1939,0.016 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,223743,1762,0.015 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,166709,1313,0.011 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",sort,sampler_tail,2064,127,0.99,162825,1282,0.011 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,161146,1269,0.011 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,160858,1267,0.011 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,157658,1241,0.010 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,156212,1230,0.010 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,154730,1218,0.010 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,153846,1211,0.010 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,36,0.28,76581,2127,0.005 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_kernel_stats.csv new file mode 100644 index 000000000..dfa75e68a --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_kernel_stats.csv @@ -0,0 +1,103 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",17760,2655703189,149532.837218,54.09,15028,569357202,5507714.230845 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",24674,418587500,16964.719948,8.53,3527,567983,30711.427384 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",338,278147724,822922.260355,5.66,110887,24443730,1923127.840669 +"Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",5328,179478753,33685.952140,3.66,1363,14873805,323207.824492 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",26496,152995482,5774.286005,3.12,1002,553276,28845.440664 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",20952,120694617,5760.529639,2.46,1242,79709,8112.533936 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",5328,117964841,22140.548236,2.40,1282,62515446,859684.233738 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",10806,78771851,7289.640107,1.60,1283,597078,48681.908330 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",17316,65653578,3791.497921,1.34,2925,67687,2598.449065 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",5920,64188709,10842.687331,1.31,7254,68528,4366.870417 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",49144,56803239,1155.852983,1.16,681,148157,1740.059634 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",11760,54409217,4626.634099,1.11,2966,117058,6576.064353 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",10732,50800949,4733.595695,1.03,1283,897360,36089.561416 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",23016,50122486,2177.723584,1.02,1362,126798,3141.061807 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",10584,35005741,3307.420729,0.7130,2524,95739,2307.177884 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",5256,29305079,5575.547755,0.5969,4569,26970,2875.821856 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",11988,27384507,2284.326577,0.5577,1402,143227,3024.103428 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",10584,24697371,2333.462868,0.5030,1563,52859,3641.431797 +"Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",5184,24017870,4633.076775,0.4892,3606,21480,2298.491261 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",17324,23886784,1378.826137,0.4865,762,26370,1416.616311 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",5992,22533784,3760.644860,0.4589,1162,138900,15644.971706 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",11856,20237258,1706.921221,0.4122,881,206306,3916.427730 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",21316,19859765,931.683477,0.4045,601,10620,706.010269 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",5920,19796615,3344.022804,0.4032,1522,416740,12937.003596 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",17552,19269164,1097.832954,0.3925,721,10379,692.134421 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",15984,17937482,1122.214840,0.3653,722,10219,696.625913 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",5840,17793329,3046.802911,0.3624,2445,10740,826.016489 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",5328,16041961,3010.878566,0.3267,1563,313788,9502.355188 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",17760,15448214,869.831869,0.3146,601,20237,622.222477 +"void mlx::core::rocm::softmax_kernel(float const*, float*, int)",5920,14109089,2383.292061,0.2874,1883,12303,1076.329888 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",10512,12734180,1211.394597,0.2594,801,9938,709.948007 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",5328,12729844,2389.234985,0.2593,1883,10941,1202.155110 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",6432,12623748,1962.647388,0.2571,1163,29856,1463.208036 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",5476,9450241,1725.756209,0.1925,1242,9698,916.810414 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",5328,9212385,1729.051239,0.1876,1282,11021,1019.818307 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",5328,9076332,1703.515766,0.1849,1163,14908,964.878230 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",5256,9046821,1721.236872,0.1843,1243,10860,923.866347 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",5256,8920054,1697.118341,0.1817,1283,10179,1085.240424 +"Cijk_Ailk_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",72,8404450,116728.472222,0.1712,111328,124953,2699.830120 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",80,8265028,103312.850000,0.1683,101710,108443,1202.890697 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",456,8120136,17807.315789,0.1654,1122,133410,18542.327141 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",5328,8092950,1518.947072,0.1648,1202,8015,789.578934 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",5256,7513337,1429.478120,0.1530,1082,10419,1010.103070 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::merge_oddeven_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::merge_sort_block_merge_impl >(unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::radix_merge_compare, ihipStream_t*, bool, std::iterator_traits::value_type*, std::iterator_traits::value_type*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::vsmem_t)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(unsigned int*, unsigned int*, unsigned int*, unsigned int*) const::{lambda(auto:1)#3})",1036,7316439,7062.199807,0.1490,5690,38632,2317.899825 +"void mlx::core::rocm::contiguous_scan(hip_bfloat16 const*, hip_bfloat16*, int)",148,7295500,49293.918919,0.1486,47448,231353,15069.963208 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",584,7011976,12006.808219,0.1428,7213,18154,1151.656348 +"void mlx::core::rocm::gemv_single(float const*, float const*, float*, int, int)",5256,6632671,1261.923706,0.1351,881,10740,866.709935 +"void mlx::core::rocm::binary_sv(float const*, float const*, float*, unsigned int)",5256,6631161,1261.636416,0.1351,842,10459,767.860854 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,5312386,36891.569444,0.1082,1562,78787,34622.362179 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",148,4996276,33758.621622,0.1018,32661,41878,1205.343722 +"void mlx::core::rocm::block_sort_kernel(unsigned int const*, unsigned int*, int, long, long, long, long)",160,4636808,28980.050000,0.0944,13345,78507,24475.260227 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,4363488,54543.600000,0.0889,46046,396902,40102.057036 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",288,4286502,14883.687500,0.0873,1683,174807,21555.049141 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",296,4162632,14062.945946,0.0848,4328,1316544,107298.784481 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,3384214,22866.310811,0.0689,21440,23885,324.982299 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_iteration >(hip_bfloat16*, std::iterator_traits::value_type*, hip_bfloat16*, unsigned int*, std::iterator_traits::value_type*, unsigned int*, unsigned int, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::detail::onesweep_lookback_state*, bool, bool, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::detail::block_id_wrapper, ihipStream_t*, bool)::{lambda(auto:1, auto:2, auto:3, auto:4)#1}::operator()(hip_bfloat16*, hip_bfloat16*, unsigned int*, unsigned int*) const::{lambda(auto:1)#1})",296,2197268,7423.202703,0.0448,5210,14187,2086.031228 +"Cijk_Alik_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",72,2044546,28396.472222,0.0416,26129,130244,12256.580517 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_strided<2, unsigned int, 16>(hip_bfloat16 const*, hip::std::array, hip_bfloat16*, hip::std::array, unsigned int)",72,1956663,27175.875000,0.0399,24606,33904,2311.834129 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_selector, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_block_sort, false, unsigned int*, unsigned int*, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity_decomposer>(unsigned int*, unsigned int*, unsigned int*, unsigned int*, unsigned int, unsigned int&, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,1836959,12411.885135,0.0374,11863,58109,3785.383248 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,1248731,8671.743056,0.0254,7735,14748,1580.351730 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_strided<2, unsigned int, 16>(hip_bfloat16 const*, hip::std::array, hip_bfloat16 const*, hip::std::array, hip_bfloat16*, hip::std::array, unsigned int)",80,1217797,15222.462500,0.0248,12663,27371,2724.148985 +"__amd_rocclr_copyBuffer",592,1156801,1954.055743,0.0236,1202,10660,1304.842617 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,1077812,11227.208333,0.0220,1483,145472,21816.333769 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",8,915554,114444.250000,0.0186,76784,370694,103543.167403 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,799378,5401.202703,0.0163,1362,34584,3405.701318 +"__amd_rocclr_fillBufferUnAligned",445,782074,1757.469663,0.0159,721,99186,4901.102752 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,771367,5211.939189,0.0157,4408,22682,1556.004946 +"Cijk_Ailk_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x112x32_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR1_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB4_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA2048_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA16_LPB16_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT1_7_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR1_PKA0_SGROB0_SIA3_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG64_2_1_WGMXCC1",8,715659,89457.375000,0.0146,61635,277520,75992.848669 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,702670,4747.770270,0.0143,1523,21681,2557.422132 +"Cijk_Alik_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x32x32_MI16x16x1_SN_LDSB1_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS1_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU2_K1_LDSTI0_LBSPPA256_LBSPPB0_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB0_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT4_1_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR1_PLR1_PKA0_SGROB0_SIA2_SS1_SPO0_SRVW0_SSO0_SVW4_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSMn1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA4_VWB1_WSGRA0_WSGRB0_WS32_WG16_4_1_WGMXCC1",8,607698,75962.250000,0.0124,52338,234479,64053.647169 +"mlx::core::rocm::iota_kernel(unsigned int*, int)",296,600717,2029.449324,0.0122,1242,7133,1164.107132 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,547543,1849.807432,0.0112,1122,12704,883.925275 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,526908,3560.189189,0.0107,3206,10980,637.579696 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::default_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::transform_impl, unsigned int*, unsigned int*, rocprim::ROCPRIM_400600_NS::identity >(unsigned int*, unsigned int*, unsigned long, rocprim::ROCPRIM_400600_NS::identity, ihipStream_t*, bool)::{lambda(auto:1)#1})",296,526017,1777.084459,0.0107,1322,9017,1005.956698 +"void mlx::core::rocm::ternary_g(bool const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",148,525026,3547.472973,0.0107,3005,4769,383.568022 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,512604,1731.770270,0.0104,1162,2765,447.132553 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,485788,1641.175676,9.894e-03,1041,3487,496.917771 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#1})",148,462706,3126.391892,9.424e-03,2324,7775,1315.490212 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,416340,1406.554054,8.480e-03,882,6211,977.938515 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,403243,1362.307432,8.213e-03,1041,2885,312.717423 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,398378,2691.743243,8.114e-03,1563,9418,1868.945875 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,384681,2599.195946,7.835e-03,2204,3526,236.232840 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,347168,1172.864865,7.071e-03,801,2445,244.685052 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, bool*, unsigned int)",148,310946,2100.986486,6.333e-03,1323,6853,625.868637 +"void mlx::core::rocm::gather_general_kernel(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, mlx::core::rocm::hip_array, unsigned int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,309096,3863.700000,6.295e-03,2004,9377,2343.219852 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",80,296108,3701.350000,6.031e-03,2484,8456,1692.269236 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",228,290704,1275.017544,5.921e-03,1001,6212,386.682000 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,286016,1932.540541,5.825e-03,1603,6893,1052.612244 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,258891,1749.263514,5.273e-03,1363,2285,135.086838 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",80,217447,2718.087500,4.429e-03,1923,11542,1790.015580 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,195922,1323.797297,3.990e-03,1042,2605,309.461062 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,187634,1267.797297,3.822e-03,1122,2244,184.844335 +"void rocprim::ROCPRIM_400600_NS::detail::trampoline_kernel, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2}, rocprim::ROCPRIM_400600_NS::detail::comp_target<(rocprim::ROCPRIM_400600_NS::detail::gen)9, (rocprim::ROCPRIM_400600_NS::detail::target_arch)1100, (rocprim::ROCPRIM_400600_NS::detail::gpu)3, (rocprim::ROCPRIM_400600_NS::detail::rep)0>, rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_histogram_config_static_selector, (rocprim::ROCPRIM_400600_NS::arch::wavefront::target)0>(rocprim::ROCPRIM_400600_NS::detail::radix_sort_onesweep_global_offsets(hip_bfloat16*, unsigned int*, unsigned int*, unsigned int, unsigned int, rocprim::ROCPRIM_400600_NS::identity_decomposer, unsigned int, unsigned int, ihipStream_t*, bool)::{lambda(auto:1)#2})",148,187393,1266.168919,3.817e-03,1122,5971,553.069289 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,186669,1261.277027,3.802e-03,1002,2365,233.187228 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,182947,1236.128378,3.726e-03,1042,1803,169.559505 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,179898,1215.527027,3.664e-03,961,2364,208.533768 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,178092,1203.324324,3.627e-03,962,1923,169.317607 +"void mlx::core::rocm::binary_sv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",8,177011,22126.375000,3.605e-03,20478,26570,2398.708935 +"void mlx::core::rocm::copy_g_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,155168,1939.600000,3.160e-03,1242,8014,1280.068972 +"void mlx::core::rocm::copy_v(int const*, float*, unsigned int)",80,155088,1938.600000,3.159e-03,1042,6813,1626.819049 +"void mlx::core::rocm::copy_v(float const*, int*, unsigned int)",80,137139,1714.237500,2.793e-03,1002,9738,1703.754813 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",2,27172,13586.000000,5.534e-04,13546,13626,56.568542 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_summary.json new file mode 100644 index 000000000..e4af892c5 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7-p0.95_summary.json @@ -0,0 +1,190 @@ +{ + "name": "granite-4.0-h-tiny-4bit_t0.7-p0.95", + "model_type": "granitemoehybrid", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 42.91, + "profiled_prefill_tok_s": 241.15, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 416294, + "dispatches_per_token": 3252.3, + "decode_wall_ms": 2983.233, + "decode_gpu_sum_ms": 1527.619, + "decode_gpu_busy_ms": 1527.619, + "gpu_ms_per_token": 11.9345, + "host_gap_ms_per_token": 11.372, + "host_gap_pct_of_wall": 48.79, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 47357.1, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 167.68, + "share_pct": 23.86 + }, + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 119.06, + "share_pct": 18.06 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)", + "class": "binary", + "calls_per_token": 142.59, + "share_pct": 6.9 + }, + { + "kernel": "void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "binary", + "calls_per_token": 178.31, + "share_pct": 5.81 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 39.69, + "share_pct": 3.58 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 116.09, + "share_pct": 3.49 + }, + { + "kernel": "void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)", + "class": "copy", + "calls_per_token": 331.39, + "share_pct": 3.02 + }, + { + "kernel": "void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 79.38, + "share_pct": 2.78 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 154.78, + "share_pct": 2.48 + }, + { + "kernel": "void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 71.44, + "share_pct": 1.91 + } + ], + "class_share_pct": { + "qmv": 24.87, + "binary": 21.9, + "gather_qmv": 18.06, + "copy": 9.79, + "compiled": 4.69, + "sort": 4.27, + "gemm": 4.22, + "rms_norm": 3.49, + "unary": 2.07, + "scan": 1.57, + "ternary": 1.36, + "softmax": 1.05, + "gather_scatter": 0.87, + "reduce": 0.7, + "conv": 0.59, + "sdpa": 0.4, + "dequantize": 0.04, + "other": 0.03, + "random": 0.03 + }, + "role_share_pct": { + "unattributed": 35.73, + "ssm_step": 31.28, + "moe_expert_gemv": 18.06, + "add_rms_join": 3.41, + "sampler_tail": 2.51, + "ssm_gated_norm": 2.1, + "ssm_silu": 1.91, + "moe_activation": 1.89, + "ssm_conv": 1.43, + "moe_gather_indices": 0.85, + "moe_weighted_sum": 0.82, + "add_rms_join_post_attn": 0.0, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 669.9, + "ssm_step": 1678.5, + "ssm_conv": 107.2, + "ssm_silu": 71.4, + "ssm_gated_norm": 107.2, + "add_rms_join": 156.8, + "moe_gather_indices": 119.1, + "moe_expert_gemv": 119.1, + "moe_activation": 79.4, + "moe_weighted_sum": 79.4, + "sampler_tail": 64.5, + "add_rms_join_post_attn": 0.0, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "granitemoehybrid never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 2.51, + "fallback_dispatches_per_token": 64.5, + "reached_default_share_pct": 2.51, + "reached_default_dispatches_per_token": 64.5, + "reached_optin_share_pct": 2.51, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 21.62, + "fallback_dispatches_per_token": 396.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "granitemoehybrid does not call forward_fused_kernel" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 31.28, + "fallback_dispatches_per_token": 1678.5, + "reached_default_share_pct": 31.28, + "reached_default_dispatches_per_token": 1678.5, + "reached_optin_share_pct": 31.28, + "note": "ssm_kernel_available() gates the decode step" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_bench.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_bench.log new file mode 100644 index 000000000..0c2ac54cc --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_bench.log @@ -0,0 +1,16 @@ +[mlx-rocm] bound HIP device 0: gfx1151 (AMD Radeon 8060S Graphics) cus=20 warp=32 lds=64KB +[Load] wall: 0.124 s MLX peak: 0.00 GB +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] warmup_start monotonic_ns=177715193907016 boottime_ns=177715193907056 +[mlx-rocm] gfx1151 native_wmma=1 cus=20 warp=32 low_cu=0 host_WARP_SIZE=32 +[hipBLASLt caps] device 0 (gfx1151): bf16=1 fp8_e4m3=0 fp8_e5m2=0 int8=1 +[long-prompt] target=512 tokens -> using 512 tokens (max_context=Some(131072), reserved 128 for generation) +[phase] measured_start monotonic_ns=177720082502603 boottime_ns=177720082502623 +[phase] decode_start monotonic_ns=177721195658847 boottime_ns=177721195658867 +[phase] measured_end monotonic_ns=177723941979446 boottime_ns=177723941979466 +[Profile Results] + Prompt tokens: 512 + Generated tokens: 128 + Prefill: 1060.89 ms (482.61 tok/s) + Decode: 2746.32 ms (46.61 tok/s) + MLX peak memory: 13.92 GB diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_decode_kernels.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_decode_kernels.csv new file mode 100644 index 000000000..ff17e754a --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_decode_kernels.csv @@ -0,0 +1,68 @@ +kernel,class,roles,port_units,calls,calls_per_token,total_ns,avg_ns,share_pct +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,21463,167.68,356297800,16601,25.100 +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",gather_qmv,moe_expert_gemv,2065,15240,119.06,267792156,17572,18.865 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,18252,142.59,96648188,5295,6.809 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,ssm_step,2067,22824,178.31,81533750,3572,5.744 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",sort,,,5080,39.69,52169137,10270,3.675 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",rms_norm,add_rms_join;ssm_gated_norm,,14859,116.09,50117013,3373,3.531 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",copy,moe_weighted_sum;sampler_tail;ssm_step,2064;2065;2067,42418,331.39,41217528,972,2.904 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,moe_activation,2065,10160,79.38,41188467,4054,2.902 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,add_rms_join;ssm_step,2067,19812,154.78,35784541,1806,2.521 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",compiled,ssm_silu,,9144,71.44,26570489,2906,1.872 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",gemm,ssm_step,2067,4572,35.72,23276650,5091,1.640 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,,,10287,80.37,19434755,1889,1.369 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",ternary,sampler_tail;ssm_step,2064;2067,9271,72.43,18744388,2022,1.320 +Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,4536,35.44,17684472,3899,1.246 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,ssm_gated_norm;ssm_step,2067,9144,71.44,17526438,1917,1.235 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",copy,moe_activation;sampler_tail;ssm_step,2064;2065;2067,14859,116.09,17307168,1165,1.219 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,ssm_gated_norm;ssm_step,2067,14994,117.14,15598670,1040,1.099 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",unary,sampler_tail;ssm_step,2064;2067,9271,72.43,15258775,1646,1.075 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",copy,ssm_step,2067,18288,142.88,15044404,823,1.060 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",qmv,,,5080,39.69,14965744,2946,1.054 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",copy,ssm_step,2067,13716,107.16,14360150,1047,1.012 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_conv,,10176,79.5,12759806,1254,0.899 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",copy,moe_gather_indices,2065,15240,119.06,12166541,798,0.857 +"void mlx::core::rocm::softmax_kernel(float const*, float*, int)",softmax,,,5080,39.69,11223683,2209,0.791 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",binary,ssm_step,2067,9144,71.44,10452113,1143,0.736 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,9897581,2165,0.697 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",reduce,,,5080,39.69,9177689,1807,0.647 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",conv,ssm_conv,,4572,35.72,8170676,1787,0.576 +Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,4572,35.72,7927124,1734,0.558 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",gather_scatter,sampler_tail,2064,5207,40.68,7908142,1519,0.557 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7238214,1583,0.510 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",binary,moe_weighted_sum,2065,5080,39.69,7182924,1414,0.506 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7180977,1571,0.506 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,7158899,1566,0.504 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,ssm_step,2067,4572,35.72,7128191,1559,0.502 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,ssm_step,2067,4572,35.72,6872259,1503,0.484 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",scan,ssm_step,2067,4572,35.72,6513279,1425,0.459 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",unary,ssm_step,2067,4572,35.72,6301627,1378,0.444 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",sdpa,,,508,3.97,6055303,11920,0.427 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,4572,35.72,5702766,1247,0.402 +"void mlx::core::rocm::binary_sv(float const*, float const*, float*, unsigned int)",binary,ssm_step,2067,4572,35.72,5512725,1206,0.388 +"void mlx::core::rocm::gemv_single(float const*, float const*, float*, int, int)",gemm,ssm_step,2067,4572,35.72,5099214,1115,0.359 +Cijk_Ailk_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8,gemm,ssm_step,2067,36,0.28,4223238,117312,0.298 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",scan,sampler_tail,2064,127,0.99,2697092,21237,0.190 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",binary,sampler_tail,2064,254,1.98,882167,3473,0.062 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",sort,sampler_tail,2064,127,0.99,415860,3274,0.029 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,411332,1619,0.029 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",reduce,sampler_tail,2064,254,1.98,411084,1618,0.029 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",gather_scatter,,,254,1.98,395261,1556,0.028 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",random,sampler_tail,2064,254,1.98,387124,1524,0.027 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,254,1.98,328336,1293,0.023 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,317001,1248,0.022 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,300240,2364,0.021 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,254,1.98,291109,1146,0.021 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",gather_scatter,sampler_tail,2064,127,0.99,231509,1823,0.016 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",copy,sampler_tail,2064,127,0.99,216649,1706,0.015 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",dequantize,,,127,0.99,207754,1636,0.015 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",gather_scatter,,,127,0.99,197571,1556,0.014 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,191605,1509,0.013 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",binary,sampler_tail,2064,127,0.99,174692,1376,0.012 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,173924,1369,0.012 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,173446,1366,0.012 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,167875,1322,0.012 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",copy,sampler_tail,2064,127,0.99,162825,1282,0.011 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",binary,sampler_tail,2064,127,0.99,161812,1274,0.011 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",binary,sampler_tail,2064,127,0.99,161501,1272,0.011 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",copy,ssm_step,2067,36,0.28,80952,2249,0.006 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_kernel_stats.csv b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_kernel_stats.csv new file mode 100644 index 000000000..cae6f9b9c --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_kernel_stats.csv @@ -0,0 +1,90 @@ +"Name","Calls","TotalDurationNs","AverageNs","Percentage","MinNs","MaxNs","StdDev" +"void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)",17760,1592604930,89673.701014,43.82,14828,87963881,1012238.452118 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",24674,409330467,16589.546365,11.26,3526,574677,30693.644037 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",338,255858123,756976.695266,7.04,113853,17631446,1377335.561728 +"Cijk_Ailk_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",5328,166802067,31306.694257,4.59,1442,2506257,253234.100174 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",26496,144625152,5458.376812,3.98,1002,550672,28624.986155 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",20952,110729549,5284.915473,3.05,1242,72937,6235.776001 +"void mlx::core::rocm::ternary_g(bool const*, float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",10806,77373830,7160.265593,2.13,1362,589666,48639.861916 +"void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)",17316,61810293,3569.547990,1.70,2925,62797,1910.442032 +"void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)",5920,61385159,10369.114696,1.69,7414,64681,1715.534334 +"void mlx::core::rocm::strided_scan(float const*, float*, int, long, long)",5328,55443766,10406.112237,1.53,1282,814086,76730.128116 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",11760,53014408,4508.027891,1.46,2965,116498,6449.654497 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)",49144,51231513,1042.477474,1.41,721,41398,1501.621707 +"void mlx::core::rocm::unary_v(float const*, float*, unsigned int)",10732,48855510,4552.321096,1.34,1322,453851,35310.049442 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",23016,47356505,2057.547141,1.30,1402,125315,2653.243379 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",10584,32053986,3028.532313,0.8820,2485,93095,1656.451641 +"void mlx::core::rocm::gemv_batched_inline(float const*, float const*, float*, int, int, mlx::core::rocm::GemvBatchParams)",5256,26808439,5100.540145,0.7377,4528,23685,665.007493 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",11988,26178299,2183.708625,0.7204,1402,136536,2874.415832 +"void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",10584,23375278,2208.548564,0.6432,1563,49733,3535.715410 +"void mlx::core::rocm::copy_v(float const*, hip_bfloat16*, unsigned int)",17324,22328641,1288.884842,0.6144,801,40516,1335.629470 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",5992,22093619,3687.186081,0.6080,1122,142707,15796.567364 +"Cijk_Alik_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",5184,20223926,3901.220293,0.5565,3327,19797,578.337597 +"void mlx::core::rocm::col_reduce_small(float const*, float*, mlx::core::rocm::ColReduceArgs, unsigned long)",5920,18592650,3140.650338,0.5116,1523,182502,11398.686373 +"void mlx::core::rocm::copy_s(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",17552,18470476,1052.328851,0.5083,721,9938,453.858426 +"void mlx::core::rocm::copy_gg_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",11856,18394106,1551.459683,0.5062,881,206347,3863.082385 +"void mlx::core::rocm::arange_kernel(int*, int, int, unsigned long)",21316,17804239,835.252346,0.4899,601,10420,448.037914 +"void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)",5840,17238927,2951.871062,0.4744,2405,11622,372.467087 +"void mlx::core::rocm::copy_s(float const*, float*, unsigned int)",15984,16906664,1057.724224,0.4652,721,10019,460.537856 +"void mlx::core::(anonymous namespace)::depthwise_conv1d_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, mlx::core::ConvParams<1>)",5328,14952109,2806.326764,0.4114,1563,82073,8583.643633 +"void mlx::core::rocm::arange_kernel(unsigned int*, unsigned int, unsigned int, unsigned long)",17760,14633618,823.964977,0.4027,601,14146,482.421279 +"void mlx::core::rocm::softmax_kernel(float const*, float*, int)",5920,13554481,2289.608277,0.3730,1924,12583,775.500302 +"void mlx::core::rocm::binary_ss(int const*, int const*, bool*, unsigned int)",10512,12026546,1144.077816,0.3309,761,10019,406.738471 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",5328,11679708,2192.137387,0.3214,1843,9658,440.281645 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",5988,9151440,1528.296593,0.2518,1163,15510,521.259004 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",5328,8580688,1610.489489,0.2361,1282,10259,505.573801 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",5328,8507232,1596.702703,0.2341,1202,14467,484.631709 +"Cijk_Ailk_Bjlk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",72,8424870,117012.083333,0.2318,112050,128842,2852.925781 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",5328,8381287,1573.064377,0.2306,1242,10179,436.995991 +"void mlx::core::rocm::qmm_wmma_dense_kernel(hip_bfloat16 const*, unsigned int const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int)",80,8279597,103494.962500,0.2278,101831,108724,1287.993115 +"void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)",5256,8230599,1565.943493,0.2265,1242,10861,385.434274 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",5256,7920176,1506.882801,0.2179,1322,9618,430.026370 +"void mlx::core::rocm::unary_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",5328,7452566,1398.754880,0.2051,1202,8095,489.439173 +"void mlx::core::rocm::kernel_sdpav_1pass(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, hip_bfloat16 const*, mlx::core::rocm::AttnParams)",584,6923695,11855.642123,0.1905,7174,27492,1056.496625 +"void mlx::core::rocm::gather_rows_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, unsigned int, int)",456,6905506,15143.653509,0.1900,1001,133450,19928.819308 +"void mlx::core::rocm::copy_g_byval(float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",5256,6619942,1259.501903,0.1822,1082,9337,484.897793 +"void mlx::core::rocm::binary_sv(float const*, float const*, float*, unsigned int)",5256,6290063,1196.739536,0.1731,841,10300,441.857010 +"void mlx::core::rocm::gemv_single(float const*, float const*, float*, int, int)",5256,5883382,1119.364916,0.1619,881,9698,384.406032 +"void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,5397594,37483.291667,0.1485,1562,81713,34999.310980 +"void mlx::core::rocm::block_sort_kernel(unsigned int const*, unsigned int*, int, long, long, long, long)",160,4592807,28705.043750,0.1264,13345,78828,24122.948747 +"void mlx::core::rocm::binary_g(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,4401528,55019.100000,0.1211,46087,396943,40152.767512 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",288,4200430,14584.826389,0.1156,1643,107682,20227.733666 +"void mlx::core::rocm::binary_vs(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",296,3660375,12366.131757,0.1007,1443,1328450,108291.074803 +"void mlx::core::rocm::contiguous_scan(float const*, float*, int)",148,3142206,21231.121622,0.0865,20799,25568,450.804887 +"Cijk_Alik_Bljk_SB_MT32x32x8_SN_1LDSB0_AMAS0_BL1_BS1_EPS0_GLVWA1_GLVWB1_GRVW1_GSU1_GSUASB_ISA1151_IU1_K1_KLA_LBSPPA0_LBSPPB0_LPA0_LPB0_LRVW1_MIAV0_MMFGLC_NLCA1_NLCB1_PGR0_PLR1_SIA1_SS0_SU32_SUS256_SVW4_TT2_2_TLDS0_UMLDSA0_UMLDSB0_USFGROn1_VAW1_VSn1_VW1_VWB1_WSGRA0_WSGRB0_WS64_WG16_16_1_WGM8",72,2073126,28793.416667,0.0570,26209,129443,12155.154770 +"void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_strided<2, unsigned int, 16>(hip_bfloat16 const*, hip::std::array, hip_bfloat16*, hip::std::array, unsigned int)",72,1934989,26874.847222,0.0532,24445,32862,1568.254162 +"void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_strided<2, unsigned int, 16>(hip_bfloat16 const*, hip::std::array, hip_bfloat16 const*, hip::std::array, hip_bfloat16*, hip::std::array, unsigned int)",80,1296749,16209.362500,0.0357,13545,73418,8339.715138 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",144,1258679,8740.826389,0.0346,7774,13626,1589.759930 +"void mlx::core::rocm::copy_g_byval(hip_bfloat16 const*, hip_bfloat16*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",96,1095013,11406.385417,0.0301,1482,145593,21948.458569 +"void mlx::core::rocm::softmax_kernel(hip_bfloat16 const*, hip_bfloat16*, int)",8,729847,91230.875000,0.0201,76744,176050,34316.810743 +"Cijk_Ailk_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x112x32_MI16x16x1_SN_LDSB0_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR1_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS0_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB4_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU1_K1_LDSTI0_LBSPPA2048_LBSPPB128_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA16_LPB16_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT1_7_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR2_PLR1_PKA0_SGROB0_SIA3_SS1_SPO0_SRVW0_SSO0_SVW1_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSM1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA1_VWB1_WSGRA0_WSGRB0_WS32_WG64_2_1_WGMXCC1",8,722232,90279.000000,0.0199,62237,276799,75373.320960 +"Cijk_Alik_Bljk_BBS_BH_Bias_HA_S_SAV_UserArgs_MT64x32x32_MI16x16x1_SN_LDSB1_AFC1_AG0_AGGSUA0_AGNTAB0_AFEM1_AFEM1_ASEM1_BL1_BS1_CD1_1_CLR0_CLS0_CADS0_DTLA0_DTLB0_DTLM0_DTVA0_DTVB0_DTVMXSA0_DTVMXSB0_DTVSM0_DPLB0_EPS1_ELFLR0_EMLLn1_FDSI0_GRPM1_GRVWA8_GRVWB8_GSUAMB_GLS0_HPLR0_ISA1151_ICIW0_IU2_K1_LDSTI0_LBSPPA256_LBSPPB0_LBSPPMXSA0_LBSPPMXSB0_LBSPPM0_LPA8_LPB0_LPMXSA0_LPMXSB0_LPM0_LRVW16_LWPMn1_MIAV1_MIWT4_1_MXLIBL_MXSFNS_MO40_MGRIPM1_NTn1_NTA0_NTB0_NTC0_NTD0_NTE0_NTMXSA0_NTMXSB0_NTM0_NTWS0_NVn1_NVA0_NVB0_NVC0_NVD0_NVE0_NVMXSA0_NVMXSB0_NVM0_NVWS0_NEPBS0_NLCA1_NLCB1_ONLL1_PAP0_PGL0_PGR1_PLR1_PKA0_SGROB0_SIA2_SS1_SPO0_SRVW0_SSO0_SVW4_SK0_SKFTR0_SKFDPO0_SKWS0_SKXCCM0_SNLL0_SIP1_SGRO0_TDMI0_TDMIM0_TDMS0_TIN0_THn1_THA0_THB0_THC0_THD0_THE0_THMXSA0_THMXSB0_THM0_THWS0_TLDS1_TLDSMn1_ULSGRO0_USL1_USLMX0_UDFMAC0_UIOFGRO0_UPLRP0_USFGROn1_USI0_VSn1_VWA4_VWB1_WSGRA0_WSGRB0_WS32_WG16_4_1_WGMXCC1",8,619440,77430.000000,0.0170,53420,237486,64679.368850 +"void mlx::core::rocm::searchsorted_kernel(float const*, float const*, unsigned int*, long, unsigned int, long)",148,492682,3328.932432,0.0136,2805,12704,815.167723 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,477491,1613.145270,0.0131,1122,7013,532.957335 +"void mlx::core::rocm::all_reduce_kernel(float const*, float*, unsigned long, unsigned long)",296,473172,1598.554054,0.0130,1042,9899,883.457302 +"mlx::core::rocm::rbitsc_kernel(unsigned int const*, unsigned char*, unsigned int, unsigned int, bool, unsigned int)",296,453404,1531.770270,0.0125,1082,9017,670.257402 +"void mlx::core::rocm::binary_ss(float const*, float const*, bool*, unsigned int)",296,379554,1282.277027,0.0104,841,5972,576.870861 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,366816,1239.243243,0.0101,801,9217,806.993530 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",148,349493,2361.439189,9.617e-03,2204,7254,421.340769 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",296,343887,1161.780405,9.463e-03,761,6733,509.498899 +"void mlx::core::rocm::affine_dequantize_packed_kernel(unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned long, int)",148,316075,2135.641892,8.698e-03,1322,38993,4133.794865 +"void mlx::core::rocm::gather_general_kernel(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, mlx::core::rocm::hip_array, unsigned int, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,309820,3872.750000,8.525e-03,2083,9017,2354.970459 +"void mlx::core::rocm::gather_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long)",80,306009,3825.112500,8.421e-03,2525,8896,1728.402628 +"void mlx::core::rocm::copy_v(unsigned int const*, float*, unsigned int)",228,299725,1314.583333,8.248e-03,922,6452,620.485569 +"void mlx::core::rocm::scatter_axis_kernel(hip_bfloat16 const*, int const*, hip_bfloat16*, long, long, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, int, long, long, long)",148,270866,1830.175676,7.453e-03,1323,7654,872.082180 +"void mlx::core::rocm::gather_rows_kernel(unsigned int const*, int const*, unsigned int*, long, unsigned int, int)",148,264856,1789.567568,7.288e-03,1162,23564,2109.338027 +"void mlx::core::rocm::copy_v(hip_bfloat16 const*, hip_bfloat16*, unsigned int)",148,252437,1705.655405,6.946e-03,1563,2084,108.678814 +"void mlx::core::rocm::copy_v(bool const*, float*, unsigned int)",148,223028,1506.945946,6.137e-03,1242,6131,450.542663 +"void mlx::core::rocm::binary_vs(float const*, float const*, float*, unsigned int)",80,221495,2768.687500,6.095e-03,1963,11021,1918.629090 +"void mlx::core::rocm::binary_vs(float const*, float const*, bool*, unsigned int)",148,200500,1354.729730,5.517e-03,1162,10339,853.559674 +"void mlx::core::rocm::binary_ss(unsigned int const*, unsigned int const*, unsigned int*, unsigned int)",148,198656,1342.270270,5.466e-03,922,8857,898.279524 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,198250,1339.527027,5.455e-03,922,7333,864.765278 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,193083,1304.614865,5.313e-03,962,8215,720.627396 +"void mlx::core::rocm::binary_ss(float const*, float const*, float*, unsigned int)",148,186777,1262.006757,5.140e-03,961,5210,356.150039 +"void mlx::core::rocm::binary_ss(bool const*, bool const*, bool*, unsigned int)",148,186628,1261.000000,5.135e-03,921,6011,533.025392 +"void mlx::core::rocm::binary_sv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)",8,178374,22296.750000,4.908e-03,20438,26410,2452.450888 +"void mlx::core::rocm::copy_g_byval(unsigned int const*, unsigned int*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",80,154328,1929.100000,4.247e-03,1243,8014,1267.334366 +"void mlx::core::rocm::copy_v(int const*, float*, unsigned int)",80,154044,1925.550000,4.239e-03,1082,6372,1575.006867 +"void mlx::core::rocm::copy_v(float const*, int*, unsigned int)",80,138056,1725.700000,3.799e-03,1002,6773,1607.848412 +"__amd_rocclr_fillBufferUnAligned",1,60834,60834.000000,1.674e-03,60834,60834,0.00000000e+00 +"void mlx::core::rocm::binary_g(int const*, int const*, bool*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)",2,26850,13425.000000,7.388e-04,13345,13505,113.137085 diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_summary.json b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_summary.json new file mode 100644 index 000000000..417601cea --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/granite-4.0-h-tiny-4bit_t0.7_summary.json @@ -0,0 +1,189 @@ +{ + "name": "granite-4.0-h-tiny-4bit_t0.7", + "model_type": "granitemoehybrid", + "temperature": 0.7, + "clock": "boottime", + "prompt_tokens": 512, + "generated_tokens": 128, + "profiled_decode_tok_s": 46.61, + "profiled_prefill_tok_s": 482.61, + "plain_decode_tok_s": null, + "plain_prefill_tok_s": null, + "profiler_decode_slowdown_pct": null, + "decode_dispatches": 412230, + "dispatches_per_token": 3220.5, + "decode_wall_ms": 2746.321, + "decode_gpu_sum_ms": 1419.512, + "decode_gpu_busy_ms": 1419.499, + "gpu_ms_per_token": 11.0898, + "host_gap_ms_per_token": 10.3658, + "host_gap_pct_of_wall": 48.31, + "plain_wall_ms_per_token": null, + "plain_host_gap_ms_per_token_est": null, + "checks": { + "kernels_straddling_decode_start": 0, + "idle_gap_before_first_decode_dispatch_us": 57338.5, + "last_dispatch_before_decode": "void mlx::core::rocm::binary_ss(unsigned int con" + }, + "top_kernels": [ + { + "kernel": "void mlx::core::rocm::qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, int, int, int, bool)", + "class": "qmv", + "calls_per_token": 167.68, + "share_pct": 25.1 + }, + { + "kernel": "void mlx::core::rocm::gather_qmv_wide_kernel(hip_bfloat16 const*, unsigned char const*, hip_bfloat16 const*, hip_bfloat16 const*, unsigned int const*, unsigned int const*, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int, hip_bfloat16*, int, int, int, int, int, bool, bool, long)", + "class": "gather_qmv", + "calls_per_token": 119.06, + "share_pct": 18.87 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(float const*, float const*, float*, unsigned int)", + "class": "binary", + "calls_per_token": 142.59, + "share_pct": 6.81 + }, + { + "kernel": "void mlx::core::rocm::binary_g(float const*, float const*, float*, long, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, mlx::core::rocm::hip_array, int)", + "class": "binary", + "calls_per_token": 178.31, + "share_pct": 5.74 + }, + { + "kernel": "void mlx::core::rocm::block_sort_kernel(hip_bfloat16 const*, unsigned int*, int, long, long, long, long)", + "class": "sort", + "calls_per_token": 39.69, + "share_pct": 3.68 + }, + { + "kernel": "void mlx::core::rocm::rms_norm_kernel(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, float, unsigned int, long)", + "class": "rms_norm", + "calls_per_token": 116.09, + "share_pct": 3.53 + }, + { + "kernel": "void mlx::core::rocm::copy_v(hip_bfloat16 const*, float*, unsigned int)", + "class": "copy", + "calls_per_token": 331.39, + "share_pct": 2.9 + }, + { + "kernel": "void mlx::core::rocm::CV2ISigmoidADV2IBroadcastACEV2IBroadcastCAFV2IMultiplyDEGV2IBroadcastFBHV2IBroadcastBFIV2OMultiplyGH_VV_V2V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 79.38, + "share_pct": 2.9 + }, + { + "kernel": "void mlx::core::rocm::binary_vv(hip_bfloat16 const*, hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "binary", + "calls_per_token": 154.78, + "share_pct": 2.52 + }, + { + "kernel": "void mlx::core::rocm::BV2ISigmoidACV2IBroadcastABDV2IBroadcastBAEV2OMultiplyCD_V_V2_6142509188972423790_contiguous(hip_bfloat16 const*, hip_bfloat16*, unsigned int)", + "class": "compiled", + "calls_per_token": 71.44, + "share_pct": 1.87 + } + ], + "class_share_pct": { + "qmv": 26.15, + "binary": 21.75, + "gather_qmv": 18.87, + "copy": 9.5, + "compiled": 4.77, + "gemm": 4.1, + "sort": 3.7, + "rms_norm": 3.53, + "unary": 2.02, + "ternary": 1.32, + "scan": 1.13, + "softmax": 0.79, + "reduce": 0.7, + "gather_scatter": 0.62, + "conv": 0.58, + "sdpa": 0.43, + "random": 0.03, + "dequantize": 0.01 + }, + "role_share_pct": { + "unattributed": 37.1, + "ssm_step": 30.75, + "moe_expert_gemv": 18.87, + "add_rms_join": 3.52, + "ssm_gated_norm": 2.16, + "moe_activation": 1.95, + "ssm_silu": 1.87, + "ssm_conv": 1.38, + "moe_gather_indices": 0.86, + "moe_weighted_sum": 0.83, + "sampler_tail": 0.72, + "add_rms_join_post_attn": 0.0, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "role_dispatches_per_token": { + "unattributed": 669.9, + "ssm_step": 1678.5, + "ssm_conv": 107.2, + "ssm_silu": 71.4, + "ssm_gated_norm": 107.2, + "add_rms_join": 156.8, + "moe_gather_indices": 119.1, + "moe_expert_gemv": 119.1, + "moe_activation": 79.4, + "moe_weighted_sum": 79.4, + "sampler_tail": 32.7, + "add_rms_join_post_attn": 0.0, + "rope_append": 0.0, + "paged_attention": 0.0 + }, + "port_units": { + "2063": { + "port": "fused_add_rms_norm + fused_rope_qk_append", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "granitemoehybrid never calls fused_add_rms_norm or forward_fused_rope_append" + }, + "2064": { + "port": "samplers (gumbel_max_sample, rejection_sample)", + "fallback_share_pct": 0.72, + "fallback_dispatches_per_token": 32.7, + "reached_default_share_pct": 0.72, + "reached_default_dispatches_per_token": 32.7, + "reached_optin_share_pct": 0.72, + "note": "sampled run: the draw is the fallback chain" + }, + "2065": { + "port": "fused MoE decode (moe_gateup, moe_down)", + "fallback_share_pct": 22.51, + "fallback_dispatches_per_token": 396.9, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "granitemoehybrid does not call forward_fused_kernel" + }, + "2067": { + "port": "SSM update (ssm_update_kernel)", + "fallback_share_pct": 30.75, + "fallback_dispatches_per_token": 1678.5, + "reached_default_share_pct": 30.75, + "reached_default_dispatches_per_token": 1678.5, + "reached_optin_share_pct": 30.75, + "note": "ssm_kernel_available() gates the decode step" + }, + "2068": { + "port": "paged attention (v1, v2, merge)", + "fallback_share_pct": 0.0, + "fallback_dispatches_per_token": 0.0, + "reached_default_share_pct": 0.0, + "reached_default_dispatches_per_token": 0.0, + "reached_optin_share_pct": 0.0, + "note": "the bench decodes into a dense KVCache; the paged path is not taken" + } + } +} diff --git a/benchmarks/rocm_profiles/gfx1151_929c80ab/guard.log b/benchmarks/rocm_profiles/gfx1151_929c80ab/guard.log new file mode 100644 index 000000000..cf0f5b146 --- /dev/null +++ b/benchmarks/rocm_profiles/gfx1151_929c80ab/guard.log @@ -0,0 +1,396 @@ +== Meta-Llama-3.1-8B-Instruct-4bit_greedy plain +2026-09-30T17:41:01+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T17:41:03+0900 idle streak reset after 1s: gpu=[3292996:mlxcel-bench-de] compilers=[] +2026-09-30T17:41:48+0900 idle streak reset after 10s: gpu=[3320768:mlxcel-bench-de] compilers=[] +2026-09-30T17:42:34+0900 idle streak reset after 10s: gpu=[3349635:mlxcel-bench-de] compilers=[] +2026-09-30T17:43:18+0900 idle streak reset after 10s: gpu=[3377917:mlxcel-bench-de] compilers=[] +2026-09-30T17:44:04+0900 idle streak reset after 11s: gpu=[3406710:mlxcel-bench-de] compilers=[] +2026-09-30T17:44:51+0900 idle streak reset after 11s: gpu=[3436308:mlxcel-bench-de] compilers=[] +2026-09-30T17:45:31+0900 idle streak reset after 10s: gpu=[3461649:mlxcel-bench-de] compilers=[] +2026-09-30T17:46:14+0900 idle streak reset after 11s: gpu=[3488911:mlxcel-bench-de] compilers=[] +2026-09-30T17:47:00+0900 idle streak reset after 12s: gpu=[] compilers=[3518044:cargo] +2026-09-30T17:58:15+0900 idle streak reset after 19s: gpu=[3935133:mlxcel-bench-de] compilers=[] +2026-09-30T17:58:37+0900 idle streak reset after 11s: gpu=[3948181:mlxcel-bench-de] compilers=[] +2026-09-30T17:58:59+0900 idle streak reset after 11s: gpu=[3961800:mlxcel-bench-de] compilers=[] +2026-09-30T17:59:21+0900 idle streak reset after 11s: gpu=[3976292:mlxcel-bench-de] compilers=[] +2026-09-30T17:59:42+0900 idle streak reset after 11s: gpu=[3989594:mlxcel-bench-de] compilers=[] +2026-09-30T18:00:04+0900 idle streak reset after 11s: gpu=[4002750:mlxcel-bench-de] compilers=[] +2026-09-30T18:00:25+0900 idle streak reset after 10s: gpu=[4016457:mlxcel-bench-de] compilers=[] +2026-09-30T18:00:47+0900 idle streak reset after 10s: gpu=[4030174:mlxcel-bench-de] compilers=[] +2026-09-30T18:01:10+0900 idle streak reset after 12s: gpu=[4043876:mlxcel-bench-de] compilers=[] +2026-09-30T18:01:55+0900 idle streak reset after 10s: gpu=[4073428:mlxcel-bench-de] compilers=[] +2026-09-30T18:02:35+0900 idle streak reset after 11s: gpu=[4098431:mlxcel-bench-de] compilers=[] +2026-09-30T18:02:56+0900 idle streak reset after 10s: gpu=[4112360:mlxcel-bench-de] compilers=[] +2026-09-30T18:03:20+0900 idle streak reset after 11s: gpu=[4127198:mlxcel-bench-de] compilers=[] +2026-09-30T18:03:43+0900 idle streak reset after 11s: gpu=[4141078:mlxcel-bench-de] compilers=[] +2026-09-30T18:04:04+0900 idle streak reset after 11s: gpu=[4155838:mlxcel-bench-de] compilers=[] +2026-09-30T18:04:26+0900 idle streak reset after 10s: gpu=[4169702:mlxcel-bench-de] compilers=[] +2026-09-30T18:04:49+0900 idle streak reset after 11s: gpu=[4184084:mlxcel-bench-de] compilers=[] +2026-09-30T18:05:11+0900 idle streak reset after 10s: gpu=[4524:mlxcel-bench-de] compilers=[] +2026-09-30T18:05:34+0900 idle streak reset after 11s: gpu=[18514:mlxcel-bench-de] compilers=[] +2026-09-30T18:05:56+0900 idle streak reset after 10s: gpu=[33064:mlxcel-bench-de] compilers=[] +2026-09-30T18:06:55+0900 idle streak reset after 36s: gpu=[70252:mlxcel-bench-de] compilers=[] +2026-09-30T18:07:17+0900 idle streak reset after 10s: gpu=[84344:mlxcel-bench-de] compilers=[] +2026-09-30T18:07:39+0900 idle streak reset after 11s: gpu=[98466:mlxcel-bench-de] compilers=[] +2026-09-30T18:08:00+0900 idle streak reset after 10s: gpu=[111681:mlxcel-bench-de] compilers=[] +2026-09-30T18:08:22+0900 idle streak reset after 10s: gpu=[125819:mlxcel-bench-de] compilers=[] +2026-09-30T18:08:44+0900 idle streak reset after 10s: gpu=[139951:mlxcel-bench-de] compilers=[] +2026-09-30T18:09:07+0900 idle streak reset after 11s: gpu=[154265:mlxcel-bench-de] compilers=[] +2026-09-30T18:09:28+0900 idle streak reset after 10s: gpu=[168248:mlxcel-bench-de] compilers=[] +2026-09-30T18:09:50+0900 idle streak reset after 10s: gpu=[182367:mlxcel-bench-de] compilers=[] +2026-09-30T18:10:13+0900 idle streak reset after 11s: gpu=[196536:mlxcel-bench-de] compilers=[] +2026-09-30T18:10:35+0900 idle streak reset after 11s: gpu=[211637:mlxcel-bench-de] compilers=[] +2026-09-30T18:10:58+0900 idle streak reset after 11s: gpu=[226794:mlxcel-bench-de] compilers=[] +2026-09-30T18:11:20+0900 idle streak reset after 10s: gpu=[241001:mlxcel-bench-de] compilers=[] +2026-09-30T18:12:20+0900 idle streak reset after 36s: gpu=[279121:mlxcel] compilers=[] +2026-09-30T18:12:45+0900 idle streak reset after 11s: gpu=[295124:mlxcel-bench-de] compilers=[] +2026-09-30T18:13:07+0900 idle streak reset after 10s: gpu=[308966:mlxcel] compilers=[] +2026-09-30T18:13:30+0900 idle streak reset after 11s: gpu=[323893:mlxcel] compilers=[] +2026-09-30T18:13:53+0900 idle streak reset after 11s: gpu=[338559:mlxcel] compilers=[] +2026-09-30T18:14:16+0900 idle streak reset after 11s: gpu=[353825:mlxcel] compilers=[] +2026-09-30T18:14:40+0900 idle streak reset after 11s: gpu=[377038:mlxcel] compilers=[] +2026-09-30T18:15:03+0900 idle streak reset after 10s: gpu=[404143:mlxcel] compilers=[] +2026-09-30T18:15:27+0900 idle streak reset after 10s: gpu=[432193:mlxcel] compilers=[] +2026-09-30T18:15:49+0900 idle streak reset after 11s: gpu=[447198:mlxcel] compilers=[] +2026-09-30T18:16:13+0900 idle streak reset after 11s: gpu=[462265:?] compilers=[] +2026-09-30T18:16:35+0900 idle streak reset after 10s: gpu=[476239:mlxcel] compilers=[] +2026-09-30T18:16:58+0900 idle streak reset after 11s: gpu=[490996:mlxcel] compilers=[] +2026-09-30T18:17:20+0900 idle streak reset after 10s: gpu=[505312:mlxcel] compilers=[] +2026-09-30T18:17:44+0900 idle streak reset after 10s: gpu=[520633:mlxcel] compilers=[] +2026-09-30T18:18:06+0900 idle streak reset after 10s: gpu=[535163:mlxcel] compilers=[] +2026-09-30T18:18:30+0900 idle streak reset after 10s: gpu=[550738:mlxcel] compilers=[] +2026-09-30T18:18:54+0900 idle streak reset after 11s: gpu=[567404:mlxcel-bench-de] compilers=[] +2026-09-30T18:19:17+0900 idle streak reset after 10s: gpu=[581810:mlxcel] compilers=[] +2026-09-30T18:19:40+0900 idle streak reset after 10s: gpu=[597299:mlxcel] compilers=[] +2026-09-30T18:20:04+0900 idle streak reset after 10s: gpu=[612787:mlxcel] compilers=[] +2026-09-30T18:21:46+0900 idle streak reset after 64s: gpu=[678718:mlxcel] compilers=[] +2026-09-30T18:22:10+0900 idle streak reset after 11s: gpu=[694166:mlxcel] compilers=[] +2026-09-30T18:22:33+0900 idle streak reset after 11s: gpu=[709598:mlxcel] compilers=[] +2026-09-30T18:22:57+0900 idle streak reset after 11s: gpu=[724238:?] compilers=[] +2026-09-30T18:23:19+0900 idle streak reset after 10s: gpu=[739325:mlxcel] compilers=[] +2026-09-30T18:23:45+0900 idle streak reset after 12s: gpu=[755854:mlxcel-bench-de] compilers=[] +2026-09-30T18:24:58+0900 idle streak reset after 46s: gpu=[] compilers=[802788:cargo] +2026-09-30T18:29:01+0900 idle streak reset after 10s: gpu=[962091:rocminfo] compilers=[962050:cargo] +2026-09-30T18:29:37+0900 idle streak reset after 1s: gpu=[] compilers=[984576:cargo] +2026-09-30T18:36:35+0900 idle streak reset after 36s: gpu=[] compilers=[1246167:cargo] +2026-09-30T18:47:26+0900 idle streak reset after 53s: gpu=[1660523:rocminfo] compilers=[1660484:cargo] +2026-09-30T18:54:27+0900 idle streak reset after 53s: gpu=[] compilers=[1935991:cargo] +2026-09-30T18:54:43+0900 idle streak reset after 10s: gpu=[1945971:rocm_inflight_b] compilers=[1945068:cargo] +2026-09-30T18:54:50+0900 idle streak reset after 4s: gpu=[1950516:rocm_inflight_b] compilers=[] +2026-09-30T18:55:27+0900 idle streak reset after 22s: gpu=[1973989:rocm_inflight_b] compilers=[1973929:cargo] +2026-09-30T18:56:54+0900 idle streak reset after 56s: gpu=[] compilers=[2028086:cargo] +2026-09-30T19:06:57+0900 idle streak reset after 27s: gpu=[2397046:mlxcel] compilers=[] +2026-09-30T19:07:16+0900 idle streak reset after 4s: gpu=[] compilers=[2410732:cargo] +2026-09-30T19:08:27+0900 idle streak reset after 11s: gpu=[2455665:mlxcel] compilers=[] +2026-09-30T19:08:49+0900 idle streak reset after 10s: gpu=[2470457:mlxcel] compilers=[] +2026-09-30T19:09:37+0900 idle streak reset after 11s: gpu=[2501066:mlxcel] compilers=[] +2026-09-30T19:10:25+0900 idle streak reset after 11s: gpu=[2531600:mlxcel] compilers=[] +2026-09-30T19:10:56+0900 idle streak reset after 12s: gpu=[2552023:mlxcel-bench-de] compilers=[] +2026-09-30T19:11:17+0900 idle streak reset after 10s: gpu=[2566474:mlxcel] compilers=[] +2026-09-30T19:12:02+0900 idle streak reset after 11s: gpu=[2595551:mlxcel] compilers=[] +2026-09-30T19:12:48+0900 idle streak reset after 10s: gpu=[2625290:mlxcel] compilers=[] +2026-09-30T19:13:21+0900 idle streak reset after 12s: gpu=[2645538:mlxcel-bench-de] compilers=[] +2026-09-30T19:13:42+0900 idle streak reset after 10s: gpu=[2659792:mlxcel] compilers=[] +2026-09-30T19:14:46+0900 idle streak reset after 11s: gpu=[2700493:mlxcel] compilers=[] +2026-09-30T19:16:42+0900 idle streak reset after 50s: gpu=[] compilers=[2772984:cargo] +2026-09-30T19:19:20+0900 idle streak reset after 1s: gpu=[] compilers=[2876056:cargo,2877134:hipcc,2877135:hipcc,2877136:hipcc,2877140:hipcc,2877150:hipcc,2877151:hipcc,2877152:hipcc,2877160:hipcc,2877167:clang++,2877168:clang++,2877169:clang++,2877170:clang++,2877172:clang++,2877175:clang++,2877176:clang++,2877178:clang++,2877179:clang-23,2877180:clang-23,2877181:clang-23,2877182:clang-23,2877183:clang-23,2877184:clang-23,2877185:clang-23,2877186:clang-23] +2026-09-30T19:36:18+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T19:36:18+0900 attempt 1/5: start: target/release/mlxcel-bench-decode -m models/mlx/Meta-Llama-3.1-8B-Instruct-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T19:36:18+0900 sample 1: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:20+0900 sample 2: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:21+0900 sample 3: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:23+0900 sample 4: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:24+0900 sample 5: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:26+0900 sample 6: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:27+0900 sample 7: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:29+0900 sample 8: clean gpu=[3506769:mlxcel-bench-de] +2026-09-30T19:36:30+0900 monitor: 8 samples +2026-09-30T19:36:30+0900 attempt 1: CLEAN, exit 0 +== Meta-Llama-3.1-8B-Instruct-4bit_greedy profiled +2026-09-30T19:36:30+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T19:38:39+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T19:38:39+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /Meta-Llama-3.1-8B-Instruct-4bit_greedy -o Meta-Llama-3.1-8B-Instruct-4bit_greedy -- target/release/mlxcel-bench-decode -m models/mlx/Meta-Llama-3.1-8B-Instruct-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T19:38:40+0900 sample 1: clean gpu=[3594851:mlxcel-bench-de] +2026-09-30T19:38:41+0900 sample 2: clean gpu=[3594851:mlxcel-bench-de] +2026-09-30T19:38:43+0900 sample 3: clean gpu=[3594851:mlxcel-bench-de] +2026-09-30T19:38:44+0900 sample 4: clean gpu=[3594851:mlxcel-bench-de] +2026-09-30T19:38:46+0900 sample 5: clean gpu=[3594851:mlxcel-bench-de] +2026-09-30T19:38:47+0900 monitor: 5 samples +2026-09-30T19:38:47+0900 attempt 1: CLEAN, exit 0 +== Qwen3-30B-A3B-4bit_greedy plain +2026-09-30T19:38:47+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T19:40:56+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T19:40:56+0900 attempt 1/5: start: target/release/mlxcel-bench-decode -m models/mlx/Qwen3-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T19:40:56+0900 sample 1: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:40:58+0900 sample 2: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:40:59+0900 sample 3: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:01+0900 sample 4: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:02+0900 sample 5: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:04+0900 sample 6: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:05+0900 sample 7: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:06+0900 sample 8: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:08+0900 sample 9: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:09+0900 sample 10: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:11+0900 sample 11: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:12+0900 sample 12: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:14+0900 sample 13: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:15+0900 sample 14: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:17+0900 sample 15: clean gpu=[3679867:mlxcel-bench-de] +2026-09-30T19:41:18+0900 monitor: 15 samples +2026-09-30T19:41:18+0900 attempt 1: CLEAN, exit 0 +== Qwen3-30B-A3B-4bit_greedy profiled +2026-09-30T19:41:18+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T19:43:28+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T19:43:28+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /Qwen3-30B-A3B-4bit_greedy -o Qwen3-30B-A3B-4bit_greedy -- target/release/mlxcel-bench-decode -m models/mlx/Qwen3-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T19:43:28+0900 sample 1: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:30+0900 sample 2: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:31+0900 sample 3: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:33+0900 sample 4: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:34+0900 sample 5: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:36+0900 sample 6: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:37+0900 sample 7: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:39+0900 sample 8: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:40+0900 sample 9: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:42+0900 sample 10: clean gpu=[3776000:mlxcel-bench-de] +2026-09-30T19:43:43+0900 monitor: 10 samples +2026-09-30T19:43:43+0900 attempt 1: CLEAN, exit 0 +== granite-4.0-h-tiny-4bit_greedy plain +2026-09-30T19:43:43+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T19:45:53+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T19:45:53+0900 attempt 1/5: start: target/release/mlxcel-bench-decode -m models/mlx/granite-4.0-h-tiny-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T19:45:54+0900 sample 1: clean gpu=[3867773:mlxcel-bench-de] +2026-09-30T19:45:55+0900 sample 2: clean gpu=[3867773:mlxcel-bench-de] +2026-09-30T19:45:56+0900 sample 3: clean gpu=[3867773:mlxcel-bench-de] +2026-09-30T19:45:58+0900 sample 4: clean gpu=[3867773:mlxcel-bench-de] +2026-09-30T19:45:59+0900 sample 5: clean gpu=[3867773:mlxcel-bench-de] +2026-09-30T19:46:01+0900 sample 6: clean gpu=[3867773:mlxcel-bench-de] +2026-09-30T19:46:02+0900 monitor: 6 samples +2026-09-30T19:46:02+0900 attempt 1: CLEAN, exit 0 +== granite-4.0-h-tiny-4bit_greedy profiled +2026-09-30T19:46:02+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T19:48:09+0900 idle streak reset after 88s: gpu=[] compilers=[3952605:cargo] +2026-09-30T19:57:49+0900 idle streak reset after 11s: gpu=[115642:mlxcel] compilers=[] +2026-09-30T19:58:14+0900 idle streak reset after 11s: gpu=[131217:mlxcel] compilers=[] +2026-09-30T19:58:37+0900 idle streak reset after 11s: gpu=[146038:mlxcel] compilers=[] +2026-09-30T19:59:00+0900 idle streak reset after 11s: gpu=[160622:mlxcel] compilers=[] +2026-09-30T19:59:35+0900 idle streak reset after 10s: gpu=[183308:mlxcel] compilers=[] +2026-09-30T20:00:00+0900 idle streak reset after 11s: gpu=[199352:mlxcel] compilers=[] +2026-09-30T20:00:23+0900 idle streak reset after 11s: gpu=[213916:mlxcel] compilers=[] +2026-09-30T20:00:45+0900 idle streak reset after 10s: gpu=[228706:mlxcel] compilers=[] +2026-09-30T20:01:10+0900 idle streak reset after 11s: gpu=[244306:mlxcel] compilers=[] +2026-09-30T20:01:35+0900 idle streak reset after 12s: gpu=[260114:?] compilers=[] +2026-09-30T20:01:57+0900 idle streak reset after 10s: gpu=[275069:mlxcel] compilers=[] +2026-09-30T20:02:21+0900 idle streak reset after 11s: gpu=[290015:mlxcel] compilers=[] +2026-09-30T20:02:44+0900 idle streak reset after 11s: gpu=[305416:mlxcel] compilers=[] +2026-09-30T20:03:41+0900 idle streak reset after 33s: gpu=[] compilers=[341168:cargo] +2026-09-30T20:28:31+0900 idle streak reset after 6s: gpu=[1286793:logit_trace] compilers=[] +2026-09-30T20:59:38+0900 idle streak reset after 16s: gpu=[2549366:mlxcel_core-480] compilers=[] +2026-09-30T21:03:05+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:03:05+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /granite-4.0-h-tiny-4bit_greedy -o granite-4.0-h-tiny-4bit_greedy -- target/release/mlxcel-bench-decode -m models/mlx/granite-4.0-h-tiny-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T21:03:05+0900 sample 1: clean gpu=[2679179:mlxcel-bench-de] +2026-09-30T21:03:07+0900 sample 2: clean gpu=[2679179:mlxcel-bench-de] +2026-09-30T21:03:08+0900 sample 3: clean gpu=[2679179:mlxcel-bench-de] +2026-09-30T21:03:10+0900 sample 4: clean gpu=[2679179:mlxcel-bench-de,2682448:mlxcel-bench-de] +2026-09-30T21:03:11+0900 sample 5: CONTENDED foreign_gpu=[2682448:mlxcel-bench-de] compilers=[] +2026-09-30T21:03:13+0900 sample 6: clean gpu=[2679179:mlxcel-bench-de] +2026-09-30T21:03:14+0900 monitor: 6 samples +2026-09-30T21:03:14+0900 attempt 1: REJECTED (contended), exit 0; rerunning +2026-09-30T21:03:14+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:04:01+0900 idle streak reset after 32s: gpu=[2713606:mlxcel-bench-de] compilers=[] +2026-09-30T21:06:20+0900 idle streak reset after 42s: gpu=[2802676:mlxcel-bench-de] compilers=[] +2026-09-30T21:10:02+0900 idle streak reset after 9s: gpu=[] compilers=[2947450:cargo] +2026-09-30T21:10:43+0900 idle streak reset after 6s: gpu=[] compilers=[2973938:cargo] +2026-09-30T21:15:18+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:15:18+0900 attempt 2/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /granite-4.0-h-tiny-4bit_greedy -o granite-4.0-h-tiny-4bit_greedy -- target/release/mlxcel-bench-decode -m models/mlx/granite-4.0-h-tiny-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T21:15:19+0900 sample 1: clean gpu=[3151933:mlxcel-bench-de] +2026-09-30T21:15:20+0900 sample 2: clean gpu=[3151933:mlxcel-bench-de] +2026-09-30T21:15:22+0900 sample 3: clean gpu=[3151933:mlxcel-bench-de] +2026-09-30T21:15:23+0900 sample 4: clean gpu=[3151933:mlxcel-bench-de] +2026-09-30T21:15:25+0900 sample 5: clean gpu=[3151933:mlxcel-bench-de] +2026-09-30T21:15:27+0900 sample 6: clean gpu=[3151933:mlxcel-bench-de] +2026-09-30T21:15:28+0900 sample 7: clean gpu=[3151933:mlxcel-bench-de] +2026-09-30T21:15:29+0900 monitor: 7 samples +2026-09-30T21:15:29+0900 attempt 2: CLEAN, exit 0 +== NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy plain +2026-09-30T21:15:29+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:17:39+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:17:39+0900 attempt 1/5: start: target/release/mlxcel-bench-decode -m models/mlx/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T21:17:40+0900 sample 1: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:41+0900 sample 2: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:43+0900 sample 3: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:44+0900 sample 4: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:46+0900 sample 5: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:47+0900 sample 6: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:49+0900 sample 7: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:50+0900 sample 8: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:52+0900 sample 9: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:53+0900 sample 10: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:55+0900 sample 11: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:56+0900 sample 12: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:58+0900 sample 13: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:17:59+0900 sample 14: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:18:01+0900 sample 15: clean gpu=[3239105:mlxcel-bench-de] +2026-09-30T21:18:02+0900 monitor: 15 samples +2026-09-30T21:18:02+0900 attempt 1: CLEAN, exit 0 +== NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy profiled +2026-09-30T21:18:02+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:20:12+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:20:12+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy -o NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_greedy -- target/release/mlxcel-bench-decode -m models/mlx/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0 --top-p 1.0 +2026-09-30T21:20:12+0900 sample 1: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:14+0900 sample 2: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:15+0900 sample 3: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:17+0900 sample 4: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:18+0900 sample 5: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:20+0900 sample 6: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:21+0900 sample 7: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:22+0900 sample 8: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:24+0900 sample 9: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:26+0900 sample 10: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:27+0900 sample 11: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:29+0900 sample 12: clean gpu=[3334557:mlxcel-bench-de] +2026-09-30T21:20:30+0900 monitor: 12 samples +2026-09-30T21:20:30+0900 attempt 1: CLEAN, exit 0 +== Meta-Llama-3.1-8B-Instruct-4bit_t0.7 profiled +2026-09-30T21:21:37+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:23:41+0900 idle streak reset after 85s: gpu=[3438825:mlxcel_core-480] compilers=[] +2026-09-30T21:24:28+0900 idle streak reset after 29s: gpu=[3467939:mlxcel_core-480] compilers=[] +2026-09-30T21:25:50+0900 idle streak reset after 7s: gpu=[] compilers=[3518517:cargo] +2026-09-30T21:36:48+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:36:48+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /Meta-Llama-3.1-8B-Instruct-4bit_t0.7 -o Meta-Llama-3.1-8B-Instruct-4bit_t0.7 -- target/release/mlxcel-bench-decode -m models/mlx/Meta-Llama-3.1-8B-Instruct-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 1.0 +2026-09-30T21:36:48+0900 sample 1: clean gpu=[] +2026-09-30T21:36:50+0900 sample 2: clean gpu=[3927313:mlxcel-bench-de] +2026-09-30T21:36:51+0900 sample 3: clean gpu=[3927313:mlxcel-bench-de] +2026-09-30T21:36:53+0900 sample 4: clean gpu=[3927313:mlxcel-bench-de] +2026-09-30T21:36:54+0900 sample 5: clean gpu=[3927313:mlxcel-bench-de] +2026-09-30T21:36:55+0900 monitor: 5 samples +2026-09-30T21:36:55+0900 attempt 1: CLEAN, exit 0 +== Qwen3-30B-A3B-4bit_t0.7 profiled +2026-09-30T21:36:58+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:37:27+0900 idle streak reset after 20s: gpu=[3950066:mlxcel-bench-de] compilers=[] +2026-09-30T21:37:33+0900 idle streak reset after 1s: gpu=[3953075:mlxcel-bench-de] compilers=[] +2026-09-30T21:40:45+0900 idle streak reset after 1s: gpu=[4080625:mlxcel-bench-de] compilers=[] +2026-09-30T21:41:18+0900 idle streak reset after 1s: gpu=[4102225:mlxcel-bench-de] compilers=[] +2026-09-30T21:42:36+0900 idle streak reset after 43s: gpu=[4153829:mlxcel-bench-de] compilers=[] +2026-09-30T21:42:51+0900 idle streak reset after 2s: gpu=[4164387:mlxcel-bench-de] compilers=[] +2026-09-30T21:43:06+0900 idle streak reset after 1s: gpu=[4173488:mlxcel-bench-de] compilers=[] +2026-09-30T21:45:00+0900 idle streak reset after 58s: gpu=[56029:rocminfo] compilers=[55997:cargo] +2026-09-30T21:48:42+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:48:42+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /Qwen3-30B-A3B-4bit_t0.7 -o Qwen3-30B-A3B-4bit_t0.7 -- target/release/mlxcel-bench-decode -m models/mlx/Qwen3-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 1.0 +2026-09-30T21:48:42+0900 sample 1: clean gpu=[] +2026-09-30T21:48:44+0900 sample 2: clean gpu=[195614:mlxcel-bench-de] +2026-09-30T21:48:45+0900 sample 3: clean gpu=[195614:mlxcel-bench-de] +2026-09-30T21:48:47+0900 sample 4: clean gpu=[195614:mlxcel-bench-de] +2026-09-30T21:48:48+0900 sample 5: clean gpu=[195614:mlxcel-bench-de] +2026-09-30T21:48:50+0900 sample 6: clean gpu=[195614:mlxcel-bench-de] +2026-09-30T21:48:51+0900 monitor: 6 samples +2026-09-30T21:48:51+0900 attempt 1: CLEAN, exit 0 +== granite-4.0-h-tiny-4bit_t0.7 profiled +2026-09-30T21:49:00+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:51:11+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:51:11+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /granite-4.0-h-tiny-4bit_t0.7 -o granite-4.0-h-tiny-4bit_t0.7 -- target/release/mlxcel-bench-decode -m models/mlx/granite-4.0-h-tiny-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 1.0 +2026-09-30T21:51:11+0900 sample 1: clean gpu=[] +2026-09-30T21:51:13+0900 sample 2: clean gpu=[283739:mlxcel-bench-de] +2026-09-30T21:51:14+0900 sample 3: clean gpu=[283739:mlxcel-bench-de] +2026-09-30T21:51:16+0900 sample 4: clean gpu=[283739:mlxcel-bench-de] +2026-09-30T21:51:17+0900 sample 5: clean gpu=[283739:mlxcel-bench-de] +2026-09-30T21:51:19+0900 sample 6: clean gpu=[283739:mlxcel-bench-de] +2026-09-30T21:51:21+0900 sample 7: clean gpu=[283739:mlxcel-bench-de] +2026-09-30T21:51:22+0900 monitor: 7 samples +2026-09-30T21:51:22+0900 attempt 1: CLEAN, exit 0 +== NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7 profiled +2026-09-30T21:51:44+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:53:55+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T21:53:55+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7 -o NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7 -- target/release/mlxcel-bench-decode -m models/mlx/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 1.0 +2026-09-30T21:53:56+0900 sample 1: clean gpu=[] +2026-09-30T21:53:57+0900 sample 2: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:53:59+0900 sample 3: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:01+0900 sample 4: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:02+0900 sample 5: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:04+0900 sample 6: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:05+0900 sample 7: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:07+0900 sample 8: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:08+0900 sample 9: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:10+0900 sample 10: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:11+0900 sample 11: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:13+0900 sample 12: clean gpu=[373616:mlxcel-bench-de] +2026-09-30T21:54:14+0900 monitor: 12 samples +2026-09-30T21:54:14+0900 attempt 1: CLEAN, exit 0 +== Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95 profiled +2026-09-30T21:54:29+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T21:55:27+0900 idle streak reset after 40s: gpu=[] compilers=[421923:cargo] +2026-09-30T22:00:22+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T22:00:22+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95 -o Meta-Llama-3.1-8B-Instruct-4bit_t0.7-p0.95 -- target/release/mlxcel-bench-decode -m models/mlx/Meta-Llama-3.1-8B-Instruct-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 0.95 +2026-09-30T22:00:23+0900 sample 1: clean gpu=[] +2026-09-30T22:00:24+0900 sample 2: clean gpu=[] +2026-09-30T22:00:26+0900 sample 3: clean gpu=[615714:mlxcel-bench-de] +2026-09-30T22:00:28+0900 sample 4: clean gpu=[615714:mlxcel-bench-de] +2026-09-30T22:00:29+0900 sample 5: clean gpu=[615714:mlxcel-bench-de] +2026-09-30T22:00:31+0900 sample 6: clean gpu=[615714:mlxcel-bench-de] +2026-09-30T22:00:32+0900 sample 7: clean gpu=[615714:mlxcel-bench-de] +2026-09-30T22:00:34+0900 sample 8: clean gpu=[615714:mlxcel-bench-de] +2026-09-30T22:00:35+0900 sample 9: clean gpu=[615714:mlxcel-bench-de] +2026-09-30T22:00:36+0900 monitor: 9 samples +2026-09-30T22:00:36+0900 attempt 1: CLEAN, exit 0 +== Qwen3-30B-A3B-4bit_t0.7-p0.95 profiled +2026-09-30T22:00:40+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T22:02:50+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T22:02:50+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /Qwen3-30B-A3B-4bit_t0.7-p0.95 -o Qwen3-30B-A3B-4bit_t0.7-p0.95 -- target/release/mlxcel-bench-decode -m models/mlx/Qwen3-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 0.95 +2026-09-30T22:02:50+0900 sample 1: clean gpu=[] +2026-09-30T22:02:52+0900 sample 2: clean gpu=[706229:mlxcel-bench-de] +2026-09-30T22:02:53+0900 sample 3: clean gpu=[706229:mlxcel-bench-de] +2026-09-30T22:02:55+0900 sample 4: clean gpu=[706229:mlxcel-bench-de] +2026-09-30T22:02:56+0900 sample 5: clean gpu=[706229:mlxcel-bench-de] +2026-09-30T22:02:58+0900 sample 6: clean gpu=[706229:mlxcel-bench-de] +2026-09-30T22:02:59+0900 sample 7: clean gpu=[706229:mlxcel-bench-de] +2026-09-30T22:03:00+0900 monitor: 7 samples +2026-09-30T22:03:00+0900 attempt 1: CLEAN, exit 0 +== granite-4.0-h-tiny-4bit_t0.7-p0.95 profiled +2026-09-30T22:03:11+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T22:05:22+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T22:05:22+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /granite-4.0-h-tiny-4bit_t0.7-p0.95 -o granite-4.0-h-tiny-4bit_t0.7-p0.95 -- target/release/mlxcel-bench-decode -m models/mlx/granite-4.0-h-tiny-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 0.95 +2026-09-30T22:05:22+0900 sample 1: clean gpu=[] +2026-09-30T22:05:24+0900 sample 2: clean gpu=[795125:mlxcel-bench-de] +2026-09-30T22:05:25+0900 sample 3: clean gpu=[795125:mlxcel-bench-de] +2026-09-30T22:05:27+0900 sample 4: clean gpu=[795125:mlxcel-bench-de] +2026-09-30T22:05:28+0900 sample 5: clean gpu=[795125:mlxcel-bench-de] +2026-09-30T22:05:30+0900 sample 6: clean gpu=[795125:mlxcel-bench-de] +2026-09-30T22:05:32+0900 sample 7: clean gpu=[795125:mlxcel-bench-de] +2026-09-30T22:05:33+0900 sample 8: clean gpu=[795125:mlxcel-bench-de] +2026-09-30T22:05:34+0900 monitor: 8 samples +2026-09-30T22:05:34+0900 attempt 1: CLEAN, exit 0 +== NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95 profiled +2026-09-30T22:05:58+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T22:06:12+0900 idle streak reset after 9s: gpu=[] compilers=[810466:cargo,811295:rustc] +2026-09-30T22:07:07+0900 idle streak reset after 21s: gpu=[] compilers=[847093:cargo] +2026-09-30T22:11:17+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T22:11:17+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95 -o NVIDIA-Nemotron-3-Nano-30B-A3B-4bit_t0.7-p0.95 -- target/release/mlxcel-bench-decode -m models/mlx/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 0.95 +2026-09-30T22:11:18+0900 sample 1: clean gpu=[] +2026-09-30T22:11:19+0900 sample 2: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:20+0900 sample 3: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:22+0900 sample 4: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:23+0900 sample 5: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:25+0900 sample 6: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:26+0900 sample 7: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:27+0900 sample 8: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:29+0900 sample 9: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:30+0900 sample 10: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:32+0900 sample 11: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:33+0900 sample 12: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:35+0900 sample 13: clean gpu=[1010629:mlxcel-bench-de] +2026-09-30T22:11:36+0900 monitor: 13 samples +2026-09-30T22:11:36+0900 attempt 1: CLEAN, exit 0 +== rerun of Qwen3-30B-A3B-4bit_t0.7: the 21:48 attempt overlapped a CPU-heavy analysis process of this unit, which the guard does not watch; replaced +== Qwen3-30B-A3B-4bit_t0.7 profiled +2026-09-30T22:12:05+0900 waiting for 90s of idle GPU and no compiler +2026-09-30T22:14:17+0900 idle for 90s: kfd proc empty and no compiler in every 1 Hz sample +2026-09-30T22:14:17+0900 attempt 1/5: start: /usr/bin/rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv -d /Qwen3-30B-A3B-4bit_t0.7 -o Qwen3-30B-A3B-4bit_t0.7 -- target/release/mlxcel-bench-decode -m models/mlx/Qwen3-30B-A3B-4bit -p profile -n 128 --warmup-tokens 20 --ignore-eos --prompt-tokens 512 --temperature 0.7 --top-p 1.0 +2026-09-30T22:14:18+0900 sample 1: clean gpu=[] +2026-09-30T22:14:19+0900 sample 2: clean gpu=[1108964:mlxcel-bench-de] +2026-09-30T22:14:21+0900 sample 3: clean gpu=[1108964:mlxcel-bench-de] +2026-09-30T22:14:22+0900 sample 4: clean gpu=[1108964:mlxcel-bench-de] +2026-09-30T22:14:24+0900 sample 5: clean gpu=[1108964:mlxcel-bench-de] +2026-09-30T22:14:25+0900 sample 6: clean gpu=[1108964:mlxcel-bench-de] +2026-09-30T22:14:27+0900 sample 7: clean gpu=[1108964:mlxcel-bench-de] +2026-09-30T22:14:29+0900 sample 8: clean gpu=[1108964:mlxcel-bench-de] +2026-09-30T22:14:30+0900 monitor: 8 samples +2026-09-30T22:14:30+0900 attempt 1: CLEAN, exit 0 diff --git a/docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md b/docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md new file mode 100644 index 000000000..0676cc86d --- /dev/null +++ b/docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md @@ -0,0 +1,162 @@ +# ROCm decode profile: Radeon 8060S (gfx1151), 2026-09-30 + +Where ROCm decode time goes, per kernel, on four checkpoints, and which of the port issues split from #1814 would take the most of it (issue #2061). It turns #1814's call-frequency hypothesis into a measured order. The same session settles whether MLX's ROCm `gather_mm` works: it does, and its tests now run on ROCm. + +Raw results, one set per run (`benchmarks/rocm_profiles/gfx1151_929c80ab/`): + +- `_kernel_stats.csv`: rocprofv3's own `--stats` output for the whole process (load, warmup, prefill and decode); +- `_decode_kernels.csv`: the measured decode only, one row per kernel, with its class, role, port unit, calls per token and share of decode GPU time; +- `_summary.json`: the numbers below, the attribution per port unit, and the window checks; +- `_bench.log`, `_plain_bench.log`: the bench's output under the profiler and without it; +- `guard.log`: every idle-GPU check and 1 Hz sample behind every run. + +The full kernel traces (20 to 60 MB each) are not committed; `scripts/rocm_decode_profile.sh` regenerates all of the above. + +## Environment + +| Item | Value | +|------|-------| +| **Hardware** | AMD Ryzen AI MAX+ 395 with Radeon 8060S (`gfx1151`, RDNA 3.5), 32 CPU threads, 96 GiB VRAM carve-out (`mem_info_vram_total` 103079215104), 31 GiB visible to the host (`MemTotal` 32493820 kB) | +| **OS** | Debian GNU/Linux 13 (trixie), kernel 6.18.12+deb13-amd64 | +| **ROCm / HIP** | ROCm 10.0.0 (`/opt/rocm/core-10.0/.info/version`), HIP runtime 7.15.26333 (`hipconfig --version`) | +| **Profiler** | rocprofv3 1.3.5 (git `6b0e43f3`), `--kernel-trace --hip-graph-trace --stats -f csv` | +| **mlxcel** | 0.7.0, `main` at `929c80ab` plus this change's harness edits (`src/bin/bench_decode.rs` flags and phase marks, which run no kernel); built with `cargo build --release --features rocm --bin mlxcel-bench-decode` | +| **MLX pin** | `81ba1c6a` | +| **MLX ROCm overlay** | NripeshN/mlx `rocm-support` at `75915908` (`src/lib/mlx-cpp/patches-rocm/UPSTREAM`), plus `LOCAL_FIXES.md` items 1 to 26 | +| **Toolchain** | Rust 1.97.1 | + +## Method + +**Shape.** `scripts/bench_decode.sh`'s default: a deterministic 512-token prompt, exactly 128 generated tokens with every end-of-generation token suppressed, a discarded 20-token warmup in the same process, batch 1, default dense KV cache. Greedy for the main runs; two sampled variants (`--temperature 0.7`, and `--temperature 0.7 --top-p 0.95`) so the sampler shows up at all. + +**Idle GPU.** Every run went through `scripts/rocm_gpu_guard.sh`: 90 consecutive seconds with `/sys/class/kfd/kfd/proc` empty and no compiler process (`cargo`, `rustc`, `clang*`, `hipcc`, `cc1`, `cc1plus`, `ld*`, `lld`, `collect2`), then the run under a 1 Hz monitor, rejected and rerun if any sample showed another GPU process or a compiler. Other units and orchestrator gates shared the host, so the 16 accepted runs spread from 19:36 to 22:14 KST. One attempt was rejected by the guard (granite profiled, 21:03: another unit's bench started four seconds in). One accepted attempt was discarded by hand and rerun (Qwen3 `t0.7`, 21:48: a trace analysis of this unit was using a CPU core, which the guard does not watch). In every accepted run's window every sample showed either no GPU process or only the run itself, and no compiler; `guard.log` holds every sample. A GPU job shorter than one second could in principle fall between samples. + +**Decode window.** `MLXCEL_BENCH_PHASE_MARKS=1` makes the bench print, on both `CLOCK_MONOTONIC` and `CLOCK_BOOTTIME`, the start of the warmup, the start of the measured pass, the start of the measured decode and its end (`src/bin/bench_decode/phase_marks.rs`). The decode start is the end minus the generator's own `decode_time_ms`. `scripts/rocm_decode_profile.py` keeps the dispatches that start inside that window. Two checks say the cut is clean: the prefill ends in a blocking `eval` of the first token, and in every run no kernel straddles the cut and the device sits idle for 2 to 57 ms before the first decode dispatch; the last dispatch before the cut is always the first token's sampling (`arg_reduce_final` in greedy runs). rocprofv3 stamps dispatches on the same clock the marks use (both clocks agree on this host, which had not suspended). + +**Numbers.** Per token means per generated token, the denominator of the bench's tok/s. GPU time per token is the union of kernel intervals in the window; host gap per token is the window's wall time minus that. Shares are of the sum of kernel durations in the window (one queue, so union and sum agree to within 0.01%). + +**Profiler cost.** Each greedy run was also taken without the profiler, under the same guard. The profiler adds host time per dispatch and barely changes kernel durations, so it hurts in proportion to dispatch count: + +| Model | Dispatches per token | Decode tok/s, plain | Decode tok/s, profiled | Profiled slower by | +|---|---:|---:|---:|---:| +| `Meta-Llama-3.1-8B-Instruct-4bit` | 492 | 37.08 | 39.78 | -6.8% (faster; within run-to-run spread) | +| `Qwen3-30B-A3B-4bit` | 1540 | 61.99 | 58.86 | 5.3% | +| `granite-4.0-h-tiny-4bit` | 3197 | 61.35 | 49.86 | 23.0% | +| `NVIDIA-Nemotron-3-Nano-30B-A3B-4bit` | 2008 | 51.74 | 43.08 | 20.1% | + +Each figure is one plain and one profiled run, so the Llama reading says only that the cost there is below the noise (the #2056 baseline and `LOCAL_FIXES.md` item 24 read 35.4 to 35.9 tok/s for the same model). Read the shares, not the profiled absolute times; for the host gap, the plain run's wall time per token minus the traced GPU time is the better estimate and is given alongside. + +## Results (greedy) + +| Model | GPU ms/token | Host gap ms/token (profiled) | Host gap ms/token (plain wall minus GPU) | Dispatches/token | +|---|---:|---:|---:|---:| +| `Meta-Llama-3.1-8B-Instruct-4bit` (dense, f16) | 23.34 | 1.80 | 3.63 | 492 | +| `Qwen3-30B-A3B-4bit` (MoE, 128 experts, 8 active) | 13.36 | 3.63 | 2.77 | 1540 | +| `granite-4.0-h-tiny-4bit` (36 Mamba2 + 4 attention layers, MoE) | 11.47 | 8.59 | 4.83 | 3197 | +| `NVIDIA-Nemotron-3-Nano-30B-A3B-4bit` (23 Mamba2, 23 MoE, 6 attention) | 14.28 | 8.93 | 5.05 | 2008 | + +Top kernels by share of decode GPU time (calls per token): + +| Model | Kernels | +|---|---| +| Llama 3.1 8B | `qmv_wide_kernel` 94.8% (160), `kernel_sdpav_1pass` 2.1% (32), `rms_norm_kernel` 1.2% (64), `binary_vv` 0.6% (64), `copy_gg_byval` 0.5% (65), `rope_single_freqs_1d` 0.4% (64) | +| Qwen3-30B-A3B | `gather_qmv_wide_kernel` 42.1% (143), `qmv_wide_kernel` 33.1% (144), `kernel_sdpav_1pass` 4.6% (48), `block_sort_kernel` 4.3% (48, router top-k), `rms_norm_kernel` 3.9% (191), compiled SwiGLU 1.6% (48) | +| granite-4.0-h-tiny | `qmv_wide_kernel` 24.3% (168), `gather_qmv_wide_kernel` 21.9% (119), `binary_vv` 6.6% (143), `binary_g` 5.6% (178), `block_sort_kernel` 3.6% (40), `rms_norm_kernel` 3.4% (116) | +| Nemotron-3-Nano | `qmv_wide_kernel` 38.6% (116), `gather_qmv_wide_kernel` 27.9% (46), `binary_vv` 4.1% (92), `binary_g` 3.4% (114), `gemv_batched_inline` 2.2% (46), `copy_v` 2.1% (229) | + +What that says: + +- The dense model is weight-bandwidth bound and nothing in the port list touches it. Its GEMVs move about 4.2 GB per token in 22.1 ms, 192 GB/s. +- Qwen3's expert GEMVs are bandwidth bound too: 48 layers x 8 experts x 3 matrices of 2048 x 768 at 4.5 bits is 1.02 GB per token in 5.63 ms, 181 GB/s, about what the dense GEMVs reach (0.69 GB in 4.42 ms, 157 GB/s). +- The two hybrids are the ones leaving the GPU idle: 3197 and 2008 dispatches per token, a host gap of 4.8 and 5.1 ms per token without the profiler (about 1.5 and 2.5 us per dispatch), and a long tail of 1 to 2 us f32 elementwise kernels that are the Mamba2 SSD step. + +## Attribution + +A kernel name says which primitive ran, not which mlxcel op asked for it: a `binary_vv` can be a residual join, part of the SSD step or part of a logit bias. MLX evaluates each decode step's graph in the same order every step, so every op leaves the same run of dispatches, and `assign_roles()` in `scripts/rocm_decode_profile.py` labels those runs. Each rule was written against a printed decode step of the model it matters for and against the mlxcel code that builds the graph; `tests/test_rocm_decode_profile.py` pins them on synthetic steps. + +| Role | Rule | Built by | +|---|---|---| +| `sampler_tail` | everything after the step's widest `qmv` (the lm_head, vocab rows) up to the next step's embedding gather | `sample_token_optimized` and the `--ignore-eos` logit bias | +| `add_rms_join_post_attn` | a `binary_vv` immediately followed by `rms_norm_kernel`, where the add follows o_proj's `qmv`, which follows `kernel_sdpav_1pass` | `graph_add_rms_norm` (`layers.rs:978`), the fallback of `fused_add_rms_norm` | +| `add_rms_join` | the other add + `rms_norm_kernel` pairs (next layer's input norm) | model code, never routed to the fused kernel | +| `rope_append` | `rope_single*` dispatches and the `copy_gg_byval` K/V cache writes directly after them | the reshape / `fast_rope` / cache-append graph `forward_fused_rope_append` replaces | +| `moe_expert_gemv` | `gather_qmv_wide_kernel` | `gather_qmm` in `SwitchLinear::forward` | +| `moe_activation` | compiled, binary, copy and unary dispatches between the gate/up and down expert GEMVs, except `Divide` (score normalisation, which stays outside the fused kernel) | `compiled_swiglu_activation` (or relu2) and its casts | +| `moe_weighted_sum` | binary, copy and reduce dispatches right after the down GEMV, up to the residual add | `moe_weighted_sum` | +| `moe_gather_indices` | `arange_kernel` (one per `gather_qmm` call: 3 per layer in both SwitchGLU models) | the overlay's `gather_qmm` | +| `ssm_step` | inside a Mamba2 mixer (between the `qmv` before a `depthwise_conv1d_kernel`, in_proj, and the `qmv` after it, out_proj), everything that is not conv, SiLU or the gated norm | `ssm_step` (`granitemoehybrid.rs:474`, `nemotron_h.rs:672`), the graph `ssm_update_kernel` replaces | +| `ssm_conv`, `ssm_silu`, `ssm_gated_norm` | the conv1d and its state copies; the compiled SiLU kernels; the last `rms_norm_kernel` of the mixer and what follows it | the rest of the mixer, which the SSM kernel does not replace | + +Checks on the rules: the SSD step comes out at 46.6 dispatches per Mamba2 layer on granite and 49.6 on Nemotron, which run the same `ssm_step` graph; the MoE roles give Qwen3 exactly 3 expert GEMVs, 3 aranges and 4 weighted-sum dispatches per layer; every rule's dispatch count is a whole multiple of its layer count per token. Router top-k (`block_sort_kernel`, `softmax_kernel`) is deliberately left out of MoE, because `forward_fused_kernel` takes `topk_indices` and `scores` from the caller. No A/B was run: on ROCm none of these paths has a kernel to switch to, so there is nothing to toggle. + +Which of those roles each port would actually take over depends on whether mlxcel calls the ported kernel for that model with its shipped settings. `reach()` in the same script encodes that from the source at `929c80ab`: + +- **#2063** (`fused_add_rms_norm`, `fused_rope_qk_append`): both paths ship off on every backend (`FUSED_ADD_RMSNORM_DEFAULT` and `FUSED_ROPE_APPEND_DEFAULT` are `false`, `layers.rs:808` and `:820`, after #905 measured no decode win on Metal), and only `llama3.rs` (with `gemma.rs` and `iquestloopcoder.rs`, not profiled) calls them. With the opt-in on, Llama 3.1 would reach only the post-attention join (`llama3.rs:1201`): its `rope_scaling` builds a frequency table, which routes around the RoPE kernel (`llama3.rs:669`). +- **#2064** (`gumbel_max_sample`, `rejection_sample`): greedy argmax calls neither. A sampled run would reach the whole sampler tail (both kernels are on by default where ported), minus the few logit-bias dispatches `--ignore-eos` adds. +- **#2065** (`moe_gateup`, `moe_down`): reached by `qwen3_moe` (`qwen3_moe.rs:223`, single-token decode). Not by granite, whose `block_sparse_moe` calls `SwitchGLU::forward` and never `forward_fused_kernel`, and not by Nemotron-H, whose `fused_moe_forward` default branch is `gather_qmm`; its kernel branch needs `MLXCEL_FUSED_MOE_RELU2` and `moe_fc1_relu2` (#2069). +- **#2067** (`ssm_update_kernel`): reached by both hybrids, since `ssm_kernel_available()` gates every single-token SSD step (`granitemoehybrid.rs:440`, `nemotron_h.rs:537` and `:635`). +- **#2068** (paged attention): not on this path at all. The bench decodes one sequence into a dense `KVCache`; no paged kernel or paged graph fallback appears in any trace. + +## Share per port unit + +Share of decode GPU time the port would replace with shipped settings; in brackets, the cost of the fallback ops the port family covers whether or not mlxcel calls the port for that model. Dispatches per token replaced in the last column group. + +| Model, run | #2063 | #2064 | #2065 | #2067 | #2068 | Dispatches/token reached | +|---|---:|---:|---:|---:|---:|---| +| Llama 3.1 8B, greedy | 0 (1.66) | 0 (0.08) | 0 | 0 | 0 | none | +| Llama 3.1 8B, t0.7 | 0 (1.65) | 0.37 | 0 | 0 | 0 | #2064: 32 | +| Llama 3.1 8B, t0.7 p0.95 | 0 (1.92) | 1.17 | 0 | 0 | 0 | #2064: 64 | +| Qwen3-30B-A3B, greedy | 0 (2.82) | 0 (0.17) | **46.79** | 0 | 0 | #2065: 524 | +| Qwen3-30B-A3B, t0.7 | 0 (2.77) | 0.69 | 46.52 | 0 | 0 | #2065: 524, #2064: 32 | +| Qwen3-30B-A3B, t0.7 p0.95 | 0 (2.67) | 2.75 | 46.03 | 0 | 0 | #2065: 524, #2064: 63 | +| granite-4.0-h-tiny, greedy | 0 | 0 (0.20) | 0 (25.42) | **29.82** | 0 | #2067: 1679 | +| granite-4.0-h-tiny, t0.7 | 0 | 0.72 | 0 (22.51) | 30.75 | 0 | #2067: 1679, #2064: 33 | +| granite-4.0-h-tiny, t0.7 p0.95 | 0 | 2.51 | 0 (21.62) | 31.28 | 0 | #2067: 1679, #2064: 65 | +| Nemotron-3-Nano, greedy | 0 (0.26) | 0 (0.17) | 0 (29.22) | **20.01** | 0 | #2067: 1141 | +| Nemotron-3-Nano, t0.7 | 0 (0.26) | 0.68 | 0 (27.48) | 20.28 | 0 | #2067: 1141, #2064: 32 | +| Nemotron-3-Nano, t0.7 p0.95 | 0 (0.25) | 3.81 | 0 (25.99) | 21.50 | 0 | #2067: 1141, #2064: 64 | + +A share is an upper bound on what a port saves, since the ported kernel costs something too. How much of each share is recoverable differs sharply: + +- **#2067** replaces about 47 dispatches per layer, mostly 1 to 2 us f32 elementwise kernels, with one kernel whose memory traffic is the SSM state: 48 heads x 64 x 128 x 4 bytes, read and written, is 3.1 MB per granite layer, 113 MB per token, about 0.6 ms at the GEMVs' 180 GB/s, against the 3.42 ms per token the graph takes now (2.9 ms on Nemotron). It also removes 1679 and 1141 dispatches per token, more than half of each model's dispatches, which is where the hybrids' 5 ms per token host gap comes from. +- **#2065** covers 46.8% of Qwen3's GPU time, but 42.1 points of it are the expert GEMVs, already running at 181 GB/s. A fused kernel still reads the same weights, so what it can recover is the rest (activation, weighted sum, index building: 4.7%, 0.62 ms per token) plus about 430 of the 524 dispatches per token. That is less than on Metal and CUDA, where `gather_qmm` left the GPU idle (`switch_layers.rs`, `FUSED_MOE_MAX_DFF_METAL`). Granite and Nemotron carry 25 to 29% of fallback MoE cost that this port as scoped does not reach; wiring granite's `block_sparse_moe` through `forward_fused_kernel` would. +- **#2064** is zero in greedy decode, 0.4 to 0.7% with temperature alone, and 1.2 to 3.8% with top-p, where a full-vocabulary sort (`rocprim` radix sort, `block_sort_kernel`) and scans run every token; the top-p runs also add 0.7 to 1.4 ms per token of host gap over the temperature-only runs (profiled). +- **#2063** is zero with shipped settings on every backend. Turned on, the most it could reach here is Llama's post-attention join, 0.83% of decode. +- **#2068** has nothing to act on in single-stream decode. Its value is the batched paged serving path, which this profile does not measure, and the 36 paged-attention test skips on ROCm. + +## Ranked port order + +By share of decode GPU time reached with shipped settings, weighed by how much of it a kernel can recover: + +1. **#2067 SSM update** (#1814 item 7): 29.8% of granite and 20.0% of Nemotron decode GPU time, more than half of their dispatches, and mostly recoverable. +2. **#2065 fused MoE decode** (item 5): 46.8% of Qwen3 decode GPU time reached, but bandwidth-bound GEMVs are most of it; about 5% plus 430 dispatches per token recoverable. Worth more if granite's MoE is wired to it. +3. **#2064 samplers** (item 4): 0.4 to 3.8% of decode GPU time plus 32 to 64 dispatches per token in sampled decode, zero in greedy. +4. **#2068 paged attention** (item 8): no share in single-stream decode; ahead of #2063 only because it has a path (batched serving) and 36 skipped tests that this profile does not measure, not because of a measured share. +5. **#2063 fused add-RMSNorm and RoPE-append** (item 3): zero with shipped settings on every backend; at most 0.83% of Llama decode if opted in. + +This reverses #1814's order for items 3 and 7: implied order 7, 5, 4, 8, 3. + +## MLX's ROCm `gather_mm` + +`src/lib/mlxcel-core/src/grouped_gemm_numeric_tests.rs` gated its three tests on Metal or CUDA, so on ROCm they returned before touching the GPU. With the gates switched to `crate::gpu_backend_available()`, each was run by exact name on gfx1151 (`cargo test --release --features rocm -p mlxcel-core --lib grouped_gemm_numeric_tests:: -- --exact --test-threads=1 --nocapture`): + +| Test | Result | Kernels in its rocprofv3 trace | +|---|---|---| +| `gather_mm_matches_dense_per_expert_reference` | pass | `gather_batched_gemm_kernel` x4, `` x1, a Tensile `Cijk_*` GEMM x4 (hipBLASLt, the sorted single-row case) | +| `gather_mm_selects_the_indexed_expert` | pass | `gather_batched_gemm_kernel` | +| `gather_mm_half_precision_matches_reference` | pass | `gather_batched_gemm_kernel`, `gather_batched_gemm_kernel<__half, ...>` | + +So the overlay implements `GatherMM` itself (`GatherMM::eval_gpu`, `matmul.cpp`) and its results match the f64 host reference within the tests' tolerances. The tests do discriminate: with the reference deliberately pointed at the wrong expert, all three failed at the value assertion (`grouped_gemm_numeric_tests.rs:111`). The three gates now read `gpu_backend_available()` (Metal and CUDA still run them), the file left `BACKEND_ENUMERATION_TODO`, and `python3 scripts/ci/check_kernel_port_dispatch.py` reports `0 awaiting a predicate`. + +## Reproducing + +```bash +cargo build --release --features rocm --bin mlxcel-bench-decode +M="models/mlx/Meta-Llama-3.1-8B-Instruct-4bit models/mlx/Qwen3-30B-A3B-4bit models/mlx/granite-4.0-h-tiny-4bit models/mlx/NVIDIA-Nemotron-3-Nano-30B-A3B-4bit" +scripts/rocm_decode_profile.sh $M +scripts/rocm_decode_profile.sh --no-plain --temperature 0.7 $M +scripts/rocm_decode_profile.sh --no-plain --temperature 0.7 --top-p 0.95 $M +python3 scripts/rocm_decode_profile.py report benchmarks/rocm_profiles/gfx1151_ +``` + +`docs/benchmarks.md` ("ROCm per-kernel decode profile") describes the options. diff --git a/docs/installation.md b/docs/installation.md index 373bbf670..b1c5848b2 100644 --- a/docs/installation.md +++ b/docs/installation.md @@ -504,6 +504,9 @@ in [`docs/benchmark_results/rocm-baseline-gfx1151-2026-09-30.md`](benchmark_results/rocm-baseline-gfx1151-2026-09-30.md); `scripts/bench_decode.sh` recognises a ROCm host on its own (see [Benchmarks](benchmarks.md#rocm-hosts-issue-1810)). +Where that decode time goes, per kernel, and the order the #1814 kernel ports +should land in, is in +[`docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md`](benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md). #### Memory footprint From 523882d1cd4241c0a83c314a058f8b00ea9a6aec Mon Sep 17 00:00:00 2001 From: Jeongkyu Shin Date: Wed, 30 Sep 2026 23:24:01 +0900 Subject: [PATCH 4/6] docs(rocm): correct per-step counts in the gfx1151 decode profile The decode window holds 127 forward passes (the first of 128 tokens comes from the prefill), while per-token figures divide by 128 to match the bench's tok/s. The attribution check quoted 46.6 and 49.6 SSD dispatches per Mamba2 layer and claimed whole multiples "per token"; per step they are exactly 47 and 50, which is the check that actually holds. The Method section now states the 127/128 relation. Also: the #2063 opt-in ceiling is 0.83 to 0.89% (the top-p run reads 0.89), the #2064 reached share counts the whole sampler tail including the --ignore-eos logit bias, so it is slightly high rather than net of it, and a guard-rejected attempt stays in the run's _bench.log (granite greedy) while only the last, accepted attempt is read. Numbers recomputed from the committed summaries and decode CSVs; no GPU run. Refs #2061 --- .../rocm-decode-profile-gfx1151-2026-09-30.md | 12 +++++++----- 1 file changed, 7 insertions(+), 5 deletions(-) diff --git a/docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md b/docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md index 0676cc86d..a273745fb 100644 --- a/docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md +++ b/docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md @@ -10,6 +10,8 @@ Raw results, one set per run (`benchmarks/rocm_profiles/gfx1151_929c80ab/`): - `_bench.log`, `_plain_bench.log`: the bench's output under the profiler and without it; - `guard.log`: every idle-GPU check and 1 Hz sample behind every run. +When the guard rejects an attempt and reruns it, the run's `_bench.log` keeps both attempts' output (granite greedy); the last one is the accepted run, and it is the only one `rocm_decode_profile.py` reads and the one rocprofv3's files hold. + The full kernel traces (20 to 60 MB each) are not committed; `scripts/rocm_decode_profile.sh` regenerates all of the above. ## Environment @@ -33,7 +35,7 @@ The full kernel traces (20 to 60 MB each) are not committed; `scripts/rocm_decod **Decode window.** `MLXCEL_BENCH_PHASE_MARKS=1` makes the bench print, on both `CLOCK_MONOTONIC` and `CLOCK_BOOTTIME`, the start of the warmup, the start of the measured pass, the start of the measured decode and its end (`src/bin/bench_decode/phase_marks.rs`). The decode start is the end minus the generator's own `decode_time_ms`. `scripts/rocm_decode_profile.py` keeps the dispatches that start inside that window. Two checks say the cut is clean: the prefill ends in a blocking `eval` of the first token, and in every run no kernel straddles the cut and the device sits idle for 2 to 57 ms before the first decode dispatch; the last dispatch before the cut is always the first token's sampling (`arg_reduce_final` in greedy runs). rocprofv3 stamps dispatches on the same clock the marks use (both clocks agree on this host, which had not suspended). -**Numbers.** Per token means per generated token, the denominator of the bench's tok/s. GPU time per token is the union of kernel intervals in the window; host gap per token is the window's wall time minus that. Shares are of the sum of kernel durations in the window (one queue, so union and sum agree to within 0.01%). +**Numbers.** Per token means per generated token, the denominator of the bench's tok/s. The first of the 128 tokens comes out of the prefill, so the decode window holds 127 forward passes, and a per-token call count reads 127/128 of the per-step count (Qwen3's 144 expert GEMVs per step show as 142.9 per token). GPU time per token is the union of kernel intervals in the window; host gap per token is the window's wall time minus that. Shares are of the sum of kernel durations in the window (one queue, so union and sum agree to within 0.01%). **Profiler cost.** Each greedy run was also taken without the profiler, under the same guard. The profiler adds host time per dispatch and barely changes kernel durations, so it hurts in proportion to dispatch count: @@ -87,12 +89,12 @@ A kernel name says which primitive ran, not which mlxcel op asked for it: a `bin | `ssm_step` | inside a Mamba2 mixer (between the `qmv` before a `depthwise_conv1d_kernel`, in_proj, and the `qmv` after it, out_proj), everything that is not conv, SiLU or the gated norm | `ssm_step` (`granitemoehybrid.rs:474`, `nemotron_h.rs:672`), the graph `ssm_update_kernel` replaces | | `ssm_conv`, `ssm_silu`, `ssm_gated_norm` | the conv1d and its state copies; the compiled SiLU kernels; the last `rms_norm_kernel` of the mixer and what follows it | the rest of the mixer, which the SSM kernel does not replace | -Checks on the rules: the SSD step comes out at 46.6 dispatches per Mamba2 layer on granite and 49.6 on Nemotron, which run the same `ssm_step` graph; the MoE roles give Qwen3 exactly 3 expert GEMVs, 3 aranges and 4 weighted-sum dispatches per layer; every rule's dispatch count is a whole multiple of its layer count per token. Router top-k (`block_sort_kernel`, `softmax_kernel`) is deliberately left out of MoE, because `forward_fused_kernel` takes `topk_indices` and `scores` from the caller. No A/B was run: on ROCm none of these paths has a kernel to switch to, so there is nothing to toggle. +Checks on the rules, per decode step (127 in the window): the SSD step comes out at exactly 47 dispatches per Mamba2 layer on granite and 50 on Nemotron (46.6 and 49.6 per generated token), whose `ssm_step` graphs are built by separate but parallel code; the MoE roles give Qwen3 exactly 3 expert GEMVs, 3 aranges and 4 weighted-sum dispatches per layer; the SSM and MoE roles' dispatch counts are whole multiples of their layer counts per step. Router top-k (`block_sort_kernel`, `softmax_kernel`) is deliberately left out of MoE, because `forward_fused_kernel` takes `topk_indices` and `scores` from the caller. No A/B was run: on ROCm none of these paths has a kernel to switch to, so there is nothing to toggle. Which of those roles each port would actually take over depends on whether mlxcel calls the ported kernel for that model with its shipped settings. `reach()` in the same script encodes that from the source at `929c80ab`: - **#2063** (`fused_add_rms_norm`, `fused_rope_qk_append`): both paths ship off on every backend (`FUSED_ADD_RMSNORM_DEFAULT` and `FUSED_ROPE_APPEND_DEFAULT` are `false`, `layers.rs:808` and `:820`, after #905 measured no decode win on Metal), and only `llama3.rs` (with `gemma.rs` and `iquestloopcoder.rs`, not profiled) calls them. With the opt-in on, Llama 3.1 would reach only the post-attention join (`llama3.rs:1201`): its `rope_scaling` builds a frequency table, which routes around the RoPE kernel (`llama3.rs:669`). -- **#2064** (`gumbel_max_sample`, `rejection_sample`): greedy argmax calls neither. A sampled run would reach the whole sampler tail (both kernels are on by default where ported), minus the few logit-bias dispatches `--ignore-eos` adds. +- **#2064** (`gumbel_max_sample`, `rejection_sample`): greedy argmax calls neither. A sampled run would reach the whole sampler tail (both kernels are on by default where ported) except the few logit-bias dispatches `--ignore-eos` adds; the table counts the whole tail, so its #2064 figures are slightly high. - **#2065** (`moe_gateup`, `moe_down`): reached by `qwen3_moe` (`qwen3_moe.rs:223`, single-token decode). Not by granite, whose `block_sparse_moe` calls `SwitchGLU::forward` and never `forward_fused_kernel`, and not by Nemotron-H, whose `fused_moe_forward` default branch is `gather_qmm`; its kernel branch needs `MLXCEL_FUSED_MOE_RELU2` and `moe_fc1_relu2` (#2069). - **#2067** (`ssm_update_kernel`): reached by both hybrids, since `ssm_kernel_available()` gates every single-token SSD step (`granitemoehybrid.rs:440`, `nemotron_h.rs:537` and `:635`). - **#2068** (paged attention): not on this path at all. The bench decodes one sequence into a dense `KVCache`; no paged kernel or paged graph fallback appears in any trace. @@ -121,7 +123,7 @@ A share is an upper bound on what a port saves, since the ported kernel costs so - **#2067** replaces about 47 dispatches per layer, mostly 1 to 2 us f32 elementwise kernels, with one kernel whose memory traffic is the SSM state: 48 heads x 64 x 128 x 4 bytes, read and written, is 3.1 MB per granite layer, 113 MB per token, about 0.6 ms at the GEMVs' 180 GB/s, against the 3.42 ms per token the graph takes now (2.9 ms on Nemotron). It also removes 1679 and 1141 dispatches per token, more than half of each model's dispatches, which is where the hybrids' 5 ms per token host gap comes from. - **#2065** covers 46.8% of Qwen3's GPU time, but 42.1 points of it are the expert GEMVs, already running at 181 GB/s. A fused kernel still reads the same weights, so what it can recover is the rest (activation, weighted sum, index building: 4.7%, 0.62 ms per token) plus about 430 of the 524 dispatches per token. That is less than on Metal and CUDA, where `gather_qmm` left the GPU idle (`switch_layers.rs`, `FUSED_MOE_MAX_DFF_METAL`). Granite and Nemotron carry 25 to 29% of fallback MoE cost that this port as scoped does not reach; wiring granite's `block_sparse_moe` through `forward_fused_kernel` would. - **#2064** is zero in greedy decode, 0.4 to 0.7% with temperature alone, and 1.2 to 3.8% with top-p, where a full-vocabulary sort (`rocprim` radix sort, `block_sort_kernel`) and scans run every token; the top-p runs also add 0.7 to 1.4 ms per token of host gap over the temperature-only runs (profiled). -- **#2063** is zero with shipped settings on every backend. Turned on, the most it could reach here is Llama's post-attention join, 0.83% of decode. +- **#2063** is zero with shipped settings on every backend. Turned on, the most it could reach here is Llama's post-attention join, 0.83% of decode (0.89% in the top-p run). - **#2068** has nothing to act on in single-stream decode. Its value is the batched paged serving path, which this profile does not measure, and the 36 paged-attention test skips on ROCm. ## Ranked port order @@ -132,7 +134,7 @@ By share of decode GPU time reached with shipped settings, weighed by how much o 2. **#2065 fused MoE decode** (item 5): 46.8% of Qwen3 decode GPU time reached, but bandwidth-bound GEMVs are most of it; about 5% plus 430 dispatches per token recoverable. Worth more if granite's MoE is wired to it. 3. **#2064 samplers** (item 4): 0.4 to 3.8% of decode GPU time plus 32 to 64 dispatches per token in sampled decode, zero in greedy. 4. **#2068 paged attention** (item 8): no share in single-stream decode; ahead of #2063 only because it has a path (batched serving) and 36 skipped tests that this profile does not measure, not because of a measured share. -5. **#2063 fused add-RMSNorm and RoPE-append** (item 3): zero with shipped settings on every backend; at most 0.83% of Llama decode if opted in. +5. **#2063 fused add-RMSNorm and RoPE-append** (item 3): zero with shipped settings on every backend; at most 0.83 to 0.89% of Llama decode if opted in. This reverses #1814's order for items 3 and 7: implied order 7, 5, 4, 8, 3. From 0d5d6db09ae5fa797815cd12b8a738273242a01b Mon Sep 17 00:00:00 2001 From: Jeongkyu Shin Date: Wed, 30 Sep 2026 23:29:23 +0900 Subject: [PATCH 5/6] fix(rocm): harden the decode profile scripts after security review rocm_decode_profile.sh now scrubs local paths from the published guard.log from an EXIT trap, so an early exit (guard 75, rocprofv3 or cp failure) no longer leaves them in, and the replacement is literal so '#' and regex characters in the paths are harmless. rocm_decode_profile.py makes Dispatch slotted to bound the memory of a full kernel trace; the report output for benchmarks/rocm_profiles/gfx1151_929c80ab is byte-identical. rocm_gpu_guard.sh reads the parent pid after the last ')' of /proc//stat, avoids glob expansion when splitting holders, rejects non-integer --idle-secs, --max-attempts and --max-wait (and a missing option value) with exit 2, and stops the command and monitor on INT or TERM (exit 130 or 143). Tests cover the validation, octal-looking values, the SIGTERM path and slots. Validation: python3 -m unittest tests/test_rocm_decode_profile.py (17 pass), bash -n on both scripts, make verify-fmt verify-kernel-port-dispatch. No GPU workload was run. Refs #2061 --- scripts/rocm_decode_profile.py | 3 +- scripts/rocm_decode_profile.sh | 17 +++++++-- scripts/rocm_gpu_guard.sh | 57 +++++++++++++++++++++++++------ tests/test_rocm_decode_profile.py | 50 +++++++++++++++++++++++++++ 4 files changed, 113 insertions(+), 14 deletions(-) diff --git a/scripts/rocm_decode_profile.py b/scripts/rocm_decode_profile.py index 3b82d3b1a..8ea9256bd 100644 --- a/scripts/rocm_decode_profile.py +++ b/scripts/rocm_decode_profile.py @@ -105,7 +105,8 @@ def classify(name: str) -> str: # -------------------------------------------------------------------------- # Inputs # -------------------------------------------------------------------------- -@dataclass +# slots: a full kernel trace is millions of rows, so keep each one small. +@dataclass(slots=True) class Dispatch: name: str start: int diff --git a/scripts/rocm_decode_profile.sh b/scripts/rocm_decode_profile.sh index 1b4595854..c346184df 100755 --- a/scripts/rocm_decode_profile.sh +++ b/scripts/rocm_decode_profile.sh @@ -89,6 +89,21 @@ if [[ -z "$TAG" ]]; then fi GUARD_LOG="$OUT/guard.log" +# guard.log is published evidence; keep local paths out of it. Runs on every +# exit (guard exit 75, a rocprofv3 or cp failure, a signal), and replaces the +# paths literally so '#' and regex characters in them are harmless. +scrub_guard_log() { + [[ -f "$GUARD_LOG" ]] || return 0 + ROOT_PFX="$ROOT/" TRACE_PFX="$TRACE_DIR" python3 - "$GUARD_LOG" <<'PY' || true +import os, sys +path = sys.argv[1] +text = open(path, errors="surrogateescape").read() +text = text.replace(os.environ["ROOT_PFX"], "").replace(os.environ["TRACE_PFX"], "") +open(path, "w", errors="surrogateescape").write(text) +PY +} +trap scrub_guard_log EXIT + bench_args() { printf '%s\n' -m "$1" -p "profile" -n "$MAX_TOKENS" --warmup-tokens "$WARMUP_TOKENS" \ --ignore-eos --prompt-tokens "$PROMPT_TOKENS" --temperature "$TEMPERATURE" --top-p "$TOP_P" @@ -127,7 +142,5 @@ for model in "${MODELS[@]}"; do --temperature "$TEMPERATURE" \ || echo "summarize failed for $name; rerun it on $run_dir/${name}_kernel_trace.csv" >&2 done -# guard.log is published evidence; keep local paths out of it. -sed -i "s#${ROOT}/##g; s#${TRACE_DIR}##g" "$GUARD_LOG" echo "traces: $TRACE_DIR" >&2 echo "results: $OUT" >&2 diff --git a/scripts/rocm_gpu_guard.sh b/scripts/rocm_gpu_guard.sh index bf8b9dc27..03b2a196c 100755 --- a/scripts/rocm_gpu_guard.sh +++ b/scripts/rocm_gpu_guard.sh @@ -21,7 +21,9 @@ # evidence is the log itself. Exit status: COMMAND's status from the first clean # attempt; 75 when every attempt was contended or --max-wait ran out, in which # case nothing COMMAND printed should be used. COMMAND's stdout and stderr pass -# through unchanged, so callers capture them as usual. +# through unchanged, so callers capture them as usual. INT and TERM stop COMMAND +# and the monitor and exit 130 / 143. The three numeric options take plain +# non-negative integers. # # The sampling interval is one second: a GPU job shorter than that can in # principle be missed. The compiler list matches /proc//comm exactly. @@ -40,20 +42,34 @@ KFD_PROC_DIR="${ROCM_GPU_GUARD_KFD_DIR:-/sys/class/kfd/kfd/proc}" COMPILER_RE="${ROCM_GPU_GUARD_COMPILER_RE:-^(cargo|rustc|clang|clang\+\+|clang-[0-9]+|hipcc|nvcc|cc1|cc1plus|ld|ld\.lld|ld\.gold|ld\.bfd|lld|collect2)$}" usage() { - sed -n '2,27p' "$0" | sed 's/^# \{0,1\}//' + sed -n '2,29p' "$0" | sed 's/^# \{0,1\}//' } while [[ $# -gt 0 ]]; do case "$1" in - --idle-secs) IDLE_SECS="$2"; shift 2 ;; - --max-attempts) MAX_ATTEMPTS="$2"; shift 2 ;; - --max-wait) MAX_WAIT="$2"; shift 2 ;; - --log) LOG="$2"; shift 2 ;; + --idle-secs|--max-attempts|--max-wait|--log) + [[ $# -ge 2 ]] || { echo "rocm_gpu_guard: $1 needs a value" >&2; exit 2; } + case "$1" in + --idle-secs) IDLE_SECS="$2" ;; + --max-attempts) MAX_ATTEMPTS="$2" ;; + --max-wait) MAX_WAIT="$2" ;; + --log) LOG="$2" ;; + esac + shift 2 ;; -h|--help) usage; exit 0 ;; --) shift; break ;; *) echo "rocm_gpu_guard: unknown option $1" >&2; usage >&2; exit 2 ;; esac done +for opt in IDLE_SECS MAX_ATTEMPTS MAX_WAIT; do + if ! [[ "${!opt}" =~ ^[0-9]+$ ]]; then + flag="--$(tr 'A-Z_' 'a-z-' <<<"$opt")" + echo "rocm_gpu_guard: $flag needs a non-negative integer, got '${!opt}'" >&2 + exit 2 + fi +done +# Force base 10 so a value such as 08 is not read as octal. +IDLE_SECS=$((10#$IDLE_SECS)); MAX_ATTEMPTS=$((10#$MAX_ATTEMPTS)); MAX_WAIT=$((10#$MAX_WAIT)) [[ $# -gt 0 ]] || { echo "rocm_gpu_guard: no command given" >&2; usage >&2; exit 2; } [[ -d "$KFD_PROC_DIR" ]] || { echo "rocm_gpu_guard: $KFD_PROC_DIR not found (no ROCm KFD driver?)" >&2; exit 2; } @@ -93,10 +109,14 @@ compilers() { # True when $1 is $2 or a descendant of it. descends_from() { - local pid="$1" root="$2" ppid + local pid="$1" root="$2" ppid stat while [[ -n "$pid" && "$pid" != 0 && "$pid" != 1 ]]; do [[ "$pid" == "$root" ]] && return 0 - ppid=$(awk '{print $4}' "/proc/$pid/stat" 2>/dev/null) || return 1 + # comm (field 2) may contain spaces and parentheses; the fields that + # follow the last ')' are "state ppid ...". + stat=$(cat "/proc/$pid/stat" 2>/dev/null) || return 1 + stat="${stat##*) }" + read -r _ ppid _ <<<"$stat" pid="$ppid" done return 1 @@ -104,12 +124,13 @@ descends_from() { # Holders in $2 (a kfd_holders reading) that are not $1 or its descendants. foreign_holders() { - local root="$1" entry pid out=() - local IFS=, - for entry in $2; do + local root="$1" entry pid out=() entries + IFS=, read -r -a entries <<<"$2" + for entry in ${entries[@]+"${entries[@]}"}; do pid="${entry%%:*}" descends_from "$pid" "$root" || out+=("$entry") done + local IFS=, echo "${out[*]}" } @@ -134,6 +155,20 @@ wait_idle() { note "idle for ${quiet}s: kfd proc empty and no compiler in every 1 Hz sample" } +cmd_pid="" mon_pid="" +# On INT or TERM stop the command and the monitor instead of orphaning them. +on_signal() { + local sig="$1" code="$2" + trap - INT TERM + note "caught SIG${sig}: stopping command and monitor" + [[ -n "$mon_pid" ]] && kill "$mon_pid" 2>/dev/null || true + [[ -n "$cmd_pid" ]] && kill -TERM "$cmd_pid" 2>/dev/null || true + wait 2>/dev/null || true + exit "$code" +} +trap 'on_signal INT 130' INT +trap 'on_signal TERM 143' TERM + attempt=0 while (( attempt < MAX_ATTEMPTS )); do attempt=$((attempt + 1)) diff --git a/tests/test_rocm_decode_profile.py b/tests/test_rocm_decode_profile.py index 1d12f6369..00adba794 100644 --- a/tests/test_rocm_decode_profile.py +++ b/tests/test_rocm_decode_profile.py @@ -27,6 +27,7 @@ import subprocess import sys import tempfile +import time import unittest ROOT = pathlib.Path(__file__).resolve().parents[1] @@ -94,6 +95,50 @@ def test_a_compiler_counts_as_contention(self): self.assertEqual(r.returncode, 75, r.stderr) self.assertIn("gave up after", r.stderr) + def test_option_values_must_be_plain_non_negative_integers(self): + with tempfile.TemporaryDirectory() as kfd: + for flag, bad in (("--idle-secs", "abc"), ("--max-attempts", "-1"), + ("--max-wait", "1.5"), ("--idle-secs", "")): + r = run_guard(kfd, flag, bad, "--", "true") + self.assertEqual(r.returncode, 2, (flag, bad, r.stderr)) + self.assertIn(f"{flag} needs a non-negative integer", r.stderr) + r = run_guard(kfd, "--idle-secs") + self.assertEqual(r.returncode, 2) + self.assertIn("needs a value", r.stderr) + + def test_a_leading_zero_is_decimal_not_octal(self): + with tempfile.TemporaryDirectory() as kfd: + r = run_guard(kfd, "--idle-secs", "08", "--max-wait", "1", "--", "true") + self.assertEqual(r.returncode, 75, r.stderr) + self.assertNotIn("value too great", r.stderr) + + def test_sigterm_stops_the_command_and_the_guard(self): + with tempfile.TemporaryDirectory() as kfd, tempfile.TemporaryDirectory() as out: + pidfile = pathlib.Path(out) / "cmd.pid" + env = dict(os.environ, ROCM_GPU_GUARD_KFD_DIR=kfd, + ROCM_GPU_GUARD_COMPILER_RE=NO_COMPILERS) + p = subprocess.Popen( + ["bash", str(GUARD), "--idle-secs", "1", "--", "bash", "-c", + f"echo $$ > {pidfile}; exec sleep 30"], + env=env, stderr=subprocess.PIPE, text=True) + for _ in range(100): + if pidfile.exists() and pidfile.read_text().strip(): + break + time.sleep(0.1) + else: + p.kill() + self.fail("guarded command never started") + cmd_pid = int(pidfile.read_text()) + p.terminate() + _, err = p.communicate(timeout=30) + self.assertEqual(p.returncode, 143, err) + self.assertIn("caught SIGTERM", err) + for _ in range(50): + if not os.path.exists(f"/proc/{cmd_pid}"): + break + time.sleep(0.1) + self.assertFalse(os.path.exists(f"/proc/{cmd_pid}"), "command outlived the guard") + def k(op: str) -> str: """A kernel name spelled the way rocprofv3 writes the overlay's kernels.""" @@ -194,6 +239,11 @@ def test_busy_is_the_union_of_overlapping_dispatches(self): self.assertEqual(rdp.busy_ns(d, 0, 100), 25) self.assertEqual(rdp.busy_ns(d, 8, 25), 12) + def test_dispatch_is_slotted(self): + d = rdp.Dispatch("a", 1, 4) + self.assertFalse(hasattr(d, "__dict__")) + self.assertEqual(d.dur, 3) + def test_summarize_cuts_the_decode_window_by_the_phase_marks(self): with tempfile.TemporaryDirectory() as tmp: tmp = pathlib.Path(tmp) From 96e69c64c52cf0cc326832ef2ebe29e4b6b574bd Mon Sep 17 00:00:00 2001 From: Jeongkyu Shin Date: Wed, 30 Sep 2026 23:53:36 +0900 Subject: [PATCH 6/6] docs: add technical report for PR #2086 --- ...m-decode-profile-port-order-20260930.en.md | 203 ++++++++++++++++++ ...m-decode-profile-port-order-20260930.ko.md | 203 ++++++++++++++++++ 2 files changed, 406 insertions(+) create mode 100644 TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.en.md create mode 100644 TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.ko.md diff --git a/TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.en.md b/TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.en.md new file mode 100644 index 000000000..c93fefce4 --- /dev/null +++ b/TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.en.md @@ -0,0 +1,203 @@ +# Technical Report: PR #2086 - Profile gfx1151 decode per kernel and rank the #1814 ports + +**Date**: 2026-09-30 + +**Status**: Implemented and measured on the gfx1151 host; head `0d5d6db0` rebased onto origin/main `f9aefa39`, pending merge. + +**Languages**: Rust (bench binary flags and phase marks, test gates), Bash (idle-GPU guard, profile driver), Python (trace cut, attribution, report; CI checker), Markdown, CSV/JSON (committed profile data) + +**Risk Level**: Low (no inference path changes. The bench binary gains two flags and an opt-in env var whose defaults keep its output unchanged; three test gates widen from Metal-or-CUDA to any GPU backend; the rest is scripts, docs and data) + +## Executive Summary + +Issue #2061 (part of #1814, epic #1801) asked for the measurement that #1814 left as a hypothesis: where ROCm decode time goes per kernel, and in which order the five port issues split from #1814 should land. The PR adds a profiling harness, runs it on four checkpoints on a Radeon 8060S (gfx1151), publishes the result as `docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md` with its raw data under `benchmarks/rocm_profiles/gfx1151_929c80ab/`, and settles a second open question from the same session: MLX's ROCm `gather_mm` works, so its numeric tests now run on ROCm. + +The measured order is **#2067 > #2065 > #2064 > #2068 > #2063** (#1814 items 7, 5, 4, 8, 3). #1814's hypothesis, based on call frequency, put item 3 (fused add-RMSNorm and RoPE-append, which run per layer per token on every model) first and item 7 (SSM update, one model family) near the end. The profile reverses those two: #2067 reaches 29.8% of granite-4.0-h-tiny and 20.0% of Nemotron-3-Nano decode GPU time and more than half of their dispatches, while #2063 reaches 0% on every model because both of its fusions ship disabled on every backend. + +The report's other findings: dense decode is one kernel (`qmv_wide_kernel`, 94.8% of Llama 3.1 8B decode GPU time) running at about 192 GB/s; most of #2065's 46.8% share on Qwen3-30B-A3B is expert GEMVs already near that bandwidth, so only about 5% plus about 430 dispatches per token is recoverable; the profiler's own cost scales with dispatch count (within noise on Llama, 5.3% on Qwen3, 20 to 23% on the hybrids), which is why the doc reports shares and not times. + +## 1. Problem Statement + +### 1.1 An order without a measurement + +On 2026-09-30 #1814 was split into nine issues. Five of them (#2063, #2064, #2065, #2067, #2068) are HIP ports that fill a `.rocm = nullptr` slot in a `KernelPorts` table, and each depends on #2061. #1814 gave only a starting hypothesis for their order: `fused_add_rms_norm` and `fused_rope_qk_append` run per layer per token on every model, the two samplers once per token on every model, the MoE and SSM kernels per token only on their families, and paged attention only on the paged path. The issue said explicitly that the profile decides the order and that the frequency argument was for a reviewer to check, not to assume. + +The only ROCm decode data before this PR was the #2056 baseline, which has end-to-end tok/s (Llama-3.1-8B-4bit at 35.41 tok/s) but no per-kernel breakdown, so nothing said whether the gap to the GEMV bandwidth ceiling was kernel time, graph-fallback time or host time. + +### 1.2 The last `BACKEND_ENUMERATION_TODO` entry + +`grouped_gemm_numeric_tests.rs` gated its three tests on `!metal_is_available() && !cuda_is_available()`, so on ROCm they returned before touching the GPU. The tests exercise MLX's own `gather_mm`, not an mlxcel port, so no `KernelPorts` table applies, and the file was the last exemption from checker rule 4 in `scripts/ci/check_kernel_port_dispatch.py`. Widening the gate without running the tests was ruled out because that is how the #1806 abort had been reached. + +## 2. Change Summary + +Five commits: + +- **`0908c888`** `test(rocm): run MLX gather_mm numeric tests on every GPU backend`: the three gates read `crate::gpu_backend_available()`; `BACKEND_ENUMERATION_TODO` becomes an empty set. +- **`10bee856`** `feat(bench): add a ROCm per-kernel decode profile harness`: phase marks, `--temperature` / `--top-p`, the guard script, the profile driver and post-processor, unit tests, `docs/benchmarks.md`. +- **`c098e02d`** `docs(rocm): publish the gfx1151 decode profile and #1814 port order`: the results page and the 53 files under `benchmarks/rocm_profiles/gfx1151_929c80ab/`. +- **`523882d1`** `docs(rocm): correct per-step counts in the gfx1151 decode profile`: review follow-up (128 generated tokens against 127 forward passes in the window; exact per-step SSD counts of 47 and 50; #2063 opt-in ceiling widened to 0.83 to 0.89%). +- **`0d5d6db0`** `fix(rocm): harden the decode profile scripts after security review`: path scrubbing of `guard.log` in an EXIT trap, safer PID parsing and option validation in the guard, INT/TERM cleanup, a slotted `Dispatch` class. `rocm_decode_profile.py report` output on the committed data is byte-identical before and after. + +### 2.1 The harness + +- **Phase marks** (`src/bin/bench_decode/phase_marks.rs`). With `MLXCEL_BENCH_PHASE_MARKS=1` the bench prints four stderr lines, `warmup_start`, `measured_start`, `decode_start` and `measured_end`, each on both `CLOCK_MONOTONIC` and `CLOCK_BOOTTIME` (tracers differ in which one they stamp). `measured_end` is read before the trailing `synchronize_default()`, because the decode loop has already waited on its last token, and `decode_start` is derived as `measured_end` minus the generator's own `decode_time_ms`, since the loop's start is inside the generator. Unset, the bench's output is unchanged for `scripts/bench_decode.sh`. +- **`--temperature` and `--top-p`** on `mlxcel-bench-decode` (defaults 0.0 and 1.0, so greedy as before). They exist so a profile can see the sampler at all: greedy argmax never dispatches it. +- **`scripts/rocm_gpu_guard.sh`**: the #2056 idle-GPU check as a reusable script (section 4). +- **`scripts/rocm_decode_profile.sh`**: per model, one plain bench run and one under `rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv`, both through the guard, at `bench_decode.sh`'s default pp512/tg128 shape with a 20-token same-process warmup and `--ignore-eos`. Full traces (20 to 60 MB each) go to a temporary directory unless `--trace-dir` keeps them. +- **`scripts/rocm_decode_profile.py`**: cuts the trace to the measured decode by the phase marks, assigns roles to dispatches (`assign_roles()`), decides which roles each port reaches with shipped settings (`reach()`), and writes `_decode_kernels.csv` and `_summary.json`; its `report` subcommand renders the tables. `tests/test_rocm_decode_profile.py` pins the role rules on synthetic decode steps. + +### 2.2 Rebase onto `f9aefa39` + +PR #2084 had added per-phase `[Memory]` lines to `src/bin/bench_decode.rs` (issue #2062) in the same functions this PR edits. merge-resolver resolved the conflict keeping both: the memory counters print at each phase boundary, and the phase marks print around them (`reset_peak_memory()` then `warmup_start`; `measured_start`, then the measured pass, then `decode_start` / `measured_end`, then `print_memory_phase("after measured pass")`). + +## 3. Measured Results + +### 3.1 Method in brief + +Four checkpoints: `Meta-Llama-3.1-8B-Instruct-4bit` (dense, f16), `Qwen3-30B-A3B-4bit` (MoE, 128 experts, 8 active), `granite-4.0-h-tiny-4bit` (36 Mamba2 + 4 attention layers, MoE), `NVIDIA-Nemotron-3-Nano-30B-A3B-4bit` (23 Mamba2, 23 MoE, 6 attention). Each greedy, plus `--temperature 0.7` and `--temperature 0.7 --top-p 0.95`. ROCm 10.0.0, HIP 7.15.26333, rocprofv3 1.3.5, mlxcel `929c80ab` plus the harness edits, overlay `75915908` with `LOCAL_FIXES.md` items 1 to 26. + +Per token means per generated token (the bench's tok/s denominator). The first of the 128 tokens comes from prefill, so the window holds 127 forward passes and a per-token call count reads 127/128 of the per-step count (Qwen3's 144 expert GEMVs per step show as 142.9 per token). GPU time is the union of kernel intervals in the window; host gap is window wall time minus that. The cut is checked: in every run no kernel straddles it, the device is idle 2 to 57 ms before the first decode dispatch, and the last dispatch before it is always the first token's sampling. + +### 3.2 Per model (greedy) + +| Model | GPU ms/token | Host gap ms/token (profiled) | Host gap ms/token (plain wall minus GPU) | Dispatches/token | +|---|---:|---:|---:|---:| +| Llama 3.1 8B | 23.34 | 1.80 | 3.63 | 492 | +| Qwen3-30B-A3B | 13.36 | 3.63 | 2.77 | 1540 | +| granite-4.0-h-tiny | 11.47 | 8.59 | 4.83 | 3197 | +| Nemotron-3-Nano | 14.28 | 8.93 | 5.05 | 2008 | + +Top kernels by share of decode GPU time (calls per token): + +| Model | Kernels | +|---|---| +| Llama 3.1 8B | `qmv_wide_kernel` 94.8% (160), `kernel_sdpav_1pass` 2.1% (32), `rms_norm_kernel` 1.2% (64), `binary_vv` 0.6% (64), `copy_gg_byval` 0.5% (65), `rope_single_freqs_1d` 0.4% (64) | +| Qwen3-30B-A3B | `gather_qmv_wide_kernel` 42.1% (143), `qmv_wide_kernel` 33.1% (144), `kernel_sdpav_1pass` 4.6% (48), `block_sort_kernel` 4.3% (48, router top-k), `rms_norm_kernel` 3.9% (191), compiled SwiGLU 1.6% (48) | +| granite-4.0-h-tiny | `qmv_wide_kernel` 24.3% (168), `gather_qmv_wide_kernel` 21.9% (119), `binary_vv` 6.6% (143), `binary_g` 5.6% (178), `block_sort_kernel` 3.6% (40), `rms_norm_kernel` 3.4% (116) | +| Nemotron-3-Nano | `qmv_wide_kernel` 38.6% (116), `gather_qmv_wide_kernel` 27.9% (46), `binary_vv` 4.1% (92), `binary_g` 3.4% (114), `gemv_batched_inline` 2.2% (46), `copy_v` 2.1% (229) | + +Three readings follow from the tables: + +- **Dense decode is `qmv`.** Llama's GEMVs move about 4.2 GB per token in 22.1 ms, 192 GB/s. Nothing in the #1814 port list touches that kernel, so no port changes dense decode materially. +- **Qwen3's expert GEMVs are near bandwidth too.** 48 layers x 8 experts x 3 matrices of 2048 x 768 at 4.5 bits is 1.02 GB per token in 5.63 ms, 181 GB/s, about what its dense GEMVs reach (0.69 GB in 4.42 ms, 157 GB/s). The fused-MoE share is therefore mostly work a fused kernel still has to do. +- **The hybrids leave the GPU idle.** 3197 and 2008 dispatches per token, a plain-run host gap of 4.8 and 5.1 ms per token (about 1.5 and 2.5 us per dispatch), and a long tail of 1 to 2 us f32 elementwise kernels, which are the Mamba2 SSD step. + +### 3.3 Profiling overhead, and why the doc reports shares + +Each greedy run was also taken without the profiler under the same guard: + +| Model | Dispatches/token | Decode tok/s, plain | Decode tok/s, profiled | Profiled slower by | +|---|---:|---:|---:|---:| +| Llama 3.1 8B | 492 | 37.08 | 39.78 | -6.8% (faster; within run-to-run spread) | +| Qwen3-30B-A3B | 1540 | 61.99 | 58.86 | 5.3% | +| granite-4.0-h-tiny | 3197 | 61.35 | 49.86 | 23.0% | +| Nemotron-3-Nano | 2008 | 51.74 | 43.08 | 20.1% | + +rocprofv3 adds host time per dispatch and barely changes kernel durations, so its cost tracks dispatch count. On the hybrids, which are the models whose ranking depends on the host gap, the profiled wall time is 20 to 23% too slow, so profiled absolute times would overstate exactly the quantity under study. Shares of decode GPU time are robust to that, and for the host gap the doc prefers plain wall time per token minus traced GPU time (the third column of 3.2). Each overhead figure is one plain and one profiled run, so the Llama reading says only that the cost there is below noise (the #2056 baseline and `LOCAL_FIXES.md` item 24 read 35.4 to 35.9 tok/s for that model). + +## 4. The Idle-GPU Guard + +On a UMA host another GPU tenant or a compiler competes for the same memory bus, and this host was shared with other units and orchestrator gates during the session. `scripts/rocm_gpu_guard.sh [--idle-secs N] [--max-attempts N] [--max-wait SECS] [--log FILE] -- COMMAND`: + +1. waits for `--idle-secs` (default 90) consecutive seconds with `/sys/class/kfd/kfd/proc` empty and no compiler process (`/proc//comm` matched exactly against `cargo`, `rustc`, `clang*`, `hipcc`, `nvcc`, `cc1`, `cc1plus`, `ld*`, `lld`, `collect2`); +2. runs COMMAND under a 1 Hz monitor; +3. rejects the attempt if any sample shows a GPU process that is not COMMAND or a descendant, or a compiler, and goes back to step 1. + +Every sample goes to `--log`. Exit status is COMMAND's from the first clean attempt, or 75 when every attempt was contended or `--max-wait` ran out. INT and TERM stop COMMAND and the monitor. The driver scrubs local paths from the log on exit. + +The committed `benchmarks/rocm_profiles/gfx1151_929c80ab/guard.log` shows the guard doing its job. The granite greedy profiled attempt at 21:03:05 was clean for three samples; sample 4 showed a second `mlxcel-bench-de` process (another unit's bench, four seconds in), sample 5 logged `CONTENDED foreign_gpu=[2682448:mlxcel-bench-de]`, and the attempt ended `REJECTED (contended), exit 0; rerunning`. The bench itself exited 0, which is the point: the output looked valid and would have been published without the guard. After further idle-streak resets, attempt 2 started at 21:15:18 and ended `CLEAN, exit 0`. A different run's attempt (Qwen3 `t0.7` at 21:48) passed the guard but was discarded by hand because a CPU-heavy trace analysis of this unit ran alongside it, which the guard does not watch; the log records the replacement. The 16 accepted runs span 19:36 to 22:14 KST, and in each one every sample showed no GPU process or only the run itself, and no compiler. When a run is rerun, its `_bench.log` keeps both attempts and the post-processor reads only the last. + +## 5. Attribution and the Port Order + +### 5.1 From kernel names to ports + +A kernel name identifies the primitive, not the mlxcel op that asked for it: a `binary_vv` can be a residual join, part of the SSD step or part of a logit bias. Because MLX evaluates each decode step's graph in the same order every step, each op leaves the same run of dispatches, and `assign_roles()` labels those runs (`sampler_tail`, `add_rms_join_post_attn`, `add_rms_join`, `rope_append`, `moe_expert_gemv`, `moe_activation`, `moe_weighted_sum`, `moe_gather_indices`, `ssm_step`, `ssm_conv`, `ssm_silu`, `ssm_gated_norm`). The rules were checked by counts: exactly 47 SSD dispatches per Mamba2 layer per step on granite and 50 on Nemotron, whose `ssm_step` graphs are built by separate but parallel code; exactly 3 expert GEMVs, 3 aranges and 4 weighted-sum dispatches per Qwen3 layer; every SSM and MoE role count a whole multiple of its layer count. Router top-k is deliberately left out of MoE because `forward_fused_kernel` takes `topk_indices` and `scores` from the caller. + +`reach()` then encodes, from the source at `929c80ab`, whether mlxcel would actually call each port for that model with shipped settings. Each table cell is the share reached; the bracket is the fallback cost the port family covers whether or not it is called. + +### 5.2 Share per port under shipped settings + +| Model, run | #2063 | #2064 | #2065 | #2067 | #2068 | +|---|---:|---:|---:|---:|---:| +| Llama 3.1 8B, greedy | 0 (1.66) | 0 (0.08) | 0 | 0 | 0 | +| Llama 3.1 8B, t0.7 | 0 (1.65) | 0.37 | 0 | 0 | 0 | +| Llama 3.1 8B, t0.7 p0.95 | 0 (1.92) | 1.17 | 0 | 0 | 0 | +| Qwen3-30B-A3B, greedy | 0 (2.82) | 0 (0.17) | **46.79** | 0 | 0 | +| Qwen3-30B-A3B, t0.7 | 0 (2.77) | 0.69 | 46.52 | 0 | 0 | +| Qwen3-30B-A3B, t0.7 p0.95 | 0 (2.67) | 2.75 | 46.03 | 0 | 0 | +| granite-4.0-h-tiny, greedy | 0 | 0 (0.20) | 0 (25.42) | **29.82** | 0 | +| granite-4.0-h-tiny, t0.7 | 0 | 0.72 | 0 (22.51) | 30.75 | 0 | +| granite-4.0-h-tiny, t0.7 p0.95 | 0 | 2.51 | 0 (21.62) | 31.28 | 0 | +| Nemotron-3-Nano, greedy | 0 (0.26) | 0 (0.17) | 0 (29.22) | **20.01** | 0 | +| Nemotron-3-Nano, t0.7 | 0 (0.26) | 0.68 | 0 (27.48) | 20.28 | 0 | +| Nemotron-3-Nano, t0.7 p0.95 | 0 (0.25) | 3.81 | 0 (25.99) | 21.50 | 0 | + +A share is an upper bound on savings; how much of it a kernel can recover decides the order: + +1. **#2067 SSM update (item 7).** 29.8% (granite) and 20.0% (Nemotron) of decode GPU time, and 1679 and 1141 dispatches per token, more than half of each model's dispatches and the source of their 5 ms per token host gap. `ssm_kernel_available()` gates every single-token SSD step, so the port is reached. One kernel's traffic is the SSM state, about 113 MB per token on granite or about 0.6 ms at 180 GB/s, against the 3.42 ms the graph takes now (2.9 ms on Nemotron). Mostly recoverable. +2. **#2065 fused MoE decode (item 5).** 46.8% of Qwen3 decode GPU time is reached (`qwen3_moe.rs:223`), but 42.1 points are the expert GEMVs at 181 GB/s. What a fused kernel can recover is the activation, weighted sum and index building (4.7%, 0.62 ms per token) plus about 430 of the 524 dispatches per token. Granite and Nemotron carry 25 to 29% fallback MoE cost the port as scoped does not reach: granite's `block_sparse_moe` calls `SwitchGLU::forward`, not `forward_fused_kernel`, and Nemotron-H's default branch is `gather_qmm` (its kernel branch needs `MLXCEL_FUSED_MOE_RELU2` and #2069). +3. **#2064 samplers (item 4).** Zero in greedy decode, 0.4 to 0.7% with temperature, 1.2 to 3.8% with top-p, where a full-vocabulary `rocprim` radix sort and scans run each token. The figures include the few `--ignore-eos` logit-bias dispatches, so they read slightly high. +4. **#2068 paged attention (item 8).** No share: the bench decodes one sequence into a dense `KVCache`, and no paged kernel or paged fallback appears in any trace. +5. **#2063 fused add-RMSNorm and RoPE-append (item 3).** Zero with shipped settings on every backend: `FUSED_ADD_RMSNORM_DEFAULT` and `FUSED_ROPE_APPEND_DEFAULT` are `false` (`layers.rs:808`, `:820`) after #905 measured no decode win on Metal, and only `llama3.rs` (plus the unprofiled `gemma.rs` and `iquestloopcoder.rs`) calls them. Opted in, Llama 3.1 would reach only the post-attention join, 0.83% (0.89% in the top-p run), because its `rope_scaling` builds a frequency table that routes around the RoPE kernel (`llama3.rs:669`). + +**Why #2068 ranks above #2063 with no measured share.** Both are zero here, so the tie is broken on scope, not data. #2068 has a real path this profile does not cover (batched paged serving) and 36 ROCm test skips that its port clears. #2063 has nothing on any backend's default path, and its best case if someone opts in is under 1% of one model. The doc states this explicitly so the placement is not read as a measurement. + +**Against #1814's hypothesis.** The frequency argument was right that items 3 and 4 run on every model and item 7 on one family. It missed two things a profile shows immediately: a kernel that runs everywhere can still reach nothing if its caller ships it off, and per-token cost differs by more than an order of magnitude between a 47-dispatch f32 graph per layer (3.42 ms per token on granite) and a single add-plus-norm join (0.83% of Llama decode, about 0.19 ms). The ranking comment on #1814 records the new order, 7, 5, 4, 8, 3, against the old 3, 4, 5, 7, 8. + +## 6. MLX's ROCm `gather_mm` + +Each test was run on gfx1151 by exact name, and each also under rocprofv3 to confirm the GPU path: + +| Test | Result | Kernels in its trace | +|---|---|---| +| `gather_mm_matches_dense_per_expert_reference` | pass | `gather_batched_gemm_kernel` x4, `` x1, a Tensile `Cijk_*` GEMM x4 (hipBLASLt, the sorted single-row case) | +| `gather_mm_selects_the_indexed_expert` | pass | `gather_batched_gemm_kernel` | +| `gather_mm_half_precision_matches_reference` | pass | `gather_batched_gemm_kernel`, `gather_batched_gemm_kernel<__half, ...>` | + +The overlay implements `GatherMM::eval_gpu` itself (`matmul.cpp`) and matches the f64 host reference within the tests' tolerances. As a negative control, with the reference deliberately pointed at the wrong expert, all three failed at the value assertion (`grouped_gemm_numeric_tests.rs:111`), so a pass is evidence and not a skip. The gates now read `gpu_backend_available()` (Metal and CUDA still run them), `BACKEND_ENUMERATION_TODO` is empty, and the checker prints `0 awaiting a predicate`, which closes the last rule 4 exemption #1814 owned. + +## 7. Technical Decisions + +- **Cut the decode window by host clock, not by kernel name.** The warmup and measured passes run the same kernels, so kernel names alone cannot say which dispatches belong to the measured decode. Marks printed by the bench, on both clocks, plus two cleanliness checks (no straddling kernel, a device-idle gap before the first dispatch) make the cut verifiable per run. +- **Attribute by position in the step, and pin it with counts.** Role rules key off the graph order of each op and were checked against exact per-layer dispatch counts. The alternative, attributing by kernel name alone, would have charged residual adds and SSD adds to the same port. +- **Report reach under shipped settings, with the fallback cost in brackets.** A port that fills a `.rocm` slot nobody calls saves nothing. Separating "reached" from "covered" is what turns #2063 from a large bracketed number into zero and exposes granite's unwired MoE as a follow-up rather than a #2065 win. +- **Weigh shares by recoverability.** Ranking by raw share alone would put #2065 first. Bandwidth arithmetic on the GEMVs shows most of that share is not recoverable, which is why #2067 leads. +- **Shares, not times.** With profiler overhead up to 23% on the models that matter most, profiled absolute times would bias the ranking toward the dispatch-heavy hybrids. +- **Guard as a script with a log.** The #2056 guard was a procedure; making it a script with an exit code and a per-sample log makes every published run auditable and reusable by the port PRs that will need before/after numbers. +- **Run the `gather_mm` tests before widening the gate, with a negative control.** This follows the #1806 lesson directly: the gate change is backed by a trace and by a test mutation that must fail. + +## 8. Validation + +From the PR body, on gfx1151 before the rebase: + +- `cargo test --release --features rocm -p mlxcel-core --lib grouped_gemm_numeric_tests -- --test-threads=1 --nocapture`: 3 passed; each rerun by exact name under rocprofv3; all three fail with a wrong-expert reference. +- `python3 scripts/ci/check_kernel_port_dispatch.py`: `0 awaiting a predicate`; `make verify-versions verify-kernel-dtype-keys verify-kernel-port-dispatch verify-llama-compat verify-fmt` pass. +- `cargo clippy -p mlxcel --features rocm --bin mlxcel-bench-decode -- -D warnings` and `cargo clippy -p mlxcel-core --features rocm --lib --tests -- -D warnings`: clean. +- `cargo test --features rocm --test dead_doc_pointers`: pass. `python3 -m unittest tests/test_rocm_decode_profile.py`: 17 passed; `bash -n` on the three scripts. + +Orchestrator verification after the rebase onto `f9aefa39`: + +- merge-resolver resolved the `src/bin/bench_decode.rs` conflict with PR #2084's per-phase `[Memory]` lines and kept both behaviors. +- Clippy on the bench binary was clean; `tests/test_rocm_decode_profile.py` passed 17; the bench harness Python tests passed 35. +- A functional run of the rebased bench printed both the `[Memory]` lines and the four phase marks in order. +- The full `make verify-rocm` was run on the rebased head by the orchestrator. + +## 9. Residual Risks and What Was Not Verified + +- **One device, one session.** All numbers are from gfx1151 at `929c80ab` with overlay `75915908`. Other RDNA or CDNA parts, a different overlay, or a future `qmv` change can move the shares. The profile is reproducible with the committed scripts. +- **Overhead figures are single runs.** One plain and one profiled run per model; the Llama figure (profiled 6.8% faster) is noise, and the hybrids' 20 to 23% has no spread attached. +- **The guard's blind spots.** Sampling is 1 Hz, so a GPU job under a second can slip between samples, and CPU load from non-compiler processes is not watched (the Qwen3 `t0.7` rerun was a manual catch). +- **Shares are upper bounds, not speedups.** No A/B was run, because on ROCm none of these paths has a kernel to switch to. Each port PR has to measure its own before/after. +- **Attribution is rule-based.** The rules are pinned by per-layer counts and synthetic-step tests, but a model whose graph order differs from the four profiled here would need its own check. `reach()` encodes source at `929c80ab`; a later change to `FUSED_*_DEFAULT`, `block_sparse_moe` or `fused_moe_forward` changes the reach column. +- **#2068 is not measured.** Batched paged serving is outside this harness; its rank rests on scope, not data. +- **Not verified here:** Metal and CUDA (the widened `gather_mm` gates still run there, unchanged); `cargo test --test dead_doc_pointers` without `--features rocm` fails to link on this host (`copy_gpu_inplace` undefined from `kv_inplace_write.cpp`), unrelated to this change. + +## 10. Learning Points + +- **"Runs everywhere" is not "costs the most".** Call frequency ignores whether the fused path is enabled at all and how much work each call does. A profile with reach encoded answered in one pass what the frequency hypothesis had inverted. +- **A share needs a recoverability estimate before it becomes a priority.** Bandwidth arithmetic on the kernels a port replaces separates work that must still happen from overhead that can disappear. +- **Measure the profiler.** Tracing cost that scales with dispatch count skews exactly the dispatch-heavy models, so the unprofiled run is part of the method, not an extra. +- **A guard earns trust by rejecting something.** The committed log contains a real rejection of a run that exited 0, which is stronger evidence for the other 16 runs than a log of only clean samples. +- **Widen a gate with a negative control.** Passing tests plus a trace show the GPU path ran; a deliberately wrong reference that fails shows the tests can tell. + +Refs: #2061, #1814, #1801, #2056, #2062, #2063, #2064, #2065, #2067, #2068, #2069, #2084, #1806, #905. diff --git a/TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.ko.md b/TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.ko.md new file mode 100644 index 000000000..6ba2ed5ba --- /dev/null +++ b/TECHNICAL_REPORTS/2086-rocm-decode-profile-port-order-20260930.ko.md @@ -0,0 +1,203 @@ +# 기술 보고서: PR #2086 - gfx1151 decode를 kernel 단위로 profile하고 #1814 port 순서를 정함 + +**날짜**: 2026-09-30 + +**상태**: gfx1151 호스트에서 구현 및 측정 완료. head `0d5d6db0`, origin/main `f9aefa39` 위로 rebase됨, 머지 대기 중. + +**언어**: Rust (bench binary flag와 phase mark, test gate), Bash (idle-GPU guard, profile driver), Python (trace 절단, attribution, report, CI checker), Markdown, CSV/JSON (커밋된 profile 데이터) + +**위험도**: 낮음 (추론 경로는 바뀌지 않습니다. bench binary에 flag 두 개와 opt-in 환경 변수가 추가되며 기본값에서는 출력이 그대로입니다. test gate 세 개가 Metal 또는 CUDA에서 모든 GPU backend로 넓어지고, 나머지는 script, 문서, 데이터입니다) + +## 요약 + +Issue #2061 (#1814의 일부, epic #1801)은 #1814가 가설로 남겨 둔 측정을 요구했습니다. ROCm decode 시간이 kernel별로 어디에 쓰이는지, 그리고 #1814에서 분리된 다섯 개 port issue를 어떤 순서로 진행해야 하는지입니다. 이 PR은 profiling harness를 추가하고, Radeon 8060S (gfx1151)에서 checkpoint 네 개로 실행한 결과를 `docs/benchmark_results/rocm-decode-profile-gfx1151-2026-09-30.md`와 `benchmarks/rocm_profiles/gfx1151_929c80ab/`의 원자료로 공개합니다. 같은 세션에서 두 번째 미해결 질문도 정리했습니다. MLX의 ROCm `gather_mm`은 동작하므로, 그 numeric test가 이제 ROCm에서도 실행됩니다. + +측정된 순서는 **#2067 > #2065 > #2064 > #2068 > #2063** (#1814 항목 7, 5, 4, 8, 3)입니다. #1814의 가설은 호출 빈도를 근거로 항목 3 (모든 모델에서 layer마다 token마다 실행되는 fused add-RMSNorm과 RoPE-append)을 맨 앞에, 항목 7 (한 모델 계열에서만 쓰이는 SSM update)을 뒤쪽에 두었습니다. Profile은 이 둘을 뒤집습니다. #2067은 granite-4.0-h-tiny decode GPU 시간의 29.8%, Nemotron-3-Nano의 20.0%와 두 모델 dispatch의 절반 이상에 닿는 반면, #2063은 두 fusion 모두 모든 backend에서 꺼진 채로 배포되기 때문에 모든 모델에서 0%입니다. + +그 밖의 결과는 다음과 같습니다. Dense decode는 사실상 kernel 하나 (`qmv_wide_kernel`, Llama 3.1 8B decode GPU 시간의 94.8%)이고 약 192 GB/s로 동작합니다. Qwen3-30B-A3B에서 #2065가 차지하는 46.8%의 대부분은 이미 그 대역폭에 가까운 expert GEMV이므로, 회수 가능한 것은 약 5%와 token당 약 430개의 dispatch뿐입니다. Profiler 자체 비용은 dispatch 수에 비례하며 (Llama는 noise 수준, Qwen3는 5.3%, hybrid는 20~23%), 그래서 문서는 시간이 아니라 비율을 보고합니다. + +## 1. 문제 정의 + +### 1.1 측정 없는 순서 + +2026-09-30에 #1814가 아홉 개 issue로 나뉘었습니다. 그중 다섯 개 (#2063, #2064, #2065, #2067, #2068)는 `KernelPorts` table의 `.rocm = nullptr` 칸을 채우는 HIP port이고, 모두 #2061에 의존합니다. #1814는 이들의 순서에 대해 출발점이 되는 가설만 제시했습니다. `fused_add_rms_norm`과 `fused_rope_qk_append`는 모든 모델에서 layer마다 token마다, sampler 두 개는 모든 모델에서 token마다 한 번, MoE와 SSM kernel은 해당 계열에서만 token마다, paged attention은 paged 경로에서만 실행된다는 것입니다. Issue는 순서를 정하는 것은 profile이며, 빈도 논리는 reviewer가 검증할 대상이지 전제로 삼을 것이 아니라고 명시했습니다. + +이 PR 이전의 ROCm decode 데이터는 #2056 baseline뿐이었습니다. End-to-end tok/s (Llama-3.1-8B-4bit 35.41 tok/s)는 있었지만 kernel별 분해가 없어서, GEMV 대역폭 상한과의 차이가 kernel 시간인지, graph fallback 시간인지, host 시간인지 알 수 없었습니다. + +### 1.2 마지막 `BACKEND_ENUMERATION_TODO` 항목 + +`grouped_gemm_numeric_tests.rs`는 세 test를 `!metal_is_available() && !cuda_is_available()`로 막고 있어서, ROCm에서는 GPU를 건드리기 전에 반환했습니다. 이 test는 mlxcel port가 아니라 MLX 자체의 `gather_mm`을 검사하므로 적용할 `KernelPorts` table이 없고, 이 파일은 `scripts/ci/check_kernel_port_dispatch.py` checker rule 4의 마지막 예외였습니다. Test를 돌려 보지 않고 gate를 넓히는 것은 #1806 abort로 이어진 바로 그 방식이라 배제되었습니다. + +## 2. 변경 요약 + +Commit 다섯 개: + +- **`0908c888`** `test(rocm): run MLX gather_mm numeric tests on every GPU backend`: gate 세 개가 `crate::gpu_backend_available()`를 읽고, `BACKEND_ENUMERATION_TODO`는 빈 set이 됩니다. +- **`10bee856`** `feat(bench): add a ROCm per-kernel decode profile harness`: phase mark, `--temperature` / `--top-p`, guard script, profile driver와 후처리기, unit test, `docs/benchmarks.md`. +- **`c098e02d`** `docs(rocm): publish the gfx1151 decode profile and #1814 port order`: 결과 문서와 `benchmarks/rocm_profiles/gfx1151_929c80ab/` 아래 파일 53개. +- **`523882d1`** `docs(rocm): correct per-step counts in the gfx1151 decode profile`: review 반영 (생성 token 128개 대 window 안 forward pass 127회, SSD step당 정확한 dispatch 수 47과 50, #2063 opt-in 상한을 0.83~0.89%로 수정). +- **`0d5d6db0`** `fix(rocm): harden the decode profile scripts after security review`: EXIT trap에서 `guard.log`의 로컬 경로 제거, guard의 더 안전한 PID parsing과 option 검증, INT/TERM 정리, slot을 쓰는 `Dispatch` class. 커밋된 데이터에 대한 `rocm_decode_profile.py report` 출력은 변경 전후 byte 단위로 같습니다. + +### 2.1 Harness + +- **Phase mark** (`src/bin/bench_decode/phase_marks.rs`). `MLXCEL_BENCH_PHASE_MARKS=1`이면 bench가 stderr에 `warmup_start`, `measured_start`, `decode_start`, `measured_end` 네 줄을 `CLOCK_MONOTONIC`과 `CLOCK_BOOTTIME` 양쪽 값으로 출력합니다 (tracer마다 쓰는 clock이 다르기 때문입니다). `measured_end`는 마지막 `synchronize_default()` 전에 읽습니다. Decode loop가 이미 마지막 token을 기다렸기 때문입니다. `decode_start`는 loop의 시작이 generator 내부에 있으므로 `measured_end`에서 generator 자신의 `decode_time_ms`를 빼서 구합니다. 설정하지 않으면 `scripts/bench_decode.sh`가 보는 bench 출력은 바뀌지 않습니다. +- **`mlxcel-bench-decode`의 `--temperature`와 `--top-p`** (기본값 0.0과 1.0이므로 이전처럼 greedy). Greedy argmax는 sampler를 dispatch하지 않기 때문에, profile에 sampler가 보이게 하려고 추가했습니다. +- **`scripts/rocm_gpu_guard.sh`**: #2056의 idle-GPU 확인을 재사용 가능한 script로 만든 것 (4절). +- **`scripts/rocm_decode_profile.sh`**: 모델마다 plain bench 실행 한 번과 `rocprofv3 --kernel-trace --hip-graph-trace --stats -f csv` 아래 실행 한 번을 모두 guard를 거쳐 수행합니다. Shape는 `bench_decode.sh` 기본값인 pp512/tg128, 같은 process 안 20-token warmup, `--ignore-eos`입니다. 전체 trace (각 20~60 MB)는 `--trace-dir`로 보존하지 않으면 임시 디렉터리에 둡니다. +- **`scripts/rocm_decode_profile.py`**: phase mark로 trace에서 측정 decode 구간을 잘라내고, dispatch에 역할을 붙이고 (`assign_roles()`), 배포 설정에서 각 port가 어떤 역할에 닿는지 판정하고 (`reach()`), `_decode_kernels.csv`와 `_summary.json`을 씁니다. `report` subcommand가 표를 만듭니다. `tests/test_rocm_decode_profile.py`가 합성 decode step으로 역할 규칙을 고정합니다. + +### 2.2 `f9aefa39` 위로의 rebase + +PR #2084가 `src/bin/bench_decode.rs`의 같은 함수들에 phase별 `[Memory]` 출력 (issue #2062)을 추가했습니다. merge-resolver가 두 동작을 모두 유지하는 방향으로 충돌을 해결했습니다. 메모리 counter는 각 phase 경계에서 출력되고, phase mark는 그 주변에 출력됩니다 (`reset_peak_memory()` 다음 `warmup_start`, `measured_start` 다음 측정 pass, 그다음 `decode_start` / `measured_end`, 그다음 `print_memory_phase("after measured pass")`). + +## 3. 측정 결과 + +### 3.1 방법 요약 + +Checkpoint 네 개: `Meta-Llama-3.1-8B-Instruct-4bit` (dense, f16), `Qwen3-30B-A3B-4bit` (MoE, expert 128개 중 8개 활성), `granite-4.0-h-tiny-4bit` (Mamba2 36 + attention 4 layer, MoE), `NVIDIA-Nemotron-3-Nano-30B-A3B-4bit` (Mamba2 23, MoE 23, attention 6). 각각 greedy와 `--temperature 0.7`, `--temperature 0.7 --top-p 0.95`로 실행했습니다. ROCm 10.0.0, HIP 7.15.26333, rocprofv3 1.3.5, mlxcel `929c80ab`에 harness 수정을 더한 것, overlay `75915908`과 `LOCAL_FIXES.md` 항목 1~26입니다. + +Token당은 생성된 token당 (bench tok/s의 분모)을 뜻합니다. 128개 token 중 첫 번째는 prefill에서 나오므로 window에는 forward pass 127회가 들어 있고, token당 호출 수는 step당 호출 수의 127/128로 읽힙니다 (Qwen3의 step당 expert GEMV 144회는 token당 142.9회로 보입니다). GPU 시간은 window 안 kernel 구간의 합집합이고, host gap은 window wall time에서 그것을 뺀 값입니다. 절단은 검증됩니다. 모든 실행에서 절단점에 걸친 kernel이 없고, 첫 decode dispatch 전에 device가 2~57 ms 쉬며, 절단 직전 dispatch는 항상 첫 token의 sampling입니다. + +### 3.2 모델별 (greedy) + +| 모델 | GPU ms/token | Host gap ms/token (profiled) | Host gap ms/token (plain wall - GPU) | Dispatch/token | +|---|---:|---:|---:|---:| +| Llama 3.1 8B | 23.34 | 1.80 | 3.63 | 492 | +| Qwen3-30B-A3B | 13.36 | 3.63 | 2.77 | 1540 | +| granite-4.0-h-tiny | 11.47 | 8.59 | 4.83 | 3197 | +| Nemotron-3-Nano | 14.28 | 8.93 | 5.05 | 2008 | + +Decode GPU 시간 비율 기준 상위 kernel (token당 호출 수): + +| 모델 | Kernel | +|---|---| +| Llama 3.1 8B | `qmv_wide_kernel` 94.8% (160), `kernel_sdpav_1pass` 2.1% (32), `rms_norm_kernel` 1.2% (64), `binary_vv` 0.6% (64), `copy_gg_byval` 0.5% (65), `rope_single_freqs_1d` 0.4% (64) | +| Qwen3-30B-A3B | `gather_qmv_wide_kernel` 42.1% (143), `qmv_wide_kernel` 33.1% (144), `kernel_sdpav_1pass` 4.6% (48), `block_sort_kernel` 4.3% (48, router top-k), `rms_norm_kernel` 3.9% (191), compiled SwiGLU 1.6% (48) | +| granite-4.0-h-tiny | `qmv_wide_kernel` 24.3% (168), `gather_qmv_wide_kernel` 21.9% (119), `binary_vv` 6.6% (143), `binary_g` 5.6% (178), `block_sort_kernel` 3.6% (40), `rms_norm_kernel` 3.4% (116) | +| Nemotron-3-Nano | `qmv_wide_kernel` 38.6% (116), `gather_qmv_wide_kernel` 27.9% (46), `binary_vv` 4.1% (92), `binary_g` 3.4% (114), `gemv_batched_inline` 2.2% (46), `copy_v` 2.1% (229) | + +표에서 세 가지가 읽힙니다. + +- **Dense decode는 `qmv`입니다.** Llama의 GEMV는 token당 약 4.2 GB를 22.1 ms에 옮기므로 192 GB/s입니다. #1814 port 목록의 어떤 항목도 이 kernel을 건드리지 않으므로, 어떤 port도 dense decode를 크게 바꾸지 못합니다. +- **Qwen3의 expert GEMV도 대역폭에 가깝습니다.** 48 layer x expert 8개 x 2048 x 768 행렬 3개, 4.5 bit이면 token당 1.02 GB를 5.63 ms에 옮기므로 181 GB/s이고, dense GEMV (0.69 GB를 4.42 ms, 157 GB/s)와 비슷합니다. 따라서 fused MoE 비율의 대부분은 fused kernel도 여전히 해야 하는 일입니다. +- **Hybrid 모델이 GPU를 놀립니다.** Token당 dispatch 3197개와 2008개, plain 실행 기준 host gap token당 4.8 ms와 5.1 ms (dispatch당 약 1.5 us와 2.5 us), 그리고 1~2 us짜리 f32 elementwise kernel이 길게 이어지는데, 이것이 Mamba2 SSD step입니다. + +### 3.3 Profiling overhead와 비율로 보고하는 이유 + +각 greedy 실행은 같은 guard 아래에서 profiler 없이도 수행했습니다. + +| 모델 | Dispatch/token | Decode tok/s, plain | Decode tok/s, profiled | Profiled가 느린 정도 | +|---|---:|---:|---:|---:| +| Llama 3.1 8B | 492 | 37.08 | 39.78 | -6.8% (더 빠름, 실행 간 편차 이내) | +| Qwen3-30B-A3B | 1540 | 61.99 | 58.86 | 5.3% | +| granite-4.0-h-tiny | 3197 | 61.35 | 49.86 | 23.0% | +| Nemotron-3-Nano | 2008 | 51.74 | 43.08 | 20.1% | + +rocprofv3는 dispatch마다 host 시간을 더하고 kernel 실행 시간은 거의 바꾸지 않으므로, 비용이 dispatch 수를 따라갑니다. 순위가 host gap에 달려 있는 hybrid 모델에서 profiled wall time이 20~23% 느리므로, profiled 절대 시간은 연구 대상인 바로 그 양을 과대평가합니다. Decode GPU 시간의 비율은 이 영향을 받지 않고, host gap은 plain wall time per token에서 trace된 GPU 시간을 뺀 값 (3.2 표의 세 번째 열)을 우선합니다. Overhead 수치는 모델당 plain 한 번과 profiled 한 번이므로, Llama 값은 비용이 noise 이하라는 뜻일 뿐입니다 (#2056 baseline과 `LOCAL_FIXES.md` 항목 24는 이 모델에서 35.4~35.9 tok/s). + +## 4. Idle-GPU guard + +UMA 호스트에서는 다른 GPU 사용자나 compiler가 같은 메모리 bus를 다툽니다. 이 세션 동안 호스트는 다른 unit과 orchestrator gate와 공유되었습니다. `scripts/rocm_gpu_guard.sh [--idle-secs N] [--max-attempts N] [--max-wait SECS] [--log FILE] -- COMMAND`는 다음과 같이 동작합니다. + +1. `/sys/class/kfd/kfd/proc`이 비어 있고 compiler process (`/proc//comm`을 `cargo`, `rustc`, `clang*`, `hipcc`, `nvcc`, `cc1`, `cc1plus`, `ld*`, `lld`, `collect2`와 정확히 비교)가 없는 상태가 `--idle-secs` (기본 90)초 연속될 때까지 기다립니다. +2. 1 Hz monitor 아래에서 COMMAND를 실행합니다. +3. 어떤 sample에서든 COMMAND나 그 자손이 아닌 GPU process, 또는 compiler가 보이면 그 시도를 거부하고 1단계로 돌아갑니다. + +모든 sample은 `--log`에 기록됩니다. 종료 상태는 첫 깨끗한 시도의 COMMAND 종료 상태이거나, 모든 시도가 경합했거나 `--max-wait`가 끝나면 75입니다. INT와 TERM은 COMMAND와 monitor를 멈춥니다. Driver는 종료 시 log에서 로컬 경로를 지웁니다. + +커밋된 `benchmarks/rocm_profiles/gfx1151_929c80ab/guard.log`에 guard가 실제로 동작한 기록이 있습니다. 21:03:05의 granite greedy profiled 시도는 세 sample 동안 깨끗했고, sample 4에 두 번째 `mlxcel-bench-de` process (4초 뒤 시작된 다른 unit의 bench)가 나타났으며, sample 5가 `CONTENDED foreign_gpu=[2682448:mlxcel-bench-de]`를 기록했고, 시도는 `REJECTED (contended), exit 0; rerunning`으로 끝났습니다. Bench 자체는 0으로 종료했다는 점이 핵심입니다. 출력은 정상으로 보였고 guard가 없었다면 그대로 공개되었을 것입니다. 이후 idle streak가 몇 번 더 초기화된 뒤 21:15:18에 시도 2가 시작되어 `CLEAN, exit 0`으로 끝났습니다. 다른 실행의 시도 하나 (21:48의 Qwen3 `t0.7`)는 guard를 통과했지만, 이 unit의 CPU를 많이 쓰는 trace 분석이 함께 돌았기 때문에 (guard는 이를 감시하지 않습니다) 수동으로 폐기했고, log에 교체 사실이 남아 있습니다. 채택된 16개 실행은 19:36부터 22:14 KST까지 분포하며, 각 실행에서 모든 sample은 GPU process가 없거나 실행 자신만 있었고 compiler는 없었습니다. 재실행된 경우 `_bench.log`에는 두 시도가 모두 남고, 후처리기는 마지막 것만 읽습니다. + +## 5. Attribution과 port 순서 + +### 5.1 Kernel 이름에서 port로 + +Kernel 이름은 primitive를 알려 줄 뿐, 그것을 요청한 mlxcel op를 알려 주지 않습니다. `binary_vv`는 residual join일 수도, SSD step의 일부일 수도, logit bias의 일부일 수도 있습니다. MLX는 매 decode step의 graph를 같은 순서로 평가하므로 op마다 같은 dispatch 묶음을 남기고, `assign_roles()`가 이 묶음에 역할을 붙입니다 (`sampler_tail`, `add_rms_join_post_attn`, `add_rms_join`, `rope_append`, `moe_expert_gemv`, `moe_activation`, `moe_weighted_sum`, `moe_gather_indices`, `ssm_step`, `ssm_conv`, `ssm_silu`, `ssm_gated_norm`). 규칙은 개수로 검증했습니다. Step당 Mamba2 layer마다 SSD dispatch가 granite에서 정확히 47개, Nemotron에서 50개이고 (두 모델의 `ssm_step` graph는 별개지만 평행한 코드가 만듭니다), Qwen3 layer마다 expert GEMV 3개, arange 3개, weighted-sum dispatch 4개이며, 모든 SSM과 MoE 역할의 개수가 layer 수의 정수배입니다. `forward_fused_kernel`이 `topk_indices`와 `scores`를 호출자에게서 받으므로 router top-k는 의도적으로 MoE에서 뺐습니다. + +그다음 `reach()`가 `929c80ab` 소스를 기준으로, 배포 설정에서 mlxcel이 각 모델에 대해 실제로 그 port를 호출하는지를 반영합니다. 표의 각 칸은 닿는 비율이고, 괄호는 호출 여부와 무관하게 해당 port 계열이 덮는 fallback 비용입니다. + +### 5.2 배포 설정에서 port별 비율 + +| 모델, 실행 | #2063 | #2064 | #2065 | #2067 | #2068 | +|---|---:|---:|---:|---:|---:| +| Llama 3.1 8B, greedy | 0 (1.66) | 0 (0.08) | 0 | 0 | 0 | +| Llama 3.1 8B, t0.7 | 0 (1.65) | 0.37 | 0 | 0 | 0 | +| Llama 3.1 8B, t0.7 p0.95 | 0 (1.92) | 1.17 | 0 | 0 | 0 | +| Qwen3-30B-A3B, greedy | 0 (2.82) | 0 (0.17) | **46.79** | 0 | 0 | +| Qwen3-30B-A3B, t0.7 | 0 (2.77) | 0.69 | 46.52 | 0 | 0 | +| Qwen3-30B-A3B, t0.7 p0.95 | 0 (2.67) | 2.75 | 46.03 | 0 | 0 | +| granite-4.0-h-tiny, greedy | 0 | 0 (0.20) | 0 (25.42) | **29.82** | 0 | +| granite-4.0-h-tiny, t0.7 | 0 | 0.72 | 0 (22.51) | 30.75 | 0 | +| granite-4.0-h-tiny, t0.7 p0.95 | 0 | 2.51 | 0 (21.62) | 31.28 | 0 | +| Nemotron-3-Nano, greedy | 0 (0.26) | 0 (0.17) | 0 (29.22) | **20.01** | 0 | +| Nemotron-3-Nano, t0.7 | 0 (0.26) | 0.68 | 0 (27.48) | 20.28 | 0 | +| Nemotron-3-Nano, t0.7 p0.95 | 0 (0.25) | 3.81 | 0 (25.99) | 21.50 | 0 | + +비율은 절감의 상한입니다. 그중 kernel이 얼마를 회수할 수 있는지가 순서를 정합니다. + +1. **#2067 SSM update (항목 7).** Decode GPU 시간의 29.8% (granite)와 20.0% (Nemotron), 그리고 token당 dispatch 1679개와 1141개로, 각 모델 dispatch의 절반 이상이며 token당 5 ms host gap의 원천입니다. `ssm_kernel_available()`가 모든 single-token SSD step을 gate하므로 port는 호출됩니다. Kernel 하나의 메모리 traffic은 SSM state로, granite 기준 token당 약 113 MB, 180 GB/s에서 약 0.6 ms이고, 현재 graph는 3.42 ms (Nemotron 2.9 ms)를 씁니다. 대부분 회수 가능합니다. +2. **#2065 fused MoE decode (항목 5).** Qwen3 decode GPU 시간의 46.8%에 닿지만 (`qwen3_moe.rs:223`), 42.1%p가 181 GB/s로 도는 expert GEMV입니다. Fused kernel이 회수할 수 있는 것은 activation, weighted sum, index 생성 (4.7%, token당 0.62 ms)과 token당 dispatch 524개 중 약 430개입니다. Granite와 Nemotron은 현재 범위의 port가 닿지 않는 MoE fallback 비용 25~29%를 갖고 있습니다. Granite의 `block_sparse_moe`는 `forward_fused_kernel`이 아니라 `SwitchGLU::forward`를 부르고, Nemotron-H의 기본 분기는 `gather_qmm`입니다 (kernel 분기는 `MLXCEL_FUSED_MOE_RELU2`와 #2069가 필요합니다). +3. **#2064 sampler (항목 4).** Greedy decode에서는 0, temperature만 쓰면 0.4~0.7%, top-p를 쓰면 전체 vocabulary에 대한 `rocprim` radix sort와 scan이 매 token 돌아서 1.2~3.8%입니다. 이 수치에는 `--ignore-eos`가 더하는 logit-bias dispatch 몇 개가 포함되어 약간 높게 나옵니다. +4. **#2068 paged attention (항목 8).** 비율 없음. Bench는 한 sequence를 dense `KVCache`로 decode하며, 어떤 trace에도 paged kernel이나 paged fallback이 나타나지 않습니다. +5. **#2063 fused add-RMSNorm과 RoPE-append (항목 3).** 모든 backend에서 배포 설정으로는 0입니다. #905가 Metal에서 decode 이득이 없다고 측정한 뒤 `FUSED_ADD_RMSNORM_DEFAULT`와 `FUSED_ROPE_APPEND_DEFAULT`가 `false`이고 (`layers.rs:808`, `:820`), 이들을 부르는 것은 `llama3.rs` (그리고 profile하지 않은 `gemma.rs`, `iquestloopcoder.rs`)뿐입니다. Opt-in하면 Llama 3.1은 post-attention join 하나에만 닿아 0.83% (top-p 실행에서 0.89%)입니다. `rope_scaling`이 frequency table을 만들어 RoPE kernel을 우회하기 때문입니다 (`llama3.rs:669`). + +**측정된 비율 없이 #2068이 #2063보다 앞서는 이유.** 둘 다 여기서는 0이므로, 동률은 데이터가 아니라 범위로 판정했습니다. #2068은 이 profile이 다루지 않는 실제 경로 (batched paged serving)와 그 port가 해소할 ROCm test skip 36개를 갖고 있습니다. #2063은 어느 backend의 기본 경로에도 없고, 누군가 opt-in해도 최선이 한 모델의 1% 미만입니다. 문서는 이 배치가 측정값으로 읽히지 않도록 이를 명시합니다. + +**#1814 가설과의 비교.** 빈도 논리는 항목 3과 4가 모든 모델에서, 항목 7이 한 계열에서 돈다는 점에서는 맞았습니다. 그러나 profile이 바로 보여 주는 두 가지를 놓쳤습니다. 모든 곳에서 도는 kernel이라도 호출자가 꺼 둔 채 배포하면 아무 데도 닿지 않으며, token당 비용은 layer마다 47개 dispatch로 된 f32 graph (granite에서 token당 3.42 ms)와 add 하나에 norm 하나로 된 join (Llama decode의 0.83%, 약 0.19 ms) 사이에서 10배 이상 차이 납니다. #1814의 ranking comment는 새 순서 7, 5, 4, 8, 3을 이전 순서 3, 4, 5, 7, 8과 함께 기록합니다. + +## 6. MLX의 ROCm `gather_mm` + +각 test를 gfx1151에서 정확한 이름으로 실행했고, GPU 경로를 확인하려고 각각 rocprofv3 아래에서도 실행했습니다. + +| Test | 결과 | Trace에 나온 kernel | +|---|---|---| +| `gather_mm_matches_dense_per_expert_reference` | 통과 | `gather_batched_gemm_kernel` x4, `` x1, Tensile `Cijk_*` GEMM x4 (hipBLASLt, 정렬된 single-row 경우) | +| `gather_mm_selects_the_indexed_expert` | 통과 | `gather_batched_gemm_kernel` | +| `gather_mm_half_precision_matches_reference` | 통과 | `gather_batched_gemm_kernel`, `gather_batched_gemm_kernel<__half, ...>` | + +Overlay는 `GatherMM::eval_gpu` (`matmul.cpp`)를 직접 구현하며, 결과가 test 허용 오차 안에서 f64 host reference와 일치합니다. Negative control로 reference를 일부러 잘못된 expert로 향하게 하자 세 test 모두 값 assertion (`grouped_gemm_numeric_tests.rs:111`)에서 실패했으므로, 통과는 skip이 아니라 증거입니다. 이제 gate는 `gpu_backend_available()`를 읽고 (Metal과 CUDA에서도 그대로 실행됩니다), `BACKEND_ENUMERATION_TODO`는 비어 있으며, checker는 `0 awaiting a predicate`를 출력합니다. #1814가 맡았던 마지막 rule 4 예외가 이로써 닫힙니다. + +## 7. 기술적 선택과 그 이유 + +- **Kernel 이름이 아니라 host clock으로 decode 구간을 자릅니다.** Warmup pass와 측정 pass는 같은 kernel을 실행하므로, kernel 이름만으로는 어떤 dispatch가 측정 decode에 속하는지 알 수 없습니다. Bench가 두 clock으로 출력하는 mark와 두 가지 청결 검사 (절단점에 걸친 kernel 없음, 첫 dispatch 전 device idle 구간)가 실행마다 절단을 검증 가능하게 만듭니다. +- **Step 안의 위치로 역할을 붙이고 개수로 고정합니다.** 역할 규칙은 각 op의 graph 순서를 기준으로 하며 layer당 정확한 dispatch 수로 검증했습니다. Kernel 이름만으로 귀속시키면 residual add와 SSD add가 같은 port로 계산되었을 것입니다. +- **배포 설정에서의 도달을 보고하고 fallback 비용은 괄호에 둡니다.** 아무도 부르지 않는 `.rocm` 칸을 채우는 port는 아무것도 아끼지 못합니다. "닿음"과 "덮음"을 나눈 덕분에 #2063이 괄호 속 큰 숫자가 아니라 0이 되고, granite의 연결되지 않은 MoE가 #2065의 이득이 아니라 후속 작업으로 드러납니다. +- **비율을 회수 가능성으로 가중합니다.** 원래 비율만으로 순위를 매기면 #2065가 1위입니다. GEMV에 대한 대역폭 계산이 그 비율 대부분이 회수 불가능함을 보여 주므로 #2067이 앞섭니다. +- **시간이 아니라 비율.** 가장 중요한 모델에서 profiler overhead가 23%에 이르므로, profiled 절대 시간은 dispatch가 많은 hybrid 쪽으로 순위를 치우치게 했을 것입니다. +- **Log를 남기는 script로서의 guard.** #2056의 guard는 절차였습니다. 종료 코드와 sample별 log가 있는 script로 만들어, 공개된 모든 실행을 감사할 수 있고 before/after 수치가 필요할 port PR들이 재사용할 수 있습니다. +- **Gate를 넓히기 전에 `gather_mm` test를 negative control과 함께 실행합니다.** #1806의 교훈을 그대로 따른 것입니다. Gate 변경은 trace와, 반드시 실패해야 하는 test 변형으로 뒷받침됩니다. + +## 8. 검증 + +PR 본문 기준, rebase 전 gfx1151에서: + +- `cargo test --release --features rocm -p mlxcel-core --lib grouped_gemm_numeric_tests -- --test-threads=1 --nocapture`: 3개 통과. 각각 rocprofv3 아래에서 정확한 이름으로 재실행. 잘못된 expert reference로 세 개 모두 실패. +- `python3 scripts/ci/check_kernel_port_dispatch.py`: `0 awaiting a predicate`. `make verify-versions verify-kernel-dtype-keys verify-kernel-port-dispatch verify-llama-compat verify-fmt` 통과. +- `cargo clippy -p mlxcel --features rocm --bin mlxcel-bench-decode -- -D warnings`와 `cargo clippy -p mlxcel-core --features rocm --lib --tests -- -D warnings`: 경고 없음. +- `cargo test --features rocm --test dead_doc_pointers`: 통과. `python3 -m unittest tests/test_rocm_decode_profile.py`: 17개 통과. Script 세 개에 `bash -n`. + +`f9aefa39` 위로 rebase한 뒤 orchestrator 검증: + +- merge-resolver가 PR #2084의 phase별 `[Memory]` 출력과의 `src/bin/bench_decode.rs` 충돌을 해결했고 두 동작을 모두 유지했습니다. +- Bench binary clippy는 경고 없음, `tests/test_rocm_decode_profile.py`는 17개 통과, bench harness Python test는 35개 통과. +- Rebase된 bench의 기능 실행에서 `[Memory]` 출력과 phase mark 네 개가 순서대로 모두 출력되었습니다. +- 전체 `make verify-rocm`은 orchestrator가 rebase된 head에서 실행했습니다. + +## 9. 남은 위험과 검증하지 않은 부분 + +- **장치 하나, 세션 하나.** 모든 수치는 `929c80ab`, overlay `75915908`의 gfx1151 결과입니다. 다른 RDNA나 CDNA 장치, 다른 overlay, 향후 `qmv` 변경은 비율을 바꿀 수 있습니다. 커밋된 script로 profile을 재현할 수 있습니다. +- **Overhead 수치는 단일 실행입니다.** 모델당 plain 한 번과 profiled 한 번입니다. Llama 값 (profiled가 6.8% 빠름)은 noise이고, hybrid의 20~23%에는 편차 정보가 없습니다. +- **Guard의 사각지대.** Sampling이 1 Hz이므로 1초 미만의 GPU 작업은 sample 사이로 빠질 수 있고, compiler가 아닌 process의 CPU 부하는 감시하지 않습니다 (Qwen3 `t0.7` 재실행은 수동으로 잡은 것입니다). +- **비율은 speedup이 아니라 상한입니다.** ROCm에서는 이 경로들 중 어느 것도 전환할 kernel이 없으므로 A/B를 실행하지 않았습니다. 각 port PR이 자신의 before/after를 측정해야 합니다. +- **Attribution은 규칙 기반입니다.** 규칙은 layer당 개수와 합성 step test로 고정되어 있지만, graph 순서가 여기서 profile한 네 모델과 다른 모델은 별도 확인이 필요합니다. `reach()`는 `929c80ab`의 소스를 반영하므로, 이후 `FUSED_*_DEFAULT`, `block_sparse_moe`, `fused_moe_forward`가 바뀌면 도달 열이 바뀝니다. +- **#2068은 측정하지 않았습니다.** Batched paged serving은 이 harness 밖에 있으며, 그 순위는 데이터가 아니라 범위에 근거합니다. +- **여기서 검증하지 않은 것:** Metal과 CUDA (넓어진 `gather_mm` gate는 그곳에서도 그대로 실행됩니다). `--features rocm` 없이 `cargo test --test dead_doc_pointers`를 실행하면 이 호스트에서 link가 실패하며 (`kv_inplace_write.cpp`에서 `copy_gpu_inplace` 미정의), 이 변경과 무관합니다. + +## 10. 학습 포인트 + +- **"어디서나 돈다"가 "가장 비싸다"는 뜻은 아닙니다.** 호출 빈도는 fused 경로가 켜져 있는지, 호출 한 번이 얼마나 일하는지를 무시합니다. 도달 여부를 반영한 profile 한 번이 빈도 가설이 뒤집어 놓은 순서를 바로잡았습니다. +- **비율은 회수 가능성 추정을 거쳐야 우선순위가 됩니다.** Port가 대체하는 kernel에 대한 대역폭 계산이, 여전히 해야 하는 일과 사라질 수 있는 overhead를 구분합니다. +- **Profiler를 측정하십시오.** Dispatch 수에 비례하는 tracing 비용은 바로 dispatch가 많은 모델을 왜곡하므로, profiler 없는 실행은 부가 작업이 아니라 방법의 일부입니다. +- **Guard는 무언가를 거부할 때 신뢰를 얻습니다.** 커밋된 log에는 0으로 종료한 실행을 실제로 거부한 기록이 있고, 이는 깨끗한 sample만 있는 log보다 나머지 16개 실행에 대한 더 강한 증거입니다. +- **Gate는 negative control과 함께 넓히십시오.** 통과한 test와 trace는 GPU 경로가 실행되었음을 보여 주고, 일부러 틀린 reference가 실패하는 것은 test가 구분할 수 있음을 보여 줍니다. + +Refs: #2061, #1814, #1801, #2056, #2062, #2063, #2064, #2065, #2067, #2068, #2069, #2084, #1806, #905.