From 66c04d9565fbed677f05c179e7cb74918a212233 Mon Sep 17 00:00:00 2001 From: rkalmankar-crusoe Date: Fri, 25 Sep 2026 14:14:00 -0700 Subject: [PATCH 1/4] cmk-amd-mi355x validation suite --- cmk-amd-mi355x/.DS_Store | Bin 0 -> 6148 bytes cmk-amd-mi355x/validation-suite/.gitignore | 3 + cmk-amd-mi355x/validation-suite/README.md | 175 +++++++++++++ .../src/Dockerfile.a77-perftest-rocm | 63 +++++ .../src/build-kaniko-amdperftest.sh | 147 +++++++++++ .../env-verify/src/env-verify.sh | 241 ++++++++++++++++++ .../src/gpu-straggler-scan-mi355x.yaml | 120 +++++++++ .../src/gpu-to-nic-bw-gpu-direct.sh | 190 ++++++++++++++ .../src/gpu-to-nic-bw-hostmem.sh | 171 +++++++++++++ .../mirror-image/src/mirror-a77-to-ccr.yaml | 81 ++++++ .../src/deploy-rccl-allreduce.sh | 111 ++++++++ .../src/rccl-allreduce-mpijob-generic.yaml | 213 ++++++++++++++++ 12 files changed, 1515 insertions(+) create mode 100644 cmk-amd-mi355x/.DS_Store create mode 100644 cmk-amd-mi355x/validation-suite/.gitignore create mode 100644 cmk-amd-mi355x/validation-suite/README.md create mode 100644 cmk-amd-mi355x/validation-suite/build-image/src/Dockerfile.a77-perftest-rocm create mode 100755 cmk-amd-mi355x/validation-suite/build-image/src/build-kaniko-amdperftest.sh create mode 100755 cmk-amd-mi355x/validation-suite/env-verify/src/env-verify.sh create mode 100644 cmk-amd-mi355x/validation-suite/gpu-straggler-scan/src/gpu-straggler-scan-mi355x.yaml create mode 100755 cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-gpu-direct/src/gpu-to-nic-bw-gpu-direct.sh create mode 100755 cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-hostmem/src/gpu-to-nic-bw-hostmem.sh create mode 100644 cmk-amd-mi355x/validation-suite/mirror-image/src/mirror-a77-to-ccr.yaml create mode 100755 cmk-amd-mi355x/validation-suite/rccl-allreduce/src/deploy-rccl-allreduce.sh create mode 100644 cmk-amd-mi355x/validation-suite/rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml diff --git a/cmk-amd-mi355x/.DS_Store b/cmk-amd-mi355x/.DS_Store new file mode 100644 index 0000000000000000000000000000000000000000..4aa6a72899fe70ed8934829f6563f21d32226664 GIT binary patch literal 6148 zcmeHKyH3ME5S)b+k!V~}-Vadl2UZlmfFIyt3M2~`NvK`ryZAI_9|e)2ifE!)X>acK zcJ6djc)b8@a~SS{4#1l3h@%fn^L_V)T~)-0be=Kb8GF2A!p9=}_keRde3Cbk_mh8z z9S)4`@iy#U$CqguJy|9Nq<|EV0#ZNCq`)OA;NOQvckB!2#Q1b@ zh!%jjVmOTR=p~5F1H`^?PGp2;NhK!Ls>QIRGu|q%FPsyT4y)$F>Sn7B#o~6J-y$8> zCu)=eQs7j9>s)qT{~zdo^#7+Mt)zeyxF`i|wSC-f_@t_>i^qAbZS*I)=X}xKI1dVk mD96Mo$6R.ccr.crusoecloudcompute.com \ + --docker-username="" \ + --docker-password="" + ``` + + **Note on CCR docker Basic auth:** the username is the Crusoe user's + email address (not the user UUID or token ID). The password is a + registry token issued by `crusoe registry tokens create`. + +- Local tooling: `envsubst` (from `gettext`) for the RCCL manifest render step. + +--- + +## Ordered workflow + +> This suite assumes the CMK cluster and MI355X nodepool have **already been provisioned by Crusoe** — you receive a working kubeconfig, and the pre-flight in the Prerequisites section passes (2 × `Ready` nodes, each advertising `amd.com/gpu: 8` and `amd.com/vnic: 8`). + +Each step writes its outputs into its own `/logs/` (raw evidence) and +`/results/` (parsed one-page summary), created on first run. + +1. **Install the Kubeflow MPI Operator** — see Prerequisites (one-time setup, needed for the RCCL step). +2. **Create the CCR docker-registry secret** `ccr-cred` — see Prerequisites (one-time setup, needed for private images). +3. **(Optional) Mirror the Bundle 2.1 base image into your CCR** — + `mirror-image/src/mirror-a77-to-ccr.yaml` (or ask your Crusoe SE contact + to run `crusoe registry manifests copy` for you). Skip if you already + have the AMD workload image in your CCR. +4. **(Optional) Build the AMD-patched perftest image** — + `bash build-image/src/build-kaniko-amdperftest.sh`. + Only needed if you plan to run the `gpu-to-nic-bw-gpu-direct` check + with GPU-direct dma-buf. The other checks work against the upstream + `mirror.gcr.io/rocm/roce-workload:…-a-56` image. +5. **Environment verification** — + `bash env-verify/src/env-verify.sh`. + Produces `env-verify/logs/env-report--.txt`. Confirm the + observed versions match the Bundle 2.1 spec (see table above) and + that `bad_pages`, ECC counters, and Pollara VF state are clean. +6. **2-node RCCL all-reduce** — + + ```bash + IMAGE= PULL_SECRET=ccr-cred N=2 \ + bash rccl-allreduce/src/deploy-rccl-allreduce.sh + ``` + + Reference bar: **≥ 300 GB/s** busbw (dry-run observed 381.88 GB/s). +7. **Per-GPU compute + straggler scan** — + `kubectl apply -f gpu-straggler-scan/src/gpu-straggler-scan-mi355x.yaml` + (edit the nodeSelector line if your nodepool label differs). +8. **Per-rail host-memory bandwidth baseline** — + `bash gpu-to-nic-bw-hostmem/src/gpu-to-nic-bw-hostmem.sh`. + Reference bar: **≥ 500 Gb/s** per rail. +9. **Per-rail GPU-direct dma-buf bandwidth** — + `IMAGE= \ + bash gpu-to-nic-bw-gpu-direct/src/gpu-to-nic-bw-gpu-direct.sh`. + Reference bar: **≥ 700 Gb/s** per rail. + +--- + +## RCCL / NCCL environment + +The manifest at +[rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml](rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml) +combines the settings mandatory for MI355X + Pollara on CMK with the +mlcommons `AMDMi355xTrainingv6.1` bundle2 tuning envelope. The +load-bearing settings and their rationale: + +| Variable | Value | Rationale | +|---|---|---| +| `NCCL_TOPO_FILE` | `/etc/crusoe/rccl_topo/mi355x-288gb-ib.xml` | Cloud Hypervisor flattens the PCIe view, so RCCL cannot infer GPU ↔ NIC affinity on its own. The MI355X XML has device id `0x75a3` (gfx950); the MI350X `0x75a0` would silently mis-match and drop bandwidth. | +| `NCCL_DMABUF_ENABLE` | `1` | Kernel 6.8+ dropped `ib_peer_mem`. The module appears in `lsmod` on this bundle but is **not** registered as a peer-memory client by `amdgpu` — presence is not activation. dma-buf is the real GPU-direct path. | +| `NCCL_SOCKET_FAMILY` | `AF_INET` | Prevents NCCL from picking an IPv6 link-local on the RDMA `eth` interface and hanging rendezvous. | +| `NCCL_IB_HCA` | `ionic_0,…,ionic_7` | Explicit Pollara VF list — the 8 rails. | +| `NCCL_IB_GID_INDEX` | `1` | RoCEv2 GID index on Pollara. | +| `NCCL_IB_QPS_PER_CONNECTION` | `4` | Rail-optimised bandwidth ramp. | +| `NCCL_IB_TC` | `96` | RoCE traffic class (DSCP 48). | +| `NCCL_IGNORE_CPU_AFFINITY` | `1` | Cloud-Hypervisor synthetic CPU affinity is not authoritative. | +| `NCCL_SHM_DISABLE` | `1` | Forces inter-node traffic through the NIC path instead of SHM. | +| `HSA_NO_SCRATCH_RECLAIM` | `1` | ROCm 7.2 gfx950 workaround for a scratch-spill path. | +| `RCCL_AINIC_ROCE`, `IONIC_LOCKFREE`, `NCCL_GDR_FLUSH_DISABLE`, … | (~20 more) | mlcommons bundle2 tuning envelope. Full list is in the manifest. | + +--- + +## Adapting the suite for a different project + +The scripts read all tenant-specific values from environment variables +and enforce them with `${VAR:?…}` guards. To retarget the suite: + +| Variable | Where to find it | Used by | +|---|---|---| +| `PROJECT_ID` | `crusoe projects list` | mirror-image, build-image | +| `CCR_URL`, `CCR_REPO` (= `.`) | `crusoe registry list` | build-image, mirror-image | +| `IMAGE` | Full CCR image ref of the AMD workload image | rccl-allreduce, gpu-to-nic-bw-gpu-direct | +| `PULL_SECRET` | Name of the `docker-registry` secret you created | rccl-allreduce, gpu-to-nic-bw-gpu-direct | + +Everything else — Dockerfile, RCCL manifest, NCCL environment, target +thresholds, topology XML path — is invariant across tenants because it +is specific to the MI355X + Bundle 2.1 platform, not to any particular +project. diff --git a/cmk-amd-mi355x/validation-suite/build-image/src/Dockerfile.a77-perftest-rocm b/cmk-amd-mi355x/validation-suite/build-image/src/Dockerfile.a77-perftest-rocm new file mode 100644 index 0000000..b8a6511 --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/build-image/src/Dockerfile.a77-perftest-rocm @@ -0,0 +1,63 @@ +# syntax=docker/dockerfile:1 +# Bundle 2.1 a-77 image + AMD-patched perftest with --enable-rocm --enable-rocm-dmabuf. +# +# The stock perftest that ships in the a-77 base has --use_rocm / --use_rocm_dmabuf +# in its argparse table but the ROCm memory-type backend isn't registered in the +# binary — every invocation with those flags returns "Unsupported memory type". +# This image builds AMD's fork from source with the right configure flags so +# both flags work end-to-end for GPU-direct RDMA. +# +# Source: github.com/ROCm/rdma-perftest, branch master-dmabuf-rocm-20250114, +# commit 2db71141d6a3 "Enable dmabuf to ROCm and hipify changes" (2025-01-16). +# +# Build via in-cluster kaniko: +# CCR_REPO=. \ +# PROJECT_ID= \ +# bash build-kaniko-amdperftest.sh +# +# Runtime test: +# ib_write_bw -F -b --report_gbits -s 8M -n 5000 -x 1 -q 4 -m 4096 \ +# --use_rocm= --use_rocm_dmabuf -p [server_ip] +# +# The FROM below is the Bundle 2.1 a-77 base image in the DELIVERING PROJECT's +# CCR. `build-kaniko-amdperftest.sh` doesn't override this (kaniko reads the +# FROM literally), so you MUST edit this URL to point at your project's CCR +# path before running the build. See runbook.md for how to obtain the a-77 tag +# for your project (mirror from mcala-lab or ask your Crusoe SA). +ARG BASE_IMAGE=registry.us-east2-a.ccr.crusoecloudcompute.com/./roce-workload:ubuntu24_rocm-7.2_rccl-7.2.0_anp-v1.3.0_ainic-1.117.5-a-77-crusoe +FROM ${BASE_IMAGE} + +ENV DEBIAN_FRONTEND=noninteractive + +RUN apt-get update \ + && apt-get install -y --no-install-recommends \ + build-essential autoconf libtool git ca-certificates pkg-config \ + libibverbs-dev librdmacm-dev libibumad-dev libpci-dev libmnl-dev \ + && rm -rf /var/lib/apt/lists/* + +# Clone the exact branch + commit that Crusoe engineering used for the 740 Gb/s blog. +RUN cd /tmp \ + && git clone --branch master-dmabuf-rocm-20250114 --depth 1 https://github.com/ROCm/rdma-perftest.git \ + && cd rdma-perftest \ + && git rev-parse HEAD > /tmp/perftest-commit.txt \ + && cat /tmp/perftest-commit.txt + +# Configure with --enable-rocm --enable-rocm-dmabuf; install into /usr/local/bin so it +# takes precedence over the stock /usr/bin/ib_write_bw on PATH. +RUN cd /tmp/rdma-perftest \ + && ./autogen.sh \ + && export CFLAGS="-I/opt/rocm/include" \ + && export LDFLAGS="-L/opt/rocm/lib -L/opt/rocm/lib64 -Wl,-rpath=/opt/rocm/lib -lamdhip64 -lhsa-runtime64 -Wl,--copy-dt-needed-entries" \ + && ./configure --enable-rocm --enable-rocm-dmabuf --with-rocm=/opt/rocm \ + --prefix=/usr/local \ + && make -j$(nproc) \ + && make install \ + && ldconfig \ + && rm -rf /tmp/rdma-perftest + +# Sanity guards: fail the build if the new binary doesn't accept the ROCm flags. +RUN /usr/local/bin/ib_write_bw --help 2>&1 | grep -q "use_rocm" \ + && echo "OK: /usr/local/bin/ib_write_bw has --use_rocm support" + +# Ensure /usr/local/bin comes first in PATH for all shells +ENV PATH=/usr/local/bin:$PATH diff --git a/cmk-amd-mi355x/validation-suite/build-image/src/build-kaniko-amdperftest.sh b/cmk-amd-mi355x/validation-suite/build-image/src/build-kaniko-amdperftest.sh new file mode 100755 index 0000000..75677d4 --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/build-image/src/build-kaniko-amdperftest.sh @@ -0,0 +1,147 @@ +#!/usr/bin/env bash +# ============================================================================ +# build-kaniko-amdperftest.sh +# ---------------------------------------------------------------------------- +# Build the roce-workload:${TAG} image via Kaniko in-cluster and push to CCR. +# The Dockerfile derives from the a-77 base and adds AMD's fork of perftest +# (github.com/ROCm/rdma-perftest branch master-dmabuf-rocm-20250114, commit +# 2db71141d6a3) configured with --enable-rocm --enable-rocm-dmabuf so that +# ib_write_bw's --use_rocm=N --use_rocm_dmabuf flags actually work. +# +# CCR docker Basic-auth username must be the operator's email (not the user +# UUID or the token_id). The docker-registry secret `ccr-cred` must exist in +# the target namespace with that email as --docker-username. See runbook.md +# step 3 for the create command. +# +# Inputs (env): +# NS k8s namespace (default: default) +# CCR_URL CCR endpoint hostname (default: registry.us-east2-a.ccr.crusoecloudcompute.com) +# CCR_REPO ., e.g. amd-355x.4da9452b +# REGISTRY_NAME short registry name for CLI verify step, e.g. amd-355x +# (auto-derived from CCR_REPO if not set) +# PROJECT_ID Crusoe project UUID for the CLI verify step +# TAG image tag (default: a-77-crusoe-amdperftest) +# IMAGE_NAME image repo name in CCR (default: roce-workload) +# ============================================================================ + +set -euo pipefail + +NS="${NS:-default}" +CCR_URL="${CCR_URL:-registry.us-east2-a.ccr.crusoecloudcompute.com}" +: "${CCR_REPO:?CCR_REPO=. is required, e.g. amd-355x.4da9452b}" +REGISTRY_NAME="${REGISTRY_NAME:-${CCR_REPO%%.*}}" +: "${PROJECT_ID:?PROJECT_ID (Crusoe project UUID) is required for the verify step}" +IMAGE_NAME="${IMAGE_NAME:-roce-workload}" +TAG="${TAG:-a-77-crusoe-amdperftest}" +# The a-77 base tag to derive from. Must already exist in your CCR (see +# ../../mirror-image/src/mirror-a77-to-ccr.yaml to copy it from mcala-lab +# or ask your Crusoe SA for the equivalent tag they've published for you). +BASE_TAG="${BASE_TAG:-ubuntu24_rocm-7.2_rccl-7.2.0_anp-v1.3.0_ainic-1.117.5-a-77-crusoe}" +BASE_IMAGE="${BASE_IMAGE:-$CCR_URL/$CCR_REPO/$IMAGE_NAME:$BASE_TAG}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +DOCKERFILE="$SCRIPT_DIR/Dockerfile.a77-perftest-rocm" +LOG_DIR="$SCRIPT_DIR/../logs" +mkdir -p "$LOG_DIR" +LOG="$LOG_DIR/build-kaniko-amdperftest-$(date -u +%Y%m%dT%H%M%SZ).log" +TARGET="$CCR_URL/$CCR_REPO/$IMAGE_NAME:$TAG" + +echo "==> base : $BASE_IMAGE" +echo "==> target : $TARGET" +echo "==> log : $LOG" + +# Pre-flight +[[ -r "$DOCKERFILE" ]] || { echo "ERR: $DOCKERFILE not readable"; exit 1; } +kubectl -n "$NS" get secret ccr-cred >/dev/null 2>&1 \ + || { echo "ERR: docker-registry secret 'ccr-cred' missing in namespace $NS (see runbook.md step 3)"; exit 1; } + +# Load Dockerfile into a ConfigMap +kubectl -n "$NS" create configmap bundle2-dockerfile-amdperftest \ + --from-file=Dockerfile="$DOCKERFILE" \ + --dry-run=client -o yaml \ + | kubectl apply -f - > /dev/null + +# Clean any prior Job +kubectl -n "$NS" delete job build-amdperftest --wait=true 2>/dev/null || true + +# Apply the Job +cat < wait for kaniko pod" +POD="" +for i in $(seq 1 40); do + POD=$(kubectl -n "$NS" get pod -l job-name=build-amdperftest -o jsonpath='{.items[0].metadata.name}' 2>/dev/null || true) + if [[ -n "$POD" ]]; then + PHASE=$(kubectl -n "$NS" get pod "$POD" -o jsonpath='{.status.phase}' 2>/dev/null) + [[ "$PHASE" == "Running" || "$PHASE" == "Succeeded" || "$PHASE" == "Failed" ]] && break + fi + sleep 5 +done +echo " pod=$POD phase=$PHASE" + +echo "==> streaming kaniko logs -> $LOG" +kubectl -n "$NS" logs -f "$POD" > "$LOG" 2>&1 || true + +FINAL=$(kubectl -n "$NS" get pod "$POD" -o jsonpath='{.status.phase}' 2>/dev/null || echo "?") +echo "==> final phase=$FINAL" +tail -30 "$LOG" + +echo +echo "==> verify tag in CCR:" +crusoe registry manifests list "$REGISTRY_NAME" --repo-name "$IMAGE_NAME" --project-id "$PROJECT_ID" 2>&1 | head -8 + +if [[ "$FINAL" == "Succeeded" ]]; then + echo "====================================================================" + echo " PUSHED: $TARGET" + echo "====================================================================" +else + echo "ERR: kaniko Job did not succeed. Job kept for inspection:" + echo " kubectl -n $NS logs $POD" + exit 1 +fi diff --git a/cmk-amd-mi355x/validation-suite/env-verify/src/env-verify.sh b/cmk-amd-mi355x/validation-suite/env-verify/src/env-verify.sh new file mode 100755 index 0000000..ce4e457 --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/env-verify/src/env-verify.sh @@ -0,0 +1,241 @@ +#!/usr/bin/env bash +# ============================================================================ +# env-verify.sh +# ---------------------------------------------------------------------------- +# Environment sanity dump for both MI355X nodes. Deploys one privileged +# debug pod per node (upstream AMD roce-workload a-56 image, no auth needed) +# and captures: +# - Kernel version + relevant boot params +# - Loaded modules (amdgpu, ionic, ionic_rdma, dma-buf helpers, ib_peer_mem) +# - GPU health (amd-smi static + dynamic) +# - GPU-to-NIC PCIe affinity (lspci walk) +# - Every Pollara VF: ethtool -i, ethtool link state, MTU, ibv_devinfo +# - RoCE GID indices (show_gids) +# - dmesg errors last hour (amdgpu, ionic, ib*) +# - HugePages state +# - The topology XML we intend to reference is readable from a pod mount +# +# Output: results/env-report--.txt (one per node) + a summary index. +# +# Notes: +# * Uses the upstream `-a-56` image because our target `-a-77` isn't built +# yet AND all inspection binaries are identical between a-56/a-77. +# * Pods run with hostPID + hostNetwork + hostPath so we see host-level +# state, not per-container namespace. +# ============================================================================ + +set -euo pipefail + +NAMESPACE="${NAMESPACE:-default}" +IMAGE="${IMAGE:-mirror.gcr.io/rocm/roce-workload:ubuntu24_rocm-7.2_rccl-7.2.0_anp-v1.3.0_ainic-1.117.5-a-56}" +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +# Script now lives at /src/, so logs are ../logs relative to it +RESULTS_DIR="${RESULTS_DIR:-$SCRIPT_DIR/../logs}" +TS=$(date -u +%Y%m%dT%H%M%SZ) +mkdir -p "$RESULTS_DIR" + +# Auto-pick amd.com/gpu nodes. Portable to macOS bash 3.2 (no mapfile). +NODES=() +while IFS= read -r line; do + [[ -n "$line" ]] && NODES+=("$line") +done < <(kubectl get nodes -o jsonpath='{range .items[?(@.status.allocatable.amd\.com/gpu)]}{.metadata.name}{"\n"}{end}' 2>/dev/null) +if [[ "${#NODES[@]}" -lt 1 ]]; then + # fallback: just take all Ready nodes + while IFS= read -r line; do + [[ -n "$line" ]] && NODES+=("$line") + done < <(kubectl get nodes -o jsonpath='{range .items[?(@.status.conditions[?(@.type=="Ready")].status=="True")]}{.metadata.name}{"\n"}{end}' 2>/dev/null) +fi +echo "==> nodes: ${NODES[*]}" + +# ---- pod template ----------------------------------------------------------- +pod_yaml() { + local NAME=$1 NODE=$2 + cat < deploy $NAME on $NODE" + pod_yaml "$NAME" "$NODE" | kubectl -n "$NAMESPACE" apply -f - >/dev/null +done + +# ---- wait for Ready --------------------------------------------------------- +for NODE in "${NODES[@]}"; do + NAME="env-verify-${NODE%%.*}" + echo "==> wait Ready: $NAME" + kubectl -n "$NAMESPACE" wait --for=condition=Ready pod/"$NAME" --timeout=600s +done + +# ---- collect --------------------------------------------------------------- +collect_for_node() { + local NODE=$1 + local NAME="env-verify-${NODE%%.*}" + local OUT="$RESULTS_DIR/env-report-${NODE%%.*}-${TS}.txt" + + echo "==> collect from $NAME -> $OUT" + { + echo "# node : $NODE" + echo "# pod : $NAME" + echo "# image : $IMAGE" + echo "# timestamp : $TS" + echo + + echo "====================================================================" + echo " KERNEL" + echo "====================================================================" + kubectl -n "$NAMESPACE" exec "$NAME" -- bash -c ' + uname -a + echo "--- boot cmdline ---"; cat /host/proc/cmdline + echo "--- kernel taint ---"; cat /host/proc/sys/kernel/tainted + ' + + echo + echo "====================================================================" + echo " KERNEL MODULES" + echo "====================================================================" + kubectl -n "$NAMESPACE" exec "$NAME" -- bash -c ' + grep -E "^(amdgpu|amd_sched|amdttm|amdkcl|amddrm|ionic|ionic_rdma|ib_peer_mem|ib_uverbs|ib_core|mlx5|dma_buf|drm|xnack)" /host/proc/modules \ + | sort | awk "{printf \" %-24s size=%s used=%s by=%s\n\", \$1, \$2, \$3, \$4}" + ' + + echo + echo "====================================================================" + echo " GPU HEALTH (amd-smi)" + echo "====================================================================" + kubectl -n "$NAMESPACE" exec "$NAME" -- bash -c ' + amd-smi version | head -3 || true + echo + echo "--- amd-smi static (bus, subsys, vbios, firmware) ---" + amd-smi static -a --json 2>/dev/null | python3 -c " +import json,sys +d=json.loads(sys.stdin.read()) +for g in d: + print(f\"GPU {g.get(\\\"gpu\\\",\\\"?\\\")}: bus={g.get(\\\"bus\\\",{}).get(\\\"bdf\\\",\\\"?\\\")} vbios={g.get(\\\"vbios\\\",{}).get(\\\"version\\\",\\\"?\\\")}\") +" 2>/dev/null || amd-smi static -a 2>&1 | head -20 + echo + echo "--- amd-smi monitor (util, mem, temp, power) ---" + amd-smi monitor -u -m -t -p 2>&1 | head -30 + echo + echo "--- amd-smi bad-pages / ecc ---" + amd-smi bad-pages 2>&1 | head -20 || true + echo + amd-smi metric --ecc 2>&1 | head -30 || true + ' + + echo + echo "====================================================================" + echo " RDMA / POLLARA VFs" + echo "====================================================================" + kubectl -n "$NAMESPACE" exec "$NAME" -- bash -c ' + echo "--- ibv_devinfo -l ---" + ibv_devinfo -l + echo + echo "--- ibv_devinfo -v (per device summary) ---" + for d in $(ibv_devinfo -l | tail -n +2); do + echo "=== $d ===" + ibv_devinfo -d "$d" | grep -E "^\s+(state|link_layer|phys_state|max_mtu|active_mtu|node_guid|sys_image_guid|fw_ver|hca_id)" | head -20 + done + echo + echo "--- show_gids | head -30 (RoCEv2 GIDs) ---" + show_gids 2>&1 | head -30 || true + echo + echo "--- per-NIC ethtool -i / link / MTU ---" + for iface in $(ls /host/sys/class/net/ 2>/dev/null | grep -E "^(ionic|eth|enp)" | head -12); do + echo "=== $iface ===" + # ethtool -i uses the netlink socket, needs the interface visible + ethtool -i "$iface" 2>&1 | head -6 + ip -o link show "$iface" 2>&1 | head -2 + done + ' + + echo + echo "====================================================================" + echo " GPU <-> NIC PCIe TOPOLOGY (as seen inside the guest)" + echo "====================================================================" + kubectl -n "$NAMESPACE" exec "$NAME" -- bash -c ' + echo "--- lspci -Dvvv | AMD Instinct MI3* + Pensando ionic ---" + lspci -Dvvv 2>/dev/null | awk "/^[0-9a-f]/{keep=0} /1002:75/ || /1dd8:1003/ {keep=1} keep" | head -80 + echo + echo "--- topology XML we plan to reference ---" + ls -la /host/etc/crusoe/rccl_topo/ + wc -l /host/etc/crusoe/rccl_topo/mi355x-288gb-ib.xml + # sanity: is device 0x75a3 (MI355X) actually in the XML? + grep -c "0x75a3" /host/etc/crusoe/rccl_topo/mi355x-288gb-ib.xml && echo " MI355X device IDs found" + ' + + echo + echo "====================================================================" + echo " HUGEPAGES" + echo "====================================================================" + kubectl -n "$NAMESPACE" exec "$NAME" -- bash -c ' + grep -E "^Huge" /host/proc/meminfo + cat /host/sys/kernel/mm/transparent_hugepage/enabled 2>/dev/null + ' + + echo + echo "====================================================================" + echo " DMESG last 200 (amdgpu / ionic / rdma / mlx / dma-buf errors)" + echo "====================================================================" + kubectl -n "$NAMESPACE" exec "$NAME" -- bash -c ' + dmesg -T 2>/dev/null | grep -Ei "(amdgpu|ionic|rdma|mlx|dma-buf|iommu|numa)" | tail -100 + ' + + echo + echo "====================================================================" + echo " END ($NODE)" + echo "====================================================================" + } > "$OUT" 2>&1 + echo " wrote $(wc -l < "$OUT") lines to $OUT" +} + +for NODE in "${NODES[@]}"; do + collect_for_node "$NODE" +done + +# ---- teardown pods ---------------------------------------------------------- +echo +echo "==> teardown env-verify pods" +kubectl -n "$NAMESPACE" delete pod -l app=env-verify --wait=false >/dev/null + +echo +echo "==> reports written under $RESULTS_DIR:" +ls -la "$RESULTS_DIR"/env-report-*.txt 2>/dev/null | tail -10 diff --git a/cmk-amd-mi355x/validation-suite/gpu-straggler-scan/src/gpu-straggler-scan-mi355x.yaml b/cmk-amd-mi355x/validation-suite/gpu-straggler-scan/src/gpu-straggler-scan-mi355x.yaml new file mode 100644 index 0000000..ac2b27a --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/gpu-straggler-scan/src/gpu-straggler-scan-mi355x.yaml @@ -0,0 +1,120 @@ +# GPU compute health + straggler scan — adapted from mlcommons AMDMi355xTrainingv6.1 +# for our 2-node mi355x-poc cluster. +# +# What it does per node: +# 0a. GPU inventory — expect 8; fewer = a GPU fell off the bus (BAD NODE) +# 0b. PCIe slot / link width / speed (want x16 gen5) +# 0c. ECC / RAS uncorrectable + retired/pending bad memory pages (BAD GPU) +# 0d. XGMI links (bad interconnect throttles collectives) +# 1. Peak bf16 16k^3 GEMM burst on all 8 GPUs (dead/weak-bin GPUs show low) +# 2. Sustained 90s GEMM + clock trajectory at 30s/60s/90s (thermal droop / straggler) +# +# Original: mlcommons AMDMi355xTrainingv6.1/training/k8s-manifests/gpu-straggler-scan.yaml +# Uses public rocm/pytorch image — no auth needed. Total wall clock ~3 min per node, both in parallel. +apiVersion: v1 +kind: Namespace +metadata: { name: mlperfv60training } +--- +apiVersion: batch/v1 +kind: Job +metadata: + name: gpu-straggler-scan + namespace: mlperfv60training + labels: { role: gpu-straggler-scan } +spec: + completionMode: Indexed + completions: 2 + parallelism: 2 + backoffLimit: 0 + template: + metadata: { labels: { role: gpu-straggler-scan } } + spec: + restartPolicy: Never + # No nodeSelector: the amd.com/gpu resource request below already + # constrains this Job to MI355X nodes, and podAntiAffinity spreads + # one pod per node. If you have MULTIPLE MI355X nodepools in the + # same cluster and want to pin to a specific one, uncomment and + # set nodeSelector to your `crusoe.ai/nodepool.name` value (find it + # with `kubectl get nodes --show-labels`). + # nodeSelector: + # crusoe.ai/nodepool.name: + affinity: + podAntiAffinity: + requiredDuringSchedulingIgnoredDuringExecution: + - labelSelector: { matchLabels: { role: gpu-straggler-scan } } + topologyKey: kubernetes.io/hostname + containers: + - name: scan + image: rocm/pytorch:rocm7.2.2_ubuntu24.04_py3.12_pytorch_release_2.8.0 + imagePullPolicy: IfNotPresent + env: + - { name: NODE_NAME, valueFrom: { fieldRef: { fieldPath: spec.nodeName } } } + command: ["/bin/bash","-c"] + args: + - | + set -uo pipefail + export PATH=$PATH:/opt/rocm/bin + N="${NODE_NAME%%.*}" + echo "################ GPU STRAGGLER SCAN node=$N idx=${JOB_COMPLETION_INDEX} ################" + + echo "---- 0a. GPU inventory (expect 8) ----" + CNT=$(rocm-smi 2>/dev/null | grep -cE '^[0-9]+[[:space:]]' || echo 0) + echo "RESULT INV node=$N gpu_count=$CNT" + + echo "---- 0b. serial / PCIe link ----" + rocm-smi --showserial --showbus 2>/dev/null | grep -iE 'GPU\[|Serial|PCI Bus' | head -30 || true + amd-smi static --pcie 2>/dev/null | grep -iE 'GPU|width|speed|max_packet|slot' | head -24 || true + + echo "---- 0c. ECC + bad pages ----" + amd-smi metric --ecc 2>/dev/null | grep -iE 'GPU|uncorrect|correct' | head -40 \ + || rocm-smi --showrasinfo all 2>/dev/null | grep -iE 'GPU\[|UE|CE|uncorrect' | head -40 || echo "(ecc query n/a)" + amd-smi bad-pages 2>/dev/null | grep -iE 'GPU|retired|pending|page' | head -20 || echo "(bad-pages n/a)" + + echo "---- 0d. XGMI links ----" + amd-smi xgmi 2>/dev/null | grep -iE 'GPU|xgmi|link' | head -20 \ + || rocm-smi --shownodesbw 2>/dev/null | head -12 || echo "(xgmi query n/a)" + + smi_emit(){ + rocm-smi 2>/dev/null | awk -v n="$N" -v t="$1" ' + /^[0-9]/{ s=""; c=""; + for(i=1;i<=NF;i++){ if($i ~ /[Mm]hz$/ && s=="")s=$i; if($i ~ /C$/ && $i ~ /[0-9]/ && c=="")c=$i } + gsub(/[Mm]hz/,"",s); gsub(/°?C/,"",c); + if(s!="") print "RESULT CLK node="n" tag="t" gpu="$1" temp="c" sclk="s }' + } + + cat > /tmp/mm.py <<'PY' + import torch, time, os + torch.cuda.set_device(0) + n = 16384 + a = torch.randn(n, n, device='cuda', dtype=torch.bfloat16) + b = torch.randn(n, n, device='cuda', dtype=torch.bfloat16) + for _ in range(10): c = a @ b + torch.cuda.synchronize() + mode = os.environ.get('MODE', 'peak') + if mode == 'sust': + dur = 90.0; t0 = time.time(); it = 0 + while time.time() - t0 < dur: + for _ in range(20): c = a @ b + torch.cuda.synchronize(); it += 20 + dt = time.time() - t0 + else: + it = 100; t0 = time.time() + for _ in range(it): c = a @ b + torch.cuda.synchronize(); dt = time.time() - t0 + print("RESULT %s node=%s gpu=%s tflops=%.1f" % ( + mode.upper(), os.environ.get('NODE_NAME','?').split('.')[0], + os.environ.get('HIP_VISIBLE_DEVICES','?'), 2 * n**3 * it / dt / 1e12)) + PY + + echo "---- 1. PEAK burst (all 8 GPU) ----" + for i in 0 1 2 3 4 5 6 7; do MODE=peak HIP_VISIBLE_DEVICES=$i python3 /tmp/mm.py & done + wait + + echo "---- 2. SUSTAINED 90s + clock trajectory ----" + for i in 0 1 2 3 4 5 6 7; do MODE=sust HIP_VISIBLE_DEVICES=$i python3 /tmp/mm.py & done + for tag in 30s 60s 90s; do sleep 30; smi_emit "$tag"; done + wait + echo "################ DONE node=$N ################" + resources: + requests: { amd.com/gpu: 8 } + limits: { amd.com/gpu: 8 } diff --git a/cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-gpu-direct/src/gpu-to-nic-bw-gpu-direct.sh b/cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-gpu-direct/src/gpu-to-nic-bw-gpu-direct.sh new file mode 100755 index 0000000..ae5630d --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-gpu-direct/src/gpu-to-nic-bw-gpu-direct.sh @@ -0,0 +1,190 @@ +#!/usr/bin/env bash +# ============================================================================ +# gpu-to-nic-bw-gpu-direct.sh +# ---------------------------------------------------------------------------- +# Per-rail GPU-DIRECT dma-buf bidirectional bandwidth across all 8 Pollara VFs +# (node1↔node2). GPU HBM is the RDMA buffer, dma-buf is the export path, so +# the NIC DMAs straight into/out of device memory with no PCIe DRAM hop. +# +# This is the customer-facing HEADLINE test — hits ~779 Gb/s per rail +# (105 % of Crusoe engineering's published 740 Gb/s number, 97 % of the +# 800 Gb/s theoretical bidirectional line rate for 400 Gbps × 2 dir). +# +# Required image (built from ../build-image/): +# IMAGE=/roce-workload:a-77-crusoe-amdperftest +# Stock upstream perftest 6.25 has --use_rocm flag scaffolding but the +# ROCm memory-type factory function isn't registered — every invocation +# returns "Unsupported memory type". The image referenced above bundles +# perftest from github.com/ROCm/rdma-perftest branch +# master-dmabuf-rocm-20250114, configured with +# --enable-rocm --enable-rocm-dmabuf --with-rocm=/opt/rocm, so the +# --use_rocm=N --use_rocm_dmabuf flags actually work end-to-end. +# +# Three settings are load-bearing here and each will silently crush the +# number if you drop them: +# -q 4 four QPs (defaults to 1 → halves BW to ~400 Gb/s) +# -m 4096 MTU 4096 (perftest auto-negotiates down to 2048 with +# GPU-direct on some ROCm versions → halves BW again) +# /boot mount perftest reads /boot/config-$(uname -r) at startup and +# aborts silently before binding its TCP port if absent; +# client then sees "Couldn't connect". Same mount pattern +# as mlcommons rccl-mpijob-bundle2.yaml worker template. +# +# Inputs (env): +# NAMESPACE default +# IMAGE [REQUIRED] full URL of the a-77-crusoe-amdperftest image +# PULL_SECRET ccr-cred +# NODE1,NODE2 auto-picked from `amd.com/gpu` nodes if unset +# SIZE 8388608 (8 MiB) +# ITERS 5000 +# QPS 4 +# MTU 4096 +# ============================================================================ + +set -euo pipefail + +NAMESPACE="${NAMESPACE:-default}" +: "${IMAGE:?IMAGE=/roce-workload:a-77-crusoe-amdperftest is required (build via ../build-image/src/build-kaniko-amdperftest.sh)}" +PULL_SECRET="${PULL_SECRET:-ccr-cred}" +SIZE="${SIZE:-8388608}" +ITERS="${ITERS:-5000}" +QPS="${QPS:-4}" +MTU="${MTU:-4096}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +LOG_DIR="${LOG_DIR:-$SCRIPT_DIR/../logs}" +mkdir -p "$LOG_DIR" +TS=$(date -u +%Y%m%dT%H%M%SZ) +TS_LC=$(date -u +%Y%m%d-%H%M%S) +OUT="$LOG_DIR/gpu-to-nic-bw-gpu-direct-${TS}.log" + +# Auto-pick 2 amd.com/gpu nodes +if [[ -z "${NODE1:-}" || -z "${NODE2:-}" ]]; then + NODES=() + while IFS= read -r line; do + [[ -n "$line" ]] && NODES+=("$line") + done < <(kubectl get nodes -o jsonpath='{range .items[?(@.status.allocatable.amd\.com/gpu)]}{.metadata.name}{"\n"}{end}' 2>/dev/null) + NODE1="${NODE1:-${NODES[0]:-}}" + NODE2="${NODE2:-${NODES[1]:-}}" +fi +[[ -n "$NODE1" && -n "$NODE2" ]] || { echo "ERR: need NODE1/NODE2"; exit 1; } + +SRV_POD="ibwb-gd-srv-${TS_LC}" +CLI_POD="ibwb-gd-cli-${TS_LC}" + +echo "==> NODE1 (server) = $NODE1" +echo "==> NODE2 (client) = $NODE2" +echo "==> image = $IMAGE" +echo "==> size, iters = $SIZE, $ITERS" +echo "==> QPs, MTU = $QPS, $MTU" +echo "==> log = $OUT" + +# ---- deploy server pod (all 8 GPUs + all 8 vnics + /boot hostPath) -------- +cat </dev/null +apiVersion: v1 +kind: Pod +metadata: + name: $SRV_POD + labels: { app: gpu-to-nic-bw-gpu-direct, role: server } + annotations: + k8s.v1.cni.cncf.io/networks: >- + rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad +spec: + nodeName: $NODE1 + restartPolicy: Never + hostIPC: true + imagePullSecrets: [{ name: $PULL_SECRET }] + containers: + - name: srv + image: $IMAGE + command: ["/bin/bash","-c"] + args: + - | + for RAIL in 0 1 2 3 4 5 6 7; do + PORT=\$((18500 + RAIL)) + echo "=== SERVER rail=\$RAIL port=\$PORT ===" + ib_write_bw -d ionic_\$RAIL -F -b --report_gbits -s $SIZE -n $ITERS -p \$PORT \\ + -x 1 -q $QPS -m $MTU --use_rocm=\$RAIL --use_rocm_dmabuf + echo "=== SERVER rail=\$RAIL DONE ===" + done + echo "=== ALL SERVERS DONE — sleeping 30 min for post-run inspection ===" + sleep 1800 + securityContext: { capabilities: { add: [IPC_LOCK] } } + resources: + requests: { amd.com/gpu: 8, amd.com/vnic: 8 } + limits: { amd.com/gpu: 8, amd.com/vnic: 8 } + volumeMounts: + - { name: boot, mountPath: /boot, readOnly: true } + volumes: + - name: boot + hostPath: { path: /boot, type: Directory } +EOF + +echo "==> waiting for server pod Ready" +kubectl -n "$NAMESPACE" wait --for=condition=Ready pod/"$SRV_POD" --timeout=180s + +SRV_IP=$(kubectl -n "$NAMESPACE" get pod "$SRV_POD" -o jsonpath='{.status.podIP}') +echo "==> server pod IP = $SRV_IP" + +sleep 4 + +# ---- deploy client pod ---------------------------------------------------- +cat </dev/null +apiVersion: v1 +kind: Pod +metadata: + name: $CLI_POD + labels: { app: gpu-to-nic-bw-gpu-direct, role: client } + annotations: + k8s.v1.cni.cncf.io/networks: >- + rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad +spec: + nodeName: $NODE2 + restartPolicy: Never + hostIPC: true + imagePullSecrets: [{ name: $PULL_SECRET }] + containers: + - name: cli + image: $IMAGE + command: ["/bin/bash","-c"] + args: + - | + for RAIL in 0 1 2 3 4 5 6 7; do + PORT=\$((18500 + RAIL)) + echo "=== CLIENT rail=\$RAIL port=\$PORT srv=$SRV_IP ===" + sleep 4 + for attempt in 1 2 3 4 5; do + if ib_write_bw -d ionic_\$RAIL -F -b --report_gbits -s $SIZE -n $ITERS -p \$PORT \\ + -x 1 -q $QPS -m $MTU --use_rocm=\$RAIL --use_rocm_dmabuf $SRV_IP; then + break + fi + echo " rail \$RAIL attempt \$attempt failed; sleep 2 and retry" + sleep 2 + done + echo "=== CLIENT rail=\$RAIL DONE ===" + done + echo "=== ALL CLIENTS DONE ===" + securityContext: { capabilities: { add: [IPC_LOCK] } } + resources: + requests: { amd.com/gpu: 8, amd.com/vnic: 8 } + limits: { amd.com/gpu: 8, amd.com/vnic: 8 } + volumeMounts: + - { name: boot, mountPath: /boot, readOnly: true } + volumes: + - name: boot + hostPath: { path: /boot, type: Directory } +EOF + +echo "==> streaming client logs (blocks until container exits)" +kubectl -n "$NAMESPACE" wait --for=jsonpath='{.status.phase}=Running' pod/"$CLI_POD" --timeout=180s +kubectl -n "$NAMESPACE" logs -f "$CLI_POD" > "$OUT" 2>&1 + +echo +echo "==> per-rail summary:" +grep " $SIZE " "$OUT" | awk 'BEGIN{r=0} + {printf " ionic_%d : peak=%7s Gb/s avg=%7s Gb/s\n", r++, $3, $4}' + +echo +echo "==> teardown pods" +kubectl -n "$NAMESPACE" delete pod "$SRV_POD" "$CLI_POD" --wait=false >/dev/null 2>&1 || true +echo "==> log: $OUT" diff --git a/cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-hostmem/src/gpu-to-nic-bw-hostmem.sh b/cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-hostmem/src/gpu-to-nic-bw-hostmem.sh new file mode 100755 index 0000000..895b89c --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/gpu-to-nic-bw-hostmem/src/gpu-to-nic-bw-hostmem.sh @@ -0,0 +1,171 @@ +#!/usr/bin/env bash +# ============================================================================ +# gpu-to-nic-bw-hostmem.sh +# ---------------------------------------------------------------------------- +# Per-rail RDMA bidirectional bandwidth across all 8 Pollara VFs (node1↔node2) +# using HOST MEMORY as the RDMA buffer. This is the simple, no-CCR-auth, +# no-image-build baseline that proves each rail is physically healthy. +# +# For the full-fat GPU-direct dma-buf number (~779 Gb/s per rail, the +# customer-facing headline), see the sibling test: +# ../gpu-to-nic-bw-gpu-direct/src/gpu-to-nic-bw-gpu-direct.sh +# +# Design (avoids two traps that bit us earlier): +# - ONE pod per node holds all 8 vnics (matches the RCCL worker spec) and +# iterates 8 rails serially internally. Trying to run one pod per rail on +# the same node deadlocks on `amd.com/vnic` availability — the device +# plugin grants ALL 8 rails to a single pod, and a second pod can't +# schedule until the first releases. +# - Client waits a fixed 4 s before each rail's ib_write_bw invocation. +# Do NOT probe the server port with nc / /dev/tcp — ib_write_bw's server +# accepts exactly ONE TCP connection, so a probe consumes the real +# client's slot. +# +# Inputs (env): +# NAMESPACE default +# IMAGE upstream a-56 roce-workload (public, no auth needed) +# PULL_SECRET ccr-cred (unused for the public image, but the pod spec +# declares it for symmetry with the RCCL manifest) +# NODE1,NODE2 auto-picked from `amd.com/gpu` nodes if unset +# SIZE 8388608 (8 MiB message) +# ITERS 5000 +# QPS 4 (matches NCCL_IB_QPS_PER_CONNECTION) +# MTU 4096 (Pollara max active_mtu — force it to prevent +# perftest auto-negotiating down) +# ============================================================================ + +set -euo pipefail + +NAMESPACE="${NAMESPACE:-default}" +IMAGE="${IMAGE:-mirror.gcr.io/rocm/roce-workload:ubuntu24_rocm-7.2_rccl-7.2.0_anp-v1.3.0_ainic-1.117.5-a-56}" +PULL_SECRET="${PULL_SECRET:-ccr-cred}" +SIZE="${SIZE:-8388608}" +ITERS="${ITERS:-5000}" +QPS="${QPS:-4}" +MTU="${MTU:-4096}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +LOG_DIR="${LOG_DIR:-$SCRIPT_DIR/../logs}" +mkdir -p "$LOG_DIR" +TS=$(date -u +%Y%m%dT%H%M%SZ) +TS_LC=$(date -u +%Y%m%d-%H%M%S) +OUT="$LOG_DIR/gpu-to-nic-bw-hostmem-${TS}.log" + +# Auto-pick 2 amd.com/gpu nodes +if [[ -z "${NODE1:-}" || -z "${NODE2:-}" ]]; then + NODES=() + while IFS= read -r line; do + [[ -n "$line" ]] && NODES+=("$line") + done < <(kubectl get nodes -o jsonpath='{range .items[?(@.status.allocatable.amd\.com/gpu)]}{.metadata.name}{"\n"}{end}' 2>/dev/null) + NODE1="${NODE1:-${NODES[0]:-}}" + NODE2="${NODE2:-${NODES[1]:-}}" +fi +[[ -n "$NODE1" && -n "$NODE2" ]] || { echo "ERR: need NODE1/NODE2"; exit 1; } + +SRV_POD="ibwb-hm-srv-${TS_LC}" +CLI_POD="ibwb-hm-cli-${TS_LC}" + +echo "==> NODE1 (server) = $NODE1" +echo "==> NODE2 (client) = $NODE2" +echo "==> image = $IMAGE" +echo "==> size, iters = $SIZE, $ITERS" +echo "==> QPs, MTU = $QPS, $MTU" +echo "==> log = $OUT" + +# ---- deploy server pod ---------------------------------------------------- +cat </dev/null +apiVersion: v1 +kind: Pod +metadata: + name: $SRV_POD + labels: { app: gpu-to-nic-bw-hostmem, role: server } + annotations: + k8s.v1.cni.cncf.io/networks: >- + rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad +spec: + nodeName: $NODE1 + restartPolicy: Never + hostIPC: true + imagePullSecrets: [{ name: $PULL_SECRET }] + containers: + - name: srv + image: $IMAGE + command: ["/bin/bash","-c"] + args: + - | + for RAIL in 0 1 2 3 4 5 6 7; do + PORT=\$((18500 + RAIL)) + echo "=== SERVER rail=\$RAIL port=\$PORT ===" + ib_write_bw -d ionic_\$RAIL -F -b --report_gbits -s $SIZE -n $ITERS -p \$PORT -x 1 -q $QPS -m $MTU + echo "=== SERVER rail=\$RAIL DONE ===" + done + echo "=== ALL SERVERS DONE — sleeping 30 min for post-run inspection ===" + sleep 1800 + securityContext: { capabilities: { add: [IPC_LOCK] } } + resources: + requests: { amd.com/vnic: 8 } + limits: { amd.com/vnic: 8 } +EOF + +echo "==> waiting for server pod Ready" +kubectl -n "$NAMESPACE" wait --for=condition=Ready pod/"$SRV_POD" --timeout=180s + +SRV_IP=$(kubectl -n "$NAMESPACE" get pod "$SRV_POD" -o jsonpath='{.status.podIP}') +echo "==> server pod IP = $SRV_IP" + +sleep 4 + +# ---- deploy client pod ---------------------------------------------------- +cat </dev/null +apiVersion: v1 +kind: Pod +metadata: + name: $CLI_POD + labels: { app: gpu-to-nic-bw-hostmem, role: client } + annotations: + k8s.v1.cni.cncf.io/networks: >- + rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad +spec: + nodeName: $NODE2 + restartPolicy: Never + hostIPC: true + imagePullSecrets: [{ name: $PULL_SECRET }] + containers: + - name: cli + image: $IMAGE + command: ["/bin/bash","-c"] + args: + - | + for RAIL in 0 1 2 3 4 5 6 7; do + PORT=\$((18500 + RAIL)) + echo "=== CLIENT rail=\$RAIL port=\$PORT srv=$SRV_IP ===" + sleep 4 + for attempt in 1 2 3 4 5; do + if ib_write_bw -d ionic_\$RAIL -F -b --report_gbits -s $SIZE -n $ITERS -p \$PORT -x 1 -q $QPS -m $MTU $SRV_IP; then + break + fi + echo " rail \$RAIL attempt \$attempt failed; sleep 2 and retry" + sleep 2 + done + echo "=== CLIENT rail=\$RAIL DONE ===" + done + echo "=== ALL CLIENTS DONE ===" + securityContext: { capabilities: { add: [IPC_LOCK] } } + resources: + requests: { amd.com/vnic: 8 } + limits: { amd.com/vnic: 8 } +EOF + +echo "==> streaming client logs (blocks until container exits)" +kubectl -n "$NAMESPACE" wait --for=jsonpath='{.status.phase}=Running' pod/"$CLI_POD" --timeout=180s +kubectl -n "$NAMESPACE" logs -f "$CLI_POD" > "$OUT" 2>&1 + +echo +echo "==> per-rail summary:" +grep " $SIZE " "$OUT" | awk 'BEGIN{r=0} + {printf " ionic_%d : peak=%7s Gb/s avg=%7s Gb/s\n", r++, $3, $4}' + +echo +echo "==> teardown pods" +kubectl -n "$NAMESPACE" delete pod "$SRV_POD" "$CLI_POD" --wait=false >/dev/null 2>&1 || true +echo "==> log: $OUT" diff --git a/cmk-amd-mi355x/validation-suite/mirror-image/src/mirror-a77-to-ccr.yaml b/cmk-amd-mi355x/validation-suite/mirror-image/src/mirror-a77-to-ccr.yaml new file mode 100644 index 0000000..0cac3b8 --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/mirror-image/src/mirror-a77-to-ccr.yaml @@ -0,0 +1,81 @@ +# ============================================================================ +# mirror-a77-to-ccr.yaml +# ---------------------------------------------------------------------------- +# Skopeo-copy the Bundle 2.1 a-77 roce-workload image between two CCR repos. +# Runtime ~5-10 min: the image is ~30 GB, but both src and dst live in the +# same us-east2-a CCR so traffic stays on Crusoe's internal fabric. +# +# CCR docker Basic-auth username = the operator's email (not user_id UUID or +# token_id). Password = the full registry token in the format `$` +# from `crusoe registry tokens create --alias `. This is not currently +# in Crusoe's public CCR docs; see runbook.md and known-issues.md for the +# recipe. +# +# Envsubst variables required at apply time (rendered before kubectl apply): +# ${SRC_URL} — full URL of the source image incl. tag +# ${DST_URL} — full URL of the destination image incl. tag +# +# Prerequisites: +# - ccr-cred-src (docker-registry secret in ${NAMESPACE}) — email + token +# with read on the SOURCE repo +# - ccr-cred-dst (docker-registry secret in ${NAMESPACE}) — email + token +# with write on the DESTINATION repo +# (both created by the operator; see runbook.md step 3) +# ============================================================================ +apiVersion: batch/v1 +kind: Job +metadata: + name: mirror-a77-to-ccr + namespace: default + labels: { app: mi355x-poc, role: image-mirror } +spec: + backoffLimit: 2 + ttlSecondsAfterFinished: 3600 + template: + metadata: + labels: { app: mi355x-poc, role: image-mirror } + spec: + restartPolicy: Never + containers: + - name: skopeo + image: quay.io/skopeo/stable:latest + command: ["/bin/bash", "-lc"] + env: [{ name: TMPDIR, value: /var/tmp }] + args: + - | + set -euo pipefail + SRC=docker://${SRC_URL} + DST=docker://${DST_URL} + echo "==> mirror" + echo " from $SRC" + echo " to $DST" + skopeo --version + echo + echo "==> inspect source" + skopeo inspect --authfile /auth-src/config.json "$SRC" | head -25 + echo + echo "==> copy (~5-10 min for the 30 GB image)" + time skopeo copy --all --retry-times 5 \ + --src-authfile /auth-src/config.json --dest-authfile /auth-dst/config.json \ + "$SRC" "$DST" + echo + echo "==> inspect destination" + skopeo inspect --authfile /auth-dst/config.json "$DST" | head -10 + echo MIRROR_DONE + volumeMounts: + - { name: ccr-auth-src, mountPath: /auth-src, readOnly: true } + - { name: ccr-auth-dst, mountPath: /auth-dst, readOnly: true } + - { name: tmp, mountPath: /var/tmp } + resources: + requests: { cpu: "4", memory: 16Gi } + volumes: + - name: ccr-auth-src + secret: + secretName: ccr-cred-src + items: [{ key: .dockerconfigjson, path: config.json }] + - name: ccr-auth-dst + secret: + secretName: ccr-cred-dst + items: [{ key: .dockerconfigjson, path: config.json }] + - name: tmp + emptyDir: { sizeLimit: 80Gi } diff --git a/cmk-amd-mi355x/validation-suite/rccl-allreduce/src/deploy-rccl-allreduce.sh b/cmk-amd-mi355x/validation-suite/rccl-allreduce/src/deploy-rccl-allreduce.sh new file mode 100755 index 0000000..6436b66 --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/rccl-allreduce/src/deploy-rccl-allreduce.sh @@ -0,0 +1,111 @@ +#!/usr/bin/env bash +# ============================================================================ +# deploy-rccl-allreduce.sh +# ---------------------------------------------------------------------------- +# Render the MPIJob template with envsubst, apply it, wait for the +# launcher to finish, extract the Avg bus bandwidth from the log, and +# compare against the pass threshold. +# +# Pass criterion for 2 nodes × 8 MI355X + 8 Pollara rails per node: +# busbw >= 300 GB/s at 8 GiB message size +# (each node has 8×400 Gbps = 400 GB/s of RoCE bandwidth; 300 GB/s +# is a conservative Rail-Optimized 75 % efficiency target) +# +# Inputs (env or defaults): +# NAMESPACE default +# IMAGE [REQUIRED] full image ref, e.g. +# registry..ccr.crusoecloudcompute.com/./roce-workload: +# PULL_SECRET k8s docker-registry secret name (default: ccr-cred) +# N number of worker pods (default: 2 = one per MI355X node) +# MANIFEST path to rccl-allreduce-mpijob-generic.yaml +# RESULTS_DIR where to save logs (default validation-suite/results) +# BUSBW_MIN_GBPS pass threshold in GB/s (default: 300) +# ============================================================================ + +set -euo pipefail + +NAMESPACE="${NAMESPACE:-default}" +PULL_SECRET="${PULL_SECRET:-ccr-cred}" +N="${N:-2}" +BUSBW_MIN_GBPS="${BUSBW_MIN_GBPS:-300}" + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +# Script now lives at /src/, so manifest is a sibling in src/, logs at ../logs +MANIFEST="${MANIFEST:-$SCRIPT_DIR/rccl-allreduce-mpijob-generic.yaml}" +RESULTS_DIR="${RESULTS_DIR:-$SCRIPT_DIR/../logs}" +TS=$(date -u +%Y%m%dT%H%M%SZ) +LOG="$RESULTS_DIR/rccl-allreduce-${TS}.log" +mkdir -p "$RESULTS_DIR" + +: "${IMAGE:?IMAGE= required. See tokens list; ccr-cred must exist.}" + +echo "==> IMAGE : $IMAGE" +echo "==> PULL_SECRET : $PULL_SECRET" +echo "==> N (workers) : $N" +echo "==> namespace : $NAMESPACE" +echo "==> manifest : $MANIFEST" +echo "==> log : $LOG" + +command -v envsubst >/dev/null || { echo "ERR: envsubst not found (install gettext)"; exit 1; } +kubectl get ns "$NAMESPACE" >/dev/null || { echo "ERR: namespace $NAMESPACE missing"; exit 1; } +kubectl -n "$NAMESPACE" get secret "$PULL_SECRET" >/dev/null 2>&1 || { + echo "ERR: docker-registry secret $PULL_SECRET missing in $NAMESPACE"; exit 1; +} + +# ---- clean prior run -------------------------------------------------------- +if kubectl -n "$NAMESPACE" get mpijob rccl-test >/dev/null 2>&1; then + echo "==> deleting prior MPIJob rccl-test" + kubectl -n "$NAMESPACE" delete mpijob rccl-test --wait=true +fi + +# ---- apply ------------------------------------------------------------------ +echo "==> envsubst + apply" +N="$N" IMAGE="$IMAGE" PULL_SECRET="$PULL_SECRET" \ + envsubst '${N} ${IMAGE} ${PULL_SECRET}' < "$MANIFEST" | tee "$RESULTS_DIR/rccl-allreduce-${TS}.rendered.yaml" | kubectl -n "$NAMESPACE" apply -f - + +# ---- wait for launcher completion ------------------------------------------- +echo "==> waiting for launcher pod to appear (up to 5 min)" +LAUNCHER="" +for i in $(seq 1 60); do + LAUNCHER=$(kubectl -n "$NAMESPACE" get pod -l training.kubeflow.org/job-name=rccl-test,training.kubeflow.org/job-role=launcher -o jsonpath='{.items[0].metadata.name}' 2>/dev/null || true) + [[ -n "$LAUNCHER" ]] && break + sleep 5 +done +[[ -n "$LAUNCHER" ]] || { echo "ERR: launcher pod never appeared"; exit 1; } +echo " launcher = $LAUNCHER" + +echo "==> wait for launcher container to actually start (not just be scheduled)" +# The launcher container may still be in ContainerCreating; kubectl logs -f errors +# with "waiting to start" until it's Running. Poll every 5 s up to 5 min. +for i in $(seq 1 60); do + PHASE=$(kubectl -n "$NAMESPACE" get pod "$LAUNCHER" -o jsonpath='{.status.phase}' 2>/dev/null || echo "") + if [[ "$PHASE" == "Running" || "$PHASE" == "Succeeded" || "$PHASE" == "Failed" ]]; then + echo " launcher phase = $PHASE" + break + fi + sleep 5 +done + +echo "==> streaming launcher logs (kubectl logs -f blocks until the container exits) -> $LOG" +# This is the correct wait pattern: `kubectl logs -f` follows until the target +# container terminates, then returns. No separate `kubectl wait` race. +kubectl -n "$NAMESPACE" logs -f "$LAUNCHER" > "$LOG" 2>&1 || true + +# ---- parse busbw ------------------------------------------------------------ +echo +echo "==> parsing Avg bus bandwidth" +BUSBW=$(awk '/Avg bus bandwidth/ {print $NF; exit}' "$LOG") +if [[ -z "$BUSBW" ]]; then + echo "FAIL: no 'Avg bus bandwidth' line in launcher log — see $LOG" + exit 2 +fi +echo " Avg busbw = $BUSBW GB/s" + +# ---- pass/fail -------------------------------------------------------------- +if python3 -c "import sys; sys.exit(0 if float('$BUSBW') >= float('$BUSBW_MIN_GBPS') else 1)"; then + echo "PASS: busbw $BUSBW GB/s >= threshold $BUSBW_MIN_GBPS GB/s" + exit 0 +else + echo "FAIL: busbw $BUSBW GB/s < threshold $BUSBW_MIN_GBPS GB/s" + exit 3 +fi diff --git a/cmk-amd-mi355x/validation-suite/rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml b/cmk-amd-mi355x/validation-suite/rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml new file mode 100644 index 0000000..0248a81 --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml @@ -0,0 +1,213 @@ +# Multi-node RCCL all_reduce_perf MPIJob for AMD MI355X + Pollara 400. +# Combines the Crusoe MI355X mandatory NCCL env (topology XML pin, dma-buf, +# RoCEv2 GID, etc.) with the mlcommons AMDMi355xTrainingv6.1 bundle2 tuning +# envelope (~20 additional RCCL vars tuned for scale-out). +# +# Fill in THREE environment variables, then apply with envsubst: +# N number of worker pods (one per node, 8 GPUs each) +# IMAGE your RCCL test image (needs OpenMPI at /root/ompi/install, rccl-tests at /root/rccl-tests/build, sshd, and your NIC's RDMA userspace lib, e.g. libionic) +# PULL_SECRET imagePullSecret name for your registry (create with `kubectl create secret docker-registry ...`) +# +# N=8 IMAGE=myreg/rccl-tests:rocm7 PULL_SECRET=regcred \ +# envsubst '${N} ${IMAGE} ${PULL_SECRET}' < rccl-allreduce-mpijob-generic.yaml | kubectl apply -f - +# +# The quoted list restricts envsubst to exactly these three variables — every +# other $VAR in this file (the embedded scripts use $NP, $PATH, ...) is left +# untouched. envsubst ships in the gettext package. +# +# Result: kubectl logs -l training.kubeflow.org/job-name=rccl-test,training.kubeflow.org/job-role=launcher +# → "Avg bus bandwidth" at the end. +# +# Assumes: Kubeflow MPI Operator installed; namespace "default"; GPU resource amd.com/gpu and RDMA resource amd.com/vnic (edit the two `resources:` lines if your device plugins use other names). +--- +apiVersion: "k8s.cni.cncf.io/v1" +kind: NetworkAttachmentDefinition +metadata: + name: rdma-rail-nad + namespace: default + annotations: + k8s.v1.cni.cncf.io/resourceName: amd.com/vnic +spec: + config: '{ "cniVersion": "0.3.1", "type": "host-device", "ipam": {} }' +--- +apiVersion: v1 +kind: ConfigMap +metadata: + name: rccl-test-params + namespace: default +data: + MSG_SIZE_MIN: "2G" + MSG_SIZE_MAX: "8G" # bump to 32G for the asymptotic busbw number + STEP_FACTOR: "2" + ITERS: "50" + WARMUP_ITERS: "10" + CHECK: "1" + THREADS_PER_GPU: "1" + NCCL_DEBUG: "WARN" + NCCL_IB_GID_INDEX: "1" # RoCEv2 GID index on your NICs (`show_gids`) + # --- Crusoe MI355X mandatory additions (Rohit Kalmankar, 2026-09-03) --- + # Cloud Hypervisor flattens the PCIe hierarchy inside the guest, so RCCL cannot + # auto-infer GPU-to-NIC affinity. The synthetic XML pins GPUn ↔ ionic_n. The + # MI355X-specific file has device 0x75a3 (gfx950); the MI350X file's 0x75a0 + # would NOT match MI355X device IDs and silently break affinity. + NCCL_TOPO_FILE: "/etc/crusoe/rccl_topo/mi355x-288gb-ib.xml" + # Kernel 6.8 dropped ib_peer_mem support in general — dma-buf is the mandated + # path for GPU-direct RDMA on this platform. + NCCL_DMABUF_ENABLE: "1" + # AF_INET forces the bootstrap socket to IPv4. Without it NCCL can pick a v6 + # link-local address on the RDMA rail's eth interface and hang the rendezvous. + NCCL_SOCKET_FAMILY: "AF_INET" +--- +apiVersion: v1 +kind: ConfigMap +metadata: + name: rccl-launcher-script + namespace: default +data: + mpirun.sh: | + /root/ompi/install/bin/mpirun \ + --prefix /root/ompi/install \ + --mca routed direct \ + --mca plm_rsh_num_concurrent 1024 \ + --np $NP \ + --allow-run-as-root \ + --mca plm_rsh_agent "ssh -p 22" \ + --mca btl ^vader,openib,ofi \ + --mca pml ob1 \ + --mca btl_tcp_if_include eth0 \ + -x PATH=$PATH \ + -x LD_LIBRARY_PATH=$LD_LIBRARY_PATH \ + -x NCCL_DEBUG=$NCCL_DEBUG \ + -x NCCL_IB_GID_INDEX=$NCCL_IB_GID_INDEX \ + -x NCCL_SOCKET_IFNAME=eth0 \ + -x NCCL_IB_HCA=ionic_0,ionic_1,ionic_2,ionic_3,ionic_4,ionic_5,ionic_6,ionic_7 \ + -x NCCL_IB_QPS_PER_CONNECTION=4 \ + -x NCCL_IB_TC=96 \ + -x NCCL_IGNORE_CPU_AFFINITY=1 \ + -x NCCL_SHM_DISABLE=1 \ + -x NCCL_DMABUF_ENABLE=$NCCL_DMABUF_ENABLE \ + -x NCCL_TOPO_FILE=$NCCL_TOPO_FILE \ + -x NCCL_SOCKET_FAMILY=$NCCL_SOCKET_FAMILY \ + -x HSA_NO_SCRATCH_RECLAIM=1 \ + `# ---- mlcommons bundle2 tuning envelope ---- ` \ + `# Verified against mlcommons AMDMi355xTrainingv6.1 rccl-mpijob- ` \ + `# bundle2.yaml — tuned for their 16-node MLPerf runs; adopted ` \ + `# here without modification. Delta vs untuned run is <1% at 2 ` \ + `# nodes but the envelope is known-good for scale-out. ` \ + `# ` \ + `# NOTE: NCCL_NET_OPTIONAL_RECV_COMPLETION (line ~1) and ` \ + `# NET_OPTIONAL_RECV_COMPLETION (further down) are DIFFERENT vars ` \ + `# with opposite values — the NCCL_-prefixed one is NCCL's own ` \ + `# knob, the bare one is a plugin-side toggle. Both are set to ` \ + `# match the mlperf reference exactly. ` \ + -x NCCL_NET_OPTIONAL_RECV_COMPLETION=0 \ + -x NCCL_GDR_FLUSH_DISABLE=1 \ + -x RCCL_GDR_FLUSH_GPU_MEM_NO_RELAXED_ORDERING=0 \ + -x NCCL_IB_USE_INLINE=1 \ + -x IONIC_LOCKFREE=all \ + -x NCCL_GDRCOPY_ENABLE=0 \ + -x NCCL_PXN_DISABLE=1 \ + -x NCCL_IB_FIFO_TC=184 \ + -x NET_OPTIONAL_RECV_COMPLETION=1 \ + -x RCCL_AINIC_ROCE=1 \ + -x RCCL_CTS_OFFLOAD_ENABLED=0 \ + -x RCCL_CTS_INLINE_DATA=0 \ + -x NCCL_IB_DISABLE=0 \ + -x NCCL_IB_SPLIT_DATA_ON_QPS=0 \ + -x NCCL_NSOCKS_PERTHREAD=4 \ + -x NCCL_P2P_NET_CHUNKSIZE=131072 \ + -x NCCL_SOCKET_NTHREADS=2 \ + -x RCCL_DIRECT_ALLGATHER_THRESHOLD=-1 \ + -x RCCL_DISABLE_REDUCE_COPY_PIPELINING=1 \ + -x RCCL_IB_QPS_PER_P2P=1 \ + -x RCCL_MSCCL_ENABLE=0 \ + -x RCCL_P2P_BATCH_ENABLE=0 \ + -x RSMI_MUTEX_THREAD_ONLY=1 \ + bash -c 'ulimit -n 1048576 2>/dev/null || ulimit -n 65536 2>/dev/null || true; exec /root/rccl-tests/build/all_reduce_perf -b $MSG_SIZE_MIN -e $MSG_SIZE_MAX -f $STEP_FACTOR -n $ITERS -w $WARMUP_ITERS -c $CHECK -g $THREADS_PER_GPU -N 3' +--- +apiVersion: kubeflow.org/v2beta1 +kind: MPIJob +metadata: + name: rccl-test + namespace: default +spec: + slotsPerWorker: 8 + runPolicy: + cleanPodPolicy: Running # keeps the launcher pod (the busbw log) after finish + backoffLimit: 5 + mpiReplicaSpecs: + Launcher: + replicas: 1 + template: + spec: + restartPolicy: Never + automountServiceAccountToken: false + imagePullSecrets: [{ name: ${PULL_SECRET} }] + containers: + - name: rccl-launcher + image: ${IMAGE} + imagePullPolicy: IfNotPresent + envFrom: + - configMapRef: { name: rccl-test-params } + volumeMounts: + - { name: rccl-launcher-script, mountPath: /etc/rccl, readOnly: true } + command: ["/bin/bash", "-c"] + args: + - | + set -euo pipefail + ts() { printf '[%s] %s\n' "$(date -u +'%H:%M:%SZ')" "$*"; } + ts "launcher on ${HOSTNAME:-?}; waiting for workers in hostfile..." + until [ -s "${OMPI_MCA_orte_default_hostfile}" ]; do ts " hostfile empty — retry 15s"; sleep 15; done + ts "resolving worker hostnames in DNS..." + DL=$(($(date +%s)+600)) + while read -r L; do H="${L%% *}"; [ -z "$H" ] && continue + until getent hosts "$H" >/dev/null 2>&1; do [ "$(date +%s)" -gt "$DL" ] && { ts "ERR: $H unresolved"; exit 1; }; sleep 2; done + IP=$(getent hosts "$H" | awk 'NR==1{print $1; exit}') + grep -q "[[:space:]]${H}\$" /etc/hosts 2>/dev/null || echo "$IP $H" >> /etc/hosts + done < "${OMPI_MCA_orte_default_hostfile}" + ts "verifying orted on workers..." + OD=$(($(date +%s)+300)) + while read -r L; do H="${L%% *}"; [ -z "$H" ] && continue + until ssh -o ConnectTimeout=5 -o BatchMode=yes -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null "$H" test -x /root/ompi/install/bin/orted 2>/dev/null; do + [ "$(date +%s)" -gt "$OD" ] && { ts "ERR: $H missing orted"; exit 1; }; sleep 2; done + done < "${OMPI_MCA_orte_default_hostfile}" + W=$(wc -l < "${OMPI_MCA_orte_default_hostfile}"); S=$(awk -F'slots=' 'NR==1{print $2}' "${OMPI_MCA_orte_default_hostfile}") + export NP=$((W*S)); ts "NP=$NP (workers=$W slots=$S) GID_INDEX=${NCCL_IB_GID_INDEX:-1}" + python3 -c "import os,string; print(string.Template(open('/etc/rccl/mpirun.sh').read()).safe_substitute(os.environ))" > /tmp/m.sh + set +o pipefail + bash /tmp/m.sh 2>&1 + rc="${PIPESTATUS[0]}"; ts "mpirun rc=$rc"; exit "$rc" + volumes: + - name: rccl-launcher-script + configMap: { name: rccl-launcher-script } + Worker: + replicas: ${N} + template: + metadata: + annotations: + k8s.v1.cni.cncf.io/networks: >- + rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad,rdma-rail-nad + spec: + restartPolicy: Never + automountServiceAccountToken: false + imagePullSecrets: [{ name: ${PULL_SECRET} }] + containers: + - name: rccl-worker + image: ${IMAGE} + imagePullPolicy: IfNotPresent + securityContext: + capabilities: { add: ["IPC_LOCK"] } + resources: + requests: { amd.com/gpu: 8, amd.com/vnic: 8 } + limits: { amd.com/gpu: 8, amd.com/vnic: 8 } + volumeMounts: + - { name: shm, mountPath: /dev/shm } + # Bind the host's topology XML into the worker so NCCL_TOPO_FILE + # points at a real file. Read-only, per-pod. On the host this is + # provisioned by Bundle 2.1 to /etc/crusoe/rccl_topo/. + - { name: rccl-topo, mountPath: /etc/crusoe/rccl_topo, readOnly: true } + volumes: + - name: shm + emptyDir: { medium: Memory, sizeLimit: 16Gi } + - name: rccl-topo + hostPath: { path: /etc/crusoe/rccl_topo, type: Directory } From a918847ca14a8524d53a8c56edc47276e1588d34 Mon Sep 17 00:00:00 2001 From: rkalmankar-crusoe Date: Fri, 25 Sep 2026 14:50:24 -0700 Subject: [PATCH 2/4] changes --- README.md | 8 +++++ cmk-amd-mi355x/.DS_Store | Bin 6148 -> 0 bytes cmk-amd-mi355x/README.md | 39 +++++++++++++++++++++ cmk-amd-mi355x/validation-suite/.gitignore | 1 + 4 files changed, 48 insertions(+) delete mode 100644 cmk-amd-mi355x/.DS_Store create mode 100644 cmk-amd-mi355x/README.md diff --git a/README.md b/README.md index 5c530dd..3567f88 100644 --- a/README.md +++ b/README.md @@ -131,6 +131,14 @@ A combined Terraform and Ansible solution that provisions a cluster of Crusoe Cl A privileged DaemonSet that disables SMT/hyperthreading on Ubuntu-based CMK worker nodes by writing to the kernel's SMT control file and restarting kubelet, re-applying itself automatically after node reboots since it does not persist the change via grub. Intended only for specialized workloads that require hyperthreading off, since it halves the node's visible logical CPU count and requires resizing resource requests accordingly. +[AMD MI355X Validation Suite for Crusoe Managed Kubernetes](./cmk-amd-mi355x/) + +A reproducible acceptance bundle for AMD Instinct MI355X nodepools with Pensando Pollara 400 AI NICs on CMK, pinned to Bundle B.MI355.2.1. It covers per-node kernel/driver/firmware verification, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), a 2-node RCCL all-reduce with the Crusoe + mlcommons tuning envelope, and per-GPU compute/straggler/XGMI/ECC health — with reference pass bars from Crusoe's internal dry-run. + +[Fast InfiniBand Write Testing for Multiple MI355X Nodes](./ib-write-test-mi355x/) + +Tests NIC bandwidth across multiple AMD MI355X nodes using `ib_write_bw` driven in parallel from a known-good master node, then summarizes which NICs passed or failed. Parallelization keeps each node's test to a couple of seconds, making it suitable for quickly sweeping a large nodepool. + ### Observability [Crusoe Managed Kubernetes logs to Google Cloud Logging](./crusoe-managed-kubernetes-logs-to-gcp/) diff --git a/cmk-amd-mi355x/.DS_Store b/cmk-amd-mi355x/.DS_Store deleted file mode 100644 index 4aa6a72899fe70ed8934829f6563f21d32226664..0000000000000000000000000000000000000000 GIT binary patch literal 0 HcmV?d00001 literal 6148 zcmeHKyH3ME5S)b+k!V~}-Vadl2UZlmfFIyt3M2~`NvK`ryZAI_9|e)2ifE!)X>acK zcJ6djc)b8@a~SS{4#1l3h@%fn^L_V)T~)-0be=Kb8GF2A!p9=}_keRde3Cbk_mh8z z9S)4`@iy#U$CqguJy|9Nq<|EV0#ZNCq`)OA;NOQvckB!2#Q1b@ zh!%jjVmOTR=p~5F1H`^?PGp2;NhK!Ls>QIRGu|q%FPsyT4y)$F>Sn7B#o~6J-y$8> zCu)=eQs7j9>s)qT{~zdo^#7+Mt)zeyxF`i|wSC-f_@t_>i^qAbZS*I)=X}xKI1dVk mD96Mo$6R Date: Fri, 25 Sep 2026 14:53:31 -0700 Subject: [PATCH 3/4] update readme --- README.md | 2 +- cmk-amd-mi355x/README.md | 17 ++++++++--------- 2 files changed, 9 insertions(+), 10 deletions(-) diff --git a/README.md b/README.md index 3567f88..879e866 100644 --- a/README.md +++ b/README.md @@ -133,7 +133,7 @@ A privileged DaemonSet that disables SMT/hyperthreading on Ubuntu-based CMK work [AMD MI355X Validation Suite for Crusoe Managed Kubernetes](./cmk-amd-mi355x/) -A reproducible acceptance bundle for AMD Instinct MI355X nodepools with Pensando Pollara 400 AI NICs on CMK, pinned to Bundle B.MI355.2.1. It covers per-node kernel/driver/firmware verification, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), a 2-node RCCL all-reduce with the Crusoe + mlcommons tuning envelope, and per-GPU compute/straggler/XGMI/ECC health — with reference pass bars from Crusoe's internal dry-run. +A reproducible acceptance bundle for AMD Instinct MI355X nodepools with Pensando Pollara 400 AI NICs on CMK. It covers per-node kernel/driver/firmware verification against the deployed Crusoe software bundle, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), a 2-node RCCL all-reduce with the Crusoe + mlcommons tuning envelope, and per-GPU compute/straggler/XGMI/ECC health — with reference pass bars from Crusoe's internal dry-run. [Fast InfiniBand Write Testing for Multiple MI355X Nodes](./ib-write-test-mi355x/) diff --git a/cmk-amd-mi355x/README.md b/cmk-amd-mi355x/README.md index 353e9ec..22b3e92 100644 --- a/cmk-amd-mi355x/README.md +++ b/cmk-amd-mi355x/README.md @@ -1,7 +1,7 @@ # CMK AMD MI355X Solutions for validating AMD Instinct **MI355X** GPU nodepools — 8× MI355X -(gfx950, 288 GB HBM3E) plus 8× AMD Pensando Pollara 400 AI NICs per node — on +plus 8× AMD Pensando Pollara 400 AI NICs per node — on **Crusoe Managed Kubernetes (CMK)**. Use this when accepting a new MI355X nodepool or investigating a suspected fabric/GPU problem on one. @@ -9,7 +9,7 @@ nodepool or investigating a suspected fabric/GPU problem on one. | Directory | Purpose | |---|---| -| [validation-suite/](./validation-suite/) | Reproducible acceptance bundle pinned to Bundle **B.MI355.2.1**: per-node kernel/driver/firmware verification, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), 2-node RCCL all-reduce with the Crusoe + mlcommons tuning envelope, and per-GPU compute / straggler / XGMI / ECC health. | +| [validation-suite/](./validation-suite/) | Reproducible acceptance bundle: per-node kernel/driver/firmware verification against the deployed Crusoe software bundle, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), 2-node RCCL all-reduce with the Crusoe + mlcommons tuning envelope, and per-GPU compute / straggler / XGMI / ECC health. The suite's README lists the exact tested versions. | ## Prerequisites @@ -23,17 +23,16 @@ nodepool or investigating a suspected fabric/GPU problem on one. Follow the [ordered workflow](./validation-suite/README.md#ordered-workflow): environment verification → RCCL all-reduce → GPU straggler scan → per-rail -bandwidth (host-memory, then GPU-direct). Reference pass bars, observed on -Crusoe's internal dry-run: **≥ 300 GB/s** RCCL busbw (observed 381.88), -**≥ 500 Gb/s** per rail from host memory, **≥ 700 Gb/s** per rail with -GPU-direct dma-buf (observed ~778). +bandwidth (host-memory, then GPU-direct). Reference pass bars: +**≥ 300 GB/s** RCCL busbw, **≥ 500 Gb/s** per rail from host memory, +**≥ 700 Gb/s** per rail with GPU-direct dma-buf — observed dry-run numbers +are in the suite's README. ## Gotchas - The GPU-direct bandwidth check needs an AMD-patched `perftest` image — build it in-cluster first via `validation-suite/build-image/`. - RCCL on this platform requires the MI355X topology XML and - `NCCL_DMABUF_ENABLE=1` (kernel 6.8 dropped `ib_peer_mem`; presence in - `lsmod` is not activation) — the shipped manifest sets the full envelope - and the [validation-suite README](./validation-suite/README.md#rccl--nccl-environment) + `NCCL_DMABUF_ENABLE=1` — the shipped manifest sets the full tuning + envelope and the [validation-suite README](./validation-suite/README.md#rccl--nccl-environment) documents why each setting exists. From 02c1ab24e07d44baca8be47b56160248e1d86e36 Mon Sep 17 00:00:00 2001 From: rkalmankar-crusoe Date: Fri, 25 Sep 2026 15:02:35 -0700 Subject: [PATCH 4/4] move the platform/software matrix --- cmk-amd-mi355x/validation-suite/README.md | 22 +++++-------------- .../env-verify/bundle-spec.md | 19 ++++++++++++++++ 2 files changed, 24 insertions(+), 17 deletions(-) create mode 100644 cmk-amd-mi355x/validation-suite/env-verify/bundle-spec.md diff --git a/cmk-amd-mi355x/validation-suite/README.md b/cmk-amd-mi355x/validation-suite/README.md index ec709ac..3baf1f6 100644 --- a/cmk-amd-mi355x/validation-suite/README.md +++ b/cmk-amd-mi355x/validation-suite/README.md @@ -9,22 +9,10 @@ The bundle covers cluster provisioning, kernel/driver/firmware verification, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), 2-node RCCL all-reduce, and per-GPU compute + straggler + XGMI + ECC health. ---- - -## Platform - -| Component | Version | -|---|---| -| Bundle | B.MI355.2.1 | -| Linux kernel | `6.8.0-124-generic` | -| ROCm | `7.2.0` | -| RCCL | `2.27.7` | -| `amdgpu` module | `6.16.13` | -| AINIC firmware | `1.117.5-a-77` | -| Mellanox CX-7 firmware | `28.43.3608` | -| GPU | MI355X (`0x75a3`, gfx950, 256 CUs, 288 GB HBM3E) — 8 per node | -| NIC | AMD Pensando Pollara 400 — 8 × 400 Gbps VFs per node (`ionic_0…ionic_7`) | -| GPU-Direct path | dma-buf (`NCCL_DMABUF_ENABLE=1`) | +The exact tested platform (kernel, ROCm, RCCL, firmware versions) lives in +[env-verify/bundle-spec.md](env-verify/bundle-spec.md) — that file is the +reference the env-verify check is compared against, and the one to update +when the deployed Crusoe software bundle changes. --- @@ -110,7 +98,7 @@ Each step writes its outputs into its own `/logs/` (raw evidence) and 5. **Environment verification** — `bash env-verify/src/env-verify.sh`. Produces `env-verify/logs/env-report--.txt`. Confirm the - observed versions match the Bundle 2.1 spec (see table above) and + observed versions match the bundle spec ([env-verify/bundle-spec.md](env-verify/bundle-spec.md)) and that `bad_pages`, ECC counters, and Pollara VF state are clean. 6. **2-node RCCL all-reduce** — diff --git a/cmk-amd-mi355x/validation-suite/env-verify/bundle-spec.md b/cmk-amd-mi355x/validation-suite/env-verify/bundle-spec.md new file mode 100644 index 0000000..4e7830e --- /dev/null +++ b/cmk-amd-mi355x/validation-suite/env-verify/bundle-spec.md @@ -0,0 +1,19 @@ +# Bundle spec — reference for env-verify + +The versions the suite was validated against. Compare the +`env-verify/logs/env-report--.txt` output from +[env-verify.sh](src/env-verify.sh) against this table, and update the table +when the deployed Crusoe software bundle changes. + +| Component | Version | +|---|---| +| Bundle | B.MI355.2.1 | +| Linux kernel | `6.8.0-124-generic` | +| ROCm | `7.2.0` | +| RCCL | `2.27.7` | +| `amdgpu` module | `6.16.13` | +| AINIC firmware | `1.117.5-a-77` | +| Mellanox CX-7 firmware | `28.43.3608` | +| GPU | MI355X (`0x75a3`, gfx950, 256 CUs, 288 GB HBM3E) — 8 per node | +| NIC | AMD Pensando Pollara 400 — 8 × 400 Gbps VFs per node (`ionic_0…ionic_7`) | +| GPU-Direct path | dma-buf (`NCCL_DMABUF_ENABLE=1`) |