Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -131,6 +131,14 @@ A combined Terraform and Ansible solution that provisions a cluster of Crusoe Cl

A privileged DaemonSet that disables SMT/hyperthreading on Ubuntu-based CMK worker nodes by writing to the kernel's SMT control file and restarting kubelet, re-applying itself automatically after node reboots since it does not persist the change via grub. Intended only for specialized workloads that require hyperthreading off, since it halves the node's visible logical CPU count and requires resizing resource requests accordingly.

[AMD MI355X Validation Suite for Crusoe Managed Kubernetes](./cmk-amd-mi355x/)

A reproducible acceptance bundle for AMD Instinct MI355X nodepools with Pensando Pollara 400 AI NICs on CMK. It covers per-node kernel/driver/firmware verification against the deployed Crusoe software bundle, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), a 2-node RCCL all-reduce with the Crusoe + mlcommons tuning envelope, and per-GPU compute/straggler/XGMI/ECC health — with reference pass bars from Crusoe's internal dry-run.

[Fast InfiniBand Write Testing for Multiple MI355X Nodes](./ib-write-test-mi355x/)

Tests NIC bandwidth across multiple AMD MI355X nodes using `ib_write_bw` driven in parallel from a known-good master node, then summarizes which NICs passed or failed. Parallelization keeps each node's test to a couple of seconds, making it suitable for quickly sweeping a large nodepool.

### Observability

[Crusoe Managed Kubernetes logs to Google Cloud Logging](./crusoe-managed-kubernetes-logs-to-gcp/)
Expand Down
38 changes: 38 additions & 0 deletions cmk-amd-mi355x/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,38 @@
# CMK AMD MI355X

Solutions for validating AMD Instinct **MI355X** GPU nodepools — 8× MI355X
plus 8× AMD Pensando Pollara 400 AI NICs per node — on
**Crusoe Managed Kubernetes (CMK)**. Use this when accepting a new MI355X
nodepool or investigating a suspected fabric/GPU problem on one.

## Contents

| Directory | Purpose |
|---|---|
| [validation-suite/](./validation-suite/) | Reproducible acceptance bundle: per-node kernel/driver/firmware verification against the deployed Crusoe software bundle, per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), 2-node RCCL all-reduce with the Crusoe + mlcommons tuning envelope, and per-GPU compute / straggler / XGMI / ECC health. The suite's README lists the exact tested versions. |

## Prerequisites

- A provisioned CMK cluster with an MI355X nodepool (2 × `Ready` nodes, each
advertising `amd.com/gpu: 8` and `amd.com/vnic: 8`)
- `kubectl`, the Kubeflow MPI Operator, and a CCR `docker-registry` pull
secret — exact commands in the
[validation-suite README](./validation-suite/README.md#prerequisites)

## Quick start

Follow the [ordered workflow](./validation-suite/README.md#ordered-workflow):
environment verification → RCCL all-reduce → GPU straggler scan → per-rail
bandwidth (host-memory, then GPU-direct). Reference pass bars:
**≥ 300 GB/s** RCCL busbw, **≥ 500 Gb/s** per rail from host memory,
**≥ 700 Gb/s** per rail with GPU-direct dma-buf — observed dry-run numbers
are in the suite's README.

## Gotchas

- The GPU-direct bandwidth check needs an AMD-patched `perftest` image —
build it in-cluster first via `validation-suite/build-image/`.
- RCCL on this platform requires the MI355X topology XML and
`NCCL_DMABUF_ENABLE=1` — the shipped manifest sets the full tuning
envelope and the [validation-suite README](./validation-suite/README.md#rccl--nccl-environment)
documents why each setting exists.
4 changes: 4 additions & 0 deletions cmk-amd-mi355x/validation-suite/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,4 @@
# Outputs generated when the validation checks run — do not commit.
**/logs/
**/results/
.DS_Store
163 changes: 163 additions & 0 deletions cmk-amd-mi355x/validation-suite/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,163 @@
# CMK AMD MI355X — Validation Suite

Reproducible validation bundle for AMD Instinct **MI355X** GPUs with the
**AMD Pensando Pollara 400** AI NIC on **Crusoe Managed Kubernetes (CMK)**.
Each check is packaged as a self-contained folder with the script or
manifest needed to run it against a live 2-node cluster.

The bundle covers cluster provisioning, kernel/driver/firmware verification,
per-rail RDMA bandwidth (host-memory and GPU-direct dma-buf), 2-node RCCL
all-reduce, and per-GPU compute + straggler + XGMI + ECC health.

The exact tested platform (kernel, ROCm, RCCL, firmware versions) lives in
[env-verify/bundle-spec.md](env-verify/bundle-spec.md) — that file is the
reference the env-verify check is compared against, and the one to update
when the deployed Crusoe software bundle changes.

---

## What this suite contains

Each folder ships a `src/` directory with the runnable artifact.
`logs/` and `results/` are created when the check runs.

| Folder | Purpose |
|---|---|
| [env-verify/](env-verify/) | Per-node kernel / driver / GPU / NIC / dmesg dump against Bundle 2.1 spec |
| [rccl-allreduce/](rccl-allreduce/) | 2-node RCCL `all_reduce_perf` with the Crusoe + mlcommons-tuned NCCL envelope |
| [gpu-straggler-scan/](gpu-straggler-scan/) | Per-GPU compute burst + sustained GEMM + XGMI mesh + ECC + bad-pages scan |
| [gpu-to-nic-bw-hostmem/](gpu-to-nic-bw-hostmem/) | Per-rail bidirectional `ib_write_bw` from host memory (baseline; no GPU-direct) |
| [gpu-to-nic-bw-gpu-direct/](gpu-to-nic-bw-gpu-direct/) | Per-rail bidirectional `ib_write_bw` with GPU-direct dma-buf (`--use_rocm --use_rocm_dmabuf`) |
| [build-image/](build-image/) | In-cluster Kaniko build of an AMD-patched `perftest` image (needed only for the GPU-direct bandwidth check) |
| [mirror-image/](mirror-image/) | One-time skopeo-based mirror of the Bundle 2.1 base image between registries |

---

## Observed results (Crusoe internal dry-run, 2 × MI355X-288GB-ROCE.8x)

Numbers below were observed on a 2-node CMK cluster during Crusoe's
internal validation run. They set the reference bar for a subsequent
customer POC on the same platform.

| Check | Result | Interpretation |
|---|---|---|
| Environment (Bundle 2.1) | ✅ PASS both nodes | All versions match the Bundle 2.1 spec exactly. Zero ECC errors, zero bad pages, all 8 Pollara VFs `PORT_ACTIVE` at MTU 4096 (RoCEv2). |
| 2-node RCCL all-reduce | ✅ **381.88 GB/s** Avg bus bandwidth | 95.5 % of the 400 GB/s theoretical peak (2 nodes × 8 rails × 400 Gbps). Every rail carrying GPU-direct dma-buf traffic; parity check clean across all cycles. |
| GPU compute + straggler + XGMI | ✅ 16/16 GPUs healthy | Sustained bf16 GEMM 1331-1399 TFLOPS (means: 1389.8 / 1372.8 TF), 92-97 % of MI355X vendor spec. Full XGMI mesh up. Zero stragglers. |
| Per-rail bandwidth — host memory | ✅ **~778 Gb/s** per rail, 0.2 % spread | Confirms each of the 16 rails saturates the NIC at bidirectional line rate with buffers in host memory. |
| Per-rail bandwidth — GPU-direct dma-buf | ✅ **~778 Gb/s** per rail (peak 779.28), 0.2 % rail-to-rail spread | 97 % of the 800 Gb/s theoretical bidirectional line rate. Confirms `--use_rocm --use_rocm_dmabuf` is landing in HBM with correct GPU-NIC affinity. NIC-limited, so throughput matches the host-memory baseline; the value here is proving the dma-buf code path is live end-to-end (RDMA into HBM). |

---

## Prerequisites

- `kubectl` pointed at the target CMK cluster (`kubectl get nodes` returns 2 × `Ready`).
- Each node advertises `amd.com/gpu: 8` and `amd.com/vnic: 8`.
- Kubeflow MPI Operator installed in the `default` namespace:

```bash
kubectl apply --server-side -f https://raw.githubusercontent.com/kubeflow/mpi-operator/v0.6.0/deploy/v2beta1/mpi-operator.yaml
```

- A Crusoe Container Registry (CCR) `docker-registry` secret in `default`,
used by all pods that pull the AMD workload image:

```bash
kubectl -n default create secret docker-registry ccr-cred \
--docker-server=registry.<region>.ccr.crusoecloudcompute.com \
--docker-username="<your-crusoe-email>" \
--docker-password="<crusoe-registry-token>"
```

**Note on CCR docker Basic auth:** the username is the Crusoe user's
email address (not the user UUID or token ID). The password is a
registry token issued by `crusoe registry tokens create`.

- Local tooling: `envsubst` (from `gettext`) for the RCCL manifest render step.

---

## Ordered workflow

> This suite assumes the CMK cluster and MI355X nodepool have **already been provisioned by Crusoe** — you receive a working kubeconfig, and the pre-flight in the Prerequisites section passes (2 × `Ready` nodes, each advertising `amd.com/gpu: 8` and `amd.com/vnic: 8`).

Each step writes its outputs into its own `<test>/logs/` (raw evidence) and
`<test>/results/` (parsed one-page summary), created on first run.

1. **Install the Kubeflow MPI Operator** — see Prerequisites (one-time setup, needed for the RCCL step).
2. **Create the CCR docker-registry secret** `ccr-cred` — see Prerequisites (one-time setup, needed for private images).
3. **(Optional) Mirror the Bundle 2.1 base image into your CCR** —
`mirror-image/src/mirror-a77-to-ccr.yaml` (or ask your Crusoe SE contact
to run `crusoe registry manifests copy` for you). Skip if you already
have the AMD workload image in your CCR.
4. **(Optional) Build the AMD-patched perftest image** —
`bash build-image/src/build-kaniko-amdperftest.sh`.
Only needed if you plan to run the `gpu-to-nic-bw-gpu-direct` check
with GPU-direct dma-buf. The other checks work against the upstream
`mirror.gcr.io/rocm/roce-workload:…-a-56` image.
5. **Environment verification** —
`bash env-verify/src/env-verify.sh`.
Produces `env-verify/logs/env-report-<node>-<ts>.txt`. Confirm the
observed versions match the bundle spec ([env-verify/bundle-spec.md](env-verify/bundle-spec.md)) and
that `bad_pages`, ECC counters, and Pollara VF state are clean.
6. **2-node RCCL all-reduce** —

```bash
IMAGE=<full CCR image ref> PULL_SECRET=ccr-cred N=2 \
bash rccl-allreduce/src/deploy-rccl-allreduce.sh
```

Reference bar: **≥ 300 GB/s** busbw (dry-run observed 381.88 GB/s).
7. **Per-GPU compute + straggler scan** —
`kubectl apply -f gpu-straggler-scan/src/gpu-straggler-scan-mi355x.yaml`
(edit the nodeSelector line if your nodepool label differs).
8. **Per-rail host-memory bandwidth baseline** —
`bash gpu-to-nic-bw-hostmem/src/gpu-to-nic-bw-hostmem.sh`.
Reference bar: **≥ 500 Gb/s** per rail.
9. **Per-rail GPU-direct dma-buf bandwidth** —
`IMAGE=<CCR image with amd-patched perftest> \
bash gpu-to-nic-bw-gpu-direct/src/gpu-to-nic-bw-gpu-direct.sh`.
Reference bar: **≥ 700 Gb/s** per rail.

---

## RCCL / NCCL environment

The manifest at
[rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml](rccl-allreduce/src/rccl-allreduce-mpijob-generic.yaml)
combines the settings mandatory for MI355X + Pollara on CMK with the
mlcommons `AMDMi355xTrainingv6.1` bundle2 tuning envelope. The
load-bearing settings and their rationale:

| Variable | Value | Rationale |
|---|---|---|
| `NCCL_TOPO_FILE` | `/etc/crusoe/rccl_topo/mi355x-288gb-ib.xml` | Cloud Hypervisor flattens the PCIe view, so RCCL cannot infer GPU ↔ NIC affinity on its own. The MI355X XML has device id `0x75a3` (gfx950); the MI350X `0x75a0` would silently mis-match and drop bandwidth. |
| `NCCL_DMABUF_ENABLE` | `1` | Kernel 6.8+ dropped `ib_peer_mem`. The module appears in `lsmod` on this bundle but is **not** registered as a peer-memory client by `amdgpu` — presence is not activation. dma-buf is the real GPU-direct path. |
| `NCCL_SOCKET_FAMILY` | `AF_INET` | Prevents NCCL from picking an IPv6 link-local on the RDMA `eth` interface and hanging rendezvous. |
| `NCCL_IB_HCA` | `ionic_0,…,ionic_7` | Explicit Pollara VF list — the 8 rails. |
| `NCCL_IB_GID_INDEX` | `1` | RoCEv2 GID index on Pollara. |
| `NCCL_IB_QPS_PER_CONNECTION` | `4` | Rail-optimised bandwidth ramp. |
| `NCCL_IB_TC` | `96` | RoCE traffic class (DSCP 48). |
| `NCCL_IGNORE_CPU_AFFINITY` | `1` | Cloud-Hypervisor synthetic CPU affinity is not authoritative. |
| `NCCL_SHM_DISABLE` | `1` | Forces inter-node traffic through the NIC path instead of SHM. |
| `HSA_NO_SCRATCH_RECLAIM` | `1` | ROCm 7.2 gfx950 workaround for a scratch-spill path. |
| `RCCL_AINIC_ROCE`, `IONIC_LOCKFREE`, `NCCL_GDR_FLUSH_DISABLE`, … | (~20 more) | mlcommons bundle2 tuning envelope. Full list is in the manifest. |

---

## Adapting the suite for a different project

The scripts read all tenant-specific values from environment variables
and enforce them with `${VAR:?…}` guards. To retarget the suite:

| Variable | Where to find it | Used by |
|---|---|---|
| `PROJECT_ID` | `crusoe projects list` | mirror-image, build-image |
| `CCR_URL`, `CCR_REPO` (= `<registry>.<project-short>`) | `crusoe registry list` | build-image, mirror-image |
| `IMAGE` | Full CCR image ref of the AMD workload image | rccl-allreduce, gpu-to-nic-bw-gpu-direct |
| `PULL_SECRET` | Name of the `docker-registry` secret you created | rccl-allreduce, gpu-to-nic-bw-gpu-direct |

Everything else — Dockerfile, RCCL manifest, NCCL environment, target
thresholds, topology XML path — is invariant across tenants because it
is specific to the MI355X + Bundle 2.1 platform, not to any particular
project.
Original file line number Diff line number Diff line change
@@ -0,0 +1,63 @@
# syntax=docker/dockerfile:1
# Bundle 2.1 a-77 image + AMD-patched perftest with --enable-rocm --enable-rocm-dmabuf.
#
# The stock perftest that ships in the a-77 base has --use_rocm / --use_rocm_dmabuf
# in its argparse table but the ROCm memory-type backend isn't registered in the
# binary — every invocation with those flags returns "Unsupported memory type".
# This image builds AMD's fork from source with the right configure flags so
# both flags work end-to-end for GPU-direct RDMA.
#
# Source: github.com/ROCm/rdma-perftest, branch master-dmabuf-rocm-20250114,
# commit 2db71141d6a3 "Enable dmabuf to ROCm and hipify changes" (2025-01-16).
#
# Build via in-cluster kaniko:
# CCR_REPO=<your-registry-name>.<project-short> \
# PROJECT_ID=<your-crusoe-project-uuid> \
# bash build-kaniko-amdperftest.sh
#
# Runtime test:
# ib_write_bw -F -b --report_gbits -s 8M -n 5000 -x 1 -q 4 -m 4096 \
# --use_rocm=<N> --use_rocm_dmabuf -p <port> [server_ip]
#
# The FROM below is the Bundle 2.1 a-77 base image in the DELIVERING PROJECT's
# CCR. `build-kaniko-amdperftest.sh` doesn't override this (kaniko reads the
# FROM literally), so you MUST edit this URL to point at your project's CCR
# path before running the build. See runbook.md for how to obtain the a-77 tag
# for your project (mirror from mcala-lab or ask your Crusoe SA).
ARG BASE_IMAGE=registry.us-east2-a.ccr.crusoecloudcompute.com/<your-registry-name>.<your-project-short>/roce-workload:ubuntu24_rocm-7.2_rccl-7.2.0_anp-v1.3.0_ainic-1.117.5-a-77-crusoe
FROM ${BASE_IMAGE}

ENV DEBIAN_FRONTEND=noninteractive

RUN apt-get update \
&& apt-get install -y --no-install-recommends \
build-essential autoconf libtool git ca-certificates pkg-config \
libibverbs-dev librdmacm-dev libibumad-dev libpci-dev libmnl-dev \
&& rm -rf /var/lib/apt/lists/*

# Clone the exact branch + commit that Crusoe engineering used for the 740 Gb/s blog.
RUN cd /tmp \
&& git clone --branch master-dmabuf-rocm-20250114 --depth 1 https://github.com/ROCm/rdma-perftest.git \
&& cd rdma-perftest \
&& git rev-parse HEAD > /tmp/perftest-commit.txt \
&& cat /tmp/perftest-commit.txt

# Configure with --enable-rocm --enable-rocm-dmabuf; install into /usr/local/bin so it
# takes precedence over the stock /usr/bin/ib_write_bw on PATH.
RUN cd /tmp/rdma-perftest \
&& ./autogen.sh \
&& export CFLAGS="-I/opt/rocm/include" \
&& export LDFLAGS="-L/opt/rocm/lib -L/opt/rocm/lib64 -Wl,-rpath=/opt/rocm/lib -lamdhip64 -lhsa-runtime64 -Wl,--copy-dt-needed-entries" \
&& ./configure --enable-rocm --enable-rocm-dmabuf --with-rocm=/opt/rocm \
--prefix=/usr/local \
&& make -j$(nproc) \
&& make install \
&& ldconfig \
&& rm -rf /tmp/rdma-perftest

# Sanity guards: fail the build if the new binary doesn't accept the ROCm flags.
RUN /usr/local/bin/ib_write_bw --help 2>&1 | grep -q "use_rocm" \
&& echo "OK: /usr/local/bin/ib_write_bw has --use_rocm support"

# Ensure /usr/local/bin comes first in PATH for all shells
ENV PATH=/usr/local/bin:$PATH
Loading
Loading