Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
131 changes: 73 additions & 58 deletions .github/workflows/performance-guard.yml
Original file line number Diff line number Diff line change
Expand Up @@ -137,11 +137,14 @@ jobs:
retention-days: 30

capture-ab:
# A real debug-heavy link, where the synthetic guard workloads above are too small to show
# write-phase work: Wild itself built with fat LTO and full debug info (one ~250 MB object),
# captured with WILD_SAVE_BASE and relinked by the merge base and the PR head. The PR head's
# output must also match upstream wild's byte for byte. Runs on PRs labelled `performance`.
name: Real full-LTO link A/B (merge base vs PR head, upstream byte identity)
# Real links, where the synthetic guard workloads above are too small to show write-phase
# work. Two of the benchmarks from benchmarks/performance_guard/README.md fit a hosted runner:
# `wild-lto-debug` (Wild built with fat LTO and full debug info, one ~250 MB object) and
# `wild` (Wild's ordinary release build). Both are captured with WILD_SAVE_BASE and relinked
# by the merge base and the PR head, interleaved, into the per-benchmark statistics table that
# upstream performance PRs carry. The PR head's output must also match upstream wild's byte for
# byte. Runs on PRs labelled `performance`.
name: Real link A/B tables (merge base vs PR head, upstream byte identity)
if: contains(github.event.pull_request.labels.*.name, 'performance')
runs-on: ubuntu-24.04
timeout-minutes: 60
Expand Down Expand Up @@ -198,77 +201,89 @@ jobs:
cp "$tree/target/release/wild" "$RUNNER_TEMP/bins/$tree-wild"
done
sha256sum "$RUNNER_TEMP/bins/"*
- name: Capture a fat-LTO debug link of Wild
- name: Test the PR's performance tooling
shell: bash
env:
PYTHONDONTWRITEBYTECODE: "1"
run: python3 -m unittest discover -s candidate/benchmarks/performance_guard -p 'test_*.py' -v
- name: Capture Wild's own links (wild-lto-debug and wild)
shell: bash
run: |
set -euo pipefail
# Plain cargo: soldr's compile cache isn't wanted for a one-off fat-LTO build.
cd candidate
env -u CLICOLOR_FORCE -u FORCE_COLOR \
RUSTUP_TOOLCHAIN="$RUST_TOOLCHAIN" \
shopt -s inherit_errexit
# Plain cargo: soldr's compile cache isn't wanted for one-off capture builds.
# Prints the save-dir of the wild binary's own link (build scripts are captured too).
capture_wild() {
local name=$1
shift
(cd candidate && env -u CLICOLOR_FORCE -u FORCE_COLOR \
RUSTUP_TOOLCHAIN="$RUST_TOOLCHAIN" \
CARGO_TARGET_DIR="$RUNNER_TEMP/target-$name" \
CARGO_TARGET_X86_64_UNKNOWN_LINUX_GNU_LINKER=clang \
RUSTFLAGS="-Clink-arg=--ld-path=$RUNNER_TEMP/bins/baseline-wild" \
WILD_SAVE_BASE="$RUNNER_TEMP/capture-$name" \
"$@" cargo build --release --package wild-linker --bin wild >&2)
local dir
dir=$(dirname "$(grep -l -E '^# Original output file: .*/wild-[0-9a-f]+$' \
"$RUNNER_TEMP/capture-$name"/*/run-with | head -1)")
du -sh "$dir" >&2
echo "$dir"
}
lto=$(capture_wild wild-lto-debug env \
CARGO_PROFILE_RELEASE_LTO=fat \
CARGO_PROFILE_RELEASE_DEBUG=full \
CARGO_PROFILE_RELEASE_CODEGEN_UNITS=1 \
CARGO_TARGET_DIR="$RUNNER_TEMP/lto-target" \
CARGO_TARGET_X86_64_UNKNOWN_LINUX_GNU_LINKER=clang \
RUSTFLAGS="-Clink-arg=--ld-path=$RUNNER_TEMP/bins/baseline-wild" \
WILD_SAVE_BASE="$RUNNER_TEMP/capture" \
cargo build --release --package wild-linker --bin wild
# The biggest capture is the wild binary's own link.
capture=$(du -s "$RUNNER_TEMP/capture"/*/ | sort -n | tail -1 | cut -f2)
echo "CAPTURE=$capture" >> "$GITHUB_ENV"
du -sh "$capture"
- name: Prepare a job-private perf collector
shell: bash
run: |
set -uo pipefail
# Follows zackees/mimalloc-pprof ci/perf_access.py: never relax perf_event_paranoid;
# grant CAP_PERFMON to a private copy of the kernel-matched perf ELF instead.
sudo apt-get install -y -qq "linux-tools-$(uname -r)" linux-tools-common >/dev/null 2>&1 || true
elf="/usr/lib/linux-tools/$(uname -r)/perf"
perf_bin=perf
if [ -x "$elf" ]; then
install -d -m 0700 "$RUNNER_TEMP/perf"
cp "$elf" "$RUNNER_TEMP/perf/perf-private"
sudo -n chown "root:$(id -g)" "$RUNNER_TEMP/perf/perf-private" &&
sudo -n chmod 0550 "$RUNNER_TEMP/perf/perf-private" &&
sudo -n setcap cap_perfmon=ep "$RUNNER_TEMP/perf/perf-private" &&
perf_bin="$RUNNER_TEMP/perf/perf-private"
fi
echo "PERF_BIN=$perf_bin" >> "$GITHUB_ENV"
echo "perf_event_paranoid=$(cat /proc/sys/kernel/perf_event_paranoid) perf=$perf_bin"
- name: Paired A/B with upstream byte identity
CARGO_PROFILE_RELEASE_CODEGEN_UNITS=1)
plain=$(capture_wild wild env)
{
echo "CAPTURE_WILD_LTO_DEBUG=$lto"
echo "CAPTURE_WILD=$plain"
} >> "$GITHUB_ENV"
- name: Paired A/B tables with upstream byte identity
shell: bash
run: |
set -euo pipefail
mkdir -p evidence
base=${{ steps.revisions.outputs.base }}
cand=${{ steps.revisions.outputs.candidate }}
{
echo "## Real full-LTO link A/B"
echo "## Real link A/B"
echo
echo "- Baseline (merge base): \`${{ steps.revisions.outputs.base }}\`"
echo "- Candidate (PR head): \`${{ steps.revisions.outputs.candidate }}\`"
echo "- Upstream reference: \`${{ steps.revisions.outputs.upstream }}\`"
echo "- #1 baseline (merge base): \`$base\`"
echo "- #2 candidate (PR head): \`$cand\`"
echo "- Upstream reference for byte identity: \`${{ steps.revisions.outputs.upstream }}\`"
echo "- Runner: $(nproc) CPUs, $(grep -m1 'model name' /proc/cpuinfo | cut -d: -f2)"
echo
echo "Tables compare against the fork's merge base, not upstream main; see"
echo "benchmarks/performance_guard/README.md for the upstream-PR table."
echo
} > evidence/summary.md
for mode in none fast; do
status=0
for bench in wild-lto-debug wild; do
case $bench in
wild-lto-debug) capture=$CAPTURE_WILD_LTO_DEBUG ;;
wild) capture=$CAPTURE_WILD ;;
esac
printf '### %s\n\n' "$bench" >> evidence/summary.md
ab=(python3 candidate/benchmarks/performance_guard/capture_ab.py
--capture "$capture" --benchmark "$bench"
--baseline "$RUNNER_TEMP/bins/baseline-wild" --baseline-label "merge-base ${base:0:12}"
--candidate "$RUNNER_TEMP/bins/candidate-wild" --candidate-label "pr-head ${cand:0:12}"
--reference "$RUNNER_TEMP/bins/upstream-wild")
# The table: the capture's own arguments, with --no-fork so rusage is the linker's.
"${ab[@]}" --extra=--no-fork --pairs 30 --max-pairs 5000 --budget-seconds 240 \
--output "evidence/table-$bench.json" --markdown evidence/summary.md || status=1
# Ubuntu's clang passes --build-id, so the capture has one; a later flag overrides it.
extra="--build-id=$mode"
reference=(--reference "$RUNNER_TEMP/bins/upstream-wild")
printf '### --build-id=%s\n\n' "$mode" >> evidence/summary.md
python3 candidate/benchmarks/performance_guard/capture_ab.py \
--capture "$CAPTURE" \
--baseline "$RUNNER_TEMP/bins/baseline-wild" \
--candidate "$RUNNER_TEMP/bins/candidate-wild" \
"${reference[@]}" \
--extra="$extra" \
--perf "$PERF_BIN" \
--pairs 30 \
--output "evidence/capture-ab-$mode.json" \
--markdown evidence/summary.md
for mode in none fast; do
printf '\nIdentity with `--build-id=%s`:\n\n' "$mode" >> evidence/summary.md
"${ab[@]}" --extra="--build-id=$mode --no-fork" --pairs 0 --warmups 1 \
--output "evidence/identity-$bench-$mode.json" --markdown evidence/summary.md || status=1
done
echo >> evidence/summary.md
done
cat evidence/summary.md >> "$GITHUB_STEP_SUMMARY"
python3 candidate/benchmarks/performance_guard/pr_table.py evidence/table-*.json \
> evidence/pr-table.md || status=1
exit $status
- name: Upload A/B evidence
if: always()
uses: actions/upload-artifact@b7c566a772e6b6bfb58ed0dc250532a479d7789f # v6.0.0
Expand Down
5 changes: 5 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,3 +15,8 @@
- `main` is the fork: its own commits rebased on `upstream`. To take new upstream work, fast-forward `upstream`, then rebase `main` onto it.
- PR branches come off `main` and target `main`.
- The performance guard requires a PR's output to match wild built from the newest `upstream` commit the PR contains, so keep `upstream` exactly equal to what `main` is rebased on.

## Performance PRs

- Every performance PR proposed to `wild-linker/wild` must include the per-benchmark statistics table described in `benchmarks/performance_guard/README.md` ("Upstream performance PRs"), comparing the exact PR head against the exact upstream `main` it is based on.
- Generate the tables with `benchmarks/performance_guard/capture_ab.py` and paste the output of `pr_table.py`; never hand-type or edit the numbers. Fork CI's tables (the `performance` label) compare against the fork's merge base and are not a substitute.
6 changes: 6 additions & 0 deletions BENCHMARKING.md
Original file line number Diff line number Diff line change
Expand Up @@ -389,3 +389,9 @@ cargo run --bin benchmark-runner -- \
bench --config benchmarks/ryzen-9955hx.toml --saves ~/save --tmp /ram/%fs/out linker1 linker2 linker3
cargo run --bin benchmark-runner -- report
```

### Tables for performance pull requests

This fork's requirements for performance PRs sent upstream (a per-benchmark statistics table
comparing the PR head against upstream `main`, and the tool that produces it) are in
[benchmarks/performance_guard/README.md](benchmarks/performance_guard/README.md#upstream-performance-prs-the-per-benchmark-table).
119 changes: 119 additions & 0 deletions benchmarks/performance_guard/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,3 +75,122 @@ python3 benchmarks/performance_guard/compare_revisions.py \
```

Results are written to `results.json` and `results.md` in `--output`.

## Upstream performance PRs: the per-benchmark table

Every performance PR we propose to `wild-linker/wild` must include, for each benchmark it
touches, a table in the format the maintainer uses to review performance changes
([wild#2621](https://github.com/wild-linker/wild/pull/2621#issuecomment-5910436060); these are
its `wild-lto-debug` numbers, with this tool's header lines):

```
#1 baseline upstream-main <upstream-sha>
#2 candidate pr-head <pr-sha>
metric mean CI ± SD n delta vs #1 (95% CI)
wall 81.263 ms 0.09% 1.87% 1771 -6.60% [-6.71%, -6.49%]
user 687.344 ms 0.18% 3.83% 1771 -16.37% [-16.56%, -16.18%]
system 173.787 ms 0.50% 10.79% 1771 -2.31% [-3.01%, -1.60%]
peak RSS 467.30 MiB 0.01% 0.16% 1771 +0.07% [+0.06%, +0.08%]
cycles 1.313 G~ 0.29% 6.31% 1771 -17.14% [-17.42%, -16.87%]
instructions 1.370 G~ 0.22% 4.76% 1771 -34.47% [-34.67%, -34.27%]
cache-references 94.601 M~ 3.42% 73.34% 1771 -23.56% [-27.09%, -20.02%]
cache-misses 18.960 M~ 2.95% 63.36% 1771 -16.94% [-20.43%, -13.44%]
branches 209.206 M~ 0.46% 9.92% 1771 -34.43% [-34.82%, -34.05%]
branch-misses 63.827 M~ 4.76% 102.23% 1771 -16.87% [-22.05%, -11.70%]
```

Rules:

- `#1` is the exact current upstream `main` the PR branch is based on, and `#2` is the exact PR
head. Rebuild both after every rebase or push; old numbers are context only. Both are built the
same way (`--release`, same toolchain, no debug info).
- The table comes from `capture_ab.py` / `pr_table.py`. Paste their output; never hand-type or
edit numbers.
- Include every benchmark below that the change can plausibly affect, and at least `wild` and
`wild-lto-debug`. For a benchmark you couldn't run, say so rather than leaving it out silently.
- Outputs must be identical apart from the `.comment` provenance record and build ID; the tool
checks this for every table.

The maintainer's own tool that prints this layout isn't published. It isn't in upstream's
`benchmarks/runner` (whose `report` uses 99% intervals and no hardware counters), and his earlier
reviews used [poop](https://github.com/andrewrk/poop). `capture_ab.py` reproduces the layout:

- Runs are warmed up, then interleaved in ABBA pairs until `--pairs` is reached and, after that,
until `--budget-seconds` of measuring is spent (capped by `--max-pairs`). Short links need
thousands of runs for CIs of about 0.1%; the maintainer used n = 10,881 for `wild`.
- wall is timed around the fork and exec of `run-with`. user, system and peak RSS come from
`wait4(2)` rusage, so pass `--no-fork` (the default in the commands below) to make the linker do
its own teardown and have peak RSS be the linker's.
- cycles, instructions, cache-references, cache-misses, branches and branch-misses come from
`perf_event_open(2)` counters attached to the child before exec (no `perf` wrapper overhead),
counting user and kernel work where `perf_event_paranoid` allows it and user-only otherwise.
Rows print `n/a` with the reason when the host has no usable PMU, as on GitHub-hosted runners.
- `mean`, `n`, `CI ±` and `SD` describe `#2`. `CI ±` is the 95% Student-t half-width of the mean
and `SD` the sample standard deviation, both as a percentage of the mean. `delta vs #1` is
`mean(#2) / mean(#1) - 1`, with a 95% interval from the delta method on that ratio (standard
error `r * sqrt((se2/m2)^2 + (se1/m1)^2)`, Welch-Satterthwaite degrees of freedom).

### Running it

```sh
# Build both binaries the same way. Upstream has no rust-toolchain.toml, so pin the toolchain.
git worktree add ../wild-base upstream/main
git worktree add ../wild-pr <pr-branch>
for tree in ../wild-base ../wild-pr; do
(cd "$tree" && RUSTUP_TOOLCHAIN=1.98.1 SOLDR_ALLOW_UNPINNED=1 \
soldr cargo build --release -p wild-linker --bin wild)
done

# One benchmark: a save-dir captured with WILD_SAVE_BASE (see below).
python3 benchmarks/performance_guard/capture_ab.py \
--capture ~/save/wild-lto-debug --benchmark wild-lto-debug \
--baseline ../wild-base/target/release/wild \
--baseline-label "upstream-main $(git -C ../wild-base rev-parse --short=12 HEAD)" \
--candidate ../wild-pr/target/release/wild \
--candidate-label "pr-head $(git -C ../wild-pr rev-parse --short=12 HEAD)" \
--extra=--no-fork --pairs 100 --max-pairs 20000 --budget-seconds 600 \
--output /tmp/pr-tables/wild-lto-debug.json

# The PR-description markdown for everything measured.
python3 benchmarks/performance_guard/pr_table.py /tmp/pr-tables/*.json
```

Measure on a quiet machine (check `/proc/loadavg`), and add `--cpus` to pin both arms to the same
cores if other work can't be stopped. The JSON keeps every raw sample; attach it or keep it with
the PR.

### Benchmarks

Captures are made as described in [BENCHMARKING.md](../../BENCHMARKING.md): build the project with
Wild as the linker (`RUSTFLAGS="-Clinker=clang -Clink-arg=--ld-path=$(which wild)"`, or
`-DCMAKE_{EXE,SHARED}_LINKER_FLAGS=--ld-path=$(which wild)` for CMake) and
`WILD_SAVE_BASE=~/save/<name>`, then keep the numbered directory whose `run-with` ends with the
main binary's `# Original output file:`.

| Benchmark | What is linked | How to capture | Peak RSS / link time (maintainer's machine) | Where |
| --- | --- | --- | --- | --- |
| `wild-lto-debug` | Wild, fat LTO + full debug info | `CARGO_PROFILE_RELEASE_LTO=fat CARGO_PROFILE_RELEASE_DEBUG=full CARGO_PROFILE_RELEASE_CODEGEN_UNITS=1 cargo build --release -p wild-linker --bin wild` | 0.5 GiB / 80 ms | CI and local |
| `wild` | Wild, ordinary release build | `cargo build --release -p wild-linker --bin wild` | 90 MiB / 30 ms | CI and local |
| `rust-analyzer` | rust-analyzer release | `cargo build --release -p rust-analyzer` in rust-lang/rust-analyzer | 0.65 GiB / 120 ms | local |
| `zed-release` | Zed editor release | `cargo build --release -p zed` in zed-industries/zed | 6.5 GiB / 0.6 s | local, big machine |
| `clang-release` | clang from LLVM, `CMAKE_BUILD_TYPE=Release` | CMake + Ninja build of `clang` with the flags above | 0.8 GiB / 100 ms | local, long build |
| `clang-debug` | clang from LLVM, `CMAKE_BUILD_TYPE=Debug` | as above with `Debug` | **25 GiB** / 2.3 s | local, needs 32+ GiB RAM |
| `chrome-android` | Chromium for Android (`target_os="android"`) | `gn gen` + `autoninja` of the main Android library | 4.2 GiB / 1.3 s | local, very long build and large disk |

Hosted CI (4 CPUs, ~16 GB RAM, 60-minute job) runs only `wild-lto-debug` and `wild`; the others
either can't be built within the job budget or don't fit in memory, so they are measured locally
and pasted from the tool's output.

### What CI produces

Labelling a fork PR `performance` runs the `capture-ab` job of
`.github/workflows/performance-guard.yml` (the label only takes effect on the next push or
reopen). It builds the merge base, the PR head and the newest upstream commit the PR contains,
captures `wild-lto-debug` and `wild`, and for each prints this table (merge base as `#1`, PR head
as `#2`, about four minutes of interleaved pairs each) to the job summary. The JSON reports and
`pr-table.md` are in the `capture-ab-<sha>` artifact. It also checks that the PR head's output is
byte-identical to upstream wild's with `--build-id=none` and `--build-id=fast`.

These CI tables compare against the fork's merge base, so they show what the fork PR changes. For
the upstream PR, rerun the tool with upstream `main` as `#1` and the upstream-bound branch as
`#2`, as above. Hosted runners have no hardware counters, so the counter rows there read `n/a`.
Loading
Loading