Skip to content

perf-guard: per-benchmark PR table (wild#2621 format), docs and CI tables - #46

Merged
zackees merged 1 commit into
mainfrom
docs/perf-pr-table
Sep 30, 2026
Merged

zackees merged 1 commit into
mainfrom
docs/perf-pr-table

Conversation

@zackees

@zackees zackees commented Sep 30, 2026

Copy link
Copy Markdown
Owner

Makes the benchmark breakdown from wild-linker#2621 (comment 5910436060) our standard for performance PRs sent upstream.

  • Tool: The maintainer's tool that prints this layout isn't published. It isn't in upstream benchmarks/runner, and his earlier reviews used poop. capture_ab.py now produces the same table through the new pr_table.py. Runs are ABBA-interleaved up to a time budget. Each run records wall time, user and system time and peak RSS (from wait4 rusage, with --no-fork), plus perf_event_open hardware counters where the host allows them. The table shows mean, 95% t CI ±, SD, n, and the delta vs docs: keep GitHub mutations on the fork #1 with a 95% CI. Output is still checked for byte identity.
  • Docs: benchmarks/performance_guard/README.md sets the rule (upstream PR head vs exact upstream main, generated by the tool), lists the benchmark set with capture instructions, and says which benchmarks only run locally (clang-debug about 25 GiB, chrome-android, zed, clang-release, rust-analyzer). BENCHMARKING.md gets a pointer and AGENTS.md gets the rule.
  • CI: the performance-labelled capture-ab job now captures wild-lto-debug and wild and runs the PR's own Python tests. It writes both tables (merge base as docs: keep GitHub mutations on the fork #1, PR head as ci: add a Linux PR performance guard using Hyperfine and Poop #2) to the job summary and the artifact, and keeps the upstream byte-identity checks for --build-id=none and --build-id=fast. Hosted runners have no PMU, so the counter rows there read n/a.

…bles

capture_ab.py now records wall, user, system, peak RSS and perf_event_open
hardware counters per run, runs ABBA pairs to a time budget, and emits the
maintainer's per-benchmark statistics table via the new pr_table.py (95% t
CI, SD, n, delta-method CI on the ratio of means). The performance-labelled
capture-ab job captures wild-lto-debug and wild and puts both tables in the
job summary and artifact, keeping the upstream byte-identity checks.

README documents the rule for upstream performance PRs, the benchmark set
and which benchmarks are local-only; AGENTS.md points at it.
@zackees zackees added the performance Performance change; runs the real-link A/B label Sep 30, 2026
@zackees zackees closed this Sep 30, 2026
@zackees zackees reopened this Sep 30, 2026
@zackees
zackees merged commit 745432d into main Sep 30, 2026
33 of 56 checks passed
@zackees
zackees deleted the docs/perf-pr-table branch September 30, 2026 19:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

performance Performance change; runs the real-link A/B

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant