perf-guard: per-benchmark PR table (wild#2621 format), docs and CI tables - #46
Merged
Merged
Conversation
…bles capture_ab.py now records wall, user, system, peak RSS and perf_event_open hardware counters per run, runs ABBA pairs to a time budget, and emits the maintainer's per-benchmark statistics table via the new pr_table.py (95% t CI, SD, n, delta-method CI on the ratio of means). The performance-labelled capture-ab job captures wild-lto-debug and wild and puts both tables in the job summary and artifact, keeping the upstream byte-identity checks. README documents the rule for upstream performance PRs, the benchmark set and which benchmarks are local-only; AGENTS.md points at it.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes the benchmark breakdown from wild-linker#2621 (comment 5910436060) our standard for performance PRs sent upstream.
benchmarks/runner, and his earlier reviews used poop.capture_ab.pynow produces the same table through the newpr_table.py. Runs are ABBA-interleaved up to a time budget. Each run records wall time, user and system time and peak RSS (fromwait4rusage, with--no-fork), plusperf_event_openhardware counters where the host allows them. The table shows mean, 95% t CI ±, SD, n, and the delta vs docs: keep GitHub mutations on the fork #1 with a 95% CI. Output is still checked for byte identity.benchmarks/performance_guard/README.mdsets the rule (upstream PR head vs exact upstreammain, generated by the tool), lists the benchmark set with capture instructions, and says which benchmarks only run locally (clang-debug about 25 GiB, chrome-android, zed, clang-release, rust-analyzer).BENCHMARKING.mdgets a pointer andAGENTS.mdgets the rule.performance-labelledcapture-abjob now captureswild-lto-debugandwildand runs the PR's own Python tests. It writes both tables (merge base as docs: keep GitHub mutations on the fork #1, PR head as ci: add a Linux PR performance guard using Hyperfine and Poop #2) to the job summary and the artifact, and keeps the upstream byte-identity checks for--build-id=noneand--build-id=fast. Hosted runners have no PMU, so the counter rows there read n/a.