pumpkin-bench is a benchmark harness for
Pumpkin. It runs repeatable bot
workloads against local binaries, Git revisions, and pull requests, with tick
data collected by
pumpkin-pulse.
Each run starts from the same world snapshot. A/B campaigns alternate between both builds and compare the observed range across repeated runs instead of treating a single result as conclusive.
- Python 3.11 or newer
- A Pumpkin server directory with
pumpkin.toml - The
pumpkin-pulseWebAssembly plugin - Git and Rust when building Pumpkin or the load bot from source
- iproute2
ssfor network measurements on Linux (optional)
Tick recording is supported on Linux, macOS, and Windows. Network recording is
disabled when ss is unavailable.
git clone https://github.com/MegalithOfficial/pumpkin-bench.git
cd pumpkin-bench
uv tool install .Without uv, install with python -m pip install .. The command can also be
run from the source tree as python -m pumpkinbench.
pumpkin-bench setupsetup takes a fresh machine to benchmark-ready. It writes a default
bench.toml when none exists, builds a Pumpkin server from upstream master,
boots it once to generate pumpkin.toml and the world, sets the configuration
keys benchmarking requires, builds and installs the pumpkin-pulse plugin,
creates the world snapshot, and finishes with a doctor check. Steps whose
results already exist are skipped, so rerunning it is safe. The first run
builds a server with the release profile and takes a while.
Cached components are refreshed later with:
pumpkin-bench update # server, bot, and pulse
pumpkin-bench update pulseEverything setup does can also be done by hand. Copy the example
configuration and set server_dir:
cp bench.toml.example bench.tomlStart Pumpkin once in that directory to generate pumpkin.toml, then configure
the Java listener and plugin permissions:
[networking.java]
online_mode = false
encryption = false
[plugins]
allow_unsigned = true
allowed_permissions = ["fs.write.data"]Authentication and encryption must be disabled because they are not implemented
by the load bot. Place the pumpkin-pulse .wasm file in the server's
plugins/ directory.
Prepare the world in the state to use for every run, stop the server, and create the baseline snapshot:
pumpkin-bench world create
pumpkin-bench doctordoctor validates the configuration, server directory, plugin, snapshot,
required tools, and benchmark port. Unless bot_bin is configured,
rust-mc-bot is cloned and built on the first scenario that requires bots.
Run a scenario against a local binary:
pumpkin-bench run scatter \
--bin ./target/release/pumpkin \
--label localGit revisions and upstream pull requests can be built directly:
pumpkin-bench run scatter --ref master
pumpkin-bench run scatter --pr 3176Use matrix for a repeated A/B campaign. The following command runs in the
order A, B, A, B and prints a Markdown comparison:
pumpkin-bench matrix scatter \
--a master \
--b pr:3176 \
--runs 2 \
--format markdownpumpkin-bench list
pumpkin-bench report
pumpkin-bench report last3
pumpkin-bench report scatter-b1 --format markdownlist orders completed runs from newest to oldest. Commands accept a full path,
last, last2, and subsequent numbered references, or a unique fragment of a
run directory name.
Compare existing groups independently of a matrix campaign:
pumpkin-bench compare \
--a scatter-a1 scatter-a2 \
--b scatter-b1 scatter-b2 \
--label-a master \
--label-b proposed-change \
--format markdownScenario markers divide a run into named phases. Pass --phase to compare one
phase only:
pumpkin-bench compare \
--a scatter-a1 scatter-a2 \
--b scatter-b1 scatter-b2 \
--phase scatteredReports support text, plain, markdown, and oneline. Comparisons also
support json. Terminal output uses color when connected to a TTY and respects
NO_COLOR.
| Scenario | Workload |
|---|---|
idle |
Empty server; baseline tick cost |
joinburst |
50 simultaneous joins; login and chunk delivery |
mosh |
200 bots at spawn; entity tracking, broadcast, and collision |
scatter |
200 bots distributed across deterministic positions; active-world costs |
scatter-nomobs |
scatter with mob spawning disabled; lower-variance comparison |
stream |
50 bots moved outward every 10 seconds; chunk generation and delivery |
Scenario definitions are TOML files in scenarios/. They control
duration, bot count, daytime, mob spawning, pulse channels, and timed actions.
Available timeline actions are console, bots, stop_bots, scatter,
teleport_rounds, and mark.
- The baseline world is restored before every run.
- Daylight is fixed at noon by default.
- Scatter positions follow the same golden-angle spiral in every run.
- Matrix campaigns alternate builds to distribute thermal and cache effects.
- The first 100 ticks are omitted from whole-run summary statistics.
- Comparisons use the minimum-to-maximum range observed for each side. Overlapping tick-mean ranges are reported as inconclusive, and single-run groups are not assigned a winner.
Mob spawning is not deterministic. Use scatter-nomobs for lower-variance
engine comparisons and scatter when mob-related work belongs in the workload.
At 20 TPS, each tick has a 50 ms budget. Reports include mean, percentile, and maximum tick durations, the number of ticks over budget, per-phase statistics, and 30-second timeline windows. Network reports include bytes sent, delivery times, and peak throughput when network recording is available.
Run directories contain metadata, server and bot logs, pulse sessions, and network samples. Tick data comes from the plugin, so the Pumpkin binary under test does not need benchmark-specific changes.
Licensed under the GNU General Public License v3.0.