Skip to content

About

Benchmark toolkit for Pumpkin servers: reproducible load scenarios, tick and delivery analysis, before/after comparisons with copy-paste reports

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

pumpkin-bench

pumpkin-bench is a benchmark harness for Pumpkin. It runs repeatable bot workloads against local binaries, Git revisions, and pull requests, with tick data collected by pumpkin-pulse.

Each run starts from the same world snapshot. A/B campaigns alternate between both builds and compare the observed range across repeated runs instead of treating a single result as conclusive.

Requirements

  • Python 3.11 or newer
  • A Pumpkin server directory with pumpkin.toml
  • The pumpkin-pulse WebAssembly plugin
  • Git and Rust when building Pumpkin or the load bot from source
  • iproute2 ss for network measurements on Linux (optional)

Tick recording is supported on Linux, macOS, and Windows. Network recording is disabled when ss is unavailable.

Installation

git clone https://github.com/MegalithOfficial/pumpkin-bench.git
cd pumpkin-bench
uv tool install .

Without uv, install with python -m pip install .. The command can also be run from the source tree as python -m pumpkinbench.

Quick start

pumpkin-bench setup

setup takes a fresh machine to benchmark-ready. It writes a default bench.toml when none exists, builds a Pumpkin server from upstream master, boots it once to generate pumpkin.toml and the world, sets the configuration keys benchmarking requires, builds and installs the pumpkin-pulse plugin, creates the world snapshot, and finishes with a doctor check. Steps whose results already exist are skipped, so rerunning it is safe. The first run builds a server with the release profile and takes a while.

Cached components are refreshed later with:

pumpkin-bench update            # server, bot, and pulse
pumpkin-bench update pulse

Manual configuration

Everything setup does can also be done by hand. Copy the example configuration and set server_dir:

cp bench.toml.example bench.toml

Start Pumpkin once in that directory to generate pumpkin.toml, then configure the Java listener and plugin permissions:

[networking.java]
online_mode = false
encryption = false

[plugins]
allow_unsigned = true
allowed_permissions = ["fs.write.data"]

Authentication and encryption must be disabled because they are not implemented by the load bot. Place the pumpkin-pulse .wasm file in the server's plugins/ directory.

Prepare the world in the state to use for every run, stop the server, and create the baseline snapshot:

pumpkin-bench world create
pumpkin-bench doctor

doctor validates the configuration, server directory, plugin, snapshot, required tools, and benchmark port. Unless bot_bin is configured, rust-mc-bot is cloned and built on the first scenario that requires bots.

Usage

Run a scenario against a local binary:

pumpkin-bench run scatter \
  --bin ./target/release/pumpkin \
  --label local

Git revisions and upstream pull requests can be built directly:

pumpkin-bench run scatter --ref master
pumpkin-bench run scatter --pr 3176

Use matrix for a repeated A/B campaign. The following command runs in the order A, B, A, B and prints a Markdown comparison:

pumpkin-bench matrix scatter \
  --a master \
  --b pr:3176 \
  --runs 2 \
  --format markdown

Runs and reports

pumpkin-bench list
pumpkin-bench report
pumpkin-bench report last3
pumpkin-bench report scatter-b1 --format markdown

list orders completed runs from newest to oldest. Commands accept a full path, last, last2, and subsequent numbered references, or a unique fragment of a run directory name.

Compare existing groups independently of a matrix campaign:

pumpkin-bench compare \
  --a scatter-a1 scatter-a2 \
  --b scatter-b1 scatter-b2 \
  --label-a master \
  --label-b proposed-change \
  --format markdown

Scenario markers divide a run into named phases. Pass --phase to compare one phase only:

pumpkin-bench compare \
  --a scatter-a1 scatter-a2 \
  --b scatter-b1 scatter-b2 \
  --phase scattered

Reports support text, plain, markdown, and oneline. Comparisons also support json. Terminal output uses color when connected to a TTY and respects NO_COLOR.

Scenarios

Scenario Workload
idle Empty server; baseline tick cost
joinburst 50 simultaneous joins; login and chunk delivery
mosh 200 bots at spawn; entity tracking, broadcast, and collision
scatter 200 bots distributed across deterministic positions; active-world costs
scatter-nomobs scatter with mob spawning disabled; lower-variance comparison
stream 50 bots moved outward every 10 seconds; chunk generation and delivery

Scenario definitions are TOML files in scenarios/. They control duration, bot count, daytime, mob spawning, pulse channels, and timed actions. Available timeline actions are console, bots, stop_bots, scatter, teleport_rounds, and mark.

Methodology

  • The baseline world is restored before every run.
  • Daylight is fixed at noon by default.
  • Scatter positions follow the same golden-angle spiral in every run.
  • Matrix campaigns alternate builds to distribute thermal and cache effects.
  • The first 100 ticks are omitted from whole-run summary statistics.
  • Comparisons use the minimum-to-maximum range observed for each side. Overlapping tick-mean ranges are reported as inconclusive, and single-run groups are not assigned a winner.

Mob spawning is not deterministic. Use scatter-nomobs for lower-variance engine comparisons and scatter when mob-related work belongs in the workload.

Results

At 20 TPS, each tick has a 50 ms budget. Reports include mean, percentile, and maximum tick durations, the number of ticks over budget, per-phase statistics, and 30-second timeline windows. Network reports include bytes sent, delivery times, and peak throughput when network recording is available.

Run directories contain metadata, server and bot logs, pulse sessions, and network samples. Tick data comes from the plugin, so the Pumpkin binary under test does not need benchmark-specific changes.

License

Licensed under the GNU General Public License v3.0.

About

Benchmark toolkit for Pumpkin servers: reproducible load scenarios, tick and delivery analysis, before/after comparisons with copy-paste reports

Topics

Resources

Stars

3 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages