From baaf5044341217caea1e20cf8fad6bc4fe0e2d6e Mon Sep 17 00:00:00 2001 From: Sung-Kyu Yoo Date: Wed, 30 Sep 2026 20:55:20 +0900 Subject: [PATCH] fix: prefill the memtier keyspace before measuring `make memtier` sent GETs at 100,000 keys nobody had written, so most reads took the miss path: by the end of a durable run about 12% of the keys existed, in a cluster about 2.5%. The script now writes every key first, a thousand to an MSET so that the fill costs a hundred writes rather than minutes of fsyncs or most of an hour of consensus rounds. The performance page is re-measured with the prefill on AC power, and its explanation of the cluster's fast reads is corrected: they are fast because the node is mostly idle waiting on writes, not because they skip consensus. --- .../Documentation-20260930-205459.yaml | 5 ++ scripts/memtier.sh | 22 ++++++- website/content/docs/performance.md | 63 ++++++++++--------- 3 files changed, 58 insertions(+), 32 deletions(-) create mode 100644 changes/unreleased/Documentation-20260930-205459.yaml diff --git a/changes/unreleased/Documentation-20260930-205459.yaml b/changes/unreleased/Documentation-20260930-205459.yaml new file mode 100644 index 0000000..b42865c --- /dev/null +++ b/changes/unreleased/Documentation-20260930-205459.yaml @@ -0,0 +1,5 @@ +kind: Documentation +body: '`make memtier` writes every key before the load starts, so its reads measure hits rather than misses, and the performance page is re-measured on AC power with that change.' +time: 2026-09-30T20:54:59.618692+09:00 +custom: + Issue: "271" diff --git a/scripts/memtier.sh b/scripts/memtier.sh index e0edbe3..ee7a5af 100755 --- a/scripts/memtier.sh +++ b/scripts/memtier.sh @@ -84,9 +84,27 @@ cluster) ;; esac -# Mostly reads, the way a cache or a config store is used; 32 byte values over 100,000 keys. +# Every key memtier can pick, memtier-0 through memtier-$keys, is written before the load starts, +# so a GET measures a hit rather than the miss path. One MSET per thousand keys keeps this to a +# hundred writes: one SET at a time would take minutes durable and most of an hour in a cluster. +keys=100000 +awk -v keys="$keys" 'BEGIN { + value = sprintf("%032d", 0) + for (first = 0; first <= keys; first += 1000) { + last = first + 999 + if (last > keys) last = keys + printf "*%d\r\n$4\r\nMSET\r\n", 1 + 2 * (last - first + 1) + for (id = first; id <= last; id++) { + key = "memtier-" id + printf "$%d\r\n%s\r\n$32\r\n%s\r\n", length(key), key, value + } + } +}' | redis-cli -p "$port" --pipe >"$work/prefill.log" 2>&1 +grep -q 'errors: 0,' "$work/prefill.log" || { cat "$work/prefill.log" >&2; exit 1; } + +# Mostly reads, the way a cache or a config store is used; 32 byte values over those keys. memtier_benchmark --server 127.0.0.1 --port "$port" --protocol redis \ --threads 4 --clients 50 --ratio 1:10 --data-size 32 \ - --key-pattern R:R --key-maximum 100000 --distinct-client-seed \ + --key-pattern R:R --key-maximum "$keys" --distinct-client-seed \ --test-time "$seconds" --hide-histogram \ --json-out-file "$root/dist/memtier-$mode.json" diff --git a/website/content/docs/performance.md b/website/content/docs/performance.md index 299c703..90b1246 100644 --- a/website/content/docs/performance.md +++ b/website/content/docs/performance.md @@ -11,24 +11,25 @@ on it can be reproduced with one command from a checkout. ## How it was measured One machine, nothing else under load: an Apple M4 (10 cores, 16 GB, internal SSD) on macOS -26.6.2, running on **battery power**, with Go 1.26.7 and memtier_benchmark 2.5.1, on the change -that added this page, on top of `2a8076a`. Laptop numbers move with temperature and power -source; the two runs per mode below are there to show by how much. +26.6.2, on AC power, with Go 1.26.7 and memtier_benchmark 2.5.1, on the change that last +updated this page, on top of `f57bc5b`. The two runs per mode below are there to show how far +a run moves on its own. -The load is `make memtier`: 4 threads × 50 connections, one `SET` to every ten `GET`s, 32-byte -values over 100,000 random keys, 60 seconds. Latencies are in milliseconds and are what the client -saw, queueing included. +The load is `make memtier`: every one of the 100,001 keys is written first, so reads hit rather +than miss, then 4 threads × 50 connections send one `SET` to every ten `GET`s, 32-byte values +at random keys, for 60 seconds. Latencies are in milliseconds and are what the client saw, +queueing included. ## Under load | Mode | Run | Ops/sec | SET p50 | SET p99 | GET p50 | GET p99 | |---|---|---:|---:|---:|---:|---:| -| memory | 1 | 213,819 | 0.87 | 2.90 | 0.85 | 2.66 | -| memory | 2 | 205,800 | 0.91 | 2.93 | 0.89 | 2.74 | -| durable | 1 | 2,289 | 758 | 2,015 | 3.98 | 12.1 | -| durable | 2 | 2,786 | 758 | 774 | 3.98 | 4.67 | -| cluster | 1 | 472 | 4,555 | 4,653 | 0.087 | 0.143 | -| cluster | 2 | 469 | 4,522 | 4,686 | 0.079 | 0.143 | +| memory | 1 | 210,137 | 0.89 | 2.83 | 0.86 | 2.70 | +| memory | 2 | 204,953 | 0.92 | 2.80 | 0.90 | 2.69 | +| durable | 1 | 2,754 | 758 | 803 | 3.98 | 5.98 | +| durable | 2 | 2,773 | 754 | 934 | 3.98 | 7.71 | +| cluster | 1 | 437 | 4,850 | 5,014 | 0.055 | 0.119 | +| cluster | 2 | 435 | 4,882 | 5,145 | 0.055 | 0.143 | **memory** is `kvs serve` with no `--data-dir`: about **210,000 operations a second** at under a millisecond at the median. @@ -37,31 +38,32 @@ millisecond at the median. under the store's one write lock. On this machine a flush takes about 3.7ms, so writes top out near 270 a second however many clients send them, and two hundred connections queue behind each other for **three quarters of a second** at the median. Reads are cheap but wait for -whichever flush holds the lock, which is the 4ms at their median. The first run's 2-second SET -tail did not repeat; the machine was on battery. +whichever flush holds the lock, which is the 4ms at their median. **cluster** is three nodes on this one machine with the load on the leader. Every write is a -consensus round, one at a time, at about 21ms each (below) — roughly **45 writes a second**, and -four and a half seconds of queueing at the median under this many connections. Reads never touch -consensus, which is why they are the fastest reads in the table; it is also why they -[may be behind](../clustering/). The total is low because each connection waits for its own write -before sending its next read. +consensus round, one at a time, at about 22ms each (below) — roughly **43 writes a second**, and +nearly five seconds of queueing at the median under this many connections. The total is low +because each connection waits for its own write before sending its next read, so the node is +idle most of the time, and that is why its reads are the fastest in the table: nothing is +queued ahead of them. They are answered locally without asking the leader, which is also why +they [may be behind](../clustering/). ## Per operation -`make bench BENCH_COUNT=6`, median of the six runs: +`make bench BENCH_COUNT=6`, median of the six runs. The RESP rows come from a second run of that +package alone, after something else on the machine ran through the first: | Benchmark | What it is | ns/op | B/op | allocs/op | |---|---|---:|---:|---:| -| `BenchmarkPut` | in-memory write | 101 | 144 | 2 | -| `BenchmarkPutParallel` | the same, from every core | 166 | 144 | 2 | -| `BenchmarkGet` | in-memory read | 61 | 32 | 1 | -| `BenchmarkGetParallel` | the same, from every core | 102 | 32 | 1 | -| `BenchmarkPutDurable` | write with the append log | 3,700,000 | 4,633 | 8 | -| `BenchmarkRESPSet` | `SET` through go-redis over loopback | 14,990 | 5,469 | 32 | +| `BenchmarkPut` | in-memory write | 103 | 144 | 2 | +| `BenchmarkPutParallel` | the same, from every core | 165 | 144 | 2 | +| `BenchmarkGet` | in-memory read | 62 | 32 | 1 | +| `BenchmarkGetParallel` | the same, from every core | 106 | 32 | 1 | +| `BenchmarkPutDurable` | write with the append log | 3,750,000 | 4,662 | 8 | +| `BenchmarkRESPSet` | `SET` through go-redis over loopback | 15,070 | 5,469 | 32 | | `BenchmarkRESPGet` | `GET` through go-redis over loopback | 14,490 | 3,560 | 22 | -| `BenchmarkRESPSetParallel` | `SET` from every core, one client pool | 7,326 | 5,480 | 32 | -| `BenchmarkClusterPut` | write through a three-node cluster | 21,000,000 | 149,000 | 624 | +| `BenchmarkRESPSetParallel` | `SET` from every core, one client pool | 9,690 | 5,481 | 32 | +| `BenchmarkClusterPut` | write through a three-node cluster | 22,200,000 | 145,000 | 620 | The store's parallel runs are slower per operation than its serial ones because every caller shares its one lock; they are there so that sharding it has something to beat. @@ -75,5 +77,6 @@ make bench BENCH_COUNT=10 # then compare two builds with benchstat ``` `make memtier` builds `dist/kvs`, starts it on port 16379 (`KVS_BENCH_PORT` moves it), waits -until it answers — in cluster mode, until the other two nodes have joined — and leaves the full -result in `dist/memtier-.json`. It stops every node it started when it exits. +until it answers — in cluster mode, until the other two nodes have joined — writes every key +once, and leaves the full result in `dist/memtier-.json`. It stops every node it started +when it exits.