Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 5 additions & 0 deletions changes/unreleased/Documentation-20260930-205459.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
kind: Documentation
body: '`make memtier` writes every key before the load starts, so its reads measure hits rather than misses, and the performance page is re-measured on AC power with that change.'
time: 2026-09-30T20:54:59.618692+09:00
custom:
Issue: "271"
22 changes: 20 additions & 2 deletions scripts/memtier.sh
Original file line number Diff line number Diff line change
Expand Up @@ -84,9 +84,27 @@ cluster)
;;
esac

# Mostly reads, the way a cache or a config store is used; 32 byte values over 100,000 keys.
# Every key memtier can pick, memtier-0 through memtier-$keys, is written before the load starts,
# so a GET measures a hit rather than the miss path. One MSET per thousand keys keeps this to a
# hundred writes: one SET at a time would take minutes durable and most of an hour in a cluster.
keys=100000
awk -v keys="$keys" 'BEGIN {
value = sprintf("%032d", 0)
for (first = 0; first <= keys; first += 1000) {
last = first + 999
if (last > keys) last = keys
printf "*%d\r\n$4\r\nMSET\r\n", 1 + 2 * (last - first + 1)
for (id = first; id <= last; id++) {
key = "memtier-" id
printf "$%d\r\n%s\r\n$32\r\n%s\r\n", length(key), key, value
}
}
}' | redis-cli -p "$port" --pipe >"$work/prefill.log" 2>&1
grep -q 'errors: 0,' "$work/prefill.log" || { cat "$work/prefill.log" >&2; exit 1; }

# Mostly reads, the way a cache or a config store is used; 32 byte values over those keys.
memtier_benchmark --server 127.0.0.1 --port "$port" --protocol redis \
--threads 4 --clients 50 --ratio 1:10 --data-size 32 \
--key-pattern R:R --key-maximum 100000 --distinct-client-seed \
--key-pattern R:R --key-maximum "$keys" --distinct-client-seed \
--test-time "$seconds" --hide-histogram \
--json-out-file "$root/dist/memtier-$mode.json"
63 changes: 33 additions & 30 deletions website/content/docs/performance.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,24 +11,25 @@ on it can be reproduced with one command from a checkout.
## How it was measured

One machine, nothing else under load: an Apple M4 (10 cores, 16 GB, internal SSD) on macOS
26.6.2, running on **battery power**, with Go 1.26.7 and memtier_benchmark 2.5.1, on the change
that added this page, on top of `2a8076a`. Laptop numbers move with temperature and power
source; the two runs per mode below are there to show by how much.
26.6.2, on AC power, with Go 1.26.7 and memtier_benchmark 2.5.1, on the change that last
updated this page, on top of `f57bc5b`. The two runs per mode below are there to show how far
a run moves on its own.

The load is `make memtier`: 4 threads × 50 connections, one `SET` to every ten `GET`s, 32-byte
values over 100,000 random keys, 60 seconds. Latencies are in milliseconds and are what the client
saw, queueing included.
The load is `make memtier`: every one of the 100,001 keys is written first, so reads hit rather
than miss, then 4 threads × 50 connections send one `SET` to every ten `GET`s, 32-byte values
at random keys, for 60 seconds. Latencies are in milliseconds and are what the client saw,
queueing included.

## Under load

| Mode | Run | Ops/sec | SET p50 | SET p99 | GET p50 | GET p99 |
|---|---|---:|---:|---:|---:|---:|
| memory | 1 | 213,819 | 0.87 | 2.90 | 0.85 | 2.66 |
| memory | 2 | 205,800 | 0.91 | 2.93 | 0.89 | 2.74 |
| durable | 1 | 2,289 | 758 | 2,015 | 3.98 | 12.1 |
| durable | 2 | 2,786 | 758 | 774 | 3.98 | 4.67 |
| cluster | 1 | 472 | 4,555 | 4,653 | 0.087 | 0.143 |
| cluster | 2 | 469 | 4,522 | 4,686 | 0.079 | 0.143 |
| memory | 1 | 210,137 | 0.89 | 2.83 | 0.86 | 2.70 |
| memory | 2 | 204,953 | 0.92 | 2.80 | 0.90 | 2.69 |
| durable | 1 | 2,754 | 758 | 803 | 3.98 | 5.98 |
| durable | 2 | 2,773 | 754 | 934 | 3.98 | 7.71 |
| cluster | 1 | 437 | 4,850 | 5,014 | 0.055 | 0.119 |
| cluster | 2 | 435 | 4,882 | 5,145 | 0.055 | 0.143 |

**memory** is `kvs serve` with no `--data-dir`: about **210,000 operations a second** at under a
millisecond at the median.
Expand All @@ -37,31 +38,32 @@ millisecond at the median.
under the store's one write lock. On this machine a flush takes about 3.7ms, so writes top out
near 270 a second however many clients send them, and two hundred connections queue behind
each other for **three quarters of a second** at the median. Reads are cheap but wait for
whichever flush holds the lock, which is the 4ms at their median. The first run's 2-second SET
tail did not repeat; the machine was on battery.
whichever flush holds the lock, which is the 4ms at their median.

**cluster** is three nodes on this one machine with the load on the leader. Every write is a
consensus round, one at a time, at about 21ms each (below) — roughly **45 writes a second**, and
four and a half seconds of queueing at the median under this many connections. Reads never touch
consensus, which is why they are the fastest reads in the table; it is also why they
[may be behind](../clustering/). The total is low because each connection waits for its own write
before sending its next read.
consensus round, one at a time, at about 22ms each (below) — roughly **43 writes a second**, and
nearly five seconds of queueing at the median under this many connections. The total is low
because each connection waits for its own write before sending its next read, so the node is
idle most of the time, and that is why its reads are the fastest in the table: nothing is
queued ahead of them. They are answered locally without asking the leader, which is also why
they [may be behind](../clustering/).

## Per operation

`make bench BENCH_COUNT=6`, median of the six runs:
`make bench BENCH_COUNT=6`, median of the six runs. The RESP rows come from a second run of that
package alone, after something else on the machine ran through the first:

| Benchmark | What it is | ns/op | B/op | allocs/op |
|---|---|---:|---:|---:|
| `BenchmarkPut` | in-memory write | 101 | 144 | 2 |
| `BenchmarkPutParallel` | the same, from every core | 166 | 144 | 2 |
| `BenchmarkGet` | in-memory read | 61 | 32 | 1 |
| `BenchmarkGetParallel` | the same, from every core | 102 | 32 | 1 |
| `BenchmarkPutDurable` | write with the append log | 3,700,000 | 4,633 | 8 |
| `BenchmarkRESPSet` | `SET` through go-redis over loopback | 14,990 | 5,469 | 32 |
| `BenchmarkPut` | in-memory write | 103 | 144 | 2 |
| `BenchmarkPutParallel` | the same, from every core | 165 | 144 | 2 |
| `BenchmarkGet` | in-memory read | 62 | 32 | 1 |
| `BenchmarkGetParallel` | the same, from every core | 106 | 32 | 1 |
| `BenchmarkPutDurable` | write with the append log | 3,750,000 | 4,662 | 8 |
| `BenchmarkRESPSet` | `SET` through go-redis over loopback | 15,070 | 5,469 | 32 |
| `BenchmarkRESPGet` | `GET` through go-redis over loopback | 14,490 | 3,560 | 22 |
| `BenchmarkRESPSetParallel` | `SET` from every core, one client pool | 7,326 | 5,480 | 32 |
| `BenchmarkClusterPut` | write through a three-node cluster | 21,000,000 | 149,000 | 624 |
| `BenchmarkRESPSetParallel` | `SET` from every core, one client pool | 9,690 | 5,481 | 32 |
| `BenchmarkClusterPut` | write through a three-node cluster | 22,200,000 | 145,000 | 620 |

The store's parallel runs are slower per operation than its serial ones because every caller
shares its one lock; they are there so that sharding it has something to beat.
Expand All @@ -75,5 +77,6 @@ make bench BENCH_COUNT=10 # then compare two builds with benchstat
```

`make memtier` builds `dist/kvs`, starts it on port 16379 (`KVS_BENCH_PORT` moves it), waits
until it answers — in cluster mode, until the other two nodes have joined — and leaves the full
result in `dist/memtier-<mode>.json`. It stops every node it started when it exits.
until it answers — in cluster mode, until the other two nodes have joined — writes every key
once, and leaves the full result in `dist/memtier-<mode>.json`. It stops every node it started
when it exits.
Loading