diff --git a/README.md b/README.md
index ace8859d..82a1eb05 100644
--- a/README.md
+++ b/README.md
@@ -9,32 +9,32 @@
[](LICENSE)
[](https://github.com/sponsors/skyoo2003)
-ACOR stores a shared Aho-Corasick pattern dictionary in Redis and exposes it
-through a Go library and CLI. Multiple application instances can update the
-same dictionary at runtime while preset engines serve matches from local memory.
+ACOR keeps one Aho-Corasick dictionary in Redis and reaches it from a Go library, a CLI,
+and an experimental server module. Every application instance shares that dictionary,
+updates it at runtime, and matches against a local copy with no Redis I/O on the hot path.
-Typical uses include content filtering, keyword extraction, intrusion detection,
-search highlighting, and real-time text classification.
+Typical uses: content filtering, keyword extraction, intrusion detection, search
+highlighting, real-time text classification.
## Highlights
-- **Shared state** — every application instance uses the same Redis-backed dictionary
+- **Shared state** — every instance reads the same Redis-backed dictionary
- **Runtime updates** — Pub/Sub invalidation, with optional polling for missed messages
-- **Fast reads** — preset engines match locally without Redis I/O on the hot path
-- **Flexible deployment** — standalone Redis, Sentinel, Cluster, Ring, and Valkey
-- **Complete matching API** — occurrences, positions, sets, streams, batches, and parallel matching
+- **Fast reads** — preset engines match locally, 0 round trips
+- **Topologies** — standalone Redis, Sentinel, Cluster, Ring, and Valkey
+- **Full matching API** — occurrences, positions, sets, streams, batches, parallel scans
-## Installation
+## Install
-ACOR requires Go 1.25 or newer and Redis 3.0 or newer, or Valkey 7.2 or newer.
+Requires Go 1.25+ and Redis 3.0+ or Valkey 7.2+.
```sh
go get github.com/skyoo2003/acor/pkg/acor@latest
```
-## Quick Start
+## Quick start
-Start Redis locally, then create a matcher:
+Start Redis locally, then:
```go
@@ -69,10 +69,9 @@ func main() {
}
```
-New collections use the optimized V2 Redis schema by default. `PresetBalanced`
-is the recommended starting point when reads should run locally.
+New collections use the optimized V2 Redis schema.
-## Choosing a Preset
+## Choosing a preset
| Goal | Preset |
| ---- | ------ |
@@ -80,18 +79,25 @@ is the recommended starting point when reads should run locally.
| Highest matching throughput | `PresetSpeed` |
| Lowest memory usage | `PresetMemoryEfficient` |
-Redis remains the source of truth for every preset. See the
-[preset guide](docs/content/guides/preset-engine.md) for trade-offs and the
-[Redis-backed engine guide](docs/content/guides/redis-backed-engine.md) for
-multi-instance invalidation safety.
+Redis stays the source of truth for every preset. Trade-offs are in the
+[preset guide](docs/content/guides/preset-engine.md); multi-instance invalidation is in the
+[Redis-backed engine guide](docs/content/guides/redis-backed-engine.md).
+
+## Large dictionaries (V3)
+
+`OpenVersioned` opens a separate V3 collection with leased snapshots, expected-version
+writes, and background engine replacement. V1/V2 APIs are unaffected. See the
+[V3 guide](docs/content/reference/versioned.md), its
+[performance report](docs/content/reference/versioned-performance.md), and
+[bounded text processing](docs/content/reference/text-processing.md) for `Scan`,
+`MaskText`, and `ReplaceText`.
## Documentation
-Everything past this point lives on the
-[documentation site](https://skyoo2003.github.io/acor/): Redis topologies, the matching and
-streaming API, batch and parallel guides, the schema-V2 migration and benchmarks, deployment
-and troubleshooting, the `acor` CLI, the experimental server module, and what the `v1` line
-promises. Package signatures are on
+The [documentation site](https://skyoo2003.github.io/acor/) covers Redis topologies, the
+matching and streaming API, batch and parallel guides, the V2 schema and benchmarks,
+deployment and troubleshooting, the `acor` CLI, the experimental server module, and what
+the `v1` line promises. Package signatures are on
[pkg.go.dev](https://pkg.go.dev/github.com/skyoo2003/acor/pkg/acor).
## Project
@@ -103,13 +109,3 @@ promises. Package signatures are on
## License
[Apache License 2.0](LICENSE) — Copyright 2016-2026 Sungkyu Yoo
-
-### Large versioned dictionaries
-
-Use `OpenVersioned` for V3 snapshots, atomic expected-version updates and background
-engine replacement. V1/V2 APIs remain compatible. See the [V3 guide](docs/content/reference/versioned.md)
-and [performance report](docs/content/reference/versioned-performance.md).
-
-R2 reuses unchanged downloaded buckets and reduces sparse-engine build memory.
-R3 adds bounded `Scan`, `MaskText`, and `ReplaceText` with original byte/rune spans;
-see [text processing](docs/content/reference/text-processing.md).
diff --git a/changes/unreleased/20260906-docs-condense-and-dedupe.yaml b/changes/unreleased/20260906-docs-condense-and-dedupe.yaml
new file mode 100644
index 00000000..b1ea11d0
--- /dev/null
+++ b/changes/unreleased/20260906-docs-condense-and-dedupe.yaml
@@ -0,0 +1,5 @@
+kind: Documentation
+body: "Condensed the documentation site and removed duplicated pages. `preset-engine` and `redis-backed-engine` each carried a Quick Start, an API Reference block, and a preset table; they now split by question — which preset to pick, versus how the engine stays in sync across instances. `reference/api` repeated the Create/Add/Find/Info/Flush/Close surface in a second preset section and gave a one-line heading to each method; methods are tables now, with prose reserved for the contracts that need it. Eight pages carried hand-written nav footers the theme already renders, `monitoring` listed the same Grafana panels twice, and `deployment` repeated the topology and health-check wiring from other pages. Three defects surfaced: `extending/custom-storage` linked to `guides/presets/`, which does not exist; `cli/_index` and `getting-started/installation` disagreed on the command count; and `reference/api` omitted `FindSetContext` from its context-variants list. Performance and verification reports keep every data table unchanged. No documented behavior changed."
+time: 2026-09-06T23:30:00+09:00
+custom:
+ Issue: "249"
diff --git a/docs/content/_index.md b/docs/content/_index.md
index f47420af..c25208cf 100644
--- a/docs/content/_index.md
+++ b/docs/content/_index.md
@@ -46,11 +46,11 @@ func main() {
Guides
- Use batch, parallel, and Redis-backed engines.
+ Batch, parallel, and preset-optimized engines.
- API Reference
- Review public APIs and storage schemas.
+ Reference
+ Public API, storage schemas, and measurements.
Operations
@@ -62,7 +62,7 @@ func main() {
CLI
- Drive a collection from the shell, twenty commands.
+ Drive a collection from the shell.
Extending
diff --git a/docs/content/cli/_index.md b/docs/content/cli/_index.md
index f98ae658..aa7d305c 100644
--- a/docs/content/cli/_index.md
+++ b/docs/content/cli/_index.md
@@ -5,30 +5,21 @@ weight: 6
# CLI
-`acor` is the third way into the same collection: one binary, twenty commands, every one of
-them a shell over the library. It is the entry point for the things a program should not have
-to be written for — seeding a dictionary, checking what is in one, running a migration,
-grepping a log against keywords that live in Redis.
-
-> **The CLI is not covered by the `v1` compatibility promise.** Quoting
-> [Compatibility](../reference/compatibility/#what-is-not-covered) directly:
->
-> **The CLI.** Flags, output format, and exit codes of the `acor` command can change in
-> any release. If you need a stable contract, call the library rather than parsing CLI
-> output.
->
-> The library (`pkg/acor`) is covered; this is the surface that is not.
-
-## Install
+`acor` is the third way into the same collection: one binary, every command a shell over
+the library. It is for the things a program should not have to be written for — seeding a
+dictionary, checking what is in one, running a migration, grepping a log against keywords
+that live in Redis.
+
+> **The CLI is not covered by the `v1` compatibility promise.** Flags, output format, and
+> exit codes can change in any release. If you need a stable contract, call the library
+> rather than parsing CLI output. See
+> [Compatibility](../reference/compatibility/#what-is-covered).
```bash
go install github.com/skyoo2003/acor/cmd/acor@latest
```
-Full instructions, including verifying the install, are on
-[Getting Started → Installation](../getting-started/installation/#cli-installation).
-
-## The twenty commands
+## The commands
| Group | Commands |
| ----- | -------- |
@@ -39,17 +30,13 @@ Full instructions, including verifying the install, are on
| Migrate | `migrate`, `migrate-rollback` |
| Dictionary (V3) | `dictionary list\|diff\|replace\|status\|copy-v2\|prune` |
-Flags are deliberately not tabulated here. `acor --help` prints the command list followed by
-the flag set's own defaults, so each flag's description lives exactly once — next to where the
-flag is registered — and cannot drift out of step with a copy on this page.
+Flags are deliberately not tabulated here. `acor --help` prints the command list followed
+by the flag set's own defaults, so each flag's description lives exactly once — next to
+where the flag is registered — and cannot drift out of step with a copy on this page.
-`acor version` is the one command that needs no Redis. It prints the version stamped in at
+`acor version` is the one command needing no Redis. It prints the version stamped in at
release build time, or `dev` for a binary you built yourself.
-## Sections
-
-- [Commands](commands/) - Option ordering, batch modes, the four matching shapes, parallel chunking, and when the local cache earns its memory
-
-## Navigation
-
-← [Server](../server/) | [Extending](../extending/) →
+[Commands](commands/) covers what those one-line flag descriptions cannot carry: option
+ordering, batch modes, the matching shapes, parallel chunking, and when the local cache
+earns its memory.
diff --git a/docs/content/cli/commands.md b/docs/content/cli/commands.md
index ed607b6d..eb0e0285 100644
--- a/docs/content/cli/commands.md
+++ b/docs/content/cli/commands.md
@@ -5,30 +5,24 @@ weight: 1
# Commands
-`acor --help` prints the command list and every flag with its default. This page covers the
-behavior those one-line flag descriptions cannot carry: the ordering rule, what each batch
-mode does on failure, how the matching commands differ from one another, and when the local
-cache is worth its memory.
+`acor --help` prints every command and flag with its default. This page covers what those
+one-line descriptions cannot carry.
## Options come before the command
-CLI options must appear before the command. Batch commands accept keywords as
-arguments, or `-` as the only argument to read one keyword per line from stdin:
+Batch commands take keywords as arguments, or `-` as the only argument to read one keyword
+per line from stdin:
```bash
acor -addr localhost:6379 -batch-mode transactional add-many foo bar "hello world"
printf 'foo\nbar\n' | acor -addr localhost:6379 remove-many -
```
-`best-effort` is the default and reports per-keyword failures in JSON while
-returning success; `transactional` fails the command if the whole batch cannot
-be committed.
+`best-effort` is the default: it reports per-keyword failures in JSON and still exits
+successfully. `transactional` fails the command if the whole batch cannot be committed.
## Matching
-Beyond `find` and `find-index`, the matching commands cover the set, span, and
-presence shapes of the same scan:
-
```bash
acor -addr localhost:6379 find-set "he is him"
acor -addr localhost:6379 contains "he is him"
@@ -37,19 +31,17 @@ acor -addr localhost:6379 -match-kind leftmost-longest -whole-word \
```
`find-set` reports each keyword once, `contains` stops at the first match, and
-`find-matches` reports each occurrence with its rune span in scan order.
-`-match-kind` and `-whole-word` apply only to `find-matches`.
+`find-matches` reports each occurrence with its rune span in scan order. `-match-kind` and
+`-whole-word` apply to `find-matches` only.
-`-whole-word` assumes a script that separates words with spaces or punctuation.
-In scripts written without inter-word boundaries (CJK, Thai, …) every adjacent
-character counts as a word character, so nearly every match is treated as
-mid-word and dropped — scan such text without `-whole-word`, or use the library's
-`MatchOptions.WordRune` to supply your own boundary rule.
+`-whole-word` assumes a script that separates words with spaces or punctuation. In scripts
+written without word boundaries (CJK, Thai, …) every adjacent character counts as a word
+character, so nearly every match is treated as mid-word and dropped — scan such text
+without `-whole-word`, or use the library's `MatchOptions.WordRune`.
## Parallel matching
-Parallel matching accepts a text argument, or `-` to read the complete text
-from stdin. `word`, `sentence`, and `line` chunk boundaries are available:
+Takes a text argument, or `-` to read the whole text from stdin:
```bash
acor -addr localhost:6379 -workers 8 -chunk-size 10000 \
@@ -59,11 +51,13 @@ acor -addr localhost:6379 -workers 8 \
find-index-parallel - < document.txt
```
+`-boundary` takes `word`, `sentence`, or `line`.
+
## Local cache and presets
-Use `-cache` with the normal V2 engine, or select a Redis-backed local preset
-engine. Preset mode already keeps a local engine, so `-cache` and `-preset`
-cannot be combined:
+`-cache` enables the local cache on the normal V2 engine; `-preset` selects a Redis-backed
+local engine instead. Preset mode already keeps a local engine, so the two cannot be
+combined:
```bash
acor -addr localhost:6379 -cache find-parallel - < document.txt
@@ -72,23 +66,20 @@ acor -addr localhost:6379 -preset balanced \
-invalidation-poll-interval 30s find-parallel - < document.txt
```
-Available presets are `speed`, `balanced`, and `memory-efficient`; `none` is
-the compatibility-preserving default. Preset mode requires an explicit Redis
-address and does not support `suggest`, `suggest-index`, or migration commands.
-The local cache is most useful for parallel matching, where every chunk shares
-one CLI process; a one-shot `find` invocation has no later lookup to reuse it.
+Presets are `speed`, `balanced`, and `memory-efficient`; `none` is the
+compatibility-preserving default. Preset mode requires an explicit Redis address and
+supports neither `suggest`/`suggest-index` nor the migration commands.
-The trade-offs behind each preset are in
+The cache pays off for parallel matching, where every chunk shares one CLI process. A
+one-shot `find` has no later lookup to reuse it. Preset trade-offs are in
[Guides → Preset-Optimized Engine](../../guides/preset-engine/).
## Versioned dictionaries
-Use `acor -name new-v3-name dictionary list|diff|replace|status|copy-v2|prune`.
-Replacement and copying require `--expected-version`; empty replacements and
-empty V2 sources require `--allow-empty`. Diff and replace read a JSON string
-array from stdin. See the [V3 guide](../../reference/versioned/) for pagination,
-case policy, cutover and command examples.
-
-## Navigation
+```bash
+acor -name new-v3-name dictionary list|diff|replace|status|copy-v2|prune
+```
-← [CLI](../) | [Extending](../../extending/) →
+Replacement and copying require `--expected-version`; empty replacements and empty V2
+sources require `--allow-empty`. Diff and replace read a JSON string array from stdin.
+Pagination, case policy, and cutover: [V3 guide](../../reference/versioned/#cli).
diff --git a/docs/content/extending/_index.md b/docs/content/extending/_index.md
index 2f70752c..bde71e1b 100644
--- a/docs/content/extending/_index.md
+++ b/docs/content/extending/_index.md
@@ -5,12 +5,5 @@ weight: 7
# Extending
-Extend ACOR with custom functionality.
-
-## Sections
-
-- [Custom Storage](custom-storage/) - Why Redis is the only backend, and how to test without one
-
-## Navigation
-
-← [CLI](../cli/)
+- [Custom Storage](custom-storage/) — why Redis is the only backend, and how to test
+ without one
diff --git a/docs/content/extending/custom-storage.md b/docs/content/extending/custom-storage.md
index b8560d8d..0e33aa63 100644
--- a/docs/content/extending/custom-storage.md
+++ b/docs/content/extending/custom-storage.md
@@ -5,38 +5,32 @@ weight: 1
# Custom Storage
-**ACOR does not support custom storage backends.** Redis is the only backend, and
-`Create()` always builds it. This page explains why the interface you may have seen in
-older releases is gone, and what to do instead for the case it was usually reached for:
-testing.
+**ACOR does not support custom storage backends.** Redis is the only one, and `Create()`
+always builds it. This page explains why the interface you may have seen in older releases
+is gone, and what to do instead for the case it was usually reached for: testing.
## What changed in v1.5.0
-Through `v1.4.0` the package exported a `KVStorage` interface, along with `Pipeliner`,
-`Subscription`, `StringMapResult`, `PubSubMessage`, and `Z`. They were exported in
-anticipation of pluggable backends.
+Through `v1.4.0` the package exported `KVStorage` along with `Pipeliner`, `Subscription`,
+`StringMapResult`, `PubSubMessage`, and `Z`, in anticipation of pluggable backends. They
+are unexported as of `v1.5.0`, for two reasons:
-They are unexported as of `v1.5.0`, for two reasons:
-
-- **Nothing could be plugged in.** No public constructor accepted a `KVStorage`, and no
- public function returned one. You could name the interface and write an
- implementation, but there was no way to hand it to ACOR — so the capability the
- interface implied never existed.
+- **Nothing could be plugged in.** No public constructor accepted a `KVStorage` and no
+ public function returned one, so the capability the interface implied never existed.
- **Freezing it would have blocked the feature it was for.** The
[compatibility policy](../../reference/compatibility/) forbids adding a method to an
- exported interface inside `v1`. `KVStorage` has 23 methods; had `v1.5.0` frozen it, a
- future backend needing one more operation would have had nowhere to put it for the
- rest of the `v1` line.
+ exported interface inside `v1`. `KVStorage` has 23 methods; freezing it would have left a
+ future backend needing one more operation nowhere to put it for the rest of `v1`.
-Leaving the shape unfrozen is what keeps pluggable storage possible. Publishing an
-interface later is an addition, which `v1` allows; growing a frozen one is not.
+Publishing an interface later is an addition, which `v1` allows; growing a frozen one is
+not.
## Testing without Redis
-This is what the interface was most often reached for, and it has a better answer that
-works today: [miniredis](https://github.com/alicebob/miniredis), an in-process
-Redis-compatible server. ACOR's own test suite uses it, so the behavior you test against
-is the behavior ACOR is tested against.
+This is what the interface was most often reached for, and
+[miniredis](https://github.com/alicebob/miniredis) answers it today — an in-process
+Redis-compatible server. ACOR's own test suite uses it, so the behavior you test against is
+the behavior ACOR is tested against.
```go
@@ -75,22 +69,18 @@ func TestWithMiniredis(t *testing.T) {
}
```
-`miniredis` covers the commands ACOR issues, including the Lua scripts the V2 schema
-uses for atomic writes. To exercise Pub/Sub-driven cache invalidation across instances,
-point two instances at the same miniredis address.
+miniredis covers the commands ACOR issues, including the Lua scripts V2 uses for atomic
+writes. To exercise Pub/Sub-driven invalidation across instances, point two instances at
+the same miniredis address.
## Reducing Redis traffic
-If the reason for a custom backend was to avoid Redis round trips on reads rather than
-to replace Redis, that is already available without one:
+If the reason for a custom backend was avoiding round trips on reads rather than replacing
+Redis, that is already available:
- **`Preset`** serves reads from a local automaton and touches Redis only on writes and
- invalidations. See [Presets](../../guides/presets/).
+ invalidations — see [Preset-Optimized Engine](../../guides/preset-engine/).
- **`EnableCache`** caches trie data locally and invalidates over Pub/Sub.
-`CacheStats()` reports whether either is working — hit rate, rebuild cost, and
-invalidation lag — without any Redis I/O of its own.
-
-## Navigation
-
-← [Extending](../)
+`CacheStats()` reports whether either is working — hit rate, rebuild cost, invalidation
+lag — with no Redis I/O of its own.
diff --git a/docs/content/getting-started/_index.md b/docs/content/getting-started/_index.md
index d5d03c3d..c66374ff 100644
--- a/docs/content/getting-started/_index.md
+++ b/docs/content/getting-started/_index.md
@@ -5,37 +5,18 @@ weight: 1
# Getting Started
-This section covers everything you need to start using ACOR.
+- [Installation](installation/) — prerequisites, the package, and the `acor` CLI
+- [Quick Start](quick-start/) — a first matcher, and every Redis topology
-## Prerequisites
+## Runnable examples
-- Go >= 1.25
-- Redis >= 3.0 or Valkey >= 7.2
-
-## Installation
-
-```bash
-go get github.com/skyoo2003/acor/pkg/acor@latest
-```
-
-## Next Steps
-
-- [Installation](installation/) - Detailed setup instructions
-- [Quick Start](quick-start/) - Your first ACOR application
-
-## Runnable Examples
-
-Three complete programs live in the repository and build against the current API:
+Three complete programs in the repository build against the current API:
[`examples/basic`](https://github.com/skyoo2003/acor/tree/main/examples/basic),
[`examples/batch`](https://github.com/skyoo2003/acor/tree/main/examples/batch), and
[`examples/parallel`](https://github.com/skyoo2003/acor/tree/main/examples/parallel).
-Each expects a reachable Redis at `localhost:6379` and uses its own collection name,
-so running one does not disturb another.
+Each expects Redis at `localhost:6379` and uses its own collection name, so one does not
+disturb another.
```bash
go run ./examples/basic
```
-
-## Continue Learning
-
-After getting started, explore the [Guides](../guides/) for advanced usage patterns.
diff --git a/docs/content/getting-started/installation.md b/docs/content/getting-started/installation.md
index b6cd173e..b740511d 100644
--- a/docs/content/getting-started/installation.md
+++ b/docs/content/getting-started/installation.md
@@ -5,70 +5,28 @@ weight: 1
# Installation
-## Prerequisites
+Requires **Go 1.25+** and **Redis 3.0+** or **Valkey 7.2+**.
-- **Go**: Version 1.25 or later
-- **Redis**: Version 3.0 or later, **or Valkey**: Version 7.2 or later
+ACOR speaks RESP through [go-redis v9](https://github.com/redis/go-redis), so any
+Redis- or Valkey-compatible server works. RESP3 is negotiated on connect and falls back
+to RESP2 on servers older than `HELLO`.
-ACOR uses the standard RESP protocol via [go-redis v9](https://github.com/redis/go-redis) and works with any Redis- or Valkey-compatible server. RESP3 is negotiated on connect and falls back to RESP2 automatically on servers that predate the `HELLO` command.
-
-## Install the Package
+## Library
```bash
go get github.com/skyoo2003/acor/pkg/acor@latest
```
-## Verify Installation
-
-Create a test file to verify ACOR is installed correctly:
-
-```go
-package main
-
-import (
- "fmt"
- "github.com/skyoo2003/acor/pkg/acor"
-)
-
-func main() {
- args := &acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "test",
- }
-
- ac, err := acor.Create(args)
- if err != nil {
- panic(err)
- }
- defer ac.Close()
-
- fmt.Println("ACOR installed successfully!")
-}
-```
+Verify it with the [Quick Start](../quick-start/) program.
-## CLI Installation
-
-Install the command-line tool:
+## CLI
```bash
go install github.com/skyoo2003/acor/cmd/acor@latest
-```
-Verify the CLI:
-
-```bash
-acor --help
-acor version
+acor --help # command list and every flag's default
+acor version # needs no Redis; prints `dev` for a locally built binary
```
-`acor version` needs no Redis and prints the version stamped at release build
-time (`dev` for a locally built binary).
-
-That is the whole of installing it. What the nineteen commands do — option ordering, batch
-modes, the four matching shapes, parallel chunking, and when the local cache earns its
-memory — is the [CLI](../../cli/) section.
-
-## Next Steps
-
-- [Quick Start](../quick-start/) - Build your first application
-- [Guides](../../guides/) - Learn about batch operations and parallel matching
+What the commands do — option ordering, batch modes, matching shapes, parallel chunking,
+the local cache — is the [CLI](../../cli/) section.
diff --git a/docs/content/getting-started/quick-start.md b/docs/content/getting-started/quick-start.md
index 5a96342f..c0176ad8 100644
--- a/docs/content/getting-started/quick-start.md
+++ b/docs/content/getting-started/quick-start.md
@@ -5,36 +5,28 @@ weight: 2
# Quick Start
-This guide walks you through building your first ACOR application.
-
-## Basic Usage
-
```go
package main
import (
"fmt"
+
"github.com/skyoo2003/acor/pkg/acor"
)
func main() {
- args := &acor.AhoCorasickArgs{
+ ac, err := acor.Create(&acor.AhoCorasickArgs{
Addr: "localhost:6379",
Name: "sample",
- }
-
- ac, err := acor.Create(args)
+ })
if err != nil {
panic(err)
}
defer ac.Close()
- keywords := []string{"he", "her", "him"}
- for _, k := range keywords {
- if _, err := ac.Add(k); err != nil {
- panic(err)
- }
+ if _, err := ac.AddMany([]string{"he", "her", "him"}, nil); err != nil {
+ panic(err)
}
matched, err := ac.Find("he is him")
@@ -42,68 +34,45 @@ func main() {
panic(err)
}
fmt.Println(matched)
-
- if err := ac.Flush(); err != nil {
- panic(err)
- }
}
```
-## Redis Topologies
+## Redis topologies
-ACOR supports multiple Redis configurations:
-
-### Standalone
+Pick exactly one set of connection fields — mixing them returns
+`ErrRedisConflictingTopology`.
```go
-args := &acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Password: "",
- DB: 0,
- Name: "sample",
-}
-```
-
-### Sentinel
+// Standalone
+args := &acor.AhoCorasickArgs{Addr: "localhost:6379", Name: "sample"}
-```go
-args := &acor.AhoCorasickArgs{
+// Sentinel — Addrs plus MasterName
+args = &acor.AhoCorasickArgs{
Addrs: []string{"localhost:26379", "localhost:26380"},
MasterName: "mymaster",
- Password: "",
- DB: 0,
Name: "sample",
}
-```
-
-### Cluster
-```go
-args := &acor.AhoCorasickArgs{
- Addrs: []string{"localhost:7000", "localhost:7001", "localhost:7002"},
- Password: "",
- Name: "sample",
+// Cluster — Addrs without MasterName
+args = &acor.AhoCorasickArgs{
+ Addrs: []string{"localhost:7000", "localhost:7001"},
+ Name: "sample",
}
-```
-
-### Ring
-```go
-args := &acor.AhoCorasickArgs{
- RingAddrs: map[string]string{
- "shard-1": "localhost:7000",
- "shard-2": "localhost:7001",
- },
- Password: "",
- DB: 0,
- Name: "sample",
+// Ring — shard name to address
+args = &acor.AhoCorasickArgs{
+ RingAddrs: map[string]string{"shard-1": "localhost:7000", "shard-2": "localhost:7001"},
+ Name: "sample",
}
```
-## Next Steps
+`Password`, `DB`, and the timeout and pool fields apply to every topology; `DB` is
+rejected together with `Addrs`. Full field list:
+[API Reference](../../reference/api/#ahocorasickargs).
+
+## Next
-- [Batch Operations](../../guides/batch-operations/) - Optimize bulk operations
-- [Parallel Matching](../../guides/parallel-matching/) - Process large texts efficiently
-- [Redis-Backed Engine](../../guides/redis-backed-engine/) - Redis persistence with local speed
-- [Match Details](../../reference/api/#findmatches) - Ordered rune spans, matching options, and streaming
-- [API Reference](../../reference/api/) - Complete API documentation
+[Batch operations](../../guides/batch-operations/) ·
+[Parallel matching](../../guides/parallel-matching/) ·
+[Preset engine](../../guides/preset-engine/) ·
+[API reference](../../reference/api/)
diff --git a/docs/content/guides/_index.md b/docs/content/guides/_index.md
index 97a8988b..7fe0718e 100644
--- a/docs/content/guides/_index.md
+++ b/docs/content/guides/_index.md
@@ -5,15 +5,7 @@ weight: 2
# Guides
-Practical guides for using ACOR effectively.
-
-## Available Guides
-
-- [Batch Operations](batch-operations/) - Optimize bulk keyword operations
-- [Parallel Matching](parallel-matching/) - Process large texts with multiple workers
-- [Redis-Backed Engine](redis-backed-engine/) - Redis persistence with local preset-optimized speed
-- [Preset-Optimized Engine](preset-engine/) - What each preset trades away, and how to pick one
-
-## Navigation
-
-← [Getting Started](../getting-started/) | [Reference](../reference/) →
+- [Batch Operations](batch-operations/) — bulk adds, removes, and multi-text scans
+- [Parallel Matching](parallel-matching/) — chunking large texts across workers
+- [Preset-Optimized Engine](preset-engine/) — what each preset trades away, and how to pick one
+- [Redis-Backed Engine](redis-backed-engine/) — architecture, invalidation safety, topologies
diff --git a/docs/content/guides/batch-operations.md b/docs/content/guides/batch-operations.md
index 9b9355cd..16324780 100644
--- a/docs/content/guides/batch-operations.md
+++ b/docs/content/guides/batch-operations.md
@@ -5,13 +5,8 @@ weight: 1
# Batch Operations
-ACOR supports batch operations for better performance when working with multiple keywords.
-
-## Overview
-
-Batch operations reduce network round-trips by grouping multiple operations together.
-
-## Adding Multiple Keywords
+`AddMany` and `RemoveMany` commit a whole batch in one transaction — two round trips
+regardless of batch size, against one per keyword for a loop over `Add`.
```go
@@ -26,52 +21,31 @@ fmt.Printf("Added: %d, Failed: %d, Skipped: %d\n",
len(result.Added), len(result.Failed), len(result.Skipped))
```
-## Batch Modes
-
-### Transactional
-
-Rolls back all changes if any error occurs:
-
-```go
-result, err := ac.AddMany(keywords, &acor.BatchOptions{
- Mode: acor.BatchModeTransactional,
-})
-```
-
-### BestEffort
+`nil` options mean best-effort, which is also what `BatchOptions` with no `Mode` selects.
-Continues on errors and returns partial results:
+## Modes
-```go
-result, err := ac.AddMany(keywords, &acor.BatchOptions{
- Mode: acor.BatchModeBestEffort,
-})
-```
+| Mode | On a per-keyword failure |
+| ---- | ------------------------ |
+| `BatchModeBestEffort` (default) | Commits the rest; failures land in `result.Failed`, and the call still returns success |
+| `BatchModeTransactional` | Rolls the whole batch back and returns an error |
-## Removing Multiple Keywords
+Inspect what best-effort dropped:
-
```go
-result, err := ac.RemoveMany([]string{"he", "her"}, nil)
-if err != nil {
- panic(err)
+for _, ke := range result.Failed {
+ fmt.Printf("%q failed: %v\n", ke.Keyword, ke.Error)
}
-fmt.Printf("Removed: %d\n", len(result.Removed))
```
-`nil` options mean best-effort, the same default `AddMany` uses. Pass options to
-control the batch mode:
+`result.Skipped` holds duplicate adds and absent removes — not errors. Field-by-field
+shapes for `BatchResult` and `KeywordError` are in the
+[API reference](../../reference/api/#batchresult).
-
-```go
-result, err := ac.RemoveMany([]string{"he", "her"}, &acor.BatchOptions{
- Mode: acor.BatchModeTransactional,
-})
-_ = result
-_ = err
-```
+## Scanning many texts
-## Finding Matches in Multiple Texts
+`FindMany` loads the automaton once and scans every text against that one snapshot, so it
+costs the same round trip as a single `Find`. The result is keyed by the input text.
```go
@@ -86,37 +60,8 @@ for text, matches := range results {
}
```
-## Batch Result Structure
-
-```go
-type BatchResult struct {
- Added []string // Keywords successfully added
- Removed []string // Keywords successfully removed
- Failed []KeywordError // Keywords that could not be processed, with their errors
- Skipped []string // Keywords skipped (e.g. duplicates in the input)
-}
-
-type KeywordError struct {
- Keyword string // The keyword that caused the error
- Error error // The error that occurred while processing it
-}
-```
-
-Inspect failures in `BatchModeBestEffort`:
-
-```go
-for _, ke := range result.Failed {
- fmt.Printf("%q failed: %v\n", ke.Keyword, ke.Error)
-}
-```
-
-## Performance Tips
-
-1. Use batch sizes between 100-1000 keywords
-2. Use `BatchModeTransactional` when data consistency is critical
-3. Use `BatchModeBestEffort` when partial success is acceptable
-
-## Next Steps
+## Sizing
-- [Parallel Matching](../parallel-matching/) - Process large texts efficiently
-- [API Reference](../../reference/api/) - Complete API documentation
+100–1,000 keywords per batch is the useful range: below it the transaction overhead
+dominates, above it the single Lua commit grows large enough to block Redis noticeably.
+Measured costs are on the [benchmarks page](../../reference/benchmarks/#bulk-load-addmany).
diff --git a/docs/content/guides/parallel-matching.md b/docs/content/guides/parallel-matching.md
index b799bf4e..da0f8dfe 100644
--- a/docs/content/guides/parallel-matching.md
+++ b/docs/content/guides/parallel-matching.md
@@ -5,21 +5,17 @@ weight: 2
# Parallel Matching
-For large texts, use parallel matching to scan with multiple goroutines.
-
-## Overview
-
-Parallel matching splits text into chunks and processes them concurrently, significantly improving performance for large inputs.
-
-## Basic Usage
+`FindParallel` and `FindIndexParallel` split one text into chunks and scan them
+concurrently. The automaton is loaded once per call, so the Redis cost is the same as a
+serial `Find` however many chunks result.
```go
matches, err := ac.FindParallel(largeText, &acor.ParallelOptions{
- Workers: 4,
- ChunkSize: 1000,
- AutoOverlap: true,
- Boundary: acor.ChunkBoundaryWord,
+ Workers: 4,
+ ChunkSize: 1000, // runes per base chunk; must be positive
+ AutoOverlap: true,
+ Boundary: acor.ChunkBoundaryWord,
})
if err != nil {
panic(err)
@@ -27,108 +23,39 @@ if err != nil {
_ = matches
```
-## Chunk Boundaries
-
-Boundary selection chooses where base chunks end. Enable `AutoOverlap` to protect
-matches across every boundary; boundary selection alone cannot guarantee this.
-
-### ChunkBoundaryWord (default)
-
-Splits at word boundaries, ideal for natural language text:
-
-```go
-opts := &acor.ParallelOptions{
- Workers: 4,
- ChunkSize: 1000,
- AutoOverlap: true,
- Boundary: acor.ChunkBoundaryWord,
-}
-```
-
-### ChunkBoundaryLine
-
-Splits at line breaks, ideal for log files:
-
-```go
-opts := &acor.ParallelOptions{
- Workers: 4,
- ChunkSize: 1000,
- AutoOverlap: true,
- Boundary: acor.ChunkBoundaryLine,
-}
-```
-
-### ChunkBoundarySentence
-
-Splits at sentence endings, ideal for document processing:
-
-```go
-opts := &acor.ParallelOptions{
- Workers: 4,
- ChunkSize: 1000,
- AutoOverlap: true,
- Boundary: acor.ChunkBoundarySentence,
-}
-```
-
-## Automatic Boundary Protection
-
-`AutoOverlap: true` uses the longest keyword in the engine loaded for this call.
-Each base chunk owns its starting positions and extends its right search range by
-`max(Overlap, longest keyword rune length - 1)`. Matches starting in the extension
-belong to the next base chunk and are excluded from the current chunk. This also
-finds keywords longer than `ChunkSize`, including Korean and emoji keywords.
-No extra Redis query is needed, and all workers use the same dictionary snapshot.
+## Boundaries
-`FindParallel` returns each keyword once, in chunk order and then first-match scan
-order within each chunk. That order need not equal serial scan order.
-`FindIndexParallel` returns sorted unique rune positions in the original text.
+Only `Boundary` changes between these; everything else stays as above.
-The default `AutoOverlap` is `false`, preserving legacy overlapping chunks.
-In that mode, an insufficient `Overlap` can miss boundary matches, and a keyword
-longer than a chunk may not fit in any chunk. Options are copied before normalization.
-`ChunkSize` must be positive; empty input performs no Redis reads.
+| Boundary | Splits at | Suited to |
+| -------- | --------- | --------- |
+| `ChunkBoundaryWord` (default) | Whitespace | Natural language |
+| `ChunkBoundaryLine` | Newlines | Log files |
+| `ChunkBoundarySentence` | `.` `!` `?` | Documents |
-## Performance Tuning
-
-### Worker Count
-
-Choose worker count based on CPU cores:
-
-```go
-workers := runtime.NumCPU()
-```
+Boundary choice alone does not protect a keyword that straddles a split. `AutoOverlap`
+does.
-For I/O-bound workloads, consider higher counts:
-
-```go
-workers := runtime.NumCPU() * 2
-```
-
-### Chunk Size
-
-Control chunk size with the `ChunkSize` option:
-
-```go
-opts := &acor.ParallelOptions{
- Workers: 4,
- ChunkSize: 10000, // 10,000 runes per base chunk
-}
-```
+## Boundary protection
-## When to Use Parallel Matching
+`AutoOverlap: true` extends each base chunk to the right by
+`max(Overlap, longest keyword rune length - 1)`, using the longest keyword in the engine
+already loaded for this call — no extra Redis read, and every worker shares one snapshot.
+Matches starting inside an extension belong to the next chunk and are not double-counted.
+This also catches keywords longer than `ChunkSize`, including Korean and emoji.
-- Text size > 100KB
-- Many pattern matches expected
-- CPU cores available for parallel work
+`AutoOverlap` defaults to `false`, which keeps the legacy behavior: an `Overlap` shorter
+than your longest keyword silently misses boundary matches, and a keyword longer than
+`ChunkSize` may fit in no chunk at all.
-## When to Avoid
+## Result order
-- Small texts (< 10KB)
-- Single-match scenarios
-- Resource-constrained environments
+`FindParallel` reports each keyword once, in chunk order and then first-match order
+within a chunk — which need not equal serial scan order. `FindIndexParallel` returns
+sorted, unique rune positions in the original text. Empty input performs no Redis read.
-## Next Steps
+## When it pays
-- [Batch Operations](../batch-operations/) - Optimize bulk keyword operations
-- [API Reference](../../reference/api/) - Complete API documentation
+Worth it above roughly 100 KB of text, with cores to spare. Below ~10 KB, or for a
+single expected match, the chunking overhead outweighs it. Start `Workers` at
+`runtime.NumCPU()`.
diff --git a/docs/content/guides/preset-engine.md b/docs/content/guides/preset-engine.md
index c56f0fa9..ae035583 100644
--- a/docs/content/guides/preset-engine.md
+++ b/docs/content/guides/preset-engine.md
@@ -5,120 +5,47 @@ weight: 3
# Preset-Optimized Engine
-ACOR provides a Redis-backed Aho-Corasick engine with selectable architecture presets. Created via the unified `Create` API with a `Preset` field. Writes go to Redis atomically (V2 Lua scripts with optimistic locking); reads hit the local engine with no Redis I/O.
-
-## When to Use
-
-- Production deployments requiring Redis persistence
-- Distributed systems with multiple instances sharing a keyword collection
-- High-throughput text matching with zero read-latency on the hot path
-- Applications needing both durability and speed
-
-## Quick Start
+Setting `Preset` on `AhoCorasickArgs` builds a local automaton next to the Redis
+collection: writes still go to Redis atomically, reads never touch it. This page is about
+picking a preset. How the engine stays in sync across instances is
+[Redis-Backed Engine](../redis-backed-engine/).
```go
-package main
-
-import (
- "fmt"
- "github.com/skyoo2003/acor/pkg/acor"
-)
-
-func main() {
- ac, _ := acor.Create(&acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-collection",
- Preset: acor.PresetBalanced,
- })
- defer ac.Close()
-
- ac.Add("he")
- ac.Add("her")
- ac.Add("him")
-
- matches, _ := ac.Find("he is him")
- fmt.Println(matches) // [he him]
-
- positions, _ := ac.FindIndex("he is him")
- fmt.Println(positions) // map[he:[0] him:[6]]
-
- info, _ := ac.Info()
- fmt.Printf("Keywords: %d, Nodes: %d, Memory: %d bytes\n",
- info.Keywords, info.Nodes, info.MemoryBytes)
-}
-```
-
-## Architecture Presets
-
-Each preset optimizes for a different trade-off between speed, memory, and feature set. The preset is fixed at creation time.
-
-| Preset | Engine | Best For | Trade-off |
-|--------|--------|----------|-----------|
-| `PresetSpeed` | Full DFA + flat array trie + compact alphabet mapping | Real-time packet inspection, high-speed log scanning, latency-critical paths | Higher memory proportional to states x alphabet size |
-| `PresetBalanced` | Double-Array Trie + Banded DFA + output link compression | General-purpose backend keyword filtering, search engines | Balanced speed and memory |
-| `PresetMemoryEfficient` | Map-based sparse trie + Bloom filter pre-filtering + standard NFA | Large-scale domain blocking, malware signature matching, millions of patterns | Slower search due to failure link traversal and map lookups |
-
-### Choosing a Preset
-
-- **Start with `PresetBalanced`** — it provides the best speed-to-memory ratio for most workloads.
-- Use `PresetSpeed` when latency is critical and memory is available.
-- Use `PresetMemoryEfficient` when you have millions of patterns and memory is constrained.
-
-`PresetSpeed` measured fastest on every query shape on the
-[benchmarks page](../../reference/benchmarks/), while `PresetBalanced` trades some of
-that for a much smaller transition table. Measure your own corpus rather than
-choosing by name.
-
-## Case Sensitivity
-
-By default, matching is case-insensitive. Enable case-sensitive matching when needed:
-
-```go
-ac, _ := acor.Create(&acor.AhoCorasickArgs{
+ac, err := acor.Create(&acor.AhoCorasickArgs{
Addr: "localhost:6379",
Name: "my-collection",
Preset: acor.PresetBalanced,
- CaseSensitive: true,
+ CaseSensitive: false, // default; matching is case-insensitive
})
+if err != nil {
+ panic(err)
+}
defer ac.Close()
```
-## API Reference
-
-```go
-// Create
-ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-collection",
- Preset: acor.PresetBalanced,
-})
-defer ac.Close()
+The preset is fixed at creation. `Info()` reports the resulting `Keywords`, `Nodes`,
+`MemoryBytes`, and `TrieDepth`.
-// Add/Remove — returns 1 if changed, 0 if no-op
-ac.Add("keyword")
-ac.Remove("keyword")
+## The three presets
-// Find (0 RTT on hot path — reads from local engine)
-matches, _ := ac.Find("text") // ([]string, error)
-positions, _ := ac.FindIndex("text") // (map[string][]int, error)
+| Preset | Engine | Best for | Costs |
+|--------|--------|----------|-------|
+| `PresetSpeed` | Full DFA + flat array trie + compact alphabet map | Packet inspection, high-rate log scanning, latency-critical paths | Memory proportional to states × alphabet size |
+| `PresetBalanced` | Double-Array Trie + Banded DFA + output link compression | General backend filtering, search | Neither extreme |
+| `PresetMemoryEfficient` | Map-based sparse trie + Bloom pre-filter + NFA | Millions of patterns under a memory cap | Slower search: failure-link traversal and map lookups |
-// Stats
-info, err := ac.Info() // (*AhoCorasickInfo, error)
+**Start with `PresetBalanced`.** Move to `PresetSpeed` when latency is the constraint and
+memory is not; to `PresetMemoryEfficient` when it is the other way round.
-// Reset
-ac.Flush()
-```
-
-## Refresh and Cancellation
-
-Stale reads share a reload but respond independently to request cancellation.
-A failed reload returns a search error and is retried by a later request. Optional
-version polling helps recover from missed Pub/Sub messages; its interval does not
-guarantee freshness during failures. See [invalidation safety](../redis-backed-engine/#invalidation-safety)
-for the refresh lifecycle and failure counters.
+`PresetSpeed` measured fastest on every query shape on the
+[benchmarks page](../../reference/benchmarks/), and `PresetBalanced` gives up some of that
+for a much smaller transition table. Measure your own corpus before choosing by name.
-## Next Steps
+## Refresh behavior
-- [Redis-Backed Engine](../redis-backed-engine/) - Redis persistence details
-- [API Reference](../../reference/api/) - Complete API documentation
+Stale reads share one reload and each observes its own cancellation. A failed reload
+returns an error to the search rather than silently serving the old engine, and a later
+search retries. Optional version polling recovers from missed Pub/Sub messages; its
+interval is not a freshness bound. Full lifecycle and failure counters:
+[invalidation safety](../redis-backed-engine/#invalidation-safety).
diff --git a/docs/content/guides/redis-backed-engine.md b/docs/content/guides/redis-backed-engine.md
index ca4f4f21..e3806446 100644
--- a/docs/content/guides/redis-backed-engine.md
+++ b/docs/content/guides/redis-backed-engine.md
@@ -5,20 +5,18 @@ weight: 4
# Redis-Backed Engine
-The preset-optimized Redis mode (enabled via `Preset` on `AhoCorasickArgs`) combines Redis persistence with a local preset-optimized automaton. Redis is the source of truth; reads hit the local engine with no Redis I/O on the hot path.
+With `Preset` set, Redis stays the source of truth while every instance serves reads from
+its own automaton. This page covers how that stays consistent. Which preset to pick is
+[Preset-Optimized Engine](../preset-engine/).
-> **Redis or Valkey:** ACOR connects with [go-redis v9](https://github.com/redis/go-redis) over the standard RESP protocol, so any Redis (>= 3.0) or Valkey (>= 7.2) server works, including Standalone, Sentinel, Cluster, and Ring topologies. Cross-instance cache invalidation uses server Pub/Sub, which behaves identically on Redis and Valkey. Both are covered by the integration test suite in CI.
-
-## When to Use
-
-- Distributed deployments across multiple instances
-- Need for Redis persistence and cross-instance synchronization
-- Want preset-optimized local speed without giving up Redis durability
-- Migrating from the original `AhoCorasick` for better read performance
+> **Redis or Valkey.** ACOR connects over RESP via
+> [go-redis v9](https://github.com/redis/go-redis), so Redis 3.0+ or Valkey 7.2+ works in
+> Standalone, Sentinel, Cluster, and Ring topologies. Cross-instance invalidation uses
+> server Pub/Sub, which behaves identically on both. CI covers both.
## Architecture
-```
+```text
Write Path
Instance A ──Add()──▶ Lua Script (optimistic lock) ──▶ Redis
│
@@ -32,16 +30,17 @@ Instance B ◀────────────────┘
Instance A ──Find()──▶ local engine (0 RTT)
```
-- **Writes**: V2 Lua scripts with optimistic locking (up to 3 retries with backoff)
-- **Reads**: Local preset-optimized automaton — no Redis I/O
-- **Invalidation**: Redis Pub/Sub notifies all instances on mutation
-- **Reload failures**: The previous engine is retained, but the waiting search returns an error. A later search retries the reload.
+| Path | Behavior |
+| ---- | -------- |
+| Writes | V2 Lua scripts with optimistic locking, up to 3 retries with backoff |
+| Reads | Local automaton, no Redis I/O |
+| Invalidation | Redis Pub/Sub on every mutation |
+| Failed reload | Previous engine is retained but the waiting search returns an error; a later search retries |
-## Invalidation Safety
+## Invalidation safety
-Redis Pub/Sub is best effort. In a multi-instance deployment, set
-`InvalidationPollInterval` to recover from dropped invalidations after a successful
-version poll and reload:
+Pub/Sub is best effort — a disconnected subscriber misses invalidations. In multi-instance
+deployments set `InvalidationPollInterval`:
```go
@@ -54,138 +53,43 @@ args := &acor.AhoCorasickArgs{
_ = args
```
-The zero value disables polling. Polling only applies to Preset mode; normal
-invalidation still uses Pub/Sub. Each poll reads only the existing `version` hash
-field; a changed version marks the engine stale, and the next search reads the
-full snapshot. The interval is not a freshness upper bound: Redis failures, query
-latency, and rebuild time can delay recovery. This mode provides eventual refresh,
-not strong consistency or a guarantee of fresh results during an outage.
+Zero disables polling, and polling applies to Preset mode only. Each poll reads just the
+`version` hash field; a change marks the engine stale and the next search loads the full
+snapshot.
-Concurrent stale reads share a reload, while each request independently observes
-its own context cancellation. Canceling one waiter leaves the others running.
-The instance lifetime owns the job; all waiters leaving or `Close` cancels it.
-Redis reads and engine builds run outside the state lock. A generation check
-rejects and retries snapshots overtaken by local writes or invalidations.
+**The interval is not a freshness bound.** Redis failures, query latency, and rebuild time
+all delay recovery. This is eventual refresh, not strong consistency.
-`CacheStats().PresetReloadFailures` counts failed shared jobs once per job;
-`PresetPollFailures` counts failed version polls. Cancellation is excluded.
-Both counters are cumulative per instance and remain zero outside Preset mode.
-Monitor these counters alongside application search errors; a retained previous
-engine does not turn a failed reload into a successful response.
-
-## Quick Start
-
-
-```go
-package main
-
-import (
- "fmt"
- "github.com/skyoo2003/acor/pkg/acor"
-)
-
-func main() {
- ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-collection",
- Preset: acor.PresetBalanced,
- CaseSensitive: false,
- })
- if err != nil {
- panic(err)
- }
- defer ac.Close()
-
- added, err := ac.Add("hello")
- if err != nil {
- panic(err)
- }
- fmt.Printf("Added: %d\n", added)
-
- matches, err := ac.Find("hello world")
- if err != nil {
- panic(err)
- }
- fmt.Println(matches) // [hello]
-}
-```
+Concurrent stale reads share one reload job while each request observes its own context
+cancellation — cancelling one waiter leaves the others running, and all waiters leaving
+(or `Close`) cancels the job. Redis reads and engine builds run outside the state lock,
+and a generation check rejects snapshots overtaken by local writes.
-## AhoCorasickArgs (Preset field)
+Watch `CacheStats().PresetReloadFailures` (once per failed shared job) and
+`PresetPollFailures` (per failed version poll) alongside your search errors. Cancellation
+is excluded from both; both stay zero outside Preset mode. A retained previous engine does
+not turn a failed reload into a successful response. See
+[Monitoring](../../operations/monitoring/).
-The `AhoCorasickArgs` struct includes a `Preset` field for selecting the local engine architecture:
+## Topologies and connection tuning
-```go
-type AhoCorasickArgs struct {
- // ... Addr, Addrs, RingAddrs, Password, DB, Name ...
- Preset Preset // Architecture preset; zero value is PresetNone (original mode). Set it (e.g. PresetBalanced) to enable preset mode
- CaseSensitive bool // Enable case-sensitive matching (default: false)
- // ... other fields ...
-}
-```
+All four topologies are configured through the connection fields on `AhoCorasickArgs` —
+see [Redis topologies](../../getting-started/quick-start/#redis-topologies).
-All standard Redis topologies are supported (Standalone, Sentinel, Cluster, Ring) via the connection fields on `AhoCorasickArgs`.
+`DialTimeout`, `ReadTimeout`, `WriteTimeout`, `MaxRetries`, and `PoolSize` pass straight
+to go-redis for every topology. Zero keeps the go-redis default; `-1` disables the
+read/write timeouts or command retries where supported.
-Connection resilience can be tuned with `DialTimeout`, `ReadTimeout`,
-`WriteTimeout`, `MaxRetries`, and `PoolSize`. These values pass directly to
-go-redis for every topology; zero keeps the go-redis default. Use `-1` to
-disable read/write timeouts or command retries where supported.
+## Preset mode versus plain V2
-## Preset Selection
-
-The same [architecture presets](../preset-engine/#architecture-presets) are available:
-
-| Preset | Use Case |
-|--------|----------|
-| `PresetSpeed` | Latency-critical, memory available |
-| `PresetBalanced` | Default — best speed-to-memory ratio |
-| `PresetMemoryEfficient` | Millions of patterns, memory constrained |
-
-If the `Preset` field is left unset, it defaults to `PresetNone`, which runs the
-original `AhoCorasick` mode (not the preset-optimized engine). You must set
-`Preset` explicitly (e.g. `PresetBalanced`) to enable this engine.
-
-## API Reference
-
-```go
-// Create
-ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-collection",
- Preset: acor.PresetBalanced,
-})
-
-// Add/Remove
-added, err := ac.Add("keyword") // (int, error)
-removed, err := ac.Remove("keyword") // (int, error)
-
-// Find (0 RTT on hot path — reads from local engine)
-matches, err := ac.Find("text") // ([]string, error)
-positions, err := ac.FindIndex("text") // (map[string][]int, error)
-
-// Stats
-info, err := ac.Info() // (*AhoCorasickInfo, error)
-
-// Flush and Close
-err := ac.Flush()
-err := ac.Close()
-```
-
-## Comparison with AhoCorasick
-
-| Feature | `AhoCorasick` (no Preset) | `AhoCorasick` (with Preset) |
-|---------|--------------|-----------------|
-| Read latency | 1 RTT (V2) or cached | 0 RTT (local engine) |
+| | No `Preset` | With `Preset` |
+|---|---|---|
+| Read latency | 1 RTT, or 0 with `EnableCache` | 0 RTT |
| Write latency | Lua script | Lua script + optimistic lock |
| Cross-instance sync | Pub/Sub cache invalidation | Pub/Sub engine rebuild |
| Schema | V1 or V2 | V2 only |
-| Presets | N/A | Speed, Balanced, MemoryEfficient |
-| Suggest/SuggestIndex | Yes | No (`ErrSuggestRequiresRedis`) |
-| Batch operations | Yes | Yes |
-| Parallel matching | Yes | Yes |
-
-Use a `Preset`-optimized `AhoCorasick` when you need the fastest possible reads in a distributed setup and can accept the V2-only constraint.
-
-## Next Steps
+| `Suggest` / `SuggestIndex` | Yes | No — `ErrSuggestRequiresRedis` |
+| Batch, parallel matching | Yes | Yes |
-- [Preset-Optimized Engine](../preset-engine/) - Redis-backed engine with local speed
-- [API Reference](../../reference/api/) - Complete API documentation
+`Preset` is unset by default (`PresetNone`), which runs the original mode. Choose it when
+you need the fastest reads across several instances and can accept V2-only, no-`Suggest`.
diff --git a/docs/content/operations/_index.md b/docs/content/operations/_index.md
index 6695f240..a9bbe2ab 100644
--- a/docs/content/operations/_index.md
+++ b/docs/content/operations/_index.md
@@ -5,14 +5,6 @@ weight: 4
# Operations
-Operational guides for running ACOR in production.
-
-## Sections
-
-- [Deployment](deployment/) - Deploy ACOR in various environments
-- [Monitoring](monitoring/) - Monitor ACOR performance
-- [Troubleshooting](troubleshooting/) - Diagnose and fix issues
-
-## Navigation
-
-← [Reference](../reference/) | [Server](../server/) →
+- [Deployment](deployment/) — topologies, Kubernetes, Docker Compose, health endpoints
+- [Monitoring](monitoring/) — cache statistics, and the `server/*` metrics, logs, and traces
+- [Troubleshooting](troubleshooting/) — the errors you are most likely to hit, and their fixes
diff --git a/docs/content/operations/deployment.md b/docs/content/operations/deployment.md
index 18a33464..677ab042 100644
--- a/docs/content/operations/deployment.md
+++ b/docs/content/operations/deployment.md
@@ -5,85 +5,26 @@ weight: 1
# Deployment
-Guides for deploying ACOR in various environments.
+ACOR is a library: you deploy your application, and it connects to Redis. Connection
+fields per topology are in
+[Redis topologies](../../getting-started/quick-start/#redis-topologies) — standalone for
+development and small workloads, Sentinel for failover, Cluster for horizontal scaling,
+Ring for client-side sharding.
-## Architecture Overview
-
-```mermaid
-graph TB
- subgraph Application
- A[ACOR Client]
- end
-
- subgraph Redis
- B[(Standalone)]
- C[(Sentinel)]
- D[(Cluster)]
- end
-
- A --> B
- A --> C
- A --> D
-```
-
-## Standalone Deployment
-
-Simplest deployment for development or small workloads:
+Read credentials from the environment rather than the source:
```go
ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "redis:6379",
+ Addr: os.Getenv("REDIS_ADDR"),
Password: os.Getenv("REDIS_PASSWORD"),
- DB: 0,
Name: "production",
})
-if err != nil {
- panic(err)
-}
-```
-
-## High Availability with Sentinel
-
-For production workloads requiring failover:
-
-```go
-ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addrs: []string{
- "sentinel-1:26379",
- "sentinel-2:26379",
- "sentinel-3:26379",
- },
- MasterName: "mymaster",
- Password: os.Getenv("REDIS_PASSWORD"),
- Name: "production",
-})
if err != nil {
panic(err)
}
```
-## Cluster Deployment
-
-For horizontal scaling:
-
-```go
-ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addrs: []string{
- "redis-node-1:7000",
- "redis-node-2:7000",
- "redis-node-3:7000",
- },
- Password: os.Getenv("REDIS_PASSWORD"),
- Name: "production",
-})
-if err != nil {
- panic(err)
-}
-```
-
-## Kubernetes Deployment
-
-### ConfigMap
+## Kubernetes
```yaml
apiVersion: v1
@@ -93,11 +34,7 @@ metadata:
data:
REDIS_ADDR: "redis-service:6379"
ACOR_COLLECTION: "production"
-```
-
-### Deployment
-
-```yaml
+---
apiVersion: apps/v1
kind: Deployment
metadata:
@@ -123,7 +60,6 @@ spec:
## Docker Compose
```yaml
-version: "3.8"
services:
redis:
image: redis:7-alpine
@@ -138,65 +74,29 @@ services:
- REDIS_ADDR=redis:6379
```
-## Health Checks
-
-> `server/health` belongs to the experimental `acor/server` module, which is
-> versioned separately and not covered by the core compatibility promise. See
-> [Server](../../server/) for what that means for your `go.mod`, and
-> [Running a Server](../../server/running/) for this wiring alongside the HTTP
-> and gRPC APIs.
-
-The `server/health` package registers Kubernetes-compatible endpoints on an
-`http.ServeMux`:
-
-- `/healthz` — liveness; always returns `200 OK` while the process is up.
-- `/readyz` — readiness; runs every registered `Checker` and returns `503` if
- any checker reports as unhealthy.
+## Health endpoints
-
-```go
-package main
-
-import (
- "log"
- "net/http"
+> `server/health` belongs to the experimental `acor/server` module, versioned separately
+> and not covered by the core compatibility promise. See [Server](../../server/).
- "github.com/skyoo2003/acor/pkg/acor"
- "github.com/skyoo2003/acor/server/health"
-)
+`health.RegisterHTTPHandlers(mux, checker)` registers two Kubernetes-compatible routes:
-// redisChecker is a readiness check that implements health.Checker.
-type redisChecker struct{ ac *acor.AhoCorasick }
+| Route | Meaning |
+| ----- | ------- |
+| `/healthz` | Liveness — `200 OK` while the process is up. Put nothing about Redis here |
+| `/readyz` | Readiness — runs every registered `Checker`, `503` if any is unhealthy |
-func (c redisChecker) Check() health.CheckResult {
- if _, err := c.ac.Info(); err != nil {
- return health.CheckResult{Status: health.StatusUnhealthy, Details: err.Error()}
- }
- return health.CheckResult{Status: health.StatusHealthy}
-}
+Keep Redis reachability out of liveness: a Redis outage would fail every replica's
+liveness probe at once, restart all of them, and repair nothing. It is a *take me out of
+the load balancer* signal.
-func main() {
- ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "redis:6379",
- Name: "production",
- })
- if err != nil {
- log.Fatal(err)
- }
-
- checker := health.NewChecker()
- checker.Register("redis", redisChecker{ac})
-
- mux := http.NewServeMux()
- health.RegisterHTTPHandlers(mux, checker) // registers /healthz and /readyz
- log.Fatal(http.ListenAndServe(":8080", mux))
-}
-```
+The complete `main` — checker with its own deadline, mux composition, graceful shutdown —
+is [Running a Server](../../server/running/), which also explains why a readiness check
+built on `Info()` costs what the dictionary costs.
-## Best Practices
+## Checklist
-1. Use connection pooling (built-in)
-2. Set appropriate timeouts
-3. Monitor Redis memory usage
-4. Use V2 schema for new collections
-5. Implement graceful shutdown
+1. Use the V2 schema for new collections (the default).
+2. Set timeouts and pool size from measurements, not from guesses.
+3. Monitor Redis memory and the [cache counters](../monitoring/).
+4. Call `Close()` on shutdown.
diff --git a/docs/content/operations/monitoring.md b/docs/content/operations/monitoring.md
index 453fed92..57754008 100644
--- a/docs/content/operations/monitoring.md
+++ b/docs/content/operations/monitoring.md
@@ -5,36 +5,21 @@ weight: 2
# Monitoring
-Monitor ACOR performance with built-in observability support.
-
-> **The `server/*` packages on this page are experimental.** They live in the
-> separate `github.com/skyoo2003/acor/server` module, which publishes no version
-> tags of its own and is **not covered by the core module's compatibility
-> promise** — its API can change in any release. The core library (`pkg/acor`) is
-> unaffected.
->
-> `go get github.com/skyoo2003/acor/server` resolves a pseudo-version from `main`.
-> Pin the core module explicitly in the same `go.mod`: Go ignores a dependency's
-> own `replace` directive, so without a pin you get whichever core version the
-> server module's `require` names.
-
-## Overview
-
-Two layers, and which one you get depends on how you run ACOR:
-
-| Layer | What it gives you | Covered by the `v1` promise |
-| ----- | ----------------- | --------------------------- |
-| `pkg/acor` — the library | `CacheStats()`: cache hit rate, rebuild cost, invalidation lag | ✅ |
+Two layers, and which you get depends on how you run ACOR:
+
+| Layer | What it gives you | Covered by `v1` |
+| ----- | ----------------- | --------------- |
+| `pkg/acor` — the library | `CacheStats()`: hit rate, rebuild cost, invalidation lag | ✅ |
| `acor/server` — the service | Prometheus metrics, structured JSON logs, OpenTelemetry traces | ❌ experimental |
-Embedding the library gets you the first row only. Everything after the next section is
-`server/*`, which means importing a separate, experimental module — reach for it when
-you run ACOR as a service, not to instrument your own process.
+Embedding the library gets you the first row. Everything from
+[Service layer](#service-layer) on means importing the separate, experimental
+`acor/server` module — see [Server](../../server/) for what that implies for your
+`go.mod`.
-## Core library: cache statistics
+## Cache statistics
-`CacheStats()` reports how often reads avoid Redis, rebuild cost, invalidation lag,
-and Preset refresh failures. It does no Redis I/O, so scraping it on a timer is cheap.
+`CacheStats()` does no Redis I/O, so scraping it on a timer is cheap.
```go
@@ -56,95 +41,75 @@ lag := stats.LastInvalidationLag
_, _, _ = hitRate, meanRebuild, lag
```
-Wire those three into whatever you already run — a Prometheus collector, an OTel
-meter, a log line. ACOR deliberately depends on no metrics library, so the choice
-stays yours.
+Wire those into whatever you already run — a Prometheus collector, an OTel meter, a log
+line. ACOR depends on no metrics library on purpose.
-### Preset refresh failures
+### Reading the numbers
-Track increases in `PresetReloadFailures` and `PresetPollFailures` alongside search
-errors. A shared failed reload increments the first counter once, even when many
-requests receive the error. A failed version poll increments the second counter;
-its next tick retries. Request cancellation is excluded from both counters.
-These counters remain zero outside Preset mode, and polling is disabled by default.
+- **Per instance, per process.** Nothing is aggregated through Redis, so scrape every
+ instance. A restart resets the counters.
+- **`Rebuilds` will not equal `Misses`.** Concurrent misses coalesce onto one build, so
+ `Misses - Rebuilds` is what coalescing saved; local writes rebuild off the read path and
+ push it the other way. In `Preset` mode `Rebuilds` starts at 1, from the build during
+ `Create`, and builds discarded after a generation conflict also count. Both are `uint64`
+ — check `Misses > Rebuilds` before subtracting, or a write-heavy instance wraps to
+ roughly 1.8e19.
+- **One scanning call is one read, whatever it scans.** `FindParallel`,
+ `FindIndexParallel`, and `FindMany` load the automaton once per call, so each adds 1 to
+ `Hits`+`Misses` and their hit rate is comparable to a serial workload's. Writes,
+ `Suggest`, and `Info` never reach the automaton and add nothing.
+- **`LastInvalidationLag` needs a listener and carries clock skew.** It is populated only
+ in `Preset` mode and in V2 with `EnableCache`; elsewhere zero means unavailable, not
+ fast. Where it is populated, the publish timestamp comes from another machine's clock,
+ so the value is the real delay plus that offset — it can understate as readily as
+ overstate. Watch it for step changes, and check NTP before blaming Pub/Sub.
+- **A zero hit rate is not always a bug.** Without `Preset` or `EnableCache` every read
+ still checks Redis for freshness; a hit means the automaton was reused, not that the
+ round trip was skipped.
-A failed reload retains the previous engine but returns an error to the search;
-it does not serve that engine as a fallback. Each waiting request responds to its
-own context cancellation without canceling other waiters. All waiters leaving or
-`Close` cancels the shared job. Redis reads and engine builds run outside the
-state lock, and snapshots overtaken by local state changes are rejected.
+Not here: match counts, keyword counts, Redis latency. Keyword and node counts come from
+`Info()`, which does read Redis.
-Polling reads only the version field. It detects missed invalidations after a
-successful poll; the next search must then fetch and build the full dictionary.
-The configured interval is not an upper bound on staleness during failures.
+### Preset refresh failures
-### Reading the numbers
+Track increases in `PresetReloadFailures` and `PresetPollFailures` alongside search errors.
+A shared failed reload increments the first counter once even when many requests receive
+the error; a failed version poll increments the second and retries on its next tick.
+Cancellation is excluded from both, and both stay zero outside `Preset` mode.
-- **The counters are per instance and per process.** Nothing is aggregated through
- Redis, so in a fleet you scrape every instance. A restart resets them.
-- **`Rebuilds` will not equal `Misses`.** Concurrent misses coalesce onto one build, so
- `Misses - Rebuilds` is what that coalescing saved; local writes rebuild off the read
- path and push the count the other way. In `Preset` mode `Rebuilds` starts at 1, from
- the build during `Create`; builds discarded after a generation conflict also count. Both counters are `uint64`, so check `Misses > Rebuilds`
- before subtracting — a write-heavy instance is routinely the other way round, and the
- difference wraps to roughly 1.8e19 rather than going negative.
-- **One scanning call is one read, whatever it scans over.** `FindParallel`,
- `FindIndexParallel`, and `FindMany` load the automaton once per call and scan every
- chunk or text against that snapshot, so each adds 1 to `Hits`+`Misses` and their hit
- rate is directly comparable to a serial workload's. Calls that never reach the
- automaton — writes, `Suggest`, `Info` — add nothing to either counter.
-- **`LastInvalidationLag` needs a listener, and carries clock skew.** It is populated
- only in `Preset` mode and in V2 with `EnableCache`; the other modes subscribe to
- nothing, so a zero there means unavailable, not fast. Where it is populated, the
- publish timestamp comes from another machine's clock, so the value is the real delay
- plus that clock's offset — it can understate the delay as readily as overstate it, and
- bounds it in neither direction. Watch it for step changes rather than trusting the
- absolute value, and check NTP before concluding Pub/Sub is slow.
-- **A zero hit rate is not always a bug.** Without `Preset` or `EnableCache` every read
- still checks Redis for freshness; a hit there means only that the automaton was
- reused, not that the round trip was skipped.
-- **What is not here**: match counts, keyword counts, and Redis latency. Keyword and
- node counts come from `Info()`, which does read Redis.
-
-## Service layer: `server/*`
-
-```mermaid
-graph LR
- A[ACOR] --> B[Metrics]
- A --> C[Logs]
- A --> D[Traces]
-
- B --> E[Prometheus]
- C --> F[Log Aggregator]
- D --> G[Jaeger/Zipkin]
-```
+A failed reload keeps the previous engine but **returns an error to the search** rather
+than serving that engine as a fallback. Each waiting request answers its own cancellation
+without cancelling other waiters; all waiters leaving, or `Close`, cancels the shared job.
+Redis reads and engine builds run outside the state lock, and snapshots overtaken by local
+state changes are rejected.
-### Metrics
+Polling reads only the version field, and detects a missed invalidation only after a
+successful poll — the next search then fetches and builds the full dictionary. The
+configured interval is not an upper bound on staleness during failures.
-Import the metrics package:
+## Service layer
+
+### Metrics
```go
import "github.com/skyoo2003/acor/server/metrics"
```
-#### Available Metrics
-
-| Metric | Type | Description |
-| --------------------------------------- | --------- | ------------------------------------------- |
-| `acor_http_requests_total` | Counter | Total HTTP requests by method, path, status |
-| `acor_http_request_duration_seconds` | Histogram | HTTP request latency |
-| `acor_redis_operations_total` | Counter | Total Redis operations by type, status |
-| `acor_redis_operation_duration_seconds` | Histogram | Redis operation latency |
-| `acor_keywords_total` | Gauge | Number of registered keywords |
-| `acor_trie_nodes_total` | Gauge | Number of trie nodes |
-| `grpc_server_handled_total` | Counter | Total gRPC requests by method, code |
-| `grpc_server_handling_seconds` | Histogram | gRPC request latency |
+| Metric | Type | Description |
+| ------ | ---- | ----------- |
+| `acor_http_requests_total` | Counter | HTTP requests by method, path, status |
+| `acor_http_request_duration_seconds` | Histogram | HTTP request latency |
+| `acor_redis_operations_total` | Counter | Redis operations by type, status |
+| `acor_redis_operation_duration_seconds` | Histogram | Redis operation latency |
+| `acor_keywords_total` | Gauge | Registered keywords |
+| `acor_trie_nodes_total` | Gauge | Trie nodes |
+| `grpc_server_handled_total` | Counter | gRPC requests by method, code |
+| `grpc_server_handling_seconds` | Histogram | gRPC request latency |
gRPC metrics use the standard `grpc_server_*` names from
-`go-grpc-middleware/providers/prometheus`, wired via
-`NewGRPCServerWithObservability`.
+`go-grpc-middleware/providers/prometheus`, wired by `NewGRPCServerWithObservability`.
-#### Exposing Metrics
+Registering them does not expose them — serve `promhttp.Handler()` yourself:
```go
import (
@@ -166,17 +131,9 @@ func main() {
### Logging
-Import the logging package:
-
-```go
-import "github.com/skyoo2003/acor/server/logging"
-```
-
-#### Structured Logging
-
-`NewLogger` takes an `io.Writer` and a level string (`debug`, `info`, `warn`,
-`error`) and returns a zerolog-based logger that always emits structured JSON.
-Attach trace/span IDs with `WithTraceID`:
+`logging.NewLogger(w, level)` takes an `io.Writer` and one of `debug`, `info`, `warn`,
+`error`, and returns a zerolog logger that always emits structured JSON. `WithTraceID`
+attaches trace and span IDs.
```go
@@ -205,23 +162,8 @@ func main() {
}
```
-#### Log Levels
-
-- `debug`: Detailed debugging info
-- `info`: General operational info
-- `warn`: Warning conditions
-- `error`: Error conditions
-
### Tracing
-Import the tracing package:
-
-```go
-import "github.com/skyoo2003/acor/server/tracing"
-```
-
-#### OpenTelemetry Setup
-
```go
tracer, err := tracing.NewTracer(&tracing.Config{
@@ -236,36 +178,15 @@ if err != nil {
defer tracer.Shutdown()
```
-#### Spans
-
-The core `pkg/acor` library does not emit its own spans. Request spans are
-created by the `server/tracing` middleware for incoming traffic:
-
-- HTTP requests — via `tracing.HTTPMiddleware`
-- gRPC calls — via the standard `otelgrpc` stats handler (`tracing.GRPCStatsHandler`)
-
-To trace individual `Add`/`Find`/`Remove` calls, wrap them in your own spans
-using the OpenTelemetry API.
-
-### Dashboards
-
-#### Key Metrics to Monitor
-
-1. **Operation Latency**: P50, P95, P99
-2. **Error Rate**: Operations failing
-3. **Keyword Count**: Collection size
-4. **Redis Connections**: Pool utilization
-
-#### Grafana Dashboard
-
-Create a Grafana dashboard using the metrics above. Key panels to include:
+`pkg/acor` emits no spans of its own. Request spans come from the middleware —
+`tracing.HTTPMiddleware` for HTTP, the standard `otelgrpc` stats handler
+(`tracing.GRPCStatsHandler`) for gRPC. To trace individual `Add`/`Find`/`Remove` calls,
+wrap them in your own spans.
-1. **Operation Latency**: P50/P95/P99 of `acor_redis_operation_duration_seconds`
-2. **Error Rate**: Rate of `acor_redis_operations_total{status="error"}`
-3. **Keyword Count**: Gauge `acor_keywords_total`
-4. **Trie Nodes**: Gauge `acor_trie_nodes_total`
+### Alerting
-### Alerting Rules
+Latency percentiles, error rate, keyword count, and pool utilization are the four worth a
+dashboard panel. As rules:
```yaml
groups:
diff --git a/docs/content/operations/troubleshooting.md b/docs/content/operations/troubleshooting.md
index faf634a1..74ab25a0 100644
--- a/docs/content/operations/troubleshooting.md
+++ b/docs/content/operations/troubleshooting.md
@@ -5,127 +5,28 @@ weight: 3
# Troubleshooting
-Common issues and their solutions.
+## Errors from `Create` and the matching API
-## Common Errors
+| Error | Cause | Fix |
+| ----- | ----- | --- |
+| `ErrRedisConflictingTopology` | More than one topology configured at once | Use exactly one of `Addr`, `Addrs` (+`MasterName` for Sentinel), or `RingAddrs` — see [Redis topologies](../../getting-started/quick-start/#redis-topologies) |
+| `ErrRedisClusterDB` | Non-zero `DB` together with `Addrs` | Cluster has no database selection; drop `DB`, or use `Addr` for a single standalone server |
+| `ErrEmptyKeyword` | Empty string passed to `Add` | Trim and reject before calling |
+| `ErrInvalidChunkSize` | Non-positive `ParallelOptions.ChunkSize` | `ChunkSize` is required and must be > 0 |
+| `ErrRedisAlreadyClosed` | Operation on a closed instance | `defer ac.Close()` once, at the owning function's exit |
+| `ErrV1ReadOnly` | `Add`/`Remove` on a V1 collection | Migrate: `acor -name mycollection migrate` |
+| `ErrSuggestRequiresRedis` | `Suggest` in `Preset` mode | Suggest needs the Redis path; use a non-preset instance for it |
-### ErrRedisConflictingTopology
+## Redis connection
-**Cause:** Multiple Redis topologies specified simultaneously.
+| Message | Check |
+| ------- | ----- |
+| `connection refused` | `redis-cli ping`, the address, the firewall, network reachability |
+| `NOAUTH Authentication required` | Set `Password` |
+| `context deadline exceeded` | Redis load, network latency, and the timeouts below |
-**Solution:** Use only one configuration:
-
-```go
-// Correct: Standalone
-args := &acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-collection",
-}
-
-// Correct: Sentinel
-args := &acor.AhoCorasickArgs{
- Addrs: []string{"localhost:26379"},
- MasterName: "mymaster",
- Name: "my-collection",
-}
-
-// Wrong: Mixing configurations
-args := &acor.AhoCorasickArgs{
- Addr: "localhost:6379", // Wrong!
- Addrs: []string{"..."}, // Wrong!
- MasterName: "mymaster", // Wrong!
- Name: "my-collection",
-}
-```
-
-### ErrEmptyKeyword
-
-**Cause:** Empty string passed to `Add()`.
-
-**Solution:** Validate input:
-
-```go
-keyword := strings.TrimSpace(input)
-if keyword == "" {
- return errors.New("keyword cannot be empty")
-}
-_, err = ac.Add(keyword)
-```
-
-### ErrInvalidChunkSize
-
-**Cause:** Non-positive chunk size in parallel matching.
-
-**Solution:** Use positive values:
-
-```go
-opts := &acor.ParallelOptions{
- Workers: 4,
- ChunkSize: 1000, // Must be > 0
-}
-```
-
-### ErrRedisAlreadyClosed
-
-**Cause:** Operation on closed AhoCorasick instance.
-
-**Solution:** Ensure `Close()` is called only once, typically with `defer`:
-
-```go
-ac, err := acor.Create(args)
-if err != nil {
- log.Fatal(err)
-}
-defer ac.Close() // Called once at function exit
-```
-
-## Redis Connection Issues
-
-### Connection Refused
-
-```text
-redis GET on key "...": connection refused
-```
-
-**Checklist:**
-1. Redis is running: `redis-cli ping`
-2. Address is correct
-3. Firewall allows connection
-4. Network connectivity
-
-### Authentication Failed
-
-```text
-redis GET on key "...": NOAUTH Authentication required
-```
-
-**Solution:** Provide password:
-
-```go
-args := &acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Password: "your-password",
- Name: "my-collection",
-}
-```
-
-### Timeout Errors
-
-```text
-redis GET on key "...": context deadline exceeded
-```
-
-**Solutions:**
-1. Tune `DialTimeout`, `ReadTimeout`, or `WriteTimeout` when measurements show
- the defaults are too short
-2. Check Redis load
-3. Check network latency
-4. Set `MaxRetries` for transient failures or adjust `PoolSize` for measured
- connection contention
-5. Scale Redis cluster
-
-Zero values keep the go-redis defaults across Standalone, Sentinel, Cluster,
-and Ring modes.
+Zero values keep the go-redis defaults across every topology. Tune from measurements, not
+from guesses:
```go
@@ -140,64 +41,44 @@ args := &acor.AhoCorasickArgs{
_ = args
```
-### Preset Cache Appears Stale
+`PoolSize` is the knob for measured connection contention.
-**Cause:** Preset mode normally reloads through best-effort Redis Pub/Sub. A
-disconnected subscriber can miss an invalidation.
+## Preset cache looks stale
-**Solution:** In multi-instance deployments, set
-`InvalidationPollInterval` to retry version checks periodically:
+Preset mode reloads through best-effort Pub/Sub, and a disconnected subscriber misses an
+invalidation. In multi-instance deployments, enable polling:
```go
args.InvalidationPollInterval = 30 * time.Second
```
-The option is disabled by default and ignored outside Preset mode. The interval
-is not a freshness bound: recovery requires a successful version poll and a
-successful reload on the next search. Inspect `CacheStats().PresetPollFailures`
-and `PresetReloadFailures` when updates remain invisible. Reload errors are
-returned to searches; the retained engine is not automatically served as fallback.
-
-## Performance Issues
-
-### Slow Find Operations
+Disabled by default, and ignored outside `Preset` mode. **The interval is not a freshness
+bound** — recovery needs a successful version poll *and* a successful reload on the next
+search. When updates stay invisible, check `CacheStats().PresetPollFailures` and
+`PresetReloadFailures`; reload errors are returned to searches rather than silently served
+from the retained engine. See
+[invalidation safety](../../guides/redis-backed-engine/#invalidation-safety).
-**Diagnostic:**
-1. Check schema version: `acor -name collection schema-version`
-2. Check collection size: `acor -name collection info`
+## Slow reads or high memory
-**Solutions:**
-- Migrate to V2 schema
-- Use parallel matching for large texts
-- Increase Redis memory
-
-### High Memory Usage
-
-**Diagnostic:**
-1. Check Redis memory: `redis-cli info memory`
-2. Check keyword count: `acor -name collection info`
+```bash
+acor -name mycollection schema-version # V1 is the usual answer to "why is Find slow"
+acor -name mycollection info # keyword and node counts
+redis-cli info memory
+```
-**Solutions:**
-- Remove unused keywords
-- Use V2 schema (lower memory)
-- Scale Redis cluster
+- On V1, migrate to V2 — see [Schema V2](../../reference/schema-v2/).
+- On V2, a read-heavy workload wants `EnableCache` or a `Preset`; the schema alone does
+ not make reads fast ([benchmarks](../../reference/benchmarks/#what-the-numbers-mean)).
+- For large texts, use [parallel matching](../../guides/parallel-matching/).
+- For high memory, remove unused keywords or move to `PresetMemoryEfficient`.
## Debugging
-### Enable Debug Logging
-
-```go
-logger := logging.NewLogger(os.Stdout, "debug")
-```
-
-### CLI Debug Mode
-
```bash
-acor -name mycollection -debug find "test text"
+acor -name mycollection -debug find "test text" # CLI debug logging
+redis-cli keys "{mycollection}:*" # what the collection actually stores
```
-### Check Redis Keys
-
-```bash
-redis-cli keys "{mycollection}:*"
-```
+In library code, set `Debug: true` for the default stdout logger, or supply your own
+`Logger`.
diff --git a/docs/content/reference/_index.md b/docs/content/reference/_index.md
index eb09bea7..e9ef0a39 100644
--- a/docs/content/reference/_index.md
+++ b/docs/content/reference/_index.md
@@ -5,20 +5,15 @@ weight: 3
# Reference
-Technical reference documentation for ACOR.
+- [API Reference](api/) — constructors, matching, batch, parallel, and the option types
+- [Compatibility](compatibility/) — what the `v1` line promises, and what it excludes
+- [Schema V2](schema-v2/) — the current storage layout
+- [Schema V1](schema-v1/) — deprecated, read-only
+- [Versioned dictionaries (V3)](versioned/) — leased snapshots, expected-version writes, cutover
+- [Bounded text processing](text-processing/) — `Scan`, `MaskText`, `ReplaceText` over source positions
-## Sections
+## Measurements
-- [API Reference](api/) - Public API documentation (unified `Create` API with `Preset` options)
-- [Compatibility](compatibility/) - What the `v1` line promises, and what it excludes
-- [Schema V1](schema-v1/) - Legacy schema details
-- [Schema V2](schema-v2/) - Optimized schema (recommended)
-- [Versioned dictionaries (V3)](versioned/) - Leased snapshots, expected-version writes, and cutover
-- [Bounded search, masking and replacement](text-processing/) - `Scan`, `MaskText`, and `ReplaceText` over original source positions
-- [Benchmarks](benchmarks/) - Measured performance and how to reproduce it
-- [V3 performance report](versioned-performance/) - Million-keyword measurements on Redis and Valkey
-- [R2/R3 verification](r2-r3-performance/) - Boundary protection and bounded-processing evidence
-
-## Navigation
-
-← [Guides](../guides/) | [Operations](../operations/) →
+- [Benchmarks](benchmarks/) — round trips and timings, with the commands that produce them
+- [V3 performance report](versioned-performance/) — the archived R1 million-keyword baseline
+- [R2/R3 verification](r2-r3-performance/) — incremental download, engine memory, bounded APIs
diff --git a/docs/content/reference/api.md b/docs/content/reference/api.md
index 46e3a6c8..cf8c0dee 100644
--- a/docs/content/reference/api.md
+++ b/docs/content/reference/api.md
@@ -5,15 +5,37 @@ weight: 1
# API Reference
-Core API documentation for ACOR. See
-[pkg.go.dev](https://pkg.go.dev/github.com/skyoo2003/acor/pkg/acor) for the
-complete generated reference.
+Contracts and behavior for `pkg/acor`. Generated signatures live on
+[pkg.go.dev](https://pkg.go.dev/github.com/skyoo2003/acor/pkg/acor).
-## Core Types
+## Creating a collection
-### AhoCorasickArgs
+```go
+ac, err := acor.Create(&acor.AhoCorasickArgs{...})
+defer ac.Close()
+```
-Configuration for creating an AhoCorasick instance.
+`CreateContext` is the same constructor with a context bounding the setup I/O — the
+schema check, the initialization write, and the initial keyword load:
+
+
+```go
+setupCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
+defer cancel()
+instance, err := acor.CreateContext(setupCtx, &acor.AhoCorasickArgs{
+ Addr: "localhost:6379",
+ Name: "default",
+})
+_ = instance
+_ = err
+```
+
+That context bounds construction only. Cancelling it later neither closes the instance
+nor stops its invalidation listener — use `Close` for that and the `*Context` methods for
+per-operation cancellation. The Pub/Sub subscribe belongs to the listener, so it runs on
+the instance's own context.
+
+### AhoCorasickArgs
```go
@@ -43,73 +65,25 @@ type AhoCorasickArgs struct {
```
-### AhoCorasick
-
-Main type for pattern matching operations.
-
-```go
-ac, err := acor.Create(&acor.AhoCorasickArgs{...})
-defer ac.Close()
-```
-
-`CreateContext` is the same constructor with a context bounding the setup I/O
-(schema check and initialization write, initial keyword load):
-
-
-```go
-setupCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second)
-defer cancel()
-instance, err := acor.CreateContext(setupCtx, &acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "default",
-})
-_ = instance
-_ = err
-```
-
-The context bounds construction only. Canceling it afterwards does not close the
-instance or stop its invalidation listener — use `Close` for that, and the
-`*Context` methods for per-operation cancellation. The Pub/Sub subscribe is part
-of that listener, so it runs on the instance's own context rather than this one.
-
-## Core Methods
-
-### Add
-
-Add a single keyword to the collection.
-
-```go
-count, err := ac.Add("keyword")
-```
-
-### AddMany
-
-Add multiple keywords in a batch.
-
-```go
-result, err := ac.AddMany([]string{"a", "b", "c"}, nil)
-// or with options:
-result, err := ac.AddMany([]string{"a", "b", "c"}, &acor.BatchOptions{
- Mode: acor.BatchModeTransactional,
-})
-```
-
-### Remove
+Setting `Preset` switches reads to a local automaton — see
+[Preset-Optimized Engine](../../guides/preset-engine/). Topology fields are covered in
+[Redis topologies](../../getting-started/quick-start/#redis-topologies).
-Remove a single keyword from the collection.
-
-```go
-count, err := ac.Remove("keyword")
-```
+## Writing
-### RemoveMany
-
-Remove multiple keywords in a batch.
+| Method | Returns | Notes |
+| ------ | ------- | ----- |
+| `Add(keyword)` | `(int, error)` | 1 if the collection changed, 0 if it was already there |
+| `Remove(keyword)` | `(int, error)` | Same convention |
+| `AddMany(keywords, *BatchOptions)` | `(*BatchResult, error)` | One transaction; `nil` options mean best-effort |
+| `RemoveMany(keywords, *BatchOptions)` | `(*BatchResult, error)` | Same |
+| `Flush()` | `error` | Deletes every key in the collection |
+| `Close()` | `error` | Closes the connection and stops the invalidation listener |
```go
result, err := ac.RemoveMany([]string{"a", "b"}, nil)
-// Pass options when transactional behavior is required.
+// Pass options when the whole batch must land or fail together.
result, err = ac.RemoveMany([]string{"c", "d"}, &acor.BatchOptions{
Mode: acor.BatchModeTransactional,
})
@@ -117,29 +91,29 @@ _ = result
_ = err
```
-### Find
-
-Find all matching keywords in text.
+See [Batch Operations](../../guides/batch-operations/) for the modes and result shape.
-```go
-matches, err := ac.Find("sample text")
-// Returns: []string{"match1", "match2", ...}
-```
+## Matching
-### FindIndex
+| Method | Returns | Shape |
+| ------ | ------- | ----- |
+| `Find(text)` | `[]string` | Which keywords occur |
+| `FindIndex(text)` | `map[string][]int` | Keyword to its rune start positions |
+| `FindSet(text)` | `[]string` | Each keyword once, in first-match order |
+| `FindMatches(text, *MatchOptions)` | `[]Match` | Every occurrence with its rune span |
+| `Contains(text)` | `bool` | Stops at the first match |
+| `FindStream(reader, callback)` | `error` | Scans an `io.Reader` without buffering it |
+| `FindMany(texts)` | `map[string][]string` | Keyed by input text |
+| `FindParallel(text, *ParallelOptions)` | `[]string` | Chunked across workers |
+| `FindIndexParallel(text, *ParallelOptions)` | `map[string][]int` | Same chunking, with positions |
-Find matches with their start positions.
-
-```go
-positions, err := ac.FindIndex("sample text")
-// Returns: map[string][]int{"keyword": {startPos, ...}, ...}
-```
+All positions are **rune** offsets, not byte offsets.
### FindMatches
-Return every occurrence in scan order with its keyword and half-open rune span
-`[Start, End)`. The default includes overlapping matches; use
-`MatchKindLeftmostLongest` for non-overlapping tokenization or replacement.
+Every occurrence in scan order with its half-open rune span `[Start, End)`. The default
+includes overlaps; `MatchKindLeftmostLongest` produces non-overlapping spans suitable for
+tokenizing or replacing.
```go
@@ -170,22 +144,14 @@ const (
)
```
-`WholeWord` uses letters, digits, combining marks, and underscores as word
-runes. Set `WordRune` when those defaults do not fit the input script.
-
-### Contains
-
-Report whether any keyword occurs, stopping at the first match.
-
-```go
-found, err := ac.Contains("sample text")
-```
+`WholeWord` treats letters, digits, combining marks, and underscores as word runes. Set
+`WordRune` for scripts those defaults do not fit — in CJK or Thai text every adjacent
+character counts as a word rune, so nearly every match is dropped as mid-word.
### FindStream
-Scan an `io.Reader` without buffering the whole input. Matches include
-overlaps, retain rune offsets across reads, and arrive in scan order. Returning
-`false` from the callback stops the scan.
+Matches keep rune offsets across reads and arrive in scan order; returning `false` from
+the callback stops the scan.
```go
@@ -196,83 +162,62 @@ err := ac.FindStream(strings.NewReader("sample text"), func(match acor.Match) bo
_ = err
```
-Streaming does not apply whole-word or leftmost-longest filtering because
-those modes require buffering. Use `FindMatches` for bounded strings that need
-those options.
-
-### FindMany
-
-Find matches in multiple texts.
-
-```go
-matches, err := ac.FindMany([]string{"text1", "text2"})
-// Returns: map[string][]string{"text1": {"kw", ...}, ...} (keyed by input text)
-```
-
-### FindParallel
+Streaming always overlaps: whole-word and leftmost-longest need buffering, so use
+`FindMatches` on a bounded string when you need them.
-Find matches using parallel processing.
+### Parallel
```go
-matches, err := ac.FindParallel(largeText, &acor.ParallelOptions{
- Workers: 4,
- Boundary: acor.ChunkBoundaryWord,
-})
-```
-
-Keywords longer than `ParallelOptions.Overlap` can be missed at chunk
-boundaries. Set `Overlap` to at least the longest expected keyword.
-
-### FindIndexParallel
-
-Find start positions using the same parallel chunking options.
-
-```go
-positions, err := ac.FindIndexParallel(largeText, acor.DefaultParallelOptions())
-```
-
-### Info
-
-Get collection statistics.
+type ParallelOptions struct {
+ Workers int // Concurrent goroutines (default: runtime.NumCPU())
+ ChunkSize int // Target chunk size in runes (required; no fallback)
+ Boundary ChunkBoundary // How chunks are split (default: ChunkBoundaryWord)
+ Overlap int // Overlap runes between chunks (unset means zero)
+ AutoOverlap bool // Extend each chunk by the dictionary's longest keyword (default: false)
+}
-```go
-info, err := ac.Info()
-// Returns: &AhoCorasickInfo{Keywords: N, Nodes: M, Preset: ..., MemoryBytes: ..., TrieDepth: ...}
+const (
+ ChunkBoundaryWord ChunkBoundary = iota // Split at whitespace (default)
+ ChunkBoundarySentence // Split at . ! ?
+ ChunkBoundaryLine // Split at newlines
+)
```
-### CacheStats
+`DefaultParallelOptions()` returns a usable starting point. Without `AutoOverlap`, a
+keyword longer than `Overlap` can be missed at a chunk boundary; with it, chunks extend
+by the dictionary's longest keyword at no extra Redis cost. Details:
+[Parallel Matching](../../guides/parallel-matching/).
-Get local cache statistics. Unlike `Info`, this performs no Redis I/O, so it is cheap
-enough to scrape on a timer.
+## Suggest
-```go
-stats := ac.CacheStats()
-// Returns: CacheStats{Hits: N, Misses: M, Rebuilds: R, RebuildDuration: ..., LastInvalidationLag: ...}
-```
-
-The counters are per instance and per process — scrape every instance in a fleet. See
-[Monitoring](../../operations/monitoring/) for how to read them, including why
-`Rebuilds` does not equal `Misses` and why `LastInvalidationLag` carries clock skew.
+`Suggest(prefix)` returns keywords starting with the prefix; `SuggestIndex(prefix)`
+returns the same keys mapped to `[0]`, since a prefix match always starts at the
+beginning. Both require Redis and are unavailable in `Preset` mode
+(`ErrSuggestRequiresRedis`).
-### Flush
-
-Clear all data from the collection.
+## Batch types
```go
-err := ac.Flush()
-```
-
-### Close
+type BatchOptions struct {
+ Mode BatchMode // BatchModeBestEffort (default) or BatchModeTransactional
+}
-Close the Redis connection.
+type BatchResult struct {
+ Added []string // Successfully added keywords
+ Removed []string // Successfully removed keywords
+ Failed []KeywordError // Keywords that failed, with their errors
+ Skipped []string // Duplicate adds or absent removes
+}
-```go
-err := ac.Close()
+type KeywordError struct {
+ Keyword string
+ Error error
+}
```
-### AhoCorasickInfo
+## Statistics
-Statistics about an Aho-Corasick instance.
+`Info()` reads Redis; `CacheStats()` does not, so it is cheap to scrape on a timer.
```go
@@ -286,11 +231,6 @@ type AhoCorasickInfo struct {
```
-### CacheStats (type)
-
-A snapshot of one instance's local cache activity. Returned by `CacheStats()`, never
-constructed by callers — fields may be added inside `v1`.
-
```go
type CacheStats struct {
PresetReloadFailures uint64 // Failed shared reload jobs, once per job (Preset only; cancellation excluded)
@@ -303,190 +243,40 @@ type CacheStats struct {
}
```
-### Preset
+`CacheStats` is returned by ACOR and never constructed by callers, so fields may be added
+inside `v1`. Counters are per instance and per process — scrape every instance in a fleet.
+How to read them, including why `Rebuilds` does not equal `Misses`:
+[Monitoring](../../operations/monitoring/).
-Architecture presets for the preset-optimized Redis engine.
+## Presets
```go
const (
- PresetNone Preset = iota // Zero value (unset) — falls through to original V1/V2 mode
+ PresetNone Preset = iota // Zero value — original V1/V2 mode
PresetSpeed // Full DFA + flat array — max speed, higher memory
PresetBalanced // Double-Array Trie + Banded DFA — best speed-to-memory ratio
PresetMemoryEfficient // Map-based + Bloom filter — min memory, slower search
)
```
-## Redis-Backed Engine with Presets
-
-Redis-backed Aho-Corasick that combines Redis persistence with a local preset-optimized automaton. Writes go to Redis atomically (V2 Lua scripts with optimistic locking); reads hit the local engine with no Redis I/O. Created via the unified `Create` API with `Preset` set.
-
-```go
-ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-collection",
- Preset: acor.PresetBalanced,
- CaseSensitive: false,
-})
-defer ac.Close()
-```
-
-### AhoCorasickArgs (Preset field)
-
-The `AhoCorasickArgs` struct includes a `Preset` field for the engine mode:
-
-```go
-type AhoCorasickArgs struct {
- // ... standard Redis connection fields ...
- Preset Preset // Architecture preset: PresetSpeed, PresetBalanced, PresetMemoryEfficient
- // ... other fields ...
-}
-```
-
-### Preset-Optimized Redis Methods
-
-```go
-// Create
-ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-collection",
- Preset: acor.PresetBalanced,
-})
-
-// Add/Remove
-added, err := ac.Add("keyword") // (int, error)
-removed, err := ac.Remove("keyword") // (int, error)
-
-// Find (0 RTT on hot path — reads from local engine)
-matches, err := ac.Find("text") // ([]string, error)
-positions, err := ac.FindIndex("text") // (map[string][]int, error)
-spans, err := ac.FindMatches("text", nil) // ([]Match, error)
-found, err := ac.Contains("text") // (bool, error)
-
-// Info
-info, err := ac.Info() // (*AhoCorasickInfo, error)
-
-// Flush
-err := ac.Flush()
-
-// Close
-err := ac.Close()
-```
-
-## Context Variants
+## Context variants
-Operations that may perform Redis I/O also accept an explicit
-`context.Context`: `AddContext`, `RemoveContext`, `FindContext`,
-`FindIndexContext`, `FindMatchesContext`, `ContainsContext`,
-`FindStreamContext`, `FlushContext`, `InfoContext`, `SuggestContext`,
-`SuggestIndexContext`, `AddManyContext`, `RemoveManyContext`,
-`FindManyContext`, `FindParallelContext`, and `FindIndexParallelContext`.
+Every method that may touch Redis has a `*Context` twin taking an explicit
+`context.Context`: `AddContext`, `RemoveContext`, `FindContext`, `FindIndexContext`,
+`FindSetContext`, `FindMatchesContext`, `ContainsContext`, `FindStreamContext`,
+`FlushContext`, `InfoContext`, `SuggestContext`, `SuggestIndexContext`,
+`AddManyContext`, `RemoveManyContext`, `FindManyContext`, `FindParallelContext`, and
+`FindIndexParallelContext`.
```go
matches, err := ac.FindMatchesContext(ctx, text, nil)
```
-## Suggest Methods
-
-### Suggest
-
-Get prefix suggestions.
-
-```go
-suggestions, err := ac.Suggest("pre")
-```
-
-### SuggestIndex
-
-Get suggestions with positions.
-
-```go
-positions, err := ac.SuggestIndex("pre")
-```
-
-## Batch Operations
-
-### BatchOptions
-
-```go
-type BatchOptions struct {
- Mode BatchMode // BestEffort (default) or Transactional
-}
-```
-
-### BatchResult
-
-```go
-type BatchResult struct {
- Added []string // Successfully added keywords
- Removed []string // Successfully removed keywords
- Failed []KeywordError // Keywords that failed with their errors
- Skipped []string // Duplicate adds or absent removes
-}
-```
-
-### KeywordError
-
-```go
-type KeywordError struct {
- Keyword string
- Error error
-}
-```
-
-## Parallel Options
-
-### ParallelOptions
-
-```go
-type ParallelOptions struct {
- Workers int // Concurrent goroutines (default: runtime.NumCPU())
- ChunkSize int // Target chunk size in characters (required; no fallback)
- Boundary ChunkBoundary // How chunks are split (default: ChunkBoundaryWord)
- Overlap int // Overlap characters between chunks (unset means zero)
- AutoOverlap bool // Extend each chunk by the dictionary's longest keyword (default: false)
-}
-```
-
-### DefaultParallelOptions
-
-Returns parallel options with sensible defaults:
-
-```go
-opts := acor.DefaultParallelOptions()
-matches, err := ac.FindParallel(text, opts)
-```
-
-### ChunkBoundary
-
-```go
-const (
- ChunkBoundaryWord ChunkBoundary = iota // Split at whitespace (default)
- ChunkBoundarySentence // Split at sentence boundaries (. ! ?)
- ChunkBoundaryLine // Split at newlines
-)
-```
-
-## Parallel Boundary Protection and Preset Refresh
-
-`ParallelOptions.AutoOverlap` (default `false`) enables dictionary-aware right
-extensions for each base chunk. It protects keywords longer than `ChunkSize`
-without additional Redis reads. Results use the existing parallel order and rune
-positions. See [parallel matching](../../guides/parallel-matching/).
-
-`CacheStats.PresetReloadFailures` and `CacheStats.PresetPollFailures` are cumulative
-`uint64` counters. Shared reload failures count once per job; cancellations are
-excluded. They remain zero in other modes. See [Preset refresh](../../guides/redis-backed-engine/#invalidation-safety).
-
-## Versioned dictionaries
-
-`OpenVersioned(ctx, *VersionedOptions)` opens a separate V3 collection. Its
-`VersionedCollection` API provides leased snapshots, paginated list/diff, expected-version
-replace/add/remove, asynchronous engine refresh, operation receipts, V2 copying
-and explicit pruning. See [the V3 API guide](../versioned/) for contracts and a
-compilable example; V3 uses the existing module version and search semantics.
-
-## Bounded source-position APIs
+## Beyond V1/V2
-`Scan`, `MaskText`, and `ReplaceText` are available on AhoCorasick and
-VersionedCollection with explicit contexts. See [bounded text processing](../text-processing/)
-for ScanOptions, RewriteOptions, original byte/rune spans and error contracts.
+- **`OpenVersioned(ctx, *VersionedOptions)`** opens a separate V3 collection with leased
+ snapshots, paginated list/diff, expected-version writes, background engine refresh, and
+ operation receipts. See [Versioned dictionaries](../versioned/).
+- **`Scan`, `MaskText`, `ReplaceText`** report original byte and rune spans under explicit
+ work limits, on both `AhoCorasick` and `VersionedCollection`. See
+ [bounded text processing](../text-processing/).
diff --git a/docs/content/reference/benchmarks.md b/docs/content/reference/benchmarks.md
index 152c0e70..a5f7ea21 100644
--- a/docs/content/reference/benchmarks.md
+++ b/docs/content/reference/benchmarks.md
@@ -5,21 +5,19 @@ weight: 5
# Benchmarks
-Every performance number ACOR publishes is on this page, with the command that
-produces it. Nothing here is an estimate.
+Every performance number ACOR publishes is here, with the command that produces it.
+Nothing is an estimate.
-The evidence comes in two kinds, and they are not interchangeable:
+The evidence comes in two kinds:
-- **Round trips** are structural. They are counted at the storage seam, so they
- are identical on miniredis and on a real server, and no hardware difference
- can move them. They are enforced by tests on every CI run.
-- **Timings** are hardware-bound. Absolute nanoseconds on your machine will not
- match ours. The reproducible quantity is the **ratio** between configurations.
+- **Round trips** are structural — counted at the storage seam, identical on miniredis and
+ on a real server, and enforced by tests on every CI run.
+- **Timings** are hardware-bound. Absolute nanoseconds will not match yours. The
+ reproducible quantity is the **ratio** between configurations.
## Round trips per operation
-Enforced by `TestRTT*` in `pkg/acor/rtt_claims_test.go`. If any of these change,
-CI fails.
+Enforced by `TestRTT*` in `pkg/acor/rtt_claims_test.go`; a change fails CI.
| Operation | V1 | V2 |
|---|---|---|
@@ -31,15 +29,14 @@ CI fails.
| `Add()`, 5-character keyword | 53 | 2 |
| `Add()`, 26-character keyword | 507 | 2 |
-Both schemas read in a single round trip: V1 issues one `SMEMBERS`, V2 pipelines
-two `HGETALL` calls into one trip. V1's round-trip cost is on **writes**, where it
-walks the trie node by node, so the cost grows with the length of the keyword being
-added rather than with the size of the dictionary.
+Both schemas read in one round trip: V1 issues one `SMEMBERS`, V2 pipelines two `HGETALL`
+calls into one trip. V1's cost is on **writes**, where it walks the trie node by node — so
+it grows with the length of the keyword being added, not with the dictionary.
-The multi-scan reads cost the same as one `Find()`, and the count does not grow with
-the chunk count or the batch size: the automaton is loaded once per call and every
-chunk or text is scanned against that one snapshot. Before `v1.5.0` each chunk loaded
-its own, so a 63-chunk text issued 63 reads.
+Multi-scan reads cost the same as one `Find()`, and do not grow with chunk count or batch
+size: the automaton is loaded once per call and every chunk or text is scanned against
+that snapshot. Before `v1.5.0` each chunk loaded its own, so a 63-chunk text issued 63
+reads.
```sh
go test -run RTT ./pkg/acor
@@ -50,9 +47,8 @@ ACOR_INTEGRATION_ADDR=localhost:6379 go test -run RTT ./pkg/acor
## Timings
-Measured on: Apple M4, `darwin/arm64`, Go 1.26, Redis 8 on loopback,
-`-benchtime=200x`. The exact command that produced the tables below, which takes
-about 20 seconds:
+Apple M4, `darwin/arm64`, Go 1.26, Redis 8 on loopback, `-benchtime=200x`. This produces
+the tables below in about 20 seconds:
```sh
redis-server --port 6379 --save "" --daemonize yes
@@ -60,8 +56,8 @@ ACOR_INTEGRATION_ADDR=localhost:6379 \
go test -bench RealServer -benchmem -benchtime=200x -run '^$' ./pkg/acor
```
-`make bench` runs the full sweep including the miniredis benchmarks. That takes
-several minutes, and its timings are not published; see the caveats below.
+`make bench` runs the full sweep including the miniredis benchmarks — several minutes, and
+its timings are not published (see [Caveats](#caveats)).
### Find, 1000 keywords
@@ -81,8 +77,8 @@ several minutes, and its timings are not published; see the caveats below.
| V2 + `EnableCache`, warm | 3,120 | 2,857 | 10 | ~29x faster |
| `PresetBalanced` | 2,214 | 2,048 | 4 | ~41x faster |
-At 100 keywords V1 and V2 are close enough that the winner changes between runs;
-only the 1000-keyword gap is stable.
+At 100 keywords V1 and V2 are close enough that the winner changes between runs; only the
+1000-keyword gap is stable.
### Add
@@ -93,10 +89,8 @@ only the 1000-keyword gap is stable.
### Bulk load, `AddMany`
-`AddMany` plans the whole batch in one pass and commits it in a single
-transaction, so it costs two round trips regardless of batch size.
-
-On the Apple M4 with Redis 8 on loopback setup above, one sample measured:
+`AddMany` plans the whole batch in one pass and commits it in a single transaction, so it
+costs two round trips regardless of batch size. One sample on the setup above:
| Keywords | ns/op | B/op | allocs/op |
|---|---|---|---|
@@ -107,49 +101,40 @@ Reproduce with `ACOR_INTEGRATION_ADDR=localhost:6379 make bench-module`.
### Run-to-run variance
-The `ns/op` columns are one run. Repeat runs on the same idle laptop moved every
-absolute number by 20-25% while the ratios held to within about 15%.
-
-That is why the ratios are stated approximately and the raw numbers are labelled a
-sample. Absolute figures differing from ours is expected; *ratios* differing
-substantially is worth reporting as an issue.
-
-## What these numbers mean
+The `ns/op` columns are one run. Repeat runs on the same idle laptop moved every absolute
+number by 20-25% while the ratios held to within about 15% — which is why ratios are
+stated approximately and raw numbers are labelled a sample. Absolute figures differing
+from ours is expected; *ratios* differing substantially is worth an issue.
-**V2 without caching is still somewhat slower than V1 on reads**, because of
-payload rather than round trips. Both cost one trip, but V1's `SMEMBERS` returns
-just the keyword set while V2 must read an outputs hash carrying one entry per
-trie state. The gap widens with the dictionary: roughly parity at 100 keywords,
-~1.7x at 1000.
+## What the numbers mean
-Both schemas memoize the automaton, so an unchanged collection is not re-parsed or
-rebuilt between reads. Uncached V2 was ~9x slower than V1 before that memoization
-landed; what remains is inherent to reading the whole outputs hash, and
-`EnableCache` is the fix for it.
+**V2 without caching is somewhat slower than V1 on reads**, because of payload rather than
+round trips: both cost one trip, but V1's `SMEMBERS` returns just the keyword set while V2
+must read an outputs hash carrying one entry per trie state. Roughly parity at 100
+keywords, ~1.7x at 1000. Both schemas memoize the automaton, so an unchanged collection is
+not re-parsed between reads; uncached V2 was ~9x slower before that memoization, and what
+remains is inherent to reading the whole outputs hash.
-**The large read speedups come from caching, not from the schema.** 15x to 59x
-belongs to `EnableCache` and the `Preset` engines. Choosing V2 and nothing else
-does not deliver them.
+**The large read speedups come from caching, not from the schema.** 15x to 59x belongs to
+`EnableCache` and the `Preset` engines. Choosing V2 alone does not deliver them.
-**V2's unambiguous win is writes.** ~14x on `Add()`, and it is the only schema
-that supports caching or preset engines at all. Use `AddMany` rather than a loop
-over `Add`: it commits in a single transaction. In the Apple M4/Redis 8 loopback
-sample, 1,000 keywords cost 3.0 ms instead of the ~350 ms the same writes cost one
+**V2's unambiguous win is writes** — ~14x on `Add()`, and it is the only schema supporting
+caching or preset engines at all. Use `AddMany` rather than a loop over `Add`: in the
+sample above, 1,000 keywords cost 3.0 ms instead of the ~350 ms the same writes cost one
at a time.
-Practical reading: choose V2, and enable `EnableCache` or a `Preset` if your
-workload is read-heavy. V2 with neither is the one configuration these numbers
-do not recommend.
+Practical reading: choose V2, and enable `EnableCache` or a `Preset` if reads dominate. V2
+with neither is the one configuration these numbers do not recommend.
## Caveats
-- Loopback Redis has almost no network latency. Over a real network both schemas
- still pay one round trip on `Find()`, so V2's larger payload matters more rather
- than less, and V1's per-node `Add()` cost grows with the added latency.
+- Loopback Redis has almost no network latency. Over a real network both schemas still pay
+ one round trip on `Find()`, so V2's larger payload matters more rather than less, and
+ V1's per-node `Add()` cost grows with the added latency.
- These figures compare ACOR configurations against each other, not against other
- Aho-Corasick implementations. A single process with a static dictionary is better
- served by an in-memory library. ACOR earns its cost when several instances share
- one dictionary that changes at runtime.
-- Every other benchmark in the repository runs on miniredis, an in-process
- emulator with no round-trip cost. Those exist for regression detection and are
- deliberately not published here.
+ Aho-Corasick implementations. A single process with a static dictionary is better served
+ by an in-memory library; ACOR earns its cost when several instances share one dictionary
+ that changes at runtime.
+- Every other benchmark in the repository runs on miniredis, an in-process emulator with
+ no round-trip cost. Those exist for regression detection and are deliberately not
+ published here.
diff --git a/docs/content/reference/compatibility.md b/docs/content/reference/compatibility.md
index 3efc17c4..76876b35 100644
--- a/docs/content/reference/compatibility.md
+++ b/docs/content/reference/compatibility.md
@@ -5,73 +5,65 @@ weight: 2
# Compatibility
-Code that imports `github.com/skyoo2003/acor/pkg/acor` keeps compiling, and keeps
-behaving as documented, across every `v1.x.y` release. Nothing is removed from that
-surface inside `v1`. Taking something back requires a `v2` import path, and no `v2` is
-scheduled.
-
-`v1.5.0` is the first supported `v1` release, and the promise is measured from its
-surface. `v1.0.0`-`v1.4.0` are retracted and were never covered — see
-[Retracted versions](https://github.com/skyoo2003/acor/blob/main/RELEASE.md#retracted-versions)
-for why the numbering starts there.
-
-**Upgrading from `v0.11.x` is a one-time exception.** `v1.5.0` removes deprecated members
-and closes paths that `v0.11.x` still allowed, so that upgrade can require changes to your
-code. From `v0.10.x` or earlier, `v0.11.0`'s own breaking changes apply first — the
-`RemoveMany` signature, the removal of `RemoveManyWithOptions`, and the removal of the
-exported V1 key constants — so upgrade through `v0.11.x` and follow the
-[changelog](https://github.com/skyoo2003/acor/blob/main/CHANGELOG.md); this page does not
-restate them.
+Code that imports `github.com/skyoo2003/acor/pkg/acor` keeps compiling, and keeps behaving
+as documented, across every `v1.x.y` release. Nothing is removed from that surface inside
+`v1`. Taking something back requires a `v2` import path, and no `v2` is scheduled.
+
+`v1.5.0` is the first supported `v1` release and the baseline the promise is measured
+from. `v1.0.0`–`v1.4.0` are retracted and were never covered — see
+[Retracted versions](https://github.com/skyoo2003/acor/blob/main/RELEASE.md#retracted-versions).
+
+## Upgrading from `v0.11.x` is the one exception
-From `v0.11.x`, specifically:
+`v1.5.0` removes deprecated members and closes paths `v0.11.x` still allowed, so that one
+upgrade can require changes to your code:
-- `PresetUltimate` is removed. It was an alias for `PresetBalanced`; rename it.
-- `InMemoryInfo` is removed. No exported function accepted or returned it.
+- **`PresetUltimate` is removed.** It aliased `PresetBalanced`; rename it.
+- **`InMemoryInfo` is removed.** No exported function accepted or returned it.
- **V1 collections are read-only.** `Add` and `Remove` return `ErrV1ReadOnly`. Reads,
`Suggest`, `Info`, and `MigrateV1ToV2` still work, so an existing V1 collection can be
- read and converted, but it takes no new keywords. Migrate to V2. `Flush` also still
- works — and still deletes every key in the collection. Read-only means keyword writes
- are refused, not that the collection cannot be destroyed.
+ read and converted, but it takes no new keywords. `Flush` also still works — and still
+ deletes every key. Read-only refuses keyword writes; it does not make the collection
+ indestructible.
-Every upgrade *after* `v1.5.0`, within `v1`, is covered by everything below. Each of
-those three would be a promise violation in `v1.6.0`; `v1.5.0` is the only release that
-can make them.
+From `v0.10.x` or earlier, `v0.11.0`'s own breaking changes apply first — the `RemoveMany`
+signature, the removal of `RemoveManyWithOptions`, and the removal of the exported V1 key
+constants. Upgrade through `v0.11.x` and follow the
+[changelog](https://github.com/skyoo2003/acor/blob/main/CHANGELOG.md); this page does not
+restate them.
+
+Every upgrade *after* `v1.5.0` is covered by everything below. Each of those three changes
+would be a promise violation in `v1.6.0`; `v1.5.0` is the only release that could make
+them.
## What is covered
-| Surface | Covered by `v1` |
-| ---------------------------------------- | ----------------- |
-| Exported identifiers of `pkg/acor` | ✅ |
-| Documented behavior of those identifiers | ✅ |
-| Sentinel error identity (`errors.Is`) | ✅ |
-| On-Redis V2 data format | ✅ additions only |
-| Everything in the next section | ❌ |
-
-## What is not covered
-
-- **The CLI.** Flags, output format, and exit codes of the `acor` command can change in
- any release. If you need a stable contract, call the library rather than parsing CLI
- output.
-- **`acor/server`.** A separate, experimental module. It publishes no tags of its own, so
- it can only be required by pseudo-version, and the core's version numbers say nothing
- about it.
-- **`internal/...`.** Not importable, and free to change in any release.
-- **The `benchmarks` module.** A measurement harness, not an API.
-- **Documentation wording.** Pages get rewritten. What they describe is pinned by this
- page, not by their phrasing.
-- **The V1 schema layout.** Deprecated, and not evolved further. See
- [Deprecation](#deprecation).
+| Surface | Covered by `v1` |
+| ------- | --------------- |
+| Exported identifiers of `pkg/acor` | ✅ |
+| Documented behavior of those identifiers | ✅ |
+| Sentinel error identity (`errors.Is`) | ✅ |
+| On-Redis V2 data format | ✅ additions only |
+
+Not covered:
+
+| Surface | Why |
+| ------- | --- |
+| The `acor` CLI | Flags, output, and exit codes can change in any release. Call the library if you need a stable contract |
+| `acor/server` | A separate, experimental module with no tags of its own; the core's version numbers say nothing about it |
+| `internal/...` | Not importable, free to change |
+| The `benchmarks` module | A measurement harness, not an API |
+| Documentation wording | Pages get rewritten; what they describe is pinned by this page, not by their phrasing |
+| The V1 schema layout | Deprecated and not evolved further — see [Deprecation](#deprecation) |
## Conditions on your code
-The promise holds for code that follows all three. None of them is something the compiler
-can enforce on your behalf.
+The promise holds for code that follows all three. None is enforceable by the compiler.
### Construct option structs with field names
`AhoCorasickArgs`, `MatchOptions`, `BatchOptions`, `ParallelOptions`, and
-`MigrationOptions` gain fields in minor releases. That is not a breaking change for a
-keyed literal:
+`MigrationOptions` gain fields in minor releases. A keyed literal survives that:
```go
@@ -83,150 +75,128 @@ args := &acor.AhoCorasickArgs{
_ = args
```
-An unkeyed literal breaks the moment any field is added, and is not covered by this
-promise:
+An unkeyed literal breaks the moment any field is added, and is not covered:
```go
// Not covered: positional fields break when a field is added.
opts := acor.MatchOptions{acor.MatchKindLeftmostLongest, true, nil}
```
-That literal compiles against `v1.5.0` and stops compiling the moment `MatchOptions` gains
-a fourth field. `go vet`'s `composites` check reports unkeyed literals of another package's
-structs, so this condition is verifiable in your own build.
+`go vet`'s `composites` check reports unkeyed literals of another package's structs, so
+this condition is verifiable in your own build.
The structs ACOR *returns* run the other way. `AhoCorasickInfo`, `MigrationResult`,
`BatchResult`, and `CacheStats` are built by ACOR and read by you, so a new field cannot
-break a reader — and they do gain fields inside `v1`. The condition on them is only that
-you not construct or whole-value compare one: assert on the fields you care about, and a
-later release adding a sixth counter stays invisible to your code.
+break a reader — and they do gain fields inside `v1`. The condition is only that you not
+construct or whole-value compare one: assert on the fields you care about, and a later
+release adding a sixth counter stays invisible.
Their JSON field names are covered too. `api/v1.txt` records struct tags, so renaming
-`json:"status"` is a breaking change even though it moves no Go signature. That applies
-to marshaling a returned struct yourself; the `acor` command's own `--json` output is a
-CLI detail and stays uncovered.
+`json:"status"` is a breaking change even though it moves no Go signature. That applies to
+marshaling a returned struct yourself; the `acor` command's own `--json` output is a CLI
+detail and stays uncovered.
### Do not expect the exported interface to grow
-One exported interface can be implemented from outside the module: `Logger`. No method
-is added to it inside `v1` — doing so would break every existing implementation.
+`Logger` is the one exported interface implementable from outside the module, and no
+method is added to it inside `v1` — that would break every existing implementation.
`KVStorage`, `StringMapResult`, `Subscription`, and `Pipeliner` were exported through
-`v1.4.0` and are not part of the `v1.5.0` surface. Nothing public ever accepted or
-returned one, so no caller could supply an implementation; and freezing them would have
-capped the pluggable-storage work they existed for, since this very rule forbids adding
-a method to them later. They are unexported now so that feature can choose its own
-shape, which `v1` permits as an addition.
+`v1.4.0` and are not part of the `v1.5.0` surface. Nothing public accepted or returned
+one, so no caller could supply an implementation; and freezing them would have capped the
+pluggable-storage work they existed for, since this very rule forbids adding a method
+later. They are unexported now so that feature can pick its own shape.
### Do not dot-import the package
`import . "github.com/skyoo2003/acor/pkg/acor"` puts every exported name into your file's
-scope, so a name added in a minor release can collide with one of your own declarations and
+scope, so a name added in a minor release can collide with one of your declarations and
stop the file compiling. A normal import cannot: additions land behind the `acor.`
-qualifier, where they collide with nothing.
+qualifier.
## The on-Redis data format
-ACOR exists so that several instances share one dictionary, which means a rolling deploy
-runs two ACOR versions against the same Redis keys at the same time. The Go API promise
-says nothing about that case. This section does.
-
-Inside `v1`, changes to the V2 format are **additive only**:
+Several instances share one dictionary, so a rolling deploy runs two ACOR versions against
+the same keys. Inside `v1`, changes to the V2 format are **additive only**:
- Key names, hash tags, and the names and meanings of existing fields do not change.
-- New fields may be added to the `{name}:trie` hash, or under a new key. An ACOR version
- that does not recognize a field there ignores it: that hash is read by looking its
- fields up by name.
+- New fields may be added to the `{name}:trie` hash or under a new key. A version that
+ does not recognize a field there ignores it — that hash is read field by field, by name.
- Nothing is added to the `{name}:outputs` hash. Its field names are automaton states and
- every value in it is read as match data, so it has no room for metadata.
-- A mixed-version fleet therefore keeps working in both directions for the whole `v1`
- line, and a rolling deploy needs no coordinated restart.
+ every value is read as match data, so it has no room for metadata.
-`Flush` rewrites `{name}:trie` in full, so a `Flush` issued by an older instance drops
-fields a newer one added. The next write from an upgraded instance restores them.
+A mixed-version fleet therefore keeps working in both directions for the whole `v1` line,
+and a rolling deploy needs no coordinated restart.
-The cost lands on new features rather than on availability: a feature that depends on a
-newly added field does not take effect until every instance writing to the collection is
-upgraded. Expect it to appear after the rollout finishes, not during it.
+Two consequences worth planning for:
+
+- `Flush` rewrites `{name}:trie` in full, so a `Flush` from an older instance drops fields
+ a newer one added. The next write from an upgraded instance restores them.
+- A feature depending on a newly added field does not take effect until every instance
+ writing to the collection is upgraded. Expect it after the rollout, not during it.
## What counts as a breaking change
A patch release (`x.y.Z`) carries bug fixes and no surface change; a minor release
-(`x.Y.0`) adds to the surface without breaking it; a breaking change requires a new major
-version.
+(`x.Y.0`) adds without breaking; a breaking change requires a new major version.
Removing or renaming an exported identifier, changing a signature, and narrowing a return
-type are the obvious cases, and tooling catches them. These count too, and tooling does
-not catch any of them:
+type are the obvious cases, and tooling catches them. These count too, and tooling catches
+none of them:
-- **Adding a field to an option struct**, for callers using unkeyed literals — which is
- why the condition above exists.
+- **Adding a field to an option struct**, for callers using unkeyed literals — hence the
+ condition above.
- **No longer returning a sentinel error** for the situation that returns it today.
- `errors.Is(err, ErrCacheWithPreset)` continuing to hold is part of the promise; the
- variable merely continuing to exist is not enough.
+ `errors.Is(err, ErrCacheWithPreset)` continuing to hold is the promise; the variable
+ merely continuing to exist is not enough.
- **Changing documented behavior** — match ordering, `MatchKind` semantics, `FindSet`'s
first-match order, `FindParallel`'s deduplication contract.
-- **Changing the meaning of an existing V2 field**, per the section above.
+- **Changing the meaning of an existing V2 field.**
-Undocumented detail is not covered. Sort order among equally-ranked matches, the exact
+Undocumented detail is not covered: sort order among equally-ranked matches, the exact
wording of an error string, allocation counts, and round-trip counts can change in any
release.
## How the promise is enforced
-Most of this page is checked by machine rather than by a reviewer noticing.
-
-`api/v1.txt` in the repository is the covered surface, one symbol per line —
-functions, methods, struct fields, and interface methods. CI regenerates it and fails
-if the result differs from what is committed. A pull request that changes the public
-API therefore has to change that file in the same diff: an addition appends a line, and
-a removal **deletes** one, in front of a reviewer. Nothing stops a removal from being
-merged; what is gone is the possibility of merging one unnoticed.
-
-Two clauses on this page are not covered by that, and both are stated here rather than
-left to be discovered:
-
-- **Documented behavior.** Match ordering, `MatchKind` semantics, `FindSet`'s
- first-match order — no tool compares these *across versions*, and none can: the
- claim lives in prose. What is checked is that somebody looked. `api/v1-audit.txt`
- carries one verdict per entry of `api/v1.txt`, and CI fails when an entry has none,
- so a *missing* verdict is a build failure rather than an omission nobody notices.
- A verdict of `unaudited` is legal, so an unreviewed entry does not fail the build —
- it is counted instead, and the tally prints on every run. Whether each verdict is
- *right*, and whether its cited `file:line` still points where it did, is review.
-
- Every entry of the `v1.5.0` surface has been read against the code it describes,
- and 38 of them said something the code did not do. Those sentences were rewritten
- before the freeze, so what `v1` promises is the corrected wording, not the wording
- that shipped in `v1.4.0`. A non-`unaudited` verdict has to cite the `file:line`
- the behavior was read at, which is what stops a future pass from marking a line
- reviewed without reviewing it.
-
- Those verdicts still describe the code as shipped. `v1.5.1` changed no entry of
- the surface they were measured against: `git diff v1.5.0 v1.5.1 -- api/v1.txt` is
- empty, and the release touched only `LICENSE`, `server/LICENSE`, and the changelog —
- no Go file at all.
-- **Cross-version Redis interop.** Tests pin the key names and hash field names, so a
- *rename* fails CI. They do not run an older ACOR against a newer one's data, so
- additive-only is verified as far as naming and no further.
-
-Retractions are checked too: CI fails if `retract [v1.0.0, v1.4.0]` leaves `go.mod`,
-because a release without it silently makes the retracted range resolvable again.
+`api/v1.txt` is the covered surface, one symbol per line — functions, methods, struct
+fields, interface methods. CI regenerates it and fails if the result differs from what is
+committed, so a PR changing the public API has to change that file in the same diff: an
+addition appends a line, a removal **deletes** one, in front of a reviewer. Nothing stops
+a removal from being merged; what is gone is merging one unnoticed.
+
+CI also fails if `retract [v1.0.0, v1.4.0]` leaves `go.mod`, because dropping it would
+silently make the retracted range resolvable again.
+
+Two clauses are not covered by that check:
+
+- **Documented behavior.** No tool compares match ordering or `MatchKind` semantics
+ *across versions* — the claim lives in prose. What is checked is that somebody looked:
+ `api/v1-audit.txt` carries one verdict per entry of `api/v1.txt`, and CI fails when an
+ entry has none. A verdict of `unaudited` is legal but counted, and the tally prints on
+ every run. A non-`unaudited` verdict must cite the `file:line` the behavior was read at.
+
+ Every entry of the `v1.5.0` surface was read against its code, and 38 said something the
+ code did not do. Those sentences were rewritten before the freeze, so `v1` promises the
+ corrected wording, not what shipped in `v1.4.0`. `v1.5.1` changed no entry of that
+ surface — `git diff v1.5.0 v1.5.1 -- api/v1.txt` is empty, and the release touched only
+ `LICENSE`, `server/LICENSE`, and the changelog.
+- **Cross-version Redis interop.** Tests pin key names and hash field names, so a *rename*
+ fails CI. They do not run an older ACOR against a newer one's data, so additive-only is
+ verified as far as naming and no further.
## Deprecation
A deprecated identifier keeps working for the rest of `v1`. It carries a `Deprecated:`
-line naming its replacement, gains no new capability, and is removed no earlier than
-`v2`.
+line naming its replacement, gains no new capability, and is removed no earlier than `v2`.
-Deprecated today: `SchemaV1`, which is read-only and whose read path is removed no
-earlier than `v2`.
+Deprecated today: `SchemaV1`, read-only, whose read path is removed no earlier than `v2`.
## `v2`
-A `v2` would ship as the module `github.com/skyoo2003/acor/v2`, imported as
+A `v2` would ship as `github.com/skyoo2003/acor/v2`, imported as
`github.com/skyoo2003/acor/v2/pkg/acor`. `v1` stays importable and unchanged at its own
-path, so nothing breaks by the act of publishing `v2`. There is no schedule.
+path, so publishing `v2` breaks nothing by itself. There is no schedule.
-Security patches follow the [security policy](https://github.com/skyoo2003/acor/blob/main/SECURITY.md),
-which defines which lines receive them.
+Security patches follow the
+[security policy](https://github.com/skyoo2003/acor/blob/main/SECURITY.md).
diff --git a/docs/content/reference/r2-r3-performance.md b/docs/content/reference/r2-r3-performance.md
index 6d107987..ce287cb9 100644
--- a/docs/content/reference/r2-r3-performance.md
+++ b/docs/content/reference/r2-r3-performance.md
@@ -3,52 +3,51 @@ title: "R2/R3 verification and R1 comparison"
description: "Incremental download, engine memory and bounded source-text API evidence."
---
-The R2/R3 implementation completed 36 real-server runs on 2026-09-06 using the
-same R1 workload: Redis/Valkey × 10,000/100,000/1,000,000 keywords × shared/diverse/
-Korean distributions × two repetitions. The R1 measurements remain unchanged in
-`benchmarks/results/v3-20260906.json`. New measurements, source hashes and safety
-results are in `benchmarks/results/r2-r3-20260906.json`.
-
-Environment: Apple M4, 10 logical CPUs, 16 GiB RAM, macOS 26.5.2, Go 1.26.7,
-Redis 8.10.1 and source-built Valkey 9.1.2. Both are standalone localhost TCP,
-with RDB/AOF disabled. The preset remains MemoryEfficient; polling remains 50 ms
-for measurement. R2 adds the default fixed 20 ms refresh debounce. Each run uses
-a fresh Go process. These are developer-workstation measurements with possible
-background host activity, not latency guarantees or confidence intervals.
-
-Valkey uses `prefetch-batch-max-size 0`, as in the successful R1 baseline. The
-local default-prefetch build previously crashed; this work does not establish
-support for that configuration or diagnose its root cause. See the
-[R1 report](../versioned-performance/) for source checksum and crash evidence.
-
-## R2 changes
-
-- The installed engine keeps verified immutable keyword slices and its manifest.
- Unchanged buckets share those slices; only changed nonempty buckets download
- chunk data. Candidate caches are installed with the engine only after success.
-- MemoryEfficient builds consume the bucket sequence directly, avoiding the
- flattened full dictionary slice and an additional keyword-set map. Nodes keep
- their only child inline; a map is allocated only when the node branches. BFS
- uses a compact integer queue. The first-rune Bloom filter sizes itself from
- distinct first runes rather than total keyword count.
-- Builds check cancellation in insertion, failure-link construction and table
- filling for all three presets. Cancellation discards the private candidate;
- no detached builder goroutine continues after return. Allocations, rune-count
- calls and sorting finish before the next checkpoint.
-- A fixed debounce window merges burst notifications. New events cannot extend
- that window; an in-flight build completes under ordinary writes, preventing
- repeated cancellation from starving refresh. Close cancels its build context.
-- Status reports the last successful build's DownloadedBuckets/ReusedBuckets and
- cumulative CompletedBuilds. Redis schema, expected-version writes, receipt
- resolution, leases and pruning semantics are unchanged.
+# R2/R3 verification and R1 comparison
+
+R2/R3 completed 36 real-server runs on 2026-09-06 using the R1 workload: Redis/Valkey ×
+10,000/100,000/1,000,000 keywords × shared/diverse/Korean distributions × two repetitions.
+The R1 measurements are unchanged in `benchmarks/results/v3-20260906.json`; new
+measurements, source hashes, and safety results are in
+`benchmarks/results/r2-r3-20260906.json`.
+
+Environment: Apple M4, 10 logical CPUs, 16 GiB RAM, macOS 26.5.2, Go 1.26.7, Redis 8.10.1,
+source-built Valkey 9.1.2. Both standalone on localhost TCP with RDB/AOF disabled. Preset
+remains MemoryEfficient and polling remains 50 ms for measurement; R2 adds the default
+fixed 20 ms refresh debounce. Each run uses a fresh Go process. These are
+developer-workstation measurements with possible background host activity — not latency
+guarantees or confidence intervals.
+
+Valkey uses `prefetch-batch-max-size 0`, as in the R1 baseline. The local default-prefetch
+build previously crashed; this work neither establishes support for that configuration nor
+diagnoses it. See the [R1 report](../versioned-performance/) for the source checksum and
+crash evidence.
+
+## What R2 changed
+
+- The installed engine keeps verified immutable keyword slices and its manifest. Unchanged
+ buckets share those slices; only changed nonempty buckets download chunk data. Candidate
+ caches install with the engine, and only after success.
+- MemoryEfficient builds consume the bucket sequence directly, avoiding the flattened
+ dictionary slice and an extra keyword-set map. Nodes keep their only child inline and
+ allocate a map only when they branch; BFS uses a compact integer queue; the first-rune
+ Bloom filter sizes itself from distinct first runes rather than total keyword count.
+- All three presets check cancellation during insertion, failure-link construction, and
+ table filling. Cancellation discards the private candidate with no detached builder
+ goroutine left running. Allocations, rune-count calls, and sorting finish before the next
+ checkpoint.
+- A fixed debounce window merges burst notifications and cannot be extended by new events,
+ so an in-flight build completes under ordinary writes rather than being starved by
+ repeated cancellation. `Close` cancels its build context.
+- `Status` reports the last successful build's `DownloadedBuckets`/`ReusedBuckets` and
+ cumulative `CompletedBuilds`. Redis schema, expected-version writes, receipt resolution,
+ leases, and pruning semantics are unchanged.
## Million-keyword comparison
-Ranges below are the minimum–maximum of two runs. Ready is write preparation,
-commit, refresh detection and engine construction combined. RSS is the process
-high-water mark across the full run, including all update scenarios. Network
-counters are server-wide and include protocol, polling and engine downloads;
-they do not isolate only keyword payloads.
+Ranges are minimum–maximum of two runs. Ready combines write preparation, commit, refresh
+detection, and engine construction. RSS is the process high-water mark across the full run.
+Network counters are server-wide and include protocol, polling, and engine downloads.
| Server | Distribution | Add 1 ready: R1 → R2 (s) | Received bytes reduction for add 1 | Peak RSS: R1 → R2 (GiB) |
|---|---|---:|---:|---:|
@@ -59,10 +58,10 @@ they do not isolate only keyword payloads.
| valkey | diverse | 2.33–2.40 → 0.92–1.27 | 91.9% | 3.02–3.80 → 2.02–2.19 |
| valkey | korean | 2.90–3.40 → 1.21–1.60 | 96.2% | 3.37–3.69 → 2.33–2.50 |
-All six million-keyword server/distribution combinations reduced single-change
-received bytes and peak RSS in these runs. A full engine is still rebuilt after
-a change; this is not an incremental automaton or a constant-memory design.
-Larger changes touch more buckets and therefore reuse less downloaded data.
+All six million-keyword combinations reduced single-change received bytes and peak RSS in
+these runs. A full engine is still rebuilt after a change — this is not an incremental
+automaton or a constant-memory design, and larger changes touch more buckets and reuse
+less.
| Server | Distribution | Add 1,000 ready: R1 → R2 (s) | Add 1% ready: R1 → R2 (s) | Full replace ready: R1 → R2 (s) |
|---|---|---:|---:|---:|
@@ -73,50 +72,48 @@ Larger changes touch more buckets and therefore reuse less downloaded data.
| valkey | diverse | 2.42–2.60 → 1.21–1.29 | 3.03–3.54 → 2.14–2.19 | 4.08–4.33 → 2.77–3.61 |
| valkey | korean | 4.03–4.38 → 1.45–1.60 | 3.98–4.49 → 2.64–2.70 | 6.78–7.31 → 4.17–4.92 |
-Improvement is not uniform: some Redis large-change repetitions overlap or
-exceed the R1 timing range. Two repetitions do not establish a universal speedup.
+Improvement is not uniform: some Redis large-change repetitions overlap or exceed the R1
+range. Two repetitions do not establish a universal speedup.
-The JSON also records 10,000/100,000-entry results, initial loading/startup,
-identical replacement, removals, search p50/p95/p99 while refreshing, Redis
-memory and pruning. The same seed 20260906 and generation functions are used.
-Prune measurements simulate the retention horizon by aging generation registry
-scores; they do not wait 24 hours. Separate million-entry safety tests passed
-against both servers after the matrix, with no concurrent workload on a measured
-endpoint.
+The JSON also records 10,000/100,000-entry results, initial loading and startup, identical
+replacement, removals, search p50/p95/p99 while refreshing, Redis memory, and pruning,
+using the same seed 20260906 and generation functions. Prune measurements simulate the
+retention horizon by aging generation registry scores rather than waiting 24 hours.
+Separate million-entry safety tests passed against both servers after the matrix, with no
+concurrent workload on a measured endpoint.
## R3 contracts and verification
-`Scan`, `MaskText` and `ReplaceText` are additive on AhoCorasick and
-VersionedCollection. See [bounded text processing](../text-processing/) for the
-complete defaults and error behavior. Existing unlimited APIs are unchanged.
+`Scan`, `MaskText`, and `ReplaceText` are additive on `AhoCorasick` and
+`VersionedCollection`; existing unlimited APIs are unchanged. Defaults and error behavior:
+[bounded text processing](../text-processing/).
| Area | Evidence |
|---|---|
-| Original positions | All presets match existing FindMatches for overlapping and leftmost-longest results; original byte slices survive Korean, emoji, `İ` case folding and malformed UTF-8 input |
-| Result limits | At most MaxMatches are retained; Truncated requires an additional eligible result |
-| Work/input bounds | Oversize input and excess raw candidates fail explicitly, including candidates later filtered by word boundaries/overlap |
-| Atomic rewrite | Match, output and work exhaustion return no partial rewrite; empty literal replacement deletes; masks preserve original rune count, including NUL masks |
+| Original positions | All presets match existing `FindMatches` for overlapping and leftmost-longest results; original byte slices survive Korean, emoji, `İ` case folding, and malformed UTF-8 |
+| Result limits | At most `MaxMatches` retained; `Truncated` requires an additional eligible result |
+| Work/input bounds | Oversize input and excess raw candidates fail explicitly, including candidates later filtered by word boundaries or overlap |
+| Atomic rewrite | Match, output, and work exhaustion return no partial rewrite; empty literal replacement deletes; masks preserve original rune count, including NUL masks |
| Overlap | A bounded pending-start heap selects leftmost-longest matches without collecting all raw matches |
-| Fuzz | FuzzScanLeftmostParity completed 30 seconds and 819,033 executions against the existing leftmost-longest implementation |
-| Refresh | Verified bucket reuse, failed-candidate cache preservation and burst coalescing tests pass |
+| Fuzz | `FuzzScanLeftmostParity` completed 30 seconds and 819,033 executions against the existing leftmost-longest implementation |
+| Refresh | Bucket reuse, failed-candidate cache preservation, and burst coalescing tests pass |
| Cancellation | All three presets preserve the previous engine on cancellation during sequence ingestion and CPU construction; unrelated panics are not swallowed |
-| Safety | Both real servers pass the million-entry conflict, pinned paging and lease-safe pruning scenario |
+| Safety | Both real servers pass the million-entry conflict, pinned paging, and lease-safe pruning scenario |
-Resource bounds are per call. Scratch input indexing is proportional to the
-bounded input size, and the pending window is bounded by input/longest-keyword
-length and candidate count. A caller's custom WordRune function controls its own
-resource use. R3 does not claim a wall-clock bound on callbacks or allocation.
+Resource bounds are per call: scratch input indexing is proportional to the bounded input
+size, and the pending window is bounded by input/longest-keyword length and candidate
+count. A custom `WordRune` controls its own resource use. R3 claims no wall-clock bound on
+callbacks or allocation.
-Root and server race tests, root/server lint, all-module vet, benchmark-module
-tests, API snapshot/audit and documentation compilation pass. Legacy Find/Add
-fuzz targets are rerun for 30 seconds each. No HTTP/gRPC surface or storage
-schema migration is introduced by R2/R3.
+Root and server race tests, root/server lint, all-module vet, benchmark-module tests, API
+snapshot/audit, and documentation compilation pass. Legacy Find/Add fuzz targets rerun for
+30 seconds each. R2/R3 introduces no HTTP/gRPC surface and no storage schema migration.
## Reproduce
Run the unchanged `scripts/benchmark-v3.sh` against each disposable endpoint with
-`ACOR_V3_SCALE_REPEATS=2` and separate output directories. See the
-[R1 reproduction instructions](../versioned-performance/#reproduce) for the exact
-environment and dictionary distributions. The R2/R3 defaults select the new
-engine and refresh behavior automatically; existing V1/V2 public contracts remain
-covered by the regression suite.
+`ACOR_V3_SCALE_REPEATS=2` and separate output directories; the exact environment and
+dictionary distributions are in the
+[R1 reproduction instructions](../versioned-performance/#reproduce). The R2/R3 defaults
+select the new engine and refresh behavior automatically, and existing V1/V2 public
+contracts remain covered by the regression suite.
diff --git a/docs/content/reference/schema-v1.md b/docs/content/reference/schema-v1.md
index 4d23f89c..ae571a7a 100644
--- a/docs/content/reference/schema-v1.md
+++ b/docs/content/reference/schema-v1.md
@@ -5,21 +5,16 @@ weight: 3
# Schema V1 (Deprecated)
-V1 is the original ACOR storage schema. It uses multiple Redis keys per collection.
-
> **V1 is deprecated and read-only as of `v1.5.0`.** Reads, `Suggest`, and `Info` work,
> and `MigrateV1ToV2` converts a collection in place, but `Add` and `Remove` return
-> `ErrV1ReadOnly`. `Flush` also still works, and still deletes every key in the
-> collection — read-only refuses keyword writes, it does not protect the collection from
-> `Flush`. It gains no features either: preset engines and
-> `EnableCache` both require V2. New collections should use the default V2 schema.
-> The read path stays for the whole `v1` line and is removed no earlier than `v2`.
-
-## Overview
+> `ErrV1ReadOnly`. `Flush` still works and still deletes every key — read-only refuses
+> keyword writes, it does not protect the collection. V1 gains no features: preset engines
+> and `EnableCache` both require V2. The read path stays for the whole `v1` line and is
+> removed no earlier than `v2`.
-V1 creates approximately 5 keys per 100 keywords:
+V1 spreads one collection across many keys — roughly 5 per 100 keywords.
-| Key Pattern | Purpose |
+| Key pattern | Purpose |
|-------------|---------|
| `{name}:keyword` | Set of keywords |
| `{name}:prefix` | Trie prefix edges |
@@ -27,53 +22,14 @@ V1 creates approximately 5 keys per 100 keywords:
| `{name}:output:{state}` | Output keywords per state |
| `{name}:node:{keyword}` | Node metadata |
-## Performance Characteristics
-
-| Operation | Complexity |
-|-----------|------------|
-| Find() | O(N×3-5) RTT |
-| Add() | O(M×3-10) RTT — no longer reachable; kept to explain the migration's value |
-
-Where:
-- N = number of trie states visited
-- M = keyword length
+`Find()` costs O(N×3-5) round trips for N visited states. `Add()` cost O(M×3-10) for a
+keyword of length M — no longer reachable, and recorded only to explain what migration is
+worth. Compare against [Schema V2](../schema-v2/#against-v1).
-## When to Use V1
-
-- Existing collections using V1
-- Small keyword sets (< 10,000)
-- Migration not feasible
-
-## Migration to V2
+## Migrating
```bash
-# Preview migration
-acor -name mycollection migrate --dry-run
-
-# Execute migration
-acor -name mycollection migrate
-
-# Rollback to V1
-acor -name mycollection migrate-rollback
+acor -name mycollection migrate --dry-run # preview
+acor -name mycollection migrate # execute
+acor -name mycollection migrate-rollback # back to V1
```
-
-## Key Structure Diagram
-
-```mermaid
-graph LR
- A[keyword set] --> B[prefix trie]
- B --> C[suffix links]
- C --> D[output sets]
- D --> E[node metadata]
-```
-
-## Limitations
-
-- Higher memory usage
-- More Redis keys to manage
-- Slower Find() operations
-- More network round-trips
-
-## Recommendation
-
-**Migrate to V2** for new collections or when performance is critical.
diff --git a/docs/content/reference/schema-v2.md b/docs/content/reference/schema-v2.md
index 643ec4c4..8695f220 100644
--- a/docs/content/reference/schema-v2.md
+++ b/docs/content/reference/schema-v2.md
@@ -5,128 +5,59 @@ weight: 4
# Schema V2 (Optimized)
-V2 is the recommended schema for ACOR. A collection occupies a fixed set of at
-most 3 keys, whatever the dictionary size.
+V2 is the default for new collections — no configuration needed. A collection occupies at
+most 3 keys whatever the dictionary size.
-## Overview
+| Key | Holds | When it exists |
+|-----|-------|----------------|
+| `{name}:trie` | Serialized trie: keywords, prefixes, version | Always, from creation |
+| `{name}:outputs` | Output mappings, state → keywords | Once the collection has a keyword |
+| `{name}:nodes` | Node metadata | Only after `MigrateV1ToV2`; cleaned up by flush |
-V2 consolidates storage into these keys:
+Most collections hold two keys; a fresh one holds a single `:trie`. Nothing but migration
+writes `:nodes`. Budget for three, expect to count fewer.
-| Key Pattern | Purpose | When it exists |
-|-------------|---------|----------------|
-| `{name}:trie` | Serialized trie structure (keywords, prefixes, version) | Always, from creation |
-| `{name}:outputs` | All output mappings (state -> keywords) | Once the collection has a keyword |
-| `{name}:nodes` | Node metadata | Only on a collection produced by `MigrateV1ToV2`; cleaned up by flush |
+## Field layout
-Most collections therefore hold two keys, and a freshly created one holds a
-single `:trie`. Nothing but migration writes `:nodes`, so a collection built with
-`Add` never has it. Budget for three; expect to count fewer.
+```text
+{name}:trie (hash)
+ keywords -> ["keyword1", "keyword2", ...]
+ prefixes -> ["", "h", "he", ...]
+ version ->
+
+{name}:outputs (hash)
+ he -> ["he"]
+ she -> ["he", "she"]
-## Performance Characteristics
+{name}:nodes (hash, migration only)
+ keyword1 -> ["s0","s1","s2"]
+```
-| Operation | Complexity |
-|-----------|------------|
-| Find() | 1 RTT (fixed), 0 RTT with EnableCache |
-| Add() | 2 RTT (flat, independent of keyword length) |
+Collections written before v0.11 also carry a `suffixes` field on `:trie`. It is never
+read, writes leave it alone, and the next `Flush()` drops it.
-## Comparison with V1
+## Against V1
-Round trips are counted by tests on every CI run; timings are a sample from the
-hardware named on the [benchmarks page](../benchmarks/).
+Round trips are counted by tests on every CI run; timings are a sample from the hardware
+named on the [benchmarks page](../benchmarks/).
| Metric | V1 | V2 |
|--------|----|----|
| Keys per 100K keywords | ~500K | 2 |
| `Find()` round trips | 1 | 1 |
-| `Add()` round trips | grows with keyword length (53 at 5 chars, 507 at 26) | 2 |
+| `Add()` round trips | Grows with keyword length: 53 at 5 chars, 507 at 26 | 2 |
| `Add()` time | baseline | ~14x faster |
| `Find()` time, no cache | baseline | ~1.7x **slower** at 1000 keywords |
-V2's win is on writes, not reads. Both schemas read in a single round trip, and
-uncached V2 is slightly slower on `Find()` because it reads an outputs hash with
-one entry per trie state where V1 reads only the keyword set. The large read
-speedups belong to `EnableCache` and the preset engines, not to the schema. See
-[Benchmarks](../benchmarks/) for the full tables and how to reproduce them.
-
-## Architecture
-
-```mermaid
-graph TB
- subgraph V2 Schema
- A[trie key] --> B[Serialized Trie]
- C[outputs key] --> D[Output Map]
- E[nodes key] --> F[Node Metadata - migration only]
- end
-
- G[Find Operation] --> A
- G --> C
- G -.-> E
-```
-
-## Enabling V2
-
-V2 is automatically used for new collections. No configuration needed.
-
-```go
-ac, err := acor.Create(&acor.AhoCorasickArgs{
- Addr: "localhost:6379",
- Name: "my-v2-collection",
-})
-if err != nil {
- log.Fatal(err)
-}
-// Automatically uses V2 schema
-```
+**V2's win is writes, not reads.** Both schemas read in one round trip, and uncached V2 is
+slightly slower on `Find()` because it reads an outputs hash with one entry per trie state
+where V1 reads only the keyword set. The large read speedups belong to `EnableCache` and
+the preset engines, not to the schema.
-## Migration from V1
+## Migrating from V1
```bash
-# Check current schema
-acor -name mycollection schema-version
-
-# Preview migration
+acor -name mycollection schema-version # what you have now
acor -name mycollection migrate --dry-run
-
-# Execute migration
acor -name mycollection migrate
```
-
-## Key Structure
-
-### trie key
-
-Stores the serialized trie as a hash with three fields:
-
-```text
-{collection}:trie
- keywords -> ["keyword1", "keyword2", ...]
- prefixes -> ["", "h", "he", ...]
- version ->
-```
-
-Collections written before v0.11 also carry a `suffixes` field. It is never
-read, is left alone by writes, and is dropped by the next `Flush()`.
-
-### outputs key
-
-Stores output keywords per trie state as a hash:
-
-```text
-{collection}:outputs
- he -> ["he"]
- she -> ["he", "she"]
-```
-
-### nodes key
-
-A hash mapping each keyword to a JSON array of its trie state strings. It is
-populated only by V1→V2 migration (fresh V2 collections do not write it):
-
-```text
-{collection}:nodes (hash)
- keyword1 -> ["s0","s1","s2"]
-```
-
-## Recommendation
-
-**Use V2 for all new collections.** It provides significantly better performance and lower resource usage.
diff --git a/docs/content/reference/text-processing.md b/docs/content/reference/text-processing.md
index ebfee5b5..f365ebfe 100644
--- a/docs/content/reference/text-processing.md
+++ b/docs/content/reference/text-processing.md
@@ -3,73 +3,71 @@ title: "Bounded search, masking and replacement"
description: "Original byte and rune positions with explicit work and output limits."
---
-R3 adds `Scan`, `MaskText`, and `ReplaceText` to both AhoCorasick and
-VersionedCollection. All take a context. Existing Find/FindMatches APIs retain
-their existing unlimited result and position contracts. V3 calls hold one serving
-engine for the entire scan or rewrite, even if a refresh finishes concurrently.
+# Bounded search, masking and replacement
-## Original positions and bounded search
+`Scan`, `MaskText`, and `ReplaceText` exist on both `AhoCorasick` and
+`VersionedCollection` and all take a context. The existing `Find`/`FindMatches` APIs keep
+their unlimited result and position contracts. A V3 call holds one serving engine for the
+whole scan or rewrite, even if a refresh finishes concurrently.
-`Scan(ctx, text, *ScanOptions)` returns SourceMatch entries containing the normalized
-Keyword, the original Text substring, and half-open Start/End rune and
-ByteStart/ByteEnd byte offsets into the original input. Unicode case folding can
-change byte lengths: `İSTANBUL` matches `istanbul`, but its byte span still selects
-the original spelling. No normalization of the original output text occurs.
-Invalid UTF-8 input bytes are decoded as RuneError for matching, while reported
-byte slices retain those exact original bytes.
+## Scan
-Zero-valued limits select these defaults; negative limits are rejected. Construct
-ScanOptions and RewriteOptions using named fields for forward compatibility.
+`Scan(ctx, text, *ScanOptions)` returns `SourceMatch` entries carrying the normalized
+`Keyword`, the original `Text` substring, and half-open `Start`/`End` rune and
+`ByteStart`/`ByteEnd` byte offsets into the **original** input.
-| Option | Default | Behavior at limit |
+Case folding can change byte lengths — `İSTANBUL` matches `istanbul`, and its byte span
+still selects the original spelling. Output text is never normalized. Invalid UTF-8 is
+decoded as `RuneError` for matching, while the reported byte slices keep those exact
+original bytes.
+
+| Option | Default | At the limit |
|---|---:|---|
-| MaxInputBytes | 1 MiB | ErrInputLimit before loading/scanning the engine |
-| MaxMatches | 1,000 | At most this many entries; Truncated when another eligible match is observed |
-| MaxCandidates | 100,000 | ErrScanWorkLimit when another raw automaton match is encountered |
-
-Kind defaults to MatchKindOverlapping. MatchKindLeftmostLongest selects the
-leftmost available start, then the longest keyword there, and discards overlaps.
-The implementation keeps the longest candidate at each pending start until the
-engine's longest keyword makes the decision safe. It does not collect all raw
-matches and sort them. Scratch memory is bounded by input length plus the pending
-start window and retained result count; it is not constant memory. Input indexing
-uses rune and byte-offset arrays. A low result limit alone does not replace the
-input-byte or candidate-work limit.
-
-WholeWord and WordRune have the same normalized-rune boundary semantics as
-MatchOptions. Letters, digits, combining marks and underscore are word characters
-by default. Scripts without spaces may require an application-specific WordRune.
-Candidate counting happens before whole-word and overlap filtering, so a dense
-rejected-match workload cannot bypass the work budget. Custom WordRune code runs
-in the caller's goroutine; its own resource use is the caller's responsibility.
-
-Input or candidate exhaustion and cancellation return an error with no result.
-Truncated concerns the eligible result count only. A caller checking safety or
-completeness must handle both errors and Truncated; neither means a clean input.
-The context is checked during input indexing, traversal and match handling.
-
-## Atomic masking and literal replacement
-
-`ReplaceText(ctx, text, replacement, *RewriteOptions)` inserts the literal
-replacement for each selected non-overlapping leftmost-longest match. It does not
-interpret regular expressions, expand captures, or search replacement text again.
-An empty replacement deletes matched spans.
-
-`MaskText(ctx, text, maskRune, *RewriteOptions)` writes one mask rune for every
-matched original rune. The result preserves rune count, not necessarily byte count.
-Any valid Unicode rune, including NUL, is accepted. Invalid runes are rejected.
-Both APIs leave unmatched original bytes untouched.
-
-RewriteOptions shares the three scan limits above and adds MaxOutputBytes, which
-defaults to 4 MiB. WholeWord and WordRune are also available. Overlap selection is
-always leftmost-longest. Exceeding MaxMatches produces ErrMatchLimit; exceeding the
-output bound produces ErrOutputLimit. Every error returns no RewriteResult, so a
-caller cannot accidentally consume a partially masked document. All output sizes
-are checked before allocating the output buffer.
-
-RewriteResult contains Text and the SourceMatch entries used for the rewrite.
-Their offsets always refer to the **input**, even when replacement changes output
-length. Match Text substrings can retain the original input string in memory.
+| `MaxInputBytes` | 1 MiB | `ErrInputLimit`, before the engine is loaded or scanned |
+| `MaxMatches` | 1,000 | Keeps at most this many; sets `Truncated` when another eligible match appears |
+| `MaxCandidates` | 100,000 | `ErrScanWorkLimit` on the next raw automaton match |
+
+Zero-valued limits select the defaults; negative limits are rejected. Construct
+`ScanOptions` and `RewriteOptions` with named fields.
+
+`Kind` defaults to `MatchKindOverlapping`. `MatchKindLeftmostLongest` takes the leftmost
+available start, then the longest keyword there, and discards overlaps — implemented by
+keeping the longest candidate at each pending start until the engine's longest keyword
+makes the decision safe, not by collecting and sorting all raw matches. Scratch memory is
+bounded by input length plus the pending-start window and retained results; it is not
+constant.
+
+`WholeWord` and `WordRune` carry the same normalized-rune semantics as `MatchOptions`:
+letters, digits, combining marks, and underscore are word characters by default, and
+scripts without spaces usually need an application-specific `WordRune`. Candidates are
+counted *before* whole-word and overlap filtering, so a dense rejected-match workload
+cannot slip past the work budget. Custom `WordRune` code runs in the caller's goroutine.
+
+**Errors and `Truncated` are different signals.** Input exhaustion, candidate exhaustion,
+and cancellation all return an error and no result; `Truncated` concerns only the eligible
+result count. Checking for a clean input means handling both. The context is checked
+during input indexing, traversal, and match handling.
+
+## Masking and replacement
+
+`ReplaceText(ctx, text, replacement, *RewriteOptions)` inserts the literal replacement for
+each selected non-overlapping leftmost-longest match. No regular expressions, no capture
+expansion, no re-searching the replacement. An empty replacement deletes matched spans.
+
+`MaskText(ctx, text, maskRune, *RewriteOptions)` writes one mask rune per matched original
+rune, preserving rune count but not necessarily byte count. Any valid Unicode rune is
+accepted, including NUL; invalid runes are rejected.
+
+Both leave unmatched original bytes untouched. `RewriteOptions` shares the three scan
+limits and adds `MaxOutputBytes`, defaulting to 4 MiB. Overlap selection is always
+leftmost-longest. Exceeding `MaxMatches` gives `ErrMatchLimit`, exceeding the output bound
+gives `ErrOutputLimit`, and **every error returns no `RewriteResult`** — a caller cannot
+accidentally consume a half-masked document. Output size is checked before the buffer is
+allocated.
+
+`RewriteResult` holds `Text` plus the `SourceMatch` entries used. Their offsets always
+refer to the **input**, even when replacement changes the output length. Match `Text`
+substrings can retain the original input string in memory.
```go
@@ -112,5 +110,5 @@ func main() {
}
```
-See the [R2/R3 report](../r2-r3-performance/) for parity, resource-bound and
-cancellation evidence and the same-environment R1 comparison.
+Parity, resource-bound, and cancellation evidence:
+[R2/R3 report](../r2-r3-performance/).
diff --git a/docs/content/reference/versioned-performance.md b/docs/content/reference/versioned-performance.md
index 7bd257ed..6bf32885 100644
--- a/docs/content/reference/versioned-performance.md
+++ b/docs/content/reference/versioned-performance.md
@@ -3,50 +3,49 @@ title: "V3 performance and acceptance report"
description: "Reproducible million-keyword measurements on Redis and Valkey."
---
-This is the archived R1 baseline. See [R2/R3 comparison](../r2-r3-performance/)
-for the subsequent implementation and measurements.
+# V3 performance and acceptance report
-Measured on 2026-09-06. The final matrix completed **36 runs**: Redis and Valkey,
-three dictionary sizes, three deterministic distributions, two repetitions each.
-The separate million-entry safety scenario passed on both configured servers.
-These are measurements of one environment, not throughput or memory guarantees.
+The archived R1 baseline. For the implementation and measurements that followed, see the
+[R2/R3 comparison](../r2-r3-performance/).
-## Environment and methodology
+Measured 2026-09-06. The final matrix completed **36 runs**: Redis and Valkey × three
+dictionary sizes × three deterministic distributions × two repetitions. The separate
+million-entry safety scenario passed on both servers. These are measurements of one
+environment, not throughput or memory guarantees.
+
+## Environment and method
- Apple M4, 10 logical CPUs, 16 GiB RAM; macOS 26.5.2, Darwin 25.5.0, arm64.
- Go 1.26.7; Redis 8.10.1; source-built Valkey 9.1.2.
-- Standalone TCP on localhost; RDB snapshots and AOF disabled. No production
- durability or cross-node network overhead is included.
-- MemoryEfficient engine; 50 ms version polling for the measurement (the library
- default is 30 seconds); five-minute leases renewed every minute.
-- Separate Go process per size/distribution/repetition; two complete repetitions.
- This was a developer workstation, with some development checks running on the
- host. Treat the ranges as baseline observations, not isolated microbenchmark
- confidence intervals. Each server ran one scale workload at a time.
-- Fixed seed 20260906. Shared: `shared-prefix-%08d`. Diverse: a random 12-bit hex
- prefix plus a unique eight-digit index. Korean: `한국어-%04x-키워드-%08d`, using a
- random 16-bit prefix. Full replacement appends `-x` to every keyword.
-
-**Valkey qualification:** an exploratory run of the local Valkey 9.1.2 build with
-its default prefetch setting exited with SIGSEGV in `hashtableIncrementalFindStep`
-via `prefetchCommandQueueKeys`. The final successful Valkey matrix used
-`prefetch-batch-max-size 0`. This report does not establish support for that local
-build's default configuration. The stack excerpt is preserved in
-`benchmarks/results/v3-20260906-valkey-crash.txt`; the library returned an error
-when the server disappeared. The root cause has not been established.
-
-The Valkey source archive SHA-256 was
-`19c23908e7d57e8d91ef85b41f5646307582f10f4f0fb999bbf89ed24ec9c983`
-(tag 9.1.2), built with `make -j4 valkey-server`. Prefetch was disabled using its
-existing configuration option, without modifying server source.
+- Standalone TCP on localhost; RDB snapshots and AOF disabled — no production durability
+ or cross-node network overhead is included.
+- MemoryEfficient engine; 50 ms version polling for measurement (the library default is 30
+ seconds); five-minute leases renewed every minute.
+- A separate Go process per size/distribution/repetition, two complete repetitions, on a
+ developer workstation with some development checks running on the host. Treat the ranges
+ as baseline observations, not isolated microbenchmark confidence intervals. Each server
+ ran one scale workload at a time.
+- Fixed seed 20260906. Shared: `shared-prefix-%08d`. Diverse: a random 12-bit hex prefix
+ plus a unique eight-digit index. Korean: `한국어-%04x-키워드-%08d` with a random 16-bit
+ prefix. Full replacement appends `-x` to every keyword.
+
+**Valkey qualification.** An exploratory run of the local Valkey 9.1.2 build with its
+default prefetch setting exited with SIGSEGV in `hashtableIncrementalFindStep` via
+`prefetchCommandQueueKeys`. The final successful Valkey matrix used
+`prefetch-batch-max-size 0`, so this report does **not** establish support for that build's
+default configuration. The stack excerpt is in
+`benchmarks/results/v3-20260906-valkey-crash.txt`; the library returned an error when the
+server disappeared, and the root cause has not been established. The Valkey source archive
+SHA-256 was `19c23908e7d57e8d91ef85b41f5646307582f10f4f0fb999bbf89ed24ec9c983` (tag
+9.1.2), built with `make -j4 valkey-server`. Prefetch was disabled through its existing
+configuration option, without modifying server source.
## Initial load and startup
-All table cells show the minimum–maximum across two runs. Load-ready is elapsed
-from Replace to the local engine serving its version. Startup opens a new
-collection instance against the existing dictionary and builds its first engine.
-Peak RSS is the process lifetime high-water mark across the entire run, including
-subsequent changes, rather than only initial load.
+Cells show minimum–maximum across two runs. Load-ready is elapsed from `Replace` to the
+local engine serving its version. Startup opens a new collection instance against the
+existing dictionary and builds its first engine. Peak RSS is the process lifetime
+high-water mark across the entire run, including subsequent changes.
| Server | Keywords | Distribution | Load-ready (s) | Startup (s) | Peak RSS (GiB) |
|---|---:|---|---:|---:|---:|
@@ -71,11 +70,11 @@ subsequent changes, rather than only initial load.
## Million-keyword changes
-Commit below includes normalization, bucket reads/preparation, manifest storage,
-and final pointer commit. Ready also includes notification/polling and engine
-rebuilding. Their difference is the calling instance's serving-version lag;
-fleet-wide simultaneous activation is not asserted. Unchanged buckets are reused,
-but every changed generation still rebuilds its entire local engine.
+Commit covers normalization, bucket reads and preparation, manifest storage, and the final
+pointer commit. Ready adds notification/polling and engine rebuilding; their difference is
+the calling instance's serving-version lag, and fleet-wide simultaneous activation is not
+asserted. Unchanged buckets are reused, but every changed generation still rebuilds its
+entire local engine.
| Server | Distribution | Add 1 commit / ready (s) | Add 1,000 ready (s) | Add 1% ready (s) | Full replace ready (s) |
|---|---|---:|---:|---:|---:|
@@ -86,19 +85,18 @@ but every changed generation still rebuilds its entire local engine.
| valkey | diverse | 0.005–0.007 / 2.33–2.40 | 2.42–2.60 | 3.03–3.54 | 4.08–4.33 |
| valkey | korean | 0.005–0.005 / 2.90–3.40 | 4.03–4.38 | 3.98–4.49 | 6.78–7.31 |
-The JSON report also includes removals of 1, 1,000 and 1%, and identical Replace.
-At 100,000 entries, 1,000 and 1% are the same count and are measured as two
-successive add/remove cycles. Identical Replace still reads and compares buckets;
-it does not create a generation or rebuild the engine.
+The JSON report also covers removals of 1, 1,000 and 1%, and identical `Replace`. At
+100,000 entries, 1,000 and 1% are the same count and are measured as two successive
+add/remove cycles. Identical `Replace` still reads and compares buckets; it creates no
+generation and rebuilds no engine.
-## Search during replacement and retained storage
+## Search during replacement, and retained storage
-Search sampling uses a short string assembled from three dictionary keywords, at
-approximately one-millisecond intervals while a write and its engine refresh run.
-The timer includes that small string construction and the Find call. The old
-engine continues serving until replacement; initial load starts with an empty
-engine. These figures are not a general text-search throughput benchmark. The
-following cells summarize full replacement at one million keywords.
+Search sampling uses a short string assembled from three dictionary keywords, taken at
+roughly one-millisecond intervals while a write and its engine refresh run. The timer
+includes that string construction and the `Find` call. The old engine keeps serving until
+replacement; initial load starts with an empty engine. These are not general text-search
+throughput figures. The cells summarize full replacement at one million keywords.
| Server | Distribution | Search p50 / p95 / p99 (µs) | Redis before prune (MiB) | Redis after prune (MiB) |
|---|---|---:|---:|---:|
@@ -109,49 +107,47 @@ following cells summarize full replacement at one million keywords.
| valkey | diverse | 2.50–2.71 / 3.96–5.46 / 12.33–14.38 | 66.82–66.82 | 23.25–23.26 |
| valkey | korean | 3.04–3.04 / 11.62–14.00 / 25.04–29.92 | 134.78–134.78 | 43.99–43.99 |
-Redis memory is server-wide `used_memory`, including server overhead, receipts,
-script cache, and retained manifests/chunks. The pruning measurement artificially
-ages inactive generation registry timestamps beyond 24 hours; the active
-generation remains protected. Actual retention is tested separately. Network byte
-counters in the JSON include both writes and engine downloads/polling, as well as
-protocol and measurement overhead; they are not just payload keyword bytes.
+Redis memory is server-wide `used_memory`, including server overhead, receipts, script
+cache, and retained manifests and chunks. The pruning measurement artificially ages
+inactive generation registry timestamps beyond 24 hours; the active generation stays
+protected, and actual retention is tested separately. Network byte counters in the JSON
+include writes, engine downloads, polling, protocol, and measurement overhead — not just
+payload keyword bytes.
## Acceptance evidence
- The 36 scale runs completed initial loading, reopening, all paginated entries,
- naive-reference search comparison, no-op/full replacements, 1/1,000/1% additions
- and removals, WaitForVersion and explicit pruning.
-- `TestVersionedMillionSafety` passed on both configured real servers: one of two
- writers using the same million-entry expected version won; the other conflicted.
- A pinned million-entry snapshot remained readable across commit and Prune,
- contained no competing-generation keywords, and was collected only after close.
-- Unit/race tests compare existing V2 engine search results with V3 for Korean,
- multilingual, nested/overlapping, case-sensitive and stream-boundary cases.
- Concurrent batch scans preserve a single engine generation.
-- Fault injection covers chunk preparation errors, abandoned prepared generations,
- lost commit responses and delayed retries, dropped Pub/Sub, failed engine loads,
- reconnect, lease expiry, and a stale pruning fence. These are deterministic
- simulations; they are not a production failover or forced-OS-process-kill test.
-- Chunk-write instrumentation confirms an identical dictionary writes no chunks
- and a single change writes only its affected bucket. The final Lua commit only
- checks fixed keys, writes a receipt/marker and swaps the pointer; it does not
- parse a manifest or iterate keywords.
-- Root and server race suites, root/server lint, all-module vet, benchmark-module
- tests, API snapshot/audit, documentation compilation, unchanged module manifests
- and license attribution passed. FuzzFind and FuzzAdd ran for 30 seconds each;
- V3 normalization fuzzing ran for 10 seconds. No production acor function remained
- at zero coverage in the normal test suite.
-
-No real Cluster/Sentinel failover matrix or production persistence recovery was
-run for V3 here. All collection keys use one hash tag and the existing connection
-implementations; do not interpret standalone measurements as multi-node results.
-R2 incremental download reuse and build-memory/cancellation improvements, and R3
-result limits/masking/replacement, remain follow-up roadmap work.
+ naive-reference search comparison, no-op and full replacements, 1/1,000/1% additions and
+ removals, `WaitForVersion`, and explicit pruning.
+- `TestVersionedMillionSafety` passed on both real servers: of two writers using the same
+ million-entry expected version, one won and the other conflicted. A pinned million-entry
+ snapshot stayed readable across commit and `Prune`, contained no competing-generation
+ keywords, and was collected only after close.
+- Unit and race tests compare V2 and V3 search results for Korean, multilingual,
+ nested/overlapping, case-sensitive, and stream-boundary cases. Concurrent batch scans
+ preserve a single engine generation.
+- Fault injection covers chunk preparation errors, abandoned prepared generations, lost
+ commit responses and delayed retries, dropped Pub/Sub, failed engine loads, reconnect,
+ lease expiry, and a stale pruning fence. These are deterministic simulations, not a
+ production failover or forced process-kill test.
+- Chunk-write instrumentation confirms an identical dictionary writes no chunks and a
+ single change writes only its affected bucket. The final Lua commit checks fixed keys,
+ writes a receipt and marker, and swaps the pointer — it parses no manifest and iterates
+ no keywords.
+- Root and server race suites, root/server lint, all-module vet, benchmark-module tests,
+ API snapshot/audit, documentation compilation, unchanged module manifests, and license
+ attribution passed. FuzzFind and FuzzAdd ran 30 seconds each; V3 normalization fuzzing
+ ran 10 seconds. No production `acor` function remained at zero coverage in the normal
+ test suite.
+
+No real Cluster/Sentinel failover matrix or production persistence recovery was run for V3
+here. All collection keys use one hash tag and the existing connection implementations —
+do not read standalone measurements as multi-node results.
## Reproduce
-Use a disposable server with the same persistence, preset and polling settings.
-From the repository root:
+Use a disposable server with the same persistence, preset, and polling settings. From the
+repository root:
```sh
ACOR_V3_SCALE_ADDR=127.0.0.1:6379 \
@@ -159,11 +155,10 @@ ACOR_V3_SCALE_OUTPUT=/tmp/acor-v3-redis \
ACOR_V3_SCALE_REPEATS=2 make bench-v3
```
-Repeat against Valkey with `prefetch-batch-max-size 0`, using a separate output
-directory. `scripts/benchmark-v3.sh` compiles one test binary and starts a fresh
-process for every size/distribution, then runs the million-entry safety scenario.
-It removes only its uniquely named collection keys, never FLUSHDB. Do not run
-another workload on the same endpoint during measurements: Redis memory/network
-counters are server-wide. The raw operation summaries, environment strings,
-startup/prune measurements and source hashes for this report are in
-`benchmarks/results/v3-20260906.json`.
+Repeat against Valkey with `prefetch-batch-max-size 0` and a separate output directory.
+`scripts/benchmark-v3.sh` compiles one test binary, starts a fresh process for every
+size/distribution, then runs the million-entry safety scenario. It removes only its own
+uniquely named collection keys, never `FLUSHDB`. Do not run another workload on the same
+endpoint during measurement — Redis memory and network counters are server-wide. Raw
+operation summaries, environment strings, startup/prune measurements, and source hashes
+are in `benchmarks/results/v3-20260906.json`.
diff --git a/docs/content/reference/versioned.md b/docs/content/reference/versioned.md
index 4bdf50e9..e2de4a1e 100644
--- a/docs/content/reference/versioned.md
+++ b/docs/content/reference/versioned.md
@@ -3,9 +3,11 @@ title: "Versioned dictionaries (V3)"
description: "Large dictionary snapshots, atomic changes, and background search refresh."
---
-V3 is an opt-in storage format in the existing Go module. Use `OpenVersioned`
-and a new collection name. Existing `Create`, V1/V2 keys, and their error and
-update contracts remain unchanged. V3 does not provide Suggest or a server API.
+# Versioned dictionaries (V3)
+
+V3 is an opt-in storage format in the same Go module. Open it with `OpenVersioned` and a
+**new** collection name. `Create`, the V1/V2 keys, and their error and update contracts
+are untouched. V3 has no `Suggest` and no server API.
```go
@@ -40,124 +42,136 @@ func main() {
}
```
-## API contracts
-
-All I/O methods take a context. `Status`, `Snapshot.Version`, and `Snapshot.Count`
-read local state. `Close` on the collection stops its background work; snapshot
-`Close(ctx)` releases its lease. Initial loading fails the open operation.
-
-`Snapshot.List(ctx, cursor, limit)` returns up to a positive limit, ordered by
-bucket number then lexical keyword order. Pass an empty cursor to start and stop
-when `NextCursor` is empty. Keep the same snapshot for a complete traversal:
-a cursor is bound to the generation. `Snapshot.Diff(ctx, target)` returns sorted
-Added, Removed, and Retained entries without writing. Diff materializes the
-comparison in memory. Count is the normalized, deduplicated keyword count.
-
-`Replace`, `Add`, `Remove`, `AddMany`, and `RemoveMany` require an expected Version.
-Treat versions as opaque equality tokens, never ordered numbers or strings.
-Concurrent writers using the same expected version cannot overwrite each other;
-use `errors.Is(err, acor.ErrConcurrencyConflict)`. Batch changes are atomic.
-Replace accepts an empty target. Reapplying an identical dictionary keeps the
-version and does not publish invalidation. Keywords are trimmed and, unless
-CaseSensitive was set at creation, lowercased. Blank keywords and invalid UTF-8
-fail the entire input. Duplicate normalized inputs are removed. The stored case
-policy must match every subsequent open.
-
-A successful write means the Redis commit completed. The calling instance and
-other instances may still search the previous engine. Use `WaitForVersion` to
-wait for that commit or a later one; it honors cancellation. Status exposes the
-observed active version, serving version, build start/duration, and recent error.
-A failed refresh preserves the old serving engine. Polling defaults to 30 seconds,
-with Pub/Sub accelerating discovery. A fixed 20 ms debounce window (RefreshDebounce) merges bursts before each
-background build. Incoming events do not extend the window. The worker finishes
-its current build, then loads the newest observed target. Every search, including a batch,
-parallel scan, or stream, uses one engine. V3 reuses existing rune positions,
-case folding, overlap and word-boundary semantics.
-
-`Find`, `FindIndex`, `FindMatches`, `FindSet`, `Contains`, `FindStream`,
-`FindBatch`, `FindParallel`, and `FindIndexParallel` are available with explicit
-contexts. Streaming keeps state across reader boundaries. The existing parallel
-options still apply; use AutoOverlap for dictionary-length boundary protection.
-
-## Storage and failure handling
+Every I/O method takes a context. `Status`, `Snapshot.Version`, and `Snapshot.Count` read
+local state only. `Close` on the collection stops its background work; `Close(ctx)` on a
+snapshot releases its lease. A failed initial load fails the open.
+
+## Reading
+
+`Snapshot.List(ctx, cursor, limit)` pages through keywords, ordered by bucket number then
+lexically. Start with an empty cursor and stop when `NextCursor` is empty. **Keep the same
+snapshot for a whole traversal** — a cursor is bound to its generation.
+
+`Snapshot.Diff(ctx, target)` returns sorted `Added`, `Removed`, and `Retained` without
+writing anything, materializing the comparison in memory. `Count` is the normalized,
+deduplicated keyword count.
+
+Search methods — `Find`, `FindIndex`, `FindMatches`, `FindSet`, `Contains`, `FindStream`,
+`FindBatch`, `FindParallel`, `FindIndexParallel` — all take explicit contexts. Streaming
+keeps state across reader boundaries, and the existing parallel options apply, including
+`AutoOverlap` for boundary protection. V3 reuses the V1/V2 rune positions, case folding,
+overlap, and word-boundary semantics. One search — batch, parallel, or streaming — always
+uses one engine.
+
+## Writing
+
+`Replace`, `Add`, `Remove`, `AddMany`, and `RemoveMany` all require an expected `Version`.
+
+- **Versions are opaque equality tokens.** Never treat them as ordered numbers or strings.
+- Concurrent writers holding the same expected version cannot overwrite each other; the
+ loser gets `errors.Is(err, acor.ErrConcurrencyConflict)`.
+- Batch changes are atomic. `Replace` accepts an empty target.
+- Reapplying an identical dictionary keeps the version and publishes no invalidation.
+- Keywords are trimmed, and lowercased unless `CaseSensitive` was set at creation.
+ Duplicates after normalization are dropped; a blank keyword or invalid UTF-8 fails the
+ entire input. The stored case policy must match every subsequent open.
+
+### Commit is not yet visibility
+
+A successful write means the Redis commit landed. The calling instance and every other
+instance may still be searching the previous engine. `WaitForVersion` waits for that
+commit or a later one and honors cancellation. `Status` exposes the observed active
+version, the serving version, build start and duration, and the most recent error.
+
+| Refresh behavior | |
+|---|---|
+| Discovery | Polling every 30 s by default, accelerated by Pub/Sub |
+| Coalescing | A fixed 20 ms debounce window (`RefreshDebounce`) merges bursts; incoming events do not extend it |
+| In-flight build | Finishes, then the worker loads the newest observed target |
+| Failure | The old serving engine is preserved |
+
+## Storage layout
SHA-256's first 12 bits select one of 4,096 fixed buckets. Each bucket is sorted,
-deduplicated, and split into JSON string-array chunks of at most 1 MiB including
-JSON escaping and delimiters. An oversized individual keyword gets its own chunk.
-Content-addressed immutable chunks, generation manifests, metadata, and the active
-pointer use separate keys. All keys share a name-digest Redis hash tag, hence one
-Cluster slot. This does not distribute one collection across multiple nodes.
-Connection options reuse existing standalone, Sentinel, Cluster and Ring clients.
-Only connection fields and Name from VersionedOptions.Redis are used.
-
-Delta writes download and prepare affected buckets only; full replacements compare
-all buckets and reuse unchanged ones. Local engine refresh reuses verified immutable keyword slices for unchanged
-buckets and downloads changed buckets only. The previous cache remains usable
-when a candidate fails. Full local engine rebuilding still requires memory for
-the old engine, retained keyword slices, and the new engine. The default
-preset is MemoryEfficient; Speed and Balanced are also selectable. No constant
-memory or unmeasured latency guarantee is made.
-
-The final Lua commit checks the expected pointer, preparation lease and maintenance
-lock, stores a receipt, and changes the pointer. It does not parse the manifest or
-iterate keywords. Prepared data is unreachable until committed. Storage durability
-and failover behavior still depend on the Redis/Valkey persistence configuration.
-
-On `ErrCommitUnknown`, retain the returned WriteResult.OperationID and use
-`ResolveOperation(ctx, id)`. A found receipt is authoritative success. A missing
-receipt is **not** proof that an in-flight request cannot commit later. Do not
-blindly reapply ambiguous writes. Operation receipts and committed-version markers
-are retained indefinitely, including no-op receipts, so they grow with write count.
-
-## Leases and explicit pruning
-
-Snapshots, builds and preparation writers use server-time leases (five minutes,
-renewed every minute by default). Lease checks fail explicitly after expiration;
-renewal cannot resurrect an expired handle. Close snapshots promptly. Existing
-local searches need no Redis lease once their engine is built.
-
-`Prune(ctx)` retains the active generation, generations prepared or committed
-within 24 hours, and generations protected by valid reader leases. It deletes
-unreferenced chunks and unretained manifests in batches of at most 64. It also
-collects abandoned preparation data. Receipts are not deleted.
-
-Prune acquires a monotonically increasing maintenance fence only when no valid
-writer lease remains. New writers and snapshots temporarily return ErrMaintenance;
-searches continue. Each deletion checks and extends lock ownership atomically.
-The maintenance lock expires after 30 seconds if its process dies. An expired
-pruner cannot delete after a successor takes ownership. Preparation itself is
-fenced so an expired writer cannot create unregistered data during cleanup.
-Large retention sets are enumerated in memory; there is no automatic collector.
-R2 builds check cancellation throughout trie insertion, failure-link construction,
-and table filling. Close cancels the active build; initialization honors its own
-context. Cancellation discards a candidate without publishing it or leaving a
-detached builder goroutine. Individual allocation, string length and sort
-operations complete before the next checkpoint. Ordinary new commits do not
-cancel an active build, avoiding starvation under continuous writes.
-
-## V2 cutover
-
-1. Use a **different** V3 name. Migrate V1 to V2 through the existing migration
- first. `CopyV2(ctx, sourceName, expected, nil)` reads the V2 version and keywords in
- one HMGET and replaces the V3 target. It reports source version, normalized
- count and SHA-256 checksum of the sorted JSON keyword array.
-2. For rehearsal, copy and compare count, checksum and representative search
- results, including case-sensitive, Korean and overlapping matches. Set the
- V3 case policy to the source application's policy (V2 does not persist it).
-3. Stop V2 writes for the final copy. Verify its source version, count, checksum
- and search results, then change application configuration to the new V3 name.
- There is no automatic dual write or simultaneous fleet engine switch.
-4. After V3 writes begin, switching configuration back to V2 loses those changes.
- Export and apply the V3 changes back to V2 before rollback; a simple name change
- is insufficient. Keep V2 untouched until the migration has been accepted.
+deduplicated, and split into JSON string-array chunks of at most 1 MiB including escaping
+and delimiters; an oversized single keyword gets its own chunk. Content-addressed
+immutable chunks, generation manifests, metadata, and the active pointer live under
+separate keys.
+
+**All keys share one name-digest hash tag, so a collection occupies one Cluster slot.**
+V3 does not distribute a collection across nodes. Connection options reuse the existing
+standalone, Sentinel, Cluster, and Ring clients; only the connection fields and `Name`
+from `VersionedOptions.Redis` are used.
+
+Delta writes download and prepare affected buckets only; a full replacement compares every
+bucket and reuses the unchanged ones. Local engine refresh reuses verified immutable
+keyword slices for unchanged buckets and downloads only changed ones, and the previous
+cache stays usable when a candidate fails. A full rebuild still needs memory for the old
+engine, the retained slices, and the new engine at once. The default preset is
+`MemoryEfficient`; `Speed` and `Balanced` are selectable. No constant-memory or latency
+guarantee is made.
+
+The final Lua commit checks the expected pointer, the preparation lease, and the
+maintenance lock, stores a receipt, and swaps the pointer. It does not parse the manifest
+or iterate keywords. Prepared data is unreachable until committed. Durability and failover
+still depend on your Redis/Valkey persistence configuration.
+
+### Ambiguous commits
+
+On `ErrCommitUnknown`, keep the returned `WriteResult.OperationID` and call
+`ResolveOperation(ctx, id)`. A found receipt is authoritative success. **A missing receipt
+is not proof that an in-flight request cannot still commit** — do not blindly reapply an
+ambiguous write. Receipts and committed-version markers are retained indefinitely,
+including no-op receipts, so they grow with write count.
+
+## Leases and pruning
+
+Snapshots, builds, and preparation writers hold server-time leases: five minutes, renewed
+every minute by default. Lease checks fail explicitly after expiration, and renewal cannot
+resurrect an expired handle — so close snapshots promptly. A local search needs no lease
+once its engine is built.
+
+`Prune(ctx)` retains the active generation, anything prepared or committed within 24
+hours, and anything protected by a valid reader lease. It deletes unreferenced chunks and
+unretained manifests in batches of at most 64, and collects abandoned preparation data.
+Receipts are never deleted, and there is no automatic collector; large retention sets are
+enumerated in memory.
+
+Pruning takes a monotonically increasing maintenance fence, and only when no valid writer
+lease remains. While it holds, new writers and snapshots get `ErrMaintenance` and searches
+continue. Every deletion checks and extends lock ownership atomically; the lock expires
+after 30 seconds if its process dies, and an expired pruner cannot delete once a successor
+takes ownership. Preparation is fenced too, so an expired writer cannot create
+unregistered data during cleanup.
+
+Builds check cancellation throughout trie insertion, failure-link construction, and table
+filling. `Close` cancels the active build; cancellation discards the candidate without
+publishing it or leaving a detached goroutine. Ordinary new commits do not cancel an
+active build, which is what stops continuous writes from starving refresh.
+
+## Cutting over from V2
+
+1. **Use a different name.** Migrate V1 to V2 through the existing migration first.
+ `CopyV2(ctx, sourceName, expected, nil)` reads the V2 version and keywords in one
+ `HMGET` and replaces the V3 target, reporting source version, normalized count, and the
+ SHA-256 checksum of the sorted JSON keyword array.
+2. **Rehearse.** Copy, then compare count, checksum, and representative search results —
+ case-sensitive, Korean, and overlapping matches. Set the V3 case policy to the source
+ application's policy; V2 does not persist it.
+3. **Cut over.** Stop V2 writes for the final copy, verify source version, count,
+ checksum, and search results, then point the application at the new V3 name. There is
+ no automatic dual write and no simultaneous fleet-wide engine switch.
+4. **Rollback is not a name change.** Once V3 writes begin, switching back to V2 loses
+ them. Export and apply the V3 changes to V2 first, and keep V2 untouched until the
+ migration is accepted.
## CLI
-Global connection/name flags precede `dictionary`; its flags follow the subcommand.
-Input to diff/replace is one JSON string array on stdin. Empty copy-v2 sources
-also require `--allow-empty`; the guard runs before replacing the target. Status and other results
-are JSON. List is one page per invocation; concurrent generation changes reject a
-previous page cursor. Library callers should retain a Snapshot for a long export.
+Global connection and name flags come **before** `dictionary`; the subcommand's own flags
+follow it. Diff and replace read one JSON string array from stdin. Empty replacements and
+empty `copy-v2` sources need `--allow-empty`, checked before the target is replaced.
+Results are JSON. `list` returns one page per invocation, and a concurrent generation
+change rejects an older page cursor — hold a `Snapshot` in library code for a long export.
```sh
acor -name filter-v3 dictionary status
@@ -169,9 +183,11 @@ acor -name filter-v3 dictionary copy-v2 --source filter-v2 --expected-version TO
acor -name filter-v3 dictionary prune
```
-R1 validation and measured resource costs are recorded in
-[the reproducible V3 performance report](../versioned-performance/). R2 adds incremental download reuse, bounded coalescing, cancellable builds, and
-more compact sparse-engine nodes. R3 adds bounded source-position search and
-atomic masking/replacement; see [bounded text processing](../text-processing/) and
-[the R2/R3 verification report](../r2-r3-performance/). R1 measurements remain
-archived as the baseline; none of these reports are performance guarantees.
+## Measurements
+
+[R1 validation and resource costs](../versioned-performance/) are the archived baseline.
+R2 added incremental download reuse, bounded coalescing, cancellable builds, and more
+compact sparse-engine nodes; R3 added bounded source-position search and atomic
+masking/replacement ([bounded text processing](../text-processing/),
+[R2/R3 verification](../r2-r3-performance/)). None of these reports is a performance
+guarantee.
diff --git a/docs/content/server/_index.md b/docs/content/server/_index.md
index 28aae984..7b88bbb4 100644
--- a/docs/content/server/_index.md
+++ b/docs/content/server/_index.md
@@ -18,27 +18,21 @@ the middleware that makes either observable.
> directive, so without a pin you get whichever core version the server module's `require`
> names.
-## ACOR ships no server binary
+## There is no server binary
-There is no `acor-server` to install, and no image to pull. `acor/server` is a library: it
-hands you an `http.Handler`, a `*grpc.Server`, and middleware, and **you** write the `main`
-that wires them to a collection and listens.
+No `acor-server` to install, no image to pull. `acor/server` is a library — it hands you an
+`http.Handler`, a `*grpc.Server`, and middleware, and **you** write the `main` that wires
+them to a collection and listens.
-That is a deliberate consequence of the module being experimental — a published binary is a
-contract, and this module does not offer one yet. What it offers instead is that the wiring
-is about forty lines, and [Running a Server](running/) is those forty lines, complete and
-copy-pasteable.
+That follows from the module being experimental: a published binary is a contract, and this
+module does not offer one yet. What it offers instead is that the wiring is about forty
+lines, and [Running a Server](running/) is those forty lines, complete and copy-pasteable.
## Sections
-- [Running a Server](running/) - Wire a collection to HTTP or gRPC, with readiness checks and clean shutdown
-- [HTTP API](http-api/) - The nine JSON endpoints, their request and response shapes, and every error they return
-- [gRPC API](grpc-api/) - The `acor.server.v1.Acor` service, its eight RPCs, and the observability constructors
+- [Running a Server](running/) — the `main` for each protocol, with readiness checks and clean shutdown
+- [HTTP API](http-api/) — nine JSON endpoints, their shapes, and every error they return
+- [gRPC API](grpc-api/) — the `acor.server.v1.Acor` service and the observability constructors
Metrics, structured logging, and tracing are configured the same way whichever protocol you
-serve, so they live together under
-[Operations → Monitoring](../operations/monitoring/).
-
-## Navigation
-
-← [Operations](../operations/) | [CLI](../cli/) →
+serve, so they live in [Operations → Monitoring](../operations/monitoring/).
diff --git a/docs/content/server/grpc-api.md b/docs/content/server/grpc-api.md
index 3d1a5ebe..b29994d4 100644
--- a/docs/content/server/grpc-api.md
+++ b/docs/content/server/grpc-api.md
@@ -7,13 +7,11 @@ weight: 3
The service is `acor.server.v1.Acor`, defined in
[`server/proto/acor/v1/acor.proto`](https://github.com/skyoo2003/acor/blob/main/server/proto/acor/v1/acor.proto).
-It mirrors the [HTTP API](../http-api/) one for one.
+It mirrors the [HTTP API](../http-api/) one for one. The `main` that serves it is
+[Running a Server](../running/).
-> **The `acor/server` module is experimental.** This service definition is **not covered by
-> the core module's compatibility promise** and can change in any release. See the
-> [section overview](../) for what that means for your `go.mod`.
-
-See [Running a Server](../running/) for the `main` that serves it.
+> The `acor/server` module is [experimental and separately versioned](../) — this service
+> definition can change in any release.
## Methods
@@ -28,93 +26,53 @@ See [Running a Server](../running/) for the `main` that serves it.
| `Info` | `EmptyRequest` | `InfoResponse{keywords, nodes}` |
| `Flush` | `EmptyRequest` | `StatusResponse{status}` |
-All eight are unary. Full method names are `/acor.server.v1.Acor/`.
-
-### Two shapes differ from HTTP
-
-**Counts are `int64`.** `CountResponse.count`, `InfoResponse.keywords`, and
-`InfoResponse.nodes` are `int64` on the wire, where the Go API and the JSON API use `int`.
+All eight are unary; full method names are `/acor.server.v1.Acor/`.
-**Match offsets are wrapped.** proto3 maps cannot hold a repeated value, so
-`MatchIndexesResponse.matches` is `map` rather than the JSON API's
-`map[string][]int`:
+Two shapes differ from HTTP:
-```proto
-message Positions {
- repeated int64 positions = 1;
-}
+- **Counts are `int64`.** `CountResponse.count`, `InfoResponse.keywords`, and
+ `InfoResponse.nodes` are `int64` on the wire, where the Go and JSON APIs use `int`.
+- **Match offsets are wrapped.** proto3 maps cannot hold a repeated value, so
+ `MatchIndexesResponse.matches` is `map`:
-message MatchIndexesResponse {
- map matches = 1;
-}
-```
+ ```proto
+ message Positions {
+ repeated int64 positions = 1;
+ }
+ ```
-A Go client reads `resp.GetMatches()["redis"].GetPositions()`.
+ A Go client reads `resp.GetMatches()["redis"].GetPositions()`.
-**Positions are rune offsets, not byte offsets**, and `SuggestIndex` always answers `[0]`
-because it matches the input as a prefix rather than searching for it. Both behaviors come
-from the collection, not the transport, so they are identical on both surfaces —
-[the HTTP page works through them with examples](../http-api/).
+Positions are rune offsets, and `SuggestIndex` always answers `[0]` because it matches the
+input as a prefix. Both behaviors come from the collection rather than the transport, so
+they are identical on both surfaces — [the HTTP page works through them with
+examples](../http-api/).
## Errors
Every error from the collection becomes `codes.Internal` with the error's own text as the
-message. This is the same flattening the HTTP API does with its blanket `500`: a caller
-mistake and a Redis outage arrive as the same code, and telling them apart means matching on
-message text, which is not part of any promise.
-
-No RPC returns `InvalidArgument`, `NotFound`, or `FailedPrecondition` — a write to a V1
-read-only collection is `Internal`, not `FailedPrecondition`.
+message — the same flattening as the HTTP blanket `500`. No RPC returns `InvalidArgument`,
+`NotFound`, or `FailedPrecondition`; a write to a read-only V1 collection is `Internal`.
+Telling a caller mistake from a Redis outage means matching on message text, which is not
+part of any promise.
## Deadlines do not cancel the work
-**A client deadline or a disconnect ends the RPC, not the Redis operation behind it.**
+**A client deadline or disconnect ends the RPC, not the Redis operation behind it.**
`Service` declares no context on any method (`server/grpc.go` calls
`s.service.Add(req.GetKeyword())`), and `(*AhoCorasick).Add` runs against the collection's
-own long-lived context, not the caller's. The request context is accepted by each adapter
-and discarded.
+own long-lived context. The request context is accepted by each adapter and discarded.
-The consequence to plan for: a client that gives up on `Add`, `Remove`, or `Flush` and sees
-`DeadlineExceeded` or `Canceled` has **not** prevented the write. It may already have
-landed, or land shortly after. `Flush` deletes every key in the collection, and abandoning
-the call does not call it off.
+So a client that gives up on `Add`, `Remove`, or `Flush` and sees `DeadlineExceeded` has
+**not** prevented the write:
- Do not treat a timeout as "it did not happen". Re-read with `Info` or `Find` before
retrying anything that mutates.
-- Client-side retries on timeout can double-apply. `Add` and `Remove` are idempotent by
- keyword, so a repeat is harmless there; a retried `Flush` is a second flush.
-- Setting a short deadline does not shed load on the server. The Redis work continues at
- full cost after the client is gone.
-
-The [HTTP API](../http-api/) behaves identically — same `Service` interface, same discard.
-
-## Generating a client
-
-There is no `buf.yaml` or `protoc` configuration in this repository. Two options:
-
-- **Go clients** import the generated package directly and skip codegen entirely:
-
- ```go
- import acorv1 "github.com/skyoo2003/acor/server/proto/acor/v1"
-
- client := acorv1.NewAcorClient(conn)
- ```
-
-- **Other languages** run their own `protoc`/`buf` against `acor.proto`. It has no imports
- beyond proto3 itself, so the single file is the whole input — but `--_out` alone
- generates only the **message** types. The service stub needs that language's gRPC plugin
- as a second output flag:
-
- ```sh
- # Python: --python_out gives acor_pb2.py only; the stub needs grpcio-tools.
- python -m grpc_tools.protoc -I. --python_out=. --grpc_python_out=. acor.proto
-
- # C++/Java/etc. take the plugin the same way:
- protoc -I. --cpp_out=. --grpc_out=. --plugin=protoc-gen-grpc=$(which grpc_cpp_plugin) acor.proto
- ```
+- Retries can double-apply. `Add` and `Remove` are idempotent by keyword; a retried `Flush`
+ is a second flush.
+- A short deadline sheds no server load.
- Omit the gRPC flag and you get the eight request/response messages with nothing able to
- call the eight RPCs.
+The [HTTP API](../http-api/) behaves identically.
## Constructors
@@ -124,52 +82,67 @@ server.NewGRPCServerWithObservability(ctx, service, obs, opts...) // + observabi
```
Both accept any `grpc.ServerOption`, including `grpc.Creds` for TLS, and both leave
-`Serve`/`Stop` to you.
-
-`Observability` bundles four pillars, and **any field may be nil to skip that one**:
+`Serve`/`Stop` to you. `Observability` bundles four pillars, and **any field may be nil to
+skip it**:
| Field | Wires in | Notes |
| ----- | -------- | ----- |
| `Tracer` | `otelgrpc` stats handler | Configure with `tracing.NewTracer` |
-| `Metrics` | `grpc_server_*` Prometheus interceptor | gRPC has no `/metrics` route — expose `promhttp.Handler()` on a separate HTTP listener |
+| `Metrics` | `grpc_server_*` Prometheus interceptor | gRPC has no `/metrics` route — expose `promhttp.Handler()` on a separate listener |
| `Logger` | zerolog unary interceptor | JSON request logs |
| `Health` | the standard `grpc.health.v1` service | See below |
-Metric names and tracing configuration live in
-[Operations → Monitoring](../../operations/monitoring/); this page does not restate them.
+Metric names and tracing configuration are in
+[Operations → Monitoring](../../operations/monitoring/).
## Health checking
Passing `Health` registers the standard `grpc.health.v1.Health` service, so
-`grpc_health_probe` and Kubernetes gRPC probes work without extra code. Behavior worth
-knowing before you set a probe interval:
+`grpc_health_probe` and Kubernetes gRPC probes work with no extra code. Four things to know
+before setting a probe interval:
-- **Status is polled, not computed per request.** An HTTP `/readyz` probe runs your checkers
- on each request; a gRPC health probe reads whatever the last poll left behind. Probing
- more often than every 5 seconds buys no extra resolution.
+- **Status is polled, not computed per request.** A gRPC probe reads whatever the last poll
+ left behind, so probing more often than every 5 seconds buys no resolution.
- **The 5-second tick bounds neither staleness nor load.** The poller calls
- `checker.Check()` inline on a single long-lived `time.Ticker`. That ticker keeps ticking
- while a check is blocked, and its channel buffers one tick, so a check that overruns the
- interval is followed *immediately* by the next one rather than after a fresh 5-second
- gap. A checker that takes 30 seconds therefore makes the status at least 30 seconds stale
- **and** polls back-to-back for as long as it stays slow — the opposite of the breather
- the interval suggests, and it lands on Redis hardest when Redis is already the slow part.
- A checker that hangs outright is worse: the blocked goroutine cannot act on context
- cancellation either, so shutdown stops marking the server `NOT_SERVING`. Give every
- checker its own deadline; the examples in [Running a Server](../running/) use a 2-second
- `context.WithTimeout` for exactly this.
-- **Both the overall server (`""`) and `acor.server.v1.Acor` are tracked**, and they carry
- the same status.
-- **A nil `Health` field means no health service at all.** The constructor only registers
- `grpc.health.v1` when the field is set (`server/grpc.go:77`), so leaving it nil makes a
- probe fail with `UNIMPLEMENTED`, not `SERVING`. To get a health service that always
- answers `SERVING`, pass an empty `health.NewChecker()` rather than nil — the
- nil-*checker*-means-`SERVING` rule lives inside `RegisterGRPCHealthServer` and is only
- reachable by calling that function yourself.
-- **The `ctx` you pass to `NewGRPCServerWithObservability` bounds the poller.** Cancel it
- and the server is marked `NOT_SERVING`, which is what lets a load balancer drain
- connections before `GracefulStop` finishes.
-
-## Navigation
-
-← [HTTP API](../http-api/) | [CLI](../../cli/) →
+ `checker.Check()` inline on one long-lived `time.Ticker`, which keeps ticking while a
+ check is blocked and buffers one tick — so a check that overruns the interval is followed
+ *immediately* by the next. A checker taking 30 seconds makes the status at least 30
+ seconds stale **and** polls back-to-back for as long as it stays slow, hitting Redis
+ hardest when Redis is already the slow part. A checker that hangs outright cannot act on
+ context cancellation either, so shutdown stops marking the server `NOT_SERVING`. Give
+ every checker its own deadline — [Running a Server](../running/) uses 2 seconds.
+- **A nil `Health` field means no health service at all**, so a probe fails with
+ `UNIMPLEMENTED`, not `SERVING` (`server/grpc.go:77`). For a service that always answers
+ `SERVING`, pass an empty `health.NewChecker()`; the nil-*checker*-means-`SERVING` rule
+ lives inside `RegisterGRPCHealthServer` and is reachable only by calling that yourself.
+- **The `ctx` passed to `NewGRPCServerWithObservability` bounds the poller.** Cancelling it
+ marks the server `NOT_SERVING`, which is what lets a load balancer drain connections
+ before `GracefulStop` finishes.
+
+Both the overall server (`""`) and `acor.server.v1.Acor` are tracked, and they carry the
+same status.
+
+## Generating a client
+
+There is no `buf.yaml` or `protoc` configuration in this repository.
+
+- **Go clients** import the generated package and skip codegen entirely:
+
+ ```go
+ import acorv1 "github.com/skyoo2003/acor/server/proto/acor/v1"
+
+ client := acorv1.NewAcorClient(conn)
+ ```
+
+- **Other languages** run their own `protoc`/`buf` against `acor.proto`. It has no imports
+ beyond proto3, so the single file is the whole input — but `--_out` alone generates
+ only the **message** types. The service stub needs that language's gRPC plugin as a
+ second output flag:
+
+ ```sh
+ # Python: --python_out gives acor_pb2.py only; the stub needs grpcio-tools.
+ python -m grpc_tools.protoc -I. --python_out=. --grpc_python_out=. acor.proto
+
+ # C++/Java/etc. take the plugin the same way:
+ protoc -I. --cpp_out=. --grpc_out=. --plugin=protoc-gen-grpc=$(which grpc_cpp_plugin) acor.proto
+ ```
diff --git a/docs/content/server/http-api.md b/docs/content/server/http-api.md
index 58bd3b5b..10a2851e 100644
--- a/docs/content/server/http-api.md
+++ b/docs/content/server/http-api.md
@@ -6,15 +6,11 @@ weight: 2
# HTTP API
`server.NewHTTPHandler(service)` returns an `http.Handler` serving nine routes. Every
-response the handler itself produces is JSON with `Content-Type: application/json`; the
-exceptions are the two `ServeMux`-level responses noted under
-[Not every response is JSON](#not-every-response-is-json).
+response the handler itself produces is JSON; the two exceptions come from `ServeMux`
+below. The `main` that mounts it is [Running a Server](../running/).
-> **The `acor/server` module is experimental.** These paths and shapes are **not covered by
-> the core module's compatibility promise** and can change in any release. See the
-> [section overview](../) for what that means for your `go.mod`.
-
-See [Running a Server](../running/) for the `main` that mounts this handler.
+> The `acor/server` module is [experimental and separately versioned](../) — these paths
+> and shapes can change in any release.
## Endpoints
@@ -26,148 +22,108 @@ See [Running a Server](../running/) for the `main` that mounts this handler.
| `POST` | `/v1/find` | `{"input":"..."}` | `{"matches":["..."]}` |
| `POST` | `/v1/find-index` | `{"input":"..."}` | `{"matches":{"kw":[0,12]}}` |
| `POST` | `/v1/suggest` | `{"input":"..."}` | `{"matches":["..."]}` |
-| `POST` | `/v1/suggest-index` | `{"input":"..."}` | `{"matches":{"kw":[0]}}` — always `[0]`, see below |
+| `POST` | `/v1/suggest-index` | `{"input":"..."}` | `{"matches":{"kw":[0]}}` — always `[0]` |
| `GET` | `/v1/info` | — | `{"keywords":3,"nodes":7}` |
| `POST` | `/v1/flush` | — | `{"status":"ok"}` |
`count` is how many keywords the operation actually changed, so a second `add` of the same
-keyword answers `{"count":0}`.
+keyword answers `{"count":0}`. `/v1/flush` takes no body, does not read one if you send it,
+and deletes every key in the collection. The method column is enforced: anything else is
+`405`.
+
+```sh
+curl -sX POST localhost:8080/v1/add -d '{"keyword":"redis"}'
+# {"count":1}
-### Offsets are rune offsets, and the two `*-index` routes do not mean the same thing
+curl -sX POST localhost:8080/v1/find -d '{"input":"redis-backed matching"}'
+# {"matches":["redis"]}
-`/v1/find-index` returns, per keyword, the positions where it matched. Those positions are
-counted in **runes, not bytes**:
+curl -sX POST localhost:8080/v1/find-index -d '{"input":"redis and redis"}'
+# {"matches":{"redis":[0,10]}}
-```sh
-curl -sX POST localhost:8080/v1/find-index -d '{"input":"한글 레디스 and redis"}'
-# {"matches":{"redis":[11],"레디스":[3]}}
+curl -s localhost:8080/v1/info
+# {"keywords":1,"nodes":6}
```
-`레디스` starts at rune 3 but byte 7; `redis` starts at rune 11 but byte 21. The response
-gives 3 and 11. A client that slices the original string by these numbers must count
-runes — Go `[]rune`, Python `str` indices, Rust `chars()`. JavaScript needs care in both
-directions, because its string indices are UTF-16 code units: `[...s]` gives code points,
-and anything outside the BMP counts as two units under `.slice()`.
+### The two `*-index` routes do not mean the same thing
-`/v1/suggest-index` looks like the same shape but is not. `Suggest` matches the input as a
-*prefix* of each keyword, so every match starts at the beginning by construction and the
-implementation assigns exactly `[0]` to every suggestion. It never reports a second
-position:
+`/v1/find-index` returns, per keyword, the positions where it matched — counted in
+**runes, not bytes**:
```sh
-curl -sX POST localhost:8080/v1/suggest-index -d '{"input":"red"}'
-# {"matches":{"redis":[0]}}
+curl -sX POST localhost:8080/v1/find-index -d '{"input":"한글 레디스 and redis"}'
+# {"matches":{"redis":[11],"레디스":[3]}}
```
-Treat it as "these keywords start with your input", not as a position list.
-
-`/v1/flush` takes no request body and does not read one if you send it. It deletes every
-key in the collection.
+`레디스` starts at rune 3 but byte 7; `redis` at rune 11 but byte 21. A client slicing the
+original string by these numbers must count runes — Go `[]rune`, Python `str` indices,
+Rust `chars()`. JavaScript needs care in both directions, since its string indices are
+UTF-16 code units: `[...s]` gives code points, and anything outside the BMP counts as two
+units under `.slice()`.
-The method column is enforced, not advisory: `/healthz` and `/v1/info` are `GET`-only and
-everything else is `POST`-only. Any other method gets `405`.
+`/v1/suggest-index` looks like the same shape but is not. `Suggest` matches the input as a
+*prefix*, so every match starts at the beginning by construction and the implementation
+assigns exactly `[0]` to every suggestion. Read it as "these keywords start with your
+input", not as a position list.
## Errors
-Every failure returns `{"error":""}` with `Content-Type: application/json`.
+Every failure the handler produces is `{"error":""}` with
+`Content-Type: application/json`.
| Status | When | Body |
| ------ | ---- | ---- |
-| `400` | The body is not valid JSON | `{"error":"unexpected EOF"}` — the message comes from `encoding/json` and is not a stable string |
-| `400` | The body holds more than one JSON value | `{"error":"request body must contain only a single JSON value"}` |
+| `400` | Body is not valid JSON | `{"error":"unexpected EOF"}` — from `encoding/json`, not a stable string |
+| `400` | Body holds more than one JSON value | `{"error":"request body must contain only a single JSON value"}` |
| `405` | Wrong method for the path | `{"error":"method not allowed"}` |
| `413` | Reading the body reaches the 1 MiB cap | `{"error":"request body must not be larger than 1048576 bytes"}` |
| `500` | Any error from the underlying collection | `{"error":""}` |
-| `404` | No such path | **`text/plain`**, body `404 page not found` |
-| `301` | The path needs canonicalizing (`/v1//info`) | **`text/html`**, Go's `Moved Permanently` page |
-
-The body is read as it is decoded, so whichever fault surfaces first is the one you get. A
-body over 1 MiB that is *also* malformed comes back as the `400`, because the decoder
-reaches the bad byte before the reader reaches the cap. `413` means the cap is what stopped
-it, not that every oversized body is reported that way — nothing short of reading the whole
-body could promise the latter, and reading it is what the cap exists to avoid.
-
-Two of these deserve more than a table row.
-
-### Every collection error is a `500`
-
-There is no error taxonomy. The handler passes any non-nil error from the collection
-straight to a `500`, so a client mistake and a Redis outage are indistinguishable by status
-code. Writing to a V1 collection — which the core rejects with `ErrV1ReadOnly`, a caller
-error — comes back as `500 {"error":"V1 collections are read-only; migrate with MigrateV1ToV2"}`, not as a `4xx`.
-
-If your client needs to tell "retry this" from "fix your request", it has to match on the
-message text, and message text is not part of any promise. Treat `5xx` here as
-*something went wrong*, not as *the server is unhealthy*.
-
-For the latter, use `/readyz` — but note that **`/readyz` is not part of this handler**.
-`NewHTTPHandler` and `NewHTTPServer` register only the routes in the table above, so
-requesting `/readyz` from either gets the `404`. The route exists only if you mount
-`health.RegisterHTTPHandlers` on an outer mux yourself, as
-[Running a Server](../running/) does.
-
-### Not every response is JSON
-
-Two cases are answered by Go's `http.ServeMux` before ACOR's handler sees them, and neither
-is JSON:
-
-- **`404 page not found`** — `text/plain`, for a path that matches nothing.
-- **`301 Moved Permanently`** — `text/html`, for a path that needs canonicalizing.
- `/v1//info`, `/v1/./info`, and `/v1/../v1/info` all redirect to `/v1/info`.
-
-A client that unconditionally parses responses as JSON breaks on both. The redirect is the
-worse of the two: a client that follows redirects automatically may downgrade the `POST`
-to a `GET` and then receive a `405`, turning a doubled slash into a confusing method error.
-Normalize paths before sending.
+| `404` | No such path | **`text/plain`**, `404 page not found` |
+| `301` | Path needs canonicalizing (`/v1//info`) | **`text/html`**, Go's `Moved Permanently` page |
+
+The body is read as it is decoded, so whichever fault surfaces first is the one you get: an
+oversized body that is *also* malformed comes back as the `400`, because the decoder
+reaches the bad byte before the reader reaches the cap.
+
+**Every collection error is a `500`.** There is no error taxonomy — a client mistake and a
+Redis outage are indistinguishable by status code. Writing to a V1 collection, which the
+core rejects with `ErrV1ReadOnly`, arrives as
+`500 {"error":"V1 collections are read-only; migrate with MigrateV1ToV2"}`, not a `4xx`.
+Telling "retry this" from "fix your request" means matching on message text, and message
+text is not part of any promise.
+
+**`/readyz` is not part of this handler.** `NewHTTPHandler` and `NewHTTPServer` register
+only the routes above, so `/readyz` returns the `404`. It exists only if you mount
+`health.RegisterHTTPHandlers` on an outer mux, as [Running a Server](../running/) does.
+
+**Not every response is JSON.** The `404` and the `301` are answered by Go's
+`http.ServeMux` before ACOR's handler sees them. `/v1//info`, `/v1/./info`, and
+`/v1/../v1/info` all redirect to `/v1/info` — and a client that follows redirects
+automatically may downgrade the `POST` to a `GET` and then receive a `405`, turning a
+doubled slash into a confusing method error. Normalize paths before sending.
## Disconnecting does not cancel the work
-**A client timeout or a dropped connection ends the request, not the Redis operation behind
+**A client timeout or dropped connection ends the request, not the Redis operation behind
it.** Each handler passes `r.Context()` into its adapter and the adapter discards it
-(`server/server.go:96` and its siblings take `_ context.Context`); `Service` declares no
+(`server/server.go:96` and siblings take `_ context.Context`); `Service` declares no
context on any method, and `(*AhoCorasick).Add` runs against the collection's own
-long-lived context rather than the caller's.
+long-lived context.
So a client that gives up on `/v1/add`, `/v1/remove`, or `/v1/flush` has **not** prevented
-the write. It may already have landed, or land shortly after. `/v1/flush` deletes every key
-in the collection, and hanging up does not call it off.
+the write:
- Do not read a timeout as "it did not happen". Re-read with `/v1/info` or `/v1/find`
before retrying anything that mutates.
-- Retrying on timeout can double-apply. `add` and `remove` are idempotent by keyword, so a
- repeat is harmless; a retried `flush` is a second flush.
-- A short client timeout sheds no server load. The Redis work continues at full cost after
- the client is gone.
+- Retrying can double-apply. `add` and `remove` are idempotent by keyword; a retried
+ `flush` is a second flush.
+- A short client timeout sheds no server load — the Redis work continues at full cost.
-The [gRPC API](../grpc-api/) behaves identically — same `Service` interface, same discard.
+The [gRPC API](../grpc-api/) behaves identically: same `Service` interface, same discard.
## What the server does not check
-- **Request `Content-Type` is ignored.** Any value is accepted as long as the body parses
- as JSON.
+- **Request `Content-Type` is ignored** — any value is accepted if the body parses as JSON.
- **Unknown JSON fields are ignored.** `{"keyword":"x","bogus":1}` succeeds.
-- **Missing fields become the zero value.** `POST /v1/add` with `{}` is an `Add("")`; the
- server does not reject it, and what happens next is whatever the collection does with an
- empty keyword.
-- **There is no authentication or authorization.** `/v1/flush` is reachable by anyone who
- can reach the port.
-
-## Example
-
-```sh
-curl -sX POST localhost:8080/v1/add -d '{"keyword":"redis"}'
-# {"count":1}
-
-curl -sX POST localhost:8080/v1/find -d '{"input":"redis-backed matching"}'
-# {"matches":["redis"]}
-
-curl -sX POST localhost:8080/v1/find-index -d '{"input":"redis and redis"}'
-# {"matches":{"redis":[0,10]}}
-
-curl -s localhost:8080/v1/info
-# {"keywords":1,"nodes":6}
-```
-
-## Navigation
-
-← [Running a Server](../running/) | [gRPC API](../grpc-api/) →
+- **Missing fields become the zero value.** `POST /v1/add` with `{}` is an `Add("")`.
+- **There is no authentication.** `/v1/flush` is reachable by anyone who can reach the port.
diff --git a/docs/content/server/running.md b/docs/content/server/running.md
index 88f1a07e..bc4b4a4e 100644
--- a/docs/content/server/running.md
+++ b/docs/content/server/running.md
@@ -6,29 +6,25 @@ weight: 1
# Running a Server
ACOR ships no server binary. `acor/server` gives you an `http.Handler` and a
-`*grpc.Server`; the `main` that wires them to a collection and listens is yours to write.
-This page is that `main`, in full, for each protocol.
+`*grpc.Server`; the `main` that wires them to a collection and listens is yours. This page
+is that `main`, in full, for each protocol.
-> **The `acor/server` module is experimental.** It publishes no version tags of its own and
-> is **not covered by the core module's compatibility promise**. See the
-> [section overview](../) for what that means for your `go.mod`.
-
-## Dependencies
+> The `acor/server` module is [experimental and separately versioned](../).
```sh
go get github.com/skyoo2003/acor/server
go get github.com/skyoo2003/acor/pkg/acor@latest
```
-Both lines matter. `acor/server` resolves to a pseudo-version from `main`, and it carries a
+Both lines matter. `acor/server` resolves to a pseudo-version from `main` and carries a
`require` on the core module that Go will not override from the dependency's own `replace`
-directive — so name the core version yourself, in your own `go.mod`.
+directive — so name the core version yourself.
## HTTP
-The collection is the service. Every method `server.Service` requires — `Add`, `Remove`,
+The collection *is* the service. Every method `server.Service` requires — `Add`, `Remove`,
`Find`, `FindIndex`, `Suggest`, `SuggestIndex`, `Flush`, `Info` — is already a method on
-`*acor.AhoCorasick`, so the collection satisfies the interface with no adapter.
+`*acor.AhoCorasick`, so no adapter is needed.
```go
@@ -51,12 +47,11 @@ import (
// redisChecker reports whether the collection can still reach Redis.
//
-// Info() is the only exported call that proves the Redis path works, but it is
-// not free: on V2 it HGETALLs the trie hash and unmarshals the whole keyword
-// and prefix arrays just to count them, so its cost grows with the dictionary.
-// See "Readiness costs what Info() costs" below before pointing a probe at it.
+// Info() is the only exported call that proves the Redis path works, and it is
+// not free: on V2 it HGETALLs the trie hash and unmarshals the whole keyword and
+// prefix arrays just to count them. See "Readiness costs what Info() costs".
//
-// The timeout is not tidiness. Check() runs inline in both probe paths, so a
+// The timeout is load-bearing: Check() runs inline in both probe paths, so a
// checker that blocks blocks the prober.
type redisChecker struct{ ac *acor.AhoCorasick }
@@ -122,13 +117,12 @@ func main() {
`server.NewHTTPServer(addr, service)` is a one-liner that does most of this, and it is the
right choice if you do not need readiness checks. What it cannot give you is `/readyz`:
-`NewHTTPHandler` builds its own private `http.ServeMux` internally and returns it as an
-`http.Handler`, so there is no mux for you to register anything else on.
+`NewHTTPHandler` builds its own private `ServeMux` internally, so there is nothing for you
+to register on.
-Composing them on an outer mux, as above, works and does not panic — the two `/healthz`
-registrations are on different muxes and never meet. On the outer mux, Go routes `/healthz`
-to the `health` package's handler because an exact pattern outranks the `/` catch-all. The
-practical effect:
+Composing them on an outer mux works and does not panic — the two `/healthz` registrations
+live on different muxes and never meet, and on the outer mux Go routes `/healthz` to the
+`health` package because an exact pattern outranks the `/` catch-all:
| Path | Served by | Body |
| ---- | --------- | ---- |
@@ -137,8 +131,8 @@ practical effect:
| `/v1/*` | `server.NewHTTPHandler` | see [HTTP API](../http-api/) |
The two `/healthz` implementations return the same body for a `GET`, so the shadowing costs
-you nothing. They differ only in how they reject a non-`GET`: `server/health` replies in
-`text/plain` via `http.Error`, the API's own replies in JSON.
+nothing. They differ only in rejecting a non-`GET`: `server/health` replies `text/plain`
+via `http.Error`, the API's own replies in JSON.
## gRPC
@@ -162,9 +156,9 @@ import (
"github.com/skyoo2003/acor/server/metrics"
)
-// The deadline matters more here than on the HTTP side: the gRPC health
-// poller calls Check() inline on its ticker, so a checker that blocks stalls
-// every later poll and the poller's own response to cancellation.
+// The deadline matters more here than over HTTP: the gRPC health poller calls
+// Check() inline on its ticker, so a checker that blocks stalls every later poll
+// and the poller's own response to cancellation.
type redisChecker struct{ ac *acor.AhoCorasick }
func (c redisChecker) Check() health.CheckResult {
@@ -232,63 +226,48 @@ func main() {
}
```
-Every field of `Observability` is optional: a nil field skips that pillar, so you can start
-with logging only and add the rest later. `server.NewGRPCServer(service, opts...)` is the
-same server with no observability at all.
+Every field of `Observability` is optional — a nil field skips that pillar, so you can
+start with logging only. `server.NewGRPCServer(service, opts...)` is the same server with
+none of it.
-Prometheus metrics registered here are *collected* but not *exposed* — gRPC has no
-`/metrics` endpoint. Serve `promhttp.Handler()` on a separate HTTP listener; see
-[Operations → Monitoring](../../operations/monitoring/).
+Prometheus metrics registered here are *collected* but not *exposed*: gRPC has no
+`/metrics` endpoint. Serve `promhttp.Handler()` on a separate HTTP listener, per
+[Monitoring](../../operations/monitoring/).
## Readiness costs what `Info()` costs
`Info()` is the only exported call that touches Redis and returns quickly *on a small
-collection*, which is why it is the readiness check here. It is not a ping. On the default
-V2 schema it runs `HGETALL` against the trie hash and JSON-unmarshals the complete keyword
-and prefix arrays in order to return their two lengths, so its cost and its allocations
-scale with the whole dictionary.
+collection*, which is why it is the readiness check here. It is not a ping: on V2 it runs
+`HGETALL` against the trie hash and unmarshals the complete keyword and prefix arrays to
+return their two lengths, so cost and allocations scale with the whole dictionary.
Two things multiply that:
- **`/readyz` runs the checkers on every request**, and it is unauthenticated.
- **The gRPC health poller runs them every 5 seconds**, whether or not anyone is probing —
- and *more* often than that once a check gets slow. The poller's `time.Ticker` keeps
- ticking while `Check` is blocked and buffers one tick, so an overrunning check is followed
- immediately by the next. Slow checks are not throttled; they compound.
+ and *more* often once a check gets slow. Its `time.Ticker` keeps ticking while `Check`
+ is blocked and buffers one tick, so an overrunning check is followed immediately by the
+ next. Slow checks are not throttled; they compound.
On a large dictionary that is significant Redis traffic and garbage, generated hardest
-exactly when the service is already struggling. If that describes your collection, make the
-**readiness** check a direct `redis.Client.Ping` against the same address instead — that is
-a `PING`, not a dictionary scan — and accept that it proves connectivity rather than
-collection health.
-
-Keep it out of liveness either way. `/healthz` answers "is this process alive", and nothing
-about Redis belongs in that answer: a Redis outage would fail every replica's liveness probe
-at once and have the orchestrator restart all of them, which cannot repair Redis and drops
-whatever the processes were still serving. Redis reachability is a readiness signal —
-take me out of the load balancer — not a restart signal.
-
-The 2-second timeout is load-bearing either way. `HealthChecker.Check` calls every
-registered checker inline, and the gRPC poller calls `Check` inline on its ticker, so one
-checker that hangs stalls every later poll **and** the poller's response to context
-cancellation. A checker without its own deadline turns a slow Redis into a stuck health
-service.
+exactly when the service is already struggling. If that describes your collection, make
+the **readiness** check a direct `redis.Client.Ping` against the same address — a `PING`,
+not a dictionary scan — and accept that it proves connectivity rather than collection
+health.
-## What you still have to decide
+Keep it out of liveness either way. `/healthz` answers "is this process alive", and a
+Redis outage that failed every replica's liveness probe would restart all of them and
+repair nothing. Redis reachability is a readiness signal.
-The examples above hard-code answers this page cannot make for you:
+## What you still have to decide
1. **Listen address.** `:8080` and `:9090` are placeholders.
-2. **Redis credentials and topology.** `Addr` is standalone. Sentinel and Cluster use
- `Addrs`, and Sentinel also needs `MasterName` — see
- [Operations → Deployment](../../operations/deployment/).
-3. **TLS.** Neither constructor configures it. For gRPC, pass `grpc.Creds(...)` as a
- `grpc.ServerOption`; for HTTP, use `ListenAndServeTLS` or terminate at your ingress.
+2. **Redis credentials and topology.** `Addr` is standalone; Sentinel and Cluster use
+ `Addrs`, and Sentinel also needs `MasterName`. See
+ [Deployment](../../operations/deployment/).
+3. **TLS.** Neither constructor configures it. For gRPC pass `grpc.Creds(...)`; for HTTP
+ use `ListenAndServeTLS` or terminate at your ingress.
4. **Authentication.** There is none. Both surfaces expose `/v1/flush` and `Flush`, which
delete every key in the collection. Do not put either on a network you do not control.
-5. **Which protocol to serve.** They are independent; run one, the other, or both on
+5. **Which protocol to serve.** They are independent — run one, the other, or both on
separate listeners.
-
-## Navigation
-
-← [Server](../) | [HTTP API](../http-api/) →