diff --git a/README.md b/README.md index ace8859d..82a1eb05 100644 --- a/README.md +++ b/README.md @@ -9,32 +9,32 @@ [![License](https://img.shields.io/github/license/skyoo2003/acor.svg)](LICENSE) [![Sponsor](https://img.shields.io/badge/sponsor-GitHub-pink)](https://github.com/sponsors/skyoo2003) -ACOR stores a shared Aho-Corasick pattern dictionary in Redis and exposes it -through a Go library and CLI. Multiple application instances can update the -same dictionary at runtime while preset engines serve matches from local memory. +ACOR keeps one Aho-Corasick dictionary in Redis and reaches it from a Go library, a CLI, +and an experimental server module. Every application instance shares that dictionary, +updates it at runtime, and matches against a local copy with no Redis I/O on the hot path. -Typical uses include content filtering, keyword extraction, intrusion detection, -search highlighting, and real-time text classification. +Typical uses: content filtering, keyword extraction, intrusion detection, search +highlighting, real-time text classification. ## Highlights -- **Shared state** — every application instance uses the same Redis-backed dictionary +- **Shared state** — every instance reads the same Redis-backed dictionary - **Runtime updates** — Pub/Sub invalidation, with optional polling for missed messages -- **Fast reads** — preset engines match locally without Redis I/O on the hot path -- **Flexible deployment** — standalone Redis, Sentinel, Cluster, Ring, and Valkey -- **Complete matching API** — occurrences, positions, sets, streams, batches, and parallel matching +- **Fast reads** — preset engines match locally, 0 round trips +- **Topologies** — standalone Redis, Sentinel, Cluster, Ring, and Valkey +- **Full matching API** — occurrences, positions, sets, streams, batches, parallel scans -## Installation +## Install -ACOR requires Go 1.25 or newer and Redis 3.0 or newer, or Valkey 7.2 or newer. +Requires Go 1.25+ and Redis 3.0+ or Valkey 7.2+. ```sh go get github.com/skyoo2003/acor/pkg/acor@latest ``` -## Quick Start +## Quick start -Start Redis locally, then create a matcher: +Start Redis locally, then: ```go @@ -69,10 +69,9 @@ func main() { } ``` -New collections use the optimized V2 Redis schema by default. `PresetBalanced` -is the recommended starting point when reads should run locally. +New collections use the optimized V2 Redis schema. -## Choosing a Preset +## Choosing a preset | Goal | Preset | | ---- | ------ | @@ -80,18 +79,25 @@ is the recommended starting point when reads should run locally. | Highest matching throughput | `PresetSpeed` | | Lowest memory usage | `PresetMemoryEfficient` | -Redis remains the source of truth for every preset. See the -[preset guide](docs/content/guides/preset-engine.md) for trade-offs and the -[Redis-backed engine guide](docs/content/guides/redis-backed-engine.md) for -multi-instance invalidation safety. +Redis stays the source of truth for every preset. Trade-offs are in the +[preset guide](docs/content/guides/preset-engine.md); multi-instance invalidation is in the +[Redis-backed engine guide](docs/content/guides/redis-backed-engine.md). + +## Large dictionaries (V3) + +`OpenVersioned` opens a separate V3 collection with leased snapshots, expected-version +writes, and background engine replacement. V1/V2 APIs are unaffected. See the +[V3 guide](docs/content/reference/versioned.md), its +[performance report](docs/content/reference/versioned-performance.md), and +[bounded text processing](docs/content/reference/text-processing.md) for `Scan`, +`MaskText`, and `ReplaceText`. ## Documentation -Everything past this point lives on the -[documentation site](https://skyoo2003.github.io/acor/): Redis topologies, the matching and -streaming API, batch and parallel guides, the schema-V2 migration and benchmarks, deployment -and troubleshooting, the `acor` CLI, the experimental server module, and what the `v1` line -promises. Package signatures are on +The [documentation site](https://skyoo2003.github.io/acor/) covers Redis topologies, the +matching and streaming API, batch and parallel guides, the V2 schema and benchmarks, +deployment and troubleshooting, the `acor` CLI, the experimental server module, and what +the `v1` line promises. Package signatures are on [pkg.go.dev](https://pkg.go.dev/github.com/skyoo2003/acor/pkg/acor). ## Project @@ -103,13 +109,3 @@ promises. Package signatures are on ## License [Apache License 2.0](LICENSE) — Copyright 2016-2026 Sungkyu Yoo - -### Large versioned dictionaries - -Use `OpenVersioned` for V3 snapshots, atomic expected-version updates and background -engine replacement. V1/V2 APIs remain compatible. See the [V3 guide](docs/content/reference/versioned.md) -and [performance report](docs/content/reference/versioned-performance.md). - -R2 reuses unchanged downloaded buckets and reduces sparse-engine build memory. -R3 adds bounded `Scan`, `MaskText`, and `ReplaceText` with original byte/rune spans; -see [text processing](docs/content/reference/text-processing.md). diff --git a/changes/unreleased/20260906-docs-condense-and-dedupe.yaml b/changes/unreleased/20260906-docs-condense-and-dedupe.yaml new file mode 100644 index 00000000..b1ea11d0 --- /dev/null +++ b/changes/unreleased/20260906-docs-condense-and-dedupe.yaml @@ -0,0 +1,5 @@ +kind: Documentation +body: "Condensed the documentation site and removed duplicated pages. `preset-engine` and `redis-backed-engine` each carried a Quick Start, an API Reference block, and a preset table; they now split by question — which preset to pick, versus how the engine stays in sync across instances. `reference/api` repeated the Create/Add/Find/Info/Flush/Close surface in a second preset section and gave a one-line heading to each method; methods are tables now, with prose reserved for the contracts that need it. Eight pages carried hand-written nav footers the theme already renders, `monitoring` listed the same Grafana panels twice, and `deployment` repeated the topology and health-check wiring from other pages. Three defects surfaced: `extending/custom-storage` linked to `guides/presets/`, which does not exist; `cli/_index` and `getting-started/installation` disagreed on the command count; and `reference/api` omitted `FindSetContext` from its context-variants list. Performance and verification reports keep every data table unchanged. No documented behavior changed." +time: 2026-09-06T23:30:00+09:00 +custom: + Issue: "249" diff --git a/docs/content/_index.md b/docs/content/_index.md index f47420af..c25208cf 100644 --- a/docs/content/_index.md +++ b/docs/content/_index.md @@ -46,11 +46,11 @@ func main() { Guides - Use batch, parallel, and Redis-backed engines. + Batch, parallel, and preset-optimized engines. - API Reference - Review public APIs and storage schemas. + Reference + Public API, storage schemas, and measurements. Operations @@ -62,7 +62,7 @@ func main() { CLI - Drive a collection from the shell, twenty commands. + Drive a collection from the shell. Extending diff --git a/docs/content/cli/_index.md b/docs/content/cli/_index.md index f98ae658..aa7d305c 100644 --- a/docs/content/cli/_index.md +++ b/docs/content/cli/_index.md @@ -5,30 +5,21 @@ weight: 6 # CLI -`acor` is the third way into the same collection: one binary, twenty commands, every one of -them a shell over the library. It is the entry point for the things a program should not have -to be written for — seeding a dictionary, checking what is in one, running a migration, -grepping a log against keywords that live in Redis. - -> **The CLI is not covered by the `v1` compatibility promise.** Quoting -> [Compatibility](../reference/compatibility/#what-is-not-covered) directly: -> -> **The CLI.** Flags, output format, and exit codes of the `acor` command can change in -> any release. If you need a stable contract, call the library rather than parsing CLI -> output. -> -> The library (`pkg/acor`) is covered; this is the surface that is not. - -## Install +`acor` is the third way into the same collection: one binary, every command a shell over +the library. It is for the things a program should not have to be written for — seeding a +dictionary, checking what is in one, running a migration, grepping a log against keywords +that live in Redis. + +> **The CLI is not covered by the `v1` compatibility promise.** Flags, output format, and +> exit codes can change in any release. If you need a stable contract, call the library +> rather than parsing CLI output. See +> [Compatibility](../reference/compatibility/#what-is-covered). ```bash go install github.com/skyoo2003/acor/cmd/acor@latest ``` -Full instructions, including verifying the install, are on -[Getting Started → Installation](../getting-started/installation/#cli-installation). - -## The twenty commands +## The commands | Group | Commands | | ----- | -------- | @@ -39,17 +30,13 @@ Full instructions, including verifying the install, are on | Migrate | `migrate`, `migrate-rollback` | | Dictionary (V3) | `dictionary list\|diff\|replace\|status\|copy-v2\|prune` | -Flags are deliberately not tabulated here. `acor --help` prints the command list followed by -the flag set's own defaults, so each flag's description lives exactly once — next to where the -flag is registered — and cannot drift out of step with a copy on this page. +Flags are deliberately not tabulated here. `acor --help` prints the command list followed +by the flag set's own defaults, so each flag's description lives exactly once — next to +where the flag is registered — and cannot drift out of step with a copy on this page. -`acor version` is the one command that needs no Redis. It prints the version stamped in at +`acor version` is the one command needing no Redis. It prints the version stamped in at release build time, or `dev` for a binary you built yourself. -## Sections - -- [Commands](commands/) - Option ordering, batch modes, the four matching shapes, parallel chunking, and when the local cache earns its memory - -## Navigation - -← [Server](../server/) | [Extending](../extending/) → +[Commands](commands/) covers what those one-line flag descriptions cannot carry: option +ordering, batch modes, the matching shapes, parallel chunking, and when the local cache +earns its memory. diff --git a/docs/content/cli/commands.md b/docs/content/cli/commands.md index ed607b6d..eb0e0285 100644 --- a/docs/content/cli/commands.md +++ b/docs/content/cli/commands.md @@ -5,30 +5,24 @@ weight: 1 # Commands -`acor --help` prints the command list and every flag with its default. This page covers the -behavior those one-line flag descriptions cannot carry: the ordering rule, what each batch -mode does on failure, how the matching commands differ from one another, and when the local -cache is worth its memory. +`acor --help` prints every command and flag with its default. This page covers what those +one-line descriptions cannot carry. ## Options come before the command -CLI options must appear before the command. Batch commands accept keywords as -arguments, or `-` as the only argument to read one keyword per line from stdin: +Batch commands take keywords as arguments, or `-` as the only argument to read one keyword +per line from stdin: ```bash acor -addr localhost:6379 -batch-mode transactional add-many foo bar "hello world" printf 'foo\nbar\n' | acor -addr localhost:6379 remove-many - ``` -`best-effort` is the default and reports per-keyword failures in JSON while -returning success; `transactional` fails the command if the whole batch cannot -be committed. +`best-effort` is the default: it reports per-keyword failures in JSON and still exits +successfully. `transactional` fails the command if the whole batch cannot be committed. ## Matching -Beyond `find` and `find-index`, the matching commands cover the set, span, and -presence shapes of the same scan: - ```bash acor -addr localhost:6379 find-set "he is him" acor -addr localhost:6379 contains "he is him" @@ -37,19 +31,17 @@ acor -addr localhost:6379 -match-kind leftmost-longest -whole-word \ ``` `find-set` reports each keyword once, `contains` stops at the first match, and -`find-matches` reports each occurrence with its rune span in scan order. -`-match-kind` and `-whole-word` apply only to `find-matches`. +`find-matches` reports each occurrence with its rune span in scan order. `-match-kind` and +`-whole-word` apply to `find-matches` only. -`-whole-word` assumes a script that separates words with spaces or punctuation. -In scripts written without inter-word boundaries (CJK, Thai, …) every adjacent -character counts as a word character, so nearly every match is treated as -mid-word and dropped — scan such text without `-whole-word`, or use the library's -`MatchOptions.WordRune` to supply your own boundary rule. +`-whole-word` assumes a script that separates words with spaces or punctuation. In scripts +written without word boundaries (CJK, Thai, …) every adjacent character counts as a word +character, so nearly every match is treated as mid-word and dropped — scan such text +without `-whole-word`, or use the library's `MatchOptions.WordRune`. ## Parallel matching -Parallel matching accepts a text argument, or `-` to read the complete text -from stdin. `word`, `sentence`, and `line` chunk boundaries are available: +Takes a text argument, or `-` to read the whole text from stdin: ```bash acor -addr localhost:6379 -workers 8 -chunk-size 10000 \ @@ -59,11 +51,13 @@ acor -addr localhost:6379 -workers 8 \ find-index-parallel - < document.txt ``` +`-boundary` takes `word`, `sentence`, or `line`. + ## Local cache and presets -Use `-cache` with the normal V2 engine, or select a Redis-backed local preset -engine. Preset mode already keeps a local engine, so `-cache` and `-preset` -cannot be combined: +`-cache` enables the local cache on the normal V2 engine; `-preset` selects a Redis-backed +local engine instead. Preset mode already keeps a local engine, so the two cannot be +combined: ```bash acor -addr localhost:6379 -cache find-parallel - < document.txt @@ -72,23 +66,20 @@ acor -addr localhost:6379 -preset balanced \ -invalidation-poll-interval 30s find-parallel - < document.txt ``` -Available presets are `speed`, `balanced`, and `memory-efficient`; `none` is -the compatibility-preserving default. Preset mode requires an explicit Redis -address and does not support `suggest`, `suggest-index`, or migration commands. -The local cache is most useful for parallel matching, where every chunk shares -one CLI process; a one-shot `find` invocation has no later lookup to reuse it. +Presets are `speed`, `balanced`, and `memory-efficient`; `none` is the +compatibility-preserving default. Preset mode requires an explicit Redis address and +supports neither `suggest`/`suggest-index` nor the migration commands. -The trade-offs behind each preset are in +The cache pays off for parallel matching, where every chunk shares one CLI process. A +one-shot `find` has no later lookup to reuse it. Preset trade-offs are in [Guides → Preset-Optimized Engine](../../guides/preset-engine/). ## Versioned dictionaries -Use `acor -name new-v3-name dictionary list|diff|replace|status|copy-v2|prune`. -Replacement and copying require `--expected-version`; empty replacements and -empty V2 sources require `--allow-empty`. Diff and replace read a JSON string -array from stdin. See the [V3 guide](../../reference/versioned/) for pagination, -case policy, cutover and command examples. - -## Navigation +```bash +acor -name new-v3-name dictionary list|diff|replace|status|copy-v2|prune +``` -← [CLI](../) | [Extending](../../extending/) → +Replacement and copying require `--expected-version`; empty replacements and empty V2 +sources require `--allow-empty`. Diff and replace read a JSON string array from stdin. +Pagination, case policy, and cutover: [V3 guide](../../reference/versioned/#cli). diff --git a/docs/content/extending/_index.md b/docs/content/extending/_index.md index 2f70752c..bde71e1b 100644 --- a/docs/content/extending/_index.md +++ b/docs/content/extending/_index.md @@ -5,12 +5,5 @@ weight: 7 # Extending -Extend ACOR with custom functionality. - -## Sections - -- [Custom Storage](custom-storage/) - Why Redis is the only backend, and how to test without one - -## Navigation - -← [CLI](../cli/) +- [Custom Storage](custom-storage/) — why Redis is the only backend, and how to test + without one diff --git a/docs/content/extending/custom-storage.md b/docs/content/extending/custom-storage.md index b8560d8d..0e33aa63 100644 --- a/docs/content/extending/custom-storage.md +++ b/docs/content/extending/custom-storage.md @@ -5,38 +5,32 @@ weight: 1 # Custom Storage -**ACOR does not support custom storage backends.** Redis is the only backend, and -`Create()` always builds it. This page explains why the interface you may have seen in -older releases is gone, and what to do instead for the case it was usually reached for: -testing. +**ACOR does not support custom storage backends.** Redis is the only one, and `Create()` +always builds it. This page explains why the interface you may have seen in older releases +is gone, and what to do instead for the case it was usually reached for: testing. ## What changed in v1.5.0 -Through `v1.4.0` the package exported a `KVStorage` interface, along with `Pipeliner`, -`Subscription`, `StringMapResult`, `PubSubMessage`, and `Z`. They were exported in -anticipation of pluggable backends. +Through `v1.4.0` the package exported `KVStorage` along with `Pipeliner`, `Subscription`, +`StringMapResult`, `PubSubMessage`, and `Z`, in anticipation of pluggable backends. They +are unexported as of `v1.5.0`, for two reasons: -They are unexported as of `v1.5.0`, for two reasons: - -- **Nothing could be plugged in.** No public constructor accepted a `KVStorage`, and no - public function returned one. You could name the interface and write an - implementation, but there was no way to hand it to ACOR — so the capability the - interface implied never existed. +- **Nothing could be plugged in.** No public constructor accepted a `KVStorage` and no + public function returned one, so the capability the interface implied never existed. - **Freezing it would have blocked the feature it was for.** The [compatibility policy](../../reference/compatibility/) forbids adding a method to an - exported interface inside `v1`. `KVStorage` has 23 methods; had `v1.5.0` frozen it, a - future backend needing one more operation would have had nowhere to put it for the - rest of the `v1` line. + exported interface inside `v1`. `KVStorage` has 23 methods; freezing it would have left a + future backend needing one more operation nowhere to put it for the rest of `v1`. -Leaving the shape unfrozen is what keeps pluggable storage possible. Publishing an -interface later is an addition, which `v1` allows; growing a frozen one is not. +Publishing an interface later is an addition, which `v1` allows; growing a frozen one is +not. ## Testing without Redis -This is what the interface was most often reached for, and it has a better answer that -works today: [miniredis](https://github.com/alicebob/miniredis), an in-process -Redis-compatible server. ACOR's own test suite uses it, so the behavior you test against -is the behavior ACOR is tested against. +This is what the interface was most often reached for, and +[miniredis](https://github.com/alicebob/miniredis) answers it today — an in-process +Redis-compatible server. ACOR's own test suite uses it, so the behavior you test against is +the behavior ACOR is tested against. ```go @@ -75,22 +69,18 @@ func TestWithMiniredis(t *testing.T) { } ``` -`miniredis` covers the commands ACOR issues, including the Lua scripts the V2 schema -uses for atomic writes. To exercise Pub/Sub-driven cache invalidation across instances, -point two instances at the same miniredis address. +miniredis covers the commands ACOR issues, including the Lua scripts V2 uses for atomic +writes. To exercise Pub/Sub-driven invalidation across instances, point two instances at +the same miniredis address. ## Reducing Redis traffic -If the reason for a custom backend was to avoid Redis round trips on reads rather than -to replace Redis, that is already available without one: +If the reason for a custom backend was avoiding round trips on reads rather than replacing +Redis, that is already available: - **`Preset`** serves reads from a local automaton and touches Redis only on writes and - invalidations. See [Presets](../../guides/presets/). + invalidations — see [Preset-Optimized Engine](../../guides/preset-engine/). - **`EnableCache`** caches trie data locally and invalidates over Pub/Sub. -`CacheStats()` reports whether either is working — hit rate, rebuild cost, and -invalidation lag — without any Redis I/O of its own. - -## Navigation - -← [Extending](../) +`CacheStats()` reports whether either is working — hit rate, rebuild cost, invalidation +lag — with no Redis I/O of its own. diff --git a/docs/content/getting-started/_index.md b/docs/content/getting-started/_index.md index d5d03c3d..c66374ff 100644 --- a/docs/content/getting-started/_index.md +++ b/docs/content/getting-started/_index.md @@ -5,37 +5,18 @@ weight: 1 # Getting Started -This section covers everything you need to start using ACOR. +- [Installation](installation/) — prerequisites, the package, and the `acor` CLI +- [Quick Start](quick-start/) — a first matcher, and every Redis topology -## Prerequisites +## Runnable examples -- Go >= 1.25 -- Redis >= 3.0 or Valkey >= 7.2 - -## Installation - -```bash -go get github.com/skyoo2003/acor/pkg/acor@latest -``` - -## Next Steps - -- [Installation](installation/) - Detailed setup instructions -- [Quick Start](quick-start/) - Your first ACOR application - -## Runnable Examples - -Three complete programs live in the repository and build against the current API: +Three complete programs in the repository build against the current API: [`examples/basic`](https://github.com/skyoo2003/acor/tree/main/examples/basic), [`examples/batch`](https://github.com/skyoo2003/acor/tree/main/examples/batch), and [`examples/parallel`](https://github.com/skyoo2003/acor/tree/main/examples/parallel). -Each expects a reachable Redis at `localhost:6379` and uses its own collection name, -so running one does not disturb another. +Each expects Redis at `localhost:6379` and uses its own collection name, so one does not +disturb another. ```bash go run ./examples/basic ``` - -## Continue Learning - -After getting started, explore the [Guides](../guides/) for advanced usage patterns. diff --git a/docs/content/getting-started/installation.md b/docs/content/getting-started/installation.md index b6cd173e..b740511d 100644 --- a/docs/content/getting-started/installation.md +++ b/docs/content/getting-started/installation.md @@ -5,70 +5,28 @@ weight: 1 # Installation -## Prerequisites +Requires **Go 1.25+** and **Redis 3.0+** or **Valkey 7.2+**. -- **Go**: Version 1.25 or later -- **Redis**: Version 3.0 or later, **or Valkey**: Version 7.2 or later +ACOR speaks RESP through [go-redis v9](https://github.com/redis/go-redis), so any +Redis- or Valkey-compatible server works. RESP3 is negotiated on connect and falls back +to RESP2 on servers older than `HELLO`. -ACOR uses the standard RESP protocol via [go-redis v9](https://github.com/redis/go-redis) and works with any Redis- or Valkey-compatible server. RESP3 is negotiated on connect and falls back to RESP2 automatically on servers that predate the `HELLO` command. - -## Install the Package +## Library ```bash go get github.com/skyoo2003/acor/pkg/acor@latest ``` -## Verify Installation - -Create a test file to verify ACOR is installed correctly: - -```go -package main - -import ( - "fmt" - "github.com/skyoo2003/acor/pkg/acor" -) - -func main() { - args := &acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "test", - } - - ac, err := acor.Create(args) - if err != nil { - panic(err) - } - defer ac.Close() - - fmt.Println("ACOR installed successfully!") -} -``` +Verify it with the [Quick Start](../quick-start/) program. -## CLI Installation - -Install the command-line tool: +## CLI ```bash go install github.com/skyoo2003/acor/cmd/acor@latest -``` -Verify the CLI: - -```bash -acor --help -acor version +acor --help # command list and every flag's default +acor version # needs no Redis; prints `dev` for a locally built binary ``` -`acor version` needs no Redis and prints the version stamped at release build -time (`dev` for a locally built binary). - -That is the whole of installing it. What the nineteen commands do — option ordering, batch -modes, the four matching shapes, parallel chunking, and when the local cache earns its -memory — is the [CLI](../../cli/) section. - -## Next Steps - -- [Quick Start](../quick-start/) - Build your first application -- [Guides](../../guides/) - Learn about batch operations and parallel matching +What the commands do — option ordering, batch modes, matching shapes, parallel chunking, +the local cache — is the [CLI](../../cli/) section. diff --git a/docs/content/getting-started/quick-start.md b/docs/content/getting-started/quick-start.md index 5a96342f..c0176ad8 100644 --- a/docs/content/getting-started/quick-start.md +++ b/docs/content/getting-started/quick-start.md @@ -5,36 +5,28 @@ weight: 2 # Quick Start -This guide walks you through building your first ACOR application. - -## Basic Usage - ```go package main import ( "fmt" + "github.com/skyoo2003/acor/pkg/acor" ) func main() { - args := &acor.AhoCorasickArgs{ + ac, err := acor.Create(&acor.AhoCorasickArgs{ Addr: "localhost:6379", Name: "sample", - } - - ac, err := acor.Create(args) + }) if err != nil { panic(err) } defer ac.Close() - keywords := []string{"he", "her", "him"} - for _, k := range keywords { - if _, err := ac.Add(k); err != nil { - panic(err) - } + if _, err := ac.AddMany([]string{"he", "her", "him"}, nil); err != nil { + panic(err) } matched, err := ac.Find("he is him") @@ -42,68 +34,45 @@ func main() { panic(err) } fmt.Println(matched) - - if err := ac.Flush(); err != nil { - panic(err) - } } ``` -## Redis Topologies +## Redis topologies -ACOR supports multiple Redis configurations: - -### Standalone +Pick exactly one set of connection fields — mixing them returns +`ErrRedisConflictingTopology`. ```go -args := &acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Password: "", - DB: 0, - Name: "sample", -} -``` - -### Sentinel +// Standalone +args := &acor.AhoCorasickArgs{Addr: "localhost:6379", Name: "sample"} -```go -args := &acor.AhoCorasickArgs{ +// Sentinel — Addrs plus MasterName +args = &acor.AhoCorasickArgs{ Addrs: []string{"localhost:26379", "localhost:26380"}, MasterName: "mymaster", - Password: "", - DB: 0, Name: "sample", } -``` - -### Cluster -```go -args := &acor.AhoCorasickArgs{ - Addrs: []string{"localhost:7000", "localhost:7001", "localhost:7002"}, - Password: "", - Name: "sample", +// Cluster — Addrs without MasterName +args = &acor.AhoCorasickArgs{ + Addrs: []string{"localhost:7000", "localhost:7001"}, + Name: "sample", } -``` - -### Ring -```go -args := &acor.AhoCorasickArgs{ - RingAddrs: map[string]string{ - "shard-1": "localhost:7000", - "shard-2": "localhost:7001", - }, - Password: "", - DB: 0, - Name: "sample", +// Ring — shard name to address +args = &acor.AhoCorasickArgs{ + RingAddrs: map[string]string{"shard-1": "localhost:7000", "shard-2": "localhost:7001"}, + Name: "sample", } ``` -## Next Steps +`Password`, `DB`, and the timeout and pool fields apply to every topology; `DB` is +rejected together with `Addrs`. Full field list: +[API Reference](../../reference/api/#ahocorasickargs). + +## Next -- [Batch Operations](../../guides/batch-operations/) - Optimize bulk operations -- [Parallel Matching](../../guides/parallel-matching/) - Process large texts efficiently -- [Redis-Backed Engine](../../guides/redis-backed-engine/) - Redis persistence with local speed -- [Match Details](../../reference/api/#findmatches) - Ordered rune spans, matching options, and streaming -- [API Reference](../../reference/api/) - Complete API documentation +[Batch operations](../../guides/batch-operations/) · +[Parallel matching](../../guides/parallel-matching/) · +[Preset engine](../../guides/preset-engine/) · +[API reference](../../reference/api/) diff --git a/docs/content/guides/_index.md b/docs/content/guides/_index.md index 97a8988b..7fe0718e 100644 --- a/docs/content/guides/_index.md +++ b/docs/content/guides/_index.md @@ -5,15 +5,7 @@ weight: 2 # Guides -Practical guides for using ACOR effectively. - -## Available Guides - -- [Batch Operations](batch-operations/) - Optimize bulk keyword operations -- [Parallel Matching](parallel-matching/) - Process large texts with multiple workers -- [Redis-Backed Engine](redis-backed-engine/) - Redis persistence with local preset-optimized speed -- [Preset-Optimized Engine](preset-engine/) - What each preset trades away, and how to pick one - -## Navigation - -← [Getting Started](../getting-started/) | [Reference](../reference/) → +- [Batch Operations](batch-operations/) — bulk adds, removes, and multi-text scans +- [Parallel Matching](parallel-matching/) — chunking large texts across workers +- [Preset-Optimized Engine](preset-engine/) — what each preset trades away, and how to pick one +- [Redis-Backed Engine](redis-backed-engine/) — architecture, invalidation safety, topologies diff --git a/docs/content/guides/batch-operations.md b/docs/content/guides/batch-operations.md index 9b9355cd..16324780 100644 --- a/docs/content/guides/batch-operations.md +++ b/docs/content/guides/batch-operations.md @@ -5,13 +5,8 @@ weight: 1 # Batch Operations -ACOR supports batch operations for better performance when working with multiple keywords. - -## Overview - -Batch operations reduce network round-trips by grouping multiple operations together. - -## Adding Multiple Keywords +`AddMany` and `RemoveMany` commit a whole batch in one transaction — two round trips +regardless of batch size, against one per keyword for a loop over `Add`. ```go @@ -26,52 +21,31 @@ fmt.Printf("Added: %d, Failed: %d, Skipped: %d\n", len(result.Added), len(result.Failed), len(result.Skipped)) ``` -## Batch Modes - -### Transactional - -Rolls back all changes if any error occurs: - -```go -result, err := ac.AddMany(keywords, &acor.BatchOptions{ - Mode: acor.BatchModeTransactional, -}) -``` - -### BestEffort +`nil` options mean best-effort, which is also what `BatchOptions` with no `Mode` selects. -Continues on errors and returns partial results: +## Modes -```go -result, err := ac.AddMany(keywords, &acor.BatchOptions{ - Mode: acor.BatchModeBestEffort, -}) -``` +| Mode | On a per-keyword failure | +| ---- | ------------------------ | +| `BatchModeBestEffort` (default) | Commits the rest; failures land in `result.Failed`, and the call still returns success | +| `BatchModeTransactional` | Rolls the whole batch back and returns an error | -## Removing Multiple Keywords +Inspect what best-effort dropped: - ```go -result, err := ac.RemoveMany([]string{"he", "her"}, nil) -if err != nil { - panic(err) +for _, ke := range result.Failed { + fmt.Printf("%q failed: %v\n", ke.Keyword, ke.Error) } -fmt.Printf("Removed: %d\n", len(result.Removed)) ``` -`nil` options mean best-effort, the same default `AddMany` uses. Pass options to -control the batch mode: +`result.Skipped` holds duplicate adds and absent removes — not errors. Field-by-field +shapes for `BatchResult` and `KeywordError` are in the +[API reference](../../reference/api/#batchresult). - -```go -result, err := ac.RemoveMany([]string{"he", "her"}, &acor.BatchOptions{ - Mode: acor.BatchModeTransactional, -}) -_ = result -_ = err -``` +## Scanning many texts -## Finding Matches in Multiple Texts +`FindMany` loads the automaton once and scans every text against that one snapshot, so it +costs the same round trip as a single `Find`. The result is keyed by the input text. ```go @@ -86,37 +60,8 @@ for text, matches := range results { } ``` -## Batch Result Structure - -```go -type BatchResult struct { - Added []string // Keywords successfully added - Removed []string // Keywords successfully removed - Failed []KeywordError // Keywords that could not be processed, with their errors - Skipped []string // Keywords skipped (e.g. duplicates in the input) -} - -type KeywordError struct { - Keyword string // The keyword that caused the error - Error error // The error that occurred while processing it -} -``` - -Inspect failures in `BatchModeBestEffort`: - -```go -for _, ke := range result.Failed { - fmt.Printf("%q failed: %v\n", ke.Keyword, ke.Error) -} -``` - -## Performance Tips - -1. Use batch sizes between 100-1000 keywords -2. Use `BatchModeTransactional` when data consistency is critical -3. Use `BatchModeBestEffort` when partial success is acceptable - -## Next Steps +## Sizing -- [Parallel Matching](../parallel-matching/) - Process large texts efficiently -- [API Reference](../../reference/api/) - Complete API documentation +100–1,000 keywords per batch is the useful range: below it the transaction overhead +dominates, above it the single Lua commit grows large enough to block Redis noticeably. +Measured costs are on the [benchmarks page](../../reference/benchmarks/#bulk-load-addmany). diff --git a/docs/content/guides/parallel-matching.md b/docs/content/guides/parallel-matching.md index b799bf4e..da0f8dfe 100644 --- a/docs/content/guides/parallel-matching.md +++ b/docs/content/guides/parallel-matching.md @@ -5,21 +5,17 @@ weight: 2 # Parallel Matching -For large texts, use parallel matching to scan with multiple goroutines. - -## Overview - -Parallel matching splits text into chunks and processes them concurrently, significantly improving performance for large inputs. - -## Basic Usage +`FindParallel` and `FindIndexParallel` split one text into chunks and scan them +concurrently. The automaton is loaded once per call, so the Redis cost is the same as a +serial `Find` however many chunks result. ```go matches, err := ac.FindParallel(largeText, &acor.ParallelOptions{ - Workers: 4, - ChunkSize: 1000, - AutoOverlap: true, - Boundary: acor.ChunkBoundaryWord, + Workers: 4, + ChunkSize: 1000, // runes per base chunk; must be positive + AutoOverlap: true, + Boundary: acor.ChunkBoundaryWord, }) if err != nil { panic(err) @@ -27,108 +23,39 @@ if err != nil { _ = matches ``` -## Chunk Boundaries - -Boundary selection chooses where base chunks end. Enable `AutoOverlap` to protect -matches across every boundary; boundary selection alone cannot guarantee this. - -### ChunkBoundaryWord (default) - -Splits at word boundaries, ideal for natural language text: - -```go -opts := &acor.ParallelOptions{ - Workers: 4, - ChunkSize: 1000, - AutoOverlap: true, - Boundary: acor.ChunkBoundaryWord, -} -``` - -### ChunkBoundaryLine - -Splits at line breaks, ideal for log files: - -```go -opts := &acor.ParallelOptions{ - Workers: 4, - ChunkSize: 1000, - AutoOverlap: true, - Boundary: acor.ChunkBoundaryLine, -} -``` - -### ChunkBoundarySentence - -Splits at sentence endings, ideal for document processing: - -```go -opts := &acor.ParallelOptions{ - Workers: 4, - ChunkSize: 1000, - AutoOverlap: true, - Boundary: acor.ChunkBoundarySentence, -} -``` - -## Automatic Boundary Protection - -`AutoOverlap: true` uses the longest keyword in the engine loaded for this call. -Each base chunk owns its starting positions and extends its right search range by -`max(Overlap, longest keyword rune length - 1)`. Matches starting in the extension -belong to the next base chunk and are excluded from the current chunk. This also -finds keywords longer than `ChunkSize`, including Korean and emoji keywords. -No extra Redis query is needed, and all workers use the same dictionary snapshot. +## Boundaries -`FindParallel` returns each keyword once, in chunk order and then first-match scan -order within each chunk. That order need not equal serial scan order. -`FindIndexParallel` returns sorted unique rune positions in the original text. +Only `Boundary` changes between these; everything else stays as above. -The default `AutoOverlap` is `false`, preserving legacy overlapping chunks. -In that mode, an insufficient `Overlap` can miss boundary matches, and a keyword -longer than a chunk may not fit in any chunk. Options are copied before normalization. -`ChunkSize` must be positive; empty input performs no Redis reads. +| Boundary | Splits at | Suited to | +| -------- | --------- | --------- | +| `ChunkBoundaryWord` (default) | Whitespace | Natural language | +| `ChunkBoundaryLine` | Newlines | Log files | +| `ChunkBoundarySentence` | `.` `!` `?` | Documents | -## Performance Tuning - -### Worker Count - -Choose worker count based on CPU cores: - -```go -workers := runtime.NumCPU() -``` +Boundary choice alone does not protect a keyword that straddles a split. `AutoOverlap` +does. -For I/O-bound workloads, consider higher counts: - -```go -workers := runtime.NumCPU() * 2 -``` - -### Chunk Size - -Control chunk size with the `ChunkSize` option: - -```go -opts := &acor.ParallelOptions{ - Workers: 4, - ChunkSize: 10000, // 10,000 runes per base chunk -} -``` +## Boundary protection -## When to Use Parallel Matching +`AutoOverlap: true` extends each base chunk to the right by +`max(Overlap, longest keyword rune length - 1)`, using the longest keyword in the engine +already loaded for this call — no extra Redis read, and every worker shares one snapshot. +Matches starting inside an extension belong to the next chunk and are not double-counted. +This also catches keywords longer than `ChunkSize`, including Korean and emoji. -- Text size > 100KB -- Many pattern matches expected -- CPU cores available for parallel work +`AutoOverlap` defaults to `false`, which keeps the legacy behavior: an `Overlap` shorter +than your longest keyword silently misses boundary matches, and a keyword longer than +`ChunkSize` may fit in no chunk at all. -## When to Avoid +## Result order -- Small texts (< 10KB) -- Single-match scenarios -- Resource-constrained environments +`FindParallel` reports each keyword once, in chunk order and then first-match order +within a chunk — which need not equal serial scan order. `FindIndexParallel` returns +sorted, unique rune positions in the original text. Empty input performs no Redis read. -## Next Steps +## When it pays -- [Batch Operations](../batch-operations/) - Optimize bulk keyword operations -- [API Reference](../../reference/api/) - Complete API documentation +Worth it above roughly 100 KB of text, with cores to spare. Below ~10 KB, or for a +single expected match, the chunking overhead outweighs it. Start `Workers` at +`runtime.NumCPU()`. diff --git a/docs/content/guides/preset-engine.md b/docs/content/guides/preset-engine.md index c56f0fa9..ae035583 100644 --- a/docs/content/guides/preset-engine.md +++ b/docs/content/guides/preset-engine.md @@ -5,120 +5,47 @@ weight: 3 # Preset-Optimized Engine -ACOR provides a Redis-backed Aho-Corasick engine with selectable architecture presets. Created via the unified `Create` API with a `Preset` field. Writes go to Redis atomically (V2 Lua scripts with optimistic locking); reads hit the local engine with no Redis I/O. - -## When to Use - -- Production deployments requiring Redis persistence -- Distributed systems with multiple instances sharing a keyword collection -- High-throughput text matching with zero read-latency on the hot path -- Applications needing both durability and speed - -## Quick Start +Setting `Preset` on `AhoCorasickArgs` builds a local automaton next to the Redis +collection: writes still go to Redis atomically, reads never touch it. This page is about +picking a preset. How the engine stays in sync across instances is +[Redis-Backed Engine](../redis-backed-engine/). ```go -package main - -import ( - "fmt" - "github.com/skyoo2003/acor/pkg/acor" -) - -func main() { - ac, _ := acor.Create(&acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-collection", - Preset: acor.PresetBalanced, - }) - defer ac.Close() - - ac.Add("he") - ac.Add("her") - ac.Add("him") - - matches, _ := ac.Find("he is him") - fmt.Println(matches) // [he him] - - positions, _ := ac.FindIndex("he is him") - fmt.Println(positions) // map[he:[0] him:[6]] - - info, _ := ac.Info() - fmt.Printf("Keywords: %d, Nodes: %d, Memory: %d bytes\n", - info.Keywords, info.Nodes, info.MemoryBytes) -} -``` - -## Architecture Presets - -Each preset optimizes for a different trade-off between speed, memory, and feature set. The preset is fixed at creation time. - -| Preset | Engine | Best For | Trade-off | -|--------|--------|----------|-----------| -| `PresetSpeed` | Full DFA + flat array trie + compact alphabet mapping | Real-time packet inspection, high-speed log scanning, latency-critical paths | Higher memory proportional to states x alphabet size | -| `PresetBalanced` | Double-Array Trie + Banded DFA + output link compression | General-purpose backend keyword filtering, search engines | Balanced speed and memory | -| `PresetMemoryEfficient` | Map-based sparse trie + Bloom filter pre-filtering + standard NFA | Large-scale domain blocking, malware signature matching, millions of patterns | Slower search due to failure link traversal and map lookups | - -### Choosing a Preset - -- **Start with `PresetBalanced`** — it provides the best speed-to-memory ratio for most workloads. -- Use `PresetSpeed` when latency is critical and memory is available. -- Use `PresetMemoryEfficient` when you have millions of patterns and memory is constrained. - -`PresetSpeed` measured fastest on every query shape on the -[benchmarks page](../../reference/benchmarks/), while `PresetBalanced` trades some of -that for a much smaller transition table. Measure your own corpus rather than -choosing by name. - -## Case Sensitivity - -By default, matching is case-insensitive. Enable case-sensitive matching when needed: - -```go -ac, _ := acor.Create(&acor.AhoCorasickArgs{ +ac, err := acor.Create(&acor.AhoCorasickArgs{ Addr: "localhost:6379", Name: "my-collection", Preset: acor.PresetBalanced, - CaseSensitive: true, + CaseSensitive: false, // default; matching is case-insensitive }) +if err != nil { + panic(err) +} defer ac.Close() ``` -## API Reference - -```go -// Create -ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-collection", - Preset: acor.PresetBalanced, -}) -defer ac.Close() +The preset is fixed at creation. `Info()` reports the resulting `Keywords`, `Nodes`, +`MemoryBytes`, and `TrieDepth`. -// Add/Remove — returns 1 if changed, 0 if no-op -ac.Add("keyword") -ac.Remove("keyword") +## The three presets -// Find (0 RTT on hot path — reads from local engine) -matches, _ := ac.Find("text") // ([]string, error) -positions, _ := ac.FindIndex("text") // (map[string][]int, error) +| Preset | Engine | Best for | Costs | +|--------|--------|----------|-------| +| `PresetSpeed` | Full DFA + flat array trie + compact alphabet map | Packet inspection, high-rate log scanning, latency-critical paths | Memory proportional to states × alphabet size | +| `PresetBalanced` | Double-Array Trie + Banded DFA + output link compression | General backend filtering, search | Neither extreme | +| `PresetMemoryEfficient` | Map-based sparse trie + Bloom pre-filter + NFA | Millions of patterns under a memory cap | Slower search: failure-link traversal and map lookups | -// Stats -info, err := ac.Info() // (*AhoCorasickInfo, error) +**Start with `PresetBalanced`.** Move to `PresetSpeed` when latency is the constraint and +memory is not; to `PresetMemoryEfficient` when it is the other way round. -// Reset -ac.Flush() -``` - -## Refresh and Cancellation - -Stale reads share a reload but respond independently to request cancellation. -A failed reload returns a search error and is retried by a later request. Optional -version polling helps recover from missed Pub/Sub messages; its interval does not -guarantee freshness during failures. See [invalidation safety](../redis-backed-engine/#invalidation-safety) -for the refresh lifecycle and failure counters. +`PresetSpeed` measured fastest on every query shape on the +[benchmarks page](../../reference/benchmarks/), and `PresetBalanced` gives up some of that +for a much smaller transition table. Measure your own corpus before choosing by name. -## Next Steps +## Refresh behavior -- [Redis-Backed Engine](../redis-backed-engine/) - Redis persistence details -- [API Reference](../../reference/api/) - Complete API documentation +Stale reads share one reload and each observes its own cancellation. A failed reload +returns an error to the search rather than silently serving the old engine, and a later +search retries. Optional version polling recovers from missed Pub/Sub messages; its +interval is not a freshness bound. Full lifecycle and failure counters: +[invalidation safety](../redis-backed-engine/#invalidation-safety). diff --git a/docs/content/guides/redis-backed-engine.md b/docs/content/guides/redis-backed-engine.md index ca4f4f21..e3806446 100644 --- a/docs/content/guides/redis-backed-engine.md +++ b/docs/content/guides/redis-backed-engine.md @@ -5,20 +5,18 @@ weight: 4 # Redis-Backed Engine -The preset-optimized Redis mode (enabled via `Preset` on `AhoCorasickArgs`) combines Redis persistence with a local preset-optimized automaton. Redis is the source of truth; reads hit the local engine with no Redis I/O on the hot path. +With `Preset` set, Redis stays the source of truth while every instance serves reads from +its own automaton. This page covers how that stays consistent. Which preset to pick is +[Preset-Optimized Engine](../preset-engine/). -> **Redis or Valkey:** ACOR connects with [go-redis v9](https://github.com/redis/go-redis) over the standard RESP protocol, so any Redis (>= 3.0) or Valkey (>= 7.2) server works, including Standalone, Sentinel, Cluster, and Ring topologies. Cross-instance cache invalidation uses server Pub/Sub, which behaves identically on Redis and Valkey. Both are covered by the integration test suite in CI. - -## When to Use - -- Distributed deployments across multiple instances -- Need for Redis persistence and cross-instance synchronization -- Want preset-optimized local speed without giving up Redis durability -- Migrating from the original `AhoCorasick` for better read performance +> **Redis or Valkey.** ACOR connects over RESP via +> [go-redis v9](https://github.com/redis/go-redis), so Redis 3.0+ or Valkey 7.2+ works in +> Standalone, Sentinel, Cluster, and Ring topologies. Cross-instance invalidation uses +> server Pub/Sub, which behaves identically on both. CI covers both. ## Architecture -``` +```text Write Path Instance A ──Add()──▶ Lua Script (optimistic lock) ──▶ Redis │ @@ -32,16 +30,17 @@ Instance B ◀────────────────┘ Instance A ──Find()──▶ local engine (0 RTT) ``` -- **Writes**: V2 Lua scripts with optimistic locking (up to 3 retries with backoff) -- **Reads**: Local preset-optimized automaton — no Redis I/O -- **Invalidation**: Redis Pub/Sub notifies all instances on mutation -- **Reload failures**: The previous engine is retained, but the waiting search returns an error. A later search retries the reload. +| Path | Behavior | +| ---- | -------- | +| Writes | V2 Lua scripts with optimistic locking, up to 3 retries with backoff | +| Reads | Local automaton, no Redis I/O | +| Invalidation | Redis Pub/Sub on every mutation | +| Failed reload | Previous engine is retained but the waiting search returns an error; a later search retries | -## Invalidation Safety +## Invalidation safety -Redis Pub/Sub is best effort. In a multi-instance deployment, set -`InvalidationPollInterval` to recover from dropped invalidations after a successful -version poll and reload: +Pub/Sub is best effort — a disconnected subscriber misses invalidations. In multi-instance +deployments set `InvalidationPollInterval`: ```go @@ -54,138 +53,43 @@ args := &acor.AhoCorasickArgs{ _ = args ``` -The zero value disables polling. Polling only applies to Preset mode; normal -invalidation still uses Pub/Sub. Each poll reads only the existing `version` hash -field; a changed version marks the engine stale, and the next search reads the -full snapshot. The interval is not a freshness upper bound: Redis failures, query -latency, and rebuild time can delay recovery. This mode provides eventual refresh, -not strong consistency or a guarantee of fresh results during an outage. +Zero disables polling, and polling applies to Preset mode only. Each poll reads just the +`version` hash field; a change marks the engine stale and the next search loads the full +snapshot. -Concurrent stale reads share a reload, while each request independently observes -its own context cancellation. Canceling one waiter leaves the others running. -The instance lifetime owns the job; all waiters leaving or `Close` cancels it. -Redis reads and engine builds run outside the state lock. A generation check -rejects and retries snapshots overtaken by local writes or invalidations. +**The interval is not a freshness bound.** Redis failures, query latency, and rebuild time +all delay recovery. This is eventual refresh, not strong consistency. -`CacheStats().PresetReloadFailures` counts failed shared jobs once per job; -`PresetPollFailures` counts failed version polls. Cancellation is excluded. -Both counters are cumulative per instance and remain zero outside Preset mode. -Monitor these counters alongside application search errors; a retained previous -engine does not turn a failed reload into a successful response. - -## Quick Start - - -```go -package main - -import ( - "fmt" - "github.com/skyoo2003/acor/pkg/acor" -) - -func main() { - ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-collection", - Preset: acor.PresetBalanced, - CaseSensitive: false, - }) - if err != nil { - panic(err) - } - defer ac.Close() - - added, err := ac.Add("hello") - if err != nil { - panic(err) - } - fmt.Printf("Added: %d\n", added) - - matches, err := ac.Find("hello world") - if err != nil { - panic(err) - } - fmt.Println(matches) // [hello] -} -``` +Concurrent stale reads share one reload job while each request observes its own context +cancellation — cancelling one waiter leaves the others running, and all waiters leaving +(or `Close`) cancels the job. Redis reads and engine builds run outside the state lock, +and a generation check rejects snapshots overtaken by local writes. -## AhoCorasickArgs (Preset field) +Watch `CacheStats().PresetReloadFailures` (once per failed shared job) and +`PresetPollFailures` (per failed version poll) alongside your search errors. Cancellation +is excluded from both; both stay zero outside Preset mode. A retained previous engine does +not turn a failed reload into a successful response. See +[Monitoring](../../operations/monitoring/). -The `AhoCorasickArgs` struct includes a `Preset` field for selecting the local engine architecture: +## Topologies and connection tuning -```go -type AhoCorasickArgs struct { - // ... Addr, Addrs, RingAddrs, Password, DB, Name ... - Preset Preset // Architecture preset; zero value is PresetNone (original mode). Set it (e.g. PresetBalanced) to enable preset mode - CaseSensitive bool // Enable case-sensitive matching (default: false) - // ... other fields ... -} -``` +All four topologies are configured through the connection fields on `AhoCorasickArgs` — +see [Redis topologies](../../getting-started/quick-start/#redis-topologies). -All standard Redis topologies are supported (Standalone, Sentinel, Cluster, Ring) via the connection fields on `AhoCorasickArgs`. +`DialTimeout`, `ReadTimeout`, `WriteTimeout`, `MaxRetries`, and `PoolSize` pass straight +to go-redis for every topology. Zero keeps the go-redis default; `-1` disables the +read/write timeouts or command retries where supported. -Connection resilience can be tuned with `DialTimeout`, `ReadTimeout`, -`WriteTimeout`, `MaxRetries`, and `PoolSize`. These values pass directly to -go-redis for every topology; zero keeps the go-redis default. Use `-1` to -disable read/write timeouts or command retries where supported. +## Preset mode versus plain V2 -## Preset Selection - -The same [architecture presets](../preset-engine/#architecture-presets) are available: - -| Preset | Use Case | -|--------|----------| -| `PresetSpeed` | Latency-critical, memory available | -| `PresetBalanced` | Default — best speed-to-memory ratio | -| `PresetMemoryEfficient` | Millions of patterns, memory constrained | - -If the `Preset` field is left unset, it defaults to `PresetNone`, which runs the -original `AhoCorasick` mode (not the preset-optimized engine). You must set -`Preset` explicitly (e.g. `PresetBalanced`) to enable this engine. - -## API Reference - -```go -// Create -ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-collection", - Preset: acor.PresetBalanced, -}) - -// Add/Remove -added, err := ac.Add("keyword") // (int, error) -removed, err := ac.Remove("keyword") // (int, error) - -// Find (0 RTT on hot path — reads from local engine) -matches, err := ac.Find("text") // ([]string, error) -positions, err := ac.FindIndex("text") // (map[string][]int, error) - -// Stats -info, err := ac.Info() // (*AhoCorasickInfo, error) - -// Flush and Close -err := ac.Flush() -err := ac.Close() -``` - -## Comparison with AhoCorasick - -| Feature | `AhoCorasick` (no Preset) | `AhoCorasick` (with Preset) | -|---------|--------------|-----------------| -| Read latency | 1 RTT (V2) or cached | 0 RTT (local engine) | +| | No `Preset` | With `Preset` | +|---|---|---| +| Read latency | 1 RTT, or 0 with `EnableCache` | 0 RTT | | Write latency | Lua script | Lua script + optimistic lock | | Cross-instance sync | Pub/Sub cache invalidation | Pub/Sub engine rebuild | | Schema | V1 or V2 | V2 only | -| Presets | N/A | Speed, Balanced, MemoryEfficient | -| Suggest/SuggestIndex | Yes | No (`ErrSuggestRequiresRedis`) | -| Batch operations | Yes | Yes | -| Parallel matching | Yes | Yes | - -Use a `Preset`-optimized `AhoCorasick` when you need the fastest possible reads in a distributed setup and can accept the V2-only constraint. - -## Next Steps +| `Suggest` / `SuggestIndex` | Yes | No — `ErrSuggestRequiresRedis` | +| Batch, parallel matching | Yes | Yes | -- [Preset-Optimized Engine](../preset-engine/) - Redis-backed engine with local speed -- [API Reference](../../reference/api/) - Complete API documentation +`Preset` is unset by default (`PresetNone`), which runs the original mode. Choose it when +you need the fastest reads across several instances and can accept V2-only, no-`Suggest`. diff --git a/docs/content/operations/_index.md b/docs/content/operations/_index.md index 6695f240..a9bbe2ab 100644 --- a/docs/content/operations/_index.md +++ b/docs/content/operations/_index.md @@ -5,14 +5,6 @@ weight: 4 # Operations -Operational guides for running ACOR in production. - -## Sections - -- [Deployment](deployment/) - Deploy ACOR in various environments -- [Monitoring](monitoring/) - Monitor ACOR performance -- [Troubleshooting](troubleshooting/) - Diagnose and fix issues - -## Navigation - -← [Reference](../reference/) | [Server](../server/) → +- [Deployment](deployment/) — topologies, Kubernetes, Docker Compose, health endpoints +- [Monitoring](monitoring/) — cache statistics, and the `server/*` metrics, logs, and traces +- [Troubleshooting](troubleshooting/) — the errors you are most likely to hit, and their fixes diff --git a/docs/content/operations/deployment.md b/docs/content/operations/deployment.md index 18a33464..677ab042 100644 --- a/docs/content/operations/deployment.md +++ b/docs/content/operations/deployment.md @@ -5,85 +5,26 @@ weight: 1 # Deployment -Guides for deploying ACOR in various environments. +ACOR is a library: you deploy your application, and it connects to Redis. Connection +fields per topology are in +[Redis topologies](../../getting-started/quick-start/#redis-topologies) — standalone for +development and small workloads, Sentinel for failover, Cluster for horizontal scaling, +Ring for client-side sharding. -## Architecture Overview - -```mermaid -graph TB - subgraph Application - A[ACOR Client] - end - - subgraph Redis - B[(Standalone)] - C[(Sentinel)] - D[(Cluster)] - end - - A --> B - A --> C - A --> D -``` - -## Standalone Deployment - -Simplest deployment for development or small workloads: +Read credentials from the environment rather than the source: ```go ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "redis:6379", + Addr: os.Getenv("REDIS_ADDR"), Password: os.Getenv("REDIS_PASSWORD"), - DB: 0, Name: "production", }) -if err != nil { - panic(err) -} -``` - -## High Availability with Sentinel - -For production workloads requiring failover: - -```go -ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addrs: []string{ - "sentinel-1:26379", - "sentinel-2:26379", - "sentinel-3:26379", - }, - MasterName: "mymaster", - Password: os.Getenv("REDIS_PASSWORD"), - Name: "production", -}) if err != nil { panic(err) } ``` -## Cluster Deployment - -For horizontal scaling: - -```go -ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addrs: []string{ - "redis-node-1:7000", - "redis-node-2:7000", - "redis-node-3:7000", - }, - Password: os.Getenv("REDIS_PASSWORD"), - Name: "production", -}) -if err != nil { - panic(err) -} -``` - -## Kubernetes Deployment - -### ConfigMap +## Kubernetes ```yaml apiVersion: v1 @@ -93,11 +34,7 @@ metadata: data: REDIS_ADDR: "redis-service:6379" ACOR_COLLECTION: "production" -``` - -### Deployment - -```yaml +--- apiVersion: apps/v1 kind: Deployment metadata: @@ -123,7 +60,6 @@ spec: ## Docker Compose ```yaml -version: "3.8" services: redis: image: redis:7-alpine @@ -138,65 +74,29 @@ services: - REDIS_ADDR=redis:6379 ``` -## Health Checks - -> `server/health` belongs to the experimental `acor/server` module, which is -> versioned separately and not covered by the core compatibility promise. See -> [Server](../../server/) for what that means for your `go.mod`, and -> [Running a Server](../../server/running/) for this wiring alongside the HTTP -> and gRPC APIs. - -The `server/health` package registers Kubernetes-compatible endpoints on an -`http.ServeMux`: - -- `/healthz` — liveness; always returns `200 OK` while the process is up. -- `/readyz` — readiness; runs every registered `Checker` and returns `503` if - any checker reports as unhealthy. +## Health endpoints - -```go -package main - -import ( - "log" - "net/http" +> `server/health` belongs to the experimental `acor/server` module, versioned separately +> and not covered by the core compatibility promise. See [Server](../../server/). - "github.com/skyoo2003/acor/pkg/acor" - "github.com/skyoo2003/acor/server/health" -) +`health.RegisterHTTPHandlers(mux, checker)` registers two Kubernetes-compatible routes: -// redisChecker is a readiness check that implements health.Checker. -type redisChecker struct{ ac *acor.AhoCorasick } +| Route | Meaning | +| ----- | ------- | +| `/healthz` | Liveness — `200 OK` while the process is up. Put nothing about Redis here | +| `/readyz` | Readiness — runs every registered `Checker`, `503` if any is unhealthy | -func (c redisChecker) Check() health.CheckResult { - if _, err := c.ac.Info(); err != nil { - return health.CheckResult{Status: health.StatusUnhealthy, Details: err.Error()} - } - return health.CheckResult{Status: health.StatusHealthy} -} +Keep Redis reachability out of liveness: a Redis outage would fail every replica's +liveness probe at once, restart all of them, and repair nothing. It is a *take me out of +the load balancer* signal. -func main() { - ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "redis:6379", - Name: "production", - }) - if err != nil { - log.Fatal(err) - } - - checker := health.NewChecker() - checker.Register("redis", redisChecker{ac}) - - mux := http.NewServeMux() - health.RegisterHTTPHandlers(mux, checker) // registers /healthz and /readyz - log.Fatal(http.ListenAndServe(":8080", mux)) -} -``` +The complete `main` — checker with its own deadline, mux composition, graceful shutdown — +is [Running a Server](../../server/running/), which also explains why a readiness check +built on `Info()` costs what the dictionary costs. -## Best Practices +## Checklist -1. Use connection pooling (built-in) -2. Set appropriate timeouts -3. Monitor Redis memory usage -4. Use V2 schema for new collections -5. Implement graceful shutdown +1. Use the V2 schema for new collections (the default). +2. Set timeouts and pool size from measurements, not from guesses. +3. Monitor Redis memory and the [cache counters](../monitoring/). +4. Call `Close()` on shutdown. diff --git a/docs/content/operations/monitoring.md b/docs/content/operations/monitoring.md index 453fed92..57754008 100644 --- a/docs/content/operations/monitoring.md +++ b/docs/content/operations/monitoring.md @@ -5,36 +5,21 @@ weight: 2 # Monitoring -Monitor ACOR performance with built-in observability support. - -> **The `server/*` packages on this page are experimental.** They live in the -> separate `github.com/skyoo2003/acor/server` module, which publishes no version -> tags of its own and is **not covered by the core module's compatibility -> promise** — its API can change in any release. The core library (`pkg/acor`) is -> unaffected. -> -> `go get github.com/skyoo2003/acor/server` resolves a pseudo-version from `main`. -> Pin the core module explicitly in the same `go.mod`: Go ignores a dependency's -> own `replace` directive, so without a pin you get whichever core version the -> server module's `require` names. - -## Overview - -Two layers, and which one you get depends on how you run ACOR: - -| Layer | What it gives you | Covered by the `v1` promise | -| ----- | ----------------- | --------------------------- | -| `pkg/acor` — the library | `CacheStats()`: cache hit rate, rebuild cost, invalidation lag | ✅ | +Two layers, and which you get depends on how you run ACOR: + +| Layer | What it gives you | Covered by `v1` | +| ----- | ----------------- | --------------- | +| `pkg/acor` — the library | `CacheStats()`: hit rate, rebuild cost, invalidation lag | ✅ | | `acor/server` — the service | Prometheus metrics, structured JSON logs, OpenTelemetry traces | ❌ experimental | -Embedding the library gets you the first row only. Everything after the next section is -`server/*`, which means importing a separate, experimental module — reach for it when -you run ACOR as a service, not to instrument your own process. +Embedding the library gets you the first row. Everything from +[Service layer](#service-layer) on means importing the separate, experimental +`acor/server` module — see [Server](../../server/) for what that implies for your +`go.mod`. -## Core library: cache statistics +## Cache statistics -`CacheStats()` reports how often reads avoid Redis, rebuild cost, invalidation lag, -and Preset refresh failures. It does no Redis I/O, so scraping it on a timer is cheap. +`CacheStats()` does no Redis I/O, so scraping it on a timer is cheap. ```go @@ -56,95 +41,75 @@ lag := stats.LastInvalidationLag _, _, _ = hitRate, meanRebuild, lag ``` -Wire those three into whatever you already run — a Prometheus collector, an OTel -meter, a log line. ACOR deliberately depends on no metrics library, so the choice -stays yours. +Wire those into whatever you already run — a Prometheus collector, an OTel meter, a log +line. ACOR depends on no metrics library on purpose. -### Preset refresh failures +### Reading the numbers -Track increases in `PresetReloadFailures` and `PresetPollFailures` alongside search -errors. A shared failed reload increments the first counter once, even when many -requests receive the error. A failed version poll increments the second counter; -its next tick retries. Request cancellation is excluded from both counters. -These counters remain zero outside Preset mode, and polling is disabled by default. +- **Per instance, per process.** Nothing is aggregated through Redis, so scrape every + instance. A restart resets the counters. +- **`Rebuilds` will not equal `Misses`.** Concurrent misses coalesce onto one build, so + `Misses - Rebuilds` is what coalescing saved; local writes rebuild off the read path and + push it the other way. In `Preset` mode `Rebuilds` starts at 1, from the build during + `Create`, and builds discarded after a generation conflict also count. Both are `uint64` + — check `Misses > Rebuilds` before subtracting, or a write-heavy instance wraps to + roughly 1.8e19. +- **One scanning call is one read, whatever it scans.** `FindParallel`, + `FindIndexParallel`, and `FindMany` load the automaton once per call, so each adds 1 to + `Hits`+`Misses` and their hit rate is comparable to a serial workload's. Writes, + `Suggest`, and `Info` never reach the automaton and add nothing. +- **`LastInvalidationLag` needs a listener and carries clock skew.** It is populated only + in `Preset` mode and in V2 with `EnableCache`; elsewhere zero means unavailable, not + fast. Where it is populated, the publish timestamp comes from another machine's clock, + so the value is the real delay plus that offset — it can understate as readily as + overstate. Watch it for step changes, and check NTP before blaming Pub/Sub. +- **A zero hit rate is not always a bug.** Without `Preset` or `EnableCache` every read + still checks Redis for freshness; a hit means the automaton was reused, not that the + round trip was skipped. -A failed reload retains the previous engine but returns an error to the search; -it does not serve that engine as a fallback. Each waiting request responds to its -own context cancellation without canceling other waiters. All waiters leaving or -`Close` cancels the shared job. Redis reads and engine builds run outside the -state lock, and snapshots overtaken by local state changes are rejected. +Not here: match counts, keyword counts, Redis latency. Keyword and node counts come from +`Info()`, which does read Redis. -Polling reads only the version field. It detects missed invalidations after a -successful poll; the next search must then fetch and build the full dictionary. -The configured interval is not an upper bound on staleness during failures. +### Preset refresh failures -### Reading the numbers +Track increases in `PresetReloadFailures` and `PresetPollFailures` alongside search errors. +A shared failed reload increments the first counter once even when many requests receive +the error; a failed version poll increments the second and retries on its next tick. +Cancellation is excluded from both, and both stay zero outside `Preset` mode. -- **The counters are per instance and per process.** Nothing is aggregated through - Redis, so in a fleet you scrape every instance. A restart resets them. -- **`Rebuilds` will not equal `Misses`.** Concurrent misses coalesce onto one build, so - `Misses - Rebuilds` is what that coalescing saved; local writes rebuild off the read - path and push the count the other way. In `Preset` mode `Rebuilds` starts at 1, from - the build during `Create`; builds discarded after a generation conflict also count. Both counters are `uint64`, so check `Misses > Rebuilds` - before subtracting — a write-heavy instance is routinely the other way round, and the - difference wraps to roughly 1.8e19 rather than going negative. -- **One scanning call is one read, whatever it scans over.** `FindParallel`, - `FindIndexParallel`, and `FindMany` load the automaton once per call and scan every - chunk or text against that snapshot, so each adds 1 to `Hits`+`Misses` and their hit - rate is directly comparable to a serial workload's. Calls that never reach the - automaton — writes, `Suggest`, `Info` — add nothing to either counter. -- **`LastInvalidationLag` needs a listener, and carries clock skew.** It is populated - only in `Preset` mode and in V2 with `EnableCache`; the other modes subscribe to - nothing, so a zero there means unavailable, not fast. Where it is populated, the - publish timestamp comes from another machine's clock, so the value is the real delay - plus that clock's offset — it can understate the delay as readily as overstate it, and - bounds it in neither direction. Watch it for step changes rather than trusting the - absolute value, and check NTP before concluding Pub/Sub is slow. -- **A zero hit rate is not always a bug.** Without `Preset` or `EnableCache` every read - still checks Redis for freshness; a hit there means only that the automaton was - reused, not that the round trip was skipped. -- **What is not here**: match counts, keyword counts, and Redis latency. Keyword and - node counts come from `Info()`, which does read Redis. - -## Service layer: `server/*` - -```mermaid -graph LR - A[ACOR] --> B[Metrics] - A --> C[Logs] - A --> D[Traces] - - B --> E[Prometheus] - C --> F[Log Aggregator] - D --> G[Jaeger/Zipkin] -``` +A failed reload keeps the previous engine but **returns an error to the search** rather +than serving that engine as a fallback. Each waiting request answers its own cancellation +without cancelling other waiters; all waiters leaving, or `Close`, cancels the shared job. +Redis reads and engine builds run outside the state lock, and snapshots overtaken by local +state changes are rejected. -### Metrics +Polling reads only the version field, and detects a missed invalidation only after a +successful poll — the next search then fetches and builds the full dictionary. The +configured interval is not an upper bound on staleness during failures. -Import the metrics package: +## Service layer + +### Metrics ```go import "github.com/skyoo2003/acor/server/metrics" ``` -#### Available Metrics - -| Metric | Type | Description | -| --------------------------------------- | --------- | ------------------------------------------- | -| `acor_http_requests_total` | Counter | Total HTTP requests by method, path, status | -| `acor_http_request_duration_seconds` | Histogram | HTTP request latency | -| `acor_redis_operations_total` | Counter | Total Redis operations by type, status | -| `acor_redis_operation_duration_seconds` | Histogram | Redis operation latency | -| `acor_keywords_total` | Gauge | Number of registered keywords | -| `acor_trie_nodes_total` | Gauge | Number of trie nodes | -| `grpc_server_handled_total` | Counter | Total gRPC requests by method, code | -| `grpc_server_handling_seconds` | Histogram | gRPC request latency | +| Metric | Type | Description | +| ------ | ---- | ----------- | +| `acor_http_requests_total` | Counter | HTTP requests by method, path, status | +| `acor_http_request_duration_seconds` | Histogram | HTTP request latency | +| `acor_redis_operations_total` | Counter | Redis operations by type, status | +| `acor_redis_operation_duration_seconds` | Histogram | Redis operation latency | +| `acor_keywords_total` | Gauge | Registered keywords | +| `acor_trie_nodes_total` | Gauge | Trie nodes | +| `grpc_server_handled_total` | Counter | gRPC requests by method, code | +| `grpc_server_handling_seconds` | Histogram | gRPC request latency | gRPC metrics use the standard `grpc_server_*` names from -`go-grpc-middleware/providers/prometheus`, wired via -`NewGRPCServerWithObservability`. +`go-grpc-middleware/providers/prometheus`, wired by `NewGRPCServerWithObservability`. -#### Exposing Metrics +Registering them does not expose them — serve `promhttp.Handler()` yourself: ```go import ( @@ -166,17 +131,9 @@ func main() { ### Logging -Import the logging package: - -```go -import "github.com/skyoo2003/acor/server/logging" -``` - -#### Structured Logging - -`NewLogger` takes an `io.Writer` and a level string (`debug`, `info`, `warn`, -`error`) and returns a zerolog-based logger that always emits structured JSON. -Attach trace/span IDs with `WithTraceID`: +`logging.NewLogger(w, level)` takes an `io.Writer` and one of `debug`, `info`, `warn`, +`error`, and returns a zerolog logger that always emits structured JSON. `WithTraceID` +attaches trace and span IDs. ```go @@ -205,23 +162,8 @@ func main() { } ``` -#### Log Levels - -- `debug`: Detailed debugging info -- `info`: General operational info -- `warn`: Warning conditions -- `error`: Error conditions - ### Tracing -Import the tracing package: - -```go -import "github.com/skyoo2003/acor/server/tracing" -``` - -#### OpenTelemetry Setup - ```go tracer, err := tracing.NewTracer(&tracing.Config{ @@ -236,36 +178,15 @@ if err != nil { defer tracer.Shutdown() ``` -#### Spans - -The core `pkg/acor` library does not emit its own spans. Request spans are -created by the `server/tracing` middleware for incoming traffic: - -- HTTP requests — via `tracing.HTTPMiddleware` -- gRPC calls — via the standard `otelgrpc` stats handler (`tracing.GRPCStatsHandler`) - -To trace individual `Add`/`Find`/`Remove` calls, wrap them in your own spans -using the OpenTelemetry API. - -### Dashboards - -#### Key Metrics to Monitor - -1. **Operation Latency**: P50, P95, P99 -2. **Error Rate**: Operations failing -3. **Keyword Count**: Collection size -4. **Redis Connections**: Pool utilization - -#### Grafana Dashboard - -Create a Grafana dashboard using the metrics above. Key panels to include: +`pkg/acor` emits no spans of its own. Request spans come from the middleware — +`tracing.HTTPMiddleware` for HTTP, the standard `otelgrpc` stats handler +(`tracing.GRPCStatsHandler`) for gRPC. To trace individual `Add`/`Find`/`Remove` calls, +wrap them in your own spans. -1. **Operation Latency**: P50/P95/P99 of `acor_redis_operation_duration_seconds` -2. **Error Rate**: Rate of `acor_redis_operations_total{status="error"}` -3. **Keyword Count**: Gauge `acor_keywords_total` -4. **Trie Nodes**: Gauge `acor_trie_nodes_total` +### Alerting -### Alerting Rules +Latency percentiles, error rate, keyword count, and pool utilization are the four worth a +dashboard panel. As rules: ```yaml groups: diff --git a/docs/content/operations/troubleshooting.md b/docs/content/operations/troubleshooting.md index faf634a1..74ab25a0 100644 --- a/docs/content/operations/troubleshooting.md +++ b/docs/content/operations/troubleshooting.md @@ -5,127 +5,28 @@ weight: 3 # Troubleshooting -Common issues and their solutions. +## Errors from `Create` and the matching API -## Common Errors +| Error | Cause | Fix | +| ----- | ----- | --- | +| `ErrRedisConflictingTopology` | More than one topology configured at once | Use exactly one of `Addr`, `Addrs` (+`MasterName` for Sentinel), or `RingAddrs` — see [Redis topologies](../../getting-started/quick-start/#redis-topologies) | +| `ErrRedisClusterDB` | Non-zero `DB` together with `Addrs` | Cluster has no database selection; drop `DB`, or use `Addr` for a single standalone server | +| `ErrEmptyKeyword` | Empty string passed to `Add` | Trim and reject before calling | +| `ErrInvalidChunkSize` | Non-positive `ParallelOptions.ChunkSize` | `ChunkSize` is required and must be > 0 | +| `ErrRedisAlreadyClosed` | Operation on a closed instance | `defer ac.Close()` once, at the owning function's exit | +| `ErrV1ReadOnly` | `Add`/`Remove` on a V1 collection | Migrate: `acor -name mycollection migrate` | +| `ErrSuggestRequiresRedis` | `Suggest` in `Preset` mode | Suggest needs the Redis path; use a non-preset instance for it | -### ErrRedisConflictingTopology +## Redis connection -**Cause:** Multiple Redis topologies specified simultaneously. +| Message | Check | +| ------- | ----- | +| `connection refused` | `redis-cli ping`, the address, the firewall, network reachability | +| `NOAUTH Authentication required` | Set `Password` | +| `context deadline exceeded` | Redis load, network latency, and the timeouts below | -**Solution:** Use only one configuration: - -```go -// Correct: Standalone -args := &acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-collection", -} - -// Correct: Sentinel -args := &acor.AhoCorasickArgs{ - Addrs: []string{"localhost:26379"}, - MasterName: "mymaster", - Name: "my-collection", -} - -// Wrong: Mixing configurations -args := &acor.AhoCorasickArgs{ - Addr: "localhost:6379", // Wrong! - Addrs: []string{"..."}, // Wrong! - MasterName: "mymaster", // Wrong! - Name: "my-collection", -} -``` - -### ErrEmptyKeyword - -**Cause:** Empty string passed to `Add()`. - -**Solution:** Validate input: - -```go -keyword := strings.TrimSpace(input) -if keyword == "" { - return errors.New("keyword cannot be empty") -} -_, err = ac.Add(keyword) -``` - -### ErrInvalidChunkSize - -**Cause:** Non-positive chunk size in parallel matching. - -**Solution:** Use positive values: - -```go -opts := &acor.ParallelOptions{ - Workers: 4, - ChunkSize: 1000, // Must be > 0 -} -``` - -### ErrRedisAlreadyClosed - -**Cause:** Operation on closed AhoCorasick instance. - -**Solution:** Ensure `Close()` is called only once, typically with `defer`: - -```go -ac, err := acor.Create(args) -if err != nil { - log.Fatal(err) -} -defer ac.Close() // Called once at function exit -``` - -## Redis Connection Issues - -### Connection Refused - -```text -redis GET on key "...": connection refused -``` - -**Checklist:** -1. Redis is running: `redis-cli ping` -2. Address is correct -3. Firewall allows connection -4. Network connectivity - -### Authentication Failed - -```text -redis GET on key "...": NOAUTH Authentication required -``` - -**Solution:** Provide password: - -```go -args := &acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Password: "your-password", - Name: "my-collection", -} -``` - -### Timeout Errors - -```text -redis GET on key "...": context deadline exceeded -``` - -**Solutions:** -1. Tune `DialTimeout`, `ReadTimeout`, or `WriteTimeout` when measurements show - the defaults are too short -2. Check Redis load -3. Check network latency -4. Set `MaxRetries` for transient failures or adjust `PoolSize` for measured - connection contention -5. Scale Redis cluster - -Zero values keep the go-redis defaults across Standalone, Sentinel, Cluster, -and Ring modes. +Zero values keep the go-redis defaults across every topology. Tune from measurements, not +from guesses: ```go @@ -140,64 +41,44 @@ args := &acor.AhoCorasickArgs{ _ = args ``` -### Preset Cache Appears Stale +`PoolSize` is the knob for measured connection contention. -**Cause:** Preset mode normally reloads through best-effort Redis Pub/Sub. A -disconnected subscriber can miss an invalidation. +## Preset cache looks stale -**Solution:** In multi-instance deployments, set -`InvalidationPollInterval` to retry version checks periodically: +Preset mode reloads through best-effort Pub/Sub, and a disconnected subscriber misses an +invalidation. In multi-instance deployments, enable polling: ```go args.InvalidationPollInterval = 30 * time.Second ``` -The option is disabled by default and ignored outside Preset mode. The interval -is not a freshness bound: recovery requires a successful version poll and a -successful reload on the next search. Inspect `CacheStats().PresetPollFailures` -and `PresetReloadFailures` when updates remain invisible. Reload errors are -returned to searches; the retained engine is not automatically served as fallback. - -## Performance Issues - -### Slow Find Operations +Disabled by default, and ignored outside `Preset` mode. **The interval is not a freshness +bound** — recovery needs a successful version poll *and* a successful reload on the next +search. When updates stay invisible, check `CacheStats().PresetPollFailures` and +`PresetReloadFailures`; reload errors are returned to searches rather than silently served +from the retained engine. See +[invalidation safety](../../guides/redis-backed-engine/#invalidation-safety). -**Diagnostic:** -1. Check schema version: `acor -name collection schema-version` -2. Check collection size: `acor -name collection info` +## Slow reads or high memory -**Solutions:** -- Migrate to V2 schema -- Use parallel matching for large texts -- Increase Redis memory - -### High Memory Usage - -**Diagnostic:** -1. Check Redis memory: `redis-cli info memory` -2. Check keyword count: `acor -name collection info` +```bash +acor -name mycollection schema-version # V1 is the usual answer to "why is Find slow" +acor -name mycollection info # keyword and node counts +redis-cli info memory +``` -**Solutions:** -- Remove unused keywords -- Use V2 schema (lower memory) -- Scale Redis cluster +- On V1, migrate to V2 — see [Schema V2](../../reference/schema-v2/). +- On V2, a read-heavy workload wants `EnableCache` or a `Preset`; the schema alone does + not make reads fast ([benchmarks](../../reference/benchmarks/#what-the-numbers-mean)). +- For large texts, use [parallel matching](../../guides/parallel-matching/). +- For high memory, remove unused keywords or move to `PresetMemoryEfficient`. ## Debugging -### Enable Debug Logging - -```go -logger := logging.NewLogger(os.Stdout, "debug") -``` - -### CLI Debug Mode - ```bash -acor -name mycollection -debug find "test text" +acor -name mycollection -debug find "test text" # CLI debug logging +redis-cli keys "{mycollection}:*" # what the collection actually stores ``` -### Check Redis Keys - -```bash -redis-cli keys "{mycollection}:*" -``` +In library code, set `Debug: true` for the default stdout logger, or supply your own +`Logger`. diff --git a/docs/content/reference/_index.md b/docs/content/reference/_index.md index eb09bea7..e9ef0a39 100644 --- a/docs/content/reference/_index.md +++ b/docs/content/reference/_index.md @@ -5,20 +5,15 @@ weight: 3 # Reference -Technical reference documentation for ACOR. +- [API Reference](api/) — constructors, matching, batch, parallel, and the option types +- [Compatibility](compatibility/) — what the `v1` line promises, and what it excludes +- [Schema V2](schema-v2/) — the current storage layout +- [Schema V1](schema-v1/) — deprecated, read-only +- [Versioned dictionaries (V3)](versioned/) — leased snapshots, expected-version writes, cutover +- [Bounded text processing](text-processing/) — `Scan`, `MaskText`, `ReplaceText` over source positions -## Sections +## Measurements -- [API Reference](api/) - Public API documentation (unified `Create` API with `Preset` options) -- [Compatibility](compatibility/) - What the `v1` line promises, and what it excludes -- [Schema V1](schema-v1/) - Legacy schema details -- [Schema V2](schema-v2/) - Optimized schema (recommended) -- [Versioned dictionaries (V3)](versioned/) - Leased snapshots, expected-version writes, and cutover -- [Bounded search, masking and replacement](text-processing/) - `Scan`, `MaskText`, and `ReplaceText` over original source positions -- [Benchmarks](benchmarks/) - Measured performance and how to reproduce it -- [V3 performance report](versioned-performance/) - Million-keyword measurements on Redis and Valkey -- [R2/R3 verification](r2-r3-performance/) - Boundary protection and bounded-processing evidence - -## Navigation - -← [Guides](../guides/) | [Operations](../operations/) → +- [Benchmarks](benchmarks/) — round trips and timings, with the commands that produce them +- [V3 performance report](versioned-performance/) — the archived R1 million-keyword baseline +- [R2/R3 verification](r2-r3-performance/) — incremental download, engine memory, bounded APIs diff --git a/docs/content/reference/api.md b/docs/content/reference/api.md index 46e3a6c8..cf8c0dee 100644 --- a/docs/content/reference/api.md +++ b/docs/content/reference/api.md @@ -5,15 +5,37 @@ weight: 1 # API Reference -Core API documentation for ACOR. See -[pkg.go.dev](https://pkg.go.dev/github.com/skyoo2003/acor/pkg/acor) for the -complete generated reference. +Contracts and behavior for `pkg/acor`. Generated signatures live on +[pkg.go.dev](https://pkg.go.dev/github.com/skyoo2003/acor/pkg/acor). -## Core Types +## Creating a collection -### AhoCorasickArgs +```go +ac, err := acor.Create(&acor.AhoCorasickArgs{...}) +defer ac.Close() +``` -Configuration for creating an AhoCorasick instance. +`CreateContext` is the same constructor with a context bounding the setup I/O — the +schema check, the initialization write, and the initial keyword load: + + +```go +setupCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) +defer cancel() +instance, err := acor.CreateContext(setupCtx, &acor.AhoCorasickArgs{ + Addr: "localhost:6379", + Name: "default", +}) +_ = instance +_ = err +``` + +That context bounds construction only. Cancelling it later neither closes the instance +nor stops its invalidation listener — use `Close` for that and the `*Context` methods for +per-operation cancellation. The Pub/Sub subscribe belongs to the listener, so it runs on +the instance's own context. + +### AhoCorasickArgs ```go @@ -43,73 +65,25 @@ type AhoCorasickArgs struct { ``` -### AhoCorasick - -Main type for pattern matching operations. - -```go -ac, err := acor.Create(&acor.AhoCorasickArgs{...}) -defer ac.Close() -``` - -`CreateContext` is the same constructor with a context bounding the setup I/O -(schema check and initialization write, initial keyword load): - - -```go -setupCtx, cancel := context.WithTimeout(context.Background(), 5*time.Second) -defer cancel() -instance, err := acor.CreateContext(setupCtx, &acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "default", -}) -_ = instance -_ = err -``` - -The context bounds construction only. Canceling it afterwards does not close the -instance or stop its invalidation listener — use `Close` for that, and the -`*Context` methods for per-operation cancellation. The Pub/Sub subscribe is part -of that listener, so it runs on the instance's own context rather than this one. - -## Core Methods - -### Add - -Add a single keyword to the collection. - -```go -count, err := ac.Add("keyword") -``` - -### AddMany - -Add multiple keywords in a batch. - -```go -result, err := ac.AddMany([]string{"a", "b", "c"}, nil) -// or with options: -result, err := ac.AddMany([]string{"a", "b", "c"}, &acor.BatchOptions{ - Mode: acor.BatchModeTransactional, -}) -``` - -### Remove +Setting `Preset` switches reads to a local automaton — see +[Preset-Optimized Engine](../../guides/preset-engine/). Topology fields are covered in +[Redis topologies](../../getting-started/quick-start/#redis-topologies). -Remove a single keyword from the collection. - -```go -count, err := ac.Remove("keyword") -``` +## Writing -### RemoveMany - -Remove multiple keywords in a batch. +| Method | Returns | Notes | +| ------ | ------- | ----- | +| `Add(keyword)` | `(int, error)` | 1 if the collection changed, 0 if it was already there | +| `Remove(keyword)` | `(int, error)` | Same convention | +| `AddMany(keywords, *BatchOptions)` | `(*BatchResult, error)` | One transaction; `nil` options mean best-effort | +| `RemoveMany(keywords, *BatchOptions)` | `(*BatchResult, error)` | Same | +| `Flush()` | `error` | Deletes every key in the collection | +| `Close()` | `error` | Closes the connection and stops the invalidation listener | ```go result, err := ac.RemoveMany([]string{"a", "b"}, nil) -// Pass options when transactional behavior is required. +// Pass options when the whole batch must land or fail together. result, err = ac.RemoveMany([]string{"c", "d"}, &acor.BatchOptions{ Mode: acor.BatchModeTransactional, }) @@ -117,29 +91,29 @@ _ = result _ = err ``` -### Find - -Find all matching keywords in text. +See [Batch Operations](../../guides/batch-operations/) for the modes and result shape. -```go -matches, err := ac.Find("sample text") -// Returns: []string{"match1", "match2", ...} -``` +## Matching -### FindIndex +| Method | Returns | Shape | +| ------ | ------- | ----- | +| `Find(text)` | `[]string` | Which keywords occur | +| `FindIndex(text)` | `map[string][]int` | Keyword to its rune start positions | +| `FindSet(text)` | `[]string` | Each keyword once, in first-match order | +| `FindMatches(text, *MatchOptions)` | `[]Match` | Every occurrence with its rune span | +| `Contains(text)` | `bool` | Stops at the first match | +| `FindStream(reader, callback)` | `error` | Scans an `io.Reader` without buffering it | +| `FindMany(texts)` | `map[string][]string` | Keyed by input text | +| `FindParallel(text, *ParallelOptions)` | `[]string` | Chunked across workers | +| `FindIndexParallel(text, *ParallelOptions)` | `map[string][]int` | Same chunking, with positions | -Find matches with their start positions. - -```go -positions, err := ac.FindIndex("sample text") -// Returns: map[string][]int{"keyword": {startPos, ...}, ...} -``` +All positions are **rune** offsets, not byte offsets. ### FindMatches -Return every occurrence in scan order with its keyword and half-open rune span -`[Start, End)`. The default includes overlapping matches; use -`MatchKindLeftmostLongest` for non-overlapping tokenization or replacement. +Every occurrence in scan order with its half-open rune span `[Start, End)`. The default +includes overlaps; `MatchKindLeftmostLongest` produces non-overlapping spans suitable for +tokenizing or replacing. ```go @@ -170,22 +144,14 @@ const ( ) ``` -`WholeWord` uses letters, digits, combining marks, and underscores as word -runes. Set `WordRune` when those defaults do not fit the input script. - -### Contains - -Report whether any keyword occurs, stopping at the first match. - -```go -found, err := ac.Contains("sample text") -``` +`WholeWord` treats letters, digits, combining marks, and underscores as word runes. Set +`WordRune` for scripts those defaults do not fit — in CJK or Thai text every adjacent +character counts as a word rune, so nearly every match is dropped as mid-word. ### FindStream -Scan an `io.Reader` without buffering the whole input. Matches include -overlaps, retain rune offsets across reads, and arrive in scan order. Returning -`false` from the callback stops the scan. +Matches keep rune offsets across reads and arrive in scan order; returning `false` from +the callback stops the scan. ```go @@ -196,83 +162,62 @@ err := ac.FindStream(strings.NewReader("sample text"), func(match acor.Match) bo _ = err ``` -Streaming does not apply whole-word or leftmost-longest filtering because -those modes require buffering. Use `FindMatches` for bounded strings that need -those options. - -### FindMany - -Find matches in multiple texts. - -```go -matches, err := ac.FindMany([]string{"text1", "text2"}) -// Returns: map[string][]string{"text1": {"kw", ...}, ...} (keyed by input text) -``` - -### FindParallel +Streaming always overlaps: whole-word and leftmost-longest need buffering, so use +`FindMatches` on a bounded string when you need them. -Find matches using parallel processing. +### Parallel ```go -matches, err := ac.FindParallel(largeText, &acor.ParallelOptions{ - Workers: 4, - Boundary: acor.ChunkBoundaryWord, -}) -``` - -Keywords longer than `ParallelOptions.Overlap` can be missed at chunk -boundaries. Set `Overlap` to at least the longest expected keyword. - -### FindIndexParallel - -Find start positions using the same parallel chunking options. - -```go -positions, err := ac.FindIndexParallel(largeText, acor.DefaultParallelOptions()) -``` - -### Info - -Get collection statistics. +type ParallelOptions struct { + Workers int // Concurrent goroutines (default: runtime.NumCPU()) + ChunkSize int // Target chunk size in runes (required; no fallback) + Boundary ChunkBoundary // How chunks are split (default: ChunkBoundaryWord) + Overlap int // Overlap runes between chunks (unset means zero) + AutoOverlap bool // Extend each chunk by the dictionary's longest keyword (default: false) +} -```go -info, err := ac.Info() -// Returns: &AhoCorasickInfo{Keywords: N, Nodes: M, Preset: ..., MemoryBytes: ..., TrieDepth: ...} +const ( + ChunkBoundaryWord ChunkBoundary = iota // Split at whitespace (default) + ChunkBoundarySentence // Split at . ! ? + ChunkBoundaryLine // Split at newlines +) ``` -### CacheStats +`DefaultParallelOptions()` returns a usable starting point. Without `AutoOverlap`, a +keyword longer than `Overlap` can be missed at a chunk boundary; with it, chunks extend +by the dictionary's longest keyword at no extra Redis cost. Details: +[Parallel Matching](../../guides/parallel-matching/). -Get local cache statistics. Unlike `Info`, this performs no Redis I/O, so it is cheap -enough to scrape on a timer. +## Suggest -```go -stats := ac.CacheStats() -// Returns: CacheStats{Hits: N, Misses: M, Rebuilds: R, RebuildDuration: ..., LastInvalidationLag: ...} -``` - -The counters are per instance and per process — scrape every instance in a fleet. See -[Monitoring](../../operations/monitoring/) for how to read them, including why -`Rebuilds` does not equal `Misses` and why `LastInvalidationLag` carries clock skew. +`Suggest(prefix)` returns keywords starting with the prefix; `SuggestIndex(prefix)` +returns the same keys mapped to `[0]`, since a prefix match always starts at the +beginning. Both require Redis and are unavailable in `Preset` mode +(`ErrSuggestRequiresRedis`). -### Flush - -Clear all data from the collection. +## Batch types ```go -err := ac.Flush() -``` - -### Close +type BatchOptions struct { + Mode BatchMode // BatchModeBestEffort (default) or BatchModeTransactional +} -Close the Redis connection. +type BatchResult struct { + Added []string // Successfully added keywords + Removed []string // Successfully removed keywords + Failed []KeywordError // Keywords that failed, with their errors + Skipped []string // Duplicate adds or absent removes +} -```go -err := ac.Close() +type KeywordError struct { + Keyword string + Error error +} ``` -### AhoCorasickInfo +## Statistics -Statistics about an Aho-Corasick instance. +`Info()` reads Redis; `CacheStats()` does not, so it is cheap to scrape on a timer. ```go @@ -286,11 +231,6 @@ type AhoCorasickInfo struct { ``` -### CacheStats (type) - -A snapshot of one instance's local cache activity. Returned by `CacheStats()`, never -constructed by callers — fields may be added inside `v1`. - ```go type CacheStats struct { PresetReloadFailures uint64 // Failed shared reload jobs, once per job (Preset only; cancellation excluded) @@ -303,190 +243,40 @@ type CacheStats struct { } ``` -### Preset +`CacheStats` is returned by ACOR and never constructed by callers, so fields may be added +inside `v1`. Counters are per instance and per process — scrape every instance in a fleet. +How to read them, including why `Rebuilds` does not equal `Misses`: +[Monitoring](../../operations/monitoring/). -Architecture presets for the preset-optimized Redis engine. +## Presets ```go const ( - PresetNone Preset = iota // Zero value (unset) — falls through to original V1/V2 mode + PresetNone Preset = iota // Zero value — original V1/V2 mode PresetSpeed // Full DFA + flat array — max speed, higher memory PresetBalanced // Double-Array Trie + Banded DFA — best speed-to-memory ratio PresetMemoryEfficient // Map-based + Bloom filter — min memory, slower search ) ``` -## Redis-Backed Engine with Presets - -Redis-backed Aho-Corasick that combines Redis persistence with a local preset-optimized automaton. Writes go to Redis atomically (V2 Lua scripts with optimistic locking); reads hit the local engine with no Redis I/O. Created via the unified `Create` API with `Preset` set. - -```go -ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-collection", - Preset: acor.PresetBalanced, - CaseSensitive: false, -}) -defer ac.Close() -``` - -### AhoCorasickArgs (Preset field) - -The `AhoCorasickArgs` struct includes a `Preset` field for the engine mode: - -```go -type AhoCorasickArgs struct { - // ... standard Redis connection fields ... - Preset Preset // Architecture preset: PresetSpeed, PresetBalanced, PresetMemoryEfficient - // ... other fields ... -} -``` - -### Preset-Optimized Redis Methods - -```go -// Create -ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-collection", - Preset: acor.PresetBalanced, -}) - -// Add/Remove -added, err := ac.Add("keyword") // (int, error) -removed, err := ac.Remove("keyword") // (int, error) - -// Find (0 RTT on hot path — reads from local engine) -matches, err := ac.Find("text") // ([]string, error) -positions, err := ac.FindIndex("text") // (map[string][]int, error) -spans, err := ac.FindMatches("text", nil) // ([]Match, error) -found, err := ac.Contains("text") // (bool, error) - -// Info -info, err := ac.Info() // (*AhoCorasickInfo, error) - -// Flush -err := ac.Flush() - -// Close -err := ac.Close() -``` - -## Context Variants +## Context variants -Operations that may perform Redis I/O also accept an explicit -`context.Context`: `AddContext`, `RemoveContext`, `FindContext`, -`FindIndexContext`, `FindMatchesContext`, `ContainsContext`, -`FindStreamContext`, `FlushContext`, `InfoContext`, `SuggestContext`, -`SuggestIndexContext`, `AddManyContext`, `RemoveManyContext`, -`FindManyContext`, `FindParallelContext`, and `FindIndexParallelContext`. +Every method that may touch Redis has a `*Context` twin taking an explicit +`context.Context`: `AddContext`, `RemoveContext`, `FindContext`, `FindIndexContext`, +`FindSetContext`, `FindMatchesContext`, `ContainsContext`, `FindStreamContext`, +`FlushContext`, `InfoContext`, `SuggestContext`, `SuggestIndexContext`, +`AddManyContext`, `RemoveManyContext`, `FindManyContext`, `FindParallelContext`, and +`FindIndexParallelContext`. ```go matches, err := ac.FindMatchesContext(ctx, text, nil) ``` -## Suggest Methods - -### Suggest - -Get prefix suggestions. - -```go -suggestions, err := ac.Suggest("pre") -``` - -### SuggestIndex - -Get suggestions with positions. - -```go -positions, err := ac.SuggestIndex("pre") -``` - -## Batch Operations - -### BatchOptions - -```go -type BatchOptions struct { - Mode BatchMode // BestEffort (default) or Transactional -} -``` - -### BatchResult - -```go -type BatchResult struct { - Added []string // Successfully added keywords - Removed []string // Successfully removed keywords - Failed []KeywordError // Keywords that failed with their errors - Skipped []string // Duplicate adds or absent removes -} -``` - -### KeywordError - -```go -type KeywordError struct { - Keyword string - Error error -} -``` - -## Parallel Options - -### ParallelOptions - -```go -type ParallelOptions struct { - Workers int // Concurrent goroutines (default: runtime.NumCPU()) - ChunkSize int // Target chunk size in characters (required; no fallback) - Boundary ChunkBoundary // How chunks are split (default: ChunkBoundaryWord) - Overlap int // Overlap characters between chunks (unset means zero) - AutoOverlap bool // Extend each chunk by the dictionary's longest keyword (default: false) -} -``` - -### DefaultParallelOptions - -Returns parallel options with sensible defaults: - -```go -opts := acor.DefaultParallelOptions() -matches, err := ac.FindParallel(text, opts) -``` - -### ChunkBoundary - -```go -const ( - ChunkBoundaryWord ChunkBoundary = iota // Split at whitespace (default) - ChunkBoundarySentence // Split at sentence boundaries (. ! ?) - ChunkBoundaryLine // Split at newlines -) -``` - -## Parallel Boundary Protection and Preset Refresh - -`ParallelOptions.AutoOverlap` (default `false`) enables dictionary-aware right -extensions for each base chunk. It protects keywords longer than `ChunkSize` -without additional Redis reads. Results use the existing parallel order and rune -positions. See [parallel matching](../../guides/parallel-matching/). - -`CacheStats.PresetReloadFailures` and `CacheStats.PresetPollFailures` are cumulative -`uint64` counters. Shared reload failures count once per job; cancellations are -excluded. They remain zero in other modes. See [Preset refresh](../../guides/redis-backed-engine/#invalidation-safety). - -## Versioned dictionaries - -`OpenVersioned(ctx, *VersionedOptions)` opens a separate V3 collection. Its -`VersionedCollection` API provides leased snapshots, paginated list/diff, expected-version -replace/add/remove, asynchronous engine refresh, operation receipts, V2 copying -and explicit pruning. See [the V3 API guide](../versioned/) for contracts and a -compilable example; V3 uses the existing module version and search semantics. - -## Bounded source-position APIs +## Beyond V1/V2 -`Scan`, `MaskText`, and `ReplaceText` are available on AhoCorasick and -VersionedCollection with explicit contexts. See [bounded text processing](../text-processing/) -for ScanOptions, RewriteOptions, original byte/rune spans and error contracts. +- **`OpenVersioned(ctx, *VersionedOptions)`** opens a separate V3 collection with leased + snapshots, paginated list/diff, expected-version writes, background engine refresh, and + operation receipts. See [Versioned dictionaries](../versioned/). +- **`Scan`, `MaskText`, `ReplaceText`** report original byte and rune spans under explicit + work limits, on both `AhoCorasick` and `VersionedCollection`. See + [bounded text processing](../text-processing/). diff --git a/docs/content/reference/benchmarks.md b/docs/content/reference/benchmarks.md index 152c0e70..a5f7ea21 100644 --- a/docs/content/reference/benchmarks.md +++ b/docs/content/reference/benchmarks.md @@ -5,21 +5,19 @@ weight: 5 # Benchmarks -Every performance number ACOR publishes is on this page, with the command that -produces it. Nothing here is an estimate. +Every performance number ACOR publishes is here, with the command that produces it. +Nothing is an estimate. -The evidence comes in two kinds, and they are not interchangeable: +The evidence comes in two kinds: -- **Round trips** are structural. They are counted at the storage seam, so they - are identical on miniredis and on a real server, and no hardware difference - can move them. They are enforced by tests on every CI run. -- **Timings** are hardware-bound. Absolute nanoseconds on your machine will not - match ours. The reproducible quantity is the **ratio** between configurations. +- **Round trips** are structural — counted at the storage seam, identical on miniredis and + on a real server, and enforced by tests on every CI run. +- **Timings** are hardware-bound. Absolute nanoseconds will not match yours. The + reproducible quantity is the **ratio** between configurations. ## Round trips per operation -Enforced by `TestRTT*` in `pkg/acor/rtt_claims_test.go`. If any of these change, -CI fails. +Enforced by `TestRTT*` in `pkg/acor/rtt_claims_test.go`; a change fails CI. | Operation | V1 | V2 | |---|---|---| @@ -31,15 +29,14 @@ CI fails. | `Add()`, 5-character keyword | 53 | 2 | | `Add()`, 26-character keyword | 507 | 2 | -Both schemas read in a single round trip: V1 issues one `SMEMBERS`, V2 pipelines -two `HGETALL` calls into one trip. V1's round-trip cost is on **writes**, where it -walks the trie node by node, so the cost grows with the length of the keyword being -added rather than with the size of the dictionary. +Both schemas read in one round trip: V1 issues one `SMEMBERS`, V2 pipelines two `HGETALL` +calls into one trip. V1's cost is on **writes**, where it walks the trie node by node — so +it grows with the length of the keyword being added, not with the dictionary. -The multi-scan reads cost the same as one `Find()`, and the count does not grow with -the chunk count or the batch size: the automaton is loaded once per call and every -chunk or text is scanned against that one snapshot. Before `v1.5.0` each chunk loaded -its own, so a 63-chunk text issued 63 reads. +Multi-scan reads cost the same as one `Find()`, and do not grow with chunk count or batch +size: the automaton is loaded once per call and every chunk or text is scanned against +that snapshot. Before `v1.5.0` each chunk loaded its own, so a 63-chunk text issued 63 +reads. ```sh go test -run RTT ./pkg/acor @@ -50,9 +47,8 @@ ACOR_INTEGRATION_ADDR=localhost:6379 go test -run RTT ./pkg/acor ## Timings -Measured on: Apple M4, `darwin/arm64`, Go 1.26, Redis 8 on loopback, -`-benchtime=200x`. The exact command that produced the tables below, which takes -about 20 seconds: +Apple M4, `darwin/arm64`, Go 1.26, Redis 8 on loopback, `-benchtime=200x`. This produces +the tables below in about 20 seconds: ```sh redis-server --port 6379 --save "" --daemonize yes @@ -60,8 +56,8 @@ ACOR_INTEGRATION_ADDR=localhost:6379 \ go test -bench RealServer -benchmem -benchtime=200x -run '^$' ./pkg/acor ``` -`make bench` runs the full sweep including the miniredis benchmarks. That takes -several minutes, and its timings are not published; see the caveats below. +`make bench` runs the full sweep including the miniredis benchmarks — several minutes, and +its timings are not published (see [Caveats](#caveats)). ### Find, 1000 keywords @@ -81,8 +77,8 @@ several minutes, and its timings are not published; see the caveats below. | V2 + `EnableCache`, warm | 3,120 | 2,857 | 10 | ~29x faster | | `PresetBalanced` | 2,214 | 2,048 | 4 | ~41x faster | -At 100 keywords V1 and V2 are close enough that the winner changes between runs; -only the 1000-keyword gap is stable. +At 100 keywords V1 and V2 are close enough that the winner changes between runs; only the +1000-keyword gap is stable. ### Add @@ -93,10 +89,8 @@ only the 1000-keyword gap is stable. ### Bulk load, `AddMany` -`AddMany` plans the whole batch in one pass and commits it in a single -transaction, so it costs two round trips regardless of batch size. - -On the Apple M4 with Redis 8 on loopback setup above, one sample measured: +`AddMany` plans the whole batch in one pass and commits it in a single transaction, so it +costs two round trips regardless of batch size. One sample on the setup above: | Keywords | ns/op | B/op | allocs/op | |---|---|---|---| @@ -107,49 +101,40 @@ Reproduce with `ACOR_INTEGRATION_ADDR=localhost:6379 make bench-module`. ### Run-to-run variance -The `ns/op` columns are one run. Repeat runs on the same idle laptop moved every -absolute number by 20-25% while the ratios held to within about 15%. - -That is why the ratios are stated approximately and the raw numbers are labelled a -sample. Absolute figures differing from ours is expected; *ratios* differing -substantially is worth reporting as an issue. - -## What these numbers mean +The `ns/op` columns are one run. Repeat runs on the same idle laptop moved every absolute +number by 20-25% while the ratios held to within about 15% — which is why ratios are +stated approximately and raw numbers are labelled a sample. Absolute figures differing +from ours is expected; *ratios* differing substantially is worth an issue. -**V2 without caching is still somewhat slower than V1 on reads**, because of -payload rather than round trips. Both cost one trip, but V1's `SMEMBERS` returns -just the keyword set while V2 must read an outputs hash carrying one entry per -trie state. The gap widens with the dictionary: roughly parity at 100 keywords, -~1.7x at 1000. +## What the numbers mean -Both schemas memoize the automaton, so an unchanged collection is not re-parsed or -rebuilt between reads. Uncached V2 was ~9x slower than V1 before that memoization -landed; what remains is inherent to reading the whole outputs hash, and -`EnableCache` is the fix for it. +**V2 without caching is somewhat slower than V1 on reads**, because of payload rather than +round trips: both cost one trip, but V1's `SMEMBERS` returns just the keyword set while V2 +must read an outputs hash carrying one entry per trie state. Roughly parity at 100 +keywords, ~1.7x at 1000. Both schemas memoize the automaton, so an unchanged collection is +not re-parsed between reads; uncached V2 was ~9x slower before that memoization, and what +remains is inherent to reading the whole outputs hash. -**The large read speedups come from caching, not from the schema.** 15x to 59x -belongs to `EnableCache` and the `Preset` engines. Choosing V2 and nothing else -does not deliver them. +**The large read speedups come from caching, not from the schema.** 15x to 59x belongs to +`EnableCache` and the `Preset` engines. Choosing V2 alone does not deliver them. -**V2's unambiguous win is writes.** ~14x on `Add()`, and it is the only schema -that supports caching or preset engines at all. Use `AddMany` rather than a loop -over `Add`: it commits in a single transaction. In the Apple M4/Redis 8 loopback -sample, 1,000 keywords cost 3.0 ms instead of the ~350 ms the same writes cost one +**V2's unambiguous win is writes** — ~14x on `Add()`, and it is the only schema supporting +caching or preset engines at all. Use `AddMany` rather than a loop over `Add`: in the +sample above, 1,000 keywords cost 3.0 ms instead of the ~350 ms the same writes cost one at a time. -Practical reading: choose V2, and enable `EnableCache` or a `Preset` if your -workload is read-heavy. V2 with neither is the one configuration these numbers -do not recommend. +Practical reading: choose V2, and enable `EnableCache` or a `Preset` if reads dominate. V2 +with neither is the one configuration these numbers do not recommend. ## Caveats -- Loopback Redis has almost no network latency. Over a real network both schemas - still pay one round trip on `Find()`, so V2's larger payload matters more rather - than less, and V1's per-node `Add()` cost grows with the added latency. +- Loopback Redis has almost no network latency. Over a real network both schemas still pay + one round trip on `Find()`, so V2's larger payload matters more rather than less, and + V1's per-node `Add()` cost grows with the added latency. - These figures compare ACOR configurations against each other, not against other - Aho-Corasick implementations. A single process with a static dictionary is better - served by an in-memory library. ACOR earns its cost when several instances share - one dictionary that changes at runtime. -- Every other benchmark in the repository runs on miniredis, an in-process - emulator with no round-trip cost. Those exist for regression detection and are - deliberately not published here. + Aho-Corasick implementations. A single process with a static dictionary is better served + by an in-memory library; ACOR earns its cost when several instances share one dictionary + that changes at runtime. +- Every other benchmark in the repository runs on miniredis, an in-process emulator with + no round-trip cost. Those exist for regression detection and are deliberately not + published here. diff --git a/docs/content/reference/compatibility.md b/docs/content/reference/compatibility.md index 3efc17c4..76876b35 100644 --- a/docs/content/reference/compatibility.md +++ b/docs/content/reference/compatibility.md @@ -5,73 +5,65 @@ weight: 2 # Compatibility -Code that imports `github.com/skyoo2003/acor/pkg/acor` keeps compiling, and keeps -behaving as documented, across every `v1.x.y` release. Nothing is removed from that -surface inside `v1`. Taking something back requires a `v2` import path, and no `v2` is -scheduled. - -`v1.5.0` is the first supported `v1` release, and the promise is measured from its -surface. `v1.0.0`-`v1.4.0` are retracted and were never covered — see -[Retracted versions](https://github.com/skyoo2003/acor/blob/main/RELEASE.md#retracted-versions) -for why the numbering starts there. - -**Upgrading from `v0.11.x` is a one-time exception.** `v1.5.0` removes deprecated members -and closes paths that `v0.11.x` still allowed, so that upgrade can require changes to your -code. From `v0.10.x` or earlier, `v0.11.0`'s own breaking changes apply first — the -`RemoveMany` signature, the removal of `RemoveManyWithOptions`, and the removal of the -exported V1 key constants — so upgrade through `v0.11.x` and follow the -[changelog](https://github.com/skyoo2003/acor/blob/main/CHANGELOG.md); this page does not -restate them. +Code that imports `github.com/skyoo2003/acor/pkg/acor` keeps compiling, and keeps behaving +as documented, across every `v1.x.y` release. Nothing is removed from that surface inside +`v1`. Taking something back requires a `v2` import path, and no `v2` is scheduled. + +`v1.5.0` is the first supported `v1` release and the baseline the promise is measured +from. `v1.0.0`–`v1.4.0` are retracted and were never covered — see +[Retracted versions](https://github.com/skyoo2003/acor/blob/main/RELEASE.md#retracted-versions). + +## Upgrading from `v0.11.x` is the one exception -From `v0.11.x`, specifically: +`v1.5.0` removes deprecated members and closes paths `v0.11.x` still allowed, so that one +upgrade can require changes to your code: -- `PresetUltimate` is removed. It was an alias for `PresetBalanced`; rename it. -- `InMemoryInfo` is removed. No exported function accepted or returned it. +- **`PresetUltimate` is removed.** It aliased `PresetBalanced`; rename it. +- **`InMemoryInfo` is removed.** No exported function accepted or returned it. - **V1 collections are read-only.** `Add` and `Remove` return `ErrV1ReadOnly`. Reads, `Suggest`, `Info`, and `MigrateV1ToV2` still work, so an existing V1 collection can be - read and converted, but it takes no new keywords. Migrate to V2. `Flush` also still - works — and still deletes every key in the collection. Read-only means keyword writes - are refused, not that the collection cannot be destroyed. + read and converted, but it takes no new keywords. `Flush` also still works — and still + deletes every key. Read-only refuses keyword writes; it does not make the collection + indestructible. -Every upgrade *after* `v1.5.0`, within `v1`, is covered by everything below. Each of -those three would be a promise violation in `v1.6.0`; `v1.5.0` is the only release that -can make them. +From `v0.10.x` or earlier, `v0.11.0`'s own breaking changes apply first — the `RemoveMany` +signature, the removal of `RemoveManyWithOptions`, and the removal of the exported V1 key +constants. Upgrade through `v0.11.x` and follow the +[changelog](https://github.com/skyoo2003/acor/blob/main/CHANGELOG.md); this page does not +restate them. + +Every upgrade *after* `v1.5.0` is covered by everything below. Each of those three changes +would be a promise violation in `v1.6.0`; `v1.5.0` is the only release that could make +them. ## What is covered -| Surface | Covered by `v1` | -| ---------------------------------------- | ----------------- | -| Exported identifiers of `pkg/acor` | ✅ | -| Documented behavior of those identifiers | ✅ | -| Sentinel error identity (`errors.Is`) | ✅ | -| On-Redis V2 data format | ✅ additions only | -| Everything in the next section | ❌ | - -## What is not covered - -- **The CLI.** Flags, output format, and exit codes of the `acor` command can change in - any release. If you need a stable contract, call the library rather than parsing CLI - output. -- **`acor/server`.** A separate, experimental module. It publishes no tags of its own, so - it can only be required by pseudo-version, and the core's version numbers say nothing - about it. -- **`internal/...`.** Not importable, and free to change in any release. -- **The `benchmarks` module.** A measurement harness, not an API. -- **Documentation wording.** Pages get rewritten. What they describe is pinned by this - page, not by their phrasing. -- **The V1 schema layout.** Deprecated, and not evolved further. See - [Deprecation](#deprecation). +| Surface | Covered by `v1` | +| ------- | --------------- | +| Exported identifiers of `pkg/acor` | ✅ | +| Documented behavior of those identifiers | ✅ | +| Sentinel error identity (`errors.Is`) | ✅ | +| On-Redis V2 data format | ✅ additions only | + +Not covered: + +| Surface | Why | +| ------- | --- | +| The `acor` CLI | Flags, output, and exit codes can change in any release. Call the library if you need a stable contract | +| `acor/server` | A separate, experimental module with no tags of its own; the core's version numbers say nothing about it | +| `internal/...` | Not importable, free to change | +| The `benchmarks` module | A measurement harness, not an API | +| Documentation wording | Pages get rewritten; what they describe is pinned by this page, not by their phrasing | +| The V1 schema layout | Deprecated and not evolved further — see [Deprecation](#deprecation) | ## Conditions on your code -The promise holds for code that follows all three. None of them is something the compiler -can enforce on your behalf. +The promise holds for code that follows all three. None is enforceable by the compiler. ### Construct option structs with field names `AhoCorasickArgs`, `MatchOptions`, `BatchOptions`, `ParallelOptions`, and -`MigrationOptions` gain fields in minor releases. That is not a breaking change for a -keyed literal: +`MigrationOptions` gain fields in minor releases. A keyed literal survives that: ```go @@ -83,150 +75,128 @@ args := &acor.AhoCorasickArgs{ _ = args ``` -An unkeyed literal breaks the moment any field is added, and is not covered by this -promise: +An unkeyed literal breaks the moment any field is added, and is not covered: ```go // Not covered: positional fields break when a field is added. opts := acor.MatchOptions{acor.MatchKindLeftmostLongest, true, nil} ``` -That literal compiles against `v1.5.0` and stops compiling the moment `MatchOptions` gains -a fourth field. `go vet`'s `composites` check reports unkeyed literals of another package's -structs, so this condition is verifiable in your own build. +`go vet`'s `composites` check reports unkeyed literals of another package's structs, so +this condition is verifiable in your own build. The structs ACOR *returns* run the other way. `AhoCorasickInfo`, `MigrationResult`, `BatchResult`, and `CacheStats` are built by ACOR and read by you, so a new field cannot -break a reader — and they do gain fields inside `v1`. The condition on them is only that -you not construct or whole-value compare one: assert on the fields you care about, and a -later release adding a sixth counter stays invisible to your code. +break a reader — and they do gain fields inside `v1`. The condition is only that you not +construct or whole-value compare one: assert on the fields you care about, and a later +release adding a sixth counter stays invisible. Their JSON field names are covered too. `api/v1.txt` records struct tags, so renaming -`json:"status"` is a breaking change even though it moves no Go signature. That applies -to marshaling a returned struct yourself; the `acor` command's own `--json` output is a -CLI detail and stays uncovered. +`json:"status"` is a breaking change even though it moves no Go signature. That applies to +marshaling a returned struct yourself; the `acor` command's own `--json` output is a CLI +detail and stays uncovered. ### Do not expect the exported interface to grow -One exported interface can be implemented from outside the module: `Logger`. No method -is added to it inside `v1` — doing so would break every existing implementation. +`Logger` is the one exported interface implementable from outside the module, and no +method is added to it inside `v1` — that would break every existing implementation. `KVStorage`, `StringMapResult`, `Subscription`, and `Pipeliner` were exported through -`v1.4.0` and are not part of the `v1.5.0` surface. Nothing public ever accepted or -returned one, so no caller could supply an implementation; and freezing them would have -capped the pluggable-storage work they existed for, since this very rule forbids adding -a method to them later. They are unexported now so that feature can choose its own -shape, which `v1` permits as an addition. +`v1.4.0` and are not part of the `v1.5.0` surface. Nothing public accepted or returned +one, so no caller could supply an implementation; and freezing them would have capped the +pluggable-storage work they existed for, since this very rule forbids adding a method +later. They are unexported now so that feature can pick its own shape. ### Do not dot-import the package `import . "github.com/skyoo2003/acor/pkg/acor"` puts every exported name into your file's -scope, so a name added in a minor release can collide with one of your own declarations and +scope, so a name added in a minor release can collide with one of your declarations and stop the file compiling. A normal import cannot: additions land behind the `acor.` -qualifier, where they collide with nothing. +qualifier. ## The on-Redis data format -ACOR exists so that several instances share one dictionary, which means a rolling deploy -runs two ACOR versions against the same Redis keys at the same time. The Go API promise -says nothing about that case. This section does. - -Inside `v1`, changes to the V2 format are **additive only**: +Several instances share one dictionary, so a rolling deploy runs two ACOR versions against +the same keys. Inside `v1`, changes to the V2 format are **additive only**: - Key names, hash tags, and the names and meanings of existing fields do not change. -- New fields may be added to the `{name}:trie` hash, or under a new key. An ACOR version - that does not recognize a field there ignores it: that hash is read by looking its - fields up by name. +- New fields may be added to the `{name}:trie` hash or under a new key. A version that + does not recognize a field there ignores it — that hash is read field by field, by name. - Nothing is added to the `{name}:outputs` hash. Its field names are automaton states and - every value in it is read as match data, so it has no room for metadata. -- A mixed-version fleet therefore keeps working in both directions for the whole `v1` - line, and a rolling deploy needs no coordinated restart. + every value is read as match data, so it has no room for metadata. -`Flush` rewrites `{name}:trie` in full, so a `Flush` issued by an older instance drops -fields a newer one added. The next write from an upgraded instance restores them. +A mixed-version fleet therefore keeps working in both directions for the whole `v1` line, +and a rolling deploy needs no coordinated restart. -The cost lands on new features rather than on availability: a feature that depends on a -newly added field does not take effect until every instance writing to the collection is -upgraded. Expect it to appear after the rollout finishes, not during it. +Two consequences worth planning for: + +- `Flush` rewrites `{name}:trie` in full, so a `Flush` from an older instance drops fields + a newer one added. The next write from an upgraded instance restores them. +- A feature depending on a newly added field does not take effect until every instance + writing to the collection is upgraded. Expect it after the rollout, not during it. ## What counts as a breaking change A patch release (`x.y.Z`) carries bug fixes and no surface change; a minor release -(`x.Y.0`) adds to the surface without breaking it; a breaking change requires a new major -version. +(`x.Y.0`) adds without breaking; a breaking change requires a new major version. Removing or renaming an exported identifier, changing a signature, and narrowing a return -type are the obvious cases, and tooling catches them. These count too, and tooling does -not catch any of them: +type are the obvious cases, and tooling catches them. These count too, and tooling catches +none of them: -- **Adding a field to an option struct**, for callers using unkeyed literals — which is - why the condition above exists. +- **Adding a field to an option struct**, for callers using unkeyed literals — hence the + condition above. - **No longer returning a sentinel error** for the situation that returns it today. - `errors.Is(err, ErrCacheWithPreset)` continuing to hold is part of the promise; the - variable merely continuing to exist is not enough. + `errors.Is(err, ErrCacheWithPreset)` continuing to hold is the promise; the variable + merely continuing to exist is not enough. - **Changing documented behavior** — match ordering, `MatchKind` semantics, `FindSet`'s first-match order, `FindParallel`'s deduplication contract. -- **Changing the meaning of an existing V2 field**, per the section above. +- **Changing the meaning of an existing V2 field.** -Undocumented detail is not covered. Sort order among equally-ranked matches, the exact +Undocumented detail is not covered: sort order among equally-ranked matches, the exact wording of an error string, allocation counts, and round-trip counts can change in any release. ## How the promise is enforced -Most of this page is checked by machine rather than by a reviewer noticing. - -`api/v1.txt` in the repository is the covered surface, one symbol per line — -functions, methods, struct fields, and interface methods. CI regenerates it and fails -if the result differs from what is committed. A pull request that changes the public -API therefore has to change that file in the same diff: an addition appends a line, and -a removal **deletes** one, in front of a reviewer. Nothing stops a removal from being -merged; what is gone is the possibility of merging one unnoticed. - -Two clauses on this page are not covered by that, and both are stated here rather than -left to be discovered: - -- **Documented behavior.** Match ordering, `MatchKind` semantics, `FindSet`'s - first-match order — no tool compares these *across versions*, and none can: the - claim lives in prose. What is checked is that somebody looked. `api/v1-audit.txt` - carries one verdict per entry of `api/v1.txt`, and CI fails when an entry has none, - so a *missing* verdict is a build failure rather than an omission nobody notices. - A verdict of `unaudited` is legal, so an unreviewed entry does not fail the build — - it is counted instead, and the tally prints on every run. Whether each verdict is - *right*, and whether its cited `file:line` still points where it did, is review. - - Every entry of the `v1.5.0` surface has been read against the code it describes, - and 38 of them said something the code did not do. Those sentences were rewritten - before the freeze, so what `v1` promises is the corrected wording, not the wording - that shipped in `v1.4.0`. A non-`unaudited` verdict has to cite the `file:line` - the behavior was read at, which is what stops a future pass from marking a line - reviewed without reviewing it. - - Those verdicts still describe the code as shipped. `v1.5.1` changed no entry of - the surface they were measured against: `git diff v1.5.0 v1.5.1 -- api/v1.txt` is - empty, and the release touched only `LICENSE`, `server/LICENSE`, and the changelog — - no Go file at all. -- **Cross-version Redis interop.** Tests pin the key names and hash field names, so a - *rename* fails CI. They do not run an older ACOR against a newer one's data, so - additive-only is verified as far as naming and no further. - -Retractions are checked too: CI fails if `retract [v1.0.0, v1.4.0]` leaves `go.mod`, -because a release without it silently makes the retracted range resolvable again. +`api/v1.txt` is the covered surface, one symbol per line — functions, methods, struct +fields, interface methods. CI regenerates it and fails if the result differs from what is +committed, so a PR changing the public API has to change that file in the same diff: an +addition appends a line, a removal **deletes** one, in front of a reviewer. Nothing stops +a removal from being merged; what is gone is merging one unnoticed. + +CI also fails if `retract [v1.0.0, v1.4.0]` leaves `go.mod`, because dropping it would +silently make the retracted range resolvable again. + +Two clauses are not covered by that check: + +- **Documented behavior.** No tool compares match ordering or `MatchKind` semantics + *across versions* — the claim lives in prose. What is checked is that somebody looked: + `api/v1-audit.txt` carries one verdict per entry of `api/v1.txt`, and CI fails when an + entry has none. A verdict of `unaudited` is legal but counted, and the tally prints on + every run. A non-`unaudited` verdict must cite the `file:line` the behavior was read at. + + Every entry of the `v1.5.0` surface was read against its code, and 38 said something the + code did not do. Those sentences were rewritten before the freeze, so `v1` promises the + corrected wording, not what shipped in `v1.4.0`. `v1.5.1` changed no entry of that + surface — `git diff v1.5.0 v1.5.1 -- api/v1.txt` is empty, and the release touched only + `LICENSE`, `server/LICENSE`, and the changelog. +- **Cross-version Redis interop.** Tests pin key names and hash field names, so a *rename* + fails CI. They do not run an older ACOR against a newer one's data, so additive-only is + verified as far as naming and no further. ## Deprecation A deprecated identifier keeps working for the rest of `v1`. It carries a `Deprecated:` -line naming its replacement, gains no new capability, and is removed no earlier than -`v2`. +line naming its replacement, gains no new capability, and is removed no earlier than `v2`. -Deprecated today: `SchemaV1`, which is read-only and whose read path is removed no -earlier than `v2`. +Deprecated today: `SchemaV1`, read-only, whose read path is removed no earlier than `v2`. ## `v2` -A `v2` would ship as the module `github.com/skyoo2003/acor/v2`, imported as +A `v2` would ship as `github.com/skyoo2003/acor/v2`, imported as `github.com/skyoo2003/acor/v2/pkg/acor`. `v1` stays importable and unchanged at its own -path, so nothing breaks by the act of publishing `v2`. There is no schedule. +path, so publishing `v2` breaks nothing by itself. There is no schedule. -Security patches follow the [security policy](https://github.com/skyoo2003/acor/blob/main/SECURITY.md), -which defines which lines receive them. +Security patches follow the +[security policy](https://github.com/skyoo2003/acor/blob/main/SECURITY.md). diff --git a/docs/content/reference/r2-r3-performance.md b/docs/content/reference/r2-r3-performance.md index 6d107987..ce287cb9 100644 --- a/docs/content/reference/r2-r3-performance.md +++ b/docs/content/reference/r2-r3-performance.md @@ -3,52 +3,51 @@ title: "R2/R3 verification and R1 comparison" description: "Incremental download, engine memory and bounded source-text API evidence." --- -The R2/R3 implementation completed 36 real-server runs on 2026-09-06 using the -same R1 workload: Redis/Valkey × 10,000/100,000/1,000,000 keywords × shared/diverse/ -Korean distributions × two repetitions. The R1 measurements remain unchanged in -`benchmarks/results/v3-20260906.json`. New measurements, source hashes and safety -results are in `benchmarks/results/r2-r3-20260906.json`. - -Environment: Apple M4, 10 logical CPUs, 16 GiB RAM, macOS 26.5.2, Go 1.26.7, -Redis 8.10.1 and source-built Valkey 9.1.2. Both are standalone localhost TCP, -with RDB/AOF disabled. The preset remains MemoryEfficient; polling remains 50 ms -for measurement. R2 adds the default fixed 20 ms refresh debounce. Each run uses -a fresh Go process. These are developer-workstation measurements with possible -background host activity, not latency guarantees or confidence intervals. - -Valkey uses `prefetch-batch-max-size 0`, as in the successful R1 baseline. The -local default-prefetch build previously crashed; this work does not establish -support for that configuration or diagnose its root cause. See the -[R1 report](../versioned-performance/) for source checksum and crash evidence. - -## R2 changes - -- The installed engine keeps verified immutable keyword slices and its manifest. - Unchanged buckets share those slices; only changed nonempty buckets download - chunk data. Candidate caches are installed with the engine only after success. -- MemoryEfficient builds consume the bucket sequence directly, avoiding the - flattened full dictionary slice and an additional keyword-set map. Nodes keep - their only child inline; a map is allocated only when the node branches. BFS - uses a compact integer queue. The first-rune Bloom filter sizes itself from - distinct first runes rather than total keyword count. -- Builds check cancellation in insertion, failure-link construction and table - filling for all three presets. Cancellation discards the private candidate; - no detached builder goroutine continues after return. Allocations, rune-count - calls and sorting finish before the next checkpoint. -- A fixed debounce window merges burst notifications. New events cannot extend - that window; an in-flight build completes under ordinary writes, preventing - repeated cancellation from starving refresh. Close cancels its build context. -- Status reports the last successful build's DownloadedBuckets/ReusedBuckets and - cumulative CompletedBuilds. Redis schema, expected-version writes, receipt - resolution, leases and pruning semantics are unchanged. +# R2/R3 verification and R1 comparison + +R2/R3 completed 36 real-server runs on 2026-09-06 using the R1 workload: Redis/Valkey × +10,000/100,000/1,000,000 keywords × shared/diverse/Korean distributions × two repetitions. +The R1 measurements are unchanged in `benchmarks/results/v3-20260906.json`; new +measurements, source hashes, and safety results are in +`benchmarks/results/r2-r3-20260906.json`. + +Environment: Apple M4, 10 logical CPUs, 16 GiB RAM, macOS 26.5.2, Go 1.26.7, Redis 8.10.1, +source-built Valkey 9.1.2. Both standalone on localhost TCP with RDB/AOF disabled. Preset +remains MemoryEfficient and polling remains 50 ms for measurement; R2 adds the default +fixed 20 ms refresh debounce. Each run uses a fresh Go process. These are +developer-workstation measurements with possible background host activity — not latency +guarantees or confidence intervals. + +Valkey uses `prefetch-batch-max-size 0`, as in the R1 baseline. The local default-prefetch +build previously crashed; this work neither establishes support for that configuration nor +diagnoses it. See the [R1 report](../versioned-performance/) for the source checksum and +crash evidence. + +## What R2 changed + +- The installed engine keeps verified immutable keyword slices and its manifest. Unchanged + buckets share those slices; only changed nonempty buckets download chunk data. Candidate + caches install with the engine, and only after success. +- MemoryEfficient builds consume the bucket sequence directly, avoiding the flattened + dictionary slice and an extra keyword-set map. Nodes keep their only child inline and + allocate a map only when they branch; BFS uses a compact integer queue; the first-rune + Bloom filter sizes itself from distinct first runes rather than total keyword count. +- All three presets check cancellation during insertion, failure-link construction, and + table filling. Cancellation discards the private candidate with no detached builder + goroutine left running. Allocations, rune-count calls, and sorting finish before the next + checkpoint. +- A fixed debounce window merges burst notifications and cannot be extended by new events, + so an in-flight build completes under ordinary writes rather than being starved by + repeated cancellation. `Close` cancels its build context. +- `Status` reports the last successful build's `DownloadedBuckets`/`ReusedBuckets` and + cumulative `CompletedBuilds`. Redis schema, expected-version writes, receipt resolution, + leases, and pruning semantics are unchanged. ## Million-keyword comparison -Ranges below are the minimum–maximum of two runs. Ready is write preparation, -commit, refresh detection and engine construction combined. RSS is the process -high-water mark across the full run, including all update scenarios. Network -counters are server-wide and include protocol, polling and engine downloads; -they do not isolate only keyword payloads. +Ranges are minimum–maximum of two runs. Ready combines write preparation, commit, refresh +detection, and engine construction. RSS is the process high-water mark across the full run. +Network counters are server-wide and include protocol, polling, and engine downloads. | Server | Distribution | Add 1 ready: R1 → R2 (s) | Received bytes reduction for add 1 | Peak RSS: R1 → R2 (GiB) | |---|---|---:|---:|---:| @@ -59,10 +58,10 @@ they do not isolate only keyword payloads. | valkey | diverse | 2.33–2.40 → 0.92–1.27 | 91.9% | 3.02–3.80 → 2.02–2.19 | | valkey | korean | 2.90–3.40 → 1.21–1.60 | 96.2% | 3.37–3.69 → 2.33–2.50 | -All six million-keyword server/distribution combinations reduced single-change -received bytes and peak RSS in these runs. A full engine is still rebuilt after -a change; this is not an incremental automaton or a constant-memory design. -Larger changes touch more buckets and therefore reuse less downloaded data. +All six million-keyword combinations reduced single-change received bytes and peak RSS in +these runs. A full engine is still rebuilt after a change — this is not an incremental +automaton or a constant-memory design, and larger changes touch more buckets and reuse +less. | Server | Distribution | Add 1,000 ready: R1 → R2 (s) | Add 1% ready: R1 → R2 (s) | Full replace ready: R1 → R2 (s) | |---|---|---:|---:|---:| @@ -73,50 +72,48 @@ Larger changes touch more buckets and therefore reuse less downloaded data. | valkey | diverse | 2.42–2.60 → 1.21–1.29 | 3.03–3.54 → 2.14–2.19 | 4.08–4.33 → 2.77–3.61 | | valkey | korean | 4.03–4.38 → 1.45–1.60 | 3.98–4.49 → 2.64–2.70 | 6.78–7.31 → 4.17–4.92 | -Improvement is not uniform: some Redis large-change repetitions overlap or -exceed the R1 timing range. Two repetitions do not establish a universal speedup. +Improvement is not uniform: some Redis large-change repetitions overlap or exceed the R1 +range. Two repetitions do not establish a universal speedup. -The JSON also records 10,000/100,000-entry results, initial loading/startup, -identical replacement, removals, search p50/p95/p99 while refreshing, Redis -memory and pruning. The same seed 20260906 and generation functions are used. -Prune measurements simulate the retention horizon by aging generation registry -scores; they do not wait 24 hours. Separate million-entry safety tests passed -against both servers after the matrix, with no concurrent workload on a measured -endpoint. +The JSON also records 10,000/100,000-entry results, initial loading and startup, identical +replacement, removals, search p50/p95/p99 while refreshing, Redis memory, and pruning, +using the same seed 20260906 and generation functions. Prune measurements simulate the +retention horizon by aging generation registry scores rather than waiting 24 hours. +Separate million-entry safety tests passed against both servers after the matrix, with no +concurrent workload on a measured endpoint. ## R3 contracts and verification -`Scan`, `MaskText` and `ReplaceText` are additive on AhoCorasick and -VersionedCollection. See [bounded text processing](../text-processing/) for the -complete defaults and error behavior. Existing unlimited APIs are unchanged. +`Scan`, `MaskText`, and `ReplaceText` are additive on `AhoCorasick` and +`VersionedCollection`; existing unlimited APIs are unchanged. Defaults and error behavior: +[bounded text processing](../text-processing/). | Area | Evidence | |---|---| -| Original positions | All presets match existing FindMatches for overlapping and leftmost-longest results; original byte slices survive Korean, emoji, `İ` case folding and malformed UTF-8 input | -| Result limits | At most MaxMatches are retained; Truncated requires an additional eligible result | -| Work/input bounds | Oversize input and excess raw candidates fail explicitly, including candidates later filtered by word boundaries/overlap | -| Atomic rewrite | Match, output and work exhaustion return no partial rewrite; empty literal replacement deletes; masks preserve original rune count, including NUL masks | +| Original positions | All presets match existing `FindMatches` for overlapping and leftmost-longest results; original byte slices survive Korean, emoji, `İ` case folding, and malformed UTF-8 | +| Result limits | At most `MaxMatches` retained; `Truncated` requires an additional eligible result | +| Work/input bounds | Oversize input and excess raw candidates fail explicitly, including candidates later filtered by word boundaries or overlap | +| Atomic rewrite | Match, output, and work exhaustion return no partial rewrite; empty literal replacement deletes; masks preserve original rune count, including NUL masks | | Overlap | A bounded pending-start heap selects leftmost-longest matches without collecting all raw matches | -| Fuzz | FuzzScanLeftmostParity completed 30 seconds and 819,033 executions against the existing leftmost-longest implementation | -| Refresh | Verified bucket reuse, failed-candidate cache preservation and burst coalescing tests pass | +| Fuzz | `FuzzScanLeftmostParity` completed 30 seconds and 819,033 executions against the existing leftmost-longest implementation | +| Refresh | Bucket reuse, failed-candidate cache preservation, and burst coalescing tests pass | | Cancellation | All three presets preserve the previous engine on cancellation during sequence ingestion and CPU construction; unrelated panics are not swallowed | -| Safety | Both real servers pass the million-entry conflict, pinned paging and lease-safe pruning scenario | +| Safety | Both real servers pass the million-entry conflict, pinned paging, and lease-safe pruning scenario | -Resource bounds are per call. Scratch input indexing is proportional to the -bounded input size, and the pending window is bounded by input/longest-keyword -length and candidate count. A caller's custom WordRune function controls its own -resource use. R3 does not claim a wall-clock bound on callbacks or allocation. +Resource bounds are per call: scratch input indexing is proportional to the bounded input +size, and the pending window is bounded by input/longest-keyword length and candidate +count. A custom `WordRune` controls its own resource use. R3 claims no wall-clock bound on +callbacks or allocation. -Root and server race tests, root/server lint, all-module vet, benchmark-module -tests, API snapshot/audit and documentation compilation pass. Legacy Find/Add -fuzz targets are rerun for 30 seconds each. No HTTP/gRPC surface or storage -schema migration is introduced by R2/R3. +Root and server race tests, root/server lint, all-module vet, benchmark-module tests, API +snapshot/audit, and documentation compilation pass. Legacy Find/Add fuzz targets rerun for +30 seconds each. R2/R3 introduces no HTTP/gRPC surface and no storage schema migration. ## Reproduce Run the unchanged `scripts/benchmark-v3.sh` against each disposable endpoint with -`ACOR_V3_SCALE_REPEATS=2` and separate output directories. See the -[R1 reproduction instructions](../versioned-performance/#reproduce) for the exact -environment and dictionary distributions. The R2/R3 defaults select the new -engine and refresh behavior automatically; existing V1/V2 public contracts remain -covered by the regression suite. +`ACOR_V3_SCALE_REPEATS=2` and separate output directories; the exact environment and +dictionary distributions are in the +[R1 reproduction instructions](../versioned-performance/#reproduce). The R2/R3 defaults +select the new engine and refresh behavior automatically, and existing V1/V2 public +contracts remain covered by the regression suite. diff --git a/docs/content/reference/schema-v1.md b/docs/content/reference/schema-v1.md index 4d23f89c..ae571a7a 100644 --- a/docs/content/reference/schema-v1.md +++ b/docs/content/reference/schema-v1.md @@ -5,21 +5,16 @@ weight: 3 # Schema V1 (Deprecated) -V1 is the original ACOR storage schema. It uses multiple Redis keys per collection. - > **V1 is deprecated and read-only as of `v1.5.0`.** Reads, `Suggest`, and `Info` work, > and `MigrateV1ToV2` converts a collection in place, but `Add` and `Remove` return -> `ErrV1ReadOnly`. `Flush` also still works, and still deletes every key in the -> collection — read-only refuses keyword writes, it does not protect the collection from -> `Flush`. It gains no features either: preset engines and -> `EnableCache` both require V2. New collections should use the default V2 schema. -> The read path stays for the whole `v1` line and is removed no earlier than `v2`. - -## Overview +> `ErrV1ReadOnly`. `Flush` still works and still deletes every key — read-only refuses +> keyword writes, it does not protect the collection. V1 gains no features: preset engines +> and `EnableCache` both require V2. The read path stays for the whole `v1` line and is +> removed no earlier than `v2`. -V1 creates approximately 5 keys per 100 keywords: +V1 spreads one collection across many keys — roughly 5 per 100 keywords. -| Key Pattern | Purpose | +| Key pattern | Purpose | |-------------|---------| | `{name}:keyword` | Set of keywords | | `{name}:prefix` | Trie prefix edges | @@ -27,53 +22,14 @@ V1 creates approximately 5 keys per 100 keywords: | `{name}:output:{state}` | Output keywords per state | | `{name}:node:{keyword}` | Node metadata | -## Performance Characteristics - -| Operation | Complexity | -|-----------|------------| -| Find() | O(N×3-5) RTT | -| Add() | O(M×3-10) RTT — no longer reachable; kept to explain the migration's value | - -Where: -- N = number of trie states visited -- M = keyword length +`Find()` costs O(N×3-5) round trips for N visited states. `Add()` cost O(M×3-10) for a +keyword of length M — no longer reachable, and recorded only to explain what migration is +worth. Compare against [Schema V2](../schema-v2/#against-v1). -## When to Use V1 - -- Existing collections using V1 -- Small keyword sets (< 10,000) -- Migration not feasible - -## Migration to V2 +## Migrating ```bash -# Preview migration -acor -name mycollection migrate --dry-run - -# Execute migration -acor -name mycollection migrate - -# Rollback to V1 -acor -name mycollection migrate-rollback +acor -name mycollection migrate --dry-run # preview +acor -name mycollection migrate # execute +acor -name mycollection migrate-rollback # back to V1 ``` - -## Key Structure Diagram - -```mermaid -graph LR - A[keyword set] --> B[prefix trie] - B --> C[suffix links] - C --> D[output sets] - D --> E[node metadata] -``` - -## Limitations - -- Higher memory usage -- More Redis keys to manage -- Slower Find() operations -- More network round-trips - -## Recommendation - -**Migrate to V2** for new collections or when performance is critical. diff --git a/docs/content/reference/schema-v2.md b/docs/content/reference/schema-v2.md index 643ec4c4..8695f220 100644 --- a/docs/content/reference/schema-v2.md +++ b/docs/content/reference/schema-v2.md @@ -5,128 +5,59 @@ weight: 4 # Schema V2 (Optimized) -V2 is the recommended schema for ACOR. A collection occupies a fixed set of at -most 3 keys, whatever the dictionary size. +V2 is the default for new collections — no configuration needed. A collection occupies at +most 3 keys whatever the dictionary size. -## Overview +| Key | Holds | When it exists | +|-----|-------|----------------| +| `{name}:trie` | Serialized trie: keywords, prefixes, version | Always, from creation | +| `{name}:outputs` | Output mappings, state → keywords | Once the collection has a keyword | +| `{name}:nodes` | Node metadata | Only after `MigrateV1ToV2`; cleaned up by flush | -V2 consolidates storage into these keys: +Most collections hold two keys; a fresh one holds a single `:trie`. Nothing but migration +writes `:nodes`. Budget for three, expect to count fewer. -| Key Pattern | Purpose | When it exists | -|-------------|---------|----------------| -| `{name}:trie` | Serialized trie structure (keywords, prefixes, version) | Always, from creation | -| `{name}:outputs` | All output mappings (state -> keywords) | Once the collection has a keyword | -| `{name}:nodes` | Node metadata | Only on a collection produced by `MigrateV1ToV2`; cleaned up by flush | +## Field layout -Most collections therefore hold two keys, and a freshly created one holds a -single `:trie`. Nothing but migration writes `:nodes`, so a collection built with -`Add` never has it. Budget for three; expect to count fewer. +```text +{name}:trie (hash) + keywords -> ["keyword1", "keyword2", ...] + prefixes -> ["", "h", "he", ...] + version -> + +{name}:outputs (hash) + he -> ["he"] + she -> ["he", "she"] -## Performance Characteristics +{name}:nodes (hash, migration only) + keyword1 -> ["s0","s1","s2"] +``` -| Operation | Complexity | -|-----------|------------| -| Find() | 1 RTT (fixed), 0 RTT with EnableCache | -| Add() | 2 RTT (flat, independent of keyword length) | +Collections written before v0.11 also carry a `suffixes` field on `:trie`. It is never +read, writes leave it alone, and the next `Flush()` drops it. -## Comparison with V1 +## Against V1 -Round trips are counted by tests on every CI run; timings are a sample from the -hardware named on the [benchmarks page](../benchmarks/). +Round trips are counted by tests on every CI run; timings are a sample from the hardware +named on the [benchmarks page](../benchmarks/). | Metric | V1 | V2 | |--------|----|----| | Keys per 100K keywords | ~500K | 2 | | `Find()` round trips | 1 | 1 | -| `Add()` round trips | grows with keyword length (53 at 5 chars, 507 at 26) | 2 | +| `Add()` round trips | Grows with keyword length: 53 at 5 chars, 507 at 26 | 2 | | `Add()` time | baseline | ~14x faster | | `Find()` time, no cache | baseline | ~1.7x **slower** at 1000 keywords | -V2's win is on writes, not reads. Both schemas read in a single round trip, and -uncached V2 is slightly slower on `Find()` because it reads an outputs hash with -one entry per trie state where V1 reads only the keyword set. The large read -speedups belong to `EnableCache` and the preset engines, not to the schema. See -[Benchmarks](../benchmarks/) for the full tables and how to reproduce them. - -## Architecture - -```mermaid -graph TB - subgraph V2 Schema - A[trie key] --> B[Serialized Trie] - C[outputs key] --> D[Output Map] - E[nodes key] --> F[Node Metadata - migration only] - end - - G[Find Operation] --> A - G --> C - G -.-> E -``` - -## Enabling V2 - -V2 is automatically used for new collections. No configuration needed. - -```go -ac, err := acor.Create(&acor.AhoCorasickArgs{ - Addr: "localhost:6379", - Name: "my-v2-collection", -}) -if err != nil { - log.Fatal(err) -} -// Automatically uses V2 schema -``` +**V2's win is writes, not reads.** Both schemas read in one round trip, and uncached V2 is +slightly slower on `Find()` because it reads an outputs hash with one entry per trie state +where V1 reads only the keyword set. The large read speedups belong to `EnableCache` and +the preset engines, not to the schema. -## Migration from V1 +## Migrating from V1 ```bash -# Check current schema -acor -name mycollection schema-version - -# Preview migration +acor -name mycollection schema-version # what you have now acor -name mycollection migrate --dry-run - -# Execute migration acor -name mycollection migrate ``` - -## Key Structure - -### trie key - -Stores the serialized trie as a hash with three fields: - -```text -{collection}:trie - keywords -> ["keyword1", "keyword2", ...] - prefixes -> ["", "h", "he", ...] - version -> -``` - -Collections written before v0.11 also carry a `suffixes` field. It is never -read, is left alone by writes, and is dropped by the next `Flush()`. - -### outputs key - -Stores output keywords per trie state as a hash: - -```text -{collection}:outputs - he -> ["he"] - she -> ["he", "she"] -``` - -### nodes key - -A hash mapping each keyword to a JSON array of its trie state strings. It is -populated only by V1→V2 migration (fresh V2 collections do not write it): - -```text -{collection}:nodes (hash) - keyword1 -> ["s0","s1","s2"] -``` - -## Recommendation - -**Use V2 for all new collections.** It provides significantly better performance and lower resource usage. diff --git a/docs/content/reference/text-processing.md b/docs/content/reference/text-processing.md index ebfee5b5..f365ebfe 100644 --- a/docs/content/reference/text-processing.md +++ b/docs/content/reference/text-processing.md @@ -3,73 +3,71 @@ title: "Bounded search, masking and replacement" description: "Original byte and rune positions with explicit work and output limits." --- -R3 adds `Scan`, `MaskText`, and `ReplaceText` to both AhoCorasick and -VersionedCollection. All take a context. Existing Find/FindMatches APIs retain -their existing unlimited result and position contracts. V3 calls hold one serving -engine for the entire scan or rewrite, even if a refresh finishes concurrently. +# Bounded search, masking and replacement -## Original positions and bounded search +`Scan`, `MaskText`, and `ReplaceText` exist on both `AhoCorasick` and +`VersionedCollection` and all take a context. The existing `Find`/`FindMatches` APIs keep +their unlimited result and position contracts. A V3 call holds one serving engine for the +whole scan or rewrite, even if a refresh finishes concurrently. -`Scan(ctx, text, *ScanOptions)` returns SourceMatch entries containing the normalized -Keyword, the original Text substring, and half-open Start/End rune and -ByteStart/ByteEnd byte offsets into the original input. Unicode case folding can -change byte lengths: `İSTANBUL` matches `istanbul`, but its byte span still selects -the original spelling. No normalization of the original output text occurs. -Invalid UTF-8 input bytes are decoded as RuneError for matching, while reported -byte slices retain those exact original bytes. +## Scan -Zero-valued limits select these defaults; negative limits are rejected. Construct -ScanOptions and RewriteOptions using named fields for forward compatibility. +`Scan(ctx, text, *ScanOptions)` returns `SourceMatch` entries carrying the normalized +`Keyword`, the original `Text` substring, and half-open `Start`/`End` rune and +`ByteStart`/`ByteEnd` byte offsets into the **original** input. -| Option | Default | Behavior at limit | +Case folding can change byte lengths — `İSTANBUL` matches `istanbul`, and its byte span +still selects the original spelling. Output text is never normalized. Invalid UTF-8 is +decoded as `RuneError` for matching, while the reported byte slices keep those exact +original bytes. + +| Option | Default | At the limit | |---|---:|---| -| MaxInputBytes | 1 MiB | ErrInputLimit before loading/scanning the engine | -| MaxMatches | 1,000 | At most this many entries; Truncated when another eligible match is observed | -| MaxCandidates | 100,000 | ErrScanWorkLimit when another raw automaton match is encountered | - -Kind defaults to MatchKindOverlapping. MatchKindLeftmostLongest selects the -leftmost available start, then the longest keyword there, and discards overlaps. -The implementation keeps the longest candidate at each pending start until the -engine's longest keyword makes the decision safe. It does not collect all raw -matches and sort them. Scratch memory is bounded by input length plus the pending -start window and retained result count; it is not constant memory. Input indexing -uses rune and byte-offset arrays. A low result limit alone does not replace the -input-byte or candidate-work limit. - -WholeWord and WordRune have the same normalized-rune boundary semantics as -MatchOptions. Letters, digits, combining marks and underscore are word characters -by default. Scripts without spaces may require an application-specific WordRune. -Candidate counting happens before whole-word and overlap filtering, so a dense -rejected-match workload cannot bypass the work budget. Custom WordRune code runs -in the caller's goroutine; its own resource use is the caller's responsibility. - -Input or candidate exhaustion and cancellation return an error with no result. -Truncated concerns the eligible result count only. A caller checking safety or -completeness must handle both errors and Truncated; neither means a clean input. -The context is checked during input indexing, traversal and match handling. - -## Atomic masking and literal replacement - -`ReplaceText(ctx, text, replacement, *RewriteOptions)` inserts the literal -replacement for each selected non-overlapping leftmost-longest match. It does not -interpret regular expressions, expand captures, or search replacement text again. -An empty replacement deletes matched spans. - -`MaskText(ctx, text, maskRune, *RewriteOptions)` writes one mask rune for every -matched original rune. The result preserves rune count, not necessarily byte count. -Any valid Unicode rune, including NUL, is accepted. Invalid runes are rejected. -Both APIs leave unmatched original bytes untouched. - -RewriteOptions shares the three scan limits above and adds MaxOutputBytes, which -defaults to 4 MiB. WholeWord and WordRune are also available. Overlap selection is -always leftmost-longest. Exceeding MaxMatches produces ErrMatchLimit; exceeding the -output bound produces ErrOutputLimit. Every error returns no RewriteResult, so a -caller cannot accidentally consume a partially masked document. All output sizes -are checked before allocating the output buffer. - -RewriteResult contains Text and the SourceMatch entries used for the rewrite. -Their offsets always refer to the **input**, even when replacement changes output -length. Match Text substrings can retain the original input string in memory. +| `MaxInputBytes` | 1 MiB | `ErrInputLimit`, before the engine is loaded or scanned | +| `MaxMatches` | 1,000 | Keeps at most this many; sets `Truncated` when another eligible match appears | +| `MaxCandidates` | 100,000 | `ErrScanWorkLimit` on the next raw automaton match | + +Zero-valued limits select the defaults; negative limits are rejected. Construct +`ScanOptions` and `RewriteOptions` with named fields. + +`Kind` defaults to `MatchKindOverlapping`. `MatchKindLeftmostLongest` takes the leftmost +available start, then the longest keyword there, and discards overlaps — implemented by +keeping the longest candidate at each pending start until the engine's longest keyword +makes the decision safe, not by collecting and sorting all raw matches. Scratch memory is +bounded by input length plus the pending-start window and retained results; it is not +constant. + +`WholeWord` and `WordRune` carry the same normalized-rune semantics as `MatchOptions`: +letters, digits, combining marks, and underscore are word characters by default, and +scripts without spaces usually need an application-specific `WordRune`. Candidates are +counted *before* whole-word and overlap filtering, so a dense rejected-match workload +cannot slip past the work budget. Custom `WordRune` code runs in the caller's goroutine. + +**Errors and `Truncated` are different signals.** Input exhaustion, candidate exhaustion, +and cancellation all return an error and no result; `Truncated` concerns only the eligible +result count. Checking for a clean input means handling both. The context is checked +during input indexing, traversal, and match handling. + +## Masking and replacement + +`ReplaceText(ctx, text, replacement, *RewriteOptions)` inserts the literal replacement for +each selected non-overlapping leftmost-longest match. No regular expressions, no capture +expansion, no re-searching the replacement. An empty replacement deletes matched spans. + +`MaskText(ctx, text, maskRune, *RewriteOptions)` writes one mask rune per matched original +rune, preserving rune count but not necessarily byte count. Any valid Unicode rune is +accepted, including NUL; invalid runes are rejected. + +Both leave unmatched original bytes untouched. `RewriteOptions` shares the three scan +limits and adds `MaxOutputBytes`, defaulting to 4 MiB. Overlap selection is always +leftmost-longest. Exceeding `MaxMatches` gives `ErrMatchLimit`, exceeding the output bound +gives `ErrOutputLimit`, and **every error returns no `RewriteResult`** — a caller cannot +accidentally consume a half-masked document. Output size is checked before the buffer is +allocated. + +`RewriteResult` holds `Text` plus the `SourceMatch` entries used. Their offsets always +refer to the **input**, even when replacement changes the output length. Match `Text` +substrings can retain the original input string in memory. ```go @@ -112,5 +110,5 @@ func main() { } ``` -See the [R2/R3 report](../r2-r3-performance/) for parity, resource-bound and -cancellation evidence and the same-environment R1 comparison. +Parity, resource-bound, and cancellation evidence: +[R2/R3 report](../r2-r3-performance/). diff --git a/docs/content/reference/versioned-performance.md b/docs/content/reference/versioned-performance.md index 7bd257ed..6bf32885 100644 --- a/docs/content/reference/versioned-performance.md +++ b/docs/content/reference/versioned-performance.md @@ -3,50 +3,49 @@ title: "V3 performance and acceptance report" description: "Reproducible million-keyword measurements on Redis and Valkey." --- -This is the archived R1 baseline. See [R2/R3 comparison](../r2-r3-performance/) -for the subsequent implementation and measurements. +# V3 performance and acceptance report -Measured on 2026-09-06. The final matrix completed **36 runs**: Redis and Valkey, -three dictionary sizes, three deterministic distributions, two repetitions each. -The separate million-entry safety scenario passed on both configured servers. -These are measurements of one environment, not throughput or memory guarantees. +The archived R1 baseline. For the implementation and measurements that followed, see the +[R2/R3 comparison](../r2-r3-performance/). -## Environment and methodology +Measured 2026-09-06. The final matrix completed **36 runs**: Redis and Valkey × three +dictionary sizes × three deterministic distributions × two repetitions. The separate +million-entry safety scenario passed on both servers. These are measurements of one +environment, not throughput or memory guarantees. + +## Environment and method - Apple M4, 10 logical CPUs, 16 GiB RAM; macOS 26.5.2, Darwin 25.5.0, arm64. - Go 1.26.7; Redis 8.10.1; source-built Valkey 9.1.2. -- Standalone TCP on localhost; RDB snapshots and AOF disabled. No production - durability or cross-node network overhead is included. -- MemoryEfficient engine; 50 ms version polling for the measurement (the library - default is 30 seconds); five-minute leases renewed every minute. -- Separate Go process per size/distribution/repetition; two complete repetitions. - This was a developer workstation, with some development checks running on the - host. Treat the ranges as baseline observations, not isolated microbenchmark - confidence intervals. Each server ran one scale workload at a time. -- Fixed seed 20260906. Shared: `shared-prefix-%08d`. Diverse: a random 12-bit hex - prefix plus a unique eight-digit index. Korean: `한국어-%04x-키워드-%08d`, using a - random 16-bit prefix. Full replacement appends `-x` to every keyword. - -**Valkey qualification:** an exploratory run of the local Valkey 9.1.2 build with -its default prefetch setting exited with SIGSEGV in `hashtableIncrementalFindStep` -via `prefetchCommandQueueKeys`. The final successful Valkey matrix used -`prefetch-batch-max-size 0`. This report does not establish support for that local -build's default configuration. The stack excerpt is preserved in -`benchmarks/results/v3-20260906-valkey-crash.txt`; the library returned an error -when the server disappeared. The root cause has not been established. - -The Valkey source archive SHA-256 was -`19c23908e7d57e8d91ef85b41f5646307582f10f4f0fb999bbf89ed24ec9c983` -(tag 9.1.2), built with `make -j4 valkey-server`. Prefetch was disabled using its -existing configuration option, without modifying server source. +- Standalone TCP on localhost; RDB snapshots and AOF disabled — no production durability + or cross-node network overhead is included. +- MemoryEfficient engine; 50 ms version polling for measurement (the library default is 30 + seconds); five-minute leases renewed every minute. +- A separate Go process per size/distribution/repetition, two complete repetitions, on a + developer workstation with some development checks running on the host. Treat the ranges + as baseline observations, not isolated microbenchmark confidence intervals. Each server + ran one scale workload at a time. +- Fixed seed 20260906. Shared: `shared-prefix-%08d`. Diverse: a random 12-bit hex prefix + plus a unique eight-digit index. Korean: `한국어-%04x-키워드-%08d` with a random 16-bit + prefix. Full replacement appends `-x` to every keyword. + +**Valkey qualification.** An exploratory run of the local Valkey 9.1.2 build with its +default prefetch setting exited with SIGSEGV in `hashtableIncrementalFindStep` via +`prefetchCommandQueueKeys`. The final successful Valkey matrix used +`prefetch-batch-max-size 0`, so this report does **not** establish support for that build's +default configuration. The stack excerpt is in +`benchmarks/results/v3-20260906-valkey-crash.txt`; the library returned an error when the +server disappeared, and the root cause has not been established. The Valkey source archive +SHA-256 was `19c23908e7d57e8d91ef85b41f5646307582f10f4f0fb999bbf89ed24ec9c983` (tag +9.1.2), built with `make -j4 valkey-server`. Prefetch was disabled through its existing +configuration option, without modifying server source. ## Initial load and startup -All table cells show the minimum–maximum across two runs. Load-ready is elapsed -from Replace to the local engine serving its version. Startup opens a new -collection instance against the existing dictionary and builds its first engine. -Peak RSS is the process lifetime high-water mark across the entire run, including -subsequent changes, rather than only initial load. +Cells show minimum–maximum across two runs. Load-ready is elapsed from `Replace` to the +local engine serving its version. Startup opens a new collection instance against the +existing dictionary and builds its first engine. Peak RSS is the process lifetime +high-water mark across the entire run, including subsequent changes. | Server | Keywords | Distribution | Load-ready (s) | Startup (s) | Peak RSS (GiB) | |---|---:|---|---:|---:|---:| @@ -71,11 +70,11 @@ subsequent changes, rather than only initial load. ## Million-keyword changes -Commit below includes normalization, bucket reads/preparation, manifest storage, -and final pointer commit. Ready also includes notification/polling and engine -rebuilding. Their difference is the calling instance's serving-version lag; -fleet-wide simultaneous activation is not asserted. Unchanged buckets are reused, -but every changed generation still rebuilds its entire local engine. +Commit covers normalization, bucket reads and preparation, manifest storage, and the final +pointer commit. Ready adds notification/polling and engine rebuilding; their difference is +the calling instance's serving-version lag, and fleet-wide simultaneous activation is not +asserted. Unchanged buckets are reused, but every changed generation still rebuilds its +entire local engine. | Server | Distribution | Add 1 commit / ready (s) | Add 1,000 ready (s) | Add 1% ready (s) | Full replace ready (s) | |---|---|---:|---:|---:|---:| @@ -86,19 +85,18 @@ but every changed generation still rebuilds its entire local engine. | valkey | diverse | 0.005–0.007 / 2.33–2.40 | 2.42–2.60 | 3.03–3.54 | 4.08–4.33 | | valkey | korean | 0.005–0.005 / 2.90–3.40 | 4.03–4.38 | 3.98–4.49 | 6.78–7.31 | -The JSON report also includes removals of 1, 1,000 and 1%, and identical Replace. -At 100,000 entries, 1,000 and 1% are the same count and are measured as two -successive add/remove cycles. Identical Replace still reads and compares buckets; -it does not create a generation or rebuild the engine. +The JSON report also covers removals of 1, 1,000 and 1%, and identical `Replace`. At +100,000 entries, 1,000 and 1% are the same count and are measured as two successive +add/remove cycles. Identical `Replace` still reads and compares buckets; it creates no +generation and rebuilds no engine. -## Search during replacement and retained storage +## Search during replacement, and retained storage -Search sampling uses a short string assembled from three dictionary keywords, at -approximately one-millisecond intervals while a write and its engine refresh run. -The timer includes that small string construction and the Find call. The old -engine continues serving until replacement; initial load starts with an empty -engine. These figures are not a general text-search throughput benchmark. The -following cells summarize full replacement at one million keywords. +Search sampling uses a short string assembled from three dictionary keywords, taken at +roughly one-millisecond intervals while a write and its engine refresh run. The timer +includes that string construction and the `Find` call. The old engine keeps serving until +replacement; initial load starts with an empty engine. These are not general text-search +throughput figures. The cells summarize full replacement at one million keywords. | Server | Distribution | Search p50 / p95 / p99 (µs) | Redis before prune (MiB) | Redis after prune (MiB) | |---|---|---:|---:|---:| @@ -109,49 +107,47 @@ following cells summarize full replacement at one million keywords. | valkey | diverse | 2.50–2.71 / 3.96–5.46 / 12.33–14.38 | 66.82–66.82 | 23.25–23.26 | | valkey | korean | 3.04–3.04 / 11.62–14.00 / 25.04–29.92 | 134.78–134.78 | 43.99–43.99 | -Redis memory is server-wide `used_memory`, including server overhead, receipts, -script cache, and retained manifests/chunks. The pruning measurement artificially -ages inactive generation registry timestamps beyond 24 hours; the active -generation remains protected. Actual retention is tested separately. Network byte -counters in the JSON include both writes and engine downloads/polling, as well as -protocol and measurement overhead; they are not just payload keyword bytes. +Redis memory is server-wide `used_memory`, including server overhead, receipts, script +cache, and retained manifests and chunks. The pruning measurement artificially ages +inactive generation registry timestamps beyond 24 hours; the active generation stays +protected, and actual retention is tested separately. Network byte counters in the JSON +include writes, engine downloads, polling, protocol, and measurement overhead — not just +payload keyword bytes. ## Acceptance evidence - The 36 scale runs completed initial loading, reopening, all paginated entries, - naive-reference search comparison, no-op/full replacements, 1/1,000/1% additions - and removals, WaitForVersion and explicit pruning. -- `TestVersionedMillionSafety` passed on both configured real servers: one of two - writers using the same million-entry expected version won; the other conflicted. - A pinned million-entry snapshot remained readable across commit and Prune, - contained no competing-generation keywords, and was collected only after close. -- Unit/race tests compare existing V2 engine search results with V3 for Korean, - multilingual, nested/overlapping, case-sensitive and stream-boundary cases. - Concurrent batch scans preserve a single engine generation. -- Fault injection covers chunk preparation errors, abandoned prepared generations, - lost commit responses and delayed retries, dropped Pub/Sub, failed engine loads, - reconnect, lease expiry, and a stale pruning fence. These are deterministic - simulations; they are not a production failover or forced-OS-process-kill test. -- Chunk-write instrumentation confirms an identical dictionary writes no chunks - and a single change writes only its affected bucket. The final Lua commit only - checks fixed keys, writes a receipt/marker and swaps the pointer; it does not - parse a manifest or iterate keywords. -- Root and server race suites, root/server lint, all-module vet, benchmark-module - tests, API snapshot/audit, documentation compilation, unchanged module manifests - and license attribution passed. FuzzFind and FuzzAdd ran for 30 seconds each; - V3 normalization fuzzing ran for 10 seconds. No production acor function remained - at zero coverage in the normal test suite. - -No real Cluster/Sentinel failover matrix or production persistence recovery was -run for V3 here. All collection keys use one hash tag and the existing connection -implementations; do not interpret standalone measurements as multi-node results. -R2 incremental download reuse and build-memory/cancellation improvements, and R3 -result limits/masking/replacement, remain follow-up roadmap work. + naive-reference search comparison, no-op and full replacements, 1/1,000/1% additions and + removals, `WaitForVersion`, and explicit pruning. +- `TestVersionedMillionSafety` passed on both real servers: of two writers using the same + million-entry expected version, one won and the other conflicted. A pinned million-entry + snapshot stayed readable across commit and `Prune`, contained no competing-generation + keywords, and was collected only after close. +- Unit and race tests compare V2 and V3 search results for Korean, multilingual, + nested/overlapping, case-sensitive, and stream-boundary cases. Concurrent batch scans + preserve a single engine generation. +- Fault injection covers chunk preparation errors, abandoned prepared generations, lost + commit responses and delayed retries, dropped Pub/Sub, failed engine loads, reconnect, + lease expiry, and a stale pruning fence. These are deterministic simulations, not a + production failover or forced process-kill test. +- Chunk-write instrumentation confirms an identical dictionary writes no chunks and a + single change writes only its affected bucket. The final Lua commit checks fixed keys, + writes a receipt and marker, and swaps the pointer — it parses no manifest and iterates + no keywords. +- Root and server race suites, root/server lint, all-module vet, benchmark-module tests, + API snapshot/audit, documentation compilation, unchanged module manifests, and license + attribution passed. FuzzFind and FuzzAdd ran 30 seconds each; V3 normalization fuzzing + ran 10 seconds. No production `acor` function remained at zero coverage in the normal + test suite. + +No real Cluster/Sentinel failover matrix or production persistence recovery was run for V3 +here. All collection keys use one hash tag and the existing connection implementations — +do not read standalone measurements as multi-node results. ## Reproduce -Use a disposable server with the same persistence, preset and polling settings. -From the repository root: +Use a disposable server with the same persistence, preset, and polling settings. From the +repository root: ```sh ACOR_V3_SCALE_ADDR=127.0.0.1:6379 \ @@ -159,11 +155,10 @@ ACOR_V3_SCALE_OUTPUT=/tmp/acor-v3-redis \ ACOR_V3_SCALE_REPEATS=2 make bench-v3 ``` -Repeat against Valkey with `prefetch-batch-max-size 0`, using a separate output -directory. `scripts/benchmark-v3.sh` compiles one test binary and starts a fresh -process for every size/distribution, then runs the million-entry safety scenario. -It removes only its uniquely named collection keys, never FLUSHDB. Do not run -another workload on the same endpoint during measurements: Redis memory/network -counters are server-wide. The raw operation summaries, environment strings, -startup/prune measurements and source hashes for this report are in -`benchmarks/results/v3-20260906.json`. +Repeat against Valkey with `prefetch-batch-max-size 0` and a separate output directory. +`scripts/benchmark-v3.sh` compiles one test binary, starts a fresh process for every +size/distribution, then runs the million-entry safety scenario. It removes only its own +uniquely named collection keys, never `FLUSHDB`. Do not run another workload on the same +endpoint during measurement — Redis memory and network counters are server-wide. Raw +operation summaries, environment strings, startup/prune measurements, and source hashes +are in `benchmarks/results/v3-20260906.json`. diff --git a/docs/content/reference/versioned.md b/docs/content/reference/versioned.md index 4bdf50e9..e2de4a1e 100644 --- a/docs/content/reference/versioned.md +++ b/docs/content/reference/versioned.md @@ -3,9 +3,11 @@ title: "Versioned dictionaries (V3)" description: "Large dictionary snapshots, atomic changes, and background search refresh." --- -V3 is an opt-in storage format in the existing Go module. Use `OpenVersioned` -and a new collection name. Existing `Create`, V1/V2 keys, and their error and -update contracts remain unchanged. V3 does not provide Suggest or a server API. +# Versioned dictionaries (V3) + +V3 is an opt-in storage format in the same Go module. Open it with `OpenVersioned` and a +**new** collection name. `Create`, the V1/V2 keys, and their error and update contracts +are untouched. V3 has no `Suggest` and no server API. ```go @@ -40,124 +42,136 @@ func main() { } ``` -## API contracts - -All I/O methods take a context. `Status`, `Snapshot.Version`, and `Snapshot.Count` -read local state. `Close` on the collection stops its background work; snapshot -`Close(ctx)` releases its lease. Initial loading fails the open operation. - -`Snapshot.List(ctx, cursor, limit)` returns up to a positive limit, ordered by -bucket number then lexical keyword order. Pass an empty cursor to start and stop -when `NextCursor` is empty. Keep the same snapshot for a complete traversal: -a cursor is bound to the generation. `Snapshot.Diff(ctx, target)` returns sorted -Added, Removed, and Retained entries without writing. Diff materializes the -comparison in memory. Count is the normalized, deduplicated keyword count. - -`Replace`, `Add`, `Remove`, `AddMany`, and `RemoveMany` require an expected Version. -Treat versions as opaque equality tokens, never ordered numbers or strings. -Concurrent writers using the same expected version cannot overwrite each other; -use `errors.Is(err, acor.ErrConcurrencyConflict)`. Batch changes are atomic. -Replace accepts an empty target. Reapplying an identical dictionary keeps the -version and does not publish invalidation. Keywords are trimmed and, unless -CaseSensitive was set at creation, lowercased. Blank keywords and invalid UTF-8 -fail the entire input. Duplicate normalized inputs are removed. The stored case -policy must match every subsequent open. - -A successful write means the Redis commit completed. The calling instance and -other instances may still search the previous engine. Use `WaitForVersion` to -wait for that commit or a later one; it honors cancellation. Status exposes the -observed active version, serving version, build start/duration, and recent error. -A failed refresh preserves the old serving engine. Polling defaults to 30 seconds, -with Pub/Sub accelerating discovery. A fixed 20 ms debounce window (RefreshDebounce) merges bursts before each -background build. Incoming events do not extend the window. The worker finishes -its current build, then loads the newest observed target. Every search, including a batch, -parallel scan, or stream, uses one engine. V3 reuses existing rune positions, -case folding, overlap and word-boundary semantics. - -`Find`, `FindIndex`, `FindMatches`, `FindSet`, `Contains`, `FindStream`, -`FindBatch`, `FindParallel`, and `FindIndexParallel` are available with explicit -contexts. Streaming keeps state across reader boundaries. The existing parallel -options still apply; use AutoOverlap for dictionary-length boundary protection. - -## Storage and failure handling +Every I/O method takes a context. `Status`, `Snapshot.Version`, and `Snapshot.Count` read +local state only. `Close` on the collection stops its background work; `Close(ctx)` on a +snapshot releases its lease. A failed initial load fails the open. + +## Reading + +`Snapshot.List(ctx, cursor, limit)` pages through keywords, ordered by bucket number then +lexically. Start with an empty cursor and stop when `NextCursor` is empty. **Keep the same +snapshot for a whole traversal** — a cursor is bound to its generation. + +`Snapshot.Diff(ctx, target)` returns sorted `Added`, `Removed`, and `Retained` without +writing anything, materializing the comparison in memory. `Count` is the normalized, +deduplicated keyword count. + +Search methods — `Find`, `FindIndex`, `FindMatches`, `FindSet`, `Contains`, `FindStream`, +`FindBatch`, `FindParallel`, `FindIndexParallel` — all take explicit contexts. Streaming +keeps state across reader boundaries, and the existing parallel options apply, including +`AutoOverlap` for boundary protection. V3 reuses the V1/V2 rune positions, case folding, +overlap, and word-boundary semantics. One search — batch, parallel, or streaming — always +uses one engine. + +## Writing + +`Replace`, `Add`, `Remove`, `AddMany`, and `RemoveMany` all require an expected `Version`. + +- **Versions are opaque equality tokens.** Never treat them as ordered numbers or strings. +- Concurrent writers holding the same expected version cannot overwrite each other; the + loser gets `errors.Is(err, acor.ErrConcurrencyConflict)`. +- Batch changes are atomic. `Replace` accepts an empty target. +- Reapplying an identical dictionary keeps the version and publishes no invalidation. +- Keywords are trimmed, and lowercased unless `CaseSensitive` was set at creation. + Duplicates after normalization are dropped; a blank keyword or invalid UTF-8 fails the + entire input. The stored case policy must match every subsequent open. + +### Commit is not yet visibility + +A successful write means the Redis commit landed. The calling instance and every other +instance may still be searching the previous engine. `WaitForVersion` waits for that +commit or a later one and honors cancellation. `Status` exposes the observed active +version, the serving version, build start and duration, and the most recent error. + +| Refresh behavior | | +|---|---| +| Discovery | Polling every 30 s by default, accelerated by Pub/Sub | +| Coalescing | A fixed 20 ms debounce window (`RefreshDebounce`) merges bursts; incoming events do not extend it | +| In-flight build | Finishes, then the worker loads the newest observed target | +| Failure | The old serving engine is preserved | + +## Storage layout SHA-256's first 12 bits select one of 4,096 fixed buckets. Each bucket is sorted, -deduplicated, and split into JSON string-array chunks of at most 1 MiB including -JSON escaping and delimiters. An oversized individual keyword gets its own chunk. -Content-addressed immutable chunks, generation manifests, metadata, and the active -pointer use separate keys. All keys share a name-digest Redis hash tag, hence one -Cluster slot. This does not distribute one collection across multiple nodes. -Connection options reuse existing standalone, Sentinel, Cluster and Ring clients. -Only connection fields and Name from VersionedOptions.Redis are used. - -Delta writes download and prepare affected buckets only; full replacements compare -all buckets and reuse unchanged ones. Local engine refresh reuses verified immutable keyword slices for unchanged -buckets and downloads changed buckets only. The previous cache remains usable -when a candidate fails. Full local engine rebuilding still requires memory for -the old engine, retained keyword slices, and the new engine. The default -preset is MemoryEfficient; Speed and Balanced are also selectable. No constant -memory or unmeasured latency guarantee is made. - -The final Lua commit checks the expected pointer, preparation lease and maintenance -lock, stores a receipt, and changes the pointer. It does not parse the manifest or -iterate keywords. Prepared data is unreachable until committed. Storage durability -and failover behavior still depend on the Redis/Valkey persistence configuration. - -On `ErrCommitUnknown`, retain the returned WriteResult.OperationID and use -`ResolveOperation(ctx, id)`. A found receipt is authoritative success. A missing -receipt is **not** proof that an in-flight request cannot commit later. Do not -blindly reapply ambiguous writes. Operation receipts and committed-version markers -are retained indefinitely, including no-op receipts, so they grow with write count. - -## Leases and explicit pruning - -Snapshots, builds and preparation writers use server-time leases (five minutes, -renewed every minute by default). Lease checks fail explicitly after expiration; -renewal cannot resurrect an expired handle. Close snapshots promptly. Existing -local searches need no Redis lease once their engine is built. - -`Prune(ctx)` retains the active generation, generations prepared or committed -within 24 hours, and generations protected by valid reader leases. It deletes -unreferenced chunks and unretained manifests in batches of at most 64. It also -collects abandoned preparation data. Receipts are not deleted. - -Prune acquires a monotonically increasing maintenance fence only when no valid -writer lease remains. New writers and snapshots temporarily return ErrMaintenance; -searches continue. Each deletion checks and extends lock ownership atomically. -The maintenance lock expires after 30 seconds if its process dies. An expired -pruner cannot delete after a successor takes ownership. Preparation itself is -fenced so an expired writer cannot create unregistered data during cleanup. -Large retention sets are enumerated in memory; there is no automatic collector. -R2 builds check cancellation throughout trie insertion, failure-link construction, -and table filling. Close cancels the active build; initialization honors its own -context. Cancellation discards a candidate without publishing it or leaving a -detached builder goroutine. Individual allocation, string length and sort -operations complete before the next checkpoint. Ordinary new commits do not -cancel an active build, avoiding starvation under continuous writes. - -## V2 cutover - -1. Use a **different** V3 name. Migrate V1 to V2 through the existing migration - first. `CopyV2(ctx, sourceName, expected, nil)` reads the V2 version and keywords in - one HMGET and replaces the V3 target. It reports source version, normalized - count and SHA-256 checksum of the sorted JSON keyword array. -2. For rehearsal, copy and compare count, checksum and representative search - results, including case-sensitive, Korean and overlapping matches. Set the - V3 case policy to the source application's policy (V2 does not persist it). -3. Stop V2 writes for the final copy. Verify its source version, count, checksum - and search results, then change application configuration to the new V3 name. - There is no automatic dual write or simultaneous fleet engine switch. -4. After V3 writes begin, switching configuration back to V2 loses those changes. - Export and apply the V3 changes back to V2 before rollback; a simple name change - is insufficient. Keep V2 untouched until the migration has been accepted. +deduplicated, and split into JSON string-array chunks of at most 1 MiB including escaping +and delimiters; an oversized single keyword gets its own chunk. Content-addressed +immutable chunks, generation manifests, metadata, and the active pointer live under +separate keys. + +**All keys share one name-digest hash tag, so a collection occupies one Cluster slot.** +V3 does not distribute a collection across nodes. Connection options reuse the existing +standalone, Sentinel, Cluster, and Ring clients; only the connection fields and `Name` +from `VersionedOptions.Redis` are used. + +Delta writes download and prepare affected buckets only; a full replacement compares every +bucket and reuses the unchanged ones. Local engine refresh reuses verified immutable +keyword slices for unchanged buckets and downloads only changed ones, and the previous +cache stays usable when a candidate fails. A full rebuild still needs memory for the old +engine, the retained slices, and the new engine at once. The default preset is +`MemoryEfficient`; `Speed` and `Balanced` are selectable. No constant-memory or latency +guarantee is made. + +The final Lua commit checks the expected pointer, the preparation lease, and the +maintenance lock, stores a receipt, and swaps the pointer. It does not parse the manifest +or iterate keywords. Prepared data is unreachable until committed. Durability and failover +still depend on your Redis/Valkey persistence configuration. + +### Ambiguous commits + +On `ErrCommitUnknown`, keep the returned `WriteResult.OperationID` and call +`ResolveOperation(ctx, id)`. A found receipt is authoritative success. **A missing receipt +is not proof that an in-flight request cannot still commit** — do not blindly reapply an +ambiguous write. Receipts and committed-version markers are retained indefinitely, +including no-op receipts, so they grow with write count. + +## Leases and pruning + +Snapshots, builds, and preparation writers hold server-time leases: five minutes, renewed +every minute by default. Lease checks fail explicitly after expiration, and renewal cannot +resurrect an expired handle — so close snapshots promptly. A local search needs no lease +once its engine is built. + +`Prune(ctx)` retains the active generation, anything prepared or committed within 24 +hours, and anything protected by a valid reader lease. It deletes unreferenced chunks and +unretained manifests in batches of at most 64, and collects abandoned preparation data. +Receipts are never deleted, and there is no automatic collector; large retention sets are +enumerated in memory. + +Pruning takes a monotonically increasing maintenance fence, and only when no valid writer +lease remains. While it holds, new writers and snapshots get `ErrMaintenance` and searches +continue. Every deletion checks and extends lock ownership atomically; the lock expires +after 30 seconds if its process dies, and an expired pruner cannot delete once a successor +takes ownership. Preparation is fenced too, so an expired writer cannot create +unregistered data during cleanup. + +Builds check cancellation throughout trie insertion, failure-link construction, and table +filling. `Close` cancels the active build; cancellation discards the candidate without +publishing it or leaving a detached goroutine. Ordinary new commits do not cancel an +active build, which is what stops continuous writes from starving refresh. + +## Cutting over from V2 + +1. **Use a different name.** Migrate V1 to V2 through the existing migration first. + `CopyV2(ctx, sourceName, expected, nil)` reads the V2 version and keywords in one + `HMGET` and replaces the V3 target, reporting source version, normalized count, and the + SHA-256 checksum of the sorted JSON keyword array. +2. **Rehearse.** Copy, then compare count, checksum, and representative search results — + case-sensitive, Korean, and overlapping matches. Set the V3 case policy to the source + application's policy; V2 does not persist it. +3. **Cut over.** Stop V2 writes for the final copy, verify source version, count, + checksum, and search results, then point the application at the new V3 name. There is + no automatic dual write and no simultaneous fleet-wide engine switch. +4. **Rollback is not a name change.** Once V3 writes begin, switching back to V2 loses + them. Export and apply the V3 changes to V2 first, and keep V2 untouched until the + migration is accepted. ## CLI -Global connection/name flags precede `dictionary`; its flags follow the subcommand. -Input to diff/replace is one JSON string array on stdin. Empty copy-v2 sources -also require `--allow-empty`; the guard runs before replacing the target. Status and other results -are JSON. List is one page per invocation; concurrent generation changes reject a -previous page cursor. Library callers should retain a Snapshot for a long export. +Global connection and name flags come **before** `dictionary`; the subcommand's own flags +follow it. Diff and replace read one JSON string array from stdin. Empty replacements and +empty `copy-v2` sources need `--allow-empty`, checked before the target is replaced. +Results are JSON. `list` returns one page per invocation, and a concurrent generation +change rejects an older page cursor — hold a `Snapshot` in library code for a long export. ```sh acor -name filter-v3 dictionary status @@ -169,9 +183,11 @@ acor -name filter-v3 dictionary copy-v2 --source filter-v2 --expected-version TO acor -name filter-v3 dictionary prune ``` -R1 validation and measured resource costs are recorded in -[the reproducible V3 performance report](../versioned-performance/). R2 adds incremental download reuse, bounded coalescing, cancellable builds, and -more compact sparse-engine nodes. R3 adds bounded source-position search and -atomic masking/replacement; see [bounded text processing](../text-processing/) and -[the R2/R3 verification report](../r2-r3-performance/). R1 measurements remain -archived as the baseline; none of these reports are performance guarantees. +## Measurements + +[R1 validation and resource costs](../versioned-performance/) are the archived baseline. +R2 added incremental download reuse, bounded coalescing, cancellable builds, and more +compact sparse-engine nodes; R3 added bounded source-position search and atomic +masking/replacement ([bounded text processing](../text-processing/), +[R2/R3 verification](../r2-r3-performance/)). None of these reports is a performance +guarantee. diff --git a/docs/content/server/_index.md b/docs/content/server/_index.md index 28aae984..7b88bbb4 100644 --- a/docs/content/server/_index.md +++ b/docs/content/server/_index.md @@ -18,27 +18,21 @@ the middleware that makes either observable. > directive, so without a pin you get whichever core version the server module's `require` > names. -## ACOR ships no server binary +## There is no server binary -There is no `acor-server` to install, and no image to pull. `acor/server` is a library: it -hands you an `http.Handler`, a `*grpc.Server`, and middleware, and **you** write the `main` -that wires them to a collection and listens. +No `acor-server` to install, no image to pull. `acor/server` is a library — it hands you an +`http.Handler`, a `*grpc.Server`, and middleware, and **you** write the `main` that wires +them to a collection and listens. -That is a deliberate consequence of the module being experimental — a published binary is a -contract, and this module does not offer one yet. What it offers instead is that the wiring -is about forty lines, and [Running a Server](running/) is those forty lines, complete and -copy-pasteable. +That follows from the module being experimental: a published binary is a contract, and this +module does not offer one yet. What it offers instead is that the wiring is about forty +lines, and [Running a Server](running/) is those forty lines, complete and copy-pasteable. ## Sections -- [Running a Server](running/) - Wire a collection to HTTP or gRPC, with readiness checks and clean shutdown -- [HTTP API](http-api/) - The nine JSON endpoints, their request and response shapes, and every error they return -- [gRPC API](grpc-api/) - The `acor.server.v1.Acor` service, its eight RPCs, and the observability constructors +- [Running a Server](running/) — the `main` for each protocol, with readiness checks and clean shutdown +- [HTTP API](http-api/) — nine JSON endpoints, their shapes, and every error they return +- [gRPC API](grpc-api/) — the `acor.server.v1.Acor` service and the observability constructors Metrics, structured logging, and tracing are configured the same way whichever protocol you -serve, so they live together under -[Operations → Monitoring](../operations/monitoring/). - -## Navigation - -← [Operations](../operations/) | [CLI](../cli/) → +serve, so they live in [Operations → Monitoring](../operations/monitoring/). diff --git a/docs/content/server/grpc-api.md b/docs/content/server/grpc-api.md index 3d1a5ebe..b29994d4 100644 --- a/docs/content/server/grpc-api.md +++ b/docs/content/server/grpc-api.md @@ -7,13 +7,11 @@ weight: 3 The service is `acor.server.v1.Acor`, defined in [`server/proto/acor/v1/acor.proto`](https://github.com/skyoo2003/acor/blob/main/server/proto/acor/v1/acor.proto). -It mirrors the [HTTP API](../http-api/) one for one. +It mirrors the [HTTP API](../http-api/) one for one. The `main` that serves it is +[Running a Server](../running/). -> **The `acor/server` module is experimental.** This service definition is **not covered by -> the core module's compatibility promise** and can change in any release. See the -> [section overview](../) for what that means for your `go.mod`. - -See [Running a Server](../running/) for the `main` that serves it. +> The `acor/server` module is [experimental and separately versioned](../) — this service +> definition can change in any release. ## Methods @@ -28,93 +26,53 @@ See [Running a Server](../running/) for the `main` that serves it. | `Info` | `EmptyRequest` | `InfoResponse{keywords, nodes}` | | `Flush` | `EmptyRequest` | `StatusResponse{status}` | -All eight are unary. Full method names are `/acor.server.v1.Acor/`. - -### Two shapes differ from HTTP - -**Counts are `int64`.** `CountResponse.count`, `InfoResponse.keywords`, and -`InfoResponse.nodes` are `int64` on the wire, where the Go API and the JSON API use `int`. +All eight are unary; full method names are `/acor.server.v1.Acor/`. -**Match offsets are wrapped.** proto3 maps cannot hold a repeated value, so -`MatchIndexesResponse.matches` is `map` rather than the JSON API's -`map[string][]int`: +Two shapes differ from HTTP: -```proto -message Positions { - repeated int64 positions = 1; -} +- **Counts are `int64`.** `CountResponse.count`, `InfoResponse.keywords`, and + `InfoResponse.nodes` are `int64` on the wire, where the Go and JSON APIs use `int`. +- **Match offsets are wrapped.** proto3 maps cannot hold a repeated value, so + `MatchIndexesResponse.matches` is `map`: -message MatchIndexesResponse { - map matches = 1; -} -``` + ```proto + message Positions { + repeated int64 positions = 1; + } + ``` -A Go client reads `resp.GetMatches()["redis"].GetPositions()`. + A Go client reads `resp.GetMatches()["redis"].GetPositions()`. -**Positions are rune offsets, not byte offsets**, and `SuggestIndex` always answers `[0]` -because it matches the input as a prefix rather than searching for it. Both behaviors come -from the collection, not the transport, so they are identical on both surfaces — -[the HTTP page works through them with examples](../http-api/). +Positions are rune offsets, and `SuggestIndex` always answers `[0]` because it matches the +input as a prefix. Both behaviors come from the collection rather than the transport, so +they are identical on both surfaces — [the HTTP page works through them with +examples](../http-api/). ## Errors Every error from the collection becomes `codes.Internal` with the error's own text as the -message. This is the same flattening the HTTP API does with its blanket `500`: a caller -mistake and a Redis outage arrive as the same code, and telling them apart means matching on -message text, which is not part of any promise. - -No RPC returns `InvalidArgument`, `NotFound`, or `FailedPrecondition` — a write to a V1 -read-only collection is `Internal`, not `FailedPrecondition`. +message — the same flattening as the HTTP blanket `500`. No RPC returns `InvalidArgument`, +`NotFound`, or `FailedPrecondition`; a write to a read-only V1 collection is `Internal`. +Telling a caller mistake from a Redis outage means matching on message text, which is not +part of any promise. ## Deadlines do not cancel the work -**A client deadline or a disconnect ends the RPC, not the Redis operation behind it.** +**A client deadline or disconnect ends the RPC, not the Redis operation behind it.** `Service` declares no context on any method (`server/grpc.go` calls `s.service.Add(req.GetKeyword())`), and `(*AhoCorasick).Add` runs against the collection's -own long-lived context, not the caller's. The request context is accepted by each adapter -and discarded. +own long-lived context. The request context is accepted by each adapter and discarded. -The consequence to plan for: a client that gives up on `Add`, `Remove`, or `Flush` and sees -`DeadlineExceeded` or `Canceled` has **not** prevented the write. It may already have -landed, or land shortly after. `Flush` deletes every key in the collection, and abandoning -the call does not call it off. +So a client that gives up on `Add`, `Remove`, or `Flush` and sees `DeadlineExceeded` has +**not** prevented the write: - Do not treat a timeout as "it did not happen". Re-read with `Info` or `Find` before retrying anything that mutates. -- Client-side retries on timeout can double-apply. `Add` and `Remove` are idempotent by - keyword, so a repeat is harmless there; a retried `Flush` is a second flush. -- Setting a short deadline does not shed load on the server. The Redis work continues at - full cost after the client is gone. - -The [HTTP API](../http-api/) behaves identically — same `Service` interface, same discard. - -## Generating a client - -There is no `buf.yaml` or `protoc` configuration in this repository. Two options: - -- **Go clients** import the generated package directly and skip codegen entirely: - - ```go - import acorv1 "github.com/skyoo2003/acor/server/proto/acor/v1" - - client := acorv1.NewAcorClient(conn) - ``` - -- **Other languages** run their own `protoc`/`buf` against `acor.proto`. It has no imports - beyond proto3 itself, so the single file is the whole input — but `--_out` alone - generates only the **message** types. The service stub needs that language's gRPC plugin - as a second output flag: - - ```sh - # Python: --python_out gives acor_pb2.py only; the stub needs grpcio-tools. - python -m grpc_tools.protoc -I. --python_out=. --grpc_python_out=. acor.proto - - # C++/Java/etc. take the plugin the same way: - protoc -I. --cpp_out=. --grpc_out=. --plugin=protoc-gen-grpc=$(which grpc_cpp_plugin) acor.proto - ``` +- Retries can double-apply. `Add` and `Remove` are idempotent by keyword; a retried `Flush` + is a second flush. +- A short deadline sheds no server load. - Omit the gRPC flag and you get the eight request/response messages with nothing able to - call the eight RPCs. +The [HTTP API](../http-api/) behaves identically. ## Constructors @@ -124,52 +82,67 @@ server.NewGRPCServerWithObservability(ctx, service, obs, opts...) // + observabi ``` Both accept any `grpc.ServerOption`, including `grpc.Creds` for TLS, and both leave -`Serve`/`Stop` to you. - -`Observability` bundles four pillars, and **any field may be nil to skip that one**: +`Serve`/`Stop` to you. `Observability` bundles four pillars, and **any field may be nil to +skip it**: | Field | Wires in | Notes | | ----- | -------- | ----- | | `Tracer` | `otelgrpc` stats handler | Configure with `tracing.NewTracer` | -| `Metrics` | `grpc_server_*` Prometheus interceptor | gRPC has no `/metrics` route — expose `promhttp.Handler()` on a separate HTTP listener | +| `Metrics` | `grpc_server_*` Prometheus interceptor | gRPC has no `/metrics` route — expose `promhttp.Handler()` on a separate listener | | `Logger` | zerolog unary interceptor | JSON request logs | | `Health` | the standard `grpc.health.v1` service | See below | -Metric names and tracing configuration live in -[Operations → Monitoring](../../operations/monitoring/); this page does not restate them. +Metric names and tracing configuration are in +[Operations → Monitoring](../../operations/monitoring/). ## Health checking Passing `Health` registers the standard `grpc.health.v1.Health` service, so -`grpc_health_probe` and Kubernetes gRPC probes work without extra code. Behavior worth -knowing before you set a probe interval: +`grpc_health_probe` and Kubernetes gRPC probes work with no extra code. Four things to know +before setting a probe interval: -- **Status is polled, not computed per request.** An HTTP `/readyz` probe runs your checkers - on each request; a gRPC health probe reads whatever the last poll left behind. Probing - more often than every 5 seconds buys no extra resolution. +- **Status is polled, not computed per request.** A gRPC probe reads whatever the last poll + left behind, so probing more often than every 5 seconds buys no resolution. - **The 5-second tick bounds neither staleness nor load.** The poller calls - `checker.Check()` inline on a single long-lived `time.Ticker`. That ticker keeps ticking - while a check is blocked, and its channel buffers one tick, so a check that overruns the - interval is followed *immediately* by the next one rather than after a fresh 5-second - gap. A checker that takes 30 seconds therefore makes the status at least 30 seconds stale - **and** polls back-to-back for as long as it stays slow — the opposite of the breather - the interval suggests, and it lands on Redis hardest when Redis is already the slow part. - A checker that hangs outright is worse: the blocked goroutine cannot act on context - cancellation either, so shutdown stops marking the server `NOT_SERVING`. Give every - checker its own deadline; the examples in [Running a Server](../running/) use a 2-second - `context.WithTimeout` for exactly this. -- **Both the overall server (`""`) and `acor.server.v1.Acor` are tracked**, and they carry - the same status. -- **A nil `Health` field means no health service at all.** The constructor only registers - `grpc.health.v1` when the field is set (`server/grpc.go:77`), so leaving it nil makes a - probe fail with `UNIMPLEMENTED`, not `SERVING`. To get a health service that always - answers `SERVING`, pass an empty `health.NewChecker()` rather than nil — the - nil-*checker*-means-`SERVING` rule lives inside `RegisterGRPCHealthServer` and is only - reachable by calling that function yourself. -- **The `ctx` you pass to `NewGRPCServerWithObservability` bounds the poller.** Cancel it - and the server is marked `NOT_SERVING`, which is what lets a load balancer drain - connections before `GracefulStop` finishes. - -## Navigation - -← [HTTP API](../http-api/) | [CLI](../../cli/) → + `checker.Check()` inline on one long-lived `time.Ticker`, which keeps ticking while a + check is blocked and buffers one tick — so a check that overruns the interval is followed + *immediately* by the next. A checker taking 30 seconds makes the status at least 30 + seconds stale **and** polls back-to-back for as long as it stays slow, hitting Redis + hardest when Redis is already the slow part. A checker that hangs outright cannot act on + context cancellation either, so shutdown stops marking the server `NOT_SERVING`. Give + every checker its own deadline — [Running a Server](../running/) uses 2 seconds. +- **A nil `Health` field means no health service at all**, so a probe fails with + `UNIMPLEMENTED`, not `SERVING` (`server/grpc.go:77`). For a service that always answers + `SERVING`, pass an empty `health.NewChecker()`; the nil-*checker*-means-`SERVING` rule + lives inside `RegisterGRPCHealthServer` and is reachable only by calling that yourself. +- **The `ctx` passed to `NewGRPCServerWithObservability` bounds the poller.** Cancelling it + marks the server `NOT_SERVING`, which is what lets a load balancer drain connections + before `GracefulStop` finishes. + +Both the overall server (`""`) and `acor.server.v1.Acor` are tracked, and they carry the +same status. + +## Generating a client + +There is no `buf.yaml` or `protoc` configuration in this repository. + +- **Go clients** import the generated package and skip codegen entirely: + + ```go + import acorv1 "github.com/skyoo2003/acor/server/proto/acor/v1" + + client := acorv1.NewAcorClient(conn) + ``` + +- **Other languages** run their own `protoc`/`buf` against `acor.proto`. It has no imports + beyond proto3, so the single file is the whole input — but `--_out` alone generates + only the **message** types. The service stub needs that language's gRPC plugin as a + second output flag: + + ```sh + # Python: --python_out gives acor_pb2.py only; the stub needs grpcio-tools. + python -m grpc_tools.protoc -I. --python_out=. --grpc_python_out=. acor.proto + + # C++/Java/etc. take the plugin the same way: + protoc -I. --cpp_out=. --grpc_out=. --plugin=protoc-gen-grpc=$(which grpc_cpp_plugin) acor.proto + ``` diff --git a/docs/content/server/http-api.md b/docs/content/server/http-api.md index 58bd3b5b..10a2851e 100644 --- a/docs/content/server/http-api.md +++ b/docs/content/server/http-api.md @@ -6,15 +6,11 @@ weight: 2 # HTTP API `server.NewHTTPHandler(service)` returns an `http.Handler` serving nine routes. Every -response the handler itself produces is JSON with `Content-Type: application/json`; the -exceptions are the two `ServeMux`-level responses noted under -[Not every response is JSON](#not-every-response-is-json). +response the handler itself produces is JSON; the two exceptions come from `ServeMux` +below. The `main` that mounts it is [Running a Server](../running/). -> **The `acor/server` module is experimental.** These paths and shapes are **not covered by -> the core module's compatibility promise** and can change in any release. See the -> [section overview](../) for what that means for your `go.mod`. - -See [Running a Server](../running/) for the `main` that mounts this handler. +> The `acor/server` module is [experimental and separately versioned](../) — these paths +> and shapes can change in any release. ## Endpoints @@ -26,148 +22,108 @@ See [Running a Server](../running/) for the `main` that mounts this handler. | `POST` | `/v1/find` | `{"input":"..."}` | `{"matches":["..."]}` | | `POST` | `/v1/find-index` | `{"input":"..."}` | `{"matches":{"kw":[0,12]}}` | | `POST` | `/v1/suggest` | `{"input":"..."}` | `{"matches":["..."]}` | -| `POST` | `/v1/suggest-index` | `{"input":"..."}` | `{"matches":{"kw":[0]}}` — always `[0]`, see below | +| `POST` | `/v1/suggest-index` | `{"input":"..."}` | `{"matches":{"kw":[0]}}` — always `[0]` | | `GET` | `/v1/info` | — | `{"keywords":3,"nodes":7}` | | `POST` | `/v1/flush` | — | `{"status":"ok"}` | `count` is how many keywords the operation actually changed, so a second `add` of the same -keyword answers `{"count":0}`. +keyword answers `{"count":0}`. `/v1/flush` takes no body, does not read one if you send it, +and deletes every key in the collection. The method column is enforced: anything else is +`405`. + +```sh +curl -sX POST localhost:8080/v1/add -d '{"keyword":"redis"}' +# {"count":1} -### Offsets are rune offsets, and the two `*-index` routes do not mean the same thing +curl -sX POST localhost:8080/v1/find -d '{"input":"redis-backed matching"}' +# {"matches":["redis"]} -`/v1/find-index` returns, per keyword, the positions where it matched. Those positions are -counted in **runes, not bytes**: +curl -sX POST localhost:8080/v1/find-index -d '{"input":"redis and redis"}' +# {"matches":{"redis":[0,10]}} -```sh -curl -sX POST localhost:8080/v1/find-index -d '{"input":"한글 레디스 and redis"}' -# {"matches":{"redis":[11],"레디스":[3]}} +curl -s localhost:8080/v1/info +# {"keywords":1,"nodes":6} ``` -`레디스` starts at rune 3 but byte 7; `redis` starts at rune 11 but byte 21. The response -gives 3 and 11. A client that slices the original string by these numbers must count -runes — Go `[]rune`, Python `str` indices, Rust `chars()`. JavaScript needs care in both -directions, because its string indices are UTF-16 code units: `[...s]` gives code points, -and anything outside the BMP counts as two units under `.slice()`. +### The two `*-index` routes do not mean the same thing -`/v1/suggest-index` looks like the same shape but is not. `Suggest` matches the input as a -*prefix* of each keyword, so every match starts at the beginning by construction and the -implementation assigns exactly `[0]` to every suggestion. It never reports a second -position: +`/v1/find-index` returns, per keyword, the positions where it matched — counted in +**runes, not bytes**: ```sh -curl -sX POST localhost:8080/v1/suggest-index -d '{"input":"red"}' -# {"matches":{"redis":[0]}} +curl -sX POST localhost:8080/v1/find-index -d '{"input":"한글 레디스 and redis"}' +# {"matches":{"redis":[11],"레디스":[3]}} ``` -Treat it as "these keywords start with your input", not as a position list. - -`/v1/flush` takes no request body and does not read one if you send it. It deletes every -key in the collection. +`레디스` starts at rune 3 but byte 7; `redis` at rune 11 but byte 21. A client slicing the +original string by these numbers must count runes — Go `[]rune`, Python `str` indices, +Rust `chars()`. JavaScript needs care in both directions, since its string indices are +UTF-16 code units: `[...s]` gives code points, and anything outside the BMP counts as two +units under `.slice()`. -The method column is enforced, not advisory: `/healthz` and `/v1/info` are `GET`-only and -everything else is `POST`-only. Any other method gets `405`. +`/v1/suggest-index` looks like the same shape but is not. `Suggest` matches the input as a +*prefix*, so every match starts at the beginning by construction and the implementation +assigns exactly `[0]` to every suggestion. Read it as "these keywords start with your +input", not as a position list. ## Errors -Every failure returns `{"error":""}` with `Content-Type: application/json`. +Every failure the handler produces is `{"error":""}` with +`Content-Type: application/json`. | Status | When | Body | | ------ | ---- | ---- | -| `400` | The body is not valid JSON | `{"error":"unexpected EOF"}` — the message comes from `encoding/json` and is not a stable string | -| `400` | The body holds more than one JSON value | `{"error":"request body must contain only a single JSON value"}` | +| `400` | Body is not valid JSON | `{"error":"unexpected EOF"}` — from `encoding/json`, not a stable string | +| `400` | Body holds more than one JSON value | `{"error":"request body must contain only a single JSON value"}` | | `405` | Wrong method for the path | `{"error":"method not allowed"}` | | `413` | Reading the body reaches the 1 MiB cap | `{"error":"request body must not be larger than 1048576 bytes"}` | | `500` | Any error from the underlying collection | `{"error":""}` | -| `404` | No such path | **`text/plain`**, body `404 page not found` | -| `301` | The path needs canonicalizing (`/v1//info`) | **`text/html`**, Go's `Moved Permanently` page | - -The body is read as it is decoded, so whichever fault surfaces first is the one you get. A -body over 1 MiB that is *also* malformed comes back as the `400`, because the decoder -reaches the bad byte before the reader reaches the cap. `413` means the cap is what stopped -it, not that every oversized body is reported that way — nothing short of reading the whole -body could promise the latter, and reading it is what the cap exists to avoid. - -Two of these deserve more than a table row. - -### Every collection error is a `500` - -There is no error taxonomy. The handler passes any non-nil error from the collection -straight to a `500`, so a client mistake and a Redis outage are indistinguishable by status -code. Writing to a V1 collection — which the core rejects with `ErrV1ReadOnly`, a caller -error — comes back as `500 {"error":"V1 collections are read-only; migrate with MigrateV1ToV2"}`, not as a `4xx`. - -If your client needs to tell "retry this" from "fix your request", it has to match on the -message text, and message text is not part of any promise. Treat `5xx` here as -*something went wrong*, not as *the server is unhealthy*. - -For the latter, use `/readyz` — but note that **`/readyz` is not part of this handler**. -`NewHTTPHandler` and `NewHTTPServer` register only the routes in the table above, so -requesting `/readyz` from either gets the `404`. The route exists only if you mount -`health.RegisterHTTPHandlers` on an outer mux yourself, as -[Running a Server](../running/) does. - -### Not every response is JSON - -Two cases are answered by Go's `http.ServeMux` before ACOR's handler sees them, and neither -is JSON: - -- **`404 page not found`** — `text/plain`, for a path that matches nothing. -- **`301 Moved Permanently`** — `text/html`, for a path that needs canonicalizing. - `/v1//info`, `/v1/./info`, and `/v1/../v1/info` all redirect to `/v1/info`. - -A client that unconditionally parses responses as JSON breaks on both. The redirect is the -worse of the two: a client that follows redirects automatically may downgrade the `POST` -to a `GET` and then receive a `405`, turning a doubled slash into a confusing method error. -Normalize paths before sending. +| `404` | No such path | **`text/plain`**, `404 page not found` | +| `301` | Path needs canonicalizing (`/v1//info`) | **`text/html`**, Go's `Moved Permanently` page | + +The body is read as it is decoded, so whichever fault surfaces first is the one you get: an +oversized body that is *also* malformed comes back as the `400`, because the decoder +reaches the bad byte before the reader reaches the cap. + +**Every collection error is a `500`.** There is no error taxonomy — a client mistake and a +Redis outage are indistinguishable by status code. Writing to a V1 collection, which the +core rejects with `ErrV1ReadOnly`, arrives as +`500 {"error":"V1 collections are read-only; migrate with MigrateV1ToV2"}`, not a `4xx`. +Telling "retry this" from "fix your request" means matching on message text, and message +text is not part of any promise. + +**`/readyz` is not part of this handler.** `NewHTTPHandler` and `NewHTTPServer` register +only the routes above, so `/readyz` returns the `404`. It exists only if you mount +`health.RegisterHTTPHandlers` on an outer mux, as [Running a Server](../running/) does. + +**Not every response is JSON.** The `404` and the `301` are answered by Go's +`http.ServeMux` before ACOR's handler sees them. `/v1//info`, `/v1/./info`, and +`/v1/../v1/info` all redirect to `/v1/info` — and a client that follows redirects +automatically may downgrade the `POST` to a `GET` and then receive a `405`, turning a +doubled slash into a confusing method error. Normalize paths before sending. ## Disconnecting does not cancel the work -**A client timeout or a dropped connection ends the request, not the Redis operation behind +**A client timeout or dropped connection ends the request, not the Redis operation behind it.** Each handler passes `r.Context()` into its adapter and the adapter discards it -(`server/server.go:96` and its siblings take `_ context.Context`); `Service` declares no +(`server/server.go:96` and siblings take `_ context.Context`); `Service` declares no context on any method, and `(*AhoCorasick).Add` runs against the collection's own -long-lived context rather than the caller's. +long-lived context. So a client that gives up on `/v1/add`, `/v1/remove`, or `/v1/flush` has **not** prevented -the write. It may already have landed, or land shortly after. `/v1/flush` deletes every key -in the collection, and hanging up does not call it off. +the write: - Do not read a timeout as "it did not happen". Re-read with `/v1/info` or `/v1/find` before retrying anything that mutates. -- Retrying on timeout can double-apply. `add` and `remove` are idempotent by keyword, so a - repeat is harmless; a retried `flush` is a second flush. -- A short client timeout sheds no server load. The Redis work continues at full cost after - the client is gone. +- Retrying can double-apply. `add` and `remove` are idempotent by keyword; a retried + `flush` is a second flush. +- A short client timeout sheds no server load — the Redis work continues at full cost. -The [gRPC API](../grpc-api/) behaves identically — same `Service` interface, same discard. +The [gRPC API](../grpc-api/) behaves identically: same `Service` interface, same discard. ## What the server does not check -- **Request `Content-Type` is ignored.** Any value is accepted as long as the body parses - as JSON. +- **Request `Content-Type` is ignored** — any value is accepted if the body parses as JSON. - **Unknown JSON fields are ignored.** `{"keyword":"x","bogus":1}` succeeds. -- **Missing fields become the zero value.** `POST /v1/add` with `{}` is an `Add("")`; the - server does not reject it, and what happens next is whatever the collection does with an - empty keyword. -- **There is no authentication or authorization.** `/v1/flush` is reachable by anyone who - can reach the port. - -## Example - -```sh -curl -sX POST localhost:8080/v1/add -d '{"keyword":"redis"}' -# {"count":1} - -curl -sX POST localhost:8080/v1/find -d '{"input":"redis-backed matching"}' -# {"matches":["redis"]} - -curl -sX POST localhost:8080/v1/find-index -d '{"input":"redis and redis"}' -# {"matches":{"redis":[0,10]}} - -curl -s localhost:8080/v1/info -# {"keywords":1,"nodes":6} -``` - -## Navigation - -← [Running a Server](../running/) | [gRPC API](../grpc-api/) → +- **Missing fields become the zero value.** `POST /v1/add` with `{}` is an `Add("")`. +- **There is no authentication.** `/v1/flush` is reachable by anyone who can reach the port. diff --git a/docs/content/server/running.md b/docs/content/server/running.md index 88f1a07e..bc4b4a4e 100644 --- a/docs/content/server/running.md +++ b/docs/content/server/running.md @@ -6,29 +6,25 @@ weight: 1 # Running a Server ACOR ships no server binary. `acor/server` gives you an `http.Handler` and a -`*grpc.Server`; the `main` that wires them to a collection and listens is yours to write. -This page is that `main`, in full, for each protocol. +`*grpc.Server`; the `main` that wires them to a collection and listens is yours. This page +is that `main`, in full, for each protocol. -> **The `acor/server` module is experimental.** It publishes no version tags of its own and -> is **not covered by the core module's compatibility promise**. See the -> [section overview](../) for what that means for your `go.mod`. - -## Dependencies +> The `acor/server` module is [experimental and separately versioned](../). ```sh go get github.com/skyoo2003/acor/server go get github.com/skyoo2003/acor/pkg/acor@latest ``` -Both lines matter. `acor/server` resolves to a pseudo-version from `main`, and it carries a +Both lines matter. `acor/server` resolves to a pseudo-version from `main` and carries a `require` on the core module that Go will not override from the dependency's own `replace` -directive — so name the core version yourself, in your own `go.mod`. +directive — so name the core version yourself. ## HTTP -The collection is the service. Every method `server.Service` requires — `Add`, `Remove`, +The collection *is* the service. Every method `server.Service` requires — `Add`, `Remove`, `Find`, `FindIndex`, `Suggest`, `SuggestIndex`, `Flush`, `Info` — is already a method on -`*acor.AhoCorasick`, so the collection satisfies the interface with no adapter. +`*acor.AhoCorasick`, so no adapter is needed. ```go @@ -51,12 +47,11 @@ import ( // redisChecker reports whether the collection can still reach Redis. // -// Info() is the only exported call that proves the Redis path works, but it is -// not free: on V2 it HGETALLs the trie hash and unmarshals the whole keyword -// and prefix arrays just to count them, so its cost grows with the dictionary. -// See "Readiness costs what Info() costs" below before pointing a probe at it. +// Info() is the only exported call that proves the Redis path works, and it is +// not free: on V2 it HGETALLs the trie hash and unmarshals the whole keyword and +// prefix arrays just to count them. See "Readiness costs what Info() costs". // -// The timeout is not tidiness. Check() runs inline in both probe paths, so a +// The timeout is load-bearing: Check() runs inline in both probe paths, so a // checker that blocks blocks the prober. type redisChecker struct{ ac *acor.AhoCorasick } @@ -122,13 +117,12 @@ func main() { `server.NewHTTPServer(addr, service)` is a one-liner that does most of this, and it is the right choice if you do not need readiness checks. What it cannot give you is `/readyz`: -`NewHTTPHandler` builds its own private `http.ServeMux` internally and returns it as an -`http.Handler`, so there is no mux for you to register anything else on. +`NewHTTPHandler` builds its own private `ServeMux` internally, so there is nothing for you +to register on. -Composing them on an outer mux, as above, works and does not panic — the two `/healthz` -registrations are on different muxes and never meet. On the outer mux, Go routes `/healthz` -to the `health` package's handler because an exact pattern outranks the `/` catch-all. The -practical effect: +Composing them on an outer mux works and does not panic — the two `/healthz` registrations +live on different muxes and never meet, and on the outer mux Go routes `/healthz` to the +`health` package because an exact pattern outranks the `/` catch-all: | Path | Served by | Body | | ---- | --------- | ---- | @@ -137,8 +131,8 @@ practical effect: | `/v1/*` | `server.NewHTTPHandler` | see [HTTP API](../http-api/) | The two `/healthz` implementations return the same body for a `GET`, so the shadowing costs -you nothing. They differ only in how they reject a non-`GET`: `server/health` replies in -`text/plain` via `http.Error`, the API's own replies in JSON. +nothing. They differ only in rejecting a non-`GET`: `server/health` replies `text/plain` +via `http.Error`, the API's own replies in JSON. ## gRPC @@ -162,9 +156,9 @@ import ( "github.com/skyoo2003/acor/server/metrics" ) -// The deadline matters more here than on the HTTP side: the gRPC health -// poller calls Check() inline on its ticker, so a checker that blocks stalls -// every later poll and the poller's own response to cancellation. +// The deadline matters more here than over HTTP: the gRPC health poller calls +// Check() inline on its ticker, so a checker that blocks stalls every later poll +// and the poller's own response to cancellation. type redisChecker struct{ ac *acor.AhoCorasick } func (c redisChecker) Check() health.CheckResult { @@ -232,63 +226,48 @@ func main() { } ``` -Every field of `Observability` is optional: a nil field skips that pillar, so you can start -with logging only and add the rest later. `server.NewGRPCServer(service, opts...)` is the -same server with no observability at all. +Every field of `Observability` is optional — a nil field skips that pillar, so you can +start with logging only. `server.NewGRPCServer(service, opts...)` is the same server with +none of it. -Prometheus metrics registered here are *collected* but not *exposed* — gRPC has no -`/metrics` endpoint. Serve `promhttp.Handler()` on a separate HTTP listener; see -[Operations → Monitoring](../../operations/monitoring/). +Prometheus metrics registered here are *collected* but not *exposed*: gRPC has no +`/metrics` endpoint. Serve `promhttp.Handler()` on a separate HTTP listener, per +[Monitoring](../../operations/monitoring/). ## Readiness costs what `Info()` costs `Info()` is the only exported call that touches Redis and returns quickly *on a small -collection*, which is why it is the readiness check here. It is not a ping. On the default -V2 schema it runs `HGETALL` against the trie hash and JSON-unmarshals the complete keyword -and prefix arrays in order to return their two lengths, so its cost and its allocations -scale with the whole dictionary. +collection*, which is why it is the readiness check here. It is not a ping: on V2 it runs +`HGETALL` against the trie hash and unmarshals the complete keyword and prefix arrays to +return their two lengths, so cost and allocations scale with the whole dictionary. Two things multiply that: - **`/readyz` runs the checkers on every request**, and it is unauthenticated. - **The gRPC health poller runs them every 5 seconds**, whether or not anyone is probing — - and *more* often than that once a check gets slow. The poller's `time.Ticker` keeps - ticking while `Check` is blocked and buffers one tick, so an overrunning check is followed - immediately by the next. Slow checks are not throttled; they compound. + and *more* often once a check gets slow. Its `time.Ticker` keeps ticking while `Check` + is blocked and buffers one tick, so an overrunning check is followed immediately by the + next. Slow checks are not throttled; they compound. On a large dictionary that is significant Redis traffic and garbage, generated hardest -exactly when the service is already struggling. If that describes your collection, make the -**readiness** check a direct `redis.Client.Ping` against the same address instead — that is -a `PING`, not a dictionary scan — and accept that it proves connectivity rather than -collection health. - -Keep it out of liveness either way. `/healthz` answers "is this process alive", and nothing -about Redis belongs in that answer: a Redis outage would fail every replica's liveness probe -at once and have the orchestrator restart all of them, which cannot repair Redis and drops -whatever the processes were still serving. Redis reachability is a readiness signal — -take me out of the load balancer — not a restart signal. - -The 2-second timeout is load-bearing either way. `HealthChecker.Check` calls every -registered checker inline, and the gRPC poller calls `Check` inline on its ticker, so one -checker that hangs stalls every later poll **and** the poller's response to context -cancellation. A checker without its own deadline turns a slow Redis into a stuck health -service. +exactly when the service is already struggling. If that describes your collection, make +the **readiness** check a direct `redis.Client.Ping` against the same address — a `PING`, +not a dictionary scan — and accept that it proves connectivity rather than collection +health. -## What you still have to decide +Keep it out of liveness either way. `/healthz` answers "is this process alive", and a +Redis outage that failed every replica's liveness probe would restart all of them and +repair nothing. Redis reachability is a readiness signal. -The examples above hard-code answers this page cannot make for you: +## What you still have to decide 1. **Listen address.** `:8080` and `:9090` are placeholders. -2. **Redis credentials and topology.** `Addr` is standalone. Sentinel and Cluster use - `Addrs`, and Sentinel also needs `MasterName` — see - [Operations → Deployment](../../operations/deployment/). -3. **TLS.** Neither constructor configures it. For gRPC, pass `grpc.Creds(...)` as a - `grpc.ServerOption`; for HTTP, use `ListenAndServeTLS` or terminate at your ingress. +2. **Redis credentials and topology.** `Addr` is standalone; Sentinel and Cluster use + `Addrs`, and Sentinel also needs `MasterName`. See + [Deployment](../../operations/deployment/). +3. **TLS.** Neither constructor configures it. For gRPC pass `grpc.Creds(...)`; for HTTP + use `ListenAndServeTLS` or terminate at your ingress. 4. **Authentication.** There is none. Both surfaces expose `/v1/flush` and `Flush`, which delete every key in the collection. Do not put either on a network you do not control. -5. **Which protocol to serve.** They are independent; run one, the other, or both on +5. **Which protocol to serve.** They are independent — run one, the other, or both on separate listeners. - -## Navigation - -← [Server](../) | [HTTP API](../http-api/) →