diff --git a/.github/workflows/catalog-integrity.yml b/.github/workflows/catalog-integrity.yml new file mode 100644 index 0000000..de8ce2a --- /dev/null +++ b/.github/workflows/catalog-integrity.yml @@ -0,0 +1,22 @@ +name: catalog-integrity + +on: + pull_request: + branches: ["main"] + push: + branches: ["main"] + workflow_dispatch: + +permissions: + contents: read + +jobs: + catalog-integrity: + runs-on: ubuntu-24.04 + timeout-minutes: 5 + steps: + - uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 + with: + persist-credentials: false + - name: Check catalog structure + run: python3 scripts/check_catalog.py diff --git a/AGENTS.md b/AGENTS.md index e3c84a0..b119ee7 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -2,14 +2,29 @@ Machine-facing rules for agents using this repository. -1. Read `README4AI.md` and `CATALOG.md` before applying an optimization elsewhere. -2. Treat optimization records as patterns, not universal parameter sets. -3. Preserve reference semantics and add/retain conformance tests for optimized paths. -4. Prefer deterministic, bounded reuse over opaque caches. -5. If equivalence enables reuse, encode the equivalence as a named invariant and test it directly. -6. For Lean caches, distinguish verified reuse from cold reconstruction in both implementation and claims. -7. For parallel work, retain deterministic output ordering and verify scalar/parallel equivalence. -8. For real-time/DSP work, separate slow control work from hot sample/block work when semantics allow it; avoid allocations and synchronization on the hot path. -9. `power_module.md` contains both implemented ideas and aspirational performance language. Check the corresponding code/evidence before promoting a claim. -10. `suxen.zip` is a source candidate, not validated evidence. Use `scripts/inventory_zip.py` in an environment with the archive bytes available before extracting optimization claims. -11. New records must state status, source identity, preserved contract, evidence, limitations, and rollback conditions. +1. Read `README4AI.md`, `OPTIMIZATION-PROBLEM.md`, and `CATALOG.md` before applying an optimization elsewhere. +2. Define the target optimization contract before selecting a mechanism: search space, feasible set, objective/direction, correctness constraints, evaluation budget and stopping rule. +3. Treat optimization records as patterns, not universal parameter sets. +4. Preserve reference semantics and add/retain conformance tests for optimized paths. +5. Prefer deterministic, bounded reuse over opaque caches. +6. If equivalence enables reuse, encode the equivalence as a named invariant and test it directly. +7. For incremental reuse, bind the complete effective input identity and persist a new reusable state only after successful completion. +8. For coalescing, merge only semantically equivalent in-flight work; define cancellation/error semantics explicitly. +9. For Lean caches, distinguish verified reuse from cold reconstruction in implementation and claims. +10. For parallel work, retain deterministic output ordering where required and verify scalar/parallel equivalence. Measure effective concurrency rather than assuming requested workers ran concurrently. +11. For adaptive search, preserve the trial ledger and evaluation budget. Remember that excessive parallel batch width can reduce information efficiency. +12. For approximation, state an explicit error/degradation contract and keep an exact/reference path where practical. Never silently weaken exact semantics. +13. For pruning, test bound soundness independently; never prune on a heuristic presented as proof. +14. For critical-path/speculative work, ensure speculation cannot expose side effects before commitment and does not starve the actual critical path. +15. For performance budgets, characterize benchmark noise/environment before enforcing a threshold. +16. For real-time/DSP work, separate slow control work from hot sample/block work when semantics allow it; avoid allocations and synchronization on the hot path. +17. For SIMD/native specialization, require reference parity plus actual code-generation evidence and repeated target-host measurement. Never infer end-to-end speedup merely from wider instructions or an isolated probe. +18. For SoA/tiling, bound temporary storage per worker, prove packed-field and reduction equivalence, and re-profile tile sizes on the target cache/memory hierarchy rather than copying donor values. +19. For persistent worker pools, test repeated-dispatch state isolation and account for startup/teardown amortization. Physical/logical worker selection is not evidence of CPU affinity or NUMA placement. +20. For automatic execution-path promotion, calibrate only already-correct candidates on workload-shaped samples, include lifecycle/tuning costs, require a material promotion margin, preserve canonical/manual control, and fail closed against an independent oracle on parity mismatch. +21. `power_module.md` contains both implemented ideas and aspirational performance language. Check corresponding code/evidence before promoting a claim. +22. `suxen.zip` is a source candidate, not validated evidence. Inventory and read relevant source before extracting optimization claims. +23. The three pinned v1 Lean model files are immutable historical formalization. New records do not become formally proved by association; version future formal modules separately. +24. New post-v1 records must state status, source identity, optimization problem contract, preserved contract, validation, limitations and rollback conditions. +25. Run `python3 scripts/check_catalog.py` after catalog changes. +26. Treat `scripts/check_catalog.py` as the public integrity entrypoint; `scripts/check_catalog_normalizer.py` is its internal CommonMark-normalization helper and must not be invoked as a substitute gate. diff --git a/CATALOG.md b/CATALOG.md index 21866ee..057e6d6 100644 --- a/CATALOG.md +++ b/CATALOG.md @@ -2,53 +2,126 @@ ## Quick decision table -| Bottleneck | First record to inspect | Core idea | +| Bottleneck / problem shape | First record to inspect | Core idea | | --- | --- | --- | -| Test suite spends most time in deterministic sweeps/simulations | `OPT-PY-001` | Reduce redundant work while keeping coverage semantics and deterministic assertions | -| Same expensive result is recomputed at a provably equivalent parameter/state | `OPT-INV-001` | Prove equivalence, then reuse the already-computed result | -| Lean CI repeatedly rebuilds an unchanged dependency closure | `OPT-LEAN-001` | Reuse only cryptographically/structurally verified dependency state; always rebuild project source | -| Independent jobs/items can execute concurrently | `OPT-PAR-001` | Bound workers, preserve deterministic ordering, prove scalar/parallel equivalence | -| Expensive state changes far slower than the sample/hot-loop rate | `OPT-DSP-001` | Move state evolution to control rate; sparse-evaluate couplings; batch/vectorize the hot path | - -## Records +| Deterministic tests/sweeps dominate runtime | [OPT-PY-001](optimizations/OPT-PY-001-deterministic-test-execution.md) | Reduce redundant/high-cost work while keeping coverage semantics | +| Same expensive result is recomputed at a proven-equivalent state | [OPT-INV-001](optimizations/OPT-INV-001-invariant-driven-reuse.md) | Prove equivalence, then reuse | +| Lean dependency reconstruction dominates CI | [OPT-LEAN-001](optimizations/OPT-LEAN-001-trust-preserving-lean-ci.md) | Verify reusable dependency state; rebuild current project source | +| Independent work can execute concurrently | [OPT-PAR-001](optimizations/OPT-PAR-001-bounded-parallel-execution.md) | Bound workers and prove scalar/parallel equivalence | +| Slow control state is inside a high-rate numerical/audio loop | [OPT-DSP-001](optimizations/OPT-DSP-001-control-rate-sparse-vector-dsp.md) | Separate rates, sparse-evaluate, vectorize | +| Inputs are unchanged but pipeline stages rerun | [OPT-INC-001](optimizations/OPT-INC-001-signature-bound-incremental-execution.md) | Bind work to complete input signatures and persist only successful state | +| Many simultaneous callers request identical not-yet-computed work | [OPT-COAL-001](optimizations/OPT-COAL-001-concurrent-duplicate-work-coalescing.md) | One in-flight computation, many waiters | +| Integer sets alternate between sparse and dense regions | [OPT-SET-001](optimizations/OPT-SET-001-density-adaptive-compact-sets.md) | Density-adaptive representation with exact set algebra | +| One global lock/counter/runtime domain serializes independent work | [OPT-CONT-001](optimizations/OPT-CONT-001-partitioned-coordination-domains.md) | Partition coordination while preserving the global invariant | +| Same deterministic transform is repeated for every consumer/replay | [OPT-FAN-001](optimizations/OPT-FAN-001-shared-materialization-fanout.md) | Materialize once, reuse many times | +| Expensive parameter evaluations are being guessed or exhaustively swept | [OPT-SEARCH-001](optimizations/OPT-SEARCH-001-budget-aware-adaptive-search.md) | Adaptive, budget-aware search over the declared problem contract | +| Exactness may be traded inside an explicit quality envelope | [OPT-APPROX-001](optimizations/OPT-APPROX-001-contract-bounded-approximation.md) | Bound the error/degradation and the resource cost together | +| Expensive stages consume candidates later discarded | [OPT-REDUCE-001](optimizations/OPT-REDUCE-001-early-working-set-reduction.md) | Reduce the working set before composition | +| Non-critical work delays the dependency chain users actually wait on | [OPT-CRIT-001](optimizations/OPT-CRIT-001-critical-path-prioritization.md) | Prioritize the critical path; speculate/defer deliberately | +| Small performance regressions accumulate unnoticed | [OPT-BUDGET-001](optimizations/OPT-BUDGET-001-performance-regression-budgets.md) | Guard stable performance expectations in CI | +| Discrete search space is huge but optimistic bounds are available | [OPT-PRUNE-001](optimizations/OPT-PRUNE-001-bound-driven-search-space-pruning.md) | Prune regions that provably cannot beat the incumbent | +| A deterministic hot loop is not exploiting useful host vector instructions | [OPT-SIMD-001](optimizations/OPT-SIMD-001-evidence-gated-native-autovectorization.md) | Reshape for autovectorization, prove parity, inspect codegen, then measure native benefit | +| Large AoS traversal wastes cache/memory and only a bounded hot subset is needed at once | [OPT-SOA-001](optimizations/OPT-SOA-001-worker-local-soa-tiling.md) | Transform bounded per-worker tiles into SoA and reuse cache-local scratch | +| Repeated parallel runs keep paying thread/buffer startup or misuse SMT topology | [OPT-POOL-001](optimizations/OPT-POOL-001-persistent-topology-aware-worker-pools.md) | Persist workers/buffers and choose physical/logical worker policy explicitly | +| Several exact execution paths trade places across hosts or workload shapes | [OPT-AUTO-001](optimizations/OPT-AUTO-001-calibrated-host-aware-path-promotion.md) | Calibrate bounded candidates, include lifecycle cost, require margin and fail-closed oracle parity | + +Before selecting a record, define the target problem using [`OPTIMIZATION-PROBLEM.md`](OPTIMIZATION-PROBLEM.md). + +## Frozen v1 records ### OPT-PY-001 — Deterministic test execution **Status:** Verified mechanism; historical performance context incomplete. -QEC combined minimal fixtures, vectorized assertions, bounded deterministic caching, convergence/cycle early exit, smaller high-cost sweeps, lower safe iteration/trial counts, and repeated-work removal. QEC v68.4.0 reports about 126 s → 46 s (~2.7×) with 3779 passed / 8 skipped; v68.4.1 reports about 40 s after hardening. The cited v68.x records do not preserve runner/CPU/Python/pytest/repetition metadata, so these are historical observations, not transferable benchmark targets. +QEC combined minimal fixtures, vectorized assertions, bounded deterministic caching, convergence/cycle early exit, smaller high-cost sweeps, lower safe iteration/trial counts, and repeated-work removal. Historical timings remain source observations, not transferable targets. ### OPT-INV-001 — Invariant-driven computation reuse **Status:** Verified mechanism; historical performance context incomplete. -QEC formalized the baseline equivalence `URW(min_sum, rho=1.0) == baseline min-sum`, tested exact equality, centralized the predicate, and reused the baseline result rather than rerunning the benchmark. The implementation commit reports about 43% speedup for that hot test, but does not preserve its runner/toolchain, exact hot-test wall times, or repetitions. Re-measure before making a target-repo speed claim. +QEC encoded a baseline equivalence, tested exact equality and reused a proven-equivalent baseline result rather than rerunning the benchmark. ### OPT-LEAN-001 — Trust-preserving Lean CI -**Status:** Verified on QSOL-GEO-REASON PR #3 source lane; timing observations are environment-scoped. +**Status:** Verified on source PR; timing observations are environment-scoped. -Separates source-state cache identity from compiled dependency artifacts, verifies both before use, rebuilds the current project source, and keeps a no-cache `cold-trust` lane for release-grade reconstruction claims. The cache policy records a 2501.52 s cold dependency build using four Lean threads on a four-CPU `ubuntu-24.04` / x86_64 lane; a later verified-cache run records an 8.47 s GeoReason project build on a four-CPU Ubuntu 24.04.4 hosted runner. These are single observations of different scopes, with no exact CPU model/repetition distribution preserved, so they must not be divided into a portable speedup. +Separates source-state identity from compiled dependency artifacts, verifies reuse, rebuilds current project source, and keeps cold reconstruction claims separate. ### OPT-PAR-001 — Bounded parallel execution -**Status:** Verified, environment-specific; performance must be re-measured before transfer. +**Status:** Verified, environment-specific. -The QEC-validated NEXUS v4.0.1 qBraid evidence compared scalar and worker-count variants, checked output invariants, and recorded observed thread behavior. Its archived seven-worker observation was made on qBraid / Ubuntu 24.04.4 / AMD EPYC 7763 with 16 logical CPUs and effective worker capacity 7. The canonical receipt does not bind the performance samples to an exact `rustc --version`, so OPT retains the numbers as historical evidence rather than a transferable performance target. +Compares scalar and worker-count variants, checks output invariants and treats measured effective parallelism as evidence rather than assuming requested workers were used. ### OPT-DSP-001 — Control-rate sparse vector DSP **Status:** Implemented reference for control-rate/sparse/vector patterns; approximation/native ideas partly proposed. -The SPECTRAL NumPy reference is pinned to commit `5265b7f130287f80b5cf0d3de5bb2953152f90cd`. It precomputes static state, evolves E8/qutrit control state at ~1 kHz, computes all root phases once per control step, uses a sparse root subset per node, and vectorizes block synthesis. The audition renderer's whole-block modulation shortcut has no defined equivalence/error contract and is therefore not promoted as a reusable correctness-preserving optimization. `power_module.md` additionally proposes block SIMD, zero-copy buffers and lock-free/native Rust structures; those native performance claims are not promoted as verified here. +Precompute static state, evolve slow control state less often, evaluate sparse couplings and batch/vectorize hot numerical work. -## Composition guidance +The five records above are the immutable v1.0.0 formalized catalog. Their Lean model remains pinned; post-v1 records below do not silently alter it. + +## Post-v1 records + +### OPT-INC-001 — Signature-bound incremental execution +Persist complete effective-input identity after successful work and skip a stage only while that identity and required outputs remain valid. + +### OPT-COAL-001 — Concurrent duplicate-work coalescing +Merge equivalent simultaneous misses into one in-flight computation instead of letting a thundering herd duplicate upstream work. + +### OPT-SET-001 — Density-adaptive compact sets +Partition an integer domain and use sparse or bitmap-like containers according to local density while keeping exact set semantics. + +### OPT-CONT-001 — Partitioned coordination domains +Replace one hot global coordination point with independently advancing domains while preserving required cross-domain invariants. + +### OPT-FAN-001 — Shared materialization for fan-out and replay +Perform a deterministic transform once at the production boundary, persist/retain it when justified, and reuse it across consumers and replay. + +### OPT-SEARCH-001 — Budget-aware adaptive parameter search +Classify the optimization problem, maintain a trial ledger, adapt future evaluations from observations, and stop under an explicit evaluation/resource budget. -Optimizations compose only when their resource models do. In particular: +### OPT-APPROX-001 — Contract-bounded approximation +Permit approximation only when the interface/scientific contract explicitly defines an error or degradation envelope and a reference path exists where practical. -- pytest process parallelism plus BLAS/NumPy threads can oversubscribe CPUs; -- Lean parallel workers plus large dependency cache restore can raise memory and I/O pressure; -- a DSP block that is vectorized but allocates an `nodes × samples` temporary may still be unsuitable for hard real-time use; -- caching an equivalent result is safe only while the equivalence invariant remains true. +### OPT-REDUCE-001 — Early working-set reduction +Push semantics-preserving filtering/culling/selection ahead of joins, rendering, simulation, DSP or other expensive composition. + +### OPT-CRIT-001 — Critical-path prioritization +Prioritize work on the true latency dependency chain; prefetch/precompute likely-soon work only when justified; defer non-critical work. + +### OPT-BUDGET-001 — Performance regression budgets +Protect a stable benchmark expectation with an environment-scoped, variance-aware CI budget rather than relying on remembered performance. + +### OPT-PRUNE-001 — Bound-driven search-space pruning +Maintain a feasible incumbent, derive optimistic bounds for subregions, and discard regions that provably cannot improve the incumbent. + +### OPT-SIMD-001 — Evidence-gated native autovectorization +Expose independent batch lanes to the compiler, compare portable/native builds from identical source, prove exact parity and inspect emitted instructions before integrating a measured specialized path. + +### OPT-SOA-001 — Worker-local SoA tiling +Transform only bounded worker-local chunks of a large AoS population into hot-field SoA tiles, reuse local scratch, and retain deterministic reference/reduction semantics. + +### OPT-POOL-001 — Persistent topology-aware worker pools +Create workers and local buffers once for repeated dispatches, expose physical/logical worker policy explicitly, and account for lifecycle amortization rather than timing only the steady-state kernel. + +### OPT-AUTO-001 — Calibrated host-aware path promotion +Choose among already-correct execution paths using workload-shaped live calibration, lifecycle-aware scoring, a material promotion margin and full-work fail-closed oracle verification. + +## Composition guidance -Prefer one measured bottleneck removal at a time, then re-profile. +Optimizations compose only when their semantic and resource models compose. + +- process parallelism plus BLAS/NumPy/native threads can oversubscribe CPUs; +- coalescing reduces duplicate identical work while adaptive search may instead need to diversify independent in-flight experiments; +- a compact representation may make a formerly remote problem feasible in memory, changing the architecture rather than merely reducing bytes; +- critical-path speculation can steal resources from the path it was intended to accelerate; +- approximation must never leak into an API whose callers still assume exact semantics; +- adaptive search can lose information efficiency when parallel batches are too wide; +- performance budgets require controlled environments or statistically defensible noise handling; +- pruning is valid only when the bound is sound; +- SIMD and thread-level parallelism can move the bottleneck to memory bandwidth or CPU frequency limits; +- SoA tiling and persistent pools multiply worker-local storage by worker count, so cache/RSS behavior must be re-measured together; +- host-auto selection must calibrate only candidates that already satisfy their own correctness contracts and must not convert a selector heuristic into a universal hardware ranking. + +Prefer one measured bottleneck removal at a time, then re-profile and reconsider the problem contract. diff --git a/OPTIMIZATION-PROBLEM.md b/OPTIMIZATION-PROBLEM.md new file mode 100644 index 0000000..6958dec --- /dev/null +++ b/OPTIMIZATION-PROBLEM.md @@ -0,0 +1,81 @@ +# Optimization Problem Contract + +OPT separates **the problem being optimized** from **the mechanism used to search for an improvement**. + +For a target repository, define the optimization problem before selecting an optimization record. + +## Canonical contract + +Represent a problem as + +\[ +P = (X, F, f, d, C, B, S) +\] + +where: + +- `X` — search space / decision-variable domain; +- `F ⊆ X` — feasible set after hard constraints; +- `f : F → R^k` — measured objective or objective vector; +- `d` — objective direction (`minimize`, `maximize`, or explicit multi-objective ordering); +- `C` — correctness and semantic contract that may not be weakened implicitly; +- `B` — evaluation/resource budget; +- `S` — stopping rule. + +A candidate is admissible only if it lies in `F` **and** satisfies `C`. A faster candidate that violates `C` is not an optimization under the same problem definition. + +## Required classification + +Record the following before tuning. The dimension names below are the canonical field names used by `templates/OPTIMIZATION-RECORD.md` and `scripts/check_catalog.py`: + +| Dimension | Typical values | +| --- | --- | +| Variables | continuous / integer / categorical / conditional / mixed | +| Search scope | local / global | +| Objective behavior | deterministic / noisy / stochastic | +| Information | gradient available / derivative-free / black-box | +| Evaluation cost | cheap / moderate / expensive | +| Constraints | bounds / equality / inequality / semantic / resource | +| Parallelism | sequential / synchronous batch / asynchronous | +| Exactness | exact / approximation permitted under an explicit error contract | + +## Examples + +### Worker-count tuning + +- `X = {1, …, 32}` +- `F = X` subject to peak-memory and platform limits +- `f(x) = median wall time` +- `d = minimize` +- `C = scalar/parallel result equivalence + deterministic required ordering` +- `B = 40 benchmark trials` +- `S = budget exhausted or improvement below the predeclared threshold` + +### Approximate visualization + +- `X = {LOD policies}` +- `F = policies satisfying frame-memory limits and error ≤ ε` +- `f = (frame latency, perceptual/error metric)` +- `d = lexicographic: first require error ≤ ε through F/C, then minimize frame latency; break equal-latency ties by lower error` +- `C = reference path remains available and declared visual/semantic invariants are preserved` +- `B = at most 200 policy evaluations over the fixed benchmark fixture set` +- `S = stop when B is exhausted or no admissible policy improves frame latency by the predeclared δ for 20 consecutive evaluations` + +## Search-mechanism selection + +Use the problem classification to select a mechanism: + +- expensive black-box continuous or mixed tuning → `OPT-SEARCH-001`; +- discrete search with provable optimistic bounds → `OPT-PRUNE-001`; +- independent work that can run concurrently → `OPT-PAR-001`; +- repeated equivalent work → `OPT-INV-001`; +- simultaneous identical in-flight work → `OPT-COAL-001`; +- approximation explicitly permitted → `OPT-APPROX-001`. + +The mechanism is subordinate to the contract. Do not reshape the problem after seeing results merely to make an optimization look successful. + +## Measurement rule + +Source-project constants and historical observations are priors, not targets. Transfer requires fresh target-context measurement and validation. + +See `FORMALIZATION.md` for the frozen v1.0.0 Lean boundary. This problem-contract layer is post-v1 catalog guidance and does not mutate the pinned v1 formal model. diff --git a/README.md b/README.md index 6275e94..dfcd9bc 100644 --- a/README.md +++ b/README.md @@ -1,44 +1,72 @@ # OPT — QSOL Optimization Catalog -Reusable, provenance-linked optimization patterns extracted from QSOL projects. +Reusable, provenance-linked optimization patterns extracted from QSOL projects and carefully bounded external donors. -The point of this repository is simple: when a future project needs to go faster, use less memory, avoid redundant work, or shorten CI without weakening correctness, point the implementing agent here first. +The point of this repository is simple: when a future project needs to go faster, use less memory, avoid redundant work, shorten CI, or tune an expensive system, point the implementing agent here first — **without weakening correctness to make a benchmark look good**. ## Rules of the vault 1. **Correctness outranks speed.** An optimization must preserve the contract it claims to preserve. -2. **Measured and proposed work are different things.** Records say which is which. -3. **Keep the reference path.** Optimized/native/parallel paths should have a deterministic reference or conformance gate whenever practical. -4. **Do not cargo-cult constants.** Trial counts, worker caps, cache keys, tolerances, hashes, and block sizes belong to their source environment until re-measured. -5. **Provenance matters.** Every promoted optimization links back to the code, release, PR, or evidence that established it. +2. **Define the problem before choosing the trick.** Use [`OPTIMIZATION-PROBLEM.md`](OPTIMIZATION-PROBLEM.md) for the search space, feasible set, objective, constraints, budget and stopping rule. +3. **Measured and proposed work are different things.** Records say which is which. +4. **Keep the reference path.** Optimized/native/parallel/approximate paths should have a deterministic reference or conformance gate whenever practical. +5. **Do not cargo-cult constants.** Trial counts, worker caps, cache keys, tolerances, hashes, block sizes, thresholds and search parameters belong to their source environment until re-measured. +6. **Provenance matters.** Every promoted optimization links back to the code, release, PR, paper, or source that established it. ## Catalog -| ID | Optimization | Status | Source | Evidence snapshot | -| --- | --- | --- | --- | --- | -| [OPT-PY-001](optimizations/OPT-PY-001-deterministic-test-execution.md) | Deterministic test execution | **Verified mechanism; benchmark context incomplete** | QEC v68.2.0–v68.4.1 | ~126 s → ~46 s in v68.4.0 release; ~40 s after v68.4.1; original runner/toolchain/repetition context not preserved, so re-benchmark before transfer | -| [OPT-INV-001](optimizations/OPT-INV-001-invariant-driven-reuse.md) | Invariant-driven computation reuse | **Verified mechanism; benchmark context incomplete** | QEC v68.4.1 cycle | Eliminated a redundant benchmark at a proven-equivalent baseline point; ~43% source-reported hot-test improvement, with original environment/timing samples unavailable | -| [OPT-LEAN-001](optimizations/OPT-LEAN-001-trust-preserving-lean-ci.md) | Trust-preserving Lean dependency reuse | **Verified on source PR; timings environment-scoped** | QSOL-GEO-REASON PR #3 | 2501.52 s cold dependency reconstruction and 8.47 s verified-cache project rebuild are different-scope single observations on four-CPU Ubuntu lanes; see record | -| [OPT-PAR-001](optimizations/OPT-PAR-001-bounded-parallel-execution.md) | Bounded deterministic parallel execution | **Verified, environment-specific** | QEC v170.2.1 / NEXUS evidence | qBraid / AMD EPYC 7763 observation: 69.694063 ns/eval scalar median → 18.310215 ns/eval at 7 workers; re-benchmark before transfer | -| [OPT-DSP-001](optimizations/OPT-DSP-001-control-rate-sparse-vector-dsp.md) | Control-rate + sparse + vectorized DSP | **Implemented reference; approximation/native ideas proposed** | SPECTRAL commit `5265b7f…` + `power_module.md` | Control-rate decimation, sparse E8 coupling, shared phase computation and NumPy vectorization are implemented; whole-block audition approximation and native SIMD/zero-copy/lock-free claims are not promoted | +| ID | Optimization | Status | Core idea | +| --- | --- | --- | --- | +| [OPT-PY-001](optimizations/OPT-PY-001-deterministic-test-execution.md) | Deterministic test execution | **Verified mechanism; benchmark context incomplete** | Reduce repeated/high-cost test work without weakening coverage semantics | +| [OPT-INV-001](optimizations/OPT-INV-001-invariant-driven-reuse.md) | Invariant-driven computation reuse | **Verified mechanism; benchmark context incomplete** | Prove equivalence, then reuse the existing result | +| [OPT-LEAN-001](optimizations/OPT-LEAN-001-trust-preserving-lean-ci.md) | Trust-preserving Lean dependency reuse | **Verified on source PR; timings environment-scoped** | Reuse verified dependency state while rebuilding current project source | +| [OPT-PAR-001](optimizations/OPT-PAR-001-bounded-parallel-execution.md) | Bounded deterministic parallel execution | **Verified, environment-specific** | Bound concurrency and prove scalar/parallel equivalence | +| [OPT-DSP-001](optimizations/OPT-DSP-001-control-rate-sparse-vector-dsp.md) | Control-rate + sparse + vectorized DSP | **Implemented reference; approximation/native ideas proposed** | Move slow state out of the hot path; sparse/vectorize repeated numerical work | +| [OPT-INC-001](optimizations/OPT-INC-001-signature-bound-incremental-execution.md) | Signature-bound incremental execution | **Implemented external reference** | Rerun work only when complete effective-input identity changes | +| [OPT-COAL-001](optimizations/OPT-COAL-001-concurrent-duplicate-work-coalescing.md) | Concurrent duplicate-work coalescing | **Implemented external reference** | Share one in-flight computation among equivalent simultaneous callers | +| [OPT-SET-001](optimizations/OPT-SET-001-density-adaptive-compact-sets.md) | Density-adaptive compact sets | **Implemented external reference** | Choose sparse/dense representation locally while retaining exact set algebra | +| [OPT-CONT-001](optimizations/OPT-CONT-001-partitioned-coordination-domains.md) | Partitioned coordination domains | **Implemented external pattern** | Split one global contention hotspot into independent domains while preserving global invariants | +| [OPT-FAN-001](optimizations/OPT-FAN-001-shared-materialization-fanout.md) | Shared materialization for fan-out/replay | **Implemented external reference** | Transform/encode once and reuse the representation for many consumers | +| [OPT-SEARCH-001](optimizations/OPT-SEARCH-001-budget-aware-adaptive-search.md) | Budget-aware adaptive parameter search | **Proposed / OPT synthesis** | Spend expensive evaluations where they are most informative | +| [OPT-APPROX-001](optimizations/OPT-APPROX-001-contract-bounded-approximation.md) | Contract-bounded approximation | **Proposed / OPT synthesis** | Trade exactness only inside an explicit measurable error/degradation envelope | +| [OPT-REDUCE-001](optimizations/OPT-REDUCE-001-early-working-set-reduction.md) | Early working-set reduction | **Implemented external pattern** | Filter/cull/limit before expensive composition | +| [OPT-CRIT-001](optimizations/OPT-CRIT-001-critical-path-prioritization.md) | Critical-path prioritization | **Proposed / OPT synthesis** | Do critical work now, speculate carefully, defer non-critical work | +| [OPT-BUDGET-001](optimizations/OPT-BUDGET-001-performance-regression-budgets.md) | Performance regression budgets | **Proposed / OPT synthesis** | Turn performance expectations into environment-scoped regression contracts | +| [OPT-PRUNE-001](optimizations/OPT-PRUNE-001-bound-driven-search-space-pruning.md) | Bound-driven search-space pruning | **Proposed / OPT synthesis** | Prove whole search regions cannot improve the incumbent and skip them | +| [OPT-SIMD-001](optimizations/OPT-SIMD-001-evidence-gated-native-autovectorization.md) | Evidence-gated native autovectorization | **Verified, environment-specific** | Reshape a hot batch for vector codegen, prove parity, inspect instructions, then require measured native benefit | +| [OPT-SOA-001](optimizations/OPT-SOA-001-worker-local-soa-tiling.md) | Worker-local SoA tiling | **Implemented external reference** | Keep only hot fields in bounded per-worker SoA tiles and reuse cache-local scratch | +| [OPT-POOL-001](optimizations/OPT-POOL-001-persistent-topology-aware-worker-pools.md) | Persistent topology-aware worker pools | **Implemented external reference** | Reuse workers/buffers across dispatches and choose physical/logical topology explicitly | +| [OPT-AUTO-001](optimizations/OPT-AUTO-001-calibrated-host-aware-path-promotion.md) | Calibrated host-aware path promotion | **Implemented external reference** | Calibrate equivalent paths on the live host/workload, include lifecycle costs, and promote only with margin + oracle parity | See [CATALOG.md](CATALOG.md) for the decision map and [README4AI.md](README4AI.md) for machine-oriented usage. -## Provenance correction: the QEC links +## Formalization boundary -The originally supplied QEC v170.2.0/v170.2.1 links are useful, but they are **not the origin of the large pytest/CI test-speed optimization**. The primary deterministic test optimization lineage is: +The immutable `v1.0.0` release and its five original records are formalized by the pinned Lean v1 model described in [`FORMALIZATION.md`](FORMALIZATION.md). This catalog expansion is **post-v1**. It does not edit the three pinned v1 Lean model files or pretend the new records are already theorem-backed. -- [QEC v68.2.0 — Deterministic Execution Engine](https://github.com/QSOLKCB/QEC/releases/tag/v68.2.0) -- [QEC v68.4.0 — Deterministic Runtime Optimization](https://github.com/QSOLKCB/QEC/releases/tag/v68.4.0) -- [QEC v68.4.1 — Invariant Hardening & Repository Cleanup](https://github.com/QSOLKCB/QEC/releases/tag/v68.4.1) +The new [`OPTIMIZATION-PROBLEM.md`](OPTIMIZATION-PROBLEM.md) supplies a canonical problem contract for future records: -The v170.2.x releases remain relevant here because v170.2.1 contains independently validated multicore benchmark evidence used by **OPT-PAR-001**. +`P = (X, F, f, d, C, B, S)` -## Existing source material +where `d` is the objective direction/order; the remaining components are search space, feasible set, objective, correctness/semantic constraints, evaluation budget and stopping rule. -- [`power_module.md`](power_module.md) contains the E8/qutrit DSP architecture that motivated **OPT-DSP-001**. -- [`suxen.zip`](suxen.zip) is retained as an opaque source archive. It is **not yet promoted as optimization evidence**. See [`sources/SUXEN.md`](sources/SUXEN.md) and run the bounded recursive [`scripts/inventory_zip.py`](scripts/inventory_zip.py) scanner in a normal checkout before promoting claims from its nested payloads. +## Source material + +- [`sources/WONDERBUILD.md`](sources/WONDERBUILD.md) — incremental execution, scheduling and rebuild-benchmark donor; GPL implementation boundary recorded. +- [`sources/JAZCO.md`](sources/JAZCO.md) — production systems case studies for coalescing, compact sets, contention, fan-out, approximation and reduction. +- [`sources/OPTIMIZATION-LIBRARIES.md`](sources/OPTIMIZATION-LIBRARIES.md) — BayesianOptimization, Hyperopt and NLopt mechanism/taxonomy notes. +- [`sources/WPO.md`](sources/WPO.md) — critical-path and performance-budget discovery source. +- [`sources/MATHEMATICAL-OPTIMIZATION.md`](sources/MATHEMATICAL-OPTIMIZATION.md) — mathematical/combinatorial problem vocabulary and pruning foundations. +- [`sources/GALAXY-CPU.md`](sources/GALAXY-CPU.md) — merged GALAXY CPU optimization phases covering SIMD/autovectorization, worker-local SoA tiling, persistent topology-aware pools and calibrated host-aware path promotion. +- [`power_module.md`](power_module.md) — E8/qutrit DSP architecture that motivated **OPT-DSP-001**. +- [`sources/SUXEN.md`](sources/SUXEN.md) — provenance and the required bounded recursive inventory procedure for the opaque `suxen.zip` source candidate. +- [`scripts/inventory_zip.py`](scripts/inventory_zip.py) — bounded recursive ZIP inventory entry point; use the explicit limits documented in `sources/SUXEN.md` rather than generic/unbounded extraction. +- [`suxen.zip`](suxen.zip) — opaque source archive, still **not promoted as optimization evidence** until the bounded inventory identifies reusable mechanisms. + +## Integrity gate + +`scripts/check_catalog.py` verifies heading/filename identity, post-v1 status vocabulary, complete contracts, complete README coverage, CATALOG coverage, and optimization-record link labels/targets. CI runs it via `.github/workflows/catalog-integrity.yml`. ## Add the next optimization -Copy [`templates/OPTIMIZATION-RECORD.md`](templates/OPTIMIZATION-RECORD.md), assign the next ID, record before/after evidence, and state exactly what correctness property was preserved. +Copy [`templates/OPTIMIZATION-RECORD.md`](templates/OPTIMIZATION-RECORD.md), define the optimization problem contract, assign the next ID, record evidence honestly, and state exactly what correctness property is preserved. diff --git a/README4AI.md b/README4AI.md index e01461b..ab26bf6 100644 --- a/README4AI.md +++ b/README4AI.md @@ -4,53 +4,110 @@ This repository is a reusable optimization knowledge base for QSOL projects. ## Start here -1. Identify the dominant bottleneck. -2. Select the closest optimization record from `CATALOG.md`. -3. Read the source record completely before modifying another repository. -4. Preserve the target repository's semantics, invariants, determinism, evidence boundaries, and public API unless the task explicitly changes them. -5. Benchmark before and after in the target environment. -6. Record any adaptation rather than pretending source-project constants are universal. +1. Read `OPTIMIZATION-PROBLEM.md`. +2. Define `P = (X,F,f,d,C,B,S)` for the target: search space, feasible set, objective, direction, correctness/semantic contract, budget and stopping rule. +3. Identify the dominant bottleneck and choose the closest record from `CATALOG.md`. +4. Read that record and its source note completely before modifying another repository. +5. Preserve target semantics, invariants, determinism, evidence boundaries and public API unless the task explicitly changes them. +6. Benchmark before/after in the target environment and retain raw/repeated observations where practical. +7. Record adaptations rather than pretending source constants are universal. + +## Problem classification + +Before choosing an optimizer, classify: + +- continuous / integer / categorical / conditional / mixed variables; +- local / global search; +- deterministic / noisy / stochastic objective; +- gradient available / derivative-free / black-box; +- cheap / expensive evaluations; +- bound/equality/inequality/semantic/resource constraints; +- sequential / synchronous batch / asynchronous execution; +- exact / approximation explicitly permitted. ## Decision map -- **Slow pytest / deterministic numerical tests** → `OPT-PY-001` -- **Repeated work known to be mathematically or bitwise equivalent** → `OPT-INV-001` -- **Lean/mathlib dependency rebuild dominates CI** → `OPT-LEAN-001` -- **Independent work can run concurrently** → `OPT-PAR-001` -- **High-rate numerical/audio loop with slower control state** → `OPT-DSP-001` - -Patterns may be composed. Example: a numerical CI job can use `OPT-PY-001` inside the test process and `OPT-PAR-001` at a higher independent-work layer, but only after nested parallelism and memory pressure are measured. +- slow deterministic tests → `OPT-PY-001` +- proven-equivalent repeated computation → `OPT-INV-001` +- Lean dependency reconstruction → `OPT-LEAN-001` +- independent parallel work → `OPT-PAR-001` +- hot numerical/audio loop with slower control state → `OPT-DSP-001` +- unchanged-input pipeline reruns → `OPT-INC-001` +- identical simultaneous in-flight work → `OPT-COAL-001` +- sparse/dense integer-set mixture → `OPT-SET-001` +- hot global coordination point → `OPT-CONT-001` +- repeated per-consumer transformation → `OPT-FAN-001` +- expensive parameter tuning → `OPT-SEARCH-001` +- approximation explicitly allowed → `OPT-APPROX-001` +- large population filtered only after expensive work → `OPT-REDUCE-001` +- latency-critical path competes with optional work → `OPT-CRIT-001` +- gradual performance drift/regression → `OPT-BUDGET-001` +- combinatorial search with valid optimistic bounds → `OPT-PRUNE-001` +- hot deterministic batch underuses vector ISA → `OPT-SIMD-001` +- large AoS traversal needs cache-local bounded hot-field chunks → `OPT-SOA-001` +- repeated parallel runs pay thread/buffer startup or need explicit physical/logical worker policy → `OPT-POOL-001` +- several exact execution paths trade places across hosts/workloads → `OPT-AUTO-001` + +## Important distinctions + +- **cache/reuse**: completed result already exists; +- **coalescing**: result does not exist yet, but equivalent callers share one in-flight evaluation; +- **async search diversification**: independent workers should intentionally avoid evaluating the same pending region; +- **parallelism**: improves throughput only when resource contention and information dependencies allow it; +- **SIMD/autovectorization**: changes instruction-level execution of equivalent batch work; code-generation evidence is not itself an end-to-end speedup; +- **SoA tiling**: changes temporary data layout/working-set shape while preserving the logical source/output contract; +- **persistent pools**: change worker lifetime and lifecycle amortization, not the kernel's semantics; +- **host-auto promotion**: selects among already-correct paths using live calibration and a fail-closed oracle; it does not make one path universally best; +- **approximation**: a contract choice, never a hidden optimization. ## Status vocabulary -- **Verified**: source project contains passing validation and measured/observed evidence for the optimization. +Post-v1 records must use one of these exact status categories. A semicolon may follow the category with a short evidence-boundary caveat; `scripts/check_catalog.py` validates the category before that semicolon. + +- **Verified**: source project contains passing validation and measured/observed evidence. - **Verified, environment-specific**: measured result is real but not a universal performance guarantee. -- **Implemented reference**: the optimization mechanism exists in code, but no general speedup claim is made. -- **Proposed**: architecture/design idea only. Do not report it as achieved performance. +- **Implemented reference**: the mechanism exists in repository code, but no general speedup claim is made. +- **Implemented external reference**: the mechanism exists in an external donor; target transfer still requires local validation. +- **Implemented external pattern**: an external donor demonstrates the pattern, but this OPT record does not claim a target implementation. +- **Proposed / OPT synthesis**: architecture/design guidance only. Do not report it as achieved performance. - **Source candidate**: material exists but has not been inspected sufficiently to promote claims. -## Non-negotiable safety rules +The frozen v1 records retain their historical release wording and are exempt from post-v1 status normalization. + +## Non-negotiable rules - Never remove tests merely to make CI faster. -- Never weaken an assertion, tolerance, theorem target, receipt, claim boundary, or validation rule without an explicit contract change. +- Never weaken an assertion, tolerance, theorem target, receipt, trust boundary or validation rule without an explicit contract change. - Never treat a cache hit as proof of a cold rebuild. -- Never equate a requested worker count with observed effective parallel execution. -- Never claim SIMD/zero-copy/lock-free speedups from `power_module.md` without controlled measurements. -- Never claim anything from `suxen.zip` until it has been inventoried and the relevant source has been read. +- Never equate requested workers with observed effective execution. +- Never copy historical worker counts, thresholds, search budgets, bit partitions, cache sizes, tile sizes, calibration repeats, promotion margins or approximation limits without target measurement. +- Never prune a search region unless the bound used for pruning is sound for the declared problem. +- Never call an approximate result exact. +- Never infer end-to-end speedup from vector instructions or an isolated kernel probe alone. +- Never publish a native/ISA-specialized path as universal if deployment compatibility is not guaranteed. +- Never treat process-wide RSS gathered across calibration as isolated selected-engine memory evidence. +- Never let an auto selector hide a parity failure by silently falling back; fail closed and preserve explicit canonical/manual control. +- Never optimize from stale workload assumptions when fresh measurements are available. +- `suxen.zip` remains unpromoted until inventoried and inspected. ## What to copy vs what to adapt -Copy the **structure** of the optimization: cache identity binding, deterministic reuse, equivalence gates, bounded workers, control-rate separation, sparse evaluation, vectorized batches. +Copy the **structure**: equivalence gates, complete signature identity, coalescing ownership, partitioned coordination, density-adaptive representation, shared materialization, adaptive trial ledgers, explicit approximation envelopes, early reduction, critical-path classification, performance budgets, sound bounds, SIMD parity/codegen gates, bounded SoA working sets, persistent-worker lifecycle accounting, and calibrated promotion with independent oracle verification. -Adapt the **numbers**: trial counts, iteration caps, cache sizes, worker caps, sample/control rates, sparse cardinalities, timing thresholds, tolerances and environment-specific hashes. +Adapt the **numbers and policies**: trial counts, worker caps, hashes, cache sizes, shard counts, bit splits, batch widths, tile sizes, domain-contraction rates, acquisition parameters, tolerances, error limits, benchmark thresholds, calibration sizes/repeats, promotion margins, topology policy and stopping budgets. -## Evidence expected in a new OPT record +## Evidence expected in a new record At minimum record: -- source repository and immutable-enough source identity (release/commit/PR head); +- source identity and licensing/provenance boundary where relevant; +- the `OPTIMIZATION-PROBLEM.md` contract; - baseline and optimized behavior; - correctness/conformance gate; - benchmark environment or an explicit statement that no benchmark exists; - failure/rollback condition; -- whether the optimization changes latency, throughput, memory, CI time, or only architecture. +- whether the change affects latency, throughput, memory, I/O, CI time, quality or only architecture. + +## Formalization boundary + +`v1.0.0` contains five immutable Lean-formalized records. Post-v1 catalog records are not theorem-backed merely because they live in the same repository. See `FORMALIZATION.md`. diff --git a/ROADMAP.md b/ROADMAP.md new file mode 100644 index 0000000..b53907d --- /dev/null +++ b/ROADMAP.md @@ -0,0 +1,93 @@ +# OPT Roadmap + +This roadmap tracks candidate optimization families that are worth promoting into the catalog after their source identity, reusable contract, evidence boundary and target-specific validation are written down. A roadmap entry is **not** a verified optimization record and must not be cited as if it were already part of the catalog. + +## Near-term donor mining: VORTEX-N v5.0.0 + +The VORTEX-N v5.0.0 Zenodo bundle contains three high-signal mechanisms that appear distinct from the existing GALAXY-derived records. The goal is to preserve the reusable mechanism while refusing to promote donor-specific dimensions, buffer layouts, workgroup sizes, matching counts, receipt sizes or benchmark numbers as universal settings. + +Before any of these become catalog records, add a pinned VORTEX-N source note with the relevant release/DOI/commit identity, licensing boundary, exact donor files/sections, and a clear statement of which observations are implementation evidence versus target-independent claims. + +### Candidate: `OPT-CONFLICT-001` — Conflict-free dependency partitioning + +**Problem:** logically independent work is forced through atomics, locks, collision arbitration or serialized mutation because some operations can touch the same state. + +**Reusable mechanism:** build or derive a conflict/dependency graph over each operation's complete read set, write set and externally observable effects; partition operations into independent sets such as matchings/color classes; and execute one phase in parallel only when it contains no write/write, write/read, read/write or observable-effect hazards. + +**Why it looks promising:** this can replace runtime contention with an explicit scheduling transform. The pattern is applicable to graph processing, mesh/constraint updates, particle or interaction systems, sparse mutation, schedulers and other workloads where write conflicts are structurally knowable. + +**Evidence gate before promotion:** + +- prove the partition removes all declared write/write, write/read, read/write and observable-effect hazards for the operation model; +- compare scalar/serial and partitioned-parallel execution against the same canonical/reference semantics, including ordering where observable; +- measure scheduling/partition overhead as well as lock/atomic reduction; +- test skewed or adversarial graphs where the number of phases grows; +- do not assume the donor's number of matchings or graph topology transfers to another target. + +### Candidate: `OPT-RESIDENT-001` — Accelerator-resident double-buffered execution + +**Problem:** iterative accelerator workloads repeatedly pay host/device transfer, allocation, synchronization or read/write hazard costs even though most state is reused from one iteration to the next. + +**Reusable mechanism:** keep the hot iterative working set resident on the accelerator and separate current/next state with explicit ping-pong or equivalent double buffering. Reuse resident allocations across rounds and cross the host/device boundary only when the public contract requires it. + +**Why it looks promising:** this attacks transfer and lifecycle overhead while also making read-state/write-next-state ownership explicit. It composes naturally with local/shared-memory staging inside a workgroup and with compact readback after the resident computation finishes. + +**Evidence gate before promotion:** + +- compare lifecycle-adjusted runtime with a transfer-heavy or reallocation baseline; +- account for retained accelerator memory as a real resource cost; +- verify exact or declared-tolerance parity across many iterations; +- validate buffer-generation ownership so stale or partially written state cannot leak across rounds; +- test small workloads where residency overhead or retained memory is not justified; +- treat workgroup size, buffer shape and donor memory layout as target-specific. + +### Candidate: `OPT-READBACK-001` — Device-side aggregation and compact readback + +**Problem:** a caller transfers a large accelerator-resident state back to the host merely to compute a much smaller verification, summary or decision payload. + +**Reusable mechanism:** aggregate evidence where the data already lives, then transfer only the smallest result needed by the external contract—for example counts, histograms, extrema, mismatch summaries, digests/receipts or other bounded evidence rather than the full state. + +**Why it looks promising:** this turns readback volume into an explicit optimization target and can remove a bandwidth/synchronization boundary without changing the underlying computation. + +**Evidence gate before promotion:** + +- define exactly which host-visible questions the compact evidence must answer; +- prove the aggregate is sufficient for that contract rather than merely convenient; +- compare full-state readback with device-side reduction plus compact transfer; +- include reduction-kernel cost, synchronization and transfer latency in the objective; +- retain full-output comparison paths where exact element-level validation is required; +- do not treat a donor-specific receipt size or reduction schema as portable. + +## Supporting mechanisms to keep as composition/validation notes for now + +### Topology-aligned cooperative workgroups + +Map one self-contained logical unit to one cooperative workgroup and stage its hot local state in workgroup/shared memory when that reduces global-memory traffic and synchronization. This is likely to compose with `OPT-RESIDENT-001`, but it should not become a separate record until there is evidence that the mapping itself—not merely residency or data layout—delivers a distinct reusable win. + +### Deterministic procedural control regeneration + +Regenerate deterministic control/schedule state from compact seeds or round/cell identities instead of storing and transferring a large precomputed schedule when arithmetic is cheaper than memory traffic. Keep this as a candidate composition technique until a controlled comparison isolates its memory/bandwidth benefit and proves replay equivalence. + +### Exact inverse/replay verification + +Where an optimized transformation is reversible, use forward-then-inverse replay as supplementary correctness evidence and require restoration of the original state under the declared exactness contract. Also require a forward-result oracle—direct parity with a trusted reference output or independent semantic invariants—so mutually consistent forward/inverse defects cannot pass merely because they round-trip. This is valuable validation guidance, but it is not automatically a performance optimization and should not be promoted as one without an independent objective win. + +## Promotion rule + +A roadmap candidate becomes a catalog record only when it has: + +1. a concrete source identity and provenance note; +2. a complete `P = (X, F, f, d, C, B, S)` contract; +3. all required problem-classification fields; +4. a preserved-contract statement and explicit failure/rollback boundaries; +5. before/after evidence or an explicit no-benchmark statement; +6. validation that tests the mechanism's own failure modes rather than only the happy path; and +7. no donor constant promoted as a universal default without target evidence. + +The intended progression for the VORTEX-N-derived work is: + +**reduce conflicts → localize/cooperate → keep state resident → reduce before transfer** + +This complements the GALAXY-derived progression already in the catalog: + +**data layout → ISA exploitation → persistent execution → empirical runtime selection** diff --git a/optimizations/OPT-APPROX-001-contract-bounded-approximation.md b/optimizations/OPT-APPROX-001-contract-bounded-approximation.md new file mode 100644 index 0000000..ad77266 --- /dev/null +++ b/optimizations/OPT-APPROX-001-contract-bounded-approximation.md @@ -0,0 +1,80 @@ +# OPT-APPROX-001 — Contract-bounded approximation + +**Status:** Proposed / OPT synthesis; external production pattern, target-specific proof/measurement required +**Domains:** visualization, search, streaming, telemetry, simulation, audition DSP + +## Source evidence + +- https://jazco.dev/2025/02/19/imperfection/ +- existing `OPT-DSP-001` approximation boundary +- `sources/JAZCO.md` + +## Problem + +Exact processing has unbounded or unacceptable cost even though the product/scientific contract permits a bounded loss of precision, completeness or freshness. + +## Optimization problem contract + +- X: target-supported approximation policies, quality/resource ceilings, sampling/culling/LOD policies, update frequencies, state-reset rules, evaluation horizons, exact-mode fallback choices, hard-envelope proof/enforcement methods, and stochastic tuning/certification procedures +- F: policies whose declared error/degradation metric remains within the target's explicit envelope over the declared state/composition horizon and whose resource/semantic constraints are satisfied; **a hard pointwise/worst-case envelope must be established by an analytic/formal bound, exhaustive checking over a finite declared domain, or runtime enforcement that proves the bound for each produced approximate result and falls back to the exact path whenever the proof/guard cannot establish it**; sample-based certification alone is admissible only for explicitly statistical contracts; when a stochastic policy is selected from multiple candidates, feasibility certification is based on independent held-out conformance data or a predeclared selection-aware simultaneous-confidence/multiple-testing procedure rather than naive reuse of the tuning samples +- f: target-measured resource or latency cost, optionally paired with the declared quality/error metric +- d: minimize resource/latency cost subject to feasibility in F, or use the target's predeclared multi-objective ordering when quality is ranked rather than hard-bounded +- C: approximation is permitted only by an explicit contract; exact callers are not silently weakened; the error norm, aggregation rule, sequence/composition horizon, reset boundaries, whether the envelope is hard worst-case or statistical/confidence/tail-based, **the proof/enforcement basis for any hard envelope**, and the stochastic tuning-versus-certification procedure are declared before evaluation; an exact reference path or exact fixture remains available where practical, and a runtime-enforced hard envelope must route any unproved/unsafe case to that exact path before an out-of-envelope approximate result becomes observable +- B: target-specific benchmark/quality-evaluation budget over predeclared ordinary, boundary, adversarial, repeated-application, and long-horizon fixtures; proof construction/exhaustive finite-domain checking/runtime-guard validation for hard envelopes is budgeted separately from statistical sampling where applicable; for stochastic policy search, tuning/selection evaluations and independent certification evaluations (or the budget used by the predeclared simultaneous-confidence procedure) are accounted separately +- S: stop when the evaluation/proof budget is exhausted or a selected policy meets the target resource objective and passes the declared conformance certification/proof while remaining inside the quality envelope over the entire declared horizon +- Variables: continuous / integer / categorical / conditional / mixed, depending on approximation policy +- Search scope: local or global, explicitly declared for the target +- Objective behavior: deterministic, noisy, or stochastic depending on the quality/resource metric +- Information: derivative-free / black-box by default +- Evaluation cost: moderate to expensive when exact references, held-out certification, hard-envelope proof, exhaustive checking, runtime guards, or long-horizon trajectories are required +- Constraints: explicit error envelope, semantic/API, resource, horizon/reset, hard-envelope proof/enforcement, selection-aware statistical certification, and exact-fallback constraints +- Parallelism: sequential, synchronous batch, or asynchronous according to target evaluation; stateful validation must preserve trajectory semantics +- Exactness: approximation explicitly permitted only inside the declared measurable/provable envelope + +## Preserved contract + +Approximation is admissible only when the contract explicitly permits it. A previously exact API cannot be silently weakened and still be called correctness-preserving. For stateful or repeatedly composed approximations, the contract applies over an explicitly declared horizon—not merely to each isolated step—so bounded per-step error is insufficient if drift can accumulate beyond the allowed envelope. The contract must also state whether compliance is pointwise/worst-case or statistical; a stochastic envelope is judged by its declared aggregation, confidence, exceedance-probability, quantile, or tail criterion rather than by silently substituting a hard per-sample limit. **A hard pointwise/worst-case guarantee is a universal claim over its declared domain: ordinary, boundary, adversarial, or randomized fixtures are evidence but cannot by themselves prove that universal claim over a non-finite domain.** Such a hard envelope requires a sound analytic/formal bound, exhaustive coverage of a finite declared domain, or runtime enforcement that proves/guards each approximate output and invokes the exact reference path whenever the guard cannot certify the bound. When multiple stochastic policies are tuned or screened, choosing the apparent winner changes the sampling distribution: the data used to optimize/select a policy cannot be treated as independent nominal-confidence certification evidence unless the declared procedure explicitly accounts for that selection. + +## Optimization + +Introduce a resource ceiling and degrade only along a declared dimension: sample/cull, lower level of detail, approximate search, bounded stale data, or reduced update frequency. Make the error surface measurable and reversible. + +For stateful streaming, simulation, DSP, iterative numerical work, or any repeatedly applied approximation, define the error model before benchmarking: the norm/metric (for example absolute, relative, L2, perceptual, state-distance, or domain-specific), how error composes or is aggregated through time, the maximum sequence length or physical/time horizon over which the envelope must hold, and any reset/checkpoint/re-synchronization boundaries that legitimately restart the horizon. If the system can run longer than the validated horizon without reset, either extend validation to that operational horizon or define a separate long-run drift bound; do not infer long-run safety from one-step ε alone. + +For a **hard pointwise or worst-case envelope**, choose the proof/enforcement mode before deployment. If the input/state domain is finite and tractable, exhaustively evaluate every declared case (including every relevant trajectory/state when the contract is stateful). Otherwise establish a sound analytic or formal bound that covers the complete declared domain/horizon, **or** install a runtime guard that derives a sound per-instance error bound/invariant before publishing the approximate result and falls back to the exact path whenever that guard cannot prove compliance. A test suite—no matter how adversarial or large—does not convert a non-finite-domain hard guarantee into a proof. Sample-based testing may still regression-test the proof/guard implementation, but it is not the certification basis for the universal claim. + +For stochastic/noisy approximations, also define the statistical compliance rule before evaluation: the sampling unit and workload distribution, aggregation statistic, confidence level or interval procedure, tolerated exceedance probability, quantile/tail bound, and the sample/evaluation budget used to decide compliance. Do not reinterpret a statistical guarantee as a pointwise worst-case guarantee, and do not weaken a declared hard worst-case envelope into an average-case claim after observing data. + +If more than one stochastic approximation policy is tuned, compared, adaptively searched, thresholded, or screened using sampled error data, **separate selection from certification**. The default pattern is to use one predeclared tuning/selection set (or stream) to choose the candidate and then evaluate that frozen candidate on an independent held-out conformance set drawn from the declared operational distribution. If independent holdout is impractical, use a predeclared selection-aware method that preserves the advertised guarantee across the entire candidate-selection procedure—for example simultaneous confidence bounds, family-wise/multiple-testing correction, valid selective-inference/e-process machinery, or another target-justified method. A nominal per-policy confidence interval computed on the same samples used to select the best-looking policy is not certification. Record exactly which evaluations influenced policy selection and which evaluations supported the final compliance claim. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: No exact target baseline has been established here. +- Optimized: No target approximation implementation has been benchmarked here. +- Speedup / memory reduction: No transferable claim; external production observations remain source evidence only. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Measure error/degradation and resource savings together across ordinary, boundary and adversarial workloads. Keep an exact reference for differential evaluation where practical. Declare and test the error norm/metric, aggregation rule, sequence/composition horizon, reset boundaries, hard-versus-statistical envelope semantics, **hard-envelope proof/enforcement mode**, and tuning-versus-certification procedure explicitly. + +For stateful/repeated use, run differential trajectories against the exact path across short, nominal, maximum-supported, and adversarially long sequences. Include biased-error fixtures where each individual step remains within the local ε but errors accumulate in the same direction; verify the cumulative/state error still respects the declared horizon envelope. Test reset/checkpoint boundaries before, at, and after the limit; verify resets actually restore the assumptions used by the next horizon. + +For a **hard pointwise/worst-case envelope**, validation must verify the claimed proof mechanism rather than merely accumulate examples. For a finite declared domain, exhaustively enumerate every input/state/trajectory covered by C and compare with the exact path. For a non-finite domain, review/check the analytic or formal bound against the implementation and all assumptions it depends on, or validate a runtime guard whose sound per-instance bound/invariant is checked before output publication. Inject cases where the runtime guard cannot establish safety and prove the approximate result is suppressed and the exact fallback is used. Random, adversarial, property-based, and boundary samples remain valuable regression tests, but passing them is not sufficient evidence for a universal hard guarantee. + +Where stochastic approximation is used, evaluate the declared expected, quantile, exceedance-probability, confidence, or tail criterion against the exact path as specified by C. Include fixtures where individual samples exceed a nominal pointwise value while the declared statistical envelope remains satisfied, and fixtures where the configured tail/confidence/exceedance criterion truly fails. Verify rollback decisions distinguish those cases rather than triggering on one sample unless the contract explicitly declares a hard single-sample/worst-case bound. **Sample-based certification is reserved for these explicitly statistical contracts; it must not be cited as proof of a non-finite-domain hard worst-case envelope.** + +Add **selection-bias fixtures** whenever multiple stochastic policies are considered. Generate several candidate policies whose apparent sampled errors vary by chance, select the best-looking candidate using the declared tuning procedure, and prove that the final compliance decision uses either fresh held-out samples unavailable to selection or the declared simultaneous/selection-aware inference procedure. Verify the tuning samples alone cannot certify the selected winner at nominal per-policy confidence. Include repeated/adaptive candidate selection, early stopping, and candidate-count changes; confirm the advertised confidence/tail/exceedance guarantee remains valid under the complete selection procedure. Persist an audit trail labeling each evaluation as tuning/selection, certification, or both only when the declared selection-aware method formally permits dual use. + +## Target-repo adaptation + +Define `ε`, the exact quality/error norm, aggregation rule, workload distribution, maximum state/composition horizon, reset/checkpoint semantics, long-run drift policy, escape hatch and exact-mode availability locally. Explicitly classify the quality envelope as **hard pointwise/worst-case** or **statistical**. For a hard envelope, declare one sound certification mechanism: analytic/formal proof over the complete domain/horizon, exhaustive checking over an explicitly finite domain, or runtime per-instance enforcement with exact fallback whenever the guard cannot prove the bound. Do not use sampled conformance data as the sole certification basis for a universal hard claim. For statistical contracts specify the confidence/tail/exceedance rule and decision sample budget. If multiple policies are searched or compared, predeclare the tuning/selection dataset or stream, the independent certification dataset/budget, **or** the exact simultaneous-confidence/multiple-testing/selective-inference method that makes data reuse valid; record which observations affected selection versus certification. If the target has no finite operational horizon, establish a justified asymptotic/stability bound or periodic re-synchronization rule instead of copying a finite benchmark horizon from another system. + +## Failure modes + +Unmeasured quality loss, biased sampling, hidden rare-case failures, **claiming a universal hard worst-case envelope from finite sampled/adversarial fixtures over a non-finite domain**, unsound analytic/formal assumptions, incomplete finite-domain enumeration, a runtime guard that can publish before proving the per-instance bound or fails to fall back exactly, cumulative drift that is invisible to one-step checks, reset boundaries that fail to restore reference assumptions, state-dependent amplification, unstable feedback loops, selecting the best-looking stochastic policy and then certifying it on the same data with naive per-policy confidence, undisclosed adaptive candidate search/early stopping that invalidates nominal error guarantees, misclassifying a statistical envelope as a hard pointwise bound (or vice versa), and callers incorrectly assuming exact semantics. + +## Rollback trigger + +Evaluate rollback against the **declared envelope semantics and certification procedure**. For a hard pointwise/worst-case contract, disable/fall back immediately on any proven bound violation, **any failure of the analytic/formal/exhaustive proof obligations, or any runtime case in which the guard cannot prove the bound before publication**; a runtime-enforced design must use the exact path for such unproved cases rather than emit the approximation. A sampled counterexample to a hard envelope is an immediate violation, but absence of sampled counterexamples is never sufficient certification of a non-finite-domain hard guarantee. For a stochastic/statistical contract, disable when the predeclared aggregation, confidence, exceedance-probability, quantile, or tail criterion fails under its stated **selection-aware certification** procedure; an isolated sample beyond a nominal pointwise value is not by itself a contract violation unless the contract says it is. Treat a selected policy as uncertified—and disable or fall back—if held-out certification fails, if tuning and certification evidence are mixed contrary to the declared procedure, or if the simultaneous/multiple-testing/selective-inference assumptions required for data reuse are violated. In all cases, disable when cumulative/state drift violates its declared bound, reset/checkpoint validation fails, reference comparisons violate C, a catastrophic semantic/safety constraint is breached, or resource savings are not material. diff --git a/optimizations/OPT-AUTO-001-calibrated-host-aware-path-promotion.md b/optimizations/OPT-AUTO-001-calibrated-host-aware-path-promotion.md new file mode 100644 index 0000000..2443425 --- /dev/null +++ b/optimizations/OPT-AUTO-001-calibrated-host-aware-path-promotion.md @@ -0,0 +1,104 @@ +# OPT-AUTO-001 — Calibrated host-aware path promotion + +**Status:** Implemented external reference; GALAXY merges a fail-closed calibrated selector, while calibration constants and projection accuracy remain host/workload-specific. +**Domains:** multi-path CPU runtimes, heterogeneous execution strategies, production tuning, adaptive dispatch + +## Source evidence + +- Repository: `QSOLKCB/GALAXY` +- PR: https://github.com/QSOLKCB/GALAXY/pull/14 +- Merge commit: `b2e860309a04d7591c86f71d2b4ab1e5eec4c4d7` +- Source note: `sources/GALAXY-CPU.md` +- Licensing boundary: Apache-2.0 donor; record promotes policy structure, not target-specific constants. + +## Problem + +Several semantically equivalent execution paths exist, but the fastest path depends on CPU topology, workload shape, tile size, startup/teardown cost and repetition count. A universal hardware/model table becomes stale, while always choosing the apparently fastest microbenchmark can regress real workloads or violate correctness. + +## Optimization problem contract + +- X: A bounded candidate set of semantically equivalent execution paths plus target-specific calibration shape, scoring policy and promotion margin. +- F: Candidates that pass an independent correctness oracle, preserve workload-relevant calibration dimensions, expose truthful lifecycle/tuning/verification costs, keep canonical/manual control available, withhold externally visible effects until parity is established, and fail closed on mismatch. +- f: Total user-paid automatic-selection cost: calibration/tuning, candidate lifecycle terms, selected full-work execution, mandatory independent full-work oracle/verification, and any staging/commit overhead. If verification is explicitly out-of-band rather than paid per invocation, state and score that different scope explicitly. +- d: Minimize expected total user-paid full-work cost under one symmetric accounting boundary, but keep the canonical path on near ties or insufficient evidence. +- C: Exact selected/oracle parity before selected-path effects become externally visible, no silent fallback after a correctness mismatch, explicit requested/effective topology policy, accurate evidence scope and no universal claim from one host calibration. +- B: A bounded end-to-end selection budget covering the calibration candidate matrix/repetitions, selected full-work execution, mandatory full-work oracle/verification, and required staging/commit overhead without exceeding the declared workload/resource envelope. +- S: Select an optimized candidate only when its expected total paid invocation cost, including mandatory verification/staging where applicable, beats canonical by a predeclared material margin and the staged full-work result passes oracle verification; otherwise select canonical. +- Variables: categorical, integer and conditional +- Search scope: global +- Objective behavior: noisy +- Information: black-box +- Evaluation cost: expensive +- Constraints: semantic and resource +- Parallelism: sequential +- Exactness: exact + +## Preserved contract + +Automatic selection may change which implementation executes, but not the externally declared result. Every candidate admitted to calibration and the final selected full workload must match an implementation-independent or sufficiently independent oracle under the target's exactness contract. + +A selected execution must either be side-effect-free until verification or stage all externally observable outputs and state changes transactionally. The staged result may become visible only after full-work oracle parity succeeds; on mismatch the staged result is discarded. A system that cannot defer irreversible effects must validate before those effects or is not feasible for this pattern. + +Manual/canonical execution surfaces remain available for audit and recovery. A selector must not hide parity failures by silently switching paths after a mismatch; correctness failure is evidence that the candidate or calibration is invalid. + +## Optimization + +1. Define a bounded set of already-validated candidate implementations. +2. Build calibration work that preserves the workload dimensions that materially affect ranking, rather than using an arbitrary tiny microbenchmark. +3. Measure each candidate under the same calibration boundary and record requested versus effective topology/configuration. +4. Project or extrapolate only under an explicit documented model. Account for startup, teardown, allocation/first-touch, tuning, and the mandatory full-work oracle/verification and staging costs whenever the user pays them for the requested invocation horizon. Compare candidates and canonical under the same cost boundary; do not promote using an asymmetric score that omits work required only by the optimized path. +5. Keep the canonical path when candidates are within a predeclared margin so noise and model error do not trigger unstable path switching; apply that margin to the total expected paid invocation cost rather than only the selected kernel. +6. Execute the selected full workload into side-effect-free or transactional staging, run or validate the independent full-work oracle within the declared budget, and publish selected-path effects only after exact parity succeeds. Discard staged results and fail closed on any mismatch. +7. Emit a receipt explaining candidate scores, lifecycle/verification accounting, selection reason, oracle result, staging/publish result and evidence scope. + +The reusable mechanism is calibrated promotion with a symmetric cost boundary, uncertainty margin and correctness oracle, not a static table mapping CPU names to implementations. + +## Before / after evidence + +- Environment: GALAXY PR #14 adds dedicated host-auto coverage on Linux x86-64, Linux ARM64, macOS ARM64 and Windows x86-64. +- Workload/fixture: canonical, spawned SoA and persistent SoA candidate families calibrated against requested resident/frame shape. +- Cold baseline: canonical BAM-LUT execution retained as an explicit candidate and fallback/manual path. +- Warm/no-op baseline where relevant: persistent candidate scores include amortized startup and teardown over requested repetitions rather than comparing only steady-state dispatch. +- Small invalidation / partial-work case where relevant: calibration is bounded and may use less resident work while preserving requested frame depth/effective tile shape. +- Large invalidation / full-work case where relevant: selected path is verified on the full requested workload against the streaming canonical oracle; the donor records full-oracle timing separately, and a target must include that verification cost in the user-paid objective/budget whenever it is mandatory per invocation. +- Optimized: host/workload-specific candidate chosen only after calibration and margin gating. +- Speedup / memory / I/O / quality change: donor uses a 5% projected promotion margin; that number is source-specific and is not promoted as a universal OPT default. +- Variance / repetitions / raw samples: donor uses three calibration repeats and versioned receipts; targets must choose their own statistically defensible budget and margin. + +## Validation + +- Verify every calibration candidate against an independent oracle before it can compete. +- Preserve workload dimensions known to affect ranking, including depth, effective tile shape and topology where relevant. +- Test projection/scoring identities and lifecycle accounting, including the selected full run, mandatory full-work oracle/verification, staging and publish costs whenever those are paid by the invocation. +- Compare the promoted path's total paid cost against canonical under the same scope; test a case where a fast selected kernel is correctly rejected because verification overhead erases the advantage. +- Include topology-detection and tuning time in the declared selection overhead when users pay that cost. +- Verify the selected full workload again against the oracle before making staged effects visible. +- Test side-effect-free and transactional staging, intentional parity failures, and prove that mismatching staged output/state is discarded rather than exposed. +- Test near ties, canonical wins and optimized wins. +- Scope RSS/memory evidence correctly; a process-wide high-water mark covering calibration plus selection cannot be presented as isolated selected-engine memory. +- Keep receipts versioned so future policy changes are distinguishable from earlier selection behavior. + +## Target-repo adaptation + +Re-profile candidate families, calibration size, repetitions, projection model, promotion margin, lifecycle amortization, full-work oracle cost, staging/publish strategy, workload-shape dimensions, topology policy and recalibration cadence. Do not copy GALAXY's 65,536-particle base, tile list, three repeats or 5% margin without target evidence. Prefer direct full-work measurement when calibration cost is affordable or projection error is material. If the target has irreversible externally visible effects, establish a pre-effect oracle/validation boundary before adopting automatic promotion. + +## Failure modes + +- Calibration shape does not preserve the feature that determines real-work ranking. +- Runtime scaling is nonlinear, making the projection misleading. +- Selection or mandatory verification overhead exceeds the saved runtime on short-lived workloads. +- The scoring boundary omits oracle/staging work paid only by promoted candidates and creates a false win. +- Selected-path output or state changes escape before oracle parity is known. +- Workload phases change after calibration and invalidate the choice. +- Near-tie noise causes path thrashing because the margin is too small. +- The oracle shares the same defect or optimized primitive as the candidate and is not genuinely independent. +- Process-wide memory evidence is mislabelled as per-candidate memory. +- Static host/model assumptions replace live evidence and age badly. + +## Rollback trigger + +Immediately reject the selected path on full-work oracle mismatch and discard all uncommitted staged results. Treat any already-exposed mismatching result as a contract failure. Revert automatic promotion to canonical/manual mode when the complete selection+execution+verification budget is exceeded, lifecycle-adjusted total paid benefit disappears, calibration becomes unstable or unrepresentative, workload drift changes rankings, or selector/verification overhead materially outweighs expected savings. Recalibrate rather than preserving a stale winner. + +## Composition notes + +This record selects among mechanisms such as `OPT-SOA-001`, `OPT-POOL-001`, `OPT-SIMD-001` and canonical paths after those mechanisms have their own correctness gates. It complements `OPT-BUDGET-001`: regression budgets can detect when a formerly promoted path stops meeting its measured advantage. \ No newline at end of file diff --git a/optimizations/OPT-BUDGET-001-performance-regression-budgets.md b/optimizations/OPT-BUDGET-001-performance-regression-budgets.md new file mode 100644 index 0000000..3cd935e --- /dev/null +++ b/optimizations/OPT-BUDGET-001-performance-regression-budgets.md @@ -0,0 +1,76 @@ +# OPT-BUDGET-001 — Performance regression budgets + +**Status:** Proposed / OPT synthesis; target calibration required +**Domains:** CI, web, numerical kernels, builds, services, DSP + +## Source evidence + +- `davidsonfellipe/awesome-wpo` inspected at `84f32948a6298456d6a94cff64551f39f2666e6f` +- performance-budget tooling and measurement resources catalogued upstream +- `sources/WPO.md` + +## Problem + +Small performance regressions accumulate because performance is measured occasionally but not guarded as an engineering contract. + +## Optimization problem contract + +- X: target-supported metric/fixture/statistic/threshold configurations for a performance-regression gate +- F: gate configurations based on a sufficiently characterized environment and workload, with statistically justified tolerance, selection-aware validation, a declared good-run/regression population or explicitly enumerated control scope, and no weakening of functional correctness or workload realism +- f: target-measured **loss vector** comprising missed-material-regression loss (for example false-negative rate and, where relevant, severity-weighted miss cost), flaky/false-failure loss (false-positive rate), and measurement/CI overhead +- d: minimize every component of the declared loss vector under the target's predeclared scalar, weighted, Pareto, or lexicographic ordering; if detection quality is reported separately, it is a diagnostic complement such as `1 - false-negative-rate`, not an oppositely oriented coordinate inside `f` +- C: the performance gate must not incentivize weakening tests, assertions, evidence, semantic coverage, or representative workload inputs; once a candidate gate is selected, its claimed false-positive/false-negative performance must be established on independent control executions or under a predeclared selection-aware procedure that accounts for every configuration tried, **and every reported detector error rate must name the population/generator and sampling scheme it estimates or be explicitly scoped to the enumerated controls only** +- B: target-specific calibration and certification budget specifying repetitions, environment samples, held-out/control executions, population/generator coverage where general error rates are claimed, and allowable CI/runtime measurement cost +- S: stop calibration when the declared sample budget is exhausted or the baseline/noise estimate is stable enough to freeze one candidate gate for independent certification; promote it only if the certification contract passes +- Variables: continuous / integer / categorical / mixed metric, statistic, fixture, and threshold choices +- Search scope: local gate/calibration tuning +- Objective behavior: noisy / stochastic measurement distributions +- Information: derivative-free statistical observations +- Evaluation cost: moderate to expensive depending on repetitions and fixture scale +- Constraints: functional correctness, representative workload, declared population/control scope, statistical tolerance, runner/environment characterization, false-positive/false-negative, selection bias, and CI-overhead constraints +- Parallelism: sequential or synchronous-batch calibration; parallel sampling only when runner interference is characterized +- Exactness: no semantic approximation; statistical tolerance/noise handling is explicit + +## Preserved contract + +A performance gate may not incentivize weakening functional tests, correctness, evidence or workload realism. The gate is valid only while its fixture, environment characterization, declared detection population/control scope, detection sensitivity and measurement overhead remain inside their declared contract. Calibration evidence used to choose among competing gates is not automatically valid certification evidence for the selected gate. A rate measured on a handpicked or finite control suite must not be generalized to unseen production regressions unless the target has declared and sampled from a population/generator that supports that inference. + +## Optimization + +Turn a stable, reproducible performance expectation into a regression gate. Compare distributions or robust summaries where noise matters; separate machine/environment drift from code regression; keep cold/warm claims distinct. Keep known-fast and known-regressed control fixtures (or equivalent calibration cases) so the gate can periodically prove it still distinguishes acceptable from materially regressed behavior. + +Before claiming false-positive or false-negative rates beyond those exact controls, define the estimand. Declare the good-run population and the material-regression population or generator, including the target workloads, regression classes, severity range, environment distribution, and any exclusions. Predeclare how certification cases are sampled or generated from that population and how repeated executions are grouped. If the repository cannot justify a broader population model, use the controls strictly as an enumerated conformance suite and report control-suite detection/failure rates without implying a general production error rate. + +When multiple metric/fixture/statistic/threshold configurations are explored, treat that search as model selection. Use calibration/tuning data to choose the candidate, then freeze its complete configuration before certification. The default certification path is an independent held-out sample of known-good and known-regressed executions drawn under the declared sampling scheme and playing no role in choosing the gate. If holding out controls is impractical, use a predeclared nested-resampling, simultaneous-confidence, multiple-testing, or other selection-aware procedure whose error guarantees cover the full configuration search—not nominal per-candidate estimates computed after selecting the best one. + +## Before / after evidence + +- Environment: No controlled target-repository budget calibration has been run for this OPT record. +- Baseline: No target baseline distribution is claimed here. +- Optimized: Not applicable until a target repository adopts and calibrates a performance budget. +- Speedup / memory reduction: This pattern protects performance; it does not itself claim a speedup. +- Variance / repetitions: Must be established in the target environment before a hard threshold is promoted. + +## Validation + +Calibrate variance before setting the threshold. During tuning, compare candidate metric/fixture/statistic/threshold configurations using explicitly designated calibration data and preserve raw samples where practical. Once one gate is selected, **freeze the entire gate configuration before measuring its claimed detection performance**. + +Before certification, write down the exact error-rate scope: either (a) a declared good-run/regression population or generator plus sampling scheme, including regression classes/severities and environment/workload strata that the rate is intended to represent, or (b) a finite enumerated control suite to which the reported rates are explicitly limited. For population claims, draw the independent certification sample according to that scheme and record coverage/counts by the predeclared strata; do not substitute a convenient handpicked set after seeing gate behavior. For control-only claims, label the result as control-suite performance and prohibit extrapolation to unrepresented regression classes. + +Certify the frozen gate on independent known-good and known-regressed executions that were not used to select it. Measure false positives, false negatives, and gate overhead against predeclared acceptance limits **within the declared scope**. If independent controls are unavailable, use a predeclared nested-resampling or selection-aware procedure that accounts for every candidate/configuration examined, and report the resulting adjusted uncertainty/error rates rather than reusing naive in-sample estimates. If the declared population/generator changes, previous rates do not automatically transfer. + +Verify objective orientation explicitly: construct one candidate with fewer missed regressions but more false alarms and another with the opposite tradeoff, compute the declared loss coordinates, and prove the configured scalar/Pareto/lexicographic ordering ranks them exactly as documented. A separately reported positive detection-quality score must never be fed into a minimization coordinate without an explicit monotone conversion to loss. + +Record which executions were used for calibration/selection versus certification, along with the population/control scope and sampling provenance for each certification case. Re-run independent controls after runner/toolchain changes and periodically enough to detect stale fixtures or sensitivity drift. Add an explicit overfitting fixture where several candidate gates are tuned on one noisy control sample set; prove the gate cannot be promoted merely because one candidate looked best on those same samples. Also include at least one deliberately omitted regression class in a test report to prove the tooling labels that class as outside the estimated scope rather than silently counting the observed controls as universal evidence. + +## Target-repo adaptation + +Never copy another project's milliseconds, bundle sizes or thresholds. Establish the target's own baseline and noise envelope, define control fixtures, and declare acceptable false-positive/false-negative rates plus a maximum measurement-overhead budget. **Define what population those rates refer to:** specify representative workloads, regression classes and severities, environment strata, exclusions, and the sampling/generation process; or explicitly limit the claim to a named finite control suite. Define `f` using consistently oriented loss coordinates and predeclare how those coordinates are ordered or scalarized; if the target also reports a positive detection-quality score, document its conversion to the minimized loss coordinate. Predeclare how calibration/selection is separated from certification: held-out controls by default, or a justified nested/selection-aware alternative. Preserve the candidate-search history and sampling provenance needed to audit the claimed certification error rates. + +## Failure modes + +Flaky gates from uncontrolled runners, benchmark gaming, stale fixtures, hardware drift, thresholds so loose they miss real regressions, thresholds so tight they block good changes, an objective vector mixing maximized quality with minimized costs without an explicit per-coordinate direction/conversion, selection bias from evaluating a chosen gate on the same controls used to tune it, **sampling bias or undefined detector populations that turn control-suite performance into an unjustified general false-positive/false-negative claim**, unrepresented regression classes/severities, unreported configuration search that invalidates nominal error rates, and measurement overhead large enough to damage CI usability or distort the workload under test. + +## Rollback trigger + +Disable or demote the gate to non-blocking and recalibrate whenever its measurement environment is invalid, its fixture is stale/nonrepresentative, its declared regression/good-run population or control scope no longer matches the deployment claim, its sampling process no longer represents the declared population, its objective orientation/scalarization is ambiguous or ranks a worse detector as better, independent/selection-aware certification no longer meets the declared false-positive/false-negative limits **within that stated scope**, known regressions are no longer detected, known-good controls fail above the declared false-positive limit, observed false negatives exceed the declared limit, or measurement overhead exceeds the predeclared budget. Do **not** disable merely because product code legitimately regressed; in that case keep the valid gate and fix or explicitly accept the regression through the target's normal review process. \ No newline at end of file diff --git a/optimizations/OPT-COAL-001-concurrent-duplicate-work-coalescing.md b/optimizations/OPT-COAL-001-concurrent-duplicate-work-coalescing.md new file mode 100644 index 0000000..17b1bef --- /dev/null +++ b/optimizations/OPT-COAL-001-concurrent-duplicate-work-coalescing.md @@ -0,0 +1,115 @@ +# OPT-COAL-001 — Concurrent duplicate-work coalescing + +**Status:** Implemented external reference; target validation required +**Domains:** services, CI, artifact generation, metadata, parsing, model/data loading + +## Source evidence + +- Jazco, "Request Coalescing": https://jazco.dev/2023/09/28/request-coalescing/ +- `sources/JAZCO.md` + +## Problem + +Many callers request the same expensive computation concurrently before any caller has populated a reusable result, producing a thundering herd. + +## Optimization problem contract + +- X: target-supported request-key canonicalizations, authorization/equivalence scopes, shared-operation lifetime and launch-state policies, per-generation waiter limits, **global/per-tenant in-flight generation and waiter budgets**, bounded-overflow/backpressure policies, **baseline-availability/successful-admission/useful-throughput floors**, per-waiter cancellation/deadline/**atomic terminal-outcome publication** policies, result-preparation/clone-failure policies, retry/error-sharing policies, and result-ownership policies +- F: policies that coalesce only requests equivalent in both computation semantics and authorization/visibility scope, preserve authorization, timeout, cancellation, result, ownership, preparation-failure, launch-cancellation, and error semantics for every joined caller, linearize cancellation/deadline against **irrevocable terminal-outcome publication** for each waiter, linearize unstarted-to-running launch against closing/last-waiter cancellation, **bound total registry/generation/waiter memory across all keys and retained closing generations**, never admit new waiters to a closing or terminal generation, and **preserve the target's predeclared normal-load availability/successful-admission and useful-throughput floor rather than satisfying the resource bound by rejecting ordinary service demand** +- f: measured **loss vector** comprising duplicate upstream evaluations, end-to-end/tail latency, overload/rejection above the target's allowed baseline envelope, useful-throughput loss, and coalescer overhead (including synchronization, global/per-tenant admission accounting, generation/waiter memory, launch-state synchronization, result preparation/cloning, atomic outcome publication, backpressure, and result-copy cost) +- d: minimize every declared loss coordinate under the target's predeclared scalar, weighted, Pareto, or lexicographic ordering while F remains satisfied +- C: every joined caller receives exactly one terminal outcome valid for its original request semantics, authorization scope, ownership contract, cancellation state, deadline, launch state, and result-preparation outcome; **a waiter may become terminal-success/error only when the corresponding outcome is already irrevocably stored/published for that waiter**; non-equivalent or authorization-distinct requests are never merged; one caller leaving cannot incorrectly cancel work still required by another caller; no upstream operation may start after its generation has already become closing due to loss of all live waiters; closing/terminal generations are not joinable; **global/per-tenant generation and waiter admission limits are never exceeded**, including under high-cardinality keys and retained closing generations; overload has an explicit bounded result; and **normal-load successful-admission/availability and useful throughput remain at or above the target's declared baseline floor** +- B: target-specific concurrent-load test budget plus explicit global/per-tenant coalescer admission limits (maximum in-flight generations, waiter records, and any bounded overflow queue); no portable request count, duration, memory cap, or availability floor is supplied here +- S: stop when the declared load-test budget is exhausted or further policy changes fail to produce a validated material improvement without violating C or the declared availability/throughput floor +- Variables: categorical / integer / mixed +- Search scope: local policy tuning within one coalescing boundary +- Objective behavior: noisy under concurrent load; semantic equivalence and admission-floor checks remain deterministic for a fixed trace +- Information: derivative-free / black-box performance measurements +- Evaluation cost: moderate to expensive concurrent-load testing +- Constraints: semantic equivalence, authorization, ownership, launch-state linearizability, result-preparation failure, **global/per-tenant registry and waiter memory**, **baseline availability/successful admission and useful throughput**, cancellation, deadline, atomic outcome publication, timeout, and resource constraints +- Parallelism: asynchronous / concurrent +- Exactness: exact request/result semantics; no approximation is introduced + +## Preserved contract + +Coalescing may merge only requests that are equivalent for the same **joinable generation** of the shared operation, including any tenant/principal/visibility context that affects whether the computation or its result may be shared. Each caller retains independent authorization, cancellation, timeout/deadline, result-ownership, preparation-failure, and error semantics. A caller abandoning its wait must not by itself terminate a shared operation that still has live waiters. Once a generation enters cancellation, closure, success, or failure handling, it becomes non-joinable before later callers can attach. **Memory/admission bounds apply across the whole coalescer, not only inside one generation:** high-cardinality keys, fresh generations, and retained closing generations must all consume explicit global/per-tenant generation and waiter capacity until their state is actually retired. No request may bypass those caps merely because it is the first waiter for a new key. + +Bounded admission is not permission to stop serving the workload. For the target's declared normal-load/reference demand envelope, the optimized coalescer must preserve at least the predeclared successful-admission/availability and useful-throughput floor of the reference path (or another explicitly approved service-level floor). Overload/backpressure may be part of the contract outside that envelope, but a zero-cap or tiny-cap policy that merely rejects ordinary requests is infeasible even if it minimizes duplicate work and observed latency among the few requests that remain. + +Each waiter has exactly one atomic completion cell/terminal record, initially `pending`. Terminal completion is not a two-step “claim then notify” protocol: the winning transition must atomically compare `pending` and **publish/store the complete terminal outcome**—for example `success(immutable-or-private-result-handle)`, `error(error-record)`, `cancelled`, or `timed-out`—before that terminal state becomes visible. A promise/future `set_result`/`set_exception`, transactional queue/envelope insert, or equivalent single linearization point is acceptable. A subsequent wakeup/signal/callback may tell the caller to inspect the terminal record, but that notification is advisory; failure/interruption of the notifier cannot erase an already published outcome or leave a waiter terminal with nothing retrievable. Cleanup may not retire the waiter/outcome until the target's delivery/acknowledgement/retention contract makes that outcome safely consumable or no longer required. + +The shared operation also has a linearized launch lifecycle: a generation closed before launch may never subsequently start ownerless upstream work. + +## Optimization + +Create an in-flight registry entry for a canonical equivalence key. The key must include every request attribute required to establish safe sharing, including authorization-relevant tenant/principal/visibility scope unless the target instead proves that the upstream result is globally shareable and independently authorizes each delivered result. + +Before tuning admission caps or overflow policy, freeze the target's **normal-load service envelope and admission/throughput floor** from the reference path or an explicitly approved service-level objective. Cap reductions are admissible only while that floor remains satisfied. Rejections/backpressure inside the declared normal-load envelope count as objective loss and, once the hard floor is crossed, make the candidate infeasible rather than artificially “fast.” + +Before creating a **new generation**, atomically reserve both (a) one generation slot from the applicable global/per-tenant in-flight-generation budget and (b) one waiter slot for the initiating caller. The reservation and registry insertion must be one linearizable admission decision. If either capacity is exhausted, do not allocate a partial generation or untracked waiter; return/block/queue according to the declared bounded overload policy. A distinct equivalence key does not get a free first waiter merely because no matching generation exists yet. + +Atomically create the joinable generation **with the initiating caller already registered as its first waiter** and with an explicit launch state such as `unstarted`. Do not invoke, schedule, or otherwise permit upstream work yet. This prevents an immediately/synchronously completing operation from reaching terminal state with an empty waiter set while also giving early cancellation a state it can close before any work exists. + +Linearize upstream launch against the generation's live-waiter and closing state. Under the same registry lock/CAS/transactional boundary used for generation state, permit `unstarted -> running` only while the generation remains joinable and has at least one live `pending` waiter. Install or bind a **sticky upstream cancellation token/handle** as part of that transition, before releasing the serialization boundary. If the last waiter cancels/times out while the generation is still `unstarted`, transition it to `closing/non-joinable` and make any later launch attempt fail; no upstream work is started. If `unstarted -> running` wins first but actual invocation/scheduling occurs immediately afterward, any last-waiter cancellation that races in that interval must set the already-bound sticky cancellation token. The launcher must check/attach that token before or atomically with invocation so a cancellation that has already won cannot be lost merely because the external operation object did not yet exist. There must be no path where the generation is closed with zero live waiters and a creator later launches uncancelled work from stale local state. + +Equivalent later callers may register as independent waiters only while the generation is joinable and **both** the per-generation waiter capacity and applicable global/per-tenant waiter capacity remain. Waiter admission is atomic with both counters. When any required capacity is exhausted, apply one explicit target policy rather than silently exceeding the bound: reject/return a documented overload or retryable-backpressure result, block/queue the caller behind a separately bounded admission mechanism, or use another bounded policy with explicit timeout/cancellation semantics. Starting an unconstrained parallel generation for the same equivalence key is not the default overflow behavior because it recreates the duplicate upstream load this pattern is intended to prevent. If a target deliberately permits overflow generations, those generations still consume the global/per-tenant generation budget, and their concurrency/duplicate-work tradeoff must be part of C/B and validated separately. Any such overload behavior must still satisfy the declared normal-load availability/throughput floor. + +**Closing and terminal-but-not-yet-retired generations continue to consume their generation slot and any still-live waiter/cleanup capacity until deterministic retirement actually releases those resources.** This prevents a churn attack from repeatedly canceling callers, leaving expensive upstream operations closing, and creating unlimited fresh generations that evade the advertised memory bound. Resource release must be atomic with retirement so admission cannot observe capacity before the corresponding registry state is gone. + +Cancellation and timeout publish terminal outcomes through the same waiter completion cell. Cancellation attempts atomically compare-and-publish `pending -> cancelled`; timeout/deadline handling atomically compare-and-publishes `pending -> timed-out`. A completion path may publish success/error only if the waiter's declared deadline has not already expired at that same linearization point. If the deadline is expired, completion must instead publish/leave the target's timed-out outcome and must not publish the shared value/error. For explicit cancellation racing completion, whichever atomic compare-and-publish wins `pending` first defines the public outcome; losing transitions are no-ops. Because the winning terminal state already contains the outcome itself, scheduler or notifier timing after that transition cannot change request semantics. + +Cancellation and timeout are otherwise per waiter: when one waiter leaves through a winning cancellation/timeout publication, remove only that waiter from the live-waiter accounting. If live waiters remain, keep the shared generation joinable. If the **last** live waiter leaves, atomically make the generation closing/non-joinable. If it is still `unstarted`, this closure permanently prevents the launch transition. If it is already `running`, set/trigger the sticky upstream cancellation token according to the declared policy. A new caller arriving after the closing transition must create a fresh generation **only if global/per-tenant admission capacity permits it** rather than attach to work being prevented/canceled. The closing generation may remain internally tracked until its launch-prevention or terminal cleanup is complete, and it continues consuming its generation budget during that interval. + +On upstream success or failure, atomically transition the generation to **terminal/non-joinable** (or remove it from the joinable map) **before** snapshotting the candidate waiter set. New callers arriving after that terminal transition must create a fresh generation only through normal bounded admission and cannot attach to the completed one. Snapshotting only identifies candidate waiter records; each waiter still resolves through its own atomic outcome publication. + +For an upstream **failure**, prepare the immutable/shared error record or per-caller wrapped error representation as required by the target API while the waiter remains `pending`, then atomically compare-and-publish that complete error outcome into the waiter completion cell. If cancellation/timeout already won, discard any per-caller wrapper and publish nothing else. There is no separate “delivered-error claim” followed by a fallible notification step. + +For an upstream **success**, define result ownership explicitly. If the terminal value is immutable/share-safe under the target API, atomically compare-and-publish `success(shared-immutable-handle)` into each eligible pending waiter. If callers normally receive mutable or caller-owned results, **prepare the independent defensive clone/copy/copy-on-write handle while the waiter is still `pending`**. Preparation is not entitlement to delivery: cancellation or timeout may win while preparation is in progress, in which case discard/release the prepared value. If preparation succeeds, atomically compare-and-publish `success(prepared-private-handle)`; if that publication loses to cancellation/timeout, discard the prepared handle. If preparation fails because of allocation, serialization, quota, or another declared preparation error, prepare the corresponding error representation and atomically compare-and-publish `error(preparation-error)` only if the waiter is still pending. Thus no waiter can become terminal-success before a deliverable value exists, and no terminal success/error can exist without its outcome already stored. + +After outcome publication, wake/signal/callback delivery may be retried independently. A lost wakeup, interrupted notifier task, callback exception after publication, or scheduler failure must not make the terminal outcome inaccessible: the waiter/future/queue record remains the source of truth. If the target API cannot provide such an irrevocable completion cell or transactional delivery record, it must keep the waiter nonterminal until delivery itself is irrevocably committed; it may not mark success/error first and hope a later notification succeeds. + +Do not silently retry for only some joined callers; if shared retry is supported, its attempt limit, backoff, budget charging, authorization scope, result-preparation behavior, and terminal error semantics must be part of the declared policy. Otherwise, a retry starts a new generation after the failed generation is retired and must pass normal global/per-tenant admission. Retire/clean up the generation deterministically only after every candidate waiter has either published one terminal outcome or been safely removed under the target delivery/retention policy, then atomically release the associated admission capacity. + +This differs from caching: the reusable result does not exist yet. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: Not established in a target repository. +- Optimized: Not established in a target repository. +- Speedup / memory reduction: No transferable claim; the Jazco implementation is source evidence for the mechanism. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Stress simultaneous identical and non-identical keys; inject upstream failures/timeouts; cancel the first caller while other waiters remain; cancel all waiters and verify the declared upstream-cancellation policy; race a new caller against the last-waiter cancellation transition and prove it never joins the closing generation; race a new caller against success/failure completion and prove the terminal generation is made non-joinable before waiter snapshot/outcome publication; test waiter-specific deadlines; verify shared failure publication and retry accounting; prove only one upstream evaluation occurs per joinable generation while all surviving callers terminate correctly. + +Add an **availability/admission-floor fixture** before accepting any tuned cap. Replay the declared normal-load/reference trace against the non-coalesced or previously accepted baseline and record successful-admission/availability plus useful completed throughput. Then evaluate candidate caps, including deliberately degenerate zero/tiny generation or waiter limits that return overload quickly. Prove such candidates are rejected as infeasible whenever they fall below the declared service floor, even if they report excellent latency for admitted requests or near-zero duplicate work. Separately measure overload/rejection above the normal-load envelope as an explicit loss coordinate rather than silently excluding rejected requests from latency statistics. + +Add a **cross-key/global-admission saturation fixture**. Generate many distinct equivalence keys so every request attempts to create its own generation, and separately churn through keys whose prior generations remain `closing` because upstream cancellation/cleanup is delayed. Fill the global and per-tenant generation/waiter budgets to their limits, race additional first callers and later waiters, and prove admission is linearizable: counts never exceed the configured bounds, retained closing generations continue to occupy capacity until retirement, no first waiter bypasses the global cap, and every non-admitted caller receives exactly the declared overload/backpressure behavior. Repeat with multiple tenants to verify one tenant cannot consume capacity reserved for another when per-tenant isolation is part of C. + +Add an explicit **registration-to-launch cancellation race**. Pause after the generation and initiating waiter have been registered but before `unstarted -> running`. Cancel or time out that initiating waiter as the last live waiter, then release the launcher. Prove the generation becomes closing/non-joinable and upstream work is never started. In the opposite interleaving, let `unstarted -> running` win but pause before the external invocation exists; then cancel the last waiter and prove the sticky token is already set/observable so the subsequent invocation is suppressed or immediately canceled according to policy. Repeat under high contention and prove no zero-waiter generation can leak a running/hung upstream operation. + +Add **terminal-publication races** for upstream success, upstream failure, cancellation, and deadline expiry. Pause immediately before each waiter compare-and-publish and prove exactly one `pending -> terminal(outcome)` transition wins. Then pause **after** success/error publication but before any wakeup/callback; kill/fail/interupt the notifier and prove the waiter can still retrieve exactly the already-published result/error from its completion cell, cancellation/timeout cannot replace it, and cleanup does not retire it prematurely. Repeat for immutable success, mutable private clones, shared upstream errors, preparation errors, lost wakeups, callback exceptions, and high concurrency. Verify no terminal waiter ever lacks a retrievable outcome and no waiter observes two outcomes. + +Add **clone/preparation-failure fixtures** for mutable results. Force allocation, serialization, copy-on-write setup, or quota failure while preparing a per-waiter value. Verify a waiter is still `pending` until preparation succeeds or a preparation-error outcome is atomically published; successful preparation followed by a winning cancellation/timeout causes the prepared value to be discarded; failed preparation can resolve to exactly one published error only if cancellation/timeout has not already won; and no clone failure can leave a waiter in terminal success without an actual value. Race clone success/failure against cancellation and deadline expiry repeatedly under load. + +Add an **immediate synchronous-completion** fixture where the upstream operation can finish inline once launch actually begins. Prove the initiating caller was already registered before launch and always obtains the terminal result/error from its completion cell unless its own cancellation/deadline publication wins under the same rules. + +Add authorization-boundary fixtures: issue syntactically identical requests under different tenants, principals, roles, ACL/visibility scopes, or other authorization context. Prove they either map to different equivalence keys **or** that the shared upstream result is explicitly safe to reuse and each caller is independently authorized before outcome publication. Verify that a result produced under one authorization scope can never leak to another merely because the resource parameters match. + +Add ownership-isolation fixtures for mutable results: publish one coalesced computation to multiple callers, mutate one caller's returned object, and prove every other caller's result remains unchanged. If the API declares the shared value immutable, attempt prohibited mutation through all exposed aliases and verify the immutability/share-safety contract. + +Add per-generation waiter-overflow races: fill one generation's waiter list to one slot below the maximum, launch multiple equivalent callers concurrently for the final slot, and prove admission is linearizable, local and global capacity are never exceeded, non-admitted callers receive exactly the documented backpressure/overflow behavior, and cancellation/timeouts of queued or rejected callers remain correct. + +## Target-repo adaptation + +Define key canonicalization, the authorization/visibility context that participates in equivalence, the **normal-load/reference demand envelope and minimum successful-admission/availability plus useful-throughput floor**, **global and per-tenant maximum in-flight generation counts, total waiter-record limits, whether closing/terminal cleanup consumes those limits, any separately bounded overflow queue**, per-generation waiter count, bounded overload/backpressure semantics, result ownership/share-safety policy, how mutable per-waiter results are prepared and how preparation failures surface, the waiter-owned atomic completion cell/transactional delivery representation, cancellation/deadline winning semantics, outcome-retention/acknowledgement and notifier retry semantics, the generation launch states and serialization primitive for `unstarted -> running` versus `closing`, the sticky cancellation-token/handle semantics used before an external operation object exists, the exact condition for canceling upstream work, the atomic create-with-first-waiter rule, the atomic closing/terminal non-joinable transitions, cleanup/retirement and capacity release, and whether failures are shared as terminal or retried under one explicit shared retry policy. + +## Failure modes + +Over-broad keys merge non-equivalent or authorization-distinct work; **admission caps can game the objective by rejecting normal demand unless baseline availability/successful admission and useful throughput are constrained**; per-generation-only waiter caps can still permit unbounded total memory under high-cardinality keys or churned closing generations; releasing admission capacity before a closing generation is truly retired can let registry state exceed the advertised bound; launching upstream work before registering the initiating waiter can strand that caller on synchronous completion; registering first but launching from stale creator state after the last waiter already closed an unstarted generation can leak ownerless work; cancellation issued before an external operation exists can be lost without a sticky token or atomic launch state; non-linearized cancellation/deadline versus terminal publication can produce late or double outcomes; **marking a waiter terminal before atomically storing/enqueuing its outcome can strand it if notification fails**; cleanup that retires a published outcome before it is safely consumable can lose delivery; claiming success before a mutable per-waiter value is successfully prepared can strand a waiter with no deliverable result; clone/preparation failure can race cancellation and create inconsistent outcomes if not published atomically; coupling shared lifetime to the first caller can terminate valid waiters; leaving a canceled or terminal generation joinable can attach new callers to doomed/completed work; omitting authorization scope can leak results across principals/tenants; sharing a mutable result object can create cross-caller aliasing; undefined overflow semantics can exceed memory bounds, drop callers, or recreate duplicate upstream load; never canceling after all waiters leave can leak work; a hung upstream operation can stall many callers; ambiguous retry/error policy can cause correlated or duplicated work. + +## Rollback trigger + +Disable if coalescing changes any caller's authorization/cancellation/deadline/result/ownership/preparation-error semantics; if normal-load successful-admission/availability or useful throughput falls below the declared baseline/service floor; if objective reporting excludes overload/rejection in a way that rewards denying service; if **global/per-tenant in-flight generation, waiter-record, or overflow-queue limits can be exceeded across many keys or retained closing generations**; if admission capacity can be reused before the corresponding generation/waiter state is actually retired; if an upstream operation can start after its generation has become closing with no live waiters; if a pre-launch cancellation can be lost because no cancellation token/operation object existed yet; if any success/error terminal state can become visible before its complete outcome is irrevocably stored/enqueued; if notifier/wakeup/callback failure after terminal publication can make the outcome inaccessible; if cleanup can retire an undelivered/unacknowledged outcome contrary to the target retention contract; if a waiter can enter terminal success before an isolated deliverable result exists; if clone/preparation failure can produce no terminal outcome or a second terminal outcome; if a cancelled/timed-out waiter can later be replaced by success/error; if one waiter can observe two terminal outcomes; if an expired deadline can lose merely because timeout processing was delayed; if authorization-distinct requests are merged without independent delivery authorization; if one caller can cancel work required by another; if the initiating caller is stranded on immediate completion; if a new caller joins a closing/terminal generation; if mutable-result aliasing is possible; if shared operations leak; or if tail latency/failure amplification becomes unacceptable. diff --git a/optimizations/OPT-CONT-001-partitioned-coordination-domains.md b/optimizations/OPT-CONT-001-partitioned-coordination-domains.md new file mode 100644 index 0000000..d21b71e --- /dev/null +++ b/optimizations/OPT-CONT-001-partitioned-coordination-domains.md @@ -0,0 +1,86 @@ +# OPT-CONT-001 — Partitioned coordination domains + +**Status:** Implemented external pattern; target validation required +**Domains:** ID allocation, runtimes, queues, counters, ingestion, schedulers + +## Source evidence + +- https://jazco.dev/2025/09/26/interning/ +- https://jazco.dev/2024/01/10/golang-and-epoll/ +- `sources/JAZCO.md` + +## Problem + +Independent workers serialize on one globally coordinated resource even though the underlying work could proceed independently. + +## Optimization problem contract + +- X: target-supported shard/domain counts, namespace splits, worker-to-domain mappings, merge/aggregation policies, ownership-lease policies, fencing-epoch schemes, allocator-incarnation fencing, restart-safe allocator-state policies, finite namespace capacity, and explicit exhaustion handling +- F: configurations that preserve the target's required uniqueness, exclusive ownership, visibility, failure-domain, and ordering guarantees through assignment, rebalance, same-owner restart, crash recovery, process replacement, pause/resume, split-brain recovery, and namespace exhaustion; any finite ID layout must prove that allocation stops before wrap, widens safely, or transitions to a collision-free epoch/namespace that cannot collide with any still-valid historical identity +- f: measured coordination contention, tail latency, and coordination overhead under the declared workload +- d: minimize under the target's predeclared objective ordering +- C: partitioning must not silently weaken any global invariant; any intentional shift from global to per-domain ordering is a separately declared contract change; every mutable ownership/allocator incarnation must be fenced at the authoritative mutation boundary so a superseded process cannot continue acting merely because its emitted IDs remain unique; allocator restart must not reuse IDs/ranges already issued before the crash; finite identity spaces must never wrap or silently recycle externally valid values when local/domain/epoch capacity is exhausted +- B: target-specific contention/scale/failover/capacity benchmark budget declared before tuning; no portable shard count, bit split, or exhaustion threshold is supplied here +- S: stop when the budget is exhausted or a validated partitioning materially reduces the target bottleneck without violating C; if capacity analysis shows the chosen identity layout cannot survive the required horizon without unsafe exhaustion, reject that candidate before deployment +- Variables: integer / categorical / mixed +- Search scope: local architecture/partition-policy tuning +- Objective behavior: noisy under concurrent load; ownership/uniqueness invariants are deterministic +- Information: derivative-free / black-box performance measurements +- Evaluation cost: moderate to expensive at target scale and during failover testing +- Constraints: uniqueness, ownership, ordering, visibility, failure-domain, lease/fencing, allocator-incarnation fencing, durable allocator-state, finite namespace capacity, exhaustion handling, and resource constraints +- Parallelism: concurrent / asynchronous by construction +- Exactness: exact ownership/uniqueness semantics; no approximation is introduced + +## Preserved contract + +Partitioning must not silently weaken uniqueness, ownership, visibility or ordering guarantees. If ordering becomes per-domain rather than global, that is a contract change and must be explicit. When a domain can be reassigned or an owner process can be replaced, only the currently authoritative fenced owner/incarnation may mutate that domain; a delayed, partitioned, resumed, or split-brain previous process must be rejected even if it still believes its old lease is valid. A same-owner process restart is also part of the ownership contract: restarting an allocator must not reset process-local state in a way that can reissue an ID/range already made externally visible, and a replacement process must not coexist as an unfenced second owner with a paused predecessor. + +For fixed-width IDs or otherwise finite namespaces, exhaustion is also part of the preserved contract. Reaching the end of a local counter, range, domain field, incarnation field, or epoch space must never cause silent wraparound or reuse of an identity that could still be externally valid. The target must define in advance whether exhaustion stops/rejects new allocation, migrates to a wider representation, or performs a collision-free epoch/namespace transition whose coexistence rules prove that old and new identities cannot collide. + +## Optimization + +Factor a global coordination space into independent domains. Encode domain identity into keys/IDs or route work so each domain can advance mostly independently. Prefer a small explicit merge/aggregation boundary to a permanently hot global lock/counter/poller. + +For dynamic assignment/rebalance, use an **exclusive handoff with fencing**. A durable coordinator grants ownership together with a monotonically increasing epoch/token. Every state-changing operation that depends on domain ownership carries that fencing identity, and the authoritative storage/queue/allocation boundary rejects operations from older identities. A lease alone is insufficient if an old process can resume after expiry; the fencing token must make stale writes/actions impossible at the mutation boundary. Do not activate the replacement owner until its new fencing identity is durably authoritative. + +Treat allocator process replacement as an ownership transition unless the target explicitly proves concurrent incarnations are harmless. A new allocator-incarnation token is safe only when it participates in the **authoritative fencing check**, not merely in the emitted ID namespace. Acceptable designs include: (1) advance the domain ownership/fencing epoch for every replacement allocator process, so the predecessor becomes stale automatically; or (2) maintain a separate monotonically increasing allocator-incarnation epoch that every mutation/allocation request carries and the authoritative boundary validates together with the domain ownership epoch. In both designs, replacing a paused/partitioned allocator invalidates the previous incarnation before the replacement may serve work. Merely embedding a fresh incarnation value in IDs prevents collisions but does **not** preserve exclusive ownership if the superseded process can still mutate queues/storage/state. + +Where local IDs/counters are used, combine this fencing rule with a restart-safe allocation policy. Acceptable durable allocation designs include: (1) a high-water mark advanced atomically **before** an ID/range becomes externally usable, or (2) durable allocation of non-overlapping ranges/blocks so a restart resumes from a fresh unissued block and may safely burn any uncertain tail. A per-incarnation namespace may additionally participate in emitted IDs, but for exclusive-owner targets it cannot substitute for fencing the superseded incarnation at the mutation boundary. + +If IDs must remain stable across allocator restarts and ownership epochs, use a durable monotonic counter/high-water mark or durable non-overlapping range allocator; do **not** reset an ephemeral counter. Persist/reserve advancement before returning the corresponding ID to the caller, or otherwise use a transaction whose crash semantics can prove that recovery never reissues an already-visible value. When commit status is uncertain after a crash, prefer skipping/burning an uncertain range over risking reuse. + +Treat finite namespace capacity as an explicit feasibility bound, not an implementation afterthought. Compute the maximum allocatable local/range/epoch space under the chosen bit layout and the required lifetime/concurrency horizon. Before the final representable value can escape, the allocator must deterministically enter one of three predeclared states: (1) stop/reject new allocations with an explicit exhaustion result, (2) migrate to a wider representation under a compatibility plan, or (3) transition to a new epoch/namespace under a proof that concurrently valid old identities cannot collide with new ones. Never use wraparound as an exhaustion policy. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: Not established in a target repository. +- Optimized: Not established in a target repository. +- Speedup / memory reduction: No transferable claim; donor observations motivate the pattern only. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Check global invariants across all domains, collision/duplicate behavior, restart behavior and target-scale contention profiles. + +Exercise **same-owner allocator restarts** independently of reassignment. Issue IDs/ranges, crash the allocator before and after each persistence/reservation boundary, restart it, and prove it never reissues an externally visible ID/range. Test crashes after durable reservation but before delivery, after delivery but before acknowledgement bookkeeping, and with uncertain commit status. For high-water counters, verify monotonic durable recovery. For block/range allocation, verify recovered allocators never enter a previously issued block and that burning an uncertain tail preserves uniqueness. + +Exercise **superseded-incarnation races** separately from ID-collision tests. Pause or partition allocator incarnation A without proving it dead, start replacement incarnation B, make B's fencing identity authoritative, then resume A. Prove the authoritative mutation/allocation boundary rejects every operation from A even if A's generated IDs would be collision-free because of a different incarnation namespace. Repeat with delayed A messages, queue acknowledgements, counter updates, and storage mutations. If the target uses a separate allocator-incarnation epoch, prove both ownership epoch and incarnation epoch are validated wherever exclusivity matters. If the target instead deliberately allows concurrent incarnations, state that as a contract change and verify all affected operations are designed for multi-writer semantics. + +Exercise **rebalance/recovery races**: pause an owner, expire/revoke it, assign a higher fencing identity to a replacement, then resume the old owner and prove every stale mutation/allocation/queue claim is rejected. Inject network partition and split-brain conditions where both old and new processes run simultaneously. Verify only the highest authoritative fencing identity can mutate state, no duplicate IDs/work claims are produced, handoff is crash-recoverable, and ownership remains unique through coordinator/storage restarts. Include delayed messages from old epochs arriving after the new owner has already committed work. + +Exercise **finite-space exhaustion** at every boundary the representation can hit. Start close enough to the maximum local/range/epoch value to cross the boundary quickly; race multiple allocators against the final available slot; crash before and after reserving the final range; and verify that the next request receives the declared stop/rejection or enters the prevalidated widening/epoch-transition path. Prove no wrap, duplicate, historical collision, or stale-owner resurrection occurs. For widening or epoch transitions, test mixed old/new readers and any period where old identities remain externally valid. + +## Target-repo adaptation + +Shard counts and bit splits are workload-specific. Measure skew, cache locality, failure domains and merge costs. Define the durable ownership source, lease timeout if used, monotonically increasing fencing epoch/token, authoritative mutation boundary that validates fencing identity, handoff sequence, and restart/recovery semantics before enabling dynamic reassignment. For allocators, separately define the same-owner restart policy and the replacement-process fencing policy. Either advance the ownership epoch for every replacement or make allocator-incarnation epochs first-class fencing tokens at the mutation boundary. Do not rely on incarnation namespacing alone unless concurrent incarnations are explicitly admissible. Specify exactly which state is made durable before an ID/range can escape and how ambiguous crash outcomes are recovered without reuse. + +For finite-width identities, document total capacity, capacity consumed by domain/incarnation/epoch fields, worst-case allocation rate, required service horizon, alert/headroom threshold, and the exact exhaustion policy. Validate any representation widening or epoch transition before deployment rather than treating exhaustion as an unreachable state. + +## Failure modes + +Hot shards merely move the bottleneck; domain proliferation raises memory/management overhead; rebalancing without fencing can allow stale and replacement owners to act concurrently; lease-only ownership can fail when an old process resumes; delayed old-epoch messages can duplicate allocations or queue work; a fresh allocator-incarnation namespace can hide ID collisions while still allowing two owners to mutate the same domain; a same-owner restart can reset an ephemeral local counter and reissue prior IDs even without any fencing race; persisting allocation state after delivery can create crash windows that reuse visible values; a finite local/domain/incarnation/epoch field can exhaust and wrap into an already-valid identity if capacity is not treated as a hard feasibility constraint; identity stability may be violated; global ordering requirements may make the pattern inadmissible. + +## Rollback trigger + +Revert if partitioning does not reduce measured contention, if any cross-domain invariant fails, if failover/rebalance/replacement testing shows a stale owner or superseded allocator incarnation can mutate state after a replacement becomes authoritative, if incarnation namespacing prevents duplicate IDs but does not fence the old process where exclusive ownership is required, if same-owner crash/restart testing can reissue any externally visible ID/range or otherwise lose durable allocator progress, or if exhaustion testing can wrap/reuse an identity, enter an undefined state, or perform an unvalidated widening/epoch transition. Reject the chosen namespace layout before deployment when projected capacity cannot satisfy the required service horizon with the declared exhaustion policy. diff --git a/optimizations/OPT-CRIT-001-critical-path-prioritization.md b/optimizations/OPT-CRIT-001-critical-path-prioritization.md new file mode 100644 index 0000000..f0312bd --- /dev/null +++ b/optimizations/OPT-CRIT-001-critical-path-prioritization.md @@ -0,0 +1,78 @@ +# OPT-CRIT-001 — Critical-path prioritization + +**Status:** Proposed / OPT synthesis; target validation required +**Domains:** UI, web, games, build systems, model/data loading, interactive pipelines + +## Source evidence + +- `davidsonfellipe/awesome-wpo` inspected at `84f32948a6298456d6a94cff64551f39f2666e6f` +- resource-hint, lazy-loading and prefetch references catalogued upstream +- `sources/WPO.md` + +## Problem + +Non-critical work competes with the dependency chain that determines user-visible or pipeline latency. + +## Optimization problem contract + +- X: target-supported task-priority, prefetch/precompute, lazy/deferred-work, speculation, speculative-input identity, mutation-control, commitment, critical-capacity reservation, preemption/cancellation, and speculation-admission policies +- F: policies that preserve all semantic deadlines, avoid externally visible speculative side effects before commitment, commit speculative results only from one stable effective-input generation, satisfy starvation/resource constraints, and enforce enough protected or promptly reclaimable capacity that speculative work cannot occupy every resource a newly arriving critical task may need +- f: measured end-to-end latency of the declared critical dependency path, including resource pressure introduced by speculation/deferment +- d: minimize +- C: critical outputs and semantic deadlines are preserved; speculative work is safely discardable; any speculative result is bound to a complete immutable snapshot or full-duration mutation witness, and validation of that witness is linearized with commitment so intervening or final-window A→B→A/input changes cannot be erased before visibility; deferred work completes before it becomes semantically required; and speculative occupancy cannot delay newly arriving critical work beyond the declared critical-start/latency bound because critical capacity is reserved, speculative work is preemptible/cancellable within a bounded reclaim latency, or speculation admission is hard-limited to leave sufficient headroom +- B: target-specific trace/benchmark budget covering cold/warm, hit/miss, wrong-speculation, stale-speculation, change/revert, and critical-arrival-under-saturation cases; no portable prediction horizon is supplied here +- S: stop when the declared budget is exhausted or a validated policy materially reduces critical-path latency without violating C +- Variables: categorical / conditional / mixed priority, deferment, prefetch, speculation, capacity-reservation, preemption, and admission policies +- Search scope: local critical-path policy tuning +- Objective behavior: noisy under realistic workload timing; semantic identity/deadline/capacity checks are deterministic +- Information: derivative-free / black-box latency measurements +- Evaluation cost: moderate to expensive end-to-end tracing/benchmarking +- Constraints: semantic deadlines, starvation, side effects, input identity/mutation freshness, commitment linearizability, protected/reclaimable critical capacity, memory/CPU/I/O, and target resource constraints +- Parallelism: asynchronous / concurrent execution is common; speculative dispatch must preserve enforceable critical headroom or bounded preemption +- Exactness: exact target semantics; speculative work may be discarded but not committed stale or allowed to violate critical-capacity guarantees + +## Preserved contract + +Deferred work must still complete before its semantic deadline. Speculative work must be discardable and must not create externally visible side effects before commitment. A speculative result may be committed/delivered only if it was produced from one coherent effective-input generation equivalent to the non-speculative reference path; endpoint equality after an intervening mutation is not sufficient, and a successful freshness check is not sufficient unless the checked identity remains authoritative through the commit that makes the result visible. Priority must also remain operational rather than nominal: speculation may not consume all capacity needed by a critical request that arrives after speculative work has started. The target must preserve a declared critical-start/latency bound using reserved capacity, bounded-latency preemption/cancellation, or hard admission limits that leave sufficient headroom for non-preemptible speculation. + +## Optimization + +Execute critical dependencies first; prefetch/precompute likely-soon work only when probability and resource policy justify it; lazily defer non-critical work; avoid work with no demonstrated demand. + +Treat "spare at dispatch" as insufficient evidence that speculation is safe. Before launching speculative work, perform an **enforceable critical-capacity admission check**. For non-preemptible speculation, reserve the worker, I/O, accelerator, memory, connection, queue, or other resource capacity required by the declared critical workload, or hard-limit speculative concurrency/occupancy so worst-case admitted speculation cannot consume that headroom. For preemptible/cancellable speculation, define and enforce a maximum reclaim latency and ensure cancellation/preemption returns enough capacity before the critical-start/deadline bound can be violated. Capacity accounting/admission must be atomic enough that concurrent speculative launches cannot each observe the same final spare slot and collectively consume protected headroom. + +Bind every speculative/precomputed result to a complete effective-input identity for the **full speculation-to-commit interval**. Prefer speculation against an immutable snapshot/version. If snapshots are unavailable, use a full-duration mutation/read lock or capture a monotonically increasing, non-reusable version/epoch for every mutable effective input. Every relevant mutation must advance its witness, including A→B→A changes that restore original bytes. A commit-time hash/identity comparison may supplement the mutation witness but must not be the sole freshness proof. + +Freshness validation and commitment must be **one linearizable operation**. For lock-based targets, hold the mutation/read lock through the exact commit/publication/delivery transition that makes the speculative result externally visible. For epoch/version-based targets, use an atomic compare-and-commit/conditional transaction that verifies the complete coherent epoch vector is still the witnessed vector and, only if that comparison succeeds in the same atomic boundary, publishes the result. A separate `check epochs; later publish` sequence is not sufficient. If the compare-and-commit loses a race, discard the speculative result and execute/recompute from the current reference identity. Any lock violation, epoch change, incoherent witness, or untrackable mutable input likewise forces discard/recompute. + +Commitment is the semantic boundary: no stale speculative result may become externally visible merely because the speculation itself had no side effects. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: No target critical-path profile has been established here. +- Optimized: No target prioritization/prefetch policy has been benchmarked here. +- Speedup / memory reduction: No transferable claim; upstream WPO material supplies patterns and measurement guidance. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Trace the true dependency path and measure end-to-end latency, not only individual task duration. Test cold/warm, cache-hit/miss and wrong-speculation cases. Explicitly test semantic deadlines, starvation, cancellation, and that speculative work cannot expose side effects before commitment. + +Add a **critical-arrival-under-speculation-saturation** race. Fill speculative work to the maximum admitted occupancy, then introduce a critical request requiring each protected resource class (for example a worker plus I/O or accelerator capacity). For reserved-capacity designs, prove the critical request can acquire the reserved capacity without waiting for non-preemptible speculative completion. For preemptible/cancellable designs, prove enough speculation is reclaimed within the declared maximum reclaim latency to satisfy the critical-start/deadline bound. For admission-bound designs, prove concurrent speculative launches cannot race past the headroom limit. Repeat with non-preemptible long-running speculation, simultaneous critical arrivals, and mixed-resource bottlenecks; a policy that merely observed spare capacity before dispatch must fail this fixture if it can later block the critical path. + +Add stale-speculation fixtures with explicit **A→B→A** races. Start speculation from identity A, mutate the effective inputs to B while speculation reads/runs, then restore original bytes before demand/commitment. For snapshot-based targets, prove speculation consumed only immutable A. For lock-based targets, prove the mutation cannot interleave. For epoch/version-based targets, prove every mutation increments the monotonic witness and that the final witness exposes the intervening change even though endpoint content equals A. Also test delayed speculative completion, version rollback, and concurrent config/schema changes. Compare every committed speculative result against the non-speculative reference path for the exact committed identity. + +Add a **final validation-to-commit race**. Pause immediately after the last ordinary witness comparison but before the result would become visible, then mutate an effective input. For lock-based designs, prove the mutation is blocked until after commitment. For epoch/version designs, prove the atomic compare-and-commit rejects the stale speculative result rather than publishing it. Repeat with A→B→A and multi-input epoch-vector changes. No fixture may pass by doing an ordinary comparison followed by a separate publication step. + +## Target-repo adaptation + +Criticality and prediction horizons are workload-specific. Re-profile after topology or user-flow changes. Define the complete effective-input identity for each speculative result and choose immutable snapshots, full-duration mutation locks, or monotonic epochs that record every intervening change. Define the linearization boundary that couples freshness validation to external commitment: lock-through-commit or atomic compare-and-commit. Specify exactly when a stale speculative result is discarded. Also define the **critical-capacity invariant** for every contended resource: how much capacity is reserved, what speculative occupancy ceiling applies, or which work is preemptible/cancellable and the maximum reclaim latency. Make admission/concurrency accounting race-safe, and do not treat currently idle capacity as sufficient if non-preemptible speculation can consume it before future critical arrivals. Do not rely on commit-time endpoint revalidation alone, or on check-then-publish epoch validation, to establish freshness. + +## Failure modes + +Speculation steals resources from critical work; non-preemptible speculation can fill every worker, I/O slot, accelerator slot, connection, or other bottleneck before a new critical request arrives; concurrent speculative launches can oversubscribe supposedly reserved headroom; preemption/cancellation can be too slow to protect the critical-start/deadline bound; lazy work causes later latency cliffs; priorities become stale; deferred tasks starve; semantic deadlines are missed; speculative side effects escape before commitment; A→B→A mutations can fool endpoint-only freshness checks; non-monotonic/reused epochs can erase intervening changes; a check-then-publish window can expose stale speculation after a successful freshness check; or stale speculative output is committed after its effective inputs changed. + +## Rollback trigger + +Immediately disable/revert the policy on any violation of C, including a required task missing its semantic deadline; a critical request being delayed beyond the declared start/latency bound because speculative work consumed protected capacity; a reservation/admission race allowing speculation to exceed its occupancy ceiling; preemption/cancellation failing to reclaim capacity within its declared bound; speculative work exposing an externally visible side effect before commitment; a speculative result being committed/delivered without an immutable snapshot/lock/monotonic mutation witness proving one coherent effective-input generation; or any test showing freshness validation can be separated from commitment so a mutation can win in between. Also disable it if critical-path latency or resource pressure worsens materially. diff --git a/optimizations/OPT-FAN-001-shared-materialization-fanout.md b/optimizations/OPT-FAN-001-shared-materialization-fanout.md new file mode 100644 index 0000000..b464d17 --- /dev/null +++ b/optimizations/OPT-FAN-001-shared-materialization-fanout.md @@ -0,0 +1,87 @@ +# OPT-FAN-001 — Shared materialization for fan-out and replay + +**Status:** Implemented external reference; target validation required +**Domains:** streaming, serialization, compression, artifact pipelines, multi-consumer services + +## Source evidence + +- Jazco, Jetstream: https://jazco.dev/2024/09/24/jetstream/ +- `sources/JAZCO.md` + +## Problem + +The same deterministic transformation is repeated independently for each consumer and again during replay. + +## Optimization problem contract + +- X: target-supported materialization boundaries, representation formats/versions, persistence policies, raw-versus-materialized retention policies, complete materialization-key definitions, immutable-source snapshot/mutation-control policies, monotonic source/config epochs, source-witness compare-and-publish policies, reuse-hit source-binding policies, artifact-version pinning/consumption policies, and crash-consistent publication schemes +- F: configurations whose materialized representation satisfies every declared consumer semantic, versioning, integrity, trust, materialization-equivalence, source-snapshot/mutation-consistency, **source-validation-to-publication linearizability**, **reuse-hit source-binding linearizability**, validation-to-consumption identity, and publication-atomicity requirement +- f: measured transformation CPU, replay CPU, fan-out latency, and storage/I/O overhead under the target's declared objective ordering +- d: minimize under the target's predeclared scalar or lexicographic ordering +- C: consumers receive the declared representation semantics exactly; reuse is allowed only when one committed state binds the artifact bytes to one coherent effective source/transform identity, no intervening mutable-input change can be erased by endpoint equality, the final source witness is linearized with the materialization commit, **each reuse-hit decision is linearized against the invocation's effective source identity before delivery**, and every consumer reads the **same immutable/versioned artifact instance that was validated** rather than re-resolving a mutable alias after validation; verification/security metadata may be removed only under an explicit contract change +- B: target-specific fan-out/replay benchmark budget declared before tuning; no portable subscriber count, replay size, or retention duration is supplied here +- S: stop when the declared budget is exhausted or a validated materialization policy materially improves the target objective without violating C +- Variables: categorical / integer / mixed +- Search scope: local materialization-boundary / representation-policy tuning +- Objective behavior: noisy for performance; transformation identity/equivalence is deterministic +- Information: derivative-free / black-box performance measurements +- Evaluation cost: moderate to expensive depending on transform/replay size +- Constraints: semantic equivalence, source-snapshot/mutation consistency, source-validation/publication linearizability, reuse-hit source binding, artifact-version pinning, integrity, versioning, trust/security, storage, and crash-consistency constraints +- Parallelism: concurrent fan-out/replay; source publication, reuse-hit binding, artifact publication, and consumption pinning must remain race-safe +- Exactness: exact representation semantics; no approximation is introduced + +## Preserved contract + +Consumers must receive the same declared representation semantics. A persisted representation is reusable only under a named **materialization-equivalence invariant** that binds the artifact to every effective input capable of changing its bytes or semantics, and that binding must survive source mutation, change-and-revert races, final-check-to-commit races, reuse-hit races, crashes, interrupted publication, and concurrent replacement of mutable aliases. Validation is meaningful only if publication, the caller's reuse-hit decision, and downstream consumption remain bound to the exact source/artifact identities that were validated. Removing verification/security metadata is **not** a correctness-preserving optimization unless the interface contract explicitly changes. + +## Optimization + +Perform an expensive deterministic transform once near production, persist or retain the reusable representation, and fan out/replay those bytes/objects rather than reconstructing them per consumer. + +Define a materialization key that covers, as applicable, source object/content identity or immutable source version, transformation/encoder implementation identity, encoder configuration and dictionaries, schema/format version, feature flags, trust/security policy, and any other effective input that can affect the materialized result. + +Bind the transform to one coherent source identity for the **entire transform-to-commit interval**. Prefer reading every mutable effective input from an immutable snapshot/version captured together with the materialization key. If immutable snapshots are unavailable, use a mechanism that records intervening mutation rather than comparing endpoint content alone: hold an appropriate mutation/read lock for the full interval, or capture a monotonically increasing, non-reusable version/epoch for each mutable source/config/transform input. Every mutation must advance its witness durably/atomically with the mutation, including A→B→A changes that restore the original bytes. A final content/key recomputation may supplement this witness but must not be the sole protection. Any lock violation, epoch change, incoherent multi-input witness, or untrackable mutable input invalidates the candidate materialization; discard/retry it rather than publishing mixed-state bytes. + +For lock/epoch-based targets, make the **final witness validation and materialization activation one linearizable transition**. A check followed later by manifest publication is insufficient because the source may mutate in that gap. Either hold the mutation/read lock through the authoritative artifact/manifest switch, or use an atomic compare-and-publish/transaction that verifies the complete witnessed epoch vector and publishes the new materialization only if those epochs are still current in the same serialization domain that advances them. If any epoch changed, the activation must fail and the candidate remains non-authoritative. Immutable snapshots satisfy this requirement only when the committed materialization remains explicitly bound to that immutable source version rather than to a mutable alias. + +Publish the artifact and its identity as **one committed state**. Acceptable designs include content-addressed storage where the artifact digest is itself part of the committed key, an atomically replaced manifest that contains both the full materialization key and the artifact digest/location, or another crash-consistent transaction that makes old state or new state visible but never a mixed pair. Do not update artifact bytes and their key independently in a way that can expose a new artifact with stale metadata or stale bytes with a new key after a crash. + +Before reuse, require: (1) exact agreement with the current effective-input materialization key and immutable snapshot/monotonic mutation identity, (2) a committed manifest/content-address relation that binds that identity to the artifact identity, and (3) artifact integrity/format validity. **The invocation must then bind the reuse hit to that same effective-input identity in one linearizable step before the artifact is delivered.** Prefer capturing an immutable invocation snapshot/version before lookup. Otherwise hold the source/config mutation lock through the hit decision, or use an atomic compare-and-bind/CAS that verifies the complete current epoch vector and records the hit only if it is still current in the same serialization domain that advances those epochs. A check-then-deliver sequence is insufficient: if the source advances from A to B after the key/epoch check but before the caller is bound to the hit, the invocation must miss/retry against B rather than receiving A merely because A's artifact remains valid for its own historical source version. A caller may deliberately consume A only when its invocation was already bound to immutable source identity A before the mutation. + +**Validation must return or retain an immutable/versioned artifact handle/snapshot that uniquely identifies the validated bytes. Every fan-out/replay consumer must read through that same pinned handle/version.** Do not validate mutable pathname/object-name A and later re-resolve that alias for consumption. If the target cannot provide immutable/versioned handles, hold an appropriate read/replacement lock from the integrity check through the complete consumer read, or first create an immutable snapshot and validate/consume that snapshot. A key mismatch, mutation-epoch mismatch, failed reuse-hit compare-and-bind, missing/incomplete publication marker, digest mismatch, unverifiable artifact, or inability to bind consumption to the validated bytes is a cache miss/fail-closed condition and requires regeneration or safe fallback. Do not use format validation, endpoint key equality, or a mutable location name alone as evidence that the bytes consumed are the bytes validated. + +For multiple consumers, each may hold its own reference to the same immutable validated artifact version, or the system may retain one immutable snapshot for the fan-out/replay lifetime. Reclamation/retention must not invalidate a pinned consumer handle before that consumer completes. A mutable alias may advance to a newer committed materialization for later callers without changing the version already pinned by an in-flight consumer. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: Not established in a target repository. +- Optimized: Not established in a target repository. +- Speedup / memory reduction: No transferable claim; Jetstream observations remain external source evidence. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Compare shared materialization against per-consumer reference output, including corruption, restart/replay and mixed consumer capabilities. Independently mutate each key component—source content/version, transform implementation, encoder options/dictionary, schema/format version, feature flags and trust policy—and prove that each output-affecting change invalidates reuse. Also test unchanged-key reuse, tampered artifacts with matching metadata, and migration/version-boundary cases. + +Exercise **concurrent source mutation**, including explicit A→B→A races. Start a transform from source identity A, mutate one or more source/config/transform inputs to B while the transform is running, then restore the original bytes before commit. For snapshot-based targets, prove the transform reads only the immutable A snapshot. For lock-based targets, prove mutation cannot interleave with the protected transform/publication interval. For epoch/version-based targets, prove every mutation advances the monotonic witness and the final witness differs even when the final content/key returns to A. Reject/discard the candidate on any mutation witness change and compare every accepted materialization with a fresh transform from the exact committed source identity. + +Add a **final source-validation-to-publication race**. Pause after the last source epoch/vector check but before the new materialization becomes authoritative, mutate one source/config/transform input, then resume publication. For lock-based targets, prove the mutation cannot occur until after the authoritative switch. For epoch/CAS-based targets, prove the compare-and-publish fails because the current epoch vector no longer equals the witnessed vector; no consumer may observe the candidate as committed. Repeat with A→B→A content restoration, multiple inputs, and a concurrent manifest reader. A successful endpoint rehash after the mutation is not sufficient evidence. + +Add a **reuse-hit source-binding race**. Start an invocation against source/config identity A and pause after the committed manifest, current epoch/key, and artifact integrity checks succeed but before the caller is irrevocably bound to that hit. Mutate one effective input so the current identity becomes B, then resume. Snapshot-based targets must prove the invocation was already bound to immutable A before lookup; lock-based targets must prove the mutation cannot pass the hit-binding boundary; epoch/CAS targets must prove the compare-and-bind fails and the invocation misses/retries against B. The test must never deliver pinned artifact A as a cache hit for an invocation whose effective identity became B before binding. Repeat with A→B→A mutation, multiple inputs, and concurrent callers straddling the identity transition. + +Add a **post-validation replacement race**. Validate committed artifact A, pause before a consumer reads it, replace the mutable alias/path/object name with a different valid artifact B, then resume consumption. Prove a pinned immutable/versioned handle still yields exactly A (or fails closed if A was invalidated by the target's retention contract), never unvalidated B. Repeat with fan-out consumers at staggered start times, concurrent manifest advancement, replay after alias replacement, reclamation pressure, and mutable object-store/version aliases. For lock-based targets, prove replacement cannot occur until the protected consumer read completes. For snapshot-based targets, prove validation and consumption address the same snapshot digest/version. + +Inject crashes/interruption at every publication boundary: after artifact write but before manifest commit, after provisional metadata write, during atomic replacement, and immediately after commit. After restart, prove that readers see either the previous valid committed materialization or the new valid committed materialization, never a mixed key/artifact state. Verify digest/key mismatch is rejected even when the artifact is otherwise parseable. + +## Target-repo adaptation + +Define the complete materialization-equivalence invariant for the target, choose the identity primitive for each effective input, specify whether mutable inputs are consumed from immutable snapshots, protected by full-duration locks, or guarded by monotonic mutation epochs, and define how a coherent multi-input witness is captured. For epoch/lock targets, define both serialization boundaries explicitly: the mechanism that makes final witness validation atomic with authoritative materialization publication, and the mechanism that makes a reuse-hit decision atomic with the invocation's current source/config identity (immutable invocation snapshot, lock-through-hit-binding, or atomic compare-and-bind against the complete epoch vector). Specify representation versioning, invalidation, integrity checking, **the immutable/versioned artifact handle or lock/snapshot that binds validation through consumption**, retention/reclamation semantics for pinned consumers, **crash-consistent publication/commit mechanics**, storage-vs-CPU trade-offs and whether both raw and materialized forms are retained. Do not advertise commit-time endpoint rehashing, check-then-deliver reuse, or mutable-path validation alone as sufficient identity protection. + +## Failure modes + +Incomplete keys can serve stale representations after source or transform changes; mutable sources can change during transformation and produce mixed-state output under a stale key; A→B→A races can defeat endpoint key comparisons; non-monotonic/reused mutation versions can erase intervening changes; incoherent epoch vectors can describe no real source state; a check-then-publish gap can authorize A-derived bytes after the source already advanced to B; a reuse check followed by a later hit/delivery can bind artifact A to an invocation whose source already advanced to B; non-atomic publication can pair new bytes with an old key or vice versa after a crash; metadata can match while artifact bytes are corrupted; a mutable alias can be replaced after validation and before consumption, delivering unvalidated bytes; reclamation can invalidate a pinned artifact prematurely; materializing unused forms wastes storage; format changes create invalidation/migration costs; mutable consumer-specific transformations cannot safely share one artifact. + +## Rollback trigger + +Disable reuse immediately if any materialization-key hit, source-mutation race, final source-validation-to-publication race, reuse-hit source-binding race, publication-recovery path, integrity check, or validation-to-consumption race can return output that differs from a fresh transform for the same exact committed effective inputs; if an A→B→A race can evade the snapshot/lock/monotonic mutation witness; if source epochs can change after a successful final check yet the candidate can still become authoritative; if a source/config identity can change after a successful reuse check yet the invocation can still be bound to the stale hit; if a mutable alias replacement can make a consumer read bytes other than the exact artifact version that passed validation; if a pinned artifact can be reclaimed before consumption completes; or if interrupted publication can expose a mixed key/artifact state. Also disable when storage/invalidations outweigh avoided transform work or representation equivalence fails. \ No newline at end of file diff --git a/optimizations/OPT-INC-001-signature-bound-incremental-execution.md b/optimizations/OPT-INC-001-signature-bound-incremental-execution.md new file mode 100644 index 0000000..2d0b1b9 --- /dev/null +++ b/optimizations/OPT-INC-001-signature-bound-incremental-execution.md @@ -0,0 +1,89 @@ +# OPT-INC-001 — Signature-bound incremental execution + +**Status:** Implemented external reference; historical donor, target validation required +**Domains:** builds, CI, generated artifacts, preprocessing, scientific pipelines + +## Source evidence + +- `psycledelics/wonderbuild` commit `021d5ed7c298c6c34b091cf5e6d9802e200028a6` +- `sources/WONDERBUILD.md` + +## Problem + +Expensive work is rerun even though every input capable of affecting its result is unchanged. + +## Optimization problem contract + +- X: target-supported signature definitions, persistence scopes, invalidation granularities, output-validity/consumption policies, immutable-output handles, immutable-input snapshot/mutation-control policies, monotonic mutation epochs, **reuse-hit input-binding policies**, compare-and-publish activation policies, and crash-consistent state-publication mechanisms +- F: configurations whose signature covers every output-affecting input, whose execution consumes one immutable effective-input snapshot or is protected by a mutation lock/monotonic mutation witness that detects every intervening change, whose **reuse-hit decision is bound to one coherent effective-input identity and linearized against that same mutation witness**, whose reuse validates required outputs and binds downstream consumption to the exact validated output versions, whose authoritative generation switch is linearized with the witnessed input state, whose persisted signature/output metadata form one committed generation, and whose failed/interrupted/raced executions never publish reusable partial or stale-current state +- f: measured repeated-work cost including stage runtime plus signature/snapshot/mutation-tracking/metadata/output-validation/output-pinning/**reuse-hit binding**/compare-and-publish/publication I/O overhead +- d: minimize +- C: every reused output consumed downstream is the same immutable/versioned output instance whose validity predicate passed, and is semantically equivalent to a fresh execution for the **coherent effective-input identity bound to that invocation** with the same failure semantics; a reuse hit cannot be accepted by comparing against identity A and then consume A after a concurrent mutation has already made B the invocation's authoritative live identity unless the invocation was explicitly and immutably snapshot-bound to A; reuse metadata cannot mix fields from different generations; a committed generation cannot bind a pre-execution signature to output produced from changed or mixed inputs; A→B→A mutations during execution or reuse-hit qualification are detected rather than erased by endpoint equality; and no input mutation may linearize between the final accepted input witness and activation of a generation or acceptance of a reuse hit for those inputs +- B: target-specific benchmark/evaluation budget declared before tuning; no portable value is supplied by this record +- S: stop when the declared budget is exhausted or a validated configuration meets the predeclared improvement threshold without violating C +- Variables: categorical / mixed policy choices for signatures, snapshots, validation, output pinning, granularity, mutation control, reuse-hit binding, compare-and-publish activation, and publication +- Search scope: local to one incremental stage or pipeline boundary +- Objective behavior: noisy for performance; correctness identity/mutation checks are deterministic +- Information: derivative-free / black-box performance measurements +- Evaluation cost: moderate to expensive depending on stage runtime and validation cost +- Constraints: semantic equivalence, input/output integrity, crash consistency, snapshot/mutation consistency, **reuse-hit input identity**, validated-output consumption, publication linearizability, and resource constraints +- Parallelism: sequential or pipeline-specific; mutation tracking, reuse-hit binding, output pinning, activation, and publication must remain race-safe under concurrent producers/consumers +- Exactness: exact reuse semantics; no approximation is introduced + +## Preserved contract + +Reused output must be semantically equivalent to a fresh execution for the same effective inputs. Failed executions must not bless a new signature, an unchanged input signature alone is insufficient when an existing output can be corrupted or overwritten externally, interrupted publication must not expose a signature paired with output identities from another generation, and mutable inputs must not change underneath execution without invalidating the candidate generation. Endpoint equality is not enough: if an input changes and later returns to its original bytes, the intervening mutation must still be observable to the publication decision. Likewise, validating a mutable output path is not enough unless downstream consumption is pinned to that exact validated version. Validating an input signature for a **reuse hit** is also not enough unless the invocation is bound to that same coherent input identity before a concurrent mutation can change which generation is authoritative for the call. Finally, validating the input witness and later switching the authoritative generation are not two independent steps: generation activation must linearize against the same input state that was validated. + +## Optimization + +Compute a deterministic signature over the effective inputs and compare it with successfully persisted prior state. Reuse is allowed only when that signature still matches **and** every required output satisfies a declared validity predicate. Depending on the target, that predicate may be a content digest/version manifest, a trusted immutable/protected artifact identity, or another reproducible integrity check strong enough to detect external mutation. Mere file presence is not sufficient unless the target explicitly guarantees that reused outputs are immutable and protected from modification. Execute when the input signature differs, any required output is missing, or any output-validity check fails. + +A **reuse hit must itself be identity-bound and linearizable**. Before treating a matching signature as permission to skip execution, the invocation must acquire one coherent effective-input identity and bind the hit to it. Preferred designs capture an immutable source snapshot/version for the invocation. Lock-based designs may instead hold the mutation/read lock while verifying that the current committed generation matches the captured identity and while binding the invocation to that generation. Epoch/version designs must use a transaction, CAS, compare-and-bind operation, or equivalent serialization boundary that atomically verifies the complete current mutation-epoch vector, verifies that the reusable generation was committed for that exact vector/signature, and records/returns the invocation's binding to that generation and its pinned output handles. A plain sequence of “signature A matches; later decide to skip/read A” is insufficient. If an input mutation wins before the hit is bound, the compare/bind must fail and the invocation must re-evaluate against the new identity. If the target's semantics are **snapshot-at-invocation**, later mutations may proceed only after the immutable A snapshot/binding is established; if the semantics require **current live inputs through consumption**, retain the appropriate lock or equivalent freshness protection through the required consumption boundary. + +Treat output validation and output consumption as one identity-bound operation. A successful validity check must yield or pin the exact immutable/versioned output handle that downstream consumers will read: for example a content-addressed object, immutable artifact/version ID, snapshot handle, open file descriptor tied to a protected inode/version where the platform guarantees the needed semantics, or another target-specific stable handle. Do **not** validate bytes at a mutable pathname/object name and then later reopen that name for consumption, because another writer may replace it between validation and read. If the storage system cannot provide an immutable/versioned handle, hold an appropriate lock from validation through the downstream read/consumption, or copy the validated bytes into an immutable snapshot and consume that snapshot. Every consumer on a reuse hit must be bound to the validated handle/version, not merely to the same logical path. + +Bind execution to one coherent effective-input identity. The preferred design is an **immutable snapshot/version** of every mutable effective input. If a snapshot is unavailable, use a mechanism that records *intervening mutation*, not merely endpoint content equality: for example, hold a read/mutation lock for the full execution-through-activation interval, or capture a monotonically increasing version/epoch for every mutable input. Every mutation must advance its epoch durably/atomically with the mutation, including a change that later restores the original bytes. For multiple inputs, capture the snapshot/epoch vector coherently under the target's transaction/locking rules so a mixed vector cannot be mistaken for one state. A content signature recomputed at publication may supplement this check, but **must not be the sole fallback** because A→B→A can make endpoint signatures equal. Any lock violation, epoch change, incoherent snapshot, or untrackable mutable input discards the candidate generation and requires retry from a fresh identity. + +The **authoritative generation switch must be linearized with that witness**. A plain sequence of “read epochs; they match; later replace the current-generation manifest” is insufficient. For lock-based designs, retain the mutation/read lock through the atomic generation-pointer/manifest switch. For epoch/version designs, use one transaction, CAS, compare-and-swap manifest operation, or equivalent serialization boundary that atomically verifies the complete witnessed epoch vector still matches **and** activates the new generation; if the comparison fails, the candidate remains non-current and must be discarded or retried. For immutable-snapshot designs, immutable candidate artifacts may be produced independently, but making one authoritative for the live mutable source still requires an atomic check that the live source identity/version remains the captured snapshot identity at the activation point. There must be no interval in which an input mutation can win after the accepted witness check but before the generation becomes authoritative. + +Publish incremental state as one crash-consistent **generation** that binds the validated input signature and immutable snapshot/epoch identity to the complete output identity/validity metadata. Do not persist the signature and output metadata as independently authoritative updates. Use an atomic rename/swap of a complete manifest, a transactional store, a content-addressed generation pointer, or another mechanism where readers observe either the previous complete generation or the new complete generation—never a mixture. Only activate the new generation after every output has been produced and validated successfully **and** the input witness plus authoritative-generation switch have succeeded in the same linearizable activation boundary; an interrupted, failed, input-raced, or compare-and-publish-failed candidate leaves the previous committed generation authoritative and the candidate generation non-reusable. + +Reuse filesystem/configuration metadata lazily only while its own validity predicate still holds. + +## Evidence boundary + +Wonderbuild demonstrates the mechanism and benchmark shapes, but its historical timings and timestamp/hash choices are not transferable targets. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: Not established in a target repository. +- Optimized: Not established in a target repository. +- Speedup / memory reduction: No transferable claim. Wonderbuild observations are historical source evidence only. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Test unchanged inputs with valid outputs, changed inputs, missing outputs, failed runs, corrupted persistent state, externally overwritten/corrupted outputs, and stale output-version metadata against a forced-fresh reference path. A mutated output must force reconstruction unless the target's immutable/protected-output contract proves such mutation impossible. + +Add a **reuse-hit input-identity race**. Arrange a committed reusable generation for input identity A. Pause after the implementation has observed what would otherwise be a signature/epoch match for A but before it has irrevocably bound the invocation to the reusable generation and pinned output handles. Mutate the effective input to B in that exact interval. For snapshot-at-invocation designs, prove the call was already bound to an immutable A snapshot before the mutation and therefore legitimately consumes A. For lock-based live-input designs, prove the mutation cannot interleave until the required reuse/consumption boundary completes. For epoch/CAS designs, prove either the reuse bind or mutation wins one serialization order; if mutation wins, compare-and-bind must fail and the call must re-evaluate/rebuild for B rather than consume A as a hit. Repeat A→B→A and require the monotonic witness to reject a stale endpoint-equal hit. Include multi-input vectors and concurrent generation publishers. + +Exercise an **output validation-to-consumption race**. Arrange a reuse hit for output A, validate A successfully, then have another writer replace the mutable path/object with B before the consumer reads. Prove the consumer still reads the pinned immutable/versioned A that was validated, or prove the lock prevents replacement until consumption completes. Repeat with delete/recreate, atomic rename, symlink/object-pointer replacement, version rollback, and multiple required outputs where one is swapped after validation. A test that merely corrupts output before validation is insufficient; the mutation must occur after the validity predicate succeeds and before/downstream consumption. + +Exercise **concurrent input mutation**, including explicit A→B→A races. Start execution from identity A, mutate one or more effective inputs during execution to B (including mixed-state multi-file/config changes), then restore the original bytes before publication. For snapshot-based targets, prove execution reads only the immutable A snapshot. For lock-based targets, prove the mutation cannot interleave with the protected execution/activation interval. For epoch/version-based targets, prove every mutation increments the monotonic witness and that the final epoch vector differs even when the final content signature returns to A. Reject/discard the candidate on any mutation witness change. Compare every accepted generation with a forced-fresh execution over the exact committed input identity. + +Add the **final-check-to-generation-switch race** explicitly. Pause after the implementation has obtained what would otherwise be its successful final input witness but before the authoritative manifest/generation pointer is switched. Attempt an input mutation in that exact interval. A lock-based implementation must block the mutation until after activation; an epoch/CAS implementation must make either the mutation or the generation switch win one serialization order, and if the mutation wins the compare-and-publish must fail rather than activate stale A as current. Repeat with multi-input epoch vectors and concurrent publishers. No test may accept a design where “check succeeded” and “manifest switched” are separate unprotected events. + +Exercise interruption/crash injection at every publication boundary: before outputs complete, after outputs complete but before final input-witness/activation transaction, during the atomic compare-and-publish or lock-protected generation switch, immediately before/after the switch, and during cleanup. After each interruption, prove readers observe only a self-consistent old or new generation and can never pair signature A with output identities/metadata from generation B or activate a generation whose input witness lost the activation race. + +## Target-repo adaptation + +Re-profile signature and output-validation cost, immutable-output handle/pinning cost, immutable-input snapshot or mutation-lock/epoch cost, **reuse-hit input-binding cost**, hash/version choice, metadata granularity, persistence format, and generation-publication mechanism. Include environment/toolchain inputs when they affect output. Define how every invocation captures one coherent input identity and how a reuse decision is atomically bound to the committed generation for that identity before mutable live inputs can invalidate the hit; explicitly choose snapshot-at-invocation semantics or live-through-consumption semantics and implement the corresponding snapshot/lock/compare-and-bind boundary. Define how a successful output-validity check returns/pins the exact version consumed downstream; if mutable storage is unavoidable, define the lock scope or immutable-copy boundary. Explicitly choose whether mutable inputs are consumed from immutable snapshots, protected by locks, or guarded by monotonic mutation epochs; define how every mutation advances the witness and how a coherent multi-input witness is captured. Define the **linearizable activation primitive**: either the mutation lock remains held through the authoritative generation switch or the system atomically compare-and-publishes against the complete witnessed epoch/version vector. Do not advertise commit-time content rehashing, an epoch read followed by a later manifest swap, a signature match followed by an unprotected reuse decision, or any other check-then-act sequence as sufficient mutation protection. Also define the crash-consistency guarantee for committing the signature plus output identities. + +## Failure modes + +Incomplete signatures create stale reuse; a signature/epoch match followed by an unprotected reuse decision can consume generation A after mutable live inputs have already advanced to B; input mutation during execution can bind an old signature to new/mixed output; A→B→A races can defeat endpoint signature comparisons; non-monotonic/reused mutation versions can erase intervening changes; incoherent per-input epoch reads can represent no real source state; a mutation that wins after the final witness read but before an unprotected manifest switch can make stale state authoritative; validating a mutable output name and reopening it later can consume different unvalidated bytes; output-version handles that are not actually immutable/pinned can create time-of-check/time-of-use reuse bugs; existence-only output checks can return corrupted artifacts; weak output-validity predicates can miss external mutation; independently persisted signature/output metadata can create cross-generation false hits after interruption; overly broad signatures erase the benefit; persistence corruption can create false hits; timestamp-only schemes may be unsuitable where timestamp semantics are weak. + +## Rollback trigger + +Disable reuse immediately if any signature/output-validity hit diverges from the forced-fresh reference; if a reuse hit can be accepted for identity A and then consume A after a concurrent mutation has made B authoritative without an explicit immutable snapshot-at-invocation binding; if a consumer can read bytes/objects different from the exact output version whose validity predicate passed; if an A→B→A or other mutable-input race can publish a generation or qualify a reuse hit without an immutable snapshot/lock/monotonic mutation witness proving one coherent effective-input identity; if an input mutation can linearize between an accepted reuse witness and reuse binding, or between the accepted final build witness and authoritative generation activation; if external output mutation can bypass the declared validity/consumption binding; if crash/interruption testing can expose mixed-generation or stale-current state; or if signature/snapshot/mutation-tracking/reuse-binding/output-pinning/integrity/activation/publication maintenance costs more than the avoided work. diff --git a/optimizations/OPT-POOL-001-persistent-topology-aware-worker-pools.md b/optimizations/OPT-POOL-001-persistent-topology-aware-worker-pools.md new file mode 100644 index 0000000..d1e036c --- /dev/null +++ b/optimizations/OPT-POOL-001-persistent-topology-aware-worker-pools.md @@ -0,0 +1,117 @@ +# OPT-POOL-001 — Persistent topology-aware worker pools + +**Status:** Implemented external reference; persistent reuse and topology-aware selection are merged in GALAXY, while scaling remains host- and workload-specific. +**Domains:** CPU batch runtimes, repeated simulation/render passes, parallel numerical pipelines, thread-pool execution + +## Source evidence + +- Repository: `QSOLKCB/GALAXY` +- PR: https://github.com/QSOLKCB/GALAXY/pull/13 +- Merge commit: `1966bc2595a402a2c653f2e392465224621e20fb` +- Source note: `sources/GALAXY-CPU.md` +- Licensing boundary: Apache-2.0 donor; mechanism promoted without requiring copied source. + +## Problem + +A validated parallel kernel is fast enough that repeatedly creating worker threads, allocating worker-local buffers and choosing an unsuitable logical/physical worker count become material overheads. Per-run spawning also adds latency variance and can hide whether SMT helps or hurts. + +## Optimization problem contract + +- X: Persistent-pool lifetime, worker count, topology policy, reusable worker-local buffer capacity and dispatch strategy. +- F: Candidates preserving exact output/checksum parity, deterministic work ownership and reduction, bounded live resources, explicit topology fallback, correct shutdown/error handling, and truthful concurrency evidence: a candidate claiming parallel execution must demonstrate observed overlapping active work under workload-shaped dispatches rather than merely reporting participating worker identities. +- f: Total or amortized runtime across the expected repetition horizon, including pool lifecycle cost where relevant, plus resource/scaling evidence and observed concurrency/overlap metrics for candidates whose performance claim depends on parallel execution. +- d: Minimize lifecycle-adjusted runtime while preserving deterministic semantics; prefer simpler scheduling when gains are negligible. +- C: Completion order must not alter observable results, topology claims must match detected evidence, persistent workers must not retain stale per-dispatch state, and requested/configured/effective worker counts must not be presented as proof of simultaneous execution without a direct or equivalent overlap measurement. +- B: Bounded worker/schedule/tile sweeps and repeated dispatches on the target execution environment. +- S: Stop when the expected repetition horizon and topology policy have a repeatable useful winner, or retain spawned/canonical execution when startup amortization is insufficient; if a purportedly parallel candidate shows no meaningful overlap, classify it as serialized for evidence purposes and do not promote a concurrency claim from worker participation alone. +- Variables: integer, categorical and conditional +- Search scope: local +- Objective behavior: noisy +- Information: black-box +- Evaluation cost: moderate +- Constraints: semantic and resource +- Parallelism: asynchronous +- Exactness: exact + +## Preserved contract + +Persistent reuse changes worker lifetime, not computation semantics. Each dispatch must process the same logical work as the reference/spawned path, and reduction must remain deterministic where required. Buffer reuse must reset or overwrite all state that can affect a later dispatch. + +A failed or cancelled dispatch may be followed by reuse only after every worker-local buffer, queue, completion flag and dispatch-generation marker is returned to a known clean state. If that reset cannot be proven complete, retire the pool. + +Retirement must use one of two explicit strategies: + +1. **Quiescent retirement:** do not start a successor pool until every worker from the retired pool has resumed or been cancelled, reached a terminal state and been joined/reaped. +2. **Generation-fenced retirement:** every externally visible publication is bound to a non-reused **pool-incarnation identity** plus the dispatch generation. Retirement and publication must linearize through the same lock, transaction, atomic compare-and-publish operation or equivalent primitive; a worker must not be able to validate its generation, pause across retirement, and then publish afterward. A replacement may begin before old workers terminate only when stale publications are rejected by that linearization rule. + +Generation-fenced retirement does not waive resource bounds. Retired pools must still have an eventual termination/reaping path, and the implementation must bound the number or total resource cost of unreaped retired generations. When that bound is reached, apply backpressure or fall back to quiescent retirement/spawned execution instead of creating unbounded replacement pools. + +Concurrency evidence is part of the claim boundary rather than the computation contract. A pool may involve multiple workers yet still execute effectively serially because of locks, queue policy, scheduler throttling, cgroup limits or runtime serialization. Such a path can still be semantically correct, but it must not be described as providing parallel execution unless overlapping active work is actually observed. + +## Optimization + +Create workers once, allocate their reusable local buffers once, and dispatch repeated jobs through the persistent pool. Give each worker a stable deterministic range or identity. Allow workers to finish independently, but collect/reduce results under a deterministic ordering rule when arithmetic or output order requires it. + +Expose topology policy explicitly. A `physical-first` policy may cap workers at detected physical cores; a `logical` policy may include SMT threads. Detection must fail softly and record the fallback instead of pretending unavailable topology data is authoritative. + +Treat dispatch completion as a state transition. Successful completion must leave all reusable state ready for the next generation. Failure or cancellation must either run the same complete reset protocol or retire the pool so partial state cannot leak into a later dispatch. + +For generation-fenced retirement, identify publications by `(pool-incarnation, dispatch-generation)`, not by a generation counter alone. The pool-incarnation identity must not be reused by a replacement pool even if a local generation counter restarts or wraps. Retirement and every externally visible publication must participate in the same linearization mechanism so there is no check-then-write interval in which a retired worker can pass validation before retirement and publish after it. Apply the rule to shared output, completion state, queues, callbacks and any other externally visible mutation. + +Track retired pools until all workers terminate and are reaped. Bound the number or retained-resource budget of generation-fenced retired pools; if the bound is exhausted, stop admitting replacements until retirement progresses or switch to a quiescent/spawned fallback. + +Separate **worker participation** from **simultaneous overlap**. Instrument workload-shaped dispatches with an active-worker counter, timestamped task intervals, scheduler/runtime tracing, or another measurement that can establish how much work actually overlapped. Record at least the observed peak simultaneous active work and, where useful, overlap duration/fraction or a concurrency histogram. A queue that eventually touches every worker but runs only one task at a time is not evidence of parallel execution. + +Separate steady-state dispatch timing from startup/teardown, then include lifecycle cost when deciding whether persistence is worthwhile for the real repetition horizon. + +## Before / after evidence + +- Environment: GALAXY PR #13 verifies Linux x86-64, Linux ARM64, macOS ARM64 and Windows x86-64 command/parity surfaces. +- Workload/fixture: repeated worker-local SoA executions over deterministic resident ranges. +- Cold baseline: spawned worker-local SoA creates worker threads for each complete execution. +- Warm/no-op baseline where relevant: persistent steady-state dispatch excludes startup but records pool startup separately. +- Small invalidation / partial-work case where relevant: repeated first/second dispatch parity verifies reused state does not leak. +- Quiescent-retirement cancellation fixture: pause an old worker after partial activity, cancel and retire its pool, then resume/cancel it to terminal state and join/reap it **before** starting the successor dispatch. Verify that successor admission is blocked until quiescence is complete and that no retired resources remain live afterward. +- Generation-fenced cancellation fixture: pause an old worker after it has validated its `(pool-incarnation, dispatch-generation)` but **before** the externally visible mutation. Retire the pool, start a successor pool with a distinct incarnation identity and deliberately collide/restart its local generation number, then resume the old worker. Verify that the publication linearization rejects the stale mutation despite the colliding generation number and that successor output, queues, completion state and callbacks remain unchanged. +- Retired-pool bound fixture: repeatedly cancel generation-fenced pools while holding old workers alive until the configured retired-generation/resource bound is reached. Verify that further replacement is backpressured or falls back rather than creating another pool, then release the held workers and verify eventual termination/reaping returns the retired count/resource budget to baseline. +- Large invalidation / full-work case where relevant: bounded production receipts exercise persistent dispatch across selected worker counts. +- Optimized: one persistent worker set and reusable tile buffers across warm-up and measured repetitions. +- Speedup / memory / I/O / quality change: donor establishes the mechanism and verification boundary but does not provide a universal scaling claim in the PR summary. +- Variance / repetitions / raw samples: target-specific receipts and repetitions are required before promotion. + +## Validation + +Require equality among canonical/reference output, spawned optimized output, first persistent dispatch and subsequent persistent dispatches. In addition to ordinary repeated-success cases, force success → failure → success and success → cancellation → success sequences after partial worker activity. Verify that every reusable buffer, queue, completion record and dispatch generation is reset before the final success. + +Validate the chosen retirement strategy with the matching fixture rather than one ambiguous ordering. For quiescent retirement, hold an old worker, prove successor admission remains blocked, then release/cancel and join/reap the worker before the successor begins. For generation-fenced retirement, pause an old worker specifically **after validation but before mutation**, retire its pool, start a successor with a distinct pool-incarnation identity and a deliberately colliding local generation value, then resume the worker and prove the atomic/locked publication step rejects it. Also stress repeated fenced retirements to the configured outstanding-retired-pool/resource bound, prove backpressure/fallback at the bound, and prove eventual worker termination/reaping releases all retired resources. Test shutdown, worker-count changes, topology fallback and completion-order independence. + +For every workload-shaped dispatch used to support a parallelism or scaling claim, instrument **observed simultaneous active work** or an equivalent overlap metric. Record requested workers, configured/effective workers, topology source, observed peak concurrent activity, and preferably overlap duration/fraction or a concurrency histogram. Verify the metric itself against a deliberately serialized control. If a queue, lock, runtime limit, scheduler policy or cgroup causes configured workers to take turns without overlapping, report the execution as serialized/limited rather than treating worker participation as concurrency evidence. Compare the overlap data with measured speedup so apparent scaling cannot be attributed to concurrency that never occurred. + +Measure startup and teardown separately, then evaluate amortized cost for the actual repetition horizon. + +## Target-repo adaptation + +Re-profile pool lifetime, worker count, SMT policy, buffer size, task granularity, expected number of dispatches, CPU allowance/cgroup constraints, failure-reset protocol, shutdown behavior, pool-incarnation generation strategy, publication-linearization primitive, maximum outstanding retired-pool/resource budget, reaping timeout/backpressure policy, and the concurrency-observation method. Do not infer CPU affinity or NUMA placement from topology-aware worker counting; those require separate mechanisms and evidence. If locks, queues, external libraries or runtime quotas can serialize the hot region, instrument overlap around the actual work rather than only around task submission. + +## Failure modes + +- The workload is too infrequent to amortize pool startup and retained resources. +- Reused buffers leak stale state between dispatches. +- A failed/cancelled dispatch leaves partial buffers, queue entries or completion state that contaminates the next generation. +- A worker from a retired generation resumes after replacement and publishes stale output, completion, queue or callback state into the succeeding dispatch. +- A check-then-write race lets a worker validate before retirement and publish after retirement because validation and mutation do not linearize atomically. +- A replacement reuses a generation value from a retired pool because the fence omits a distinct pool-incarnation identity. +- Repeated generation-fenced cancellations leave too many unreaped pools, threads or worker-local buffers alive and exhaust the bounded live-resource budget. +- SMT/logical workers increase contention or memory pressure. +- Container CPU allowance or topology changes after pool creation. +- Multiple workers participate but a lock, queue, runtime limit, scheduler or cgroup serializes the hot work, creating false concurrency evidence. +- Long-lived workers hold scarce memory/resources during idle periods. +- Async completion accidentally changes reduction/output order. + +## Rollback trigger + +Use spawned/canonical execution on any parity failure, stale-state leak, failed-dispatch reset failure, late retired-generation publication, non-linearizable publication fence, pool-incarnation reuse, retired-pool bound violation, shutdown/resource leak, topology mismatch, or lifecycle-adjusted slowdown for the target repetition horizon. Retire a pool immediately when a failure/cancellation leaves its reusable state uncertain. Start a replacement only after quiescence/join, or under a generation-fenced strategy whose `(pool-incarnation, dispatch-generation)` publication operation linearizes with retirement and remains within the configured outstanding-retired-resource bound. If reaping stalls at that bound, apply backpressure or fall back to spawned/canonical execution. Disable physical-first selection when topology detection is unreliable and record the fallback. If the optimization depends on parallel execution but workload-shaped measurements show no meaningful simultaneous overlap, withdraw the parallelism claim and re-profile or fall back rather than promoting the configured worker count as effective concurrency. + +## Composition notes + +Composes with `OPT-SOA-001` when workers own reusable local tiles and with `OPT-SIMD-001` inside each worker kernel. Re-measure with `OPT-PAR-001` because persistent worker counts can still oversubscribe libraries or nested parallel regions. \ No newline at end of file diff --git a/optimizations/OPT-PRUNE-001-bound-driven-search-space-pruning.md b/optimizations/OPT-PRUNE-001-bound-driven-search-space-pruning.md new file mode 100644 index 0000000..6b89627 --- /dev/null +++ b/optimizations/OPT-PRUNE-001-bound-driven-search-space-pruning.md @@ -0,0 +1,110 @@ +# OPT-PRUNE-001 — Bound-driven search-space pruning + +**Status:** Proposed / OPT synthesis; classical mechanism, target adaptation required +**Domains:** combinatorial optimization, scheduling, assignment, configuration search, resource allocation + +## Source evidence + +- `sources/MATHEMATICAL-OPTIMIZATION.md` +- combinatorial optimization / branch-and-bound literature referenced there +- NLopt taxonomy as supporting optimizer-selection context + +## Problem + +A discrete or mixed search space is too large for exhaustive evaluation, but whole subregions can sometimes be proven unable to beat the best feasible solution already found. + +## Optimization problem contract + +- X: the target's explicitly defined discrete or mixed candidate space together with a partition of unexplored candidates into searchable subregions +- F: candidates in X satisfying every original hard constraint; relaxed/bounding solutions are not feasible final answers unless they also lie in F +- f: a scalar real-valued target objective `f : F → R` evaluated on feasible candidates only +- d: exactly one of scalar `minimize` or scalar `maximize`; vector, Pareto, lexicographic, or other partial-order objectives are outside this record unless a separately specified and validated frontier-bound mechanism is introduced +- C: every returned incumbent satisfies the original feasibility/semantic contract, every pruning decision is justified by a separately defined sound scalar region-bound function `b`, **exact pruning is authorized only when the soundness argument for `b` covers the deployed search domain through an analytic/formal proof, exhaustive verification of the complete finite deployed domain, or a conservative runtime proof/certificate checked for each pruned region**, the target's observable tie semantics are preserved, parallel dispatch cannot oversubscribe the declared hard budget, every unit of resource consumption covered by a hard wall-time/compute B is accounted for or enclosed by an enforceable whole-search cap, every hard peak-memory B is enforced over live allocated/reserved memory rather than cumulative historical allocation, frontier exhaustion is declared only after all queued **and leased/in-flight** regions are accounted for, and **when deterministic budget-limited output is part of C, worker timing may not decide which logical regions receive the final budget entitlement or which completed results become the returned anytime state** +- B: a finite, predeclared target-specific **enforceable** cap with its accounting semantics declared explicitly. Evaluation count, money/provider spend, CPU/GPU-seconds, energy, bytes transferred, or other cumulative-flow resources use cumulative accounting. Elapsed wall time uses one shared whole-search deadline. **Peak memory is a stock constraint, not a cumulative flow:** enforce `live_allocated + live_reserved + proposed <= B`, release live capacity when memory is freed, and retain only a recorded `peak_observed` for evidence. If a target instead wants cumulative allocation traffic, it must declare that as a distinct cumulative metric rather than calling it peak memory. Any resource dimension that cannot be hard-capped under its declared semantics must be labeled observational/best-effort rather than advertised as hard B. If deterministic anytime output is required, also declare the deterministic logical work/budget boundary used to select the returned result; elapsed wall time alone is not a deterministic selection boundary under variable worker timing +- S: stop immediately when the required optimality/tie contract is proven, or when the **global frontier is exhausted**, meaning there are no queued regions, no leased/in-flight regions still capable of producing candidates/children, and no unpublished child/frontier updates owned by active work. Otherwise stop when B is exhausted. If a validated incumbent exists, return it plus any remaining valid global bound/optimality gap. If no feasible incumbent exists, return `no-incumbent / feasibility-unknown` and only a separately valid global bound if one is available; do not report an optimality gap that requires an incumbent, and do not claim infeasibility or optimality. **For deterministic targets, budget exhaustion must expose the result of a canonical logical prefix/batch of search work rather than an arbitrary prefix induced by completion order; if the only stopping boundary is physical elapsed time, budget-limited anytime results must be explicitly declared nondeterministic or selected from a separately declared deterministic checkpoint that was committed before the deadline** +- Variables: integer / categorical / discrete / mixed +- Search scope: global over the declared candidate space +- Objective behavior: deterministic unless uncertainty/noise is incorporated into a separately sound bound model +- Information: derivative-free; bound/relaxation information is target-specific +- Evaluation cost: moderate to expensive when exhaustive evaluation is infeasible +- Constraints: feasibility, semantic correctness, **deployed-domain scalar-bound soundness**, tie semantics, global-frontier accounting, deterministic logical budget-boundary semantics when required, and enforceable finite-resource constraints including correctly typed cumulative, deadline, and peak/live-capacity budgets +- Parallelism: sequential, or parallel only with synchronized incumbent/frontier/bound state, **leased/in-flight region accounting**, linearizable reservation/completion accounting for cumulative resources, a shared absolute deadline for elapsed wall time, live-allocation reservation accounting for hard peak-memory caps, and—when deterministic output is required—**stable region/work IDs plus a deterministic leasing/reservation/assimilation boundary whose logical order is independent of worker completion timing** +- Exactness: exact only when the declared optimality and observable-tie contract is proven within B **and the pruning bound's soundness is justified over the complete deployed domain or certified conservatively at runtime for every pruned region**, including proof that no queued or leased region can still affect the answer; otherwise anytime/incomplete result semantics apply + +For each unexplored region `R`, define a bound `b(R)` separately from `f`: + +- minimizing: `b(R) ≤ inf { f(x) | x ∈ F ∩ R }`; +- maximizing: `b(R) ≥ sup { f(x) | x ∈ F ∩ R }`. + +**The inequality above is a proof obligation, not merely a test expectation.** Before using `b(R)` to make an exact pruning decision in a deployed search, establish one target-specific soundness basis that covers the actual deployed domain: (1) an analytic or formal derivation proving the bound relation for every admissible region/candidate under the target assumptions; (2) exhaustive verification over the complete declared finite deployed domain and every region shape the implementation may prune; or (3) a conservative runtime proof/certificate/guard whose premises are checked before each prune and which falls back to retaining/evaluating the region whenever the certificate cannot be established. Small exhaustive fixtures, randomized tests, and adversarial examples remain required regression evidence, but **they are not by themselves certification of global bound soundness** when the deployed domain is larger or non-finite. Without one of these deployed-domain soundness bases, `b(R)` may prioritize search order only and exact pruning must be disabled or explicitly treated as heuristic/incomplete. + +If the target contract accepts **any one scalar optimum** and equal-objective candidates are not observably distinct, minimization may prune `R` when `b(R) >= f(x_incumbent)` and maximization may prune when `b(R) <= f(x_incumbent)`. + +If equal-objective candidates remain observable—for example the target requires a deterministic tie winner, a secondary total ordering, or enumeration of all optimal candidates—equality is not enough to discard a region under the scalar bound alone. In that case either: + +- use strict objective pruning while unresolved ties remain (`b(R) > f(x_incumbent)` for minimization; `b(R) < f(x_incumbent)` for maximization), and continue exploring equality-bound regions as required by C; or +- define a separately sound bound over the **complete declared tie ordering/frontier** and validate that stronger bound independently. + +An independently proven infeasible region may also be pruned. A heuristic estimate that does not satisfy the declared bound relation is search-ordering evidence at most, not a pruning proof. This record does not authorize scalar bounds to prune vector/Pareto or partially ordered objectives. + +## Preserved contract + +A region may be discarded only when its sound bound proves it cannot contain any candidate that remains observably preferable or required under the target's scalar objective **and tie contract**, and the mechanism used to justify that bound is valid over the deployed domain/region being pruned. Passing a finite regression suite does not turn a heuristic estimate into a proof bound. Exhausting B without an optimality proof does not permit an exactness claim, exhausting B without a feasible incumbent does not permit an infeasibility claim, and parallel execution must preserve the same hard resource ceiling as sequential execution rather than oversubscribing work in flight. A temporarily empty shared queue is **not** frontier exhaustion while any worker owns a leased region that may still produce a candidate, proof obligation, or child region. Likewise, a hard wall-time/compute B applies to the whole search, while a hard peak-memory B applies to current live/reserved memory and must not be converted into irreversible historical consumption after memory is freed. + +When the target requires deterministic output, that requirement also applies to an **incomplete/anytime result produced at B**. The parallel implementation must therefore make the logical search prefix independent of worker speed: the same input, seed, contract and logical budget must commit the same ordered work set, incumbent/tie winner, valid global bound/gap and stop reason as the declared canonical reference. Physical completion may occur out of order, but it cannot grant a later region the last logical budget entitlement merely because that worker finished first. If deterministic budget-limited output is not required, the target must state that nondeterminism explicitly instead of inheriting a deterministic objective classification by implication. + +## Optimization + +Maintain an incumbent when one exists, partition the search space, compute a cheap **sound and deployed-domain-justified** `b(R)` for each region (often from a relaxation), prioritize promising regions, and prune only when the direction-specific bound plus the target's tie semantics prove the region cannot affect the required answer. Before the first incumbent exists, sound bounds may prioritize regions or prove individual regions infeasible, but incumbent-based objective pruning is unavailable. + +For **parallel** search, define one global frontier lifecycle. A region remains part of the frontier from enqueue until it is either (a) soundly pruned/closed, or (b) replaced by its child regions through an atomic/linearizable completion transition. Dequeuing for worker ownership therefore changes a region from `queued` to `leased/in-flight`; it does **not** remove that region from the global frontier. A worker that branches a leased region must publish all resulting children and close/release the parent as one frontier-accounting transition, or use another protocol that cannot expose a moment where the queue is empty even though unpublished descendants still exist. Worker failure/cancellation must return or recover the lease so unexplored work is not silently lost. + +For targets requiring **deterministic budget-limited results**, give every frontier item/work operation a stable logical ID and define a canonical total leasing/commit order, including deterministic tie-breaks. Reserve logical budget entitlement in that canonical order **before dispatch**, or dispatch fixed deterministic batches whose membership is independent of completion timing. Results may finish physically out of order, but logical incumbent/bound/frontier assimilation must occur through the same canonical ordered prefix (or an equivalent deterministic batch barrier). A later completed result is buffered or remains speculative until every earlier entitled item needed for the prefix has been resolved; a free worker does not opportunistically admit a new logical item if doing so would change which work fits inside B. For variable-cost cumulative resources, conservative enforceable reservations participate in the canonical entitlement decision, so the set of admitted work is determined by the declared ordering and reservation amounts rather than by which prior worker happened to finish first. + +A hard **elapsed wall-time** deadline is different: scheduler and worker speed determine which physical operations finish before the clock expires. If the target nevertheless requires deterministic anytime output, use the deadline only as a safety envelope and select the public result from the most recent fully committed **deterministic logical checkpoint/prefix** defined by a separate logical work budget or batch schedule. Otherwise explicitly classify the wall-time-cutoff anytime result as nondeterministic while retaining deterministic exact/full-exhaustion semantics where those are otherwise required. Do not claim scalar/parallel budget-exhaustion equivalence when the public cutoff itself is completion-time-driven. + +Treat hard resources according to their physical/accounting semantics rather than forcing them through one ledger shape: + +- **Cumulative-flow budgets** such as evaluation count, money/provider spend, CPU/GPU-seconds, energy, or declared cumulative bytes use linearizable reservations. Atomically reserve before covered work starts; if `consumed + reserved + proposed_reservation > B`, do not start it. Completion/failure/cancellation moves actual consumed usage into permanent `consumed` and releases only demonstrably unconsumed reservation. +- **Elapsed wall time** uses one enforceable monotonic whole-search deadline shared by workers and coordinator paths. Parallel overlap is not double-counted as if seconds were additive reservations, and retries/workers cannot reset or extend the deadline. +- **Peak/live-capacity resources** such as peak memory use live state rather than permanent consumption. Atomically reserve enough live capacity before an allocation-producing operation starts; require `live_allocated + live_reserved + proposed <= B`; on successful allocation move reserved capacity to `live_allocated`; on free/reclamation release the live allocation so later non-overlapping work may reuse the capacity. Maintain `peak_observed = max(peak_observed, live_allocated + live_reserved)` for audit/evidence. Freed memory must not remain in cumulative `consumed` unless the separately declared metric is cumulative allocation traffic rather than peak memory. + +The relevant hard cap must cover all search work that can consume that resource, including candidate/bound evaluation, branching, child construction, queue/frontier operations, serialization, incumbent updates, synchronization and cleanup. Per-evaluation controls do not by themselves prove a whole-search wall-time, compute, spend, or peak-memory bound. + +If the target meters only candidate/bound evaluations, then only **evaluation count** may be claimed as a hard B from that accounting. Nominal wall-time/compute/memory targets in that design are observational/best-effort and cannot justify finite-cap correctness claims. Similarly, if branching/frontier/serialization overhead can escape an otherwise claimed resource cap, that resource dimension is not hard-bounded. + +A relaxed solution is evidence for a bound, not automatically a feasible final answer. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: No target exhaustive or unpruned search baseline has been established here. +- Optimized: No target branch-and-bound/pruned search result has been established here. +- Speedup / memory reduction: No transferable claim; this record captures a classical mechanism and adaptation rules. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +For small fixtures, compare with exhaustive enumeration and test `b(R)` independently by checking the direction-specific inequality against exhaustive feasible values inside each test region. These small-case tests are **regression/adversarial evidence only** unless they exhaust the complete deployed finite domain. Before enabling exact pruning on a larger or non-finite deployed domain, additionally verify the declared soundness basis: inspect/check the analytic or formal derivation and its assumptions; or exhaustively enumerate the complete finite deployed domain and every supported region construction; or exercise the conservative runtime certificate/guard and prove that every actual prune is accompanied by a valid certificate and that certificate failure retains/evaluates the region rather than pruning it. Deliberately inject a defective bound that passes the small fixtures but violates one held-out/deployed region and prove exact mode rejects it or the runtime guard refuses that prune. Test pruning separately from search ordering. Include fixtures where the first feasible candidate is found late and where B expires before any feasible candidate exists; verify that the latter returns `no-incumbent / feasibility-unknown`, reports only independently valid global-bound information, and makes no infeasibility, optimality, or incumbent-based gap claim. Verify that budget exhaustion with an incumbent returns an anytime result without an exactness claim. + +Add **parallel frontier-exhaustion races**. Use a fixture where the last queued region is leased by one worker, making the shared queue empty, then pause that worker before it publishes one or more child regions. Prove the coordinator does not declare exhaustion or exact optimality while that lease remains live. Resume the worker and verify the children become searchable and the final result matches exhaustive/scalar search. Also inject worker failure/cancellation while holding the last lease and verify the region is recovered/requeued or otherwise completed without losing unexplored work. Test simultaneous parent-close/child-publish transitions and prove there is no observation in which both queued and leased frontier counts reach zero before all descendants are durably accounted for. + +Add **parallel budget-boundary fixtures**. Race multiple workers against one remaining evaluation slot and prove only one reservation succeeds for an evaluation-count B. Race bound evaluations and candidate evaluations against the same final cumulative capacity and prove both charge the declared ledger. For elapsed wall time, run overlapping workers to the same absolute deadline and prove overlap is not double-counted while no worker/coordinator survives past the enforced deadline. For cumulative compute/spend, deliberately make branching/frontier/serialization work expensive and prove it is charged before the cap is exceeded. + +For a target that declares **deterministic anytime output**, repeat budget-exhaustion fixtures under deliberately permuted worker speeds, completion orders, pauses and wakeups. Hold the same input, seed, logical B and worker-count policy constant. Prove the same stable region/work IDs receive logical budget entitlement in the same canonical order or deterministic batches, the same logical prefix is assimilated, and the returned incumbent/tie winner, valid bound/gap and stop reason match the scalar/canonical reference. Include the exact last-slot race where a later canonical region finishes before an earlier one: the later completion must not steal the final logical entitlement or become publicly assimilated ahead of the declared prefix. For a hard elapsed deadline, verify either (a) the public result comes from the same last fully committed deterministic checkpoint despite completion-order changes, or (b) the target explicitly declares wall-time-cutoff anytime results nondeterministic and tests only the deterministic guarantees it actually claims. + +Add a **peak-memory reuse fixture**. Under a hard peak-memory B, run many sequential/non-overlapping branches that each allocate and then free a large work buffer. Prove each live allocation/reservation is admitted only while `live_allocated + live_reserved <= B`, freed capacity becomes reusable, `peak_observed <= B`, and the search does **not** exhaust merely because the sum of historical allocations exceeds B. Then overlap enough workers to exceed the peak if all allocations were admitted and prove the final reservation is rejected/blocked before live memory can cross B. Race allocation, free, cancellation and cleanup to verify live-memory accounting remains linearizable and no capacity is released before the corresponding memory is actually reclaimable. + +Add **equal-objective tie fixtures**. For an any-one-optimum contract, prove equality pruning cannot alter any observable result. For deterministic tie-winner contracts, construct regions containing equal-objective candidates with better/worse tie ranks and prove equality-bound regions are retained until the declared tie winner is established. For all-optima contracts, prove every equal-objective optimum is enumerated. If using a stronger total-order bound, validate its soundness independently against exhaustive fixtures and establish the same deployed-domain proof basis before using it for exact pruning. + +## Target-repo adaptation + +The quality/cost of bounds determines whether pruning helps. Develop target-specific scalar relaxations, branch ordering, feasible-candidate discovery strategy, **tie/secondary-order semantics**, and a finite resource cap before execution; do not assume one bound or budget is universally appropriate. **Document the deployed-domain soundness basis for every bound used to prune in exact mode:** identify the analytic/formal theorem and assumptions, the complete finite domain exhausted by verification, or the runtime certificate/guard and its conservative fallback semantics. Treat small fixture results as regression evidence, not as the proof basis for a larger domain. For parallel implementations, define the global frontier state machine, lease ownership/recovery rules, parent-close/child-publish atomicity, and the exact exhaustion predicate over queued plus leased/in-flight work. Also define each budget's accounting type and enforcement boundary: cumulative flow (`consumed + reserved`), elapsed deadline, or peak/live capacity (`live_allocated + live_reserved`, plus `peak_observed`). If a target says “memory budget,” state whether it means peak live memory or cumulative allocation traffic. **Declare whether budget-limited anytime output is required to be deterministic.** If it is, specify stable region/work IDs, the canonical lease/reservation and assimilation order (or deterministic batches), the scalar/canonical reference used for equivalence, and—when elapsed wall time is a hard safety envelope—the separate deterministic logical checkpoint/budget from which the public result is selected. If wall-time-cutoff anytime results are intentionally nondeterministic, state that contract change explicitly. Downgrade any dimension that can escape its correct enforcement boundary to best-effort/observational rather than calling it hard B. + +## Failure modes + +Unsound bounds can remove the true optimum; **small exhaustive fixtures can all pass while a bound remains unsound elsewhere in a larger deployed domain**; a runtime certificate that is not conservative or whose failure path still prunes destroys exactness; weak bounds provide little pruning; expensive bounds can cost more than evaluation; numeric tolerance errors can create incorrect pruning; heuristic scores mislabeled as bounds invalidate the proof obligation; equality pruning can discard a required deterministic tie winner or additional optimum; applying scalar pruning logic to vector/Pareto objectives can discard nondominated candidates; treating queue-empty as frontier-empty can declare exact completion while a leased region still owns unexplored descendants; losing a worker lease can silently drop search regions; non-atomic parent-close/child-publication can create false exhaustion; parallel workers without linearizable evaluation reservations can oversubscribe the last evaluation slot; **worker completion timing can choose the last admitted region or anytime incumbent even though C requires deterministic output**; completion-driven assimilation can make the budget-limited parallel prefix differ from the scalar/canonical prefix; using elapsed wall time itself as a deterministic result-selection boundary can make output scheduler-dependent; branching/child/frontier/serialization/incumbent overhead can exceed a nominal wall-time/compute B if only evaluations are charged; an unenforced coordinator/cleanup path can outlive a claimed whole-search deadline; treating peak memory as cumulative consumed spend can falsely exhaust a valid search and prevent reuse of freed capacity; releasing live-memory capacity before actual reclamation can instead oversubscribe the peak; treating an unenforceable resource target as hard B makes the stopping contract false; treating budget exhaustion without an incumbent as evidence of infeasibility is unsound. + +## Rollback trigger + +Disable any pruning rule that lacks a valid deployed-domain soundness basis for exact mode, whose analytic/formal assumptions fail, whose exhaustive finite-domain verification no longer covers the deployed domain/region construction, whose runtime certificate can authorize an unsound prune or fails open, that fails exhaustive small-case validation, violates the declared scalar/tie-bound relation, is applied to an unsupported objective ordering, discards an equal-objective candidate required by C, or whose bound cost exceeds the work it eliminates. Abort parallel/exact mode if frontier exhaustion can be observed while any leased/in-flight region may still produce work, if parent-close/child-publication or lease recovery can lose unexplored regions, if workers can oversubscribe an evaluation-count/cumulative budget, if an elapsed deadline can be reset/escaped, if peak live memory can exceed B, if freed peak-memory capacity is incorrectly made permanently unavailable, or if any resource-consumption path can escape a dimension advertised as hard. **Abort deterministic parallel anytime mode if worker timing can change logical reservation/lease entitlement, assimilation order, the committed budget prefix, returned incumbent/tie winner, valid bound/gap, or stop reason relative to the declared canonical reference.** Abort exact-mode claims whenever B is exhausted before the full objective/tie/frontier contract is proven, and reject any implementation that converts a no-incumbent budget timeout into an infeasibility or optimality claim without a separate proof. diff --git a/optimizations/OPT-REDUCE-001-early-working-set-reduction.md b/optimizations/OPT-REDUCE-001-early-working-set-reduction.md new file mode 100644 index 0000000..4be8dca --- /dev/null +++ b/optimizations/OPT-REDUCE-001-early-working-set-reduction.md @@ -0,0 +1,67 @@ +# OPT-REDUCE-001 — Early working-set reduction + +**Status:** Implemented external pattern; broadly applicable mechanism +**Domains:** databases, graphs, simulation, DSP, rendering, data pipelines + +## Source evidence + +- https://jazco.dev/2023/08/10/query-optimization/ +- supporting sparse-evaluation pattern in `OPT-DSP-001` + +## Problem + +An expensive operation is applied to a large population even though only a small subset can affect the final result. + +## Optimization problem contract + +- X: semantically legal placements and implementations of filtering, culling, limiting, candidate selection, or other working-set reductions in the target pipeline, including any target-specific predicate rewrite required to move a reduction across a stage +- F: placements for which every reduction predicate is proven to commute with every crossed stage or is replaced by a semantics-preserving pre-stage predicate, while also preserving every candidate and every contractually observable behavior required by the reference pipeline—including output, ordering/tie/join semantics, errors/exceptions, writes, mutations, auditing/telemetry, and other side effects. Purity of a crossed stage proves only that displaced side effects are absent; it does not by itself prove that moving the predicate preserves values or membership +- f: measured end-to-end pipeline cost and cardinality presented to the expensive stage +- d: minimize under the target's predeclared objective ordering +- C: the reordered/reduced pipeline is semantically equivalent to the reference for all declared outputs **and observable effects**; for every crossed transformation `T` and post-stage predicate `p`, the early form must either use an equivalent predicate `p'` satisfying the target's declared commutation/rewrite law (for example `p(T(x)) = p'(x)` for every relevant `x`) or otherwise prove equivalent candidate membership and downstream semantics; an effectful stage may be bypassed for discarded candidates only when those effects/errors are explicitly proven irrelevant by the target contract +- B: target-specific benchmark budget over representative and adversarial selectivity distributions; no portable selectivity threshold is supplied here +- S: stop when the declared budget is exhausted or a validated early-reduction placement materially lowers total cost without violating C +- Variables: categorical / conditional / mixed placement and predicate choices +- Search scope: local pipeline-reordering / working-set-reduction decisions +- Objective behavior: noisy for performance; semantic equivalence is deterministic +- Information: derivative-free / black-box performance measurements +- Evaluation cost: moderate to expensive depending on downstream stage cost and workload size +- Constraints: predicate-commutation/rewrite equivalence, output, ordering/tie/join, side-effect/error, purity, and resource constraints +- Parallelism: sequential pipeline semantics with target-specific parallel execution only where equivalence remains valid +- Exactness: exact observable semantics; no approximation is introduced + +## Preserved contract + +Moving a reduction earlier is valid only when the reduction predicate remains semantically equivalent across **every crossed stage**. Purity is necessary only to establish that moving past the stage does not displace observable effects; purity alone is not sufficient to justify the reorder. For a transformation `T` followed by predicate `p`, the early placement must either prove that `p` commutes with `T` or use a correctly rewritten pre-stage predicate `p'` with an invariant such as `p(T(x)) = p'(x)` for every relevant input, together with preservation of any downstream values/order required after `T`. The full observable contract also includes final values, ordering/top-k/tie/join semantics, exceptions/error checks, writes/mutations, audit events, metrics, and other externally visible side effects where they are significant. + +## Optimization + +Push selective operations toward the input boundary only when doing so is semantics-preserving. Before crossing a stage, derive and validate the predicate relation for that stage: retain the same predicate only when it provably commutes, otherwise rewrite it to an equivalent pre-stage predicate, or do not push it across the stage. For example, a post-transform filter `value > 10` cannot be naively moved before a pure `value = input * 2` transform; over a compatible numeric domain it would require the proven rewrite `input > 5` (with boundary, overflow, NaN, rounding, and type semantics handled according to the target contract). Then filter before a pure expensive join/transform, cull before pure rendering work, select candidate roots before pure DSP calculation, or eliminate simulations proven unable to affect any required result/effect only when that relation is established. + +If the expensive stage is effectful, either keep the effectful portion on every candidate that would have reached it in the reference path, split the stage into a pure expensive computation and a required effect layer, or prove those skipped effects/errors are outside the declared contract. Do not optimize away observable behavior merely because the final data rows match. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: Not established in a target repository. +- Optimized: Not established in a target repository. +- Speedup / memory reduction: No transferable claim; external query observations and existing sparse-evaluation patterns are source evidence only. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Differential-test reordered pipelines against the reference, with emphasis on ties, null/missing values, boundary ordering and rare candidates. For **each crossed stage**, validate the declared commutation law or rewritten predicate over representative, boundary, randomized, and adversarial inputs; compare both candidate membership and final downstream values. Include a negative fixture where `T(x) = 2*x` and the reference applies `value > 10`: prove that naively applying `input > 10` before `T` is rejected because it drops values such as `x = 6`, and prove that any proposed `input > 5` rewrite is accepted only for a domain whose overflow, numeric, and boundary semantics make the equivalence valid. + +Also compare observable side effects and error behavior: writes/mutations, audit/log/metric events, callbacks, exception/error surfaces and their relevant ordering/counts. Include a deliberately effectful fixture to prove the optimization is rejected or preserves the effects, and a pure-but-noncommuting fixture to prove purity alone never authorizes predicate motion. + +## Target-repo adaptation + +Measure selectivity and reduction cost. For every candidate reorder, enumerate the crossed stages, classify each stage as pure or effectful, and record the predicate-commutation proof or explicit predicate rewrite required for that stage. Inventory contractually significant side effects/errors and define how each is preserved. Do not infer predicate mobility from purity alone. A cheap filter with low selectivity may simply add another pass. + +## Failure modes + +Illegal predicate reordering, assuming purity implies predicate commutation, an incorrect or domain-incomplete predicate rewrite, changed top-k semantics, underestimated filtering cost, loss of vectorization, duplicated scans, skipped writes/audit events/mutations, changed exceptions or validation failures, and reordered side effects can all make an apparently equivalent final result semantically wrong. + +## Rollback trigger + +Immediately revert if any crossed stage lacks a valid commutation/rewrite proof, if differential testing finds different candidate membership or downstream values, or if any output, ordering, error, or contractually significant side effect differs from the reference path. Also revert if total measured cost does not fall on representative workloads. diff --git a/optimizations/OPT-SEARCH-001-budget-aware-adaptive-search.md b/optimizations/OPT-SEARCH-001-budget-aware-adaptive-search.md new file mode 100644 index 0000000..1ec474c --- /dev/null +++ b/optimizations/OPT-SEARCH-001-budget-aware-adaptive-search.md @@ -0,0 +1,85 @@ +# OPT-SEARCH-001 — Budget-aware adaptive parameter search + +**Status:** Proposed / OPT synthesis; upstream mechanisms are implemented externally +**Domains:** expensive black-box tuning, CI/runtime parameters, simulation, numerical kernels + +## Source evidence + +- `bayesian-optimization/BayesianOptimization` inspected at `af8b928212f0eacd1ce20c20be72c1a7b1d8d421` +- `hyperopt/hyperopt` inspected at `9834314879c09c13e0b8e93eb678408ba46441a8` +- `stevengj/nlopt` inspected at `6e6593f131ba3a38bc9edbed0a357bc01526e54b` +- `sources/OPTIMIZATION-LIBRARIES.md` + +## Problem + +Optimization knobs are selected by folklore, exhaustive sweeps, or a few arbitrary values even when each benchmark evaluation is expensive. + +## Optimization problem contract + +- X: the target's explicitly bounded continuous, integer, categorical, conditional, or mixed parameter search space +- F: candidates in X that satisfy all hard resource, platform, semantic, and correctness constraints before objective ranking +- f: the target-measured objective or objective vector for each feasible candidate, including declared noise/statistical treatment +- d: the target's predeclared minimize, maximize, lexicographic, or Pareto ordering +- C: search may choose where to evaluate but may not weaken correctness, evidence, API, trust, or other target semantics to improve f; asynchronous dispatch must preserve the declared budget model under concurrency; additive resources such as evaluation count, compute, and spend use linearizable consumed/reserved accounting with enforceable caps, **and any hard compute-unit budget advertised for the search must also cover coordinator/search-control work that consumes that compute resource—including surrogate fitting, acquisition/proposal generation, model updates, frontier/portfolio coordination, assimilation, and stopping logic—through the same ledger or an enforceable whole-search compute quota**; elapsed wall-time budgets use one enforceable absolute search deadline shared by every worker; and targets that require deterministic search outcomes must use deterministic observation assimilation **and deterministic proposal/dispatch/refill scheduling** independent of wall-clock completion order +- B: an explicit target-specific hard maximum declared before the search starts together with its **budget semantics**: additive resources (for example evaluation count, compute units, or money) use conservative enforceable reservations from one shared ledger, but a hard compute-unit B is valid only if **all search work consuming the bounded compute resource, including coordinator/model/proposal work outside trials, is charged to that ledger or enclosed by one enforceable whole-search compute quota**; elapsed wall time uses one absolute monotonic search deadline that bounds the whole concurrent search rather than summing overlapping worker seconds; the accounting unit, enforcement mechanism, covered work, atomic boundary, and failure/cancellation charging policy are fixed before dispatch begins +- S: stop proposing/dispatching when no additional work is admissible under B, when the absolute wall-time deadline has arrived, when a predeclared objective/quality target is met, or when a predeclared stagnation/convergence rule fires; preserve the reason for stopping in the trial/search ledger and apply proposal, dispatch/refill, assimilation, and stopping decisions to the declared deterministic schedule when determinism is required +- Variables: mixed search spaces; may include continuous, integer, categorical, and conditional dimensions as explicitly declared by the target +- Search scope: local or global, explicitly declared for the target +- Objective behavior: deterministic, noisy, or stochastic as declared by the target; noise treatment must be explicit +- Information: derivative-free / black-box by default; gradient information may be used only when the selected target mechanism supports it +- Evaluation cost: typically expensive +- Constraints: bounds, semantic correctness, resource, platform, and target-specific equality/inequality constraints +- Parallelism: sequential / synchronous batch / asynchronous, explicitly declared +- Exactness: target evaluations must satisfy C exactly; the search itself need not prove a global optimum unless the target contract requires it + +## Preserved contract + +Search may choose *where to evaluate* but may not weaken correctness constraints to improve the objective. Under asynchronous execution, the declared maximum budget remains a hard bound according to its declared semantics. For **additive** resources, actual consumed resources plus all still-reserved in-flight capacity must remain within B, no covered operation may consume beyond its enforceable reservation/quota, and concurrent dispatch/completion transitions must not transiently expose phantom free capacity. If the additive resource is **compute**, the bounded search includes not only trial execution but every coordinator/search-control operation that consumes that compute resource; per-trial reservations alone do not prove a hard whole-search compute cap. For **elapsed wall time**, all workers share one absolute search deadline; overlapping trials do not consume duplicate elapsed seconds, but no proposal, trial, retry, assimilation step, or cleanup that is part of the bounded search may continue past the enforceable deadline except target-declared bounded termination cleanup. If the target requires deterministic selected configurations or trial traces, **both the observation prefix used to create each proposal and the schedule that decides when a new proposal may be generated/dispatched must be deterministic**; worker completion timing may not change the proposal sequence. + +## Optimization + +Use observations to adapt future evaluations: surrogate/acquisition search for expensive black-box objectives, conditional spaces where parameters only exist under certain choices, progressive domain contraction where justified, and explicit stopping/evaluation budgets. For asynchronous workers, reserve pending regions or otherwise diversify proposals so workers do not redundantly evaluate the same neighborhood. + +First classify each hard budget dimension. **Additive budgets**—for example evaluation slots, billable compute, accelerator-seconds, or monetary spend—use one atomic/serializable accounting ledger. Before dispatching a trial against an additive budget, reserve a conservative amount and record the pending trial. If `consumed + reserved + proposed_reservation > B`, do not dispatch. The reservation must be an **enforceable upper limit** for that trial, not merely an estimate: use an evaluation-slot token, provider spending cap, cgroup/job compute quota, or another mechanism that prevents actual additive consumption from exceeding the reservation. If the target cannot enforce such a cap for an additive resource dimension, that dimension cannot be advertised as a hard maximum B; define a different enforceable budget or classify the quantity as observational. + +A **hard compute-unit budget for the search is a whole-search resource contract**, not merely a sum of trial quotas. Any surrogate/model fit, acquisition optimization, candidate/proposal generation, portfolio/frontier update, observation assimilation, scheduler/coordination step, stopping-rule computation, serialization, or other coordinator work that consumes the bounded compute unit must be inside the same enforcement boundary. Two acceptable designs are: (1) an enforceable whole-search process/job/cgroup/provider compute quota that contains both workers and coordinator, or (2) complete shared-ledger accounting where coordinator operations reserve/charge compute before execution just as trials do, with no unmetered control path. If coordinator compute cannot be bounded or completely metered under the claimed unit, do not call that dimension a hard search B; scope the hard claim more narrowly (for example trial accelerator-seconds only) and label coordinator/total compute observational. + +An **elapsed wall-time budget is different**. At search start, compute one absolute deadline from a monotonic clock and make every worker, trial, retry, proposal, model update, and stopping decision subordinate to that same deadline. Do not add overlapping worker durations into `consumed + reserved`; two trials that run concurrently until the same ten-minute deadline consume at most ten minutes of elapsed search time, not twenty. A trial-specific timeout may be shorter, but never later than the remaining global deadline. Dispatch must stop when insufficient time remains for the target's declared safe launch/termination policy, and the runtime must be able to cancel/terminate in-flight work at the global deadline if wall time is claimed as hard. + +Completion, failure, cancellation, and forced termination for **additive** resources use the same atomic accounting boundary as dispatch reservation. For one terminal trial transition, atomically: (1) read the trial's reservation, (2) meter/record the amount actually consumed, (3) move that consumed amount into permanent `consumed`, (4) release only the demonstrably unconsumed remainder from `reserved`, and (5) mark the trial terminal. Coordinator/search-control compute charged through the ledger follows the same principle: reserve or otherwise atomically debit the bounded compute before covered work begins, then commit actual consumption and release only provably unused capacity. No dispatcher or coordinator may observe released capacity before the corresponding consumed charge is committed, and concurrent updates must not lose increments. Completion must not double-charge the same usage. A failed or cancelled trial never erases additive resources already consumed. For an evaluation-count budget, dispatch consumes the evaluation slot and it is not refunded merely because the trial later fails or is cancelled. For money/compute budgets, release only the measured or otherwise provable unused portion of the enforceable reservation. If unconsumed capacity cannot be established safely, retain the conservative charge. For elapsed wall time, record start/finish/cancellation times for audit but enforce B via the shared absolute deadline rather than a refundable additive reservation. Every reservation, coordinator charge, whole-search quota action, deadline/cap enforcement action, consumption adjustment, release, failure, cancellation, forced termination, and terminal accounting transaction is recorded in the ledger/audit trail. + +For targets that require deterministic search behavior, assign deterministic trial IDs and define a **deterministic proposal frontier**. A new proposal may be generated only from a declared ordered observation prefix that is the same in every replay. Buffer out-of-order completions until that prefix is available. Do **not** immediately refill whichever worker happens to become free if doing so would let wall-clock completion order choose the model state used for the next proposal. Acceptable deterministic designs include fixed deterministic batches/barriers, or an ordered-prefix scheduler where proposal `k+1` is generated only after the exact predeclared prefix needed for that proposal has been assimilated and its dispatch slot/order is determined independently of worker-speed races. Surrogate/model updates, acquisition decisions, domain contraction, portfolio-selection state, proposal generation, dispatch/refill decisions, and stopping criteria must consume the same deterministic state sequence. A fixed random seed plus buffered assimilation alone is not sufficient if worker availability can still change which proposal is generated next. If a target chooses immediate completion-driven refill for throughput, declare the resulting nondeterminism as an explicit contract change rather than claiming deterministic replay. + +Parallelism has an information cost: very wide batches receive less feedback between suggestions and can degenerate toward non-adaptive/random search. + +## Before / after evidence + +- Environment: No controlled target-repository tuning study has been run for this OPT synthesis. +- Baseline: No target comparison against manual/exhaustive/random tuning has been established. +- Optimized: No target adaptive-search result has been established. +- Speedup / memory reduction: No transferable claim; upstream libraries establish mechanisms, not a QSOL target win. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Keep a deterministic search seed where practical, preserve the full trial/search ledger, re-evaluate finalists, and validate the selected candidate against the reference contract on held-out/repeated workloads. For asynchronous search with **additive** budgets, test the boundary with multiple workers contending for the last remaining reservation and prove no dispatch can make `consumed + reserved` exceed B. Deliberately run trials that attempt to exceed their per-trial money/compute reservation and prove the quota mechanism prevents the overrun. Inject early failures, late failures, partial consumption, and cancellation after measurable work; verify that only demonstrably unconsumed reservation is released, evaluation-count slots are not resurrected after dispatch, and repeated failures cannot create extra budget capacity. + +For a **hard compute-unit B**, add coordinator-dominant fixtures rather than validating only trial quotas. Make trial evaluations cheap while surrogate fitting, acquisition optimization, proposal generation, model/portfolio updates, assimilation, serialization, and stopping logic deliberately consume most of the compute. Under the ledger design, prove those operations reserve/charge the same compute budget and cannot start when capacity is unavailable; under a whole-search-quota design, prove coordinator and workers are all contained by the same enforceable quota. Stress repeated model refits and failed proposal attempts and verify they cannot create uncharged compute. The run must remain within B even when coordinator work dominates total compute. + +For **elapsed wall-time** B, use a controlled monotonic clock and launch multiple workers concurrently under one shared deadline. Verify two trials each permitted to run until the same ten-minute deadline are admissible without requiring twenty minutes of additive reservation. Permute worker count, start order, and completion times; prove the search stops launching work as the deadline approaches, every in-flight worker observes the same deadline, forced termination completes within the declared enforcement bound, and total elapsed search lifetime never exceeds B plus only the explicitly declared bounded termination-cleanup allowance. Ensure no retry or model/proposal step can reset or extend the original deadline. + +Race multiple additive-resource trial/coordinator completions and cancellations against one another and against workers attempting the final dispatch slot. Verify the accounting transaction is linearizable: no consumed increment is lost, no reservation is released before its corresponding consumption is charged, and no participant observes capacity that would make the post-transaction invariant `consumed + reserved <= B` false. + +For deterministic targets, run the same seeded search with deliberately permuted worker speeds and completion orders, including the case where trial 2 finishes before trial 1 and frees a worker first. Verify out-of-order completion **does not permit proposal 3 to be generated from a different observation prefix**. The complete proposal sequence, parameter values, deterministic trial IDs, logical dispatch/refill order, surrogate/search states, selected candidate, and stopping reason must match the deterministic reference. Test both fixed-batch/barrier scheduling and any ordered-prefix scheduler the target claims to support. Where sequential/parallel equivalence is part of C, compare the asynchronous execution with its deterministic sequential or batch replay. If completion-driven refill is intentionally retained, verify the target explicitly labels the search trace nondeterministic instead of claiming replay equivalence. + +## Target-repo adaptation + +Do not copy acquisition constants, trial counts, domain contraction rates or parallel widths. Treat them as optimizer parameters with their own evidence boundary. Define each budget dimension as **additive**, **elapsed wall time**, or another explicitly modeled resource. For additive resources, define the accounting unit, conservative per-operation/trial reservation amount, enforcement mechanism, one atomic/serializable reservation/completion ledger, metering source, and failure/cancellation charging policy. If compute is advertised as a hard whole-search B, explicitly inventory the coordinator/search-control paths that consume it and either place the complete optimizer (workers plus coordinator) under one enforceable quota or define how surrogate fits, acquisition/proposal work, model updates, coordination, assimilation, serialization, and stopping logic reserve/charge the shared ledger. For elapsed wall time, define the monotonic absolute search deadline, maximum bounded termination-cleanup interval, worker cancellation/termination mechanism, and the minimum remaining-time rule for new dispatch. Also define the deterministic observation-assimilation policy, **deterministic proposal frontier and dispatch/refill schedule** (when required), and the exact condition under which a freed worker may receive new work before enabling asynchronous dispatch. + +## Failure modes + +Noisy objectives, nonstationary machines, weak surrogates, excessive dimensionality and too much concurrency can waste evaluations or overfit benchmark noise. Non-atomic reservation can oversubscribe an additive evaluation or monetary cap; non-atomic completion/release can transiently undercount consumed plus reserved or lose concurrent increments; an unenforced additive reservation can let a single trial exceed B before accounting observes it; refunding consumed resources can let repeated late failures exceed B; **charging only trials while leaving surrogate fitting, proposal/acquisition work, model updates, coordination, assimilation, or stopping logic outside a claimed hard compute budget can exceed B even when every trial quota is correct**; treating elapsed wall time as an additive per-worker resource can falsely reject valid overlapping trials and serialize the search; conversely, a nominal wall-time limit without one enforceable shared deadline can let work continue past B; wall-clock completion-order assimilation can make supposedly deterministic search traces irreproducible; **immediate worker refill can also make proposals nondeterministic even when assimilation itself is buffered**. + +## Rollback trigger + +Stop adaptive search when its overhead exceeds evaluation savings, the declared budget is exhausted, repeated validation does not confirm the selected improvement, any additive covered operation can consume beyond its enforceable reservation/quota, additive accounting/concurrency tests can violate B, **coordinator/search-control work can escape a hard whole-search compute ledger/quota**, any hard elapsed-wall-time run can exceed its shared absolute deadline beyond the declared bounded cleanup allowance, a retry/worker can extend or reset that deadline, or any target that requires deterministic search produces different proposals, logical dispatch/refill order, model states, selected candidates, or stopping reasons under permuted asynchronous completion orders. diff --git a/optimizations/OPT-SET-001-density-adaptive-compact-sets.md b/optimizations/OPT-SET-001-density-adaptive-compact-sets.md new file mode 100644 index 0000000..8650414 --- /dev/null +++ b/optimizations/OPT-SET-001-density-adaptive-compact-sets.md @@ -0,0 +1,76 @@ +# OPT-SET-001 — Density-adaptive compact set representation + +**Status:** Implemented external reference; target validation required +**Domains:** graphs, indexes, membership sets, telemetry, integer identifiers + +## Source evidence + +- https://jazco.dev/2024/04/20/roaring-bitmaps/ +- https://jazco.dev/2024/04/15/in-memory-graphs/ +- `sources/JAZCO.md` + +## Problem + +A single representation performs poorly across regions with very different density: sparse bitmaps waste memory, while list-like sparse structures make dense set algebra expensive. + +## Optimization problem contract + +- X: target-supported partition widths, sparse/dense container choices, switching thresholds, serialization layouts/versions, and migration policies +- F: representations that preserve exact membership and set-operation semantics, preserve contractually significant iteration ordering when one exists, and satisfy target memory/serialization compatibility constraints +- f: measured memory footprint plus target-relevant set-operation and serialization latency +- d: minimize under the target's predeclared scalar, lexicographic, or Pareto ordering +- C: membership, union, intersection, difference, iteration behavior where observable, mutation semantics, and persistence/compatibility round trips match the canonical reference contract exactly +- B: target-specific benchmark budget over declared sparse, dense, mixed, transition-boundary, mutation, and compatibility datasets; no portable trial count is supplied here +- S: stop when the declared budget is exhausted or a validated representation meets the target objective without violating C +- Variables: integer / categorical / mixed +- Search scope: local representation/threshold/layout tuning +- Objective behavior: noisy for performance; set semantics are deterministic +- Information: derivative-free performance measurements +- Evaluation cost: cheap to moderate per fixture; may become expensive at production scale +- Constraints: exact set semantics, ordering where observable, memory, serialization, version compatibility, and migration constraints +- Parallelism: sequential for representation transitions unless the target separately defines safe concurrent mutation semantics +- Exactness: exact set semantics; no approximation is introduced + +## Preserved contract + +Membership and set operations must match the reference set exactly. If iteration order is part of the target API/serialization contract, representation changes must preserve that order exactly. Persisted or exchanged sets must remain readable/writable according to the target's explicit version-compatibility policy; otherwise the layout change requires an explicit migration/version boundary rather than being treated as transparent. + +## Optimization + +Partition the identifier space and choose a representation per partition according to local density. Keep sparse regions compact while using bitmap-like containers where dense boolean algebra is advantageous. Prefer representations that can be serialized without expanding to a larger intermediate form. + +For mutable sets, representation switching is part of the state machine: insertion/deletion may cross thresholds in either direction. Conversions must be semantics-preserving and idempotent with respect to the canonical set state, including any observable iteration order and persistence metadata. + +For persisted/mixed-version use, define a versioned serialization contract. Either retain backward/forward compatibility for the required reader/writer matrix or provide an explicit migration/version bump before emitting an incompatible layout. Do not infer external compatibility from a same-version self-round-trip. + +## Evidence boundary + +Jazco reports strong production-scale graph results, but OPT treats the numbers as source observations only. The portable claim is density-adaptive representation. + +## Before / after evidence + +- Environment: No controlled target-repository benchmark has been run for this OPT record. +- Baseline: Not established in a target repository. +- Optimized: Not established in a target repository. +- Speedup / memory reduction: No transferable claim; reported graph results remain historical external observations. +- Variance / repetitions: Not available for a controlled OPT target benchmark. + +## Validation + +Differential-test membership, union, intersection, difference, iteration (when observable), and persistence against a simple canonical set implementation over sparse, dense and transition-boundary fixtures. + +Exercise **mutable transition sequences**: insert/delete elements so each switching threshold is crossed repeatedly sparse→dense and dense→sparse, including oscillation directly around thresholds. After every mutation and conversion, compare membership, cardinality, union/intersection/difference, and—when part of C—the exact iteration sequence with the canonical reference. Serialize and reload after every transition, then repeat the same comparisons. + +Exercise **serialization compatibility** independently from same-version round trips. Keep golden fixtures from every required older format/version and prove the new reader accepts them without semantic loss. Where backward writing or forward reading is required, exercise those reader/writer combinations explicitly. If compatibility is intentionally broken, require a versioned migration that converts old persisted state before the new layout becomes authoritative, and verify old consumers cannot silently misinterpret new bytes. + +## Target-repo adaptation + +Benchmark partition sizes and switching thresholds on the real identifier distribution and CPU/cache hierarchy. Determine whether iteration order is observable. Define the serialization-version matrix, migration policy, and any hysteresis needed to avoid conversion churn around thresholds. + +## Failure modes + +Conversion defects can drop/duplicate members; repeated threshold crossing can cause churn; sparse↔dense transitions can reorder iteration; new layouts can make old persisted state unreadable or emit bytes older consumers reject; pathological distributions, serialization incompatibility and hidden temporary allocations can erase the benefit. + +## Rollback trigger + +Revert when target data does not show a memory/latency win, any mutable-transition differential test fails, observable iteration order changes, or any required persistence/version-compatibility fixture fails. diff --git a/optimizations/OPT-SIMD-001-evidence-gated-native-autovectorization.md b/optimizations/OPT-SIMD-001-evidence-gated-native-autovectorization.md new file mode 100644 index 0000000..0f92ac2 --- /dev/null +++ b/optimizations/OPT-SIMD-001-evidence-gated-native-autovectorization.md @@ -0,0 +1,98 @@ +# OPT-SIMD-001 — Evidence-gated native autovectorization + +**Status:** Verified, environment-specific; donor measurements are isolated-kernel evidence, not a transferable end-to-end speedup claim. +**Domains:** CPU numerical kernels, hashing, simulation, DSP, batch transforms, compiler specialization + +## Source evidence + +- Repository: `QSOLKCB/GALAXY` +- PR: https://github.com/QSOLKCB/GALAXY/pull/10 +- Merge commit: `02f26f0f3630487b7db9522e9db3a5e5536f5765` +- Source note: `sources/GALAXY-CPU.md` +- Key donor files: `cpu-runtime/src/bin/simd_probe.rs`, `scripts/bench-cpu-simd.sh`, `docs/CPU-SIMD.md` +- Licensing boundary: donor and this repository are Apache-2.0 at the pinned revisions; this record promotes the mechanism, not copied implementation code. + +## Problem + +A deterministic hot loop performs the same scalar or narrowly vectorized operation over a large batch, but the source layout or loop shape prevents the compiler from generating the widest useful instructions for the host. Hand-written ISA intrinsics may be premature, non-portable, or unsupported by the project's compiler baseline, while a generic build may leave substantial throughput unused. + +## Optimization problem contract + +- X: Semantically equivalent loop shapes, batch layouts, compiler target settings and optional native-specialized build variants for the identified hot kernel. +- F: Candidates that preserve exact element results and aggregate checksums, compile on the supported toolchain, never expose unsupported instructions through the portable/default path, and retain an auditable reference implementation. +- f: Measured kernel runtime together with code-generation evidence and portability/deployment cost for each validated candidate. +- d: Minimize repeatable runtime after correctness gates; prefer the simpler portable candidate when performance is a near tie or code-generation evidence is ambiguous. +- C: Exact output/checksum parity, deterministic repeat stability, truthful ISA evidence, and no promotion of an isolated-kernel result into an end-to-end claim without separate full-path measurement. +- B: A bounded set of generic/native builds, benchmark repetitions, representative batch sizes and assembly inspections on the target hardware envelope. +- S: Stop after all declared candidates have passed parity and repeated measurement; retain the generic/reference path unless a native/vectorized candidate shows a useful repeatable advantage without violating deployment constraints. +- Variables: categorical and conditional +- Search scope: local +- Objective behavior: noisy +- Information: black-box +- Evaluation cost: moderate +- Constraints: semantic and resource +- Parallelism: sequential +- Exactness: exact + +## Preserved contract + +The optimized build must produce exactly the same declared element outputs and aggregate checksum as the reference path. Build-target specialization may change generated instructions, but it must not silently weaken arithmetic, hashing, ordering, determinism, or supported-machine behavior. A native binary must not become the universal default unless deployment guarantees the required ISA. + +Microbenchmark or probe wins remain probe evidence. They do not establish whole-application speedup until the transformed loop is measured inside the real memory, scheduling, setup and reduction path. + +## Optimization + +1. Isolate the suspected hot kernel behind a deterministic reference function. +2. Reshape data into a compiler-friendly batch form, commonly structure-of-arrays or separate primitive arrays, so independent iterations are visible to the optimizer. +3. Build the exact same source with a portable target and with target-specific/native code generation. +4. Compare every element and aggregate checksum against the scalar/reference path before timing. +5. Repeat timings under a controlled workload and inspect emitted assembly or equivalent compiler evidence to verify that the expected vector form actually exists. +6. Treat ISA width as evidence, not the objective: wider instructions are valuable only when the measured target workload improves. +7. Integrate only after the isolated gain survives the relevant production path, retaining a portable/reference fallback. + +The reusable idea is not “turn on AVX2/AVX-512.” It is to make the loop vectorizable, verify what the compiler emitted, and require measured semantic-preserving benefit on the machine that will run it. + +## Before / after evidence + +- Environment: GALAXY donor run on AMD Ryzen 9 5950X / Zen 3, where AVX2/FMA are available and AVX-512 is not. +- Workload/fixture: deterministic contribution/hash batch from the GALAXY CPU runtime probe. +- Cold baseline: generic x86-64 build median reported as 34,942,130 ns in the subsequent merged phase's summary. +- Warm/no-op baseline where relevant: not applicable; the donor compared identical probe work under generic and host-native compilation. +- Small invalidation / partial-work case where relevant: not applicable. +- Large invalidation / full-work case where relevant: donor documentation includes an 8M-item confirmation procedure. +- Optimized: host-native build median reported as 10,742,437 ns. +- Speedup / memory / I/O / quality change: 3.252719099x isolated speedup and 69.256491% median reduction; exact checksum parity retained. Generic code used SSE2 packed operations while the native Zen 3 build showed AVX2/VEX packed operations. +- Variance / repetitions / raw samples: the donor retains machine-readable receipts and repeated runs; this OPT record does not promote the source timings as a target expectation. + +## Validation + +Require all of the following before promotion: + +- element-by-element equality with the canonical/reference function; +- identical aggregate checksum across generic/native/vectorized candidates; +- repeat checksum stability; +- explicit host feature reporting or deployment capability guarantees; +- assembly/compiler evidence that the tested optimized build actually uses the intended vector form; +- repeated timing under the same workload and timing boundary; +- full-path validation after integration, including memory traffic, scheduling, reduction and RSS where those can erase the isolated gain. + +## Target-repo adaptation + +Re-profile the compiler version, minimum supported CPU, target-feature flags, batch size, data alignment, aliasing assumptions, integer/float semantics, hot-loop shape and deployment model. Do not copy `target-cpu=native` into distributed binaries unless the execution fleet guarantees compatibility. For floating-point kernels, separately decide whether reassociation, contraction/FMA or altered rounding is allowed; exact integer/hash evidence does not authorize floating-point semantic changes. + +## Failure modes + +- The loop is memory-bound, so wider arithmetic does not improve elapsed time. +- The compiler cannot vectorize because of aliasing, branches, gathers or unsupported operations. +- A native build improves the probe but regresses the integrated path due to cache pressure, downclocking or changed instruction mix. +- The optimized binary reaches hardware lacking the required ISA. +- Floating-point vectorization changes observable numerical behavior that the target contract requires to remain stable. +- Benchmark noise or frequency scaling makes a small apparent win non-repeatable. + +## Rollback trigger + +Disable or decline the specialized/vectorized path immediately on any parity/checksum failure, unsupported-instruction risk, reproducible full-path regression, or loss of deterministic behavior. Revert to the portable/reference implementation when the measured target workload does not retain a useful advantage after integration. + +## Composition notes + +Composes naturally with `OPT-SOA-001` when a data-layout change exposes independent batch lanes, and with `OPT-BUDGET-001` for environment-scoped regression protection after a configuration is promoted. Combine with `OPT-PAR-001` carefully: thread-level parallelism plus SIMD can shift the bottleneck to memory bandwidth or oversubscribe shared resources, so re-measure the complete path. \ No newline at end of file diff --git a/optimizations/OPT-SOA-001-worker-local-soa-tiling.md b/optimizations/OPT-SOA-001-worker-local-soa-tiling.md new file mode 100644 index 0000000..ba9a819 --- /dev/null +++ b/optimizations/OPT-SOA-001-worker-local-soa-tiling.md @@ -0,0 +1,102 @@ +# OPT-SOA-001 — Worker-local SoA tiling + +**Status:** Implemented external reference; GALAXY demonstrates exact cross-platform parity and guarded production integration, while target tile/worker settings remain environment-specific. +**Domains:** CPU simulation, rendering, numerical kernels, batch transforms, memory-bound pipelines + +## Source evidence + +- Repository: `QSOLKCB/GALAXY` +- Prototype PR: https://github.com/QSOLKCB/GALAXY/pull/11 +- Guarded production integration PR: https://github.com/QSOLKCB/GALAXY/pull/12 +- Merge commits: `8c69f87278d071db47d85a787905778eeec64ded` and `0e389cb4179902fae93f4a77d6ab2a539e73bc72` +- Source note: `sources/GALAXY-CPU.md` +- Licensing boundary: Apache-2.0 donor; this record describes the reusable data-layout/execution pattern rather than importing source code. + +## Problem + +A large array-of-structures working set is repeatedly traversed by multiple workers and frames/stages. The representation carries fields not needed by the hot path, causes poor cache/vector access, inflates resident memory, and forces worker execution to touch more data than necessary. A whole-population SoA conversion may itself be too large or expensive. + +## Optimization problem contract + +- X: Worker-local layout, packing, tile-capacity and deterministic partition choices that transform only the hot fields required for a bounded chunk of the source population. +- F: Candidates that preserve exact source-to-output semantics, deterministic partitioning and reduction, represent every required field without lossy reinterpretation, keep worker-local memory bounded, and retain an unchanged canonical/reference path. +- f: End-to-end runtime, peak working-memory/RSS evidence and useful worker scaling over the declared workload matrix. +- d: Pareto-minimize runtime and memory footprint subject to exact parity; reject candidates whose timing gain requires unacceptable RSS growth or unstable worker scaling. +- C: Exact canonical-output equality with the reference path, deterministic worker-count behavior where required, no dropped/duplicated elements, observable ordering preserved where required, and bounded worker-local storage independent of total resident population. Aggregate checksums are supplementary evidence unless the aggregate is itself the complete public output contract. +- B: A bounded worker-count × tile-size benchmark matrix with repeated runs and representative workload sizes on the target machines. +- S: Stop after a practical winning tile/worker region is identified or all candidates fail; do not promote a single pathological fast point without surrounding evidence. +- Variables: integer and categorical +- Search scope: global +- Objective behavior: noisy +- Information: black-box +- Evaluation cost: expensive +- Constraints: semantic and resource +- Parallelism: synchronous batch +- Exactness: exact + +## Preserved contract + +The tiled SoA path must compute the same declared outputs as the reference AoS path for every processed element and supported worker count. When the interface exposes per-element values or ordering, validation must compare those complete outputs element-by-element and preserve observable order. Reordering storage is allowed only when observable output order, tie behavior, reduction semantics and deterministic identity remain unchanged. + +An aggregate checksum may remain as an additional repeatability/corruption signal, but it is not a substitute for full-output comparison when richer output is observable. If the target contract exposes only the aggregate itself, state that boundary explicitly. + +Packing fields into narrower representations is permitted only when the representation is proven exact for the target domain or when the target contract explicitly allows approximation. This record is exact by default. + +## Optimization + +1. Keep the authoritative source representation or generation contract unchanged. +2. Partition the source population into deterministic contiguous or otherwise contract-safe worker ranges. +3. Give each worker a bounded tile containing only hot primitive fields in structure-of-arrays form. +4. Fill one tile, execute all profitable work for that tile while it is cache-resident, and reuse worker-local scratch rather than allocating a full transformed population. +5. Batch the resulting primitive lanes so compiler/native vectorization can operate on independent values. +6. Reduce worker results in a deterministic order when arithmetic/order semantics require it. +7. Sweep tile sizes and worker counts because the useful tile is a cache/memory/scheduling property of the target, not a universal constant. +8. Integrate behind an explicit guarded path until production evidence justifies any default change; preserve the canonical path as oracle/fallback. + +The key scaling property is that optimized working storage grows roughly with workers × tile capacity, not with the complete resident population. + +## Before / after evidence + +- Environment: GALAXY PRs #11–#12 exercised Linux x86-64, Linux ARM64, macOS ARM64 and Windows x86-64 parity/CI surfaces; performance evidence remained host-scoped. +- Workload/fixture: resident particle generation, BAM-LUT projection, contribution hashing and deterministic worker reduction across multiple frames. +- Cold baseline: full-resident 40-byte AoS reference shape in the donor experiment. +- Warm/no-op baseline where relevant: not promoted as a portable metric. +- Small invalidation / partial-work case where relevant: small tile/worker matrix cells validate capacity and partition behavior. +- Large invalidation / full-work case where relevant: donor sweep supports resident populations up to the full experimental workload and multiple tile sizes/workers. +- Optimized: bounded worker-local compact SoA tiles with worker-local x/y/output scratch and SIMD-friendly batch hashing. +- Speedup / memory / I/O / quality change: donor PR #12 states that PR #11 established exact cross-platform checksum parity and strong multi-host performance/memory evidence; this portable record additionally requires full canonical-output comparison whenever a target exposes per-element values or order. +- Variance / repetitions / raw samples: donor sweep emits raw matrix data, comparison tables and receipts; targets must repeat the sweep locally. + +## Validation + +- When the target exposes per-element values, compare the complete canonical/reference and SoA outputs element-by-element for every matrix cell, including observable order. Retain exact checksum equality as an additional repeatability signal rather than the sole equivalence proof. If the aggregate is the complete public output, state that explicitly. +- Verify deterministic partitioning covers every source element exactly once. +- Verify worker-count invariance when the target contract requires it. +- Test tile-capacity boundaries, partial final tiles and minimum/maximum supported sizes. +- Validate packed-field encode/decode or direct arithmetic equivalence independently. +- Measure end-to-end timing with the same setup/generation boundary for baseline and optimized paths. +- Record peak-memory/RSS scope honestly; process-wide high-water marks are not per-engine measurements unless isolated. +- Retain cross-platform CI for the guarded optimized path before changing defaults. + +## Target-repo adaptation + +Re-profile hot-field selection, field widths, tile capacity, cache hierarchy, worker count, scratch-array size, alignment, source-generation cost, frame/stage reuse depth and memory-bandwidth limits. Copy neither GALAXY's tile sizes nor its compact encodings without proving they fit the target domain exactly. Consider NUMA placement separately; ordinary worker-local tiling does not imply NUMA locality. + +## Failure modes + +- Tile fill/conversion overhead dominates the saved traversal cost. +- Tiles are too small to amortize setup or too large for useful cache residency. +- Narrow packing truncates or aliases values outside the donor's domain. +- Worker-local scratch multiplies memory enough to erase the AoS savings at high worker counts. +- Memory bandwidth becomes the bottleneck after vectorization/parallelism. +- Aggregate checksum parity masks an element-level or ordering defect on an interface that exposes richer output. +- Deterministic reduction is replaced by completion-order reduction and changes results. +- A guarded experimental win is promoted globally without evidence across the supported hardware/workload envelope. + +## Rollback trigger + +Fall back to the canonical representation/path on any full-output parity failure, missing/duplicated work, observable ordering mismatch, representation-range violation, unstable worker-count behavior, unacceptable RSS growth, or reproducible end-to-end slowdown. Keep the optimized path opt-in when the winning region is narrow or host-specific. + +## Composition notes + +Composes strongly with `OPT-SIMD-001` because SoA/batched primitives often unlock vector code generation, and with `OPT-POOL-001` when worker-local tiles can persist across repeated runs. Re-measure with `OPT-PAR-001`: more workers can increase local scratch and memory-bandwidth pressure even when each worker is individually faster. \ No newline at end of file diff --git a/scripts/check_catalog.py b/scripts/check_catalog.py new file mode 100644 index 0000000..24e2b4e --- /dev/null +++ b/scripts/check_catalog.py @@ -0,0 +1,1380 @@ +#!/usr/bin/env python3 +"""Run the catalog normalizer with definition-aware rendered validation.""" + +from __future__ import annotations + +import html +import posixpath +import re +import string + +import check_catalog_normalizer as normalizer + +REFERENCE_DEFINITION_RE = re.compile( + r"(?m)^ {0,3}\[(?P