Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
# Candidate Depth and Near-Duplicate Diversity Ablation

Date: 2026-08-23

Status: research result; no deployment or production data changes

## Question

The lexical coverage floor raised the noisy-prefix corpus to 14/20 gold at
rank one. All six remaining gold memories were at 0-based pool rank 10: the
first item excluded from a final top-10 page. Each was preceded by ten
almost-identical archived neighboring-service notes.

This ablation separated two hypotheses:

1. A deeper candidate pool alone lets the existing ranker resolve the cluster.
2. The result page needs an explicit diversity constraint so one
near-duplicate cluster cannot consume every slot.

## Controls

- Base commit: `47e8c82`
- Required flags in every arm: `RECALL_RERANK_SCALE_FIX`,
`RECALL_LEXICAL_COVERAGE_FLOOR`, and `RECALL_RESCUE_SQL_TIEBREAK`
- Final caller-visible top-K: 10
- Candidate depths: 10, 12, 20, 50
- Existing 20-task `mnemon-v02-noise-prefix` receipt and tasks
- Usage counters restored from `zz_probe_usage_snapshot` before every page and
deep-pool query, preventing cross-query usage-weight feedback
- Two independent server restarts/runs per arm; no LLM generation

An initial pilot reset usage only once per arm. It exposed a harness confound:
arms returning different memories produced different usage state for later
queries. Those pilot numbers are excluded below.

## Depth-only result (negative)

| Candidate depth | gold@1 | gold@5 | gold@10 | MRR@10 | Former misses |
| ---: | ---: | ---: | ---: | ---: | --- |
| 10 | 14/20 | 14/20 | 14/20 | 0.700 | absent |
| 12 | 14/20 | 14/20 | 14/20 | 0.700 | pool rank 10 |
| 20 | 14/20 | 14/20 | 14/20 | 0.700 | pool rank 10 |
| 50 | 14/20 | 14/20 | 14/20 | 0.700 | pool rank 10 |

Depth makes gold available but does not change final ordering. More candidates
alone buy zero retrieval quality here.

## Diversity intervention

The research prototype greedily preserves relevance order while allowing at
most two members from any token-set Jaccard cluster at similarity >= 0.90.
Suppressed rows are then backfilled in original relevance order so the API
still returns a complete top-10 page. It does not delete or mutate memories.

The candidate controls explicitly depend on `RECALL_RERANK_SCALE_FIX=true`:
diversifying a mixture of raw rescue-band and post-rerank scores is undefined,
so both controls are ignored when the scale fix is off. Configured candidate
depth is a **minimum**, not a maximum: the effective depth is
`max(caller limit, configured depth)`, so a caller asking for more than 12 rows
is never truncated. If the cluster cap is enabled without an explicit depth,
the service uses a bounded 50-row scan (or the caller limit when larger).

The threshold is conservative for this failure mode: sampled decoy-to-decoy
similarity is 0.951 while gold-to-decoy similarity is 0.178. The mechanism
does not inspect benchmark IDs, `archived` wording, or gold labels.

| Candidate depth | cluster cap | gold@1 | gold@5 | gold@10 | MRR@10 | Page size |
| ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| 12 | 2 | 14/20 | **20/20** | **20/20** | **0.800** | 10 |
| 20 | 2 | 14/20 | **20/20** | **20/20** | **0.800** | 10 |

At the smallest winning depth (12), all six former misses moved from 0-based
pool rank 10 to page rank 2. No former hit regressed. Both repeats were
bit-identical: 0/20 query rankings changed.

The wider matrix tested caps 4 and 8. Cap 4 moved all former misses to rank 4
(20/20 gold@5, MRR 0.760). Cap 8 moved them to rank 8 (20/20 gold@10, MRR
0.7333). Cap 2 is strongest while retaining two cluster representatives.

## Latency

- depth-10 control: p50 39-43 ms; p95 69-73 ms
- depth-12 + cap-2: p50 30-42 ms; p95 77-83 ms

There is no demonstrated material latency cost at this scale. A larger corpus
and production load test are still required before making a latency claim.

## Recommendation

The smallest winner is candidate depth 12, final top-K 10, Jaccard threshold
0.90, and at most two initially selected rows per cluster followed by
relevance-order backfill.

Keep it feature-flagged. It improves gold@5 from 14/20 to 20/20 and MRR@10 from
0.700 to 0.800 with no observed regression, no model call, and stable repeats.

Do not ship from this benchmark alone. Boilerplate-heavy decoys are deliberate,
and near-identical text can encode real version/environment distinctions. Next
validate ordinary noise, stale/superseded facts, contradictions, ambiguous
cross-project facts, false-positive injection, and downstream executable
quality.

## Artifacts

Raw JSON, server logs, repeats, former-miss ranks, page sizes, and latency
summaries are under `artifacts/candidate-matrix/`. Final backfill validation is
in `isolated-backfill/`; the wider isolated sweep is in `isolated/`.
Original file line number Diff line number Diff line change
@@ -0,0 +1,113 @@
# Candidate-diversity production gate

Date: 2026-08-23

This gate independently validates commit `08be7a4` against the 81-query,
1,210-memory real-embedding benchmark. It compares the combined recall fixes
with candidate diversity disabled and enabled. Nothing was pushed, deployed, or
run against production data.

## Configuration

Both arms enabled:

- `RECALL_RERANK_SCALE_FIX=true`
- `RECALL_LEXICAL_COVERAGE_FLOOR=true`
- `RECALL_RESCUE_SQL_TIEBREAK=true`
- final result limit 10

The treatment additionally enabled:

- `RECALL_CANDIDATE_POOL_DEPTH=12`
- `RECALL_NEAR_DUPLICATE_CLUSTER_LIMIT=2`
- token-set Jaccard threshold 0.90

The corpus used real `bge-base` 768-dimensional embeddings in a dedicated local
database. The raw artifact is
[`artifacts/diversity-production-gate/real-embedding-81.json`](artifacts/diversity-production-gate/real-embedding-81.json).

## Benchmark validity controls

Issue #326 identified two confounds in the existing harness. This run controls
both:

1. `seedCorpus` silently drops fixture tags, metadata, and memory type. The gate
restored all declared values, then compared the stored values with the fixture
definitions for all 1,210 memories. All 1,210 memories had their declared tags.
2. Recall mutates usage fields that can change later rankings. The gate took a
persistent snapshot of `retrieval_count`, `used_count`, `unused_count`,
`last_retrieved_at`, and `last_used_at`, then restored it before every query
and every repeat.

Every query was executed twice in each arm. No query changed ranking between
identical repeats in either arm.

## Results

| metric | diversity off | diversity on | delta |
| --- | ---: | ---: | ---: |
| required hit rate@5 | 60/62 (96.77%) | 60/62 (96.77%) | 0 |
| relevant recall@10 | 71/73 (97.26%) | 71/73 (97.26%) | 0 |
| MRR@10 | 0.91262 | 0.91282 | +0.00020 |
| forbidden-memory injection | 0/91 | 0/91 | 0 |
| full-page rate | 73/80 (91.25%) | 73/80 (91.25%) | 0 |
| mean page size | 9.4375 | 9.4375 | 0 |
| latency p50 | 23 ms | 18 ms | -5 ms |
| latency p95 | 31 ms | 38 ms | +7 ms |
| latency mean | 21.83 ms | 21.46 ms | -0.36 ms |

There were no top-five gains and no top-five losses. Eight emotional-query
rankings changed. Only `emotional_003` changed a relevant rank: the second
required memory moved from rank 9 to rank 7, producing the small MRR increase.
The other seven changes reordered non-required results. Page sizes were
identical query by query.

The latency sample is an execution sanity check, not a performance claim. Arms
ran sequentially rather than in an interleaved latency experiment; the mixed
p50/p95 movement does not establish a speedup or regression.

## Relationship to the targeted near-duplicate result

The controlled mnemon near-duplicate matrix remains the positive efficacy
evidence. Depth alone did not move the 14/20 ceiling. Depth 12 plus a two-member
near-duplicate cluster cap retained 14/20 at rank 1 and moved all six remaining
gold memories into the returned page: 20/20 hit@5, 20/20 hit@10, and MRR 0.800.
Two independent repeats were exact.

The 81-query gate supplies the complementary safety evidence: the same setting
did not trade away required top-five hits, recall, tenant isolation, stability,
or page fullness on the ordinary real-embedding suite.

## Graduated-corpus gap

Phase 1 defines deterministic generators for the requested conditions in
mnemon:

- G1: 20 targets plus 400 ordinary unrelated memories
- G3: 20 targets plus 200 stale/superseded memories
- G4: 20 targets plus 100 directly contradictory memories
- G5: 20 targets plus 200 ambiguous cross-project memories

No materialized corpus with a successful embedding-gate receipt or completed
retrieval report exists for these conditions. The only attempted smoke run was
invalid: it produced no usable embeddings because the local embedding path
returned NaNs and mixed 384-dimensional output with the 1,536-dimensional test
configuration. The seeder cleaned that attempt up. These conditions therefore
have **no result**, not a passing or failing result.

The 81-query suite contains temporal, emotional, adversarial, cross-feature,
edge-case, and tenant-isolation queries, but it is not a causal substitute for
G1/G3/G4/G5.

## Verdict

Commit `08be7a4` earns a merge recommendation **behind its existing default-off
flags**. The targeted near-duplicate condition shows a large, replicated gain,
and the corrected 81-query real-embedding gate finds no quality, isolation,
stability, or page-fullness regression.

It does not yet earn default-on production rollout. Before changing defaults,
materialize and embedding-gate G1, G3, G4, and G5; run the same isolated on/off
comparison; then run an interleaved latency/load check. A small internal canary
with decision telemetry is reasonable after merge because rollback is a flag
change.
3 changes: 3 additions & 0 deletions docs/research/memory-formation-query-transform/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,9 @@ interventions in Engram's memory lifecycle:
| 2 | [`02-research-memo.md`](02-research-memo.md) | Literature review (formation + query-transform), blind-spot pass, promising-vs-hype read, cited primary sources |
| 3 | [`03-experiment-spec.md`](03-experiment-spec.md) | A–E ablation matrix, graduated corpora, retrieval + downstream metrics, reproducibility controls, budgets |
| 4 | [`04-finding-tie-domination.md`](04-finding-tie-domination.md) | **Blocking finding.** Retrieval ranking is decided by array order among tied scores, not by score. Changes the required sequencing. |
| 5 | [`05-finding-band-inversion.md`](05-finding-band-inversion.md) | Rerank-scale mixing finding and scale-fix ablation. |
| 6 | [`06-candidate-depth-diversity-results.md`](06-candidate-depth-diversity-results.md) | Candidate-depth control and near-duplicate diversity ablation that breaks the 14/20 ceiling. |
| 7 | [`07-diversity-production-gate.md`](07-diversity-production-gate.md) | Corrected 81-query real-embedding safety gate for candidate diversity, including issue #326 controls and graduated-corpus gaps. |

## Executive summary

Expand Down
Loading
Loading