Skip to content

Recall benchmark: preserve fixture tags and isolate per-query usage state #326

Description

@heybeaux

Two benchmark-harness confounds were proven during recall-floor/diversity ablations.

1. Fixture tags are silently discarded

seedCorpus declares tagged fixtures but its INSERT path omits tags, so tag-aware retrieval experiments become no-ops while appearing valid. Seed every declared retrieval field and add a post-seed data-integrity assertion.

2. Queries mutate later-query ranking

Recall usage updates (retrievalCount, lastRetrievedAt, and downstream weights) feed later queries in the same arm. Full-suite order can therefore erase or create rank changes. A targeted temporal_006 replay moved rank 6→1 with tag evidence, while the same arm in full-suite order stayed rank 6. Candidate-diversity experiments also changed until snapshots were restored before every page/deep query.

Required fix

  • Restore a clean state before every query, or disable usage mutations in benchmark mode.
  • Randomize query order as an explicit stability arm, not accidental state.
  • Assert seeded tags/metadata equal fixtures.
  • Rebaseline the 81-query suite and document whether the 95% gate changes.
  • Preserve raw per-query state/version in reports.

No production behavior change is implied by this issue; this is benchmark validity and reproducibility debt.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions