Skip to content

Honest semantic metrics, flexible judge providers, robust gates - #31

Merged
reaatech merged 1 commit into
mainfrom
enhance-metrics-judge-gate
May 23, 2026
Merged

reaatech merged 1 commit into
mainfrom
enhance-metrics-judge-gate

Conversation

@reaatech

Copy link
Copy Markdown
Owner

Summary

Functionality enhancements ahead of the first npm publish, grounded in a read of the actual scorer/judge/gate implementations. Net +570 / −776 lines excluding the lockfile (which shed 456 lines of unused deps).

Metrics

  • Shared text-utils — the duplicated stopword list, stemmer, and tokenizer (4 copies) are collapsed into one module; behavior preserved.
  • Dropped dead deps — compromise and natural were never imported; removed from package.json and the lockfile.
  • Honest naming + embeddings — the relevance scorer's misnamed semantic_similarity (actually char-bigram + Jaccard) is now lexical_similarity. A new optional EmbeddingProvider gives true cosine-of-embeddings semantic scoring; semantic_similarity only appears when one is supplied.
  • New scorers — RetrievalScorer (MRR, nDCG, precision/recall/hit@k from retrieved_chunk_ids vs. relevant_chunk_ids) and AnswerCorrectnessScorer (answer vs. ground truth, lexical or embeddings).

Judge

  • Explicit provider/endpoint — JudgeConfig now accepts provider, base_url, and api_key so OpenAI-compatible gateways, proxies, and self-hosted/local models work instead of breaking the model-name keyword inference.
  • Real confidence — replaced the circular distance-from-0.5 heuristic with the judge's self-reported confidence (parsed from the response) and, for consensus, inter-judge agreement.

Gate

  • Tolerance band — baseline-comparison gates absorb sampling noise so small dips don't fail CI.
  • Severity — all gates gain warn vs. fail; GateResult.warnings is surfaced in JUnit (<skipped>) and Markdown reports without failing the build.

Docs

  • Updated README.md (scoring-modes table, gateway example), ARCHITECTURE.md, and AGENTS.md.
  • Removed all references to DEV_PLAN.md.

A changeset is included (core/metrics/judge/gate minor, cli patch).

Test plan

  • pnpm build — all 10 packages
  • pnpm typecheck — clean
  • pnpm lint — clean (115 files)
  • pnpm test — 20/20 tasks pass (metrics 115, gate 39, judge 45; 20 new tests added)
  • pnpm install --frozen-lockfile — passes (lockfile synced)
  • pnpm pack on metrics — tarball clean (workspace:* → 0.1.0, no stale deps)

🤖 Generated with Claude Code

Enhancements ahead of first npm publish:

- metrics: extract duplicated stopword/stemmer/tokenizer logic into a
  shared text-utils module; drop the unused compromise and natural deps
- metrics/core: rename relevance lexical output to lexical_similarity;
  semantic_similarity now only populated via an EmbeddingProvider
  (true paraphrase-aware scoring on RelevanceScorer + AnswerCorrectnessScorer)
- metrics/core: add RetrievalScorer (MRR, nDCG, precision/recall/hit@k)
  and AnswerCorrectnessScorer (answer vs. ground truth)
- judge/core: support explicit provider, base_url, api_key for
  OpenAI-compatible gateways, proxies, and local models
- judge: confidence is now self-reported or consensus-agreement based,
  replacing the circular distance-from-0.5 heuristic
- gate/core: baseline gates gain a tolerance band; all gates gain
  warn/fail severity; GateResult carries warnings surfaced in CI output
- docs: update README/ARCHITECTURE/AGENTS for the above; remove all
  references to DEV_PLAN.md

All 20 test tasks pass (metrics 115, gate 39, judge 45); typecheck and
lint clean; lockfile synced (frozen install verified).

Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
@reaatech
reaatech merged commit 1c7c3f7 into main May 23, 2026
12 checks passed
@reaatech
reaatech deleted the enhance-metrics-judge-gate branch May 23, 2026 18:34
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant