Honest semantic metrics, flexible judge providers, robust gates - #31
Merged
Merged
Conversation
Enhancements ahead of first npm publish: - metrics: extract duplicated stopword/stemmer/tokenizer logic into a shared text-utils module; drop the unused compromise and natural deps - metrics/core: rename relevance lexical output to lexical_similarity; semantic_similarity now only populated via an EmbeddingProvider (true paraphrase-aware scoring on RelevanceScorer + AnswerCorrectnessScorer) - metrics/core: add RetrievalScorer (MRR, nDCG, precision/recall/hit@k) and AnswerCorrectnessScorer (answer vs. ground truth) - judge/core: support explicit provider, base_url, api_key for OpenAI-compatible gateways, proxies, and local models - judge: confidence is now self-reported or consensus-agreement based, replacing the circular distance-from-0.5 heuristic - gate/core: baseline gates gain a tolerance band; all gates gain warn/fail severity; GateResult carries warnings surfaced in CI output - docs: update README/ARCHITECTURE/AGENTS for the above; remove all references to DEV_PLAN.md All 20 test tasks pass (metrics 115, gate 39, judge 45); typecheck and lint clean; lockfile synced (frozen install verified). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Functionality enhancements ahead of the first npm publish, grounded in a read of the actual scorer/judge/gate implementations. Net +570 / −776 lines excluding the lockfile (which shed 456 lines of unused deps).
Metrics
compromiseandnaturalwere never imported; removed frompackage.jsonand the lockfile.semantic_similarity(actually char-bigram + Jaccard) is nowlexical_similarity. A new optionalEmbeddingProvidergives true cosine-of-embeddings semantic scoring;semantic_similarityonly appears when one is supplied.RetrievalScorer(MRR, nDCG, precision/recall/hit@k fromretrieved_chunk_idsvs.relevant_chunk_ids) andAnswerCorrectnessScorer(answer vs. ground truth, lexical or embeddings).Judge
JudgeConfignow acceptsprovider,base_url, andapi_keyso OpenAI-compatible gateways, proxies, and self-hosted/local models work instead of breaking the model-name keyword inference.Gate
warnvs.fail;GateResult.warningsis surfaced in JUnit (<skipped>) and Markdown reports without failing the build.Docs
README.md(scoring-modes table, gateway example),ARCHITECTURE.md, andAGENTS.md.DEV_PLAN.md.A changeset is included (core/metrics/judge/gate
minor, clipatch).Test plan
pnpm build— all 10 packagespnpm typecheck— cleanpnpm lint— clean (115 files)pnpm test— 20/20 tasks pass (metrics 115, gate 39, judge 45; 20 new tests added)pnpm install --frozen-lockfile— passes (lockfile synced)pnpm packon metrics — tarball clean (workspace:*→0.1.0, no stale deps)🤖 Generated with Claude Code