From 2c8acd2f3b8d9fd774e7a5498e4dbff0b0e490df Mon Sep 17 00:00:00 2001 From: reaatech <138725666+reaatech@users.noreply.github.com> Date: Sat, 23 May 2026 11:31:24 -0700 Subject: [PATCH] feat: honest semantic metrics, flexible judge providers, robust gates Enhancements ahead of first npm publish: - metrics: extract duplicated stopword/stemmer/tokenizer logic into a shared text-utils module; drop the unused compromise and natural deps - metrics/core: rename relevance lexical output to lexical_similarity; semantic_similarity now only populated via an EmbeddingProvider (true paraphrase-aware scoring on RelevanceScorer + AnswerCorrectnessScorer) - metrics/core: add RetrievalScorer (MRR, nDCG, precision/recall/hit@k) and AnswerCorrectnessScorer (answer vs. ground truth) - judge/core: support explicit provider, base_url, api_key for OpenAI-compatible gateways, proxies, and local models - judge: confidence is now self-reported or consensus-agreement based, replacing the circular distance-from-0.5 heuristic - gate/core: baseline gates gain a tolerance band; all gates gain warn/fail severity; GateResult carries warnings surfaced in CI output - docs: update README/ARCHITECTURE/AGENTS for the above; remove all references to DEV_PLAN.md All 20 test tasks pass (metrics 115, gate 39, judge 45); typecheck and lint clean; lockfile synced (frozen install verified). Co-Authored-By: Claude Opus 4.7 --- .changeset/enhance-metrics-judge-gate.md | 16 + AGENTS.md | 22 +- ARCHITECTURE.md | 57 ++- README.md | 39 +- packages/cli/src/commands/report.command.ts | 8 +- packages/core/src/domain.ts | 84 +++- packages/core/src/schemas.ts | 29 ++ packages/gate/src/baseline-gates.ts | 16 +- packages/gate/src/ci-integration.ts | 16 +- packages/gate/src/engine.ts | 33 +- packages/gate/tests/baseline-gates.test.ts | 30 ++ packages/gate/tests/ci-integration.test.ts | 2 + packages/gate/tests/gate.test.ts | 42 ++ packages/judge/src/engine.ts | 91 ++-- packages/judge/src/prompts.ts | 32 +- packages/judge/tests/judge.test.ts | 61 +++ packages/metrics/package.json | 2 - packages/metrics/src/answer-correctness.ts | 106 ++++ packages/metrics/src/context-precision.ts | 163 +------ packages/metrics/src/context-recall.ts | 168 +------ packages/metrics/src/faithfulness.ts | 177 +------ packages/metrics/src/index.ts | 17 +- packages/metrics/src/relevance.ts | 290 +++-------- packages/metrics/src/retrieval.ts | 110 +++++ packages/metrics/src/text-utils.ts | 219 +++++++++ .../metrics/tests/answer-correctness.test.ts | 50 ++ packages/metrics/tests/metrics.test.ts | 3 +- packages/metrics/tests/relevance.test.ts | 44 +- packages/metrics/tests/retrieval.test.ts | 74 +++ pnpm-lock.yaml | 456 ------------------ 30 files changed, 1208 insertions(+), 1249 deletions(-) create mode 100644 .changeset/enhance-metrics-judge-gate.md create mode 100644 packages/metrics/src/answer-correctness.ts create mode 100644 packages/metrics/src/retrieval.ts create mode 100644 packages/metrics/src/text-utils.ts create mode 100644 packages/metrics/tests/answer-correctness.test.ts create mode 100644 packages/metrics/tests/retrieval.test.ts diff --git a/.changeset/enhance-metrics-judge-gate.md b/.changeset/enhance-metrics-judge-gate.md new file mode 100644 index 0000000..5173c69 --- /dev/null +++ b/.changeset/enhance-metrics-judge-gate.md @@ -0,0 +1,16 @@ +--- +"@reaatech/rag-eval-core": minor +"@reaatech/rag-eval-metrics": minor +"@reaatech/rag-eval-judge": minor +"@reaatech/rag-eval-gate": minor +"@reaatech/rag-eval-cli": patch +--- + +Enhance metric honesty, judge flexibility, and gate robustness ahead of first publish. + +- **metrics**: extract the duplicated stop-word/stemmer/tokenizer logic into a shared `text-utils` module and drop the unused `compromise` and `natural` dependencies. +- **metrics/core**: rename the relevance scorer's lexical output to `lexical_similarity` (it was misleadingly called `semantic_similarity`). `semantic_similarity` is now only populated when an `EmbeddingProvider` is supplied, enabling true (paraphrase-aware) semantic scoring on `RelevanceScorer` and the new `AnswerCorrectnessScorer`. +- **metrics/core**: add `RetrievalScorer` (MRR, nDCG, precision/recall/hit@k from `retrieved_chunk_ids` vs. `relevant_chunk_ids`) and `AnswerCorrectnessScorer` (generated answer vs. ground truth). +- **judge/core**: support explicit `provider`, `base_url`, and `api_key` in `JudgeConfig` so OpenAI-compatible gateways, proxies, and self-hosted/local models work without relying on model-name keyword inference. +- **judge**: confidence is now the judge's self-reported certainty (parsed from the response) or, for consensus, derived from inter-judge agreement — instead of the previous circular distance-from-0.5 heuristic. +- **gate/core**: baseline-comparison gates gain a `tolerance` band so sampling noise doesn't trip CI, and all gates gain a `warn` vs. `fail` `severity`. `GateResult` now carries a `warnings` array; CI reports surface warnings without failing the build. diff --git a/AGENTS.md b/AGENTS.md index 812f6c9..561e95c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -57,11 +57,13 @@ core ← metrics ← suite ← mcp-server, cli | **Types & Schemas** | `@reaatech/rag-eval-core` | Domain types, Zod schemas | | **Faithfulness Scorer** | `@reaatech/rag-eval-metrics` | Measure answer grounding in context | | **Relevance Scorer** | `@reaatech/rag-eval-metrics` | Measure answer relevance to query | -| **Context Precision** | `@reaatech/rag-eval-metrics` | Measure retrieval ranking quality | +| **Context Precision** | `@reaatech/rag-eval-metrics` | Measure retrieval ranking quality (text-based) | | **Context Recall** | `@reaatech/rag-eval-metrics` | Measure ground truth coverage | -| **LLM Judge** | `@reaatech/rag-eval-judge` | Calibrated quality scoring with multi-provider support | +| **Retrieval Scorer** | `@reaatech/rag-eval-metrics` | Ranking metrics (MRR, nDCG, precision/recall/hit@k) from chunk IDs | +| **Answer Correctness** | `@reaatech/rag-eval-metrics` | Measure generated answer vs. ground truth | +| **LLM Judge** | `@reaatech/rag-eval-judge` | Calibrated quality scoring across any provider or OpenAI-compatible gateway | | **Cost Tracker** | `@reaatech/rag-eval-cost` | Per-evaluation cost calculation and budget enforcement | -| **Gate Engine** | `@reaatech/rag-eval-gate` | CI regression gates and threshold checks | +| **Gate Engine** | `@reaatech/rag-eval-gate` | CI regression gates with tolerance bands and warn/fail severity | | **Dataset Manager** | `@reaatech/rag-eval-dataset` | Dataset loading, validation, generation | | **Observability** | `@reaatech/rag-eval-observability` | Structured logging, OTel tracing, metrics | | **Evaluation Suite** | `@reaatech/rag-eval-suite` | Central orchestrator tying all modules together | @@ -126,7 +128,7 @@ Fast, stateless, composable operations for mid-task self-evaluation: | Tool | Input | Output | Use Case | |------|-------|--------|----------| | `rag_eval.judge.faithfulness` | `{ context, generated_answer }` | `{ score, statements, supported_count }` | Check if answer is faithful to context | -| `rag_eval.judge.relevance` | `{ query, generated_answer }` | `{ score, semantic_similarity, intent_score }` | Check if answer addresses query | +| `rag_eval.judge.relevance` | `{ query, generated_answer }` | `{ score, lexical_similarity, intent_score }` | Check if answer addresses query (semantic_similarity added when an embedding provider is configured) | | `rag_eval.judge.context_precision` | `{ query, context[], ground_truth }` | `{ score, map, ndcg }` | Check context ranking quality | | `rag_eval.judge.context_recall` | `{ query, context[], ground_truth }` | `{ score, total_facts, covered_facts }` | Check ground truth coverage | | `rag_eval.judge.cost_check` | `{ eval_result, budget }` | `{ within_budget, cost }` | Verify cost within budget | @@ -230,6 +232,13 @@ judge: # Primary judge model (any provider) model: claude-opus + # Explicit provider/endpoint — required for OpenAI-compatible gateways, + # proxies, or self-hosted/local models whose names lack a provider keyword. + # Omit to infer the provider from the model name. + # provider: openai + # base_url: http://localhost:11434/v1 + # api_key: ${OPENAI_API_KEY} + # Fallback models for resilience fallback_models: - gpt-4-turbo @@ -433,11 +442,13 @@ gates: operator: ">=" threshold: 0.85 + # severity: warn records a non-blocking warning instead of failing the run - name: min-relevance type: threshold metric: avg_relevance operator: ">=" threshold: 0.80 + severity: warn - name: min-context-precision type: threshold @@ -461,6 +472,8 @@ gates: type: baseline-comparison metric: overall_score allow_regression: false + # tolerance absorbs sampling noise so small dips don't fail CI + tolerance: 0.01 ``` --- @@ -678,7 +691,6 @@ Before deploying a RAG evaluation pipeline to production: ## References - **ARCHITECTURE.md** — System design deep dive and package relationships -- **DEV_PLAN.md** — Development checklist - **README.md** — Quick start and overview - **datasets/examples/** — Example evaluation datasets - **MCP Specification** — https://modelcontextprotocol.io/ diff --git a/ARCHITECTURE.md b/ARCHITECTURE.md index ec06021..9f390e2 100644 --- a/ARCHITECTURE.md +++ b/ARCHITECTURE.md @@ -77,7 +77,7 @@ core ← metrics ← suite ← mcp-server, cli | Package | Role | Depends On | Key Exports | |---------|------|------------|-------------| | `@reaatech/rag-eval-core` | Foundation types + Zod schemas | (leaf) | `EvaluationSample`, `EvalSuiteConfig`, `GateConfig`, `JudgeConfig`, `CostBreakdown`, schemas | -| `@reaatech/rag-eval-metrics` | Heuristic metric scorers | core | `FaithfulnessScorer`, `RelevanceScorer`, `ContextPrecisionScorer`, `ContextRecallScorer`, `MetricsEngine` | +| `@reaatech/rag-eval-metrics` | Metric scorers (lexical by default; semantic via an `EmbeddingProvider`) | core | `FaithfulnessScorer`, `RelevanceScorer`, `ContextPrecisionScorer`, `ContextRecallScorer`, `RetrievalScorer`, `AnswerCorrectnessScorer`, `MetricsEngine`, `text-utils` | | `@reaatech/rag-eval-cost` | Cost tracking infrastructure | core | `CostTracker`, `Pricing`, `BudgetManager`, `CostReporter` | | `@reaatech/rag-eval-judge` | LLM-as-judge | core, cost | `JudgeEngine`, `JudgeCalibrator`, `JudgeCostTracker`, prompts | | `@reaatech/rag-eval-gate` | Quality gates | core | `GateEngine`, `ThresholdGates`, `BaselineGates`, `CIIntegration` | @@ -207,17 +207,19 @@ core ← metrics ← suite ← mcp-server, cli │ Input: { query, generated_answer } │ │ │ │ ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ │ -│ │ Semantic │ │ Intent │ │ Score │ │ -│ │ Similarity │ │ Coverage │ │ Aggregation │ │ -│ │ │ │ │ │ │ │ -│ │ - Embedding- │ │ - Decompose │ │ - Weighted │ │ -│ │ based │ │ query into │ │ combination │ │ -│ │ similarity │ │ intents │ │ - Semantic │ │ -│ │ (cosine) │ │ - Check answer │ │ similarity + │ │ -│ │ │ │ coverage │ │ intent score │ │ +│ │ Similarity │ │ Intent │ │ Score │ │ +│ │ │ │ Coverage │ │ Aggregation │ │ +│ │ - Lexical │ │ │ │ │ │ +│ │ (word/bigram) │ │ - Decompose │ │ - Weighted │ │ +│ │ by default │ │ query into │ │ combination │ │ +│ │ - Semantic │ │ intents │ │ - Similarity + │ │ +│ │ (cosine) when │ │ - Check answer │ │ intent score │ │ +│ │ embeddings │ │ coverage │ │ │ │ +│ │ supplied │ │ │ │ │ │ │ └─────────────────┘ └─────────────────┘ └─────────────────┘ │ │ │ -│ Output: { score, semantic_similarity, intent_score, intents } │ +│ Output: { score, lexical_similarity, intent_score, │ +│ semantic_similarity? } │ └─────────────────────────────────────────────────────────────────────┘ ``` @@ -268,6 +270,40 @@ core ← metrics ← suite ← mcp-server, cli └─────────────────────────────────────────────────────────────────────┘ ``` +### Retrieval Scorer + +``` +┌─────────────────────────────────────────────────────────────────────┐ +│ Retrieval Scorer │ +│ Package: @reaatech/rag-eval-metrics │ +│ │ +│ Input: { retrieved_chunk_ids[], relevant_chunk_ids[], k? } │ +│ │ +│ Ranking metrics over the retrieved order vs. the relevant set: │ +│ - MRR (reciprocal rank of first relevant chunk) │ +│ - nDCG (binary-relevance, log-discounted) │ +│ - precision@k, recall@k, hit@k │ +│ │ +│ Output: { mrr, ndcg, precision_at_k, recall_at_k, hit_at_k, k } │ +└─────────────────────────────────────────────────────────────────────┘ +``` + +### Answer Correctness Scorer + +``` +┌─────────────────────────────────────────────────────────────────────┐ +│ Answer Correctness Scorer │ +│ Package: @reaatech/rag-eval-metrics │ +│ │ +│ Input: { generated_answer, ground_truth } │ +│ │ +│ - Lexical: token F1 + character bigram Dice (default) │ +│ - Semantic: cosine of embeddings when an EmbeddingProvider is set │ +│ │ +│ Output: { score, lexical_similarity, semantic_similarity? } │ +└─────────────────────────────────────────────────────────────────────┘ +``` + ### LLM Judge with Calibration ``` @@ -494,7 +530,6 @@ All logs are structured JSON with standard fields: ## References - **AGENTS.md** — Agent development guide -- **DEV_PLAN.md** — Development checklist - **README.md** — Quick start and overview - **datasets/examples/** — Example evaluation datasets - **MCP Specification** — https://modelcontextprotocol.io/ diff --git a/README.md b/README.md index c209544..71cf5d0 100644 --- a/README.md +++ b/README.md @@ -10,10 +10,11 @@ This monorepo provides a composable suite of packages for evaluating Retrieval-A ## Features -- **Four evaluation metrics** — faithfulness, relevance, context precision, and context recall with heuristic scorers -- **LLM-as-judge** — multi-provider judging (Anthropic, OpenAI, Google) with calibration and consensus voting +- **Generation metrics** — faithfulness, relevance, and answer-correctness with fast lexical scorers; supply an `EmbeddingProvider` for paraphrase-aware semantic scoring +- **Retrieval metrics** — context precision/recall plus ranking metrics (MRR, nDCG, precision/recall/hit@k) from retrieved vs. relevant chunk IDs +- **LLM-as-judge** — multi-provider judging (Anthropic, OpenAI, Google, and any OpenAI-compatible gateway or local model) with calibration, consensus voting, and agreement-based confidence - **Cost accounting** — per-sample and per-run token tracking with budget enforcement and alert thresholds -- **Quality gates** — threshold and baseline-comparison gates with formatted CI output and exit codes +- **Quality gates** — threshold and baseline-comparison gates with noise-tolerance bands, `warn`/`fail` severity, and formatted CI output with exit codes - **MCP server** — three-layer tool API (`judge.*`, `suite.*`, `gate.*`) for agent-driven evaluation - **Dataset management** — multi-format loading, Zod validation, synthetic generation, and version tracking - **Observability** — structured Pino logging, OpenTelemetry tracing, and Prometheus-compatible metrics @@ -111,12 +112,41 @@ rag-eval-pack report --results results.json --output report.md See [`datasets/examples/`](./datasets/examples/) for sample datasets and configuration files. +## Scoring modes + +Each generation metric can run at three levels of fidelity and cost — pick per use case: + +| Mode | How it works | Cost | Catches paraphrase? | +| ---- | ------------ | ---- | ------------------- | +| **Lexical** (default) | Word/character overlap + intent heuristics. Deterministic, no network, no key. | Free | No — surface form only | +| **Semantic** | Cosine similarity of embeddings via an `EmbeddingProvider` you supply. | Embedding API cost | Yes | +| **LLM judge** | An LLM rates the sample with calibration and optional consensus. | Token cost | Yes, with reasoning | + +Lexical scoring is reported as `lexical_similarity`; `semantic_similarity` is populated only when an embedding provider is configured. Use lexical for fast pre-commit smoke checks, semantic for paraphrase-sensitive metrics, and the judge for the final quality bar. + +```typescript +import { RelevanceScorer } from "@reaatech/rag-eval-metrics"; + +const scorer = new RelevanceScorer({ + embeddingProvider: { embed: async (texts) => myEmbedAPI(texts) }, +}); +``` + +Point the judge at any OpenAI-compatible gateway or local model via explicit config: + +```typescript +const suite = new EvaluationSuite({ + metrics: ["faithfulness", "relevance"], + judge: { provider: "openai", base_url: "http://localhost:11434/v1", model: "llama3.1" }, +}); +``` + ## Packages | Package | Description | | ------- | ----------- | | [`@reaatech/rag-eval-core`](./packages/core) | Canonical types, Zod schemas, and domain models | -| [`@reaatech/rag-eval-metrics`](./packages/metrics) | Heuristic metric scorers (faithfulness, relevance, precision, recall) | +| [`@reaatech/rag-eval-metrics`](./packages/metrics) | Metric scorers: faithfulness, relevance, context precision/recall, retrieval ranking, and answer correctness | | [`@reaatech/rag-eval-judge`](./packages/judge) | LLM-as-judge with calibration, consensus, and cost tracking | | [`@reaatech/rag-eval-cost`](./packages/cost) | Pricing, budgeting, and cost reporting | | [`@reaatech/rag-eval-gate`](./packages/gate) | Quality gates and CI regression checks | @@ -131,7 +161,6 @@ See [`datasets/examples/`](./datasets/examples/) for sample datasets and configu - [`ARCHITECTURE.md`](./ARCHITECTURE.md) — System design, package relationships, and data flows - [`AGENTS.md`](./AGENTS.md) — Coding conventions, tool architecture, and development guidelines - [`CONTRIBUTING.md`](./CONTRIBUTING.md) — Contribution workflow and release process -- [`DEV_PLAN.md`](./DEV_PLAN.md) — Development checklist and roadmap ## License diff --git a/packages/cli/src/commands/report.command.ts b/packages/cli/src/commands/report.command.ts index 61f33da..176c95e 100644 --- a/packages/cli/src/commands/report.command.ts +++ b/packages/cli/src/commands/report.command.ts @@ -31,7 +31,13 @@ export function createReportCommand(): Command { report = generateBasicMarkdownReport(results); } } else if (options.format === 'junit') { - const gateResult = results.gate_result || { passed: true, gates: [], failures: [] }; + const gateResult = results.gate_result || { + passed: true, + gates: [], + failures: [], + warnings: [], + evaluated_at: new Date().toISOString(), + }; report = ci.generateJUnitXml(gateResult, results); } else { report = JSON.stringify(results, null, 2); diff --git a/packages/core/src/domain.ts b/packages/core/src/domain.ts index 37c1ab4..0c83ccb 100644 --- a/packages/core/src/domain.ts +++ b/packages/core/src/domain.ts @@ -23,12 +23,23 @@ export interface EvaluationSample { ground_truth: string; /** RAG system's generated answer */ generated_answer: string; - /** IDs of retrieved chunks (optional) */ + /** IDs of retrieved chunks, in retrieval rank order (optional) */ retrieved_chunk_ids?: string[]; + /** IDs of the chunks that are actually relevant (ground-truth labels for retrieval metrics) */ + relevant_chunk_ids?: string[]; /** Additional metadata */ metadata?: Record; } +/** + * Provider of text embeddings, used for true semantic similarity scoring. + * Implementations may wrap a provider SDK, a local model, or a cache. + */ +export interface EmbeddingProvider { + /** Embed a batch of texts, returning one vector per input (in order). */ + embed(texts: string[]): Promise; +} + /** Faithfulness evaluation result */ export interface FaithfulnessResult { /** Score from 0 to 1 */ @@ -60,7 +71,9 @@ export interface StatementSupport { export interface RelevanceResult { /** Score from 0 to 1 */ score: number; - /** Semantic similarity score (if applicable) */ + /** Lexical similarity score (word/character overlap; surface-form only) */ + lexical_similarity?: number; + /** Semantic similarity score (cosine of embeddings; only when an EmbeddingProvider is configured) */ semantic_similarity?: number; /** Intent coverage score */ intent_score?: number; @@ -107,6 +120,41 @@ export interface FactCoverage { matching_context?: string; } +/** + * Retrieval ranking result. + * + * Standard information-retrieval metrics computed by comparing the ranked + * `retrieved_chunk_ids` against the ground-truth `relevant_chunk_ids`. + */ +export interface RetrievalResult { + /** Reciprocal rank of the first relevant chunk (0 if none retrieved) */ + mrr: number; + /** Normalized discounted cumulative gain over the ranking */ + ndcg: number; + /** Precision@k: fraction of the top-k retrieved chunks that are relevant */ + precision_at_k: number; + /** Recall@k: fraction of all relevant chunks present in the top-k */ + recall_at_k: number; + /** Hit@k: 1 if any relevant chunk appears in the top-k, else 0 */ + hit_at_k: number; + /** The cutoff k used for the @k metrics */ + k: number; + /** Explanation of the score */ + explanation?: string; +} + +/** Answer correctness result (generated answer vs. ground truth). */ +export interface AnswerCorrectnessResult { + /** Overall correctness score from 0 to 1 */ + score: number; + /** Lexical similarity to the ground truth (surface-form only) */ + lexical_similarity?: number; + /** Semantic similarity to the ground truth (only when an EmbeddingProvider is configured) */ + semantic_similarity?: number; + /** Explanation of the score */ + explanation?: string; +} + /** Complete evaluation result for a single sample */ export interface SampleEvalResult { /** Sample identifier */ @@ -239,16 +287,28 @@ export interface GateConfig { allow_regression?: boolean; /** Minimum improvement required */ min_improvement?: number; + /** + * Tolerance band (absolute metric units). A regression smaller than this is + * treated as noise and does not fail the gate. Defaults to 0 (strict). + */ + tolerance?: number; + /** + * Severity when the gate's check is not satisfied. `fail` (default) fails the + * overall run; `warn` records a warning but lets the run pass. + */ + severity?: 'warn' | 'fail'; } /** Gate evaluation result */ export interface GateResult { - /** Whether all gates passed */ + /** Whether all `fail`-severity gates passed */ passed: boolean; /** Individual gate results */ gates: IndividualGateResult[]; - /** Failures summary */ + /** Failures summary (fail-severity gates that did not pass) */ failures: GateFailure[]; + /** Warnings summary (warn-severity gates that did not pass) */ + warnings: GateFailure[]; /** Evaluation timestamp */ evaluated_at: string; } @@ -257,8 +317,12 @@ export interface GateResult { export interface IndividualGateResult { /** Gate name */ name: string; - /** Whether this gate passed */ + /** Whether this gate's check was satisfied */ passed: boolean; + /** Severity of this gate (defaults to `fail`) */ + severity?: 'warn' | 'fail'; + /** Whether an unsatisfied check was downgraded to a warning */ + warning?: boolean; /** Actual metric value */ actual_value: number; /** Expected value/threshold */ @@ -305,6 +369,16 @@ export interface EvalSuiteConfig { export interface JudgeConfig { /** Primary judge model */ model?: string; + /** + * Explicit provider. Overrides inference from the model name — required for + * OpenAI-compatible gateways, proxies, and self-hosted/local models whose + * names don't contain a recognizable provider keyword. + */ + provider?: LLMProvider; + /** Base URL for an OpenAI-/provider-compatible endpoint (gateway, proxy, local server) */ + base_url?: string; + /** API key override (falls back to provider-specific environment variables) */ + api_key?: string; /** Fallback models */ fallback_models?: string[]; /** Calibration settings */ diff --git a/packages/core/src/schemas.ts b/packages/core/src/schemas.ts index 0fba6e0..022330b 100644 --- a/packages/core/src/schemas.ts +++ b/packages/core/src/schemas.ts @@ -12,6 +12,7 @@ export const EvaluationSampleSchema = z.object({ ground_truth: z.string().min(1, 'Ground truth cannot be empty'), generated_answer: z.string().min(1, 'Generated answer cannot be empty'), retrieved_chunk_ids: z.array(z.string()).optional(), + relevant_chunk_ids: z.array(z.string()).optional(), metadata: z.record(z.string(), z.unknown()).optional(), }); @@ -36,11 +37,31 @@ export const FaithfulnessResultSchema = z.object({ /** Relevance result schema */ export const RelevanceResultSchema = z.object({ score: z.number().min(0).max(1), + lexical_similarity: z.number().min(0).max(1).optional(), semantic_similarity: z.number().min(0).max(1).optional(), intent_score: z.number().min(0).max(1).optional(), explanation: z.string().optional(), }); +/** Retrieval ranking result schema */ +export const RetrievalResultSchema = z.object({ + mrr: z.number().min(0).max(1), + ndcg: z.number().min(0).max(1), + precision_at_k: z.number().min(0).max(1), + recall_at_k: z.number().min(0).max(1), + hit_at_k: z.number().min(0).max(1), + k: z.number().min(0), + explanation: z.string().optional(), +}); + +/** Answer correctness result schema */ +export const AnswerCorrectnessResultSchema = z.object({ + score: z.number().min(0).max(1), + lexical_similarity: z.number().min(0).max(1).optional(), + semantic_similarity: z.number().min(0).max(1).optional(), + explanation: z.string().optional(), +}); + /** Context precision result schema */ export const ContextPrecisionResultSchema = z.object({ score: z.number().min(0).max(1), @@ -135,12 +156,16 @@ export const GateConfigSchema = z.object({ baseline: z.string().optional(), allow_regression: z.boolean().optional(), min_improvement: z.number().optional(), + tolerance: z.number().min(0).optional(), + severity: z.enum(['warn', 'fail']).optional(), }); /** Individual gate result schema */ export const IndividualGateResultSchema = z.object({ name: z.string(), passed: z.boolean(), + severity: z.enum(['warn', 'fail']).optional(), + warning: z.boolean().optional(), actual_value: z.number(), expected_value: z.number().optional(), baseline_diff: z.number().optional(), @@ -161,6 +186,7 @@ export const GateResultSchema = z.object({ passed: z.boolean(), gates: z.array(IndividualGateResultSchema), failures: z.array(GateFailureSchema), + warnings: z.array(GateFailureSchema).default([]), evaluated_at: z.string(), }); @@ -188,6 +214,9 @@ export const ConsensusConfigSchema = z.object({ /** Judge configuration schema */ export const JudgeConfigSchema = z.object({ model: z.string().optional(), + provider: z.enum(['anthropic', 'openai', 'google', 'mock']).optional(), + base_url: z.string().optional(), + api_key: z.string().optional(), fallback_models: z.array(z.string()).optional(), calibration: CalibrationConfigSchema.optional(), consensus: ConsensusConfigSchema.optional(), diff --git a/packages/gate/src/baseline-gates.ts b/packages/gate/src/baseline-gates.ts index 17c5706..d63b42d 100644 --- a/packages/gate/src/baseline-gates.ts +++ b/packages/gate/src/baseline-gates.ts @@ -17,19 +17,27 @@ export class BaselineGates { const diff = candidateValue - baselineValue; const allowRegression = gate.allow_regression ?? false; const minImprovement = gate.min_improvement ?? 0; + const tolerance = gate.tolerance ?? 0; + // The tolerance band absorbs small regressions (noise). + const requiredDiff = minImprovement - tolerance; let passed = false; let message = ''; + const transition = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)}`; + const within = tolerance > 0 ? ` (tolerance: ${tolerance.toFixed(3)})` : ''; if (allowRegression) { passed = true; - message = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)} (diff: ${diff.toFixed(3)})`; - } else if (diff >= minImprovement) { + message = `${transition} (diff: ${diff.toFixed(3)})`; + } else if (diff >= requiredDiff) { passed = true; - message = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)} (improved by ${diff.toFixed(3)})`; + message = + diff >= minImprovement + ? `${transition} (improved by ${diff.toFixed(3)})` + : `${transition} (within tolerance: ${diff.toFixed(3)})${within}`; } else { passed = false; - message = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)} (regression of ${Math.abs(diff).toFixed(3)}, minimum improvement required: ${minImprovement})`; + message = `${transition} (regression of ${Math.abs(diff).toFixed(3)}, minimum improvement required: ${minImprovement}${within})`; } return { diff --git a/packages/gate/src/ci-integration.ts b/packages/gate/src/ci-integration.ts index 9ff8c2b..d0e9f70 100644 --- a/packages/gate/src/ci-integration.ts +++ b/packages/gate/src/ci-integration.ts @@ -42,14 +42,16 @@ export class CIIntegration { } generateJUnitXml(gateResult: GateResult, _evalResults?: EvalResults): string { - const failedTests = gateResult.gates.filter((g) => !g.passed).length; + // `warn`-severity gates are not hard failures. + const isHardFailure = (g: { passed: boolean; warning?: boolean }) => !g.passed && !g.warning; + const failedTests = gateResult.gates.filter(isHardFailure).length; const testCases = gateResult.gates .map( ( gate, ) => ` - ${!gate.passed ? `${this.escapeXml(gate.message)}` : ''} + ${isHardFailure(gate) ? `${this.escapeXml(gate.message)}` : ''}${gate.warning ? `` : ''} `, ) .join('\n'); @@ -110,11 +112,19 @@ ${testCases} '| Gate | Result | Details |', '|------|--------|---------|', ...gateResult.gates.map( - (g) => `| ${g.name} | ${g.passed ? '✅ Pass' : '❌ Fail'} | ${g.message} |`, + (g) => + `| ${g.name} | ${g.passed ? '✅ Pass' : g.warning ? '⚠️ Warn' : '❌ Fail'} | ${g.message} |`, ), '', ]; + if (gateResult.warnings && gateResult.warnings.length > 0) { + lines.push( + `> ⚠️ ${gateResult.warnings.length} warning gate(s) did not pass (non-blocking).`, + '', + ); + } + return lines.join('\n'); } diff --git a/packages/gate/src/engine.ts b/packages/gate/src/engine.ts index 4f891f4..9e8bba6 100644 --- a/packages/gate/src/engine.ts +++ b/packages/gate/src/engine.ts @@ -43,19 +43,29 @@ export class GateEngine { const baselineToUse = baseline ?? this.baselineResults; const individualResults: IndividualGateResult[] = []; const failures: GateFailure[] = []; + const warnings: GateFailure[] = []; for (const gate of this.gates) { const result = this.evaluateGate(gate, results, baselineToUse); + const severity = gate.severity ?? 'fail'; + result.severity = severity; individualResults.push(result); if (!result.passed) { - failures.push({ + const failure: GateFailure = { gate_name: gate.name, metric: gate.metric, actual: result.actual_value, expected: result.expected_value ?? 0, difference: result.actual_value - (result.expected_value ?? 0), - }); + }; + // `warn`-severity gates record a warning but don't fail the run. + if (severity === 'warn') { + result.warning = true; + warnings.push(failure); + } else { + failures.push(failure); + } } } @@ -63,6 +73,7 @@ export class GateEngine { passed: failures.length === 0, gates: individualResults, failures, + warnings, evaluated_at: new Date().toISOString(), }; } @@ -157,21 +168,29 @@ export class GateEngine { const diff = candidateValue - baselineValue; const allowRegression = gate.allow_regression ?? false; const minImprovement = gate.min_improvement ?? 0; + const tolerance = gate.tolerance ?? 0; + // A diff at or above this threshold passes; the tolerance band absorbs + // small regressions (noise) so the gate doesn't trip on sampling jitter. + const requiredDiff = minImprovement - tolerance; let passed = false; let message = ''; + const transition = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)}`; + const within = tolerance > 0 ? ` (tolerance: ${tolerance.toFixed(3)})` : ''; if (allowRegression) { // Any change is allowed passed = true; - message = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)} (diff: ${diff.toFixed(3)})`; - } else if (diff >= minImprovement) { - // Must improve by at least minImprovement + message = `${transition} (diff: ${diff.toFixed(3)})`; + } else if (diff >= requiredDiff) { passed = true; - message = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)} (improved by ${diff.toFixed(3)})`; + message = + diff >= minImprovement + ? `${transition} (improved by ${diff.toFixed(3)})` + : `${transition} (within tolerance: ${diff.toFixed(3)})${within}`; } else { passed = false; - message = `${gate.metric}: ${baselineValue.toFixed(3)} -> ${candidateValue.toFixed(3)} (regression of ${Math.abs(diff).toFixed(3)}, minimum improvement required: ${minImprovement})`; + message = `${transition} (regression of ${Math.abs(diff).toFixed(3)}, minimum improvement required: ${minImprovement}${within})`; } return { diff --git a/packages/gate/tests/baseline-gates.test.ts b/packages/gate/tests/baseline-gates.test.ts index 113a25a..c089c4e 100644 --- a/packages/gate/tests/baseline-gates.test.ts +++ b/packages/gate/tests/baseline-gates.test.ts @@ -57,6 +57,36 @@ describe('BaselineGates', () => { expect(result.passed).toBe(true); }); + it('should pass a small regression within the tolerance band', () => { + const gate: BaselineGateConfig = { + name: 'tolerant', + type: 'baseline-comparison', + metric: 'avg_faithfulness', + baseline: 'baseline-1', + allow_regression: false, + tolerance: 0.02, + }; + // baseline 0.8, candidate 0.79 → diff -0.01, within tolerance 0.02 + const result = gates.evaluate(gate, 0.79, mockBaselineResults); + expect(result.passed).toBe(true); + expect(result.message).toContain('tolerance'); + }); + + it('should fail a regression larger than the tolerance band', () => { + const gate: BaselineGateConfig = { + name: 'tolerant', + type: 'baseline-comparison', + metric: 'avg_faithfulness', + baseline: 'baseline-1', + allow_regression: false, + tolerance: 0.02, + }; + // baseline 0.8, candidate 0.75 → diff -0.05, exceeds tolerance 0.02 + const result = gates.evaluate(gate, 0.75, mockBaselineResults); + expect(result.passed).toBe(false); + expect(result.message).toContain('regression'); + }); + it('should fail when regressed without allow_regression', () => { const gate: BaselineGateConfig = { name: 'no-regression', diff --git a/packages/gate/tests/ci-integration.test.ts b/packages/gate/tests/ci-integration.test.ts index da48713..645d0b2 100644 --- a/packages/gate/tests/ci-integration.test.ts +++ b/packages/gate/tests/ci-integration.test.ts @@ -26,6 +26,7 @@ describe('CIIntegration', () => { difference: -0.1, }, ], + warnings: [], evaluated_at: '2024-01-01T00:00:00Z', }; @@ -88,6 +89,7 @@ describe('CIIntegration', () => { { name: 'relevance', passed: true, actual_value: 0.85, message: 'passed' }, ], failures: [], + warnings: [], evaluated_at: mockGateResult.evaluated_at, }; const xml = ci.generateJUnitXml(passedResult); diff --git a/packages/gate/tests/gate.test.ts b/packages/gate/tests/gate.test.ts index 32c4d24..a0943fe 100644 --- a/packages/gate/tests/gate.test.ts +++ b/packages/gate/tests/gate.test.ts @@ -31,6 +31,48 @@ describe('GateEngine', () => { ...overrides, }); + describe('gate severity', () => { + it('records a warning instead of failing for warn-severity gates', () => { + engine.loadGates([ + { + name: 'soft-faithfulness', + type: 'threshold', + metric: 'avg_faithfulness', + operator: '>=', + threshold: 0.9, // 0.87 < 0.9 → unsatisfied + severity: 'warn', + }, + ]); + + const gateResult = engine.evaluate(createMockResults()); + + expect(gateResult.passed).toBe(true); // warn does not fail the run + expect(gateResult.failures).toHaveLength(0); + expect(gateResult.warnings).toHaveLength(1); + expect(gateResult.warnings[0]?.gate_name).toBe('soft-faithfulness'); + expect(gateResult.gates[0]?.warning).toBe(true); + expect(gateResult.gates[0]?.severity).toBe('warn'); + }); + + it('still fails the run for fail-severity gates (default)', () => { + engine.loadGates([ + { + name: 'hard-faithfulness', + type: 'threshold', + metric: 'avg_faithfulness', + operator: '>=', + threshold: 0.9, + }, + ]); + + const gateResult = engine.evaluate(createMockResults()); + + expect(gateResult.passed).toBe(false); + expect(gateResult.failures).toHaveLength(1); + expect(gateResult.warnings).toHaveLength(0); + }); + }); + describe('threshold gates', () => { it('should pass when metric meets threshold', () => { engine.loadGates([ diff --git a/packages/judge/src/engine.ts b/packages/judge/src/engine.ts index 00e1072..ab9c4a4 100644 --- a/packages/judge/src/engine.ts +++ b/packages/judge/src/engine.ts @@ -59,26 +59,39 @@ export class JudgeEngine { } /** - * Get API key for a model from environment + * Get API key for a provider. + * + * Precedence: explicit `config.api_key`, then the provider-specific + * environment variable. Lets self-hosted/gateway setups supply credentials + * directly instead of relying on model-name keyword inference. */ - private getApiKeyForModel(model: string): string { - const modelLower = model.toLowerCase(); - if (modelLower.includes('claude') || modelLower.includes('anthropic')) { - return process.env.ANTHROPIC_API_KEY ?? ''; + private getApiKeyForProvider(provider: LLMProvider): string { + if (this.config.api_key) { + return this.config.api_key; } - if (modelLower.includes('gpt') || modelLower.includes('openai')) { - return process.env.OPENAI_API_KEY ?? ''; - } - if (modelLower.includes('gemini') || modelLower.includes('google')) { - return process.env.GOOGLE_API_KEY ?? ''; + switch (provider) { + case 'anthropic': + return process.env.ANTHROPIC_API_KEY ?? ''; + case 'openai': + return process.env.OPENAI_API_KEY ?? ''; + case 'google': + return process.env.GOOGLE_API_KEY ?? ''; + default: + return ''; } - return ''; } /** - * Determine provider from model name + * Determine the provider for a model. + * + * An explicit `config.provider` always wins — required for OpenAI-compatible + * gateways, proxies, and local models whose names lack a provider keyword. + * Otherwise the provider is inferred from the model name. */ private getProvider(model: string): LLMProvider { + if (this.config.provider) { + return this.config.provider; + } const modelLower = model.toLowerCase(); if (modelLower.includes('claude') || modelLower.includes('anthropic')) { return 'anthropic'; @@ -125,7 +138,8 @@ export class JudgeEngine { score: Math.round(calibratedScore * 1000) / 1000, raw_score: parsed.score, explanation: parsed.explanation, - confidence: this.calculateConfidence(parsed.score), + // Prefer the judge's self-reported confidence; fall back to neutral. + confidence: parsed.confidence ?? 0.5, calibrated, model: targetModel, provider, @@ -200,13 +214,14 @@ export class JudgeEngine { return sum + r.score * weight; }, 0) / totalWeight; - const avgConfidence = results.reduce((sum, r) => sum + r.confidence, 0) / results.length; + // Confidence reflects how much the judges agreed, not the score itself. + const agreement = this.agreementConfidence(results.map((r) => r.score)); const explanations = results.map((r) => r.explanation).join('\n\n'); return { score: Math.round(weightedScore * 1000) / 1000, explanation: `Consensus (weighted): ${explanations}`, - confidence: Math.round(avgConfidence * 1000) / 1000, + confidence: Math.round(agreement * 1000) / 1000, calibrated: results.some((r) => r.calibrated), raw_score: weightedScore, model: config.models.map((m) => m.id).join(', '), @@ -329,9 +344,12 @@ export class JudgeEngine { system: string, user: string, ): Promise { - const apiKey = this.getApiKeyForModel(model); + const apiKey = this.getApiKeyForProvider(provider); + const baseUrl = this.config.base_url; - if (provider === 'mock' || !apiKey) { + // A configured base URL (gateway / local server) is enough to attempt a + // real call even when no API key is set. + if (provider === 'mock' || (!apiKey && !baseUrl)) { return this.mockLLMResponse(system, user); } @@ -342,11 +360,11 @@ export class JudgeEngine { try { switch (provider) { case 'anthropic': - return await this.callAnthropic(model, apiKey, system, user); + return await this.callAnthropic(model, apiKey, system, user, baseUrl); case 'openai': - return await this.callOpenAI(model, apiKey, system, user); + return await this.callOpenAI(model, apiKey, system, user, baseUrl); case 'google': - return await this.callGoogle(model, apiKey, system, user); + return await this.callGoogle(model, apiKey, system, user, baseUrl); default: return this.mockLLMResponse(system, user); } @@ -391,10 +409,14 @@ export class JudgeEngine { apiKey: string, system: string, user: string, + baseUrl?: string, ): Promise { // Dynamic import to avoid requiring the package at runtime if not used const { Anthropic } = await import('@anthropic-ai/sdk'); - const client = new Anthropic({ apiKey }); + const client = new Anthropic({ + apiKey: apiKey || 'not-required', + ...(baseUrl ? { baseURL: baseUrl } : {}), + }); const response = await client.messages.create({ model, @@ -414,9 +436,13 @@ export class JudgeEngine { apiKey: string, system: string, user: string, + baseUrl?: string, ): Promise { const { OpenAI } = await import('openai'); - const client = new OpenAI({ apiKey }); + const client = new OpenAI({ + apiKey: apiKey || 'not-required', + ...(baseUrl ? { baseURL: baseUrl } : {}), + }); const response = await client.chat.completions.create({ model, @@ -438,10 +464,11 @@ export class JudgeEngine { apiKey: string, system: string, user: string, + baseUrl?: string, ): Promise { const { GoogleGenerativeAI } = await import('@google/generative-ai'); - const genAI = new GoogleGenerativeAI(apiKey); - const generativeModel = genAI.getGenerativeModel({ model }); + const genAI = new GoogleGenerativeAI(apiKey || 'not-required'); + const generativeModel = genAI.getGenerativeModel({ model }, baseUrl ? { baseUrl } : undefined); const response = await generativeModel.generateContent(`${system}\n\n${user}`); return response.response.text(); @@ -463,6 +490,7 @@ export class JudgeEngine { const score = 0.5 + (normalizedScore - 0.5) * 0.3; return `Score: ${score.toFixed(2)} +Confidence: 0.50 Explanation: Mock evaluation - in production, this would be evaluated by an LLM judge.`; } @@ -477,11 +505,18 @@ Explanation: Mock evaluation - in production, this would be evaluated by an LLM } /** - * Calculate confidence in the score + * Confidence derived from inter-judge agreement. + * + * Tight clustering of scores → high confidence; wide disagreement → low. + * This is a real uncertainty signal, unlike deriving confidence from a + * single score's distance from 0.5. Uses 1 - 2·stddev, clamped to [0, 1]. */ - private calculateConfidence(score: number): number { - // Higher confidence when score is far from 0.5 (uncertain) - return Math.abs(score - 0.5) * 2; + private agreementConfidence(scores: number[]): number { + if (scores.length <= 1) return 1; + const mean = scores.reduce((sum, s) => sum + s, 0) / scores.length; + const variance = scores.reduce((sum, s) => sum + (s - mean) ** 2, 0) / scores.length; + const stdDev = Math.sqrt(variance); + return Math.max(0, Math.min(1, 1 - 2 * stdDev)); } /** diff --git a/packages/judge/src/prompts.ts b/packages/judge/src/prompts.ts index 06f0208..7e5dae0 100644 --- a/packages/judge/src/prompts.ts +++ b/packages/judge/src/prompts.ts @@ -45,6 +45,7 @@ Generated Answer: Rate the faithfulness (0-1) and provide your explanation. Format your response as: Score: [0.0-1.0] +Confidence: [0.0-1.0, how certain you are in this rating] Explanation: [your explanation]`, }; @@ -77,6 +78,7 @@ Generated Answer: Rate the relevance (0-1) and provide your explanation. Format your response as: Score: [0.0-1.0] +Confidence: [0.0-1.0, how certain you are in this rating] Explanation: [your explanation]`, }; @@ -111,6 +113,7 @@ Retrieved Context Chunks: Rate the context precision (0-1) and provide your explanation. Format your response as: Score: [0.0-1.0] +Confidence: [0.0-1.0, how certain you are in this rating] Explanation: [your explanation]`, }; @@ -147,6 +150,7 @@ Retrieved Context: Rate the context recall (0-1) and provide your explanation. Format your response as: Score: [0.0-1.0] +Confidence: [0.0-1.0, how certain you are in this rating] Explanation: [your explanation]`, }; @@ -188,6 +192,7 @@ Generated Answer: Rate the overall quality (0-1) and provide your explanation. Format your response as: Score: [0.0-1.0] +Confidence: [0.0-1.0, how certain you are in this rating] Explanation: [your explanation]`, }; @@ -217,10 +222,19 @@ export function applyPromptTemplate( } /** - * Parse judge response to extract score and explanation + * Parse judge response to extract score, self-reported confidence, and explanation. + * + * `confidence` is the judge's own stated certainty when present — a more honest + * signal than deriving confidence from the score itself. It is `undefined` when + * the model did not provide one. */ -export function parseJudgeResponse(response: string): { score: number; explanation: string } { +export function parseJudgeResponse(response: string): { + score: number; + confidence?: number; + explanation: string; +} { const scoreMatch = response.match(/Score:\s*([0-9]*\.?[0-9]+)/i); + const confidenceMatch = response.match(/Confidence:\s*([0-9]*\.?[0-9]+)/i); const explanationMatch = response.match(/Explanation:\s*(.*)/is); const score = scoreMatch ? Number.parseFloat(scoreMatch[1] ?? '0.5') : 0.5; @@ -229,5 +243,17 @@ export function parseJudgeResponse(response: string): { score: number; explanati // Clamp score to [0, 1] const clampedScore = Math.max(0, Math.min(1, score)); - return { score: clampedScore, explanation }; + const result: { score: number; confidence?: number; explanation: string } = { + score: clampedScore, + explanation: explanation ?? response, + }; + + if (confidenceMatch) { + const confidence = Number.parseFloat(confidenceMatch[1] ?? ''); + if (!Number.isNaN(confidence)) { + result.confidence = Math.max(0, Math.min(1, confidence)); + } + } + + return result; } diff --git a/packages/judge/tests/judge.test.ts b/packages/judge/tests/judge.test.ts index f470653..d30ce12 100644 --- a/packages/judge/tests/judge.test.ts +++ b/packages/judge/tests/judge.test.ts @@ -121,6 +121,67 @@ describe('JudgeEngine', () => { }); }); + describe('confidence', () => { + it('uses the judge self-reported confidence when present', async () => { + const engine = new JudgeEngine({ model: 'gpt-4o', provider: 'openai', api_key: 'k' }); + vi.spyOn(engine as never, 'callOpenAI' as never).mockResolvedValue( + 'Score: 0.80\nConfidence: 0.95\nExplanation: confident.', + ); + + const result = await engine.evaluate(sampleData, 'faithfulness'); + expect(result.score).toBe(0.8); + expect(result.confidence).toBe(0.95); + }); + + it('falls back to neutral confidence when none is reported', async () => { + const engine = new JudgeEngine({ model: 'gpt-4o', provider: 'openai', api_key: 'k' }); + vi.spyOn(engine as never, 'callOpenAI' as never).mockResolvedValue( + 'Score: 0.80\nExplanation: no confidence given.', + ); + + const result = await engine.evaluate(sampleData, 'faithfulness'); + expect(result.confidence).toBe(0.5); + }); + + it('derives consensus confidence from inter-judge agreement', async () => { + const config: JudgeConfig = { + provider: 'openai', + api_key: 'k', + consensus: { + enabled: true, + models: [{ id: 'gpt-4o' }, { id: 'gpt-4o-mini' }], + voting_strategy: 'weighted', + }, + }; + const engine = new JudgeEngine(config); + // Both judges fully agree on 0.8 → high agreement → confidence 1. + vi.spyOn(engine as never, 'callOpenAI' as never).mockResolvedValue( + 'Score: 0.80\nConfidence: 0.40\nExplanation: x.', + ); + + const result = await engine.evaluateWithConsensus(sampleData, 'faithfulness'); + expect(result.confidence).toBe(1); + }); + }); + + describe('explicit provider config', () => { + it('uses config.provider and config.base_url to route to a gateway', async () => { + const engine = new JudgeEngine({ + model: 'local-model', + provider: 'openai', + base_url: 'http://localhost:11434/v1', + }); + const spy = vi + .spyOn(engine as never, 'callOpenAI' as never) + .mockResolvedValue('Score: 0.70\nExplanation: gateway.'); + + const result = await engine.evaluate(sampleData, 'relevance', 'local-model'); + expect(spy).toHaveBeenCalledOnce(); + expect(result.provider).toBe('openai'); + expect(result.score).toBe(0.7); + }); + }); + describe('evaluateWithConsensus', () => { it('should fall back to single evaluation when consensus disabled', async () => { const engine = new JudgeEngine(defaultConfig); diff --git a/packages/metrics/package.json b/packages/metrics/package.json index a11201e..3c86dcd 100644 --- a/packages/metrics/package.json +++ b/packages/metrics/package.json @@ -42,8 +42,6 @@ }, "dependencies": { "@reaatech/rag-eval-core": "workspace:*", - "compromise": "^14.15.0", - "natural": "^8.1.1", "p-limit": "^7.3.0" }, "devDependencies": { diff --git a/packages/metrics/src/answer-correctness.ts b/packages/metrics/src/answer-correctness.ts new file mode 100644 index 0000000..914ca99 --- /dev/null +++ b/packages/metrics/src/answer-correctness.ts @@ -0,0 +1,106 @@ +import type { + AnswerCorrectnessResult, + EmbeddingProvider, + EvaluationSample, +} from '@reaatech/rag-eval-core'; +import { + cosineSimilarity, + diceCoefficient, + getBigrams, + getSignificantWords, +} from './text-utils.js'; + +/** Options for the answer-correctness scorer. */ +export interface AnswerCorrectnessScorerOptions { + /** + * Optional embedding provider. When supplied, semantic similarity to the + * ground truth is computed and used as the primary signal; otherwise scoring + * is lexical only. + */ + embeddingProvider?: EmbeddingProvider; +} + +/** + * Answer Correctness Scorer + * + * Measures how well the generated answer matches the ground-truth answer. + * Unlike faithfulness (answer vs. context) and relevance (answer vs. query), + * this compares the answer directly against the reference answer. + */ +export class AnswerCorrectnessScorer { + private embeddingProvider?: EmbeddingProvider; + + constructor(options: AnswerCorrectnessScorerOptions = {}) { + this.embeddingProvider = options.embeddingProvider; + } + + /** + * Score answer correctness for a single sample. + */ + async score(sample: EvaluationSample): Promise { + const { generated_answer, ground_truth } = sample; + + const lexicalSimilarity = this.lexicalSimilarity(generated_answer, ground_truth); + const semanticSimilarity = this.embeddingProvider + ? await this.semanticSimilarity(generated_answer, ground_truth) + : undefined; + + const primary = semanticSimilarity ?? lexicalSimilarity; + const score = Math.round(primary * 1000) / 1000; + + const result: AnswerCorrectnessResult = { + score, + lexical_similarity: Math.round(lexicalSimilarity * 1000) / 1000, + explanation: this.explain(score), + }; + if (semanticSimilarity !== undefined) { + result.semantic_similarity = Math.round(semanticSimilarity * 1000) / 1000; + } + return result; + } + + /** + * Score answer correctness for multiple samples. + */ + async scoreBatch(samples: EvaluationSample[]): Promise { + return Promise.all(samples.map((sample) => this.score(sample))); + } + + private async semanticSimilarity(answer: string, groundTruth: string): Promise { + if (!this.embeddingProvider) return 0; + const [answerVec, truthVec] = await this.embeddingProvider.embed([answer, groundTruth]); + if (!answerVec || !truthVec) return 0; + return Math.max(0, Math.min(1, cosineSimilarity(answerVec, truthVec))); + } + + /** Token-level F1 over significant words, blended with character bigram Dice. */ + private lexicalSimilarity(answer: string, groundTruth: string): number { + const answerLower = answer.toLowerCase(); + const truthLower = groundTruth.toLowerCase(); + + const answerTokens = getSignificantWords(answerLower); + const truthTokens = getSignificantWords(truthLower); + const truthSet = new Set(truthTokens); + const answerSet = new Set(answerTokens); + + let f1 = 0; + if (answerSet.size > 0 && truthSet.size > 0) { + const overlap = [...truthSet].filter((t) => answerSet.has(t)).length; + const precision = overlap / answerSet.size; + const recall = overlap / truthSet.size; + f1 = precision + recall > 0 ? (2 * precision * recall) / (precision + recall) : 0; + } + + const bigramSimilarity = diceCoefficient(getBigrams(answerLower), getBigrams(truthLower)); + + // Token F1 carries the bulk of the signal; bigrams reward surface overlap. + return f1 * 0.7 + bigramSimilarity * 0.3; + } + + private explain(score: number): string { + if (score >= 0.8) return 'Answer closely matches the ground truth'; + if (score >= 0.6) return 'Answer substantially matches the ground truth'; + if (score >= 0.4) return 'Answer partially matches the ground truth'; + return 'Answer diverges from the ground truth'; + } +} diff --git a/packages/metrics/src/context-precision.ts b/packages/metrics/src/context-precision.ts index d6531d3..3df12c9 100644 --- a/packages/metrics/src/context-precision.ts +++ b/packages/metrics/src/context-precision.ts @@ -1,4 +1,5 @@ import type { ContextPrecisionResult, EvaluationSample } from '@reaatech/rag-eval-core'; +import { getSignificantWords } from './text-utils.js'; /** * Context Precision Scorer @@ -61,8 +62,8 @@ export class ContextPrecisionScorer { } // Check keyword overlap with ground truth - const groundTruthWords = this.getSignificantWords(groundTruthLower); - const chunkWords = new Set(this.getSignificantWords(chunkLower)); + const groundTruthWords = getSignificantWords(groundTruthLower); + const chunkWords = new Set(getSignificantWords(chunkLower)); if (groundTruthWords.length === 0) { return 0.5; @@ -72,7 +73,7 @@ export class ContextPrecisionScorer { const groundTruthCoverage = matchedWords.length / groundTruthWords.length; // Also check overlap with query - const queryWords = this.getSignificantWords(query.toLowerCase()); + const queryWords = getSignificantWords(query.toLowerCase()); const queryMatched = queryWords.filter((w) => chunkWords.has(w)); const queryCoverage = queryWords.length > 0 ? queryMatched.length / queryWords.length : 0; @@ -138,162 +139,6 @@ export class ContextPrecisionScorer { return dcg / idcg; } - /** - * Get words from text - */ - private getWords(text: string): string[] { - return text.match(/[a-z]+/g) ?? []; - } - - /** - * Get significant words (filtering stop words) - */ - private getSignificantWords(text: string): string[] { - const stopWords = new Set([ - 'a', - 'an', - 'the', - 'is', - 'are', - 'was', - 'were', - 'be', - 'been', - 'being', - 'have', - 'has', - 'had', - 'do', - 'does', - 'did', - 'will', - 'would', - 'could', - 'should', - 'may', - 'might', - 'shall', - 'can', - 'to', - 'of', - 'in', - 'for', - 'on', - 'with', - 'at', - 'by', - 'from', - 'as', - 'into', - 'through', - 'during', - 'before', - 'after', - 'above', - 'below', - 'between', - 'out', - 'off', - 'over', - 'under', - 'again', - 'further', - 'then', - 'once', - 'here', - 'there', - 'when', - 'where', - 'why', - 'how', - 'all', - 'each', - 'every', - 'both', - 'few', - 'more', - 'most', - 'other', - 'some', - 'such', - 'no', - 'nor', - 'not', - 'only', - 'own', - 'same', - 'so', - 'than', - 'too', - 'very', - 'just', - 'and', - 'but', - 'or', - 'if', - 'while', - 'because', - 'until', - 'about', - 'against', - 'up', - 'down', - 'it', - 'its', - 'i', - 'me', - 'my', - 'myself', - 'we', - 'our', - 'ours', - 'ourselves', - 'you', - 'your', - 'yours', - 'yourself', - 'yourselves', - 'he', - 'him', - 'his', - 'himself', - 'she', - 'her', - 'hers', - 'herself', - 'they', - 'them', - 'their', - 'theirs', - 'themselves', - 'what', - 'which', - 'who', - 'whom', - 'this', - 'that', - 'these', - 'those', - ]); - - const words = this.getWords(text); - const numbers = text.match(/\d{2,}/g) ?? []; - const allTokens = [...words, ...numbers].map((w) => this.normalizeWord(w)); - return allTokens.filter((word) => word.length > 2 && !stopWords.has(word)); - } - - /** - * Normalize a word by stripping common English inflections - */ - private normalizeWord(word: string): string { - const len = word.length; - if (len <= 3) return word; - if (word.endsWith('ing') && len > 4) return word.slice(0, -3); - if (word.endsWith('ed') && len > 4) return word.slice(0, -2); - if (word.endsWith('s') && !word.endsWith('ss') && len > 3) return word.slice(0, -1); - return word; - } - /** * Generate explanation for the context precision score */ diff --git a/packages/metrics/src/context-recall.ts b/packages/metrics/src/context-recall.ts index d352aae..6108c70 100644 --- a/packages/metrics/src/context-recall.ts +++ b/packages/metrics/src/context-recall.ts @@ -1,4 +1,5 @@ import type { ContextRecallResult, EvaluationSample, FactCoverage } from '@reaatech/rag-eval-core'; +import { getSignificantWords, splitSentences } from './text-utils.js'; /** * Context Recall Scorer @@ -49,10 +50,7 @@ export class ContextRecallScorer { */ private extractFacts(text: string): string[] { // Split by common fact separators - const sentences = text - .split(/(?<=[.!?])\s+/) - .map((s) => s.trim()) - .filter((s) => s.length > 0); + const sentences = splitSentences(text); // For each sentence, extract key information const facts: string[] = []; @@ -97,7 +95,7 @@ export class ContextRecallScorer { // If no pattern matches, extract significant word groups if (phrases.length === 0) { - const words = this.getSignificantWords(processed.toLowerCase()); + const words = getSignificantWords(processed.toLowerCase()); if (words.length > 0) { // Group words into chunks of 2-3 for (let i = 0; i < words.length; i += 2) { @@ -134,8 +132,8 @@ export class ContextRecallScorer { } // Check keyword overlap - const factWords = this.getSignificantWords(factLower); - const contextWords = new Set(this.getSignificantWords(contextLower)); + const factWords = getSignificantWords(factLower); + const contextWords = new Set(getSignificantWords(contextLower)); if (factWords.length === 0) { return { @@ -195,162 +193,6 @@ export class ContextRecallScorer { return bestMatch || context.substring(0, 100); } - /** - * Get words from text - */ - private getWords(text: string): string[] { - return text.match(/[a-z]+/g) ?? []; - } - - /** - * Get significant words (filtering stop words) - */ - private getSignificantWords(text: string): string[] { - const stopWords = new Set([ - 'a', - 'an', - 'the', - 'is', - 'are', - 'was', - 'were', - 'be', - 'been', - 'being', - 'have', - 'has', - 'had', - 'do', - 'does', - 'did', - 'will', - 'would', - 'could', - 'should', - 'may', - 'might', - 'shall', - 'can', - 'to', - 'of', - 'in', - 'for', - 'on', - 'with', - 'at', - 'by', - 'from', - 'as', - 'into', - 'through', - 'during', - 'before', - 'after', - 'above', - 'below', - 'between', - 'out', - 'off', - 'over', - 'under', - 'again', - 'further', - 'then', - 'once', - 'here', - 'there', - 'when', - 'where', - 'why', - 'how', - 'all', - 'each', - 'every', - 'both', - 'few', - 'more', - 'most', - 'other', - 'some', - 'such', - 'no', - 'nor', - 'not', - 'only', - 'own', - 'same', - 'so', - 'than', - 'too', - 'very', - 'just', - 'and', - 'but', - 'or', - 'if', - 'while', - 'because', - 'until', - 'about', - 'against', - 'up', - 'down', - 'it', - 'its', - 'i', - 'me', - 'my', - 'myself', - 'we', - 'our', - 'ours', - 'ourselves', - 'you', - 'your', - 'yours', - 'yourself', - 'yourselves', - 'he', - 'him', - 'his', - 'himself', - 'she', - 'her', - 'hers', - 'herself', - 'they', - 'them', - 'their', - 'theirs', - 'themselves', - 'what', - 'which', - 'who', - 'whom', - 'this', - 'that', - 'these', - 'those', - ]); - - const words = this.getWords(text); - const numbers = text.match(/\d{2,}/g) ?? []; - const allTokens = [...words, ...numbers].map((w) => this.normalizeWord(w)); - return allTokens.filter((word) => word.length > 2 && !stopWords.has(word)); - } - - /** - * Normalize a word by stripping common English inflections - */ - private normalizeWord(word: string): string { - const len = word.length; - if (len <= 3) return word; - if (word.endsWith('ing') && len > 4) return word.slice(0, -3); - if (word.endsWith('ed') && len > 4) return word.slice(0, -2); - if (word.endsWith('s') && !word.endsWith('ss') && len > 3) return word.slice(0, -1); - return word; - } - /** * Generate explanation for the context recall score */ diff --git a/packages/metrics/src/faithfulness.ts b/packages/metrics/src/faithfulness.ts index f592f61..989301b 100644 --- a/packages/metrics/src/faithfulness.ts +++ b/packages/metrics/src/faithfulness.ts @@ -3,6 +3,7 @@ import type { FaithfulnessResult, StatementSupport, } from '@reaatech/rag-eval-core'; +import { getSignificantWords, splitSentences } from './text-utils.js'; /** * Faithfulness Scorer @@ -52,11 +53,7 @@ export class FaithfulnessScorer { * Uses simple sentence splitting - can be enhanced with NLP */ private extractStatements(text: string): string[] { - // Split by sentence-ending punctuation, filter empty strings - const sentences = text - .split(/(?<=[.!?])\s+/) - .map((s) => s.trim()) - .filter((s) => s.length > 0); + const sentences = splitSentences(text); // If no sentences found, treat the whole text as one statement if (sentences.length === 0 && text.trim().length > 0) { @@ -87,8 +84,8 @@ export class FaithfulnessScorer { } // Check for keyword overlap - const statementWords = this.getSignificantWords(statementLower); - const contextWords = new Set(this.getSignificantWords(contextLower)); + const statementWords = getSignificantWords(statementLower); + const contextWords = new Set(getSignificantWords(contextLower)); if (statementWords.length === 0) { return { @@ -113,172 +110,6 @@ export class FaithfulnessScorer { }; } - /** - * Get significant words (nouns, verbs) from text - * Filters out stop words - */ - private getSignificantWords(text: string): string[] { - const stopWords = new Set([ - 'a', - 'an', - 'the', - 'is', - 'are', - 'was', - 'were', - 'be', - 'been', - 'being', - 'have', - 'has', - 'had', - 'do', - 'does', - 'did', - 'will', - 'would', - 'could', - 'should', - 'may', - 'might', - 'shall', - 'can', - 'to', - 'of', - 'in', - 'for', - 'on', - 'with', - 'at', - 'by', - 'from', - 'as', - 'into', - 'through', - 'during', - 'before', - 'after', - 'above', - 'below', - 'between', - 'out', - 'off', - 'over', - 'under', - 'again', - 'further', - 'then', - 'once', - 'here', - 'there', - 'when', - 'where', - 'why', - 'how', - 'all', - 'each', - 'every', - 'both', - 'few', - 'more', - 'most', - 'other', - 'some', - 'such', - 'no', - 'nor', - 'not', - 'only', - 'own', - 'same', - 'so', - 'than', - 'too', - 'very', - 'just', - 'and', - 'but', - 'or', - 'if', - 'while', - 'because', - 'until', - 'about', - 'against', - 'up', - 'down', - 'it', - 'its', - 'i', - 'me', - 'my', - 'myself', - 'we', - 'our', - 'ours', - 'ourselves', - 'you', - 'your', - 'yours', - 'yourself', - 'yourselves', - 'he', - 'him', - 'his', - 'himself', - 'she', - 'her', - 'hers', - 'herself', - 'they', - 'them', - 'their', - 'theirs', - 'themselves', - 'what', - 'which', - 'who', - 'whom', - 'this', - 'that', - 'these', - 'those', - 'am', - 'been', - 'being', - 'have', - 'has', - 'had', - 'having', - 'do', - 'does', - 'did', - 'doing', - 'would', - 'should', - 'could', - 'ought', - ]); - - // Extract words and numbers, filter stop words and short words - const words = text.match(/[a-z]+/g) ?? []; - const numbers = text.match(/\d{2,}/g) ?? []; - const allTokens = [...words, ...numbers].map((w) => this.normalizeWord(w)); - return allTokens.filter((word) => word.length > 2 && !stopWords.has(word)); - } - - /** - * Normalize a word by stripping common English inflections - */ - private normalizeWord(word: string): string { - const len = word.length; - if (len <= 3) return word; - if (word.endsWith('ing') && len > 4) return word.slice(0, -3); - if (word.endsWith('ed') && len > 4) return word.slice(0, -2); - if (word.endsWith('s') && !word.endsWith('ss') && len > 3) return word.slice(0, -1); - return word; - } - /** * Generate explanation for the faithfulness score */ diff --git a/packages/metrics/src/index.ts b/packages/metrics/src/index.ts index 48f84fb..2a77535 100644 --- a/packages/metrics/src/index.ts +++ b/packages/metrics/src/index.ts @@ -2,8 +2,23 @@ * Metrics module exports */ +export { + AnswerCorrectnessScorer, + type AnswerCorrectnessScorerOptions, +} from './answer-correctness.js'; export { ContextPrecisionScorer } from './context-precision.js'; export { ContextRecallScorer } from './context-recall.js'; export { MetricsEngine } from './engine.js'; export { FaithfulnessScorer } from './faithfulness.js'; -export { RelevanceScorer } from './relevance.js'; +export { RelevanceScorer, type RelevanceScorerOptions } from './relevance.js'; +export { RetrievalScorer, type RetrievalScorerOptions } from './retrieval.js'; +export { + cosineSimilarity, + diceCoefficient, + getBigrams, + getSignificantWords, + getWords, + normalizeWord, + STOP_WORDS, + splitSentences, +} from './text-utils.js'; diff --git a/packages/metrics/src/relevance.ts b/packages/metrics/src/relevance.ts index e7d9ac7..159480c 100644 --- a/packages/metrics/src/relevance.ts +++ b/packages/metrics/src/relevance.ts @@ -1,33 +1,70 @@ -import type { EvaluationSample, RelevanceResult } from '@reaatech/rag-eval-core'; +import type { EmbeddingProvider, EvaluationSample, RelevanceResult } from '@reaatech/rag-eval-core'; +import { + cosineSimilarity, + diceCoefficient, + getBigrams, + getSignificantWords, + getWords, + normalizeWord, +} from './text-utils.js'; + +/** Options for the relevance scorer. */ +export interface RelevanceScorerOptions { + /** + * Optional embedding provider. When supplied, true semantic similarity + * (cosine of embeddings) is computed and used as the primary signal, and + * exposed as `semantic_similarity`. Without it, scoring is purely lexical + * and only `lexical_similarity` is populated. + */ + embeddingProvider?: EmbeddingProvider; +} /** * Relevance Scorer * - * Measures whether a RAG system's generated answer actually addresses the user's query. - * Assesses semantic similarity and intent coverage. + * Measures whether a RAG system's generated answer actually addresses the + * user's query. By default this uses lexical heuristics (word/character + * overlap + intent coverage) — fast and dependency-free, but blind to + * paraphrase. Supply an `embeddingProvider` for true semantic similarity. */ export class RelevanceScorer { + private embeddingProvider?: EmbeddingProvider; + + constructor(options: RelevanceScorerOptions = {}) { + this.embeddingProvider = options.embeddingProvider; + } + /** * Score relevance for a single sample */ async score(sample: EvaluationSample): Promise { const { query, generated_answer } = sample; - // Calculate semantic similarity - const semanticSimilarity = this.calculateSemanticSimilarity(query, generated_answer); + // Lexical similarity is always computed (cheap, deterministic). + const lexicalSimilarity = this.calculateLexicalSimilarity(query, generated_answer); - // Calculate intent coverage + // Semantic similarity only when an embedding provider is configured. + const semanticSimilarity = this.embeddingProvider + ? await this.calculateSemanticSimilarity(query, generated_answer) + : undefined; + + // Intent coverage: does the answer address the parts of the query? const intentScore = this.calculateIntentCoverage(query, generated_answer); - // Weighted combination: 60% semantic, 40% intent - const score = Math.round((semanticSimilarity * 0.6 + intentScore * 0.4) * 1000) / 1000; + // Prefer the semantic signal when available, else fall back to lexical. + const primarySimilarity = semanticSimilarity ?? lexicalSimilarity; + const score = Math.round((primarySimilarity * 0.6 + intentScore * 0.4) * 1000) / 1000; - return { + const result: RelevanceResult = { score, - semantic_similarity: Math.round(semanticSimilarity * 1000) / 1000, + lexical_similarity: Math.round(lexicalSimilarity * 1000) / 1000, intent_score: Math.round(intentScore * 1000) / 1000, - explanation: this.generateExplanation(score, semanticSimilarity, intentScore), + explanation: this.generateExplanation(score, primarySimilarity, intentScore), }; + if (semanticSimilarity !== undefined) { + result.semantic_similarity = Math.round(semanticSimilarity * 1000) / 1000; + } + return result; } /** @@ -38,9 +75,21 @@ export class RelevanceScorer { } /** - * Calculate semantic similarity using word overlap and character n-gram similarity + * True semantic similarity via embeddings (cosine of query/answer vectors). + */ + private async calculateSemanticSimilarity(query: string, answer: string): Promise { + if (!this.embeddingProvider) return 0; + const [queryVec, answerVec] = await this.embeddingProvider.embed([query, answer]); + if (!queryVec || !answerVec) return 0; + // Cosine can be negative; clamp to [0, 1] to stay on the metric scale. + return Math.max(0, Math.min(1, cosineSimilarity(queryVec, answerVec))); + } + + /** + * Lexical similarity using word overlap (Jaccard) and character bigram + * similarity (Dice). Surface-form only — does not capture paraphrase. */ - private calculateSemanticSimilarity(query: string, answer: string): number { + private calculateLexicalSimilarity(query: string, answer: string): number { const queryLower = query.toLowerCase(); const answerLower = answer.toLowerCase(); @@ -50,8 +99,8 @@ export class RelevanceScorer { } // Word-level Jaccard similarity - const queryWords = new Set(this.getWords(queryLower)); - const answerWords = new Set(this.getWords(answerLower)); + const queryWords = new Set(getWords(queryLower)); + const answerWords = new Set(getWords(answerLower)); const intersection = [...queryWords].filter((w) => answerWords.has(w)); const union = new Set([...queryWords, ...answerWords]); @@ -59,9 +108,9 @@ export class RelevanceScorer { const jaccard = union.size > 0 ? intersection.length / union.size : 0; // Character bigram similarity (Dice coefficient) - const queryBigrams = this.getBigrams(queryLower); - const answerBigrams = this.getBigrams(answerLower); - const bigramSimilarity = this.diceCoefficient(queryBigrams, answerBigrams); + const queryBigrams = getBigrams(queryLower); + const answerBigrams = getBigrams(answerLower); + const bigramSimilarity = diceCoefficient(queryBigrams, answerBigrams); // Weighted average of word and character similarity return jaccard * 0.4 + bigramSimilarity * 0.6; @@ -90,21 +139,21 @@ export class RelevanceScorer { if (!isQuestion) { // For non-questions, use simple keyword matching - const queryWords = this.getSignificantWords(queryLower); - const answerWords = new Set(this.getWords(answerLower).map((w) => this.normalizeWord(w))); + const queryWords = getSignificantWords(queryLower); + const answerWords = new Set(getWords(answerLower).map((w) => normalizeWord(w))); const matchedWords = queryWords.filter((w) => answerWords.has(w)); return queryWords.length > 0 ? matchedWords.length / queryWords.length : 0.5; } // Extract key topics/entities from query - const queryWords = this.getSignificantWords(queryLower); + const queryWords = getSignificantWords(queryLower); if (queryWords.length === 0) { return 0.5; } // Check how many query topics are addressed in the answer - const answerWords = new Set(this.getWords(answerLower).map((w) => this.normalizeWord(w))); + const answerWords = new Set(getWords(answerLower).map((w) => normalizeWord(w))); const matchedTopics = queryWords.filter((w) => answerWords.has(w)); // Also check for synonyms/common responses @@ -117,191 +166,6 @@ export class RelevanceScorer { return Math.min(1, topicCoverage + bonusScore); } - /** - * Get words from text - */ - private getWords(text: string): string[] { - return text.match(/[a-z]+/g) ?? []; - } - - /** - * Get significant words (filtering stop words) - */ - private getSignificantWords(text: string): string[] { - const stopWords = new Set([ - 'a', - 'an', - 'the', - 'is', - 'are', - 'was', - 'were', - 'be', - 'been', - 'being', - 'have', - 'has', - 'had', - 'do', - 'does', - 'did', - 'will', - 'would', - 'could', - 'should', - 'may', - 'might', - 'shall', - 'can', - 'to', - 'of', - 'in', - 'for', - 'on', - 'with', - 'at', - 'by', - 'from', - 'as', - 'into', - 'through', - 'during', - 'before', - 'after', - 'above', - 'below', - 'between', - 'out', - 'off', - 'over', - 'under', - 'again', - 'further', - 'then', - 'once', - 'here', - 'there', - 'when', - 'where', - 'why', - 'how', - 'all', - 'each', - 'every', - 'both', - 'few', - 'more', - 'most', - 'other', - 'some', - 'such', - 'no', - 'nor', - 'not', - 'only', - 'own', - 'same', - 'so', - 'than', - 'too', - 'very', - 'just', - 'and', - 'but', - 'or', - 'if', - 'while', - 'because', - 'until', - 'about', - 'against', - 'up', - 'down', - 'it', - 'its', - 'i', - 'me', - 'my', - 'myself', - 'we', - 'our', - 'ours', - 'ourselves', - 'you', - 'your', - 'yours', - 'yourself', - 'yourselves', - 'he', - 'him', - 'his', - 'himself', - 'she', - 'her', - 'hers', - 'herself', - 'they', - 'them', - 'their', - 'theirs', - 'themselves', - 'what', - 'which', - 'who', - 'whom', - 'this', - 'that', - 'these', - 'those', - ]); - - const words = this.getWords(text); - const numbers = text.match(/\d{2,}/g) ?? []; - const allTokens = [...words, ...numbers].map((w) => this.normalizeWord(w)); - return allTokens.filter((word) => word.length > 2 && !stopWords.has(word)); - } - - /** - * Normalize a word by stripping common English inflections - */ - private normalizeWord(word: string): string { - const len = word.length; - if (len <= 3) return word; - if (word.endsWith('ing') && len > 4) return word.slice(0, -3); - if (word.endsWith('ed') && len > 4) return word.slice(0, -2); - if (word.endsWith('s') && !word.endsWith('ss') && len > 3) return word.slice(0, -1); - return word; - } - - /** - * Get character bigrams from text - */ - private getBigrams(text: string): Set { - const bigrams = new Set(); - for (let i = 0; i < text.length - 1; i++) { - bigrams.add(text.substring(i, i + 2)); - } - return bigrams; - } - - /** - * Calculate Dice coefficient between two sets of bigrams - */ - private diceCoefficient(bigrams1: Set, bigrams2: Set): number { - if (bigrams1.size === 0 || bigrams2.size === 0) { - return 0; - } - - let intersection = 0; - for (const bigram of bigrams1) { - if (bigrams2.has(bigram)) { - intersection++; - } - } - - return (2 * intersection) / (bigrams1.size + bigrams2.size); - } - /** * Check if text contains action words (indicating a helpful response) */ @@ -334,7 +198,7 @@ export class RelevanceScorer { 'follow', ]); - const words = this.getWords(text); + const words = getWords(text); return words.some((word) => actionWords.has(word)); } @@ -343,18 +207,18 @@ export class RelevanceScorer { */ private generateExplanation( score: number, - semanticSimilarity: number, + primarySimilarity: number, intentScore: number, ): string { if (score >= 0.8) { - return `Answer is highly relevant (semantic: ${semanticSimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; + return `Answer is highly relevant (similarity: ${primarySimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; } if (score >= 0.6) { - return `Answer is moderately relevant (semantic: ${semanticSimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; + return `Answer is moderately relevant (similarity: ${primarySimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; } if (score >= 0.4) { - return `Answer may not fully address the query (semantic: ${semanticSimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; + return `Answer may not fully address the query (similarity: ${primarySimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; } - return `Answer appears irrelevant to the query (semantic: ${semanticSimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; + return `Answer appears irrelevant to the query (similarity: ${primarySimilarity.toFixed(2)}, intent: ${intentScore.toFixed(2)})`; } } diff --git a/packages/metrics/src/retrieval.ts b/packages/metrics/src/retrieval.ts new file mode 100644 index 0000000..5d4f5de --- /dev/null +++ b/packages/metrics/src/retrieval.ts @@ -0,0 +1,110 @@ +import type { EvaluationSample, RetrievalResult } from '@reaatech/rag-eval-core'; + +/** Options for the retrieval scorer. */ +export interface RetrievalScorerOptions { + /** Cutoff rank for @k metrics. Defaults to the number of retrieved chunks. */ + k?: number; +} + +/** + * Retrieval Scorer + * + * Standard information-retrieval ranking metrics — MRR, nDCG, precision@k, + * recall@k, and hit@k — computed from the ranked `retrieved_chunk_ids` and the + * ground-truth `relevant_chunk_ids` on a sample. These require chunk-ID labels + * and complement the text-based context precision/recall scorers. + */ +export class RetrievalScorer { + private defaultK?: number; + + constructor(options: RetrievalScorerOptions = {}) { + this.defaultK = options.k; + } + + /** + * Score retrieval ranking for a single sample. + * Async to match the other scorers' interface (no awaited work here). + */ + async score(sample: EvaluationSample): Promise { + const retrieved = sample.retrieved_chunk_ids ?? []; + const relevant = new Set(sample.relevant_chunk_ids ?? []); + const k = Math.max(0, this.defaultK ?? retrieved.length); + const topK = retrieved.slice(0, k); + + const mrr = this.reciprocalRank(retrieved, relevant); + const ndcg = this.ndcg(retrieved, relevant); + const relevantInTopK = topK.filter((id) => relevant.has(id)).length; + const precisionAtK = topK.length > 0 ? relevantInTopK / topK.length : 0; + const recallAtK = relevant.size > 0 ? relevantInTopK / relevant.size : 0; + const hitAtK = relevantInTopK > 0 ? 1 : 0; + + return { + mrr: round(mrr), + ndcg: round(ndcg), + precision_at_k: round(precisionAtK), + recall_at_k: round(recallAtK), + hit_at_k: hitAtK, + k, + explanation: this.explain(relevant.size, retrieved.length, k, relevantInTopK), + }; + } + + /** + * Score retrieval ranking for multiple samples. + */ + async scoreBatch(samples: EvaluationSample[]): Promise { + return Promise.all(samples.map((sample) => this.score(sample))); + } + + /** Reciprocal rank of the first relevant chunk. */ + private reciprocalRank(retrieved: string[], relevant: Set): number { + for (let i = 0; i < retrieved.length; i++) { + const id = retrieved[i]; + if (id !== undefined && relevant.has(id)) { + return 1 / (i + 1); + } + } + return 0; + } + + /** Binary-relevance nDCG over the retrieved ranking. */ + private ndcg(retrieved: string[], relevant: Set): number { + if (retrieved.length === 0 || relevant.size === 0) return 0; + + let dcg = 0; + for (let i = 0; i < retrieved.length; i++) { + const id = retrieved[i]; + if (id !== undefined && relevant.has(id)) { + dcg += 1 / Math.log2(i + 2); + } + } + + // Ideal DCG: all relevant chunks ranked first (capped at list length). + const idealHits = Math.min(relevant.size, retrieved.length); + let idcg = 0; + for (let i = 0; i < idealHits; i++) { + idcg += 1 / Math.log2(i + 2); + } + + return idcg === 0 ? 0 : dcg / idcg; + } + + private explain( + relevantCount: number, + retrievedCount: number, + k: number, + relevantInTopK: number, + ): string { + if (retrievedCount === 0) { + return 'No retrieved_chunk_ids provided — retrieval metrics are 0'; + } + if (relevantCount === 0) { + return 'No relevant_chunk_ids provided — cannot judge retrieval relevance'; + } + return `${relevantInTopK}/${k} top-${k} chunks relevant (${relevantCount} relevant overall, ${retrievedCount} retrieved)`; + } +} + +function round(value: number): number { + return Math.round(value * 1000) / 1000; +} diff --git a/packages/metrics/src/text-utils.ts b/packages/metrics/src/text-utils.ts new file mode 100644 index 0000000..68f107b --- /dev/null +++ b/packages/metrics/src/text-utils.ts @@ -0,0 +1,219 @@ +/** + * Shared text utilities for heuristic (lexical) metric scorers. + * + * These helpers were previously duplicated across the faithfulness, relevance, + * context-precision, and context-recall scorers. They are intentionally + * lightweight and English-oriented — see the README for the limitations of + * lexical scoring and when to prefer the LLM judge or an embedding provider. + */ + +/** English stop words filtered out before computing lexical overlap. */ +export const STOP_WORDS: ReadonlySet = new Set([ + 'a', + 'an', + 'the', + 'is', + 'are', + 'was', + 'were', + 'be', + 'been', + 'being', + 'have', + 'has', + 'had', + 'having', + 'do', + 'does', + 'did', + 'doing', + 'will', + 'would', + 'could', + 'should', + 'may', + 'might', + 'shall', + 'can', + 'to', + 'of', + 'in', + 'for', + 'on', + 'with', + 'at', + 'by', + 'from', + 'as', + 'into', + 'through', + 'during', + 'before', + 'after', + 'above', + 'below', + 'between', + 'out', + 'off', + 'over', + 'under', + 'again', + 'further', + 'then', + 'once', + 'here', + 'there', + 'when', + 'where', + 'why', + 'how', + 'all', + 'each', + 'every', + 'both', + 'few', + 'more', + 'most', + 'other', + 'some', + 'such', + 'no', + 'nor', + 'not', + 'only', + 'own', + 'same', + 'so', + 'than', + 'too', + 'very', + 'just', + 'and', + 'but', + 'or', + 'if', + 'while', + 'because', + 'until', + 'about', + 'against', + 'up', + 'down', + 'it', + 'its', + 'i', + 'me', + 'my', + 'myself', + 'we', + 'our', + 'ours', + 'ourselves', + 'you', + 'your', + 'yours', + 'yourself', + 'yourselves', + 'he', + 'him', + 'his', + 'himself', + 'she', + 'her', + 'hers', + 'herself', + 'they', + 'them', + 'their', + 'theirs', + 'themselves', + 'what', + 'which', + 'who', + 'whom', + 'this', + 'that', + 'these', + 'those', + 'am', + 'ought', +]); + +/** Extract lowercase alphabetic word tokens from text. */ +export function getWords(text: string): string[] { + return text.match(/[a-z]+/g) ?? []; +} + +/** + * Normalize a word by stripping common English inflections. + * A deliberately cheap stemmer — good enough for keyword-overlap heuristics. + */ +export function normalizeWord(word: string): string { + const len = word.length; + if (len <= 3) return word; + if (word.endsWith('ing') && len > 4) return word.slice(0, -3); + if (word.endsWith('ed') && len > 4) return word.slice(0, -2); + if (word.endsWith('s') && !word.endsWith('ss') && len > 3) return word.slice(0, -1); + return word; +} + +/** + * Get significant words from text: alphabetic tokens and multi-digit numbers, + * normalized, with stop words and very short tokens removed. + */ +export function getSignificantWords(text: string): string[] { + const words = getWords(text); + const numbers = text.match(/\d{2,}/g) ?? []; + const allTokens = [...words, ...numbers].map((w) => normalizeWord(w)); + return allTokens.filter((word) => word.length > 2 && !STOP_WORDS.has(word)); +} + +/** Split text into sentence-like statements, preserving terminal punctuation. */ +export function splitSentences(text: string): string[] { + return text + .split(/(?<=[.!?])\s+/) + .map((s) => s.trim()) + .filter((s) => s.length > 0); +} + +/** Build the set of character bigrams for a string. */ +export function getBigrams(text: string): Set { + const bigrams = new Set(); + for (let i = 0; i < text.length - 1; i++) { + bigrams.add(text.substring(i, i + 2)); + } + return bigrams; +} + +/** Sørensen–Dice coefficient between two bigram sets (0–1). */ +export function diceCoefficient(bigrams1: Set, bigrams2: Set): number { + if (bigrams1.size === 0 || bigrams2.size === 0) { + return 0; + } + + let intersection = 0; + for (const bigram of bigrams1) { + if (bigrams2.has(bigram)) { + intersection++; + } + } + + return (2 * intersection) / (bigrams1.size + bigrams2.size); +} + +/** Cosine similarity between two equal-length numeric vectors (0–1 for non-negative inputs). */ +export function cosineSimilarity(a: number[], b: number[]): number { + if (a.length === 0 || a.length !== b.length) return 0; + let dot = 0; + let normA = 0; + let normB = 0; + for (let i = 0; i < a.length; i++) { + const x = a[i] ?? 0; + const y = b[i] ?? 0; + dot += x * y; + normA += x * x; + normB += y * y; + } + if (normA === 0 || normB === 0) return 0; + return dot / (Math.sqrt(normA) * Math.sqrt(normB)); +} diff --git a/packages/metrics/tests/answer-correctness.test.ts b/packages/metrics/tests/answer-correctness.test.ts new file mode 100644 index 0000000..5f5498e --- /dev/null +++ b/packages/metrics/tests/answer-correctness.test.ts @@ -0,0 +1,50 @@ +import type { EmbeddingProvider, EvaluationSample } from '@reaatech/rag-eval-core'; +import { AnswerCorrectnessScorer } from '@reaatech/rag-eval-metrics'; +import { describe, expect, it } from 'vitest'; + +function makeSample(overrides: Partial = {}): EvaluationSample { + return { + query: 'What is the refund window?', + context: ['c'], + ground_truth: 'Refunds are available within 14 days of purchase.', + generated_answer: 'You can get a refund within 14 days of purchase.', + ...overrides, + }; +} + +describe('AnswerCorrectnessScorer', () => { + it('scores high when the answer matches the ground truth', async () => { + const scorer = new AnswerCorrectnessScorer(); + const result = await scorer.score(makeSample()); + expect(result.score).toBeGreaterThan(0.5); + expect(result.lexical_similarity).toBeDefined(); + expect(result.semantic_similarity).toBeUndefined(); + }); + + it('scores low when the answer diverges from the ground truth', async () => { + const scorer = new AnswerCorrectnessScorer(); + const result = await scorer.score( + makeSample({ generated_answer: 'The store opens at 9am on weekdays.' }), + ); + expect(result.score).toBeLessThan(0.5); + }); + + it('uses an embedding provider for semantic similarity when supplied', async () => { + const identity: EmbeddingProvider = { + embed: async (texts) => texts.map(() => [1, 0, 0]), + }; + const scorer = new AnswerCorrectnessScorer({ embeddingProvider: identity }); + const result = await scorer.score( + makeSample({ generated_answer: 'totally different surface text' }), + ); + // Identical vectors → cosine 1, so semantic dominates the lexical mismatch. + expect(result.semantic_similarity).toBe(1); + expect(result.score).toBe(1); + }); + + it('scores a batch', async () => { + const scorer = new AnswerCorrectnessScorer(); + const results = await scorer.scoreBatch([makeSample(), makeSample()]); + expect(results).toHaveLength(2); + }); +}); diff --git a/packages/metrics/tests/metrics.test.ts b/packages/metrics/tests/metrics.test.ts index 6142b32..94c8002 100644 --- a/packages/metrics/tests/metrics.test.ts +++ b/packages/metrics/tests/metrics.test.ts @@ -62,7 +62,8 @@ describe('Metrics', () => { const result = await scorer.score(sample); expect(result.score).toBeGreaterThanOrEqual(0); expect(result.score).toBeLessThanOrEqual(1); - expect(result.semantic_similarity).toBeDefined(); + expect(result.lexical_similarity).toBeDefined(); + expect(result.semantic_similarity).toBeUndefined(); expect(result.intent_score).toBeDefined(); }); diff --git a/packages/metrics/tests/relevance.test.ts b/packages/metrics/tests/relevance.test.ts index 28decd0..965ffce 100644 --- a/packages/metrics/tests/relevance.test.ts +++ b/packages/metrics/tests/relevance.test.ts @@ -1,4 +1,4 @@ -import type { EvaluationSample } from '@reaatech/rag-eval-core'; +import type { EmbeddingProvider, EvaluationSample } from '@reaatech/rag-eval-core'; import { RelevanceScorer } from '@reaatech/rag-eval-metrics'; import { describe, expect, it } from 'vitest'; @@ -24,7 +24,7 @@ describe('RelevanceScorer', () => { ); expect(result.score).toBeGreaterThan(0.5); - expect(result.semantic_similarity).toBeDefined(); + expect(result.lexical_similarity).toBeDefined(); expect(result.intent_score).toBeDefined(); }); @@ -49,7 +49,7 @@ describe('RelevanceScorer', () => { }), ); - expect(result.semantic_similarity).toBe(1); + expect(result.lexical_similarity).toBe(1); }); it('handles query that contains the answer verbatim', async () => { @@ -61,7 +61,7 @@ describe('RelevanceScorer', () => { }), ); - expect(result.semantic_similarity).toBe(1); + expect(result.lexical_similarity).toBe(1); }); it('returns scores rounded to 3 decimal places', async () => { @@ -71,11 +71,37 @@ describe('RelevanceScorer', () => { const decimals = (s: number) => s.toString().includes('.') ? (s.toString().split('.')[1] ?? '').length : 0; expect(decimals(result.score)).toBeLessThanOrEqual(3); - expect(decimals(result.semantic_similarity ?? 0)).toBeLessThanOrEqual(3); + expect(decimals(result.lexical_similarity ?? 0)).toBeLessThanOrEqual(3); expect(decimals(result.intent_score ?? 0)).toBeLessThanOrEqual(3); }); }); + describe('with an embedding provider', () => { + it('populates semantic_similarity from embeddings', async () => { + const identity: EmbeddingProvider = { + embed: async (texts) => texts.map(() => [1, 0, 0]), + }; + const scorer = new RelevanceScorer({ embeddingProvider: identity }); + const result = await scorer.score(makeSample({ query: 'apple', generated_answer: 'orange' })); + // Identical vectors → cosine 1, regardless of lexical mismatch. + expect(result.semantic_similarity).toBe(1); + expect(result.lexical_similarity).toBeDefined(); + }); + + it('still computes lexical_similarity alongside semantic', async () => { + const orthogonal: EmbeddingProvider = { + embed: async () => [ + [1, 0], + [0, 1], + ], + }; + const scorer = new RelevanceScorer({ embeddingProvider: orthogonal }); + const result = await scorer.score(makeSample()); + expect(result.semantic_similarity).toBe(0); // orthogonal vectors + expect(result.lexical_similarity).toBeGreaterThanOrEqual(0); + }); + }); + describe('scoreBatch', () => { it('scores multiple samples', async () => { const scorer = new RelevanceScorer(); @@ -98,8 +124,8 @@ describe('RelevanceScorer', () => { }), ); - expect(result.semantic_similarity).toBeGreaterThanOrEqual(0); - expect(result.semantic_similarity).toBeLessThanOrEqual(1); + expect(result.lexical_similarity).toBeGreaterThanOrEqual(0); + expect(result.lexical_similarity).toBeLessThanOrEqual(1); }); it('handles empty query words', async () => { @@ -111,7 +137,7 @@ describe('RelevanceScorer', () => { }), ); - expect(result.semantic_similarity).toBeDefined(); + expect(result.lexical_similarity).toBeDefined(); }); }); @@ -323,7 +349,7 @@ describe('RelevanceScorer', () => { }), ); - expect(result.semantic_similarity).toBeDefined(); + expect(result.lexical_similarity).toBeDefined(); }); }); diff --git a/packages/metrics/tests/retrieval.test.ts b/packages/metrics/tests/retrieval.test.ts new file mode 100644 index 0000000..092fe63 --- /dev/null +++ b/packages/metrics/tests/retrieval.test.ts @@ -0,0 +1,74 @@ +import type { EvaluationSample } from '@reaatech/rag-eval-core'; +import { RetrievalScorer } from '@reaatech/rag-eval-metrics'; +import { describe, expect, it } from 'vitest'; + +function makeSample(overrides: Partial = {}): EvaluationSample { + return { + query: 'q', + context: ['c'], + ground_truth: 'gt', + generated_answer: 'a', + ...overrides, + }; +} + +describe('RetrievalScorer', () => { + it('computes MRR, nDCG, precision/recall/hit@k from chunk ids', async () => { + const scorer = new RetrievalScorer(); + const result = await scorer.score( + makeSample({ + retrieved_chunk_ids: ['a', 'b', 'c'], + relevant_chunk_ids: ['b'], + }), + ); + + expect(result.k).toBe(3); + expect(result.mrr).toBeCloseTo(0.5, 3); // first relevant at rank 2 + expect(result.ndcg).toBeCloseTo(0.631, 2); + expect(result.precision_at_k).toBeCloseTo(1 / 3, 3); + expect(result.recall_at_k).toBe(1); + expect(result.hit_at_k).toBe(1); + }); + + it('rewards relevant chunks ranked higher (nDCG)', async () => { + const scorer = new RetrievalScorer(); + const high = await scorer.score( + makeSample({ retrieved_chunk_ids: ['x', 'y', 'z'], relevant_chunk_ids: ['x'] }), + ); + const low = await scorer.score( + makeSample({ retrieved_chunk_ids: ['x', 'y', 'z'], relevant_chunk_ids: ['z'] }), + ); + expect(high.ndcg).toBeGreaterThan(low.ndcg); + expect(high.mrr).toBeGreaterThan(low.mrr); + }); + + it('honors an explicit k cutoff', async () => { + const scorer = new RetrievalScorer({ k: 1 }); + const result = await scorer.score( + makeSample({ retrieved_chunk_ids: ['a', 'b'], relevant_chunk_ids: ['b'] }), + ); + expect(result.k).toBe(1); + expect(result.hit_at_k).toBe(0); // relevant 'b' is outside top-1 + expect(result.precision_at_k).toBe(0); + }); + + it('returns zeros when no ids are provided', async () => { + const scorer = new RetrievalScorer(); + const result = await scorer.score(makeSample()); + expect(result.mrr).toBe(0); + expect(result.ndcg).toBe(0); + expect(result.hit_at_k).toBe(0); + expect(result.explanation).toContain('No retrieved_chunk_ids'); + }); + + it('scores a batch', async () => { + const scorer = new RetrievalScorer(); + const results = await scorer.scoreBatch([ + makeSample({ retrieved_chunk_ids: ['a'], relevant_chunk_ids: ['a'] }), + makeSample({ retrieved_chunk_ids: ['a'], relevant_chunk_ids: ['b'] }), + ]); + expect(results).toHaveLength(2); + expect(results[0]?.hit_at_k).toBe(1); + expect(results[1]?.hit_at_k).toBe(0); + }); +}); diff --git a/pnpm-lock.yaml b/pnpm-lock.yaml index 193c938..847553e 100644 --- a/pnpm-lock.yaml +++ b/pnpm-lock.yaml @@ -246,12 +246,6 @@ importers: '@reaatech/rag-eval-core': specifier: workspace:* version: link:../core - compromise: - specifier: ^14.15.0 - version: 14.15.0 - natural: - specifier: ^8.1.1 - version: 8.1.1(@opentelemetry/api@1.9.1) p-limit: specifier: ^7.3.0 version: 7.3.0 @@ -762,9 +756,6 @@ packages: '@cfworker/json-schema': optional: true - '@mongodb-js/saslprep@1.4.10': - resolution: {integrity: sha512-DDb3OAw8ezai9p2i1F7R5wMVmyrMIuT8ixjV56R4Hl4cazCo2tOMTDyPR5rrCUcHiMbWHzCXOdife5OD6H1C8w==} - '@nodelib/fs.scandir@2.1.5': resolution: {integrity: sha512-vq24Bq3ym5HEQm2NKCr3yXDwjc7vTsEThRDnkp2DK9p1uqLR+DHurm/NOTo0KG7HYHU7eppKZj3MyqYuMBf62g==} engines: {node: '>= 8'} @@ -832,42 +823,6 @@ packages: resolution: {integrity: sha512-+1VkjdD0QBLPodGrJUeqarH8VAIvQODIbwh9XpP5Syisf7YoQgsJKPNFoqqLQlu+VQ/tVSshMR6loPMn8U+dPg==} engines: {node: '>=14'} - '@redis/bloom@5.12.1': - resolution: {integrity: sha512-PUUfv+ms7jgPSBVoo/DN4AkPHj4D5TZSd6SbJX7egzBplkYUcKmHRE8RKia7UtZ8bSQbLguLvxVO+asKtQfZWA==} - engines: {node: '>= 18.19.0'} - peerDependencies: - '@redis/client': ^5.12.1 - - '@redis/client@5.12.1': - resolution: {integrity: sha512-7aPGWeqA3uFm43o19umzdl16CEjK/JQGtSXVPevplTaOU3VJA/rseBC1QvYUz9lLDIMBimc4SW/zrW4S89BaCA==} - engines: {node: '>= 18.19.0'} - peerDependencies: - '@node-rs/xxhash': ^1.1.0 - '@opentelemetry/api': '>=1 <2' - peerDependenciesMeta: - '@node-rs/xxhash': - optional: true - '@opentelemetry/api': - optional: true - - '@redis/json@5.12.1': - resolution: {integrity: sha512-eOze75esLve4vfqDel7aMX08CNaiLLQS2fV8mpRN9NxPe1rVR4vQyYiW/OgtGUysF6QOr9ANhfxABKNOJfXdKg==} - engines: {node: '>= 18.19.0'} - peerDependencies: - '@redis/client': ^5.12.1 - - '@redis/search@5.12.1': - resolution: {integrity: sha512-ItlxbxC9cKI6IU1TLWoczwJCRb6TdmkEpWv05UrPawqaAnWGRu3rcIqsc5vN483T2fSociuyV1UkWIL5I4//2w==} - engines: {node: '>= 18.19.0'} - peerDependencies: - '@redis/client': ^5.12.1 - - '@redis/time-series@5.12.1': - resolution: {integrity: sha512-c6JL6E3EcZJuNqKFz+KM+l9l5mpcQiKvTwgA3blt5glWJ8hjDk0yeHN3beE/MpqYIQ8UEX44ItQzgkE/gCBELQ==} - engines: {node: '>= 18.19.0'} - peerDependencies: - '@redis/client': ^5.12.1 - '@rollup/rollup-android-arm-eabi@4.60.2': resolution: {integrity: sha512-dnlp69efPPg6Uaw2dVqzWRfAWRnYVb1XJ8CyyhIbZeaq4CA5/mLeZ1IEt9QqQxmbdvagjLIm2ZL8BxXv5lH4Yw==} cpu: [arm] @@ -1041,12 +996,6 @@ packages: '@types/node@25.6.0': resolution: {integrity: sha512-+qIYRKdNYJwY3vRCZMdJbPLJAtGjQBudzZzdzwQYkEPQd+PJGixUL5QfvCLDaULoLv+RhT3LDkwEfKaAkgSmNQ==} - '@types/webidl-conversions@7.0.3': - resolution: {integrity: sha512-CiJJvcRtIgzadHCYXw7dqEnMNRjhGZlYK05Mj9OyktqV8uVT8fD2BFOB7S1uwBE3Kj2Z+4UyPmFw/Ixgw/LAlA==} - - '@types/whatwg-url@13.0.0': - resolution: {integrity: sha512-N8WXpbE6Wgri7KUSvrmQcqrMllKZ9uxkYWMt+mCSGwNc0Hsw9VQTW7ApqI4XNrx6/SaM2QQJCzMPDEXE058s+Q==} - '@vitest/coverage-istanbul@3.2.4': resolution: {integrity: sha512-IDlpuFJiWU9rhcKLkpzj8mFu/lpe64gVgnV15ZOrYx1iFzxxrxCzbExiUEKtwwXRvEiEMUS6iZeYgnMxgbqbxQ==} peerDependencies: @@ -1099,12 +1048,6 @@ packages: engines: {node: '>=0.4.0'} hasBin: true - afinn-165-financialmarketnews@3.0.0: - resolution: {integrity: sha512-0g9A1S3ZomFIGDTzZ0t6xmv4AuokBvBmpes8htiyHpH7N4xDmvSQL6UxL/Zcs2ypRb3VwgCscaD8Q3zEawKYhw==} - - afinn-165@2.0.2: - resolution: {integrity: sha512-mJ/RLUfpXfQA6bzugv+bBsc/QYkVrKaLYeS8fWBpKbTCsonv4iuV9ET0fgReEunm9vKLkaNgnekuSNlTC3WQ1Q==} - ajv-formats@3.0.1: resolution: {integrity: sha512-8iUql50EUR+uUcdRQ3HDqa6EVyo3docL8g5WJ3FNcWmu62IbkGUue/pEyLBW8VGKKucTPgqeks4fIU1DA4yowQ==} peerDependencies: @@ -1139,10 +1082,6 @@ packages: any-promise@1.3.0: resolution: {integrity: sha512-7UvmKalWRt1wgjL1RrGxoSJW/0QZFIegpeGvZG9kjp8vrRu55XTHbwnqq2GpXm9uLbcuhxm3IqX9OB4MZR1b2A==} - apparatus@0.0.10: - resolution: {integrity: sha512-KLy/ugo33KZA7nugtQ7O0E1c8kQ52N3IvD/XgIh4w/Nr28ypfkwDfA67F1ev4N1m5D+BOk1+b2dEJDfpj/VvZg==} - engines: {node: '>=0.2.6'} - argparse@1.0.10: resolution: {integrity: sha512-o5Roy6tNG4SL/FOkCAN6RzjiakZS25RLYFrcMttJqbdd8BWrnA+fGz57iN5Pb06pvBGvl5gQ0B48dJlslXvoTg==} @@ -1197,10 +1136,6 @@ packages: engines: {node: ^6 || ^7 || ^8 || ^9 || ^10 || ^11 || ^12 || >=13.7} hasBin: true - bson@7.2.0: - resolution: {integrity: sha512-YCEo7KjMlbNlyHhz7zAZNDpIpQbd+wOEHJYezv0nMYTn4x31eIUM2yomNNubclAt63dObUzKHWsBLJ9QcZNSnQ==} - engines: {node: '>=20.19.0'} - bundle-require@5.1.0: resolution: {integrity: sha512-3WrrOuZiyaaZPWiEt4G3+IffISVC9HYlWueJEBWED4ZH4aIAC2PnkdnuRrR94M+w6yGWn4AglWtJtBI8YqvgoA==} engines: {node: ^12.20.0 || ^14.13.1 || >=16.0.0} @@ -1249,10 +1184,6 @@ packages: resolution: {integrity: sha512-tRkV3HJ1ASwm19THiiLIXLO7Im7wlTuKnvkYaTkyoAPefqjNg7W7DHKUlGRxy9vxDvbyCYQkQozvptuMkGCg8A==} engines: {node: '>=4'} - cluster-key-slot@1.1.2: - resolution: {integrity: sha512-RMr0FhtfXemyinomL4hrWcYJxmX6deFdCxpJzhDttxgO1+bcCnkk+9drydLVDmAMG7NE6aN/fl4F7ucU/90gAA==} - engines: {node: '>=0.10.0'} - color-convert@2.0.1: resolution: {integrity: sha512-RRECPsj7iu/xb5oKYcsFHSppFNnsj/52OVTRKb4zP5onXwVF3zVmmToNcOfGC+CRDpfK/U584fMg38ZHCaElKQ==} engines: {node: '>=7.0.0'} @@ -1271,10 +1202,6 @@ packages: resolution: {integrity: sha512-NOKm8xhkzAjzFx8B2v5OAHT+u5pRQc2UCa2Vq9jYL/31o2wi9mxBA7LIFs3sV5VSC49z6pEhfbMULvShKj26WA==} engines: {node: '>= 6'} - compromise@14.15.0: - resolution: {integrity: sha512-YEMv5JGWyqRJw5hdZqDVQF3MMlHA6TRiXreR8IYffk6xB7GA5p/8DeDzvg0Jy2tHNGpD+qJGl0+oJwA+5R/sVA==} - engines: {node: '>=12.0.0'} - confbox@0.1.8: resolution: {integrity: sha512-RMtmw0iFkeR4YV+fUOSucriAQNb9g8zFR52MWCtl+cCZOFRNL6zeB395vPzFhEjjn4fMxXudmELnl/KF/WrK6w==} @@ -1340,10 +1267,6 @@ packages: resolution: {integrity: sha512-WkrWp9GR4KXfKGYzOLmTuGVi1UWFfws377n9cc55/tb6DuqyF6pcQ5AbiHEshaDpY9v6oaSr2XCDidGmMwdzIA==} engines: {node: '>=8'} - dotenv@17.4.2: - resolution: {integrity: sha512-nI4U3TottKAcAD9LLud4Cb7b2QztQMUEfHbvhTH09bqXTxnSie8WnjPALV/WMCrJZ6UV/qHJ6L03OqO3LcdYZw==} - engines: {node: '>=12'} - dotenv@8.6.0: resolution: {integrity: sha512-IrPdXQsk2BbzvCBGBOTmmSH5SodmqZNt4ERAZDmW4CT+tL8VtvinqywuANaFu4bOMWki16nqf0e4oC0QIaDr/g==} engines: {node: '>=10'} @@ -1358,10 +1281,6 @@ packages: ee-first@1.1.1: resolution: {integrity: sha512-WMwm9LhRUo+WUaRN+vRuETqG89IgZphVSNkdFgeb6sS/E4OrDIN7t48CAewSHXc6C8lefD8KKfr5vY61brQlow==} - efrt@2.7.0: - resolution: {integrity: sha512-/RInbCy1d4P6Zdfa+TMVsf/ufZVotat5hCw3QXmWtjU+3pFEOvOQ7ibo3aIxyCJw2leIeAMjmPj+1SLJiCpdrQ==} - engines: {node: '>=12.0.0'} - electron-to-chromium@1.5.360: resolution: {integrity: sha512-GkcBt6YYAw9SxFWn+xVar4cLVGlXVuswwtRLBozi2zp0GjXs4ZnOrqV4zbXzg35n7w81hCkyJNYicgXlVHAmBA==} @@ -1552,10 +1471,6 @@ packages: graceful-fs@4.2.11: resolution: {integrity: sha512-RbJ5/jmFcNNCcDV5o9eTnBLJ/HszWV0P73bc+Ff4nS/rJj+YaS6IGyiOL0VoBYX+l1Wrl3k63h/KrH+nhJ0XvQ==} - grad-school@0.0.5: - resolution: {integrity: sha512-rXunEHF9M9EkMydTBux7+IryYXEZinRk6g8OBOGDBzo/qWJjhTxy86i5q7lQYpCLHN8Sqv1XX3OIOc7ka2gtvQ==} - engines: {node: '>=8.0.0'} - has-flag@4.0.0: resolution: {integrity: sha512-EykJT/Q1KjTWctppgIAgfSO0tKVuZUjhgMr17kqTumMl6Afv3EISleU7qZUzoXDFTAHTDC4NOoG/ZxU3EvlMPQ==} engines: {node: '>=8'} @@ -1708,10 +1623,6 @@ packages: jsonfile@4.0.0: resolution: {integrity: sha512-m6F1R3z8jjlf2imQHS2Qez5sjKWQzbuuhuJ/FKYFRZvPE3PuHcSMVZzfsLhGVOkfd20obL5SWEBew5ShlquNxg==} - kareem@3.3.0: - resolution: {integrity: sha512-kpSuLD3/7RenBnjnJdOHXCKC8dTd1JzeOiJhN0necWWci6cC+qX+VuwPnMVgb+a4+KNJSfgqahpnfWaeDXCimw==} - engines: {node: '>=18.0.0'} - lilconfig@3.1.3: resolution: {integrity: sha512-/vlFKAoH5Cgt3Ie+JLhRbwOsCQePABiU3tJ1egGvyQ+33R/vcwM2Zl2QR/LzjsBeItPt3oSVXapn+m4nQDvpzw==} engines: {node: '>=14'} @@ -1757,13 +1668,6 @@ packages: resolution: {integrity: sha512-aisnrDP4GNe06UcKFnV5bfMNPBUw4jsLGaWwWfnH3v02GnBuXX2MCVn5RbrWo0j3pczUilYblq7fQ7Nw2t5XKw==} engines: {node: '>= 0.8'} - memjs@1.3.2: - resolution: {integrity: sha512-qUEg2g8vxPe+zPn09KidjIStHPtoBO8Cttm8bgJFWWabbsjQ9Av9Ky+6UcvKx6ue0LLb/LEhtcyQpRyKfzeXcg==} - engines: {node: '>=0.10.0'} - - memory-pager@1.5.0: - resolution: {integrity: sha512-ZS4Bp4r/Zoeq6+NLJpP+0Zzm0pR8whtGPf1XExKLJBAczGMnSi3It14OiNCStjQjM6NU1okjQGSxgEZN8eBYKg==} - merge-descriptors@2.0.0: resolution: {integrity: sha512-Snk314V5ayFLhp3fkUREub6WtjBfPdCPY1Ln8/8munuLuiYhsABgBVWsozAG+MWMbVEvcdcpbi9R7ww22l9Q3g==} engines: {node: '>=18'} @@ -1798,49 +1702,6 @@ packages: mlly@1.8.2: resolution: {integrity: sha512-d+ObxMQFmbt10sretNDytwt85VrbkhhUA/JBGm1MPaWJ65Cl4wOgLaB1NYvJSZ0Ef03MMEU/0xpPMXUIQ29UfA==} - mongodb-connection-string-url@7.0.1: - resolution: {integrity: sha512-h0AZ9A7IDVwwHyMxmdMXKy+9oNlF0zFoahHiX3vQ8e3KFcSP3VmsmfvtRSuLPxmyv2vjIDxqty8smTgie/SNRQ==} - engines: {node: '>=20.19.0'} - - mongodb@7.2.0: - resolution: {integrity: sha512-F/2+BMZtLVhY30ioZp0dAmZ+IRZMBqI+nrv6t5+9/1AIwCa8sMRC3jBf81lpxMhnZgqq8CoUD503Z1oZWq1/sw==} - engines: {node: '>=20.19.0'} - peerDependencies: - '@aws-sdk/credential-providers': ^3.806.0 - '@mongodb-js/zstd': ^7.0.0 - gcp-metadata: ^7.0.1 - kerberos: ^7.0.0 - mongodb-client-encryption: '>=7.0.0 <7.1.0' - snappy: ^7.3.2 - socks: ^2.8.6 - peerDependenciesMeta: - '@aws-sdk/credential-providers': - optional: true - '@mongodb-js/zstd': - optional: true - gcp-metadata: - optional: true - kerberos: - optional: true - mongodb-client-encryption: - optional: true - snappy: - optional: true - socks: - optional: true - - mongoose@9.6.1: - resolution: {integrity: sha512-3T8/b0plM3ZJPW3WjlzVMIGJEYYTjgDPQ05Qzru3xu3/wOPSFKWYxdwUF2dl8h3NG5dVkzIuOkZdLacnlLf/sA==} - engines: {node: '>=20.19.0'} - - mpath@0.9.0: - resolution: {integrity: sha512-ikJRQTk8hw5DEoFVxHG1Gn9T/xcjtdnOKIU1JTmGjZZlg9LST2mBLmcX3/ICIbgJydT2GOc15RnNy5mHmzfSew==} - engines: {node: '>=4.0.0'} - - mquery@6.0.0: - resolution: {integrity: sha512-b2KQNsmgtkscfeDgkYMcWGn9vZI9YoXh802VDEwE6qc50zxBFQ0Oo8ROkawbPAsXCY1/Z1yp0MagqsZStPWJjw==} - engines: {node: '>=20.19.0'} - mri@1.2.0: resolution: {integrity: sha512-tzzskb3bG8LvYGFF/mDTpq3jpI6Q9wc3LEmBaghu+DdCssd1FakN7Bc0hVNmEyGq1bq3RgfkCb3cmQLpNPOroA==} engines: {node: '>=4'} @@ -1856,10 +1717,6 @@ packages: engines: {node: ^10 || ^12 || ^13.7 || ^14 || >=15.0.1} hasBin: true - natural@8.1.1: - resolution: {integrity: sha512-Ucb+lsUcGxUqu3rn8cwHjT6gJQosO63nIX/aBQXB3+IDkNbFV7PuviysO+Rzz3aKn7PZhPj3bNF4PS9gDVjYCQ==} - engines: {node: '>=0.4.10'} - negotiator@1.0.0: resolution: {integrity: sha512-8Ofs/AUQh8MaEcrlq5xOX0CQ9ypTF5dl78mjlMNfOK08fzpgTHQRQPBxcPlEtIw0yRpws+Zo/3r+5WRby7u3Gg==} engines: {node: '>= 0.6'} @@ -1971,40 +1828,6 @@ packages: resolution: {integrity: sha512-//nshmD55c46FuFw26xV/xFAaB5HF9Xdap7HJBBnrKdAd6/GxDBaNA1870O79+9ueg61cZLSVc+OaFlfmObYVQ==} engines: {node: '>= 14.16'} - pg-cloudflare@1.3.0: - resolution: {integrity: sha512-6lswVVSztmHiRtD6I8hw4qP/nDm1EJbKMRhf3HCYaqud7frGysPv7FYJ5noZQdhQtN2xJnimfMtvQq21pdbzyQ==} - - pg-connection-string@2.12.0: - resolution: {integrity: sha512-U7qg+bpswf3Cs5xLzRqbXbQl85ng0mfSV/J0nnA31MCLgvEaAo7CIhmeyrmJpOr7o+zm0rXK+hNnT5l9RHkCkQ==} - - pg-int8@1.0.1: - resolution: {integrity: sha512-WCtabS6t3c8SkpDBUlb1kjOs7l66xsGdKpIPZsg4wR+B3+u9UAum2odSsF9tnvxg80h4ZxLWMy4pRjOsFIqQpw==} - engines: {node: '>=4.0.0'} - - pg-pool@3.13.0: - resolution: {integrity: sha512-gB+R+Xud1gLFuRD/QgOIgGOBE2KCQPaPwkzBBGC9oG69pHTkhQeIuejVIk3/cnDyX39av2AxomQiyPT13WKHQA==} - peerDependencies: - pg: '>=8.0' - - pg-protocol@1.13.0: - resolution: {integrity: sha512-zzdvXfS6v89r6v7OcFCHfHlyG/wvry1ALxZo4LqgUoy7W9xhBDMaqOuMiF3qEV45VqsN6rdlcehHrfDtlCPc8w==} - - pg-types@2.2.0: - resolution: {integrity: sha512-qTAAlrEsl8s4OiEQY69wDvcMIdQN6wdz5ojQiOy6YRMuynxenON0O5oCpJI6lshc6scgAY8qvJ2On/p+CXY0GA==} - engines: {node: '>=4'} - - pg@8.20.0: - resolution: {integrity: sha512-ldhMxz2r8fl/6QkXnBD3CR9/xg694oT6DZQ2s6c/RI28OjtSOpxnPrUCGOBJ46RCUxcWdx3p6kw/xnDHjKvaRA==} - engines: {node: '>= 16.0.0'} - peerDependencies: - pg-native: '>=3.0.1' - peerDependenciesMeta: - pg-native: - optional: true - - pgpass@1.0.5: - resolution: {integrity: sha512-FdW9r/jQZhSeohs1Z3sI1yxFQNFvMcnmfuj4WBMUTxOrAyLMaTcE1aAMBiTlbMNaXvBCQuVi0R7hd8udDSP7ug==} - picocolors@1.1.1: resolution: {integrity: sha512-xceH2snhtb5M9liqDsmEw56le376mTZkEX/jEb/RxNFyegNul7eNslCXP9FDj/Lcu0X8KEyMceP2ntpaHrDEVA==} @@ -2067,22 +1890,6 @@ packages: resolution: {integrity: sha512-qif0+jGGZoLWdHey3UFHHWP0H7Gbmsk8T5VEqyYFbWqPr1XqvLGBbk/sl8V5exGmcYJklJOhOQq1pV9IcsiFag==} engines: {node: ^10 || ^12 || >=14} - postgres-array@2.0.0: - resolution: {integrity: sha512-VpZrUqU5A69eQyW2c5CA1jtLecCsN2U/bD6VilrFDWq5+5UIEVO7nazS3TEcHf1zuPYO/sqGvUvW62g86RXZuA==} - engines: {node: '>=4'} - - postgres-bytea@1.0.1: - resolution: {integrity: sha512-5+5HqXnsZPE65IJZSMkZtURARZelel2oXUEO8rH83VS/hxH5vv1uHquPg5wZs8yMAfdv971IU+kcPUczi7NVBQ==} - engines: {node: '>=0.10.0'} - - postgres-date@1.0.7: - resolution: {integrity: sha512-suDmjLVQg78nMK2UZ454hAG+OAW+HQPZ6n++TNDUX+L0+uUlLywnoxJKDou51Zm+zTCjrCl0Nq6J9C5hP9vK/Q==} - engines: {node: '>=0.10.0'} - - postgres-interval@1.2.0: - resolution: {integrity: sha512-9ZhXKM/rw350N1ovuWHbGxnGh/SNJ4cnxHiM0rxE4VN41wsg8P8zWn9hv/buK00RP4WvlOyr/RBDiptyxVbkZQ==} - engines: {node: '>=0.10.0'} - prettier@2.8.8: resolution: {integrity: sha512-tdN8qQGvNjw4CHbY+XXk0JgCXn9QiF21a55rBe5LJAU+kDyC4WQn4+awm2Xfk2lQMk5fKup9XgzTZtGkjBdP9Q==} engines: {node: '>=10.13.0'} @@ -2098,10 +1905,6 @@ packages: pump@3.0.4: resolution: {integrity: sha512-VS7sjc6KR7e1ukRFhQSY5LM2uBWAUPiOPa/A3mkKmiMwSmRFUITt0xuj+/lesgnCv+dPIEYlkzrcyXgquIHMcA==} - punycode@2.3.1: - resolution: {integrity: sha512-vYt7UD1U9Wg6138shLtLOvdAu+8DsC/ilFtEVHcH+wydcSpNE20AfSOduf6MkRFahL5FY7X1oU7nKVZFtfq8Fg==} - engines: {node: '>=6'} - qs@6.15.2: resolution: {integrity: sha512-Rzq0KEyX/w/tEybncDgdkZrJgVUsUMk3xjh3t5bv3S1HTAtg+uOYt72+ZfwiQwKdysThkTBdL/rTi6HDmX9Ddw==} engines: {node: '>=0.6'} @@ -2135,10 +1938,6 @@ packages: resolution: {integrity: sha512-57frrGM/OCTLqLOAh0mhVA9VBMHd+9U7Zb2THMGdBUoZVOtGbJzjxsYGDJ3A9AYYCP4hn6y1TVbaOfzWtm5GFg==} engines: {node: '>= 12.13.0'} - redis@5.12.1: - resolution: {integrity: sha512-LDsoVvb/CpoV9EN3FXvgvSHNJWuCIzl9MiO3ppOevuGLpSGJhwfQjpEwfFJcQvNSddHADDdZaWx0HnmMxRXG7g==} - engines: {node: '>= 18.19.0'} - require-from-string@2.0.2: resolution: {integrity: sha512-Xf0nWe6RseziFMu+Ap9biiUbmplq6S9/p+7w7YXP/JBHhrUDDUhwa+vANyubuqfZWTveU//DYVGsDG7RKL/vEw==} engines: {node: '>=0.10.0'} @@ -2217,9 +2016,6 @@ packages: resolution: {integrity: sha512-ZX99e6tRweoUXqR+VBrslhda51Nh5MTQwou5tnUDgbtyM0dBgmhEDtWGP/xbKn6hqfPRHujUNwz5fy/wbbhnpw==} engines: {node: '>= 0.4'} - sift@17.1.3: - resolution: {integrity: sha512-Rtlj66/b0ICeFzYTuNvX/EF1igRbbnGSvEyT79McoZa/DeGhMyC5pWKOEsZKnpkqtSeovd5FL/bjHWC3CIIvCQ==} - siginfo@2.0.0: resolution: {integrity: sha512-ybx0WO1/8bSBLEWXZvEd7gMW3Sn3JFlW3TvX1nREbDLRNQNaeNN8WK0meBwPdAaOI7TtRRRJn/Es1zhrrCHu7g==} @@ -2242,9 +2038,6 @@ packages: resolution: {integrity: sha512-i5uvt8C3ikiWeNZSVZNWcfZPItFQOsYTUAOkcUPGd8DqDy1uOUikjt5dG+uRlwyvR108Fb9DOd4GvXfT0N2/uQ==} engines: {node: '>= 12'} - sparse-bitfield@3.0.3: - resolution: {integrity: sha512-kvzhi7vqKTfkh0PZU+2D2PIllw2ymqJKujUcyPMd9Y75Nv4nPbGJZXNhxsgdQab2BmlDct1YnfQCguEvHr7VsQ==} - spawndamnit@3.0.1: resolution: {integrity: sha512-MmnduQUuHCoFckZoWnXsTg7JaiLBJrKFj9UI2MbRPGaJeVpsLcVBu6P/IGZovziM/YBsellCmsprgNA+w0CzVg==} @@ -2268,10 +2061,6 @@ packages: std-env@3.10.0: resolution: {integrity: sha512-5GS12FdOZNliM5mAOxFRg7Ir0pWz8MdpYm6AY6VPkGpbA7ZzmbzNcBJQ0GPvvyWgcY7QAhCgf9Uy89I03faLkg==} - stopwords-iso@1.1.0: - resolution: {integrity: sha512-I6GPS/E0zyieHehMRPQcqkiBMJKGgLta+1hREixhoLPqEA0AlVFiC43dl8uPpmkkeRdDMzYRWFWk5/l9x7nmNg==} - engines: {node: '>=0.10.0'} - string-width@4.2.3: resolution: {integrity: sha512-wKyQRQpjJ0sIp62ErSZdGsjMJWsap5oRNihHhu6G7JVO/9jIB6UyevL+tXuOqrng8j/cxKTWyWUwvSTriiZz/g==} engines: {node: '>=8'} @@ -2304,17 +2093,10 @@ packages: engines: {node: '>=16 || 14 >=14.17'} hasBin: true - suffix-thumb@5.0.2: - resolution: {integrity: sha512-I5PWXAFKx3FYnI9a+dQMWNqTxoRt6vdBdb0O+BJ1sxXCWtSoQCusc13E58f+9p4MYx/qCnEMkD5jac6K2j3dgA==} - supports-color@7.2.0: resolution: {integrity: sha512-qpCAvRl9stuOHveKsn7HncJRvv501qIacKzQlO/+Lwxc9+0q2wLyv4Dfvt80/DPn2pqOBsJdDiogXGR9+OvwRw==} engines: {node: '>=8'} - sylvester@0.0.21: - resolution: {integrity: sha512-yUT0ukFkFEt4nb+NY+n2ag51aS/u9UHXoZw+A4jgD77/jzZsBoSDHuqysrVCBC4CYR4TYvUJq54ONpXgDBH8tA==} - engines: {node: '>=0.2.6'} - term-size@2.2.1: resolution: {integrity: sha512-wK0Ri4fOGjv/XPy8SBHZChl8CM7uMc5VML7SqiQ0zG7+J5Vr+RMQDoHa2CNT6KHUnTGIXH34UDMkPzAUyapBZg==} engines: {node: '>=8'} @@ -2367,10 +2149,6 @@ packages: tr46@0.0.3: resolution: {integrity: sha512-N3WMsuqV66lT30CrXNbEjx4GEwlow3v6rr4mCcv6prnfwhS01rkgyFdjPNBYd9br7LpXV1+Emh01fHnq2Gdgrw==} - tr46@5.1.1: - resolution: {integrity: sha512-hdF5ZgjTqgAntKkklYw0R03MG2x/bSzTtkxmIRw/sTNV8YXsCJ1tfLAX23lhxhHJlEf3CRCOCGGWw3vI3GaSPw==} - engines: {node: '>=18'} - tree-kill@1.2.2: resolution: {integrity: sha512-L0Orpi8qGpRG//Nd+H90vFB+3iHnue1zSSGmNOOCh1GLJ7rUKVwV2HvijphGQS2UmhUZewS9VgvxYIdgr+fG1A==} hasBin: true @@ -2416,9 +2194,6 @@ packages: ufo@1.6.4: resolution: {integrity: sha512-JFNbkD1Svwe0KvGi8GOeLcP4kAWQ609twvCdcHxq1oSL8svv39ZuSvajcD8B+5D0eL4+s1Is2D/O6KN3qcTeRA==} - underscore@1.13.8: - resolution: {integrity: sha512-DXtD3ZtEQzc7M8m4cXotyHR+FAS18C64asBYY5vqZexfYryNNnDc02W4hKg3rdQuqOYas1jkseX0+nZXjTXnvQ==} - undici-types@7.19.2: resolution: {integrity: sha512-qYVnV5OEm2AW8cJMCpdV20CDyaN3g0AjDlOGf1OW4iaDEx8MwdtChUp4zu4H0VP3nDRF/8RKWH+IPp9uW0YGZg==} @@ -2436,10 +2211,6 @@ packages: peerDependencies: browserslist: '>= 4.21.0' - uuid@14.0.0: - resolution: {integrity: sha512-Qo+uWgilfSmAhXCMav1uYFynlQO7fMFiMVZsQqZRMIXp0O7rR7qjkj+cPvBHLgBqi960QCoo/PH2/6ZtVqKvrg==} - hasBin: true - vary@1.1.2: resolution: {integrity: sha512-BNGbWLfd0eUPabhkXUVm0j8uuvREyTh5ovRa/dyow/BqAbZJyC+5fU+IzQOzmAKzYqYRAISoRhdQr3eIZ/PXqg==} engines: {node: '>= 0.8'} @@ -2520,14 +2291,6 @@ packages: webidl-conversions@3.0.1: resolution: {integrity: sha512-2JAn3z8AR6rjK8Sm8orRC0h/bcl/DqL7tRPdGZ4I1CjdF+EaMLmYxBHyXuKL849eucPFhvBoxMsflfOb8kxaeQ==} - webidl-conversions@7.0.0: - resolution: {integrity: sha512-VwddBukDzu71offAQR975unBIGqfKZpM+8ZX6ySk8nYhVoo5CYaZyzt3YBvYtRtO+aoGlqxPg/B87NGVZ/fu6g==} - engines: {node: '>=12'} - - whatwg-url@14.2.0: - resolution: {integrity: sha512-De72GdQZzNTUBBChsXueQUnPKDkg/5A5zp7pFDuQAj5UFoENpiACU0wlCvzpAGnTkj++ihpKwKyYewn/XNUbKw==} - engines: {node: '>=18'} - whatwg-url@5.0.0: resolution: {integrity: sha512-saE57nupxk6v3HY35+jzBwYa0rKSy0XR8JSxZPwgLr7ys0IBzhGviA1/TUGJLmSVqs8pb9AnvICXEuOHLprYTw==} @@ -2541,10 +2304,6 @@ packages: engines: {node: '>=8'} hasBin: true - wordnet-db@3.1.14: - resolution: {integrity: sha512-zVyFsvE+mq9MCmwXUWHIcpfbrHHClZWZiVOzKSxNJruIcFn2RbY55zkhiAMMxM8zCVSmtNiViq8FsAZSFpMYag==} - engines: {node: '>=0.6.0'} - wrap-ansi@7.0.0: resolution: {integrity: sha512-YVGIj2kamLSTxw6NsZjoBxfSwsn0ycdesmc4p+Q21c5zPuZ1pl+NfxVdxPtdHvmNVOQ6XSYG4AUtyt/Fi7D16Q==} engines: {node: '>=10'} @@ -2556,10 +2315,6 @@ packages: wrappy@1.0.2: resolution: {integrity: sha512-l4Sp/DRseor9wL6EvV2+TuQn63dMkPjZ/sp9XkghTEbV9KlPS1xUsZ3u7/IQO4wxtcFB4bgpQPRcR3QCvezPcQ==} - xtend@4.0.2: - resolution: {integrity: sha512-LKYU1iAXJXUgAXn9URjiu+MWhyUXHsvfp7mcuYm9dSUKK0/CjtrUwFAxD82/mCWbtLsGjFIad0wIsod4zrTAEQ==} - engines: {node: '>=0.4'} - yallist@3.1.1: resolution: {integrity: sha512-a4UGQaWPH59mOXUYnAG2ewncQS4i4F43Tv3JoAM+s2VDAmS9NsK8GpDMLrCHPksFT7h3K6TOoUNn2pb7RoXx4g==} @@ -3050,10 +2805,6 @@ snapshots: transitivePeerDependencies: - supports-color - '@mongodb-js/saslprep@1.4.10': - dependencies: - sparse-bitfield: 3.0.3 - '@nodelib/fs.scandir@2.1.5': dependencies: '@nodelib/fs.stat': 2.0.5 @@ -3112,28 +2863,6 @@ snapshots: '@pkgjs/parseargs@0.11.0': optional: true - '@redis/bloom@5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1))': - dependencies: - '@redis/client': 5.12.1(@opentelemetry/api@1.9.1) - - '@redis/client@5.12.1(@opentelemetry/api@1.9.1)': - dependencies: - cluster-key-slot: 1.1.2 - optionalDependencies: - '@opentelemetry/api': 1.9.1 - - '@redis/json@5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1))': - dependencies: - '@redis/client': 5.12.1(@opentelemetry/api@1.9.1) - - '@redis/search@5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1))': - dependencies: - '@redis/client': 5.12.1(@opentelemetry/api@1.9.1) - - '@redis/time-series@5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1))': - dependencies: - '@redis/client': 5.12.1(@opentelemetry/api@1.9.1) - '@rollup/rollup-android-arm-eabi@4.60.2': optional: true @@ -3244,12 +2973,6 @@ snapshots: dependencies: undici-types: 7.19.2 - '@types/webidl-conversions@7.0.3': {} - - '@types/whatwg-url@13.0.0': - dependencies: - '@types/webidl-conversions': 7.0.3 - '@vitest/coverage-istanbul@3.2.4(vitest@3.2.4(@types/node@25.6.0)(yaml@2.8.4))': dependencies: '@istanbuljs/schema': 0.1.6 @@ -3334,10 +3057,6 @@ snapshots: acorn@8.16.0: {} - afinn-165-financialmarketnews@3.0.0: {} - - afinn-165@2.0.2: {} - ajv-formats@3.0.1(ajv@8.20.0): optionalDependencies: ajv: 8.20.0 @@ -3363,10 +3082,6 @@ snapshots: any-promise@1.3.0: {} - apparatus@0.0.10: - dependencies: - sylvester: 0.0.21 - argparse@1.0.10: dependencies: sprintf-js: 1.0.3 @@ -3425,8 +3140,6 @@ snapshots: node-releases: 2.0.45 update-browserslist-db: 1.2.3(browserslist@4.28.2) - bson@7.2.0: {} - bundle-require@5.1.0(esbuild@0.27.7): dependencies: esbuild: 0.27.7 @@ -3470,8 +3183,6 @@ snapshots: dependencies: string-width: 4.2.3 - cluster-key-slot@1.1.2: {} - color-convert@2.0.1: dependencies: color-name: 1.1.4 @@ -3484,12 +3195,6 @@ snapshots: commander@4.1.1: {} - compromise@14.15.0: - dependencies: - efrt: 2.7.0 - grad-school: 0.0.5 - suffix-thumb: 5.0.2 - confbox@0.1.8: {} consola@3.4.2: {} @@ -3533,8 +3238,6 @@ snapshots: dependencies: path-type: 4.0.0 - dotenv@17.4.2: {} - dotenv@8.6.0: {} dunder-proto@1.0.1: @@ -3547,8 +3250,6 @@ snapshots: ee-first@1.1.1: {} - efrt@2.7.0: {} - electron-to-chromium@1.5.360: {} emoji-regex@8.0.0: {} @@ -3789,8 +3490,6 @@ snapshots: graceful-fs@4.2.11: {} - grad-school@0.0.5: {} - has-flag@4.0.0: {} has-symbols@1.1.0: {} @@ -3924,8 +3623,6 @@ snapshots: optionalDependencies: graceful-fs: 4.2.11 - kareem@3.3.0: {} - lilconfig@3.1.3: {} lines-and-columns@1.2.4: {} @@ -3964,10 +3661,6 @@ snapshots: media-typer@1.1.0: {} - memjs@1.3.2: {} - - memory-pager@1.5.0: {} - merge-descriptors@2.0.0: {} merge2@1.4.1: {} @@ -3998,38 +3691,6 @@ snapshots: pkg-types: 1.3.1 ufo: 1.6.4 - mongodb-connection-string-url@7.0.1: - dependencies: - '@types/whatwg-url': 13.0.0 - whatwg-url: 14.2.0 - - mongodb@7.2.0: - dependencies: - '@mongodb-js/saslprep': 1.4.10 - bson: 7.2.0 - mongodb-connection-string-url: 7.0.1 - - mongoose@9.6.1: - dependencies: - kareem: 3.3.0 - mongodb: 7.2.0 - mpath: 0.9.0 - mquery: 6.0.0 - ms: 2.1.3 - sift: 17.1.3 - transitivePeerDependencies: - - '@aws-sdk/credential-providers' - - '@mongodb-js/zstd' - - gcp-metadata - - kerberos - - mongodb-client-encryption - - snappy - - socks - - mpath@0.9.0: {} - - mquery@6.0.0: {} - mri@1.2.0: {} ms@2.1.3: {} @@ -4042,34 +3703,6 @@ snapshots: nanoid@3.3.12: {} - natural@8.1.1(@opentelemetry/api@1.9.1): - dependencies: - afinn-165: 2.0.2 - afinn-165-financialmarketnews: 3.0.0 - apparatus: 0.0.10 - dotenv: 17.4.2 - memjs: 1.3.2 - mongoose: 9.6.1 - pg: 8.20.0 - redis: 5.12.1(@opentelemetry/api@1.9.1) - safe-stable-stringify: 2.5.0 - stopwords-iso: 1.1.0 - sylvester: 0.0.21 - underscore: 1.13.8 - uuid: 14.0.0 - wordnet-db: 3.1.14 - transitivePeerDependencies: - - '@aws-sdk/credential-providers' - - '@mongodb-js/zstd' - - '@node-rs/xxhash' - - '@opentelemetry/api' - - gcp-metadata - - kerberos - - mongodb-client-encryption - - pg-native - - snappy - - socks - negotiator@1.0.0: {} node-fetch@2.7.0: @@ -4143,41 +3776,6 @@ snapshots: pathval@2.0.1: {} - pg-cloudflare@1.3.0: - optional: true - - pg-connection-string@2.12.0: {} - - pg-int8@1.0.1: {} - - pg-pool@3.13.0(pg@8.20.0): - dependencies: - pg: 8.20.0 - - pg-protocol@1.13.0: {} - - pg-types@2.2.0: - dependencies: - pg-int8: 1.0.1 - postgres-array: 2.0.0 - postgres-bytea: 1.0.1 - postgres-date: 1.0.7 - postgres-interval: 1.2.0 - - pg@8.20.0: - dependencies: - pg-connection-string: 2.12.0 - pg-pool: 3.13.0(pg@8.20.0) - pg-protocol: 1.13.0 - pg-types: 2.2.0 - pgpass: 1.0.5 - optionalDependencies: - pg-cloudflare: 1.3.0 - - pgpass@1.0.5: - dependencies: - split2: 4.2.0 - picocolors@1.1.1: {} picomatch@2.3.2: {} @@ -4245,16 +3843,6 @@ snapshots: picocolors: 1.1.1 source-map-js: 1.2.1 - postgres-array@2.0.0: {} - - postgres-bytea@1.0.1: {} - - postgres-date@1.0.7: {} - - postgres-interval@1.2.0: - dependencies: - xtend: 4.0.2 - prettier@2.8.8: {} process-warning@5.0.0: {} @@ -4269,8 +3857,6 @@ snapshots: end-of-stream: 1.4.5 once: 1.4.0 - punycode@2.3.1: {} - qs@6.15.2: dependencies: side-channel: 1.1.0 @@ -4301,17 +3887,6 @@ snapshots: real-require@0.2.0: {} - redis@5.12.1(@opentelemetry/api@1.9.1): - dependencies: - '@redis/bloom': 5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1)) - '@redis/client': 5.12.1(@opentelemetry/api@1.9.1) - '@redis/json': 5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1)) - '@redis/search': 5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1)) - '@redis/time-series': 5.12.1(@redis/client@5.12.1(@opentelemetry/api@1.9.1)) - transitivePeerDependencies: - - '@node-rs/xxhash' - - '@opentelemetry/api' - require-from-string@2.0.2: {} resolve-from@5.0.0: {} @@ -4434,8 +4009,6 @@ snapshots: side-channel-map: 1.0.1 side-channel-weakmap: 1.0.2 - sift@17.1.3: {} - siginfo@2.0.0: {} signal-exit@4.1.0: {} @@ -4450,10 +4023,6 @@ snapshots: source-map@0.7.6: {} - sparse-bitfield@3.0.3: - dependencies: - memory-pager: 1.5.0 - spawndamnit@3.0.1: dependencies: cross-spawn: 7.0.6 @@ -4474,8 +4043,6 @@ snapshots: std-env@3.10.0: {} - stopwords-iso@1.1.0: {} - string-width@4.2.3: dependencies: emoji-regex: 8.0.0 @@ -4514,14 +4081,10 @@ snapshots: tinyglobby: 0.2.16 ts-interface-checker: 0.1.13 - suffix-thumb@5.0.2: {} - supports-color@7.2.0: dependencies: has-flag: 4.0.0 - sylvester@0.0.21: {} - term-size@2.2.1: {} test-exclude@7.0.2: @@ -4565,10 +4128,6 @@ snapshots: tr46@0.0.3: {} - tr46@5.1.1: - dependencies: - punycode: 2.3.1 - tree-kill@1.2.2: {} ts-algebra@2.0.0: {} @@ -4622,8 +4181,6 @@ snapshots: ufo@1.6.4: {} - underscore@1.13.8: {} - undici-types@7.19.2: {} universalify@0.1.2: {} @@ -4636,8 +4193,6 @@ snapshots: escalade: 3.2.0 picocolors: 1.1.1 - uuid@14.0.0: {} - vary@1.1.2: {} vite-node@3.2.4(@types/node@25.6.0)(yaml@2.8.4): @@ -4717,13 +4272,6 @@ snapshots: webidl-conversions@3.0.1: {} - webidl-conversions@7.0.0: {} - - whatwg-url@14.2.0: - dependencies: - tr46: 5.1.1 - webidl-conversions: 7.0.0 - whatwg-url@5.0.0: dependencies: tr46: 0.0.3 @@ -4738,8 +4286,6 @@ snapshots: siginfo: 2.0.0 stackback: 0.0.2 - wordnet-db@3.1.14: {} - wrap-ansi@7.0.0: dependencies: ansi-styles: 4.3.0 @@ -4754,8 +4300,6 @@ snapshots: wrappy@1.0.2: {} - xtend@4.0.2: {} - yallist@3.1.1: {} yaml@2.8.4: {}