Conversation
Implements five functions (compute_run_id, save_run, load_run, list_runs, delete_run) that persist and retrieve eval run artifacts as plain JSON/JSONL files under EVAL_RUNS_DIR. Env-overridable root dir enables isolated tmp-path testing. Path-traversal guard on delete_run validates run_id before any filesystem call. All 10 unit tests pass.
Extract token counting and tiktoken fallback logic into a dedicated
_telemetry.py module to bring pipeline_factory.py under the 250-line
ceiling.
Changes:
- New src/eval/_telemetry.py (60 lines): count_tokens function + tiktoken
import with word-count fallback. Kept separate to signal it's internal
to the eval package (leading underscore).
- Updated src/eval/pipeline_factory.py (334 lines, was 375):
- Remove tiktoken import + _tiktoken_warned global
- Remove _count_tokens helper
- Import count_tokens from _telemetry
- Update call sites (count_tokens instead of _count_tokens)
Public API unchanged: build_pipeline() and EvalPipeline are still
exported from pipeline_factory and can be imported by callers.
All 3 tests still pass.
Implements aggregate() which turns list[EvalResult] into list[AggregatedMetric] with bootstrap CIs. Produces one row per (metric, dataset) pair plus a combined row (dataset=None). Skips combos with <3 samples and emits warning strings. Errored results are excluded before grouping.
Implement compare_runs() which loads two eval runs, validates their eval_set_versions match, pairs per-question scores by question_id, and runs a paired permutation test per (metric, dataset) combo to produce MetricDelta objects with p-values and significance flags. Also picks top-10 per-question regressions/wins by |delta| on the headline metric.
Implements four subcommands (run/list/show/compare) driven by argparse. Supports EVAL_LLM_OVERRIDE_DUMMY and EVAL_SQUAD_PATH env vars for test injection without touching production code paths.
1 of 3 tasks
Owner
Author
|
Closing as superseded/stale. This sub-plan's work (eval harness runner/CLI/aggregator, /api/eval routes + React eval UI, OpenTelemetry/Phoenix + telemetry footer) already landed in |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Phase 1 / Sub-plan 1B — runner + CLI on top of #1.
EvalConfigYAML loader + baseline configEvalPipelinefactory bridgingEvalConfigto the existingRAGBackend(telemetry helpers split out)EvalRunnerorchestrator end-to-endrun/list/show/compare)Deps added:
pyyaml,jinja2,datasets(gitignoreseval_runs/).Stack
Targets
feature/eval-harness-1a(#1). Will retarget tomainafter 1A merges.Test plan
python -m pytest tests/eval/ -v— green including new runner/CLI testspython -m src.eval.cli run --config configs/baseline.yaml