Skip to content

feat(eval): add runner, CLI, aggregator, and HTML reports - #2

Closed
elkaix wants to merge 12 commits into
feature/eval-harness-1afrom
feature/eval-harness-1b
Closed

elkaix wants to merge 12 commits into
feature/eval-harness-1afrom
feature/eval-harness-1b

Conversation

@elkaix

@elkaix elkaix commented Apr 27, 2026

Copy link
Copy Markdown
Owner

Summary

Phase 1 / Sub-plan 1B — runner + CLI on top of #1.

  • EvalConfig YAML loader + baseline config
  • Run-directory storage layer
  • EvalPipeline factory bridging EvalConfig to the existing RAGBackend (telemetry helpers split out)
  • Aggregator producing per-dataset and combined rows with bootstrap CIs
  • EvalRunner orchestrator end-to-end
  • Two-run comparison with paired significance tests
  • Jinja2-based HTML report renderer for runs and comparisons
  • Argparse CLI (run / list / show / compare)
  • Public exports + end-to-end integration smoke test

Deps added: pyyaml, jinja2, datasets (gitignores eval_runs/).

Stack

Targets feature/eval-harness-1a (#1). Will retarget to main after 1A merges.

Test plan

  • python -m pytest tests/eval/ -v — green including new runner/CLI tests
  • End-to-end integration smoke runs through factory → runner → aggregator → reports
  • Reviewer: dry-run CLI python -m src.eval.cli run --config configs/baseline.yaml
  • Reviewer: open generated HTML report in browser, confirm CI whiskers render

elkaix added 11 commits April 26, 2026 21:34
Implements five functions (compute_run_id, save_run, load_run,
list_runs, delete_run) that persist and retrieve eval run artifacts as
plain JSON/JSONL files under EVAL_RUNS_DIR. Env-overridable root dir
enables isolated tmp-path testing. Path-traversal guard on delete_run
validates run_id before any filesystem call.

All 10 unit tests pass.
Extract token counting and tiktoken fallback logic into a dedicated
_telemetry.py module to bring pipeline_factory.py under the 250-line
ceiling.

Changes:
  - New src/eval/_telemetry.py (60 lines): count_tokens function + tiktoken
    import with word-count fallback. Kept separate to signal it's internal
    to the eval package (leading underscore).
  - Updated src/eval/pipeline_factory.py (334 lines, was 375):
    - Remove tiktoken import + _tiktoken_warned global
    - Remove _count_tokens helper
    - Import count_tokens from _telemetry
    - Update call sites (count_tokens instead of _count_tokens)

Public API unchanged: build_pipeline() and EvalPipeline are still
exported from pipeline_factory and can be imported by callers.

All 3 tests still pass.
Implements aggregate() which turns list[EvalResult] into
list[AggregatedMetric] with bootstrap CIs. Produces one row per
(metric, dataset) pair plus a combined row (dataset=None). Skips
combos with <3 samples and emits warning strings. Errored results
are excluded before grouping.
Implement compare_runs() which loads two eval runs, validates their
eval_set_versions match, pairs per-question scores by question_id, and
runs a paired permutation test per (metric, dataset) combo to produce
MetricDelta objects with p-values and significance flags. Also picks
top-10 per-question regressions/wins by |delta| on the headline metric.
Implements four subcommands (run/list/show/compare) driven by argparse.
Supports EVAL_LLM_OVERRIDE_DUMMY and EVAL_SQUAD_PATH env vars for
test injection without touching production code paths.
@elkaix

elkaix commented Jul 12, 2026

Copy link
Copy Markdown
Owner Author

Closing as superseded/stale. This sub-plan's work (eval harness runner/CLI/aggregator, /api/eval routes + React eval UI, OpenTelemetry/Phoenix + telemetry footer) already landed in main through the actual merge path; this branch is ~40–50 commits behind and targets another eval-harness branch rather than main. Reopen if any specific commit here still needs to be cherry-picked.

@elkaix elkaix closed this Jul 12, 2026
@elkaix
elkaix deleted the feature/eval-harness-1b branch July 12, 2026 18:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant