Skip to content

Pull requests: benchflow-ai/awesome-evals

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Add Bifrost to observability and eval platforms
#94 opened Sep 4, 2026 by Swpn0neel Loading…
Add public benchmarks and leaderboards
#93 opened Sep 4, 2026 by crypblizz8 Loading…
Add Agent Amusement Park trajectory benchmark
#89 opened Sep 1, 2026 by axay-oxo Loading…
Add the 44%-vs-85% cross-agent harness audit to §3, with deep note
#88 opened Aug 28, 2026 by uipreliga Contributor Loading…
Add Mendmark mutation-testing benchmark
#87 opened Aug 21, 2026 by danielgaskins Loading…
Swap scan/responder model to claude-opus-4-6
#86 opened Aug 20, 2026 by xdotli Member Loading…
Add CHANGELOG.md + biweekly changelog-rollup workflow
#84 opened Aug 20, 2026 by xdotli Member Loading…
Add CrossSource to 5c · RAG / retrieval evaluation
#83 opened Aug 20, 2026 by zoeb-nomi Loading…
Add Proofline to agent-specific evaluation
#82 opened Aug 18, 2026 by ceodaradigu Loading…
Add AgentRunProof to eval frameworks and harnesses
#81 opened Aug 16, 2026 by FU-max-boop Loading…
Add ClawBench to agent benchmark resources
#64 opened Jul 29, 2026 by reacher-z Contributor Loading…
Add StructEval benchmark
#57 opened Jul 27, 2026 by reacher-z Contributor Loading…
Add greenproof to §8 (verifiers)
#50 opened Jul 21, 2026 by zxyasfas Loading…
ProTip! Type g i on any issue or pull request to go back to the issue listing page.