diff --git a/README.md b/README.md index 2aec1acc..a0aee2b6 100644 --- a/README.md +++ b/README.md @@ -390,6 +390,7 @@ The [independent model research](community/projects/tools/README.md#independent- - [jevcache](https://github.com/kushals256/jevcache) - OpenAI-compatible local cache proxy: TypeSafe Jev (OpenRouter Decisions) admits same-intent paraphrases so expensive chat completions can be skipped (fail-open on Jev errors). [Project guide](community/projects/tools/jevcache.md). - [jeval](https://github.com/rlaope/jeval) - Measures probabilistic classifier calibration (incl. Jev confidence) and cost-optimal human hand-off thresholds; offline demo included. [Project guide](community/projects/tools/jeval.md). - [jevals](https://github.com/dayhaysoos/jevals) - Local TypeSafe Jev evaluation workbench: author Noul/Choice/Score cases with expected answers, run them, and compare saved results. [Project guide](community/projects/tools/jevals.md). +- [Jevals.com](https://jevals.com/) - Hosted independent boards: TypeSafe Jev vs LLMs on PubMedQA/Banking77/HelpSteer2 (calibration/cost/latency); data CC BY 4.0 (distinct from dayhaysoos/jevals). [Project guide](community/projects/tools/jevals-com.md). - [jevseek](https://github.com/blingdivinity/jevseek) - DeepSeek proposes next tokens; TypeSafe Jev via OpenRouter System One chooses which to append. [Project guide](community/projects/tools/jevseek.md). - [jevbus](https://github.com/zkjoie/jevbus) - Rust streaming event bus: TypeSafe Jev (or any Judge) decides route/subscribe/deliver dispositions with policy thresholds. [Project guide](community/projects/tools/jevbus.md). - [jev-agent-browser](https://github.com/mhingston/jev-agent-browser) - Confidence-gated browser actions for `agent-browser` with TypeSafe Jev (or Gateway/Cloudflare/custom providers). [Project guide](community/projects/tools/jev-agent-browser.md). diff --git a/community/projects/tools/README.md b/community/projects/tools/README.md index 02b3bb96..85856c6c 100644 --- a/community/projects/tools/README.md +++ b/community/projects/tools/README.md @@ -206,6 +206,7 @@ See the [computer-use guide](../../../docs/computer-use.md) for a comparison, fo | [jev-agent-browser](jev-agent-browser.md) | Confidence-gated next browser action for `agent-browser` via TypeSafe Jev (or Gateway/Cloudflare/custom). | TypeScript · npm (`@mhingston5/jev-agent-browser` 0.3.1) | | [jev-agent-failure-benchmark](jev-agent-failure-benchmark.md) | Score TypeSafe Jev on Who&When Pro text traces for responsible agent, step, and error type; compare to paper LLMs. | Python · CLI (`jevbench`, Apache-2.0) | | [jev-agent-kit](jev-agent-kit.md) | Zero-dependency CLI + MCP tools (check/choose/score/judge/route/triage/guard/grep/rank/compact) on TypeSafe Jev — distinct from the Rust jevkit CLI. | Node.js ≥ 18 · npm (`@walidboulanouar/jevkit` 0.2.0) | +| [Jevals.com](jevals-com.md) | Hosted independent Jev vs LLM boards (accuracy/calibration/cost/latency); open data, private harness (distinct from local jevals). | Hosted boards + [jevals-data](https://github.com/Jevals/jevals-data) (CC BY 4.0) | | [Jevaluate](jevaluate.md) | Confidence-gated web walkthroughs with TypeSafe Jev; optional DeepSeek vision; eval/judge scripts and skill. | Node/Python · Playwright scripts (MIT) | | [jevbus](jevbus.md) | Route/subscribe/deliver streaming events with TypeSafe Jev (or any Judge) and policy thresholds. | Rust · crate (`jevbus` 0.1.0) | | [jevtok](jevtok.md) | Count Jev tokens and estimate billed request input_tokens offline before calling TypeSafe. | Python · library/CLI (`jevtok` 0.1.0) | diff --git a/community/projects/tools/jevals-com.md b/community/projects/tools/jevals-com.md new file mode 100644 index 00000000..57d09a62 --- /dev/null +++ b/community/projects/tools/jevals-com.md @@ -0,0 +1,43 @@ +# Jevals.com + +[All projects](../README.md) · [Developer tools](README.md#developer-tools) + +Hosted independent benchmark boards: TypeSafe Jev (`typesafe-ai/jev` via Vercel AI Gateway) vs six LLMs on PubMedQA (Noul), Banking77 (Choice), and HelpSteer2 (Score)—accuracy, calibration, coverage@threshold, cost, and latency. Distinct from the local [jevals](jevals.md) workbench (dayhaysoos). + +| At a glance | Details | +| --- | --- | +| Source | [Source](https://jevals.com/) | +| Maintainer | [Jevals](https://github.com/Jevals) (community issue [#372](https://github.com/AppitStudio/awesome-jev/issues/372)). Independently curated listing of a maintainer-submitted resource; not a TypeSafe endorsement. | +| Format | Hosted release boards + methodology pages; open data at [Jevals/jevals-data](https://github.com/Jevals/jevals-data) (CC BY 4.0). Harness code is private. | +| Requirements | Browser only to read boards. No account required for the public site (checked 2026-09-23). Recomputing numbers needs the published data repo—not the private harness. | +| License | Site: free to read (no paid tier stated). Data: [CC BY 4.0](https://github.com/Jevals/jevals-data). Harness: closed / not published. | +| Disclosure | AI-assisted catalog review of the public site, issue #372, and jevals-data metadata. Implementation of the private harness was **not** inspected. Metrics and methodology are vendor/maintainer-reported; this listing did not re-run the suite. Listing is not an endorsement. | + +## When to use + +Use it to inspect how hosted Jev confidence trades off against accuracy/cost/latency on labelled decision tasks vs adapter-prompted LLMs. Prefer [jevals](jevals.md) when you want to author and run local Noul/Choice/Score cases yourself. + +## How it works + +Per maintainer description: same typed questions graded against human labels (300 items × 5 runs). Boards report accuracy, calibration error, coverage and accuracy at confidence thresholds, cost per 1k decisions, and p95 latency. Per-decision logs and suite files ship in jevals-data so numbers can be recomputed without the harness. + +## Get started + +1. Open [jevals.com](https://jevals.com/) and the [methodology](https://jevals.com/methodology/) pages. +2. Optional: browse [Jevals/jevals-data](https://github.com/Jevals/jevals-data) for CC BY 4.0 boards and logs. + +## Examples and demos + +- Public boards on the homepage (Noul/Choice/Score). +- Atom feed linked from the site for releases. +- Issue #372 details pinning by release date (e.g. 2026-09-18) when Gateway omits a version string. + +## Limits and data handling + +One published release slice, three English tasks; harness not reproducible as-is. Visiting the site may involve analytics (Google tags observed in HTML). This listing did not verify every board cell against the data repo. + +## Review and maintenance + +Reviewed on **2026-09-23**: site HTTP 200; issue #372; jevals-data CC BY 4.0. AI-assisted public-artifact review only. + +Related: [jevals](jevals.md), [agent-evals](agent-evals.md), [chinese-workflow-decision-bench](chinese-workflow-decision-bench.md).