Fast, calibrated System One decision platform powered by TypeSafe Jev via OpenRouter. Sub-second logprob scoring, transfer curves, zero hallucinations.
-
Updated
Sep 25, 2026 - TypeScript
Fast, calibrated System One decision platform powered by TypeSafe Jev via OpenRouter. Sub-second logprob scoring, transfer curves, zero hallucinations.
Sales coaching runs on whoever happened to sit in on the call. This scores every call against the rubric instead: a verbatim quote behind every point given or withheld, and no score at all on calls that were never a pitch. Reviewing 15 calls a week by hand is 5 hours of someone's attention — this is one unattended run.
Rubric-based engine that scores spoken self-introduction transcripts on content, speech rate, grammar, vocabulary, clarity and engagement. Python + Flask.
Blind side-by-side human eval of two model responses on a weighted rubric — randomized panes to kill position bias, bootstrap CIs on the margin, inter-rater reliability. An eval you can't audit is a vote, not a measurement.
Automated hackathon submission judging — an async Celery pipeline verifies each repo, analyzes its structure, and scores it against the organizer's rubric with an LLM. Live leaderboard.
A structured portfolio showcasing AI evaluation samples, RLHF reasoning analysis, hallucination detection, rubric scoring, safety evaluation, and automation tooling. Built for platforms hiring AI evaluators and RLHF contractors.
A self-hosted hackathon submission and judging portal (Dogfood 2026)
To associate your repository with the rubric-scoring topic, visit your repo's landing page and select "manage topics."