Release: 2026-09-18, suite 0.1.0. Tabs: noul, choice, score.
Jebadiah is an open-weights (Apache-2.0) decision model in the same shape as Jev. A typed question goes in, and the probability of every option label comes out, read at one position with nothing generated. Sizes: 27B (on Qwen3.8-27B), 9B v2 and 4B v2 (on Qwen3.5). Weights, code and run records: https://huggingface.co/frontier-infra/jebadiah-27b and https://github.com/getainode/jebadiah
To run it with your harness, self-hosted: the repo ships a standalone server that speaks Jev's /v1/systemone wire (server/ in the repo), so an existing Jev client works by changing its endpoint. It also runs with llama.cpp via the GGUF builds (for example frontier-infra/jebadiah-27b-GGUF, where scripts/decide_gguf.py shows the exact prompt and the logprob read), or on AINode, an open-source local AI platform.
Our own numbers, run with your suite files and our reimplementation of your Decision Score (not your harness): 27B PubMedQA 70.4, Banking77 68.4, HelpSteer2 19.9. Disclosure: HelpSteer2 is not zero-shot for Jebadiah, because its training pool includes the HelpSteer2 train split. The 300 board items are from your selection, and we have not checked them against that split. List it or mark it however you think fits.
Question: your policy prices every row at list price, and there is no list price for self-hosted open weights. How would you like that handled: a GPU-hour figure, a hosted endpoint's price, or unpriced?
I maintain Jebadiah.
Release: 2026-09-18, suite 0.1.0. Tabs: noul, choice, score.
Jebadiah is an open-weights (Apache-2.0) decision model in the same shape as Jev. A typed question goes in, and the probability of every option label comes out, read at one position with nothing generated. Sizes: 27B (on Qwen3.8-27B), 9B v2 and 4B v2 (on Qwen3.5). Weights, code and run records: https://huggingface.co/frontier-infra/jebadiah-27b and https://github.com/getainode/jebadiah
To run it with your harness, self-hosted: the repo ships a standalone server that speaks Jev's
/v1/systemonewire (server/in the repo), so an existing Jev client works by changing its endpoint. It also runs with llama.cpp via the GGUF builds (for examplefrontier-infra/jebadiah-27b-GGUF, wherescripts/decide_gguf.pyshows the exact prompt and the logprob read), or on AINode, an open-source local AI platform.Our own numbers, run with your suite files and our reimplementation of your Decision Score (not your harness): 27B PubMedQA 70.4, Banking77 68.4, HelpSteer2 19.9. Disclosure: HelpSteer2 is not zero-shot for Jebadiah, because its training pool includes the HelpSteer2 train split. The 300 board items are from your selection, and we have not checked them against that split. List it or mark it however you think fits.
Question: your policy prices every row at list price, and there is no list price for self-hosted open weights. How would you like that handled: a GPU-hour figure, a hosted endpoint's price, or unpriced?
I maintain Jebadiah.