Skip to content

A model deserves a row: Jebadiah (open-weights System One replica) #1

Description

@webdevtodayjason

Release: 2026-09-18, suite 0.1.0. Tabs: noul, choice, score.

Jebadiah is an open-weights (Apache-2.0) decision model in the same shape as Jev. A typed question goes in, and the probability of every option label comes out, read at one position with nothing generated. Sizes: 27B (on Qwen3.8-27B), 9B v2 and 4B v2 (on Qwen3.5). Weights, code and run records: https://huggingface.co/frontier-infra/jebadiah-27b and https://github.com/getainode/jebadiah

To run it with your harness, self-hosted: the repo ships a standalone server that speaks Jev's /v1/systemone wire (server/ in the repo), so an existing Jev client works by changing its endpoint. It also runs with llama.cpp via the GGUF builds (for example frontier-infra/jebadiah-27b-GGUF, where scripts/decide_gguf.py shows the exact prompt and the logprob read), or on AINode, an open-source local AI platform.

Our own numbers, run with your suite files and our reimplementation of your Decision Score (not your harness): 27B PubMedQA 70.4, Banking77 68.4, HelpSteer2 19.9. Disclosure: HelpSteer2 is not zero-shot for Jebadiah, because its training pool includes the HelpSteer2 train split. The 300 board items are from your selection, and we have not checked them against that split. List it or mark it however you think fits.

Question: your policy prices every row at list price, and there is no list price for self-hosted open weights. How would you like that handled: a GPU-hour figure, a hosted endpoint's price, or unpriced?

I maintain Jebadiah.

Activity

  1. Jevals commented on Sep 29, 2026

    @Jevals
    Owner

    Thanks for putting this together, and for running the suite yourselves first. Yes, I'd like to add Jebadiah.

    How I'd handle it:

    • I'll run it through our own harness before it goes on the board, same as every other row, so the site numbers come from our run and not a reimplementation. Your server speaking /v1/systemone makes that easy.
    • Pricing: self-hosted rows will show as unpriced, marked self-hosted, with the hardware I ran it on. I'd rather not invent a GPU-hour number. Latency will be from our run too, with the hardware next to it so nobody compares it straight against a hosted API.
    • HelpSteer2: thanks for flagging the train split. I'll mark that cell as trained-on so it doesn't read as zero-shot.

    If you have an endpoint I can point the run at, that's the fastest route. Otherwise I'll bring up the server from the repo. Should I start with the 27B?

  2. webdevtodayjason commented on Sep 29, 2026

    @webdevtodayjason
    Author

    Thanks, that all sounds right to me, and I'm glad you're running it through your own harness.

    Yes, start with the 27B. If you have room afterwards, I'd love the 9B v2 on there too, since that's the size most people actually run.

    I'd rather you bring it up yourself so the numbers are fully yours. The standalone server in the repo speaks /v1/systemone: server/ in getainode/jebadiah, with the model pinned to frontier-infra/jebadiah-27b at revision 54b89f6a1be4c4b260eb91ea7caf8cfed831a2d3. It applies the per-type temperatures from temperatures.json by default; send "calibration": "raw" if you want the untempered distribution. The 27B weights are about 56 GB in bf16, so a single 80 GB GPU is the comfortable fit.

    If setting it up is a hassle, tell me and I'll stand up a temporary private endpoint for the run instead.

    Unpriced and self-hosted with the hardware listed is exactly how I'd want it shown. Thanks again.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions