Release: 2026-09-18, suite 0.1.0. Tabs: noul, choice, score.
Eikos is an open-weights (MIT) typed-decision model in the same shape as Jev. A state and a typed question (noul, choice or score) go in, and the probability of every option comes out of one forward pass, read from the option-label logits with nothing generated.
Sizes: 27B (fine-tuned from Qwen3.8-27B) and 4B (Qwen3.5-4B), each with FP8 and INT4 builds and MLX builds for Apple Silicon.
To run it with your harness, self-hosted: serve_vllm.sh + serve.py from the model repo start a server that speaks Jev's /v1/systemone wire, so a Jev client works by changing its endpoint. It needs one GPU and vLLM ≥ 0.30.
Our own numbers, for Eikos-27B-FP8. We used your suite files and our reimplementation of your Decision Score, not your harness: 300 items × 5 repeats, one request at a time.
| task |
Decision Score (95% CI) |
accuracy |
| PubMedQA |
71.4 [63.4, 78.5] |
0.927 |
| Banking77 |
66.8 [61.7, 71.8] |
0.789 |
| HelpSteer2 |
8.3 [−4.8, 20.1] |
0.400 |
Run records in your log format, the scripts, and the protocol we fixed before running: https://gist.github.com/caiovicentino/b5969ca7a71f34a9fc50519915328ef7
- The states we sent are yours byte for byte: all 900
state_sha256 values match. Choice orders come from the code in your README.
- The reimplementation recomputes all 21 rows of the 2026-09-18 board, with a maximum Decision Score difference of 0.0005.
- It is zero-shot on all three tasks. None of these datasets is in the training data, and an 8-gram check against the released training set flags nothing. Banking77 is in our own evaluation battery, not in training.
- Weak spot: it is underconfident here. On PubMedQA its mean confidence is 0.85 at 0.93 accuracy, and it never goes above ~0.96, so its coverage at your 0.96 choice gate is 0.
On price, the same question as #1: there is no list price for self-hosted weights. Our latencies (p50 124–247 ms) are on one RTX PRO 6000 over loopback. Happy with whatever you decide there.
I maintain Eikos.
Release: 2026-09-18, suite 0.1.0. Tabs: noul, choice, score.
Eikos is an open-weights (MIT) typed-decision model in the same shape as Jev. A state and a typed question (noul, choice or score) go in, and the probability of every option comes out of one forward pass, read from the option-label logits with nothing generated.
Sizes: 27B (fine-tuned from Qwen3.8-27B) and 4B (Qwen3.5-4B), each with FP8 and INT4 builds and MLX builds for Apple Silicon.
To run it with your harness, self-hosted:
serve_vllm.sh+serve.pyfrom the model repo start a server that speaks Jev's/v1/systemonewire, so a Jev client works by changing its endpoint. It needs one GPU and vLLM ≥ 0.30.Our own numbers, for Eikos-27B-FP8. We used your suite files and our reimplementation of your Decision Score, not your harness: 300 items × 5 repeats, one request at a time.
Run records in your log format, the scripts, and the protocol we fixed before running: https://gist.github.com/caiovicentino/b5969ca7a71f34a9fc50519915328ef7
state_sha256values match. Choice orders come from the code in your README.On price, the same question as #1: there is no list price for self-hosted weights. Our latencies (p50 124–247 ms) are on one RTX PRO 6000 over loopback. Happy with whatever you decide there.
I maintain Eikos.