Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 1 addition & 2 deletions docs/results/dashboard.html

Large diffs are not rendered by default.

2 changes: 1 addition & 1 deletion docs/results/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,4 +17,4 @@ For model and backend recommendations, see [Model Guide](../MODEL_GUIDE.md).
- [native-vs-prompt.md](raw/native-vs-prompt.md) — llama-server native FC vs prompt-injected, reforged only
- [reasoning-replay.md](raw/reasoning-replay.md) — reasoning_replay policy comparison (none / keep-last / full) per config

*Generated 2026-08-15 15:53*
*Generated 2026-08-17 17:14*
2 changes: 1 addition & 1 deletion docs/results/raw/native-vs-prompt.md
Original file line number Diff line number Diff line change
Expand Up @@ -324,4 +324,4 @@ Eval generations (older runs carried forward, superscript-tagged):
¹ gen 1 — v0.6.0 suite — incl. Anthropic ablation (commit 2b05dc4, 2026-05-08)
² gen 2 — v0.7.0 lineup refresh (8–14B) + 32GB tier debut (v0.7.4) (commit 655e1f6, 2026-05-22)

*Generated 2026-08-15 15:53*
*Generated 2026-08-17 17:14*
2 changes: 1 addition & 1 deletion docs/results/raw/reasoning-replay.md
Original file line number Diff line number Diff line change
Expand Up @@ -693,4 +693,4 @@ Eval generations (older runs carried forward, superscript-tagged):
¹ gen 1 — v0.6.0 suite — incl. Anthropic ablation (commit 2b05dc4, 2026-05-08)
² gen 2 — v0.7.0 lineup refresh (8–14B) + 32GB tier debut (v0.7.4) (commit 655e1f6, 2026-05-22)

*Generated 2026-08-15 15:53*
*Generated 2026-08-17 17:14*
13 changes: 12 additions & 1 deletion docs/results/raw/reforged-vs-bare.md
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,17 @@ Qwen3.5-35B-A3B-Q4_K_M LS/N [bare:full]² 12.2% 97.5% 12.5% 100%
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
```

## Inkling-Small-UD-IQ4_XS (llamaserver/native)

```
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Model/Backend Scr VAcc Cmp Eff Wst Spd N rel arg tsl b2s s3s crt srn err dgr dge art grs iar rel_s arg_s tsl_s b2s_s s3s_s crt_s srn_s err_s dgr_s dge_s art_s grs_s iar_s
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Inkling-Small-UD-IQ4_XS LS/N [reforged@max] 91.3% 94.6% 96.5% 76% 1.6 124.4s 50 100 100 100 100 98 100 100 100 98 96 50 52 92 100 100 100 100 100 100 100 98 96 98 36 62 98
Inkling-Small-UD-IQ4_XS LS/N [bare@max] 57.6% 93.3% 61.8% 84% 1.1 113.6s 50 2 80 66 88 86 88 90 0 74 60 28 24 70 0 86 68 78 88 86 88 0 80 60 34 20 54
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
```

## gemma-4-31B-it-Q4_K_M (llamaserver/native)

```
Expand Down Expand Up @@ -891,4 +902,4 @@ Eval generations (older runs carried forward, superscript-tagged):
¹ gen 1 — v0.6.0 suite — incl. Anthropic ablation (commit 2b05dc4, 2026-05-08)
² gen 2 — v0.7.0 lineup refresh (8–14B) + 32GB tier debut (v0.7.4) (commit 655e1f6, 2026-05-22)

*Generated 2026-08-15 15:53*
*Generated 2026-08-17 17:14*
3 changes: 2 additions & 1 deletion docs/results/raw/reforged/all.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,6 +31,7 @@ gemma-4-31B-it-Q4_K_M LS/N [reforged:full] 90.9%
gemma-4-31B-it-Q4_K_M LS/N [reforged:keep-last] 90.9% 90.9% 100.0% 100% 0.1 13.4s 50 100 100 100 100 100 98 100 100 100 0 96 96 100 100 100 100 100 100 100 100 100 100 0 78 96 100
Qwen3.5-122B-A10B-Q4_K_M LS/N [reforged] 91.3% 91.4% 99.9% 100% 0.2 42.4s 50 100 100 100 100 100 100 100 98 100 94 16 82 100 100 100 100 100 100 100 100 100 100 84 26 74 100
gpt-oss-120b-Q4_K_M LS/N [reforged:full@high] 90.8% 90.9% 99.9% 76% 1.4 28.6s 50 100 100 100 100 98 100 100 100 98 38 72 78 100 100 100 100 100 96 100 100 100 94 36 68 84 100
Inkling-Small-UD-IQ4_XS LS/N [reforged@max] 91.3% 94.6% 96.5% 76% 1.6 124.4s 50 100 100 100 100 98 100 100 100 98 96 50 52 92 100 100 100 100 100 100 100 98 96 98 36 62 98
gemma-4-31B-it-Q4_K_M LS/N [reforged] 88.7% 88.7% 100.0% 100% 0.1 15.7s 50 100 100 100 100 100 100 100 100 100 0 100 70 100 100 100 100 100 100 100 100 100 100 0 70 66 100
Qwen3.6-35B-A3B-UD-Q4_K_M LS/N [reforged] 88.0% 88.0% 100.0% 100% 0.3 5.4s 50 100 100 100 100 100 98 98 100 96 18 86 58 100 100 100 100 96 100 100 100 100 92 18 82 46 100
Qwen3.5-27B-Q4_K_M LS/P [reforged:full]² 86.8% 86.8% 100.0% 100% 0.1 24.4s 50 100 100 100 100 100 100 100 100 100 42 10 78 100 100 100 100 100 100 100 100 100 100 36 10 80 100
Expand Down Expand Up @@ -149,4 +150,4 @@ Eval generations (older runs carried forward, superscript-tagged):
¹ gen 1 — v0.6.0 suite — incl. Anthropic ablation (commit 2b05dc4, 2026-05-08)
² gen 2 — v0.7.0 lineup refresh (8–14B) + 32GB tier debut (v0.7.4) (commit 655e1f6, 2026-05-22)

*Generated 2026-08-15 15:53*
*Generated 2026-08-17 17:14*
2 changes: 1 addition & 1 deletion docs/results/raw/reforged/by-backend.md
Original file line number Diff line number Diff line change
Expand Up @@ -151,4 +151,4 @@ Eval generations (older runs carried forward, superscript-tagged):
¹ gen 1 — v0.6.0 suite — incl. Anthropic ablation (commit 2b05dc4, 2026-05-08)
² gen 2 — v0.7.0 lineup refresh (8–14B) + 32GB tier debut (v0.7.4) (commit 655e1f6, 2026-05-22)

*Generated 2026-08-15 15:53*
*Generated 2026-08-17 17:14*
12 changes: 11 additions & 1 deletion docs/results/raw/reforged/by-family.md
Original file line number Diff line number Diff line change
Expand Up @@ -111,6 +111,16 @@ Qwen3.5-35B-A3B-Q4_K_M LS/P [reforged:full]² 82.8% 82.8% 100.0% 100%
--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
```

## Inkling-Small-UD-IQ4_XS

```
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Model/Backend Scr VAcc Cmp Eff Wst Spd N rel arg tsl b2s s3s crt srn err dgr dge art grs iar rel_s arg_s tsl_s b2s_s s3s_s crt_s srn_s err_s dgr_s dge_s art_s grs_s iar_s
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
Inkling-Small-UD-IQ4_XS LS/N [reforged@max] 91.3% 94.6% 96.5% 76% 1.6 124.4s 50 100 100 100 100 98 100 100 100 98 96 50 52 92 100 100 100 100 100 100 100 98 96 98 36 62 98
-------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------
```

## gemma4-31b

```
Expand Down Expand Up @@ -358,4 +368,4 @@ Eval generations (older runs carried forward, superscript-tagged):
¹ gen 1 — v0.6.0 suite — incl. Anthropic ablation (commit 2b05dc4, 2026-05-08)
² gen 2 — v0.7.0 lineup refresh (8–14B) + 32GB tier debut (v0.7.4) (commit 655e1f6, 2026-05-22)

*Generated 2026-08-15 15:53*
*Generated 2026-08-17 17:14*
4 changes: 2 additions & 2 deletions eval_results_v0.9.0.jsonl
Git LFS file not shown
6 changes: 3 additions & 3 deletions src/forge/clients/sampling_defaults.py
Original file line number Diff line number Diff line change
Expand Up @@ -116,9 +116,9 @@
# deliberate campaign selector; change only it between effort campaigns.
"DeepSeek-V4-Flash-0731-UD-Q4_K_XL": {"temperature": 1.0, "top_p": 0.95, "chat_template_kwargs": {"reasoning_effort": "low"}}, # https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash-0731
# Inkling-Small — Unsloth recommends T=1.0/top_p=1.0 for most tasks and
# includes min_p=0.0 in its llama.cpp recipes. Use high effort (0.9) for
# the DGE pressure smoke; xhigh/max both map to the benchmark's 0.99.
"Inkling-Small-UD-IQ4_XS": {"temperature": 1.0, "top_p": 1.0, "min_p": 0.0, "chat_template_kwargs": {"reasoning_effort": "high"}}, # https://unsloth.ai/docs/models/inkling
# includes min_p=0.0 in its llama.cpp recipes. This smoke uses max effort
# (0.99); xhigh and max map to the same template value.
"Inkling-Small-UD-IQ4_XS": {"temperature": 1.0, "top_p": 1.0, "min_p": 0.0, "chat_template_kwargs": {"reasoning_effort": "max"}}, # https://unsloth.ai/docs/models/inkling
# gpt-oss-120b — OpenAI open-weight MoE (117B total, 5.1B active). Reasoning model with three
# discrete levels: "low" / "medium" / "high", controlled via chat_template_kwargs.reasoning_effort
# (per llama.cpp guide: `--chat-template-kwargs '{"reasoning_effort": "high"}'`). The registry
Expand Down
65 changes: 48 additions & 17 deletions tests/eval/batch_eval.py
Original file line number Diff line number Diff line change
Expand Up @@ -117,13 +117,26 @@ def _glimmer_server_recipe(reasoning_strength: str) -> _BatchServerRecipe:
"--reasoning-budget", "32768", "--reasoning-format", "auto",
"--no-prefill-assistant",
))
_INKLING_SMALL_RPC_SERVER_RECIPE = _BatchServerRecipe((
"--fit", "off",
"-b", "512", "-ub", "128",
"--cache-type-k", "f16", "--cache-type-v", "f16",
"--no-mmap", "-fa", "on",
"--parallel", "1",
))

_DEEPSEEK_V4_MODEL = "DeepSeek-V4-Flash-0731-UD-Q4_K_XL"
_DEEPSEEK_V4_GGUF = f"{_DEEPSEEK_V4_MODEL}-00001-of-00005.gguf"
_DEEPSEEK_V4_SAMPLING: dict[str, Any] = get_sampling_defaults(_DEEPSEEK_V4_MODEL)
_DEEPSEEK_V4_REASONING_LEVEL: str = _DEEPSEEK_V4_SAMPLING[
"chat_template_kwargs"
]["reasoning_effort"]
_INKLING_SMALL_MODEL = "Inkling-Small-UD-IQ4_XS"
_INKLING_SMALL_GGUF = f"{_INKLING_SMALL_MODEL}-00001-of-00004.gguf"
_INKLING_SMALL_SAMPLING: dict[str, Any] = get_sampling_defaults(_INKLING_SMALL_MODEL)
_INKLING_SMALL_REASONING_LEVEL: str = _INKLING_SMALL_SAMPLING[
"chat_template_kwargs"
]["reasoning_effort"]

# Effective reasoning levels for model configurations with an explicitly
# controlled effort axis. "default" remains reserved for configurations that
Expand Down Expand Up @@ -235,6 +248,7 @@ class BatchConfig:
# explicit param set (its keys match LlamafileClient kwargs).
reasoning_level: str = "default"
sampling_override: dict[str, Any] | None = None
requires_rpc: bool = False


# Ollama configs: 10 instruct models, native FC, stream
Expand Down Expand Up @@ -385,6 +399,20 @@ class BatchConfig:
gguf_filename=_DEEPSEEK_V4_GGUF,
server_recipe=_DEEPSEEK_V4_RPC_SERVER_RECIPE,
reasoning_level=_DEEPSEEK_V4_REASONING_LEVEL,
requires_rpc=True,
),
]

INKLING_SMALL_RPC_CONFIGS: list[BatchConfig] = [
BatchConfig(
model=_INKLING_SMALL_MODEL,
backend="llamaserver",
mode="native",
think=None,
gguf_filename=_INKLING_SMALL_GGUF,
server_recipe=_INKLING_SMALL_RPC_SERVER_RECIPE,
reasoning_level=_INKLING_SMALL_REASONING_LEVEL,
requires_rpc=True,
),
]

Expand All @@ -404,6 +432,7 @@ class BatchConfig:
"qwen38": QWEN38_CONFIGS,
"qwen38-medium": _QWEN38_EFFORT_CONFIGS["medium"],
"qwen38-low": _QWEN38_EFFORT_CONFIGS["low"],
"inkling-small-rpc": INKLING_SMALL_RPC_CONFIGS,
"new-models": NEW_MODEL_CONFIGS,
"new-models-native": [c for c in NEW_MODEL_CONFIGS if c.mode == "native"],
"new-models-prompt": [c for c in NEW_MODEL_CONFIGS if c.mode == "prompt"],
Expand Down Expand Up @@ -432,7 +461,7 @@ def _load_rpc_topology(path: Path) -> LlamaCppRpcConfig:
return LlamaCppRpcConfig(**topology_data)


def _attach_deepseek_rpc_topology(
def _attach_rpc_topology(
configs: list[BatchConfig], rpc: LlamaCppRpcConfig,
) -> list[BatchConfig]:
"""Attach machine-local RPC values without mutating the registry config."""
Expand All @@ -441,7 +470,7 @@ def _attach_deepseek_rpc_topology(
config,
server_recipe=replace(config.server_recipe, rpc=rpc),
)
if config.model == _DEEPSEEK_V4_MODEL else config
if config.requires_rpc else config
for config in configs
]

Expand Down Expand Up @@ -1068,13 +1097,14 @@ async def run_batch(
"batch_eval supports only managed backends "
f"{sorted(supported_backends)}; unsupported: {unsupported_backends}"
)
if any(
config.model == _DEEPSEEK_V4_MODEL
and config.server_recipe.rpc is None
for config in configs
):
missing_rpc = [
config.model for config in configs
if config.requires_rpc and config.server_recipe.rpc is None
]
if missing_rpc:
raise ValueError(
"DeepSeek V4 RPC batch config requires an attached RPC topology"
"RPC batch config requires an attached RPC topology: "
+ ", ".join(missing_rpc)
)

if scenario_names:
Expand Down Expand Up @@ -1134,10 +1164,7 @@ async def run_batch(
f"{'='*70}",
flush=True,
)
if (
config.model == _DEEPSEEK_V4_MODEL
and config.server_recipe.rpc is not None
):
if config.requires_rpc and config.server_recipe.rpc is not None:
_print_rpc_recipe(config, models_dir, budget_mode, manual_tokens)

# ── Dry run ───────────────────────────────────────
Expand Down Expand Up @@ -1380,7 +1407,7 @@ async def main() -> None:
"--rpc-topology",
type=str,
default=None,
help="JSON topology file required by --config deepseek-v4-rpc.",
help="JSON topology file required by a named RPC config set.",
)
parser.add_argument(
"--scenario", nargs="*",
Expand Down Expand Up @@ -1453,16 +1480,20 @@ async def main() -> None:
configs = [c for c in configs if args.model in c.model]
if not configs:
parser.error(f"No configs match --model '{args.model}' in set '{args.config}'")
if args.config == "deepseek-v4-rpc":
rpc_config_sets = {"deepseek-v4-rpc", "inkling-small-rpc"}
if args.config in rpc_config_sets:
if args.rpc_topology is None:
parser.error("--config deepseek-v4-rpc requires --rpc-topology")
parser.error(f"--config {args.config} requires --rpc-topology")
try:
rpc_topology = _load_rpc_topology(Path(args.rpc_topology))
except (OSError, ValueError, TypeError, KeyError) as exc:
parser.error(f"cannot load --rpc-topology: {exc}")
configs = _attach_deepseek_rpc_topology(configs, rpc_topology)
configs = _attach_rpc_topology(configs, rpc_topology)
elif args.rpc_topology is not None:
parser.error("--rpc-topology is only valid with --config deepseek-v4-rpc")
parser.error(
"--rpc-topology is only valid with --config "
+ " or ".join(sorted(rpc_config_sets))
)
output_path = Path(args.output) if args.output else Path("eval_results.jsonl")

if args.scenario:
Expand Down
20 changes: 10 additions & 10 deletions tests/eval/publication.py
Original file line number Diff line number Diff line change
Expand Up @@ -57,20 +57,20 @@
VIEW_NAMES = ("history", "snapshot", "latest")

PINNED_VIEW_COUNTS = {
"history": 544_700,
"snapshot": 404_300,
"latest": 286_000,
"history": 547_300,
"snapshot": 406_900,
"latest": 288_600,
}
PINNED_VIEW_METRIC_COUNTS = {
"history": (339_256, 454_184, 454_184),
"snapshot": (250_309, 337_641, 337_641),
"latest": (199_156, 252_956, 252_956),
"history": (341_192, 456_242, 456_242),
"snapshot": (252_245, 339_699, 339_699),
"latest": (201_092, 255_014, 255_014),
}
PINNED_SNAPSHOT_SOURCE_IDENTITY_SHA256 = (
"97060a3cac048a05b371f7bfb25ab6ba5b294394a8edc53c16a1d5bc1bd3d4b8"
"f19679e627ff3ca5e339c38f6564dfc9f215f95adbe824036dad58667f0bfe39"
)
PINNED_SNAPSHOT_SOURCE_PAYLOAD_SHA256 = (
"934ea927f9d9c8a2754404bca079ded815c901c53508a682e9fb27ede60fab75"
"6b0a9e85b6d4b863c19117b549f92e7127b41e564945d722f0c6e5c037717c98"
)


Expand Down Expand Up @@ -337,8 +337,8 @@ class SourceSpec:
"v0.9.0",
3,
CANONICAL_DIALECT,
26_000,
"84f2e8ae97136c65b17eb79c0228ada324a10f15d77f921074ef4e413c998aad",
28_600,
"50b2ae49a8b9161711289c60dc060a3403dbf2534ec1cbb6d5834e3f384fa8c3",
),
)

Expand Down
Loading
Loading