Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
364 commits
Select commit Hold shift + click to select a range
c7445cc
feat(glm): continue an open assistant turn on the GLM-5.2 base renderer
enitimeago Sep 10, 2026
f9efa23
feat(olmoe): continue an open assistant turn
enitimeago Sep 10, 2026
e62dbed
feat(v4): continue an open assistant turn on deepseek_v4
enitimeago Sep 10, 2026
87c4156
feat(inkling): continue an open assistant turn
enitimeago Sep 10, 2026
82a3e5c
test(serve): only the trailing assistant turn is opened, every family
enitimeago Sep 10, 2026
9b1151d
test(serve): continuation reaches the engine and returns as content, …
enitimeago Sep 10, 2026
f461050
feat(kimi): continue an open assistant turn, framed in the engine
enitimeago Sep 10, 2026
2327856
feat(v41): continue an open assistant turn on deepseek_v41
enitimeago Sep 21, 2026
7d1fa95
test(glm53): pin the continuation across every template-expressible e…
enitimeago Sep 21, 2026
e57b487
tests(qwen38): normalize fallback test line endings
Stamina9 Sep 21, 2026
1212b0b
ci: upgrade Node.js 20 to 22 in release pipeline (Node 20 is EOL sinc…
Suraj2105-1 Sep 21, 2026
e514e60
ci: upgrade Alpine 3.21 to 3.24 in musl job (3.21 EOL Nov 1 2026)
Suraj2105-1 Sep 21, 2026
024bef5
fix(qwen38): reject pins whose attention rows were overwritten
bokiko Sep 21, 2026
be958c1
fix: resolve static analysis bugs in coli, autotune, family_registry,…
Suraj2105-1 Sep 21, 2026
682f62d
qwen36: split added tokens before the pre-tokenizer, and the three wh…
Sep 21, 2026
6f152cd
deepseek_v4: the numeric channel (SUBMIT logprobs=k, pin=1), so brio …
Sep 21, 2026
d99eb3f
Merge pull request #1654 from JustVugg/fix/qwen36-special-after-punct
JustVugg Sep 21, 2026
26e000b
tok.h: the GPT-2 pre-tokenizer family, so OLMoE tokenizes like HF
Sep 21, 2026
2edf748
Merge pull request #1655 from JustVugg/feat/deepseek-v4-brio
JustVugg Sep 21, 2026
6203a14
qwen36: offer the rest of the dense trunk to the VRAM placer (dnout, …
Sep 21, 2026
6719b2c
tests: two reads of available memory may differ, compare within a qua…
Sep 21, 2026
472450f
Merge pull request #1658 from JustVugg/fix/test-mem-available-two-reads
JustVugg Sep 21, 2026
4d96d7c
Merge pull request #1656 from JustVugg/fix/tok-gpt2-family
JustVugg Sep 21, 2026
5c6a318
qwen36, qwen38: max_tokens is a ceiling, not a target (#1641)
Sep 21, 2026
12e7101
Merge pull request #1659 from JustVugg/fix/serve-max-tokens-ceiling
JustVugg Sep 21, 2026
0e84b37
qwen36: measure the trunk placement at startup, withdraw it when the …
Sep 21, 2026
f9674b6
fix(comment): stale "qdw entry" reference after the g_qdw removal
jtinbergen Sep 21, 2026
3b2ebb3
Merge pull request #1657 from JustVugg/feat/qwen36-trunk-vram
JustVugg Sep 21, 2026
da7bc4f
release: 1.12.1
Sep 21, 2026
48c974d
gateway: POST /v1/systemone, the request and reply of Jev, served by …
Sep 21, 2026
66ffa1d
Merge pull request #1660 from JustVugg/release/1.12.1
JustVugg Sep 21, 2026
2a61e58
Merge remote-tracking branch 'origin/dev' into feat/systemone-jev-api
Sep 21, 2026
6954935
changelog: /v1/systemone under 1.12.1
Sep 21, 2026
dfec3a4
qwen36: integer dot products for the dense trunk and the routed experts
Sep 21, 2026
24dad80
Merge pull request #1662 from JustVugg/feat/systemone-jev-api
JustVugg Sep 21, 2026
391168c
Merge pull request #1664 from JustVugg/perf/qwen36-dense-idot
JustVugg Sep 21, 2026
3c164f6
changelog: 1.12.1 counts 32 pull requests, dated the day of the tag
Sep 21, 2026
87eb980
Merge pull request #1665 from JustVugg/release/1.12.1-notes
JustVugg Sep 22, 2026
e958c0e
fix(qwen38): merge dev while preserving prefix and dot-product depend…
bokiko Sep 22, 2026
4fb195c
fix(oracle): reconcile validation with current dev
ekalb81 Sep 22, 2026
0137b8f
qwen38: int8 trunk with integer dot products, vector FP8 expert kernel
Sep 22, 2026
b8e23ac
test_qwen38_native_weights: pin the FP8 dispatch to the table kernel
Sep 22, 2026
85b7943
qwen38: rebuild the trunk table per load; the fixture checks run the …
Sep 22, 2026
dd10808
qwen38: the trunk threshold and skip list are read per load, not cach…
Sep 22, 2026
6c78834
fix(v41): size gateway expert cache from the RAM budget
ZacharyZcR Sep 22, 2026
71622eb
fix(v41): clamp output budgets and reject oversized prompts
ZacharyZcR Sep 22, 2026
8d546fa
Merge pull request #1650 from bokiko/fix/qwen38-stale-pin
JustVugg Sep 22, 2026
42f640d
test(v41): give the protocol mock an explicit cache cap
ZacharyZcR Sep 22, 2026
0d8dd39
Merge origin/dev into perf/qwen38-idot-trunk
Sep 22, 2026
e82567a
test(systemone): resolve the scoring fixture during discovery
ZacharyZcR Sep 22, 2026
b6ac01e
cuda: reuse streaming MXFP4 weight scratch
ZacharyZcR Sep 22, 2026
8e8447c
qwen36: offload batched attention projections to CUDA
ZacharyZcR Sep 22, 2026
cde709c
docs: qwen38 int8 trunk and vector FP8 kernel, measured; 1.12.1 notes…
Sep 22, 2026
a65ea15
cuda: fix resident MXFP4 scale layout and accounting
ZacharyZcR Sep 22, 2026
cc27d98
kimi: keep CUDA SiTU expert intermediates on device
ZacharyZcR Sep 22, 2026
d99dc06
Merge pull request #1668 from JustVugg/perf/qwen38-idot-trunk
JustVugg Sep 22, 2026
f2c305f
Merge remote-tracking branch 'upstream/dev' into feat/kimi-cuda-situ-…
ZacharyZcR Sep 22, 2026
0fc47cc
qwen36: batch resident DeltaNet input projections
ZacharyZcR Sep 22, 2026
ba5ca29
cuda: reuse SiTU expert staging across calls
ZacharyZcR Sep 22, 2026
9d1c441
test: account for optional MXFP4 expert loader symbol
ZacharyZcR Sep 22, 2026
25507df
Merge dev and retain V4.1 budget and Qwen test targets
ZacharyZcR Sep 22, 2026
0977650
fix(cuda): release Qwen tier resources on shutdown
ZacharyZcR Sep 22, 2026
81bbae1
fix(cuda): reclaim partially uploaded expert tensors
ZacharyZcR Sep 22, 2026
1d6bfe2
fix(cuda): budget expert scale allocations separately
ZacharyZcR Sep 22, 2026
b9401a9
docs/qwen38: remove stale "no GPU backend" line (GPU tier merged in 1…
XBold Sep 22, 2026
bff0f12
README: align Qwen3.8 entries with the shipped CUDA tier (1.12.0)
XBold Sep 22, 2026
e9c45f6
fix(cuda): propagate Qwen expert collection failures
ZacharyZcR Sep 22, 2026
fe29def
fix(cuda): unwind failed Qwen tier initialization
ZacharyZcR Sep 22, 2026
8745651
qwen38: make the prefill chunk and its workspace runtime knobs
DebugSultan Sep 14, 2026
9c0c6bc
qwen38: lift the top-k ceiling off the expert load batch
DebugSultan Sep 22, 2026
c7f5b7f
qwen38: run QSA per position at prefill, and time it by the clock
DebugSultan Sep 22, 2026
33900c2
fix(cuda): unwind partially initialized device contexts
ZacharyZcR Sep 22, 2026
90ff72b
feat(serve): expose Prometheus scheduler metrics
ZacharyZcR Sep 22, 2026
6567e76
feat(serve): measure first output and engine call durations
ZacharyZcR Sep 22, 2026
9885e89
feat(tools): add reproducible HTTP serving benchmark
ZacharyZcR Sep 22, 2026
510cfa5
fix(cuda): stop parked issuers before teardown
ZacharyZcR Sep 22, 2026
2eb1975
test(cuda): group tier regression rules with related tests
ZacharyZcR Sep 22, 2026
f5ec92f
test(cuda): group tier regression rules with related tests
ZacharyZcR Sep 22, 2026
39adcc6
test(cuda): group tier regression rules with related tests
ZacharyZcR Sep 22, 2026
09f1403
test(cuda): group tier regression rules with related tests
ZacharyZcR Sep 22, 2026
127b077
test(cuda): group tier regression rules with related tests
ZacharyZcR Sep 22, 2026
b9c679a
test(cuda): cover FP8 and per-row int4 upload rollback
ZacharyZcR Sep 22, 2026
0e939fc
feat(bench): report latency-constrained request goodput
ZacharyZcR Sep 22, 2026
df6cf53
feat(bench): schedule fixed-rate arrivals and account for client backlog
ZacharyZcR Sep 22, 2026
f0f8dfc
fix(cuda): preserve active contexts on repeated initialization
ZacharyZcR Sep 22, 2026
cba5f31
fix(cuda): reject initialization of an active Qwen tier
ZacharyZcR Sep 22, 2026
2c442da
fix(bench): reject malformed choices and unsuccessful stream endings
ZacharyZcR Sep 22, 2026
f17038d
test(qwen36): verify each attention projection failure with poisoned …
ZacharyZcR Sep 22, 2026
aee54ce
bench(cuda): measure paired resident projection batching
ZacharyZcR Sep 22, 2026
35c39ca
test(qwen36): verify split prefill state and continued decode
ZacharyZcR Sep 22, 2026
24788ac
fix(serve): check cancellation before admitting an available slot
ZacharyZcR Sep 22, 2026
5ef694d
fix(serve): enforce zero queue limit for busy pinned slots
ZacharyZcR Sep 22, 2026
282b743
feat(bench): separate warmup from measured serving requests
ZacharyZcR Sep 22, 2026
44150b6
build: avoid conflicts between CUDA benchmark and test targets
ZacharyZcR Sep 22, 2026
de27ba0
fix(serve): enforce queue deadlines before slot admission
ZacharyZcR Sep 22, 2026
03f6aa8
fix(bench): reject choice chunks after stream completion
ZacharyZcR Sep 22, 2026
3059afb
test(bench): cover partial engine failure and recovery
ZacharyZcR Sep 22, 2026
48795d4
test(cuda): verify attention prefill continuation and KV state
ZacharyZcR Sep 22, 2026
1a05695
fix(serve): admit requests to unreserved free slots
ZacharyZcR Sep 22, 2026
b1c9d52
fix(bench): account for legacy function call output
ZacharyZcR Sep 22, 2026
965d175
Add reproducible Poisson arrivals to HTTP benchmark
ZacharyZcR Sep 22, 2026
2608b0f
Admit unreserved free slots when the waiting queue is full
ZacharyZcR Sep 22, 2026
0d91cdc
Stop stream keepalive threads on generation errors and cancellation
ZacharyZcR Sep 22, 2026
73f6a6b
Merge pull request #1669 from ZacharyZcR/fix/systemone-test-discovery
JustVugg Sep 22, 2026
9f49a80
Merge pull request #1670 from ZacharyZcR/fix/1666-v41-ram-budget
JustVugg Sep 22, 2026
f24abbc
Merge pull request #1671 from ZacharyZcR/fix/v41-output-budget
JustVugg Sep 22, 2026
3c0b387
Merge pull request #1673 from ZacharyZcR/perf/cuda-mxfp4-scratch
JustVugg Sep 22, 2026
8ca90cd
Merge pull request #1675 from ZacharyZcR/fix/cuda-mxfp4-scale-layout
JustVugg Sep 22, 2026
ccf32ec
Merge pull request #1679 from ZacharyZcR/fix/qwen-cuda-upload-rollback
JustVugg Sep 22, 2026
86f4567
Merge pull request #1680 from ZacharyZcR/fix/qwen-cuda-scale-budget
JustVugg Sep 22, 2026
62f6a47
Merge pull request #1682 from ZacharyZcR/fix/qwen-cuda-take-error
JustVugg Sep 22, 2026
6a25042
Merge pull request #1683 from ZacharyZcR/fix/qwen-cuda-init-cleanup
JustVugg Sep 22, 2026
bc54327
Merge pull request #1684 from ZacharyZcR/fix/cuda-context-init-unwind
JustVugg Sep 22, 2026
f969118
Merge pull request #1678 from ZacharyZcR/fix/qwen-cuda-tier-shutdown
JustVugg Sep 22, 2026
ae48037
Merge pull request #1674 from ZacharyZcR/perf/qwen36-batch-dense
JustVugg Sep 22, 2026
94cef8c
Merge pull request #1677 from ZacharyZcR/perf/qwen36-dnproj-prefill
JustVugg Sep 22, 2026
7edf6eb
Merge dev into feat/kimi-cuda-situ-pipeline: keep both MXFP4 scratch …
Sep 22, 2026
31b0905
CHANGELOG 1.12.1: the eighteen pull requests merged today
Sep 22, 2026
b197c56
coli: resolve the launcher's own symlink before deriving libexec (#1689)
Sep 22, 2026
b5944eb
CHANGELOG 1.12.1: #1691
Sep 22, 2026
52be8e1
Merge pull request #1686 from DebugSultan/prefill-batching
JustVugg Sep 22, 2026
a4e18f9
Merge pull request #1687 from ZacharyZcR/feat/serve-prometheus-metrics
JustVugg Sep 22, 2026
ed7f535
Merge pull request #1688 from ZacharyZcR/feat/http-serving-benchmark
JustVugg Sep 22, 2026
3961574
Merge pull request #1681 from XBold/docs/qwen38-remove-stale-gpu-back…
JustVugg Sep 22, 2026
5fbf3a6
tools(qwen36): --down-bits, the mixed expert layout (int4 gate/up, in…
kreuzzelg Sep 15, 2026
87d4139
qwen36: read the mixed expert layout (int4 gate/up, int8 down) on the…
kreuzzelg Sep 15, 2026
e1d51be
ci(qwen36), docs: the mixed expert layout on the tiny fixture, and th…
kreuzzelg Sep 15, 2026
2b35814
CHANGELOG 1.12.1: #1559
Sep 22, 2026
40abd32
fix(serve): check every server->engine stdin write
monotophic Sep 22, 2026
d1c2bca
Merge pull request #1691 from JustVugg/fix/launcher-realpath
JustVugg Sep 22, 2026
a926d1b
Merge pull request #1676 from ZacharyZcR/feat/kimi-cuda-situ-pipeline
JustVugg Sep 22, 2026
9ddd00a
Merge pull request #1690 from JustVugg/release/1.12.1-notes-batch
JustVugg Sep 22, 2026
2051c18
Merge pull request #1559 from kreuzzelg/qwen36-down-bits
JustVugg Sep 22, 2026
9e2555a
Merge pull request #1651 from Suraj2105-1/fix/pyflakes-code-bugs
JustVugg Sep 22, 2026
ac94bb0
Merge pull request #1649 from Suraj2105-1/fix/alpine-3.21-eol-upgrade…
JustVugg Sep 22, 2026
5480ea5
Merge dev into fix/serve-tool-call-arguments-non-object: keep both to…
Sep 22, 2026
96ef114
Merge pull request #1647 from Suraj2105-1/fix/node20-eol-release-pipe…
JustVugg Sep 22, 2026
9fc4ca3
Merge dev into fix/windows-qwen36-cuda-dll: keep dev's qwen36 prerequ…
Sep 22, 2026
232b282
Merge dev into fix/macos-metal-gpu-discovery: keep the int8 trunk lin…
Sep 22, 2026
11106c9
Merge pull request #1624 from namespaceMarcello/fix/build-warnings
JustVugg Sep 22, 2026
0eb3a87
Merge pull request #1618 from bherald/docs/qwen38-vision-overview-202…
JustVugg Sep 22, 2026
9e85af7
CHANGELOG 1.12.1: the third wave of contributor pull requests
Sep 22, 2026
b6955b6
Merge pull request #1568 from Yoruxyv/feat/indonesian-dashboard-trans…
JustVugg Sep 22, 2026
4efbb8d
Merge pull request #1622 from Frank-zhu0404/fix/issue-1615-per-matrix…
JustVugg Sep 22, 2026
9e152d4
Merge pull request #1605 from ekalb81/fix/glm-oracle-validation
JustVugg Sep 22, 2026
6826c00
refactor(quant): move FP8_BLOCK/fp8_nblk into fp8_format.h
monotophic Sep 3, 2026
0ad3eef
feat(colibri): add fmt=8 arms to the CPU absorb path
monotophic Sep 3, 2026
9c7ae38
feat(cuda): add fmt=8 (fp8-e4m3) decode to the absorb kernels
monotophic Sep 3, 2026
ea8a7e4
test(fp8): e2e serve-batch pin, refusal canary, and mint-to-load regr…
monotophic Sep 3, 2026
7d79ad1
feat(api): accept and ignore a per-request seed
monotophic Sep 22, 2026
d8c4209
Merge pull request #1597 from kevin9327/fix/serve-tool-call-arguments…
JustVugg Sep 22, 2026
2020aec
fix(olmoe): wait for an expert already being read instead of reading …
namespaceMarcello Sep 22, 2026
5f1cb69
Merge pull request #1580 from LTCjRet/fix/windows-qwen36-cuda-dll
JustVugg Sep 22, 2026
2cc1040
Merge pull request #1556 from karlem/fix/macos-metal-gpu-discovery
JustVugg Sep 22, 2026
b94accc
Merge pull request #1692 from JustVugg/release/1.12.1-notes-wave3
JustVugg Sep 22, 2026
836310f
Merge dev into feat/glm53-assistant-continuation: keep the max_tokens…
Sep 22, 2026
7dbe7e3
CHANGELOG 1.12.1: #1402
Sep 22, 2026
4d35744
perf(sse41): vectorize grouped-int4 GEMV
Aug 30, 2026
e09b3db
fix(sse41): preserve grouped-int4 scalar order
Aug 30, 2026
a3ec030
fix(build): track SSE4.1 kernel header dependencies
Sep 20, 2026
8212a39
qwen36: keep the prefill echo out of the segment build
namespaceMarcello Sep 22, 2026
6497713
Merge pull request #1694 from JustVugg/release/1.12.1-notes-1402
JustVugg Sep 22, 2026
b02b801
Merge pull request #1402 from enitimeago/feat/glm53-assistant-continu…
JustVugg Sep 22, 2026
11eb32e
fix(web): restore the reasoning stream the redesign dropped
kevin9327 Sep 22, 2026
b95b942
docs(fp8): correct stale anchors, an unreachable test rule and a wron…
monotophic Sep 23, 2026
4159a23
docs(fp8): FORMATS.md names symbols, so stop promising line numbers
monotophic Sep 23, 2026
049d444
Opt-in exact verify mode for speculative batches (COLI_EXACT_VERIFY=1…
rybruscoe Sep 7, 2026
78018c3
Makefile: list exact_dot.h as a colibri prerequisite
rybruscoe Sep 21, 2026
a13b3cd
exact verify (#689): document the quantised-KV fallback and the untes…
rybruscoe Sep 23, 2026
6035bcd
CHANGELOG 1.12.1: #1693, #1695, #1697
Sep 23, 2026
bf7c767
Merge pull request #1695 from namespaceMarcello/fix/qwen36-echo-guard
JustVugg Sep 23, 2026
4399eef
Merge pull request #1697 from kevin9327/fix/web-reasoning-stream-dropped
JustVugg Sep 23, 2026
aa2d295
Merge pull request #1693 from namespaceMarcello/fix/olmoe-inflight-dedup
JustVugg Sep 23, 2026
577d8eb
feat(tools): add three-engine serving baseline campaigns
ZacharyZcR Sep 23, 2026
5d6671f
Merge pull request #1704 from JustVugg/release/1.12.1-notes-wave4
JustVugg Sep 23, 2026
1f82285
CHANGELOG 1.12.1: #1705
Sep 23, 2026
926be8e
Merge pull request #1705 from ZacharyZcR/feat/three-engine-serving-ba…
JustVugg Sep 23, 2026
30c70ad
Merge pull request #1706 from JustVugg/release/1.12.1-notes-1705
JustVugg Sep 23, 2026
4e8ef20
fix(v4): rebuild the unit objects when the build flags change
namespaceMarcello Sep 23, 2026
1165bcd
fix(glm53): close the Vulkan status fprintf after the ternary
wittchen Sep 23, 2026
4653c34
fix(web): regenerate the last turn with its pictures
kevin9327 Sep 23, 2026
c64c8d6
fix(serve): a tool whose function is not an object answered 500
kevin9327 Sep 23, 2026
6d616c5
fix(coli): Windows GPU probe missed Kimi CUDA_DLL and HIP hosts
kevin9327 Sep 23, 2026
1627377
fix(serve): kimi, inkling, olmoe clamp max_tokens to the context
kevin9327 Sep 23, 2026
b1b9c60
fix(planner): export K3_EXPERT_GB and GLM53_EXPERT_GB from the plan
kevin9327 Sep 23, 2026
69c3b0f
test(glm53): rename the harnesses out of the unittest glob, run the t…
namespaceMarcello Sep 23, 2026
e1feb15
fix(brio): pin the shared state in the options form so per-question r…
benmaster82 Sep 23, 2026
dd1cf4e
Merge branch 'dev' into sse41-tier-upstream-next-v2 (+25 commits)
jtinbergen Sep 23, 2026
d13dec7
feat(brio): let the web cancel a scoring run, dedupe the CLI options
benmaster82 Sep 23, 2026
2bc08d0
kimi_k3: keep K3_CUDA the only engine switch; coli maps --gpu onto it
Sep 23, 2026
b8869a2
Merge pull request #1710 from kevin9327/fix/serve-tool-function-non-o…
JustVugg Sep 23, 2026
3431e87
Merge pull request #1712 from kevin9327/fix/k3-ink-olmoe-max-tokens-c…
JustVugg Sep 23, 2026
72e49d2
Merge pull request #1713 from kevin9327/fix/planner-kimi-expert-gb
JustVugg Sep 23, 2026
c0a3fb4
Merge pull request #1707 from namespaceMarcello/fix/v4-unit-flags-stamp
JustVugg Sep 23, 2026
c1746e0
Merge pull request #1708 from wittchen/fix/glm53-vk-fprintf
JustVugg Sep 23, 2026
77e6a0d
Merge dev: keep dev's build-flags stamp (#1707) and serve_budget.h (#…
Sep 23, 2026
bf4ac29
Merge dev: keep dev's build-flags stamp (#1707) and serve_budget.h (#…
Sep 23, 2026
c07f62c
Merge dev: fallback diagnostics on the heap-allocated batch of #1686
Sep 23, 2026
f89040b
Merge pull request #1711 from kevin9327/r8b-win-cuda-dll-probe
JustVugg Sep 23, 2026
7a82be4
Merge pull request #1709 from kevin9327/fix/web-regenerate-drops-images
JustVugg Sep 23, 2026
d9e7bea
Merge pull request #1395 from rybruscoe/exact-verify-689
JustVugg Sep 23, 2026
b4e51d5
Merge dev: colibri prerequisites keep exact_dot.h (#1395) and fp8_for…
Sep 23, 2026
1d7defa
Merge pull request #1102 from monotophic/f8/absorb-fmt8
JustVugg Sep 23, 2026
621ae67
Merge dev: keep #1102's fp8_format.h, #1395's exact_dot.h and the new…
Sep 23, 2026
66167b6
Merge pull request #1646 from Stamina9/codex/fix-qwen38-parallel-fall…
JustVugg Sep 23, 2026
97622ad
Merge pull request #1721 from monotophic/serve/checked-engine-writes
JustVugg Sep 23, 2026
de563da
Merge pull request #1720 from monotophic/api/seed-accepted
JustVugg Sep 23, 2026
017fe1b
Merge pull request #1719 from benmaster82/feat/brio-web-cancel-and-cl…
JustVugg Sep 23, 2026
e2de147
Merge pull request #1714 from benmaster82/fix/brio-options-state-pin
JustVugg Sep 23, 2026
7a5bdee
Merge pull request #1715 from namespaceMarcello/test/glm53-harness-di…
JustVugg Sep 23, 2026
8db487f
CHANGELOG 1.12.1: sixteen more contributor pull requests
Sep 23, 2026
3b24221
refactor(sse41): collect grouped-int4 kernels in the shared header
Sep 23, 2026
4aac292
Merge pull request #1286 from cameron/perf/sse41-grouped-int4-upstream
JustVugg Sep 23, 2026
33422c3
Merge pull request #1722 from JustVugg/release/1.12.1-notes-wave5
JustVugg Sep 23, 2026
9342b47
qwen36: decode added tokenizer tokens
GenericRikka Sep 23, 2026
93bc084
qwen36: decode only the non-special added tokens
Sep 23, 2026
20183ac
Merge pull request #1724 from JustVugg/fix/qwen36-decode-added-tokens
JustVugg Sep 23, 2026
09b78b1
CHANGELOG 1.12.1: #1724
Sep 23, 2026
a2e578c
Merge pull request #1725 from JustVugg/release/1.12.1-notes-1724
JustVugg Sep 23, 2026
c703d26
fix(inkling): measure RAM on Windows so auto-cap is not 16 experts
kevin9327 Sep 23, 2026
37d9367
tests: link backend_vulkan.o into the engine-including tests under VK=1
crichalchemist Sep 24, 2026
2c7979f
Merge branch 'dev' into sse41-tier-upstream-next-v2 (+214 commits)
jtinbergen Sep 24, 2026
f4ee2e0
feat(web): continue a trailing assistant turn from the chat UI
enitimeago Sep 21, 2026
6923b7d
feat(web): show Continue only on an incomplete assistant turn
enitimeago Sep 21, 2026
eb592cc
feat(api): report continue_assistant in the authed /health
enitimeago Sep 23, 2026
7743fb4
feat(web): offer Continue only when /health reports continue_assistant
enitimeago Sep 23, 2026
96467ec
feat(web): keep each reply's finish reason on the message itself
enitimeago Sep 23, 2026
fa86687
tools: add FP4 expert matmul microbench (SIMD vs scalar arm)
mfethe1 Sep 24, 2026
da68961
dsv4: NEON arm for coli_fp4_matmul_batch_rows16_order (#1696)
Sep 24, 2026
9ba9ed8
fix(v4): rebuild backend_cuda_dsv4.o when the nvcc command changes
bokiko Sep 24, 2026
75d2677
dsv4: NEON arm for coli_fp4_matvec_rows16_order (decode path, #1696)
Sep 24, 2026
49e72cd
serve: give the engine its graceful exit — stdin drain, SIGBREAK, gro…
tarazum Sep 24, 2026
b6fbb65
make: write .build-config through printf on GNU Make 3.x (#1732)
Sep 24, 2026
4087d5a
ci: prove the NEON FP4 arms bit-exact against the scalar arm on the A…
Sep 24, 2026
6b5671c
Merge pull request #1717 from enitimeago/feat/web-continue
JustVugg Sep 24, 2026
a8de807
Merge pull request #1726 from kevin9327/fix/inkling-windows-ram-probe
JustVugg Sep 24, 2026
d23e362
Merge pull request #1728 from crichalchemist/vk-link-colibri-tests
JustVugg Sep 24, 2026
2516472
Merge pull request #1731 from bokiko/fix/v4-cuda-arch-stamp
JustVugg Sep 24, 2026
f5497c8
close(): decide the hard-stop ladder by poll(), not by TimeoutExpired
tarazum Sep 24, 2026
78d56c2
Merge pull request #1735 from JustVugg/fix/build-config-make381
JustVugg Sep 24, 2026
50c9e16
Merge pull request #1734 from tarazum/fix/heat-save-stdin-drain
JustVugg Sep 24, 2026
20e4582
Merge pull request #1730 from mfethe1/perf/dsv4-fp4-neon-arm
JustVugg Sep 24, 2026
53ca388
Merge pull request #1716 from jtinbergen/sse41-tier-upstream-next-v2
JustVugg Sep 24, 2026
5e11f46
CHANGELOG 1.12.1: eight more pull requests
Sep 24, 2026
4f119c0
ci: give the macOS make check 25 minutes
Sep 24, 2026
73fcbcf
ci: the Linux make check needs the same room
Sep 24, 2026
25b4a9f
docs(changelog): date 1.12.1 to the release day
Sep 24, 2026
a44865f
Merge pull request #1737 from JustVugg/ci/macos-timeout
JustVugg Sep 24, 2026
384beb4
Merge branch 'dev' into release/1.12.1-notes-wave6
JustVugg Sep 24, 2026
1f42b14
Merge pull request #1736 from JustVugg/release/1.12.1-notes-wave6
JustVugg Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 12 additions & 2 deletions .github/workflows/check.yml
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,9 @@ jobs:
linux:
name: Linux
runs-on: ubuntu-latest
timeout-minutes: 15
# make check takes about 14 minutes here now (14.9 at worst this week);
# at 15 the job was cut off with nothing failed.
timeout-minutes: 25
steps:
- uses: actions/checkout@v4
- name: make check
Expand Down Expand Up @@ -128,13 +130,21 @@ jobs:
python c/tools/make_edge_tiny_tokenizer.py --vocab-size 281 /tmp/glm53_stream-i4
cd c && COLI_GLM53_FIXTURE=/tmp/glm53_stream-i4 python -m unittest -v tests.test_glm53_dashboard
COLI_GLM53_FIXTURE=/tmp/glm53_stream-i4 python -m unittest -v tests.test_glm53_context_exceeded
# The two stdlib-only oracles used to be argparse scripts that `make
# test-python` collected as zero tests (#1700). They run here, on the
# fixtures generated above, token-exact against transformers in f32.
- name: Run the GLM-5.3 tiny oracles
run: |
cd c && GLM53_TINY=/tmp/glm53_tiny GLM53_MM_TINY=/tmp/glm53_mm python -m unittest -v tests.test_glm53_oracles

macos:
# clang; libomp for the threaded path (Makefile falls back to
# single-threaded automatically if it's ever missing).
name: macOS (colibri + V4 platform gate)
runs-on: macos-latest
timeout-minutes: 15
# `make check` here now takes about 14.5 minutes; at 15 the job was cut
# off inside the Python suite with no test failed, and showed as red.
timeout-minutes: 25
steps:
- uses: actions/checkout@v4
- name: install libomp
Expand Down
85 changes: 76 additions & 9 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -147,7 +147,7 @@ jobs:
musl:
name: musl libc (Alpine)
runs-on: ubuntu-latest
container: alpine:3.21
container: alpine:3.24
steps:
# #1430: colibri did not compile against musl, because malloc_trim is a
# glibc extension guarded on __linux__ rather than on __GLIBC__. Nothing in
Expand Down Expand Up @@ -547,6 +547,31 @@ jobs:
echo "cap=$cap: identical ($(cat ids_${cap}_1.txt))"
done
make tests/test_expert_ffn && ./tests/test_expert_ffn
- name: Mixed expert layout (int4 gs64 gate/up, int8 down) loads and is cap-independent
run: |
cd c
# The #1370 experiment knob: one slab per expert with int4 gate/up and int8
# down (2*inter*hidden bytes). The engine must tell it apart from int4 and
# int8 by size, read every matrix in its own format, and give the same ids
# at every cache capacity (cap=1 recycles the single slot after each expert,
# which is where a wrong slab offset would show). The ids are not compared to
# the torch oracle: int4 gate/up drift from f32 by design, as in the A/B above.
python3 tools/convert_qwen36.py --model qwen36_tiny64 --out qwen36_tiny64_d8 --ebits 4 --gs 64 --down-bits 8
python3 - <<'PY'
import json; m = json.load(open("qwen36_tiny64_d8/qwen36_meta.json"))
assert m["expert_down_bits"] == 8 and m["expert_down_gs"] == 0 and m["expert_gs"] == 64, m
PY
for cap in 1 2 8; do
COLI_DENSE_I8=0 SNAP=qwen36_tiny64_d8 ./qwen36 "$cap" 4 qwen36_tiny64/ref_full.json > mixed_$cap.log 2>&1 || true
grep -q "expert format on disk: int4 gate/up + int8 down" mixed_$cap.log || { echo "FAIL: mixed layout not detected at cap=$cap"; cat mixed_$cap.log; exit 1; }
grep -E "^C engine" mixed_$cap.log > mixed_ids_$cap.txt
test -s mixed_ids_$cap.txt || { echo "FAIL: no ids at cap=$cap"; cat mixed_$cap.log; exit 1; }
done
cmp mixed_ids_1.txt mixed_ids_8.txt && cmp mixed_ids_2.txt mixed_ids_8.txt || { echo "FAIL: ids differ across caps"; cat mixed_ids_*.txt; exit 1; }
echo "mixed layout: identical ids at cap 1/2/8 ($(cat mixed_ids_8.txt))"
# COLI_CUDA=1 on a mixed container is refused with a line, the CPU path stands
COLI_CUDA=1 COLI_DENSE_I8=0 SNAP=qwen36_tiny64_d8 ./qwen36 8 4 qwen36_tiny64/ref_full.json > mixed_cuda.log 2>&1 || true
grep -q "COLI_CUDA=1 ignored: the VRAM expert tier does not take the mixed layout" mixed_cuda.log || { echo "FAIL: tier refusal line missing"; cat mixed_cuda.log; exit 1; }
- name: A malformed container is refused, not read
run: |
cd c
Expand Down Expand Up @@ -843,7 +868,9 @@ jobs:
# the oracle fixture has no tokenizer; serve mode needs one
python3 tools/make_edge_tiny_tokenizer.py --vocab-size "$(python3 -c 'import json;print(json.load(open("qwen38_tiny/config.json"))["vocab_size"])')" qwen38_tiny
make qwen38 >/dev/null
QWEN38_TINY=qwen38_tiny python3 -m unittest -v tests.test_qwen38_dashboard
QWEN38_TINY=qwen38_tiny python3 -m unittest -v tests.test_qwen38_dashboard tests.test_qwen38_brio
python3 tools/make_edge_tiny_tokenizer.py --vocab-size 64 qwen38_tiny_fp8
QWEN38_TINY=qwen38_tiny_fp8 python3 -m unittest -v tests.test_qwen38_brio

inkling-oracle:
name: Inkling oracle (token-exact vs transformers)
Expand Down Expand Up @@ -923,10 +950,15 @@ jobs:
/tmp/tike
- name: Generate the glm_tiny fixture
run: cd c && python3 tools/make_glm_oracle.py
- name: Token-exact oracle (teacher forcing)
# The fixture is only trustworthy if the engine reproduces it, so assert
# that before reading anything else out of a run.
run: cd c && SNAP=./glm_tiny TF=1 COLI_TEMP=0 ./colibri 64 16 16
- name: Oracle (30–32/32 teacher forcing + exact greedy)
# Preserve CONTRIBUTING.md's two TF near-tie allowance. Greedy remains
# exact, and invalid/non-finite results cannot consume the allowance.
run: |
cd c
SNAP=./glm_tiny TF=1 COLI_TEMP=0 ORACLE_STRICT=1 ORACLE_TF_MAX_MISMATCHES=2 ./colibri 64 16 16
SNAP=./glm_tiny COLI_TEMP=0 ORACLE_STRICT=1 ./colibri 64 16 16
- name: Oracle allowance and failure regressions
run: cd c && python3 tests/test_glm_oracle.py
- name: Structural efficiency tests
# test_cpu_vs_cpu_tok_s_stability is NOT in this list: it is a tok/s
# bound (two runs within 25%) over a ~15 ms tiny replay -- it measured
Expand Down Expand Up @@ -1204,6 +1236,30 @@ jobs:
- name: Python test suite
run: cd c && python3 -m unittest discover -s tests -p 'test_*.py'

gguf:
name: GGUF reader + converter tests
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: '3.12'
cache: pip
cache-dependency-path: c/tools/requirements-gguf.txt
# This job is what makes the numerical GGUF evidence real: without these
# deps those tests skip at import, so the generic `python` job alone only
# exercises test_gguf_reader.py (pure stdlib).
- name: Install GGUF test dependencies
run: pip install -r c/tools/requirements-gguf.txt
- name: GGUF reader, dequant, profile and converter tests
run: |
cd c
python3 -m unittest -v \
tests.test_gguf_reader \
tests.test_gguf_dequant \
tests.test_gguf_olmoe_profile \
tests.test_convert_gguf_to_olmoe

windows-python-focused:
name: Python tests (Windows focused)
runs-on: windows-latest
Expand All @@ -1229,7 +1285,7 @@ jobs:
# This job builds every engine on arm64 (NEON compile coverage), runs the
# integer-kernel bit-exactness gate with the NEON branches live, and replays
# the glm_tiny teacher-forcing oracle against a fixture generated on THIS
# runner (same-machine torch reference, so no cross-ISA float excuses).
# runner. As on x86, allow two TF near ties; greedy remains exact.
oracle-arm:
name: ARM (engines + NEON kernel exactness + tiny oracle)
runs-on: ubuntu-24.04-arm
Expand All @@ -1245,12 +1301,23 @@ jobs:
run: pip install -r c/tools/oracle-requirements.txt
- name: Build every engine (NEON branches must compile)
run: make -C c colibri inkling kimi_k3 olmoe
- name: "DeepSeek V4 FP4 expert kernels: NEON arm bit-exact against the scalar arm (#1696)"
run: |
cd c
# batch (prefill) and S=1 (the matvec the decode uses)
bash tools/bench_fp4_matmul.sh 32 256 128 1
bash tools/bench_fp4_matmul.sh 1 256 128 1
- name: Integer-kernel exactness, NEON branches live (#1081)
run: |
cd c
gcc -O3 -mcpu=native -fopenmp -I. tests/test_int_kernel_exact.c -o /tmp/tike -lm
/tmp/tike
- name: Generate the glm_tiny fixture on this runner
run: cd c && python3 tools/make_glm_oracle.py
- name: Token-exact oracle on ARM (teacher forcing)
run: cd c && SNAP=./glm_tiny TF=1 COLI_TEMP=0 ./colibri 64 16 16
- name: Oracle on ARM (30–32/32 teacher forcing + exact greedy)
run: |
cd c
SNAP=./glm_tiny TF=1 COLI_TEMP=0 ORACLE_STRICT=1 ORACLE_TF_MAX_MISMATCHES=2 ./colibri 64 16 16
SNAP=./glm_tiny COLI_TEMP=0 ORACLE_STRICT=1 ./colibri 64 16 16
- name: Oracle allowance and failure regressions on ARM
run: cd c && python3 tests/test_glm_oracle.py
2 changes: 1 addition & 1 deletion .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -84,7 +84,7 @@ jobs:
- uses: actions/setup-node@v4
with:
node-version: '20'
node-version: '22'
cache: npm
cache-dependency-path: web/package-lock.json

Expand Down
2 changes: 2 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,8 @@ c/deepseek_v4.exe
c/deepseek_v41
c/deepseek_v41.exe
c/COLI_V4_UNIT_*.o
c/deepseek_v4.cflags
c/deepseek_v4.cudaflags
# ...and the ownership-test objects, which the same Makefile puts in a build/
# subdirectory (V4_OWN_DIR) rather than next to the sources. #868 caught the
# twelve in c/, not the four in here, so `make check` still left `?? c/build/`.
Expand Down
Loading
Loading