Skip to content

qwen36: port CACHE_ROUTE with the VRAM tier as the first residency level - #1612

Merged
JustVugg merged 1 commit into
JustVugg:devfrom
crichalchemist:qwen36-cache-route
Sep 21, 2026
Merged

JustVugg merged 1 commit into
JustVugg:devfrom
crichalchemist:qwen36-cache-route

Conversation

@crichalchemist

Copy link
Copy Markdown
Contributor

qwen36: port CACHE_ROUTE with the VRAM tier as the first residency level

Follow-up to the SERE/REAP discussion: the GLM engine's opt-in cache-aware re-routing (arXiv 2412.00099, docs/CACHE_ROUTE.md) now reads the same knobs in the Qwen3.6 engine. Inside the top-M window a slot past the sacred top-J prefers an expert already resident in the VRAM expert tier, then one in the RAM LRU cache, then the plain ranking. Two residency levels because qwen36 has two tiers above disk (GLM has one).

What changes

  • CACHE_ROUTE=1 with ROUTE_J(2) ROUTE_M(12) ROUTE_P(0) ROUTE_ALPHA(1) ROUTE_AGREE(auto): same names and defaults as GLM. Residency is qt_is_resident() (level 2) then slot_indexed() under g_pilot_mx (level 1).
  • route_select() is a pure function of the softmax row, the group mask and a residency callback. tests/test_qwen36_cache_route.c pins it without a model: sacred top-J never displaced, residency past the window ignored, VRAM before RAM, masked experts never chosen, ROUTE_ALPHA scaling only the substitute, ROUTE_P bounding the window by ranked mass, and the swap / agreement / KL meters. Picked up by TEST_RULES, so it runs in make check and the sanitizer job.
  • Footer: CACHE_ROUTE J=.. M=.. P=.. alpha=.. | swap x% (a/b, c to VRAM) and route_agree x% | route_kl y on both the generate and PPL=1 paths.
  • Docs: docs/CACHE_ROUTE.md gets an engines/residency-levels table and the N to VRAM stat, docs/qwen36.md a section, docs/ENVIRONMENT.md regenerated rows.

Semantics contract

  • Unset: the router is byte-identical. Tiny oracle ids match the baseline md5; on the real container the lever-off run reproduces its own reference 128/128 across the scalar and ARCH=native builds.
  • ROUTE_AGREE=1 alone only prints the meters (128/128, agree 100%, KL 0).
  • Only ranks J..K-1 are substitutable. On the top-2 tiny fixture the default ROUTE_J=2 therefore reports swap 0.0%; the docs say so. Real Qwen3.6 routes top-8, so six slots are eligible by default.

Tiny fixture (make qwen36-tiny-check model, CPU, cache 4 of 8 experts/layer)

ROUTE_J RAM hit route_agree
off 44.1% –
1 67.8% 75.9%
0 95.0% 31.9%

A/B on the real container

RAM-level regime (cache 64 of 256 slots/layer, ./qwen36 64 4 ref.json, misses read from the container). Routing meters from the first run (deterministic); tok/s, fetch and compute are the median of 3 timed runs after a discarded warm-up, page cache warm:

row tok/s fetch ms/tok compute ms/tok RAM hit swap agree KL ids match TF-NLL
lever off 4.60 26.3 28.0 76.8% – – – 128/128 0.1288
CACHE_ROUTE=1 (J=2 M=12) 4.90 12.7 28.0 85.3% 10.6% 89.4% 1.85 54/128 0.1606
ROUTE_J=4 5.39 14.6 25.7 83.9% 9.0% 91.0% 1.40 51/128 0.1360
ROUTE_M=24 5.18 6.4 28.5 89.2% 16.2% 83.8% 2.92 54/128 0.1635
ROUTE_J=1 4.89 12.2 28.3 85.7% 11.6% 88.4% 2.11 2/128 0.1767

Reading: the lever halves the fetch time per token (26 to 6–15 ms) for +8 to +12 points of hit rate, and the quality cost tracks agreement/KL: ROUTE_J=4 keeps 91% of the true top-8 picks for +0.007 nats/token, the default and ROUTE_M=24 cost about +0.03, ROUTE_J=1 +0.05. Throughput moves only +6 to +17% here because a warm page cache serves a miss in ~2 ms; the same hit-rate gain is worth far more when misses really come from disk (a cold-cache lever-off run measured 369 ms/token of fetch), but I could not produce a controlled cold row without sudo purge between runs.

VRAM tier regime (COLI_VULKAN=1 VK_EXPERT_GB=auto QWEN_EXPERT_KERNEL=0, all 10240 experts in RAM, 3302 in VRAM, single runs). Without a HEAT_FILE the warmstart fills VRAM in index order, so layers 0–12 are fully resident and the rest empty: inside any layer every candidate has the same residency level and the lever has nothing to prefer. Measured: swap 0.2–0.4% (all "to VRAM", at the one partially resident layer), agreement 99.6–99.8%, VRAM hit rate 32.1% → 32.3–32.5%, TF-NLL 0.1288 → 0.1306–0.1334. That is the expected null result for a layer-contiguous plan, and it is why the tier rows below use a heat profile.

VRAM tier regime with a heat profile (same tier settings plus HEAT_FILE=heat.bin, written by one lever-off run on this prompt and then made read-only so every row plans from the same bytes; that is the best case for a profile). Medians of 3 timed runs; cache 256 so the RAM level holds everything and the only miss is VRAM → CPU:

row tok/s moe ms/tok VRAM hit CPU-computed experts swap (all to VRAM) agree KL ids match TF-NLL
lever off 4.16 81.8 91.3% 4711 – – – 128/128 0.1288
CACHE_ROUTE=1 (J=2 M=12) 4.82 67.1 97.6% 1325 6.9% 93.1% 1.15 49/128 0.1397
ROUTE_J=4 4.36 72.9 96.8% 1746 5.7% 94.3% 0.87 56/128 0.1351
ROUTE_M=24 5.09 63.7 98.8% 650 8.2% 91.8% 1.37 61/128 0.1345
ROUTE_J=1 4.40 72.9 97.9% 1164 7.2% 92.8% 1.25 61/128 0.1405

Reading: every swap goes to VRAM (the footer's N to VRAM equals the swap count), the CPU fallback shrinks 2.7–7x, and the decode is +5 to +22% tok/s on this GPU for +0.006 to +0.012 nats/token. The profile alone is worth more than the lever (32% → 91% VRAM hit rate, 3.0 → 4.2 tok/s); the lever is what takes the last 9% of misses off the CPU. For comparison the CPU-only lever-off run with everything in RAM is 4.80 tok/s, so the tier on this 8 GB MoltenVK card only breaks even with the CPU path once the lever is on; that is a statement about the Radeon Pro 580 through MoltenVK, not about the lever.

Protocol: commit ec99181 (rebased on the #1338 tier head bcdee83 as 55c9afe for the runs), make qwen36 VK=1 ARCH=native with MacPorts libomp, 2017 iMac i7-7700K (8 threads, 64 GB), Radeon Pro 580 8 GB via MoltenVK, container Kreuzzelg/qwen36-35b-a3b-colibri-i4-gs64 on an SD card (~25–60 MB/s cold; after the first runs the misses are served from the page cache, which st.h's F_NOCACHE descriptor does not bypass for pages already cached). 43-token prompt, 128 greedy tokens, reference generated by the lever-off run, COLI_NO_OMP_TUNE=1 COLI_TIMERS=1, routing meters from one run (they are deterministic), timing as the median of 3 runs after a discarded warm-up, PPL=1 teacher-forcing the reference for TF-NLL. Commands:

./qwen36 64 4 ref_cpu.json                          # lever off
CACHE_ROUTE=1 ./qwen36 64 4 ref_cpu.json            # default lever
CACHE_ROUTE=1 ROUTE_J=4 ./qwen36 64 4 ref_cpu.json
CACHE_ROUTE=1 ROUTE_M=24 ./qwen36 64 4 ref_cpu.json
CACHE_ROUTE=1 ROUTE_J=1 ./qwen36 64 4 ref_cpu.json
PPL=1 <same env> ./qwen36 64 4 ref_cpu.json         # TF-NLL of the reference

Hit rate, swap, agreement, KL and TF-NLL are hardware-independent; tok/s and the [timers] fetch/compute split are this host's. Scripts and logs: ab.sh, rows_cpu.sh, rows_vk_heat.sh, rerun_gen.sh, table.py (available on request, not committed).

Found on the way, not part of this PR

  • #1338 at bcdee83 segfaults on the first token with COLI_VULKAN=1 VK_EXPERT_GB=auto on this container (pristine branch reproduces): xf_mode() disables the planar-int4 kernel only under COLI_CUDA=1, so under Vulkan every slot is allocated pw-only, the warmstart offers NULL weights, nothing uploads, and the tier's int8 CPU fallback dereferences e->g == NULL (qwen36.c:2266 → matmul_q_gs at :1062). With QWEN_EXPERT_KERNEL=0 the same run uploads 3302/10240 experts, reproduces the CPU ids 128/128 and decodes at 2.42 tok/s with a 32.1% VRAM hit rate, which is the regime the tier rows below use. The fix (the same env gate for COLI_VULKAN=1) is pushed to qwen36: Vulkan expert tier, and staged device-local uploads for cards without Resizable BAR #1338 as 48ac375, verified on the same run (3302 uploads, 128/128, exit 0); dev carries the CUDA-only gate today, so nothing changes there until that PR merges.
  • The Darwin x86 default build carries no -march: ARCH=native takes expert compute from 880 to 28 ms/token here.

🤖 Generated with Claude Code

https://claude.ai/code/session_01NJcbVYkasdfXg5DEsHdSCt

The GLM engine's opt-in cache-aware re-routing (arXiv 2412.00099,
docs/CACHE_ROUTE.md) now reads the same knobs in the Qwen3.6 engine:
inside the top-M window a slot past the sacred top-J prefers an expert
already resident in the VRAM expert tier, then one in the RAM LRU cache,
then the plain ranking. The router is untouched when CACHE_ROUTE is unset
(byte-identical token ids on the tiny oracle) and ROUTE_AGREE alone only
prints the meters.

route_select() is a pure function of the softmax row, the group mask and a
residency callback, so tests/test_qwen36_cache_route.c pins the contract
without a model: sacred top-J never displaced, residency past the window
ignored, VRAM before RAM, masked experts never chosen, ROUTE_ALPHA scaling
only the substitute, ROUTE_P bounding the window by mass, and the swap /
agreement / KL meters. The footer splits the swap count into "N to VRAM".

Only ranks J..K-1 are substitutable; on the top-2 tiny fixture the default
ROUTE_J=2 therefore reports swap 0.0%, and lowering it moves the RAM hit
rate from 44% to 68% (J=1) and 95% (J=0) with the agreement meter showing
what that cost. The docs say so, and the qwen36 doc gets a section.
Copilot AI lite review requested due to automatic review settings September 18, 2026 19:05

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

Edo771977 added a commit to Edo771977/colibri that referenced this pull request Sep 19, 2026
Cherry-picked from JustVugg#1612 by Courtney Andrew Richardson, with
the two conflicts against this fork resolved and the no-op claim re-verified
here rather than taken on trust.

Why this one, out of the nine upstream PRs opened in the last day: it is the
only one aimed at what this fork's own instrumentation says is costing it the
most. The 19 September run reported

    [timers]   qtier: issue 1.95 | cpu-miss 3.69 | take 1.37 ms/token

and cpu-miss -- routed experts absent from the VRAM tier, computed on CPU --
is 11% of a 32.0 ms token, against a 4.98 ms deficit to llama.cpp's curve at
equal bytes. It also moves between a cold and a warm run of the same binary
(3.69 -> 3.21), which is what a cache policy looks like and not a floor.

What it does: inside the top-M window, a slot past the sacred top-J prefers an
expert already resident in the VRAM tier, then one in the RAM LRU cache, then
the plain ranking. Two residency levels, VRAM first -- the GLM engine has one.
route_select() is a pure function of the softmax row, the group mask and a
residency callback, so tests/test_qwen36_cache_route.c pins the contract with
no model at all.

It is LOSSY and says so: it changes which experts run. The footer prints what
that cost, route_agree (overlap with the true top-K) and route_kl (mass KL),
beside the swap and hit rates. This is the part worth having -- a speed/quality
dial with its own meter attached, rather than a speedup whose price shows up
later in the answers.

Conflicts, both additive, nothing dropped:

  c/qwen36.c     main's g_cap_explicit and the upstream CACHE_ROUTE globals
                 landed on the same line. Both kept.
  docs/qwen36.md upstream adds its section where this fork has grown 300 lines
                 of sync/stream analysis. Ours kept whole, theirs appended at
                 the same anchor, before "Which container?".

One edit beyond the resolution: upstream's section opens "The residual misses
above are the lever's target", pointing at its own preceding text. Here the
thing above has a name and a number, so the sentence now names them --
cpu-miss, 3.7 ms/token, 11% of a token. Same claim, checkable.

Verified on this tree, because this fork's qwen36.c is not upstream's:

  CACHE_ROUTE unset, patched vs origin/main, tiny fixture, caps 1 and 8:
      ids identical, dumped logits BYTE-identical, oracle 9/16 on both
      (the absolute figure is -march-dependent on this host, as ci.yml's own
      comment says; what matters is that the two sides agree exactly)

  CACHE_ROUTE=1 engages, and the meters move as documented:
      ROUTE_J=1   swap 13.4% (43/320)   route_agree 86.6%   route_kl  3.4623
      ROUTE_J=0   swap 65.9% (211/320)  route_agree 34.1%   route_kl 17.7093

  "0 to VRAM" in both: the tiny fixture has no VRAM tier, so every substitution
  came from the RAM level. The VRAM level needs a real CUDA run to exercise,
  which is a measurement for the 7950X box and not for CI.

tests/test_qwen36_cache_route, _dense_batch, _ctx, _tier_int8_decode and
_cache_index all pass; qwen36 builds warning-free.

Co-Authored-By: Courtney Andrew Richardson <courtneyrichardson2027@u.northwestern.edu>
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@JustVugg
JustVugg merged commit dc3b879 into JustVugg:dev Sep 21, 2026
28 checks passed
@crichalchemist
crichalchemist deleted the qwen36-cache-route branch September 23, 2026 19:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants