Skip to content

fix(planner): export K3_EXPERT_GB and GLM53_EXPERT_GB from the plan - #1713

Merged
JustVugg merged 1 commit into
JustVugg:devfrom
kevin9327:fix/planner-kimi-expert-gb
Sep 23, 2026
Merged

JustVugg merged 1 commit into
JustVugg:devfrom
kevin9327:fix/planner-kimi-expert-gb

Conversation

@kevin9327

Copy link
Copy Markdown
Contributor

Symptom

coli plan, coli doctor and --auto-tier size a RAM expert cache, then export RAM_GB and COLI_PLAN_CAP. Two engines do not read those as the cache size:

  • Kimi K3 sizes the expert LRU from K3_EXPERT_GB (default 8 GB). RAM_GB is only a ceiling. main() never takes argv as a cap, so the gateway's COLI_PLAN_CAP is ignored.
  • GLM-5.3 sizes the cache from GLM53_EXPERT_GB and never reads RAM_GB.

Unfixed dev, a Kimi plan with tens or hundreds of GB of warm experts still started with the 8 GB default. --auto-tier looked applied. The cache was not.

Root cause

environment_for_plan always set RAM_GB from the whole-process budget. That is the right variable for colibri, OLMoE, V4. It is the wrong variable for Kimi and GLM-5.3. #1585 stopped advising colibri.c-only knobs on those engines; it did not start advising the knobs they actually read.

Fix

After the colibri-only auto-tune block, export the family cache knob from expert_cache_bytes:

  • kimi -> K3_EXPERT_GB
  • glm53 -> GLM53_EXPERT_GB

coli plan / coli doctor print it. --auto-tier applies it (setdefault, so an explicit value still wins). coli tune then starts its Kimi resource sweep from the planned cache instead of 8 GB.

Tests

In c/tests/test_resource_plan.py:

  • test_kimi_plan_exports_the_expert_cache_knob_the_engine_reads: a two-layer Kimi fixture gets K3_EXPERT_GB in tune, in format_plan, and in environment_for_plan; an explicit K3_EXPERT_GB=3.5 is kept.
  • test_glm53_plan_exports_the_expert_cache_knob_the_engine_reads: the same for GLM53_EXPERT_GB, and no Kimi key leaks onto GLM-5.3.

Fail-before on unfixed dev:

AssertionError: 'K3_EXPERT_GB' not found in {}
AssertionError: 'GLM53_EXPERT_GB' not found in {}

The existing OLMoE test still requires tune == {} (no colibri knobs, and no *EXPERT_GB either: OLMoE reads RAM_GB).

Validation

  • Targeted: the two new tests fail on unfixed dev, pass after
  • Baseline: python -m unittest tests.test_resource_plan tests.test_doctor tests.test_autotune tests.test_env_defaults tests.test_doctor_engine_arch tests.test_v4_cli -- 200 OK
  • make -C c check -- Python-only change; the modules above are the planner/doctor/tune surface
  • CUDA changes were tested with make -C c cuda-test -- not applicable
  • Performance claims include hardware, commands, and repeatable measurements -- none

Compatibility

  • The default CPU build remains dependency-free
  • No model files, generated binaries, or benchmark artifacts are included

AI-assisted (Grok)

Kimi K3 sizes the expert LRU from K3_EXPERT_GB (default 8 GB). RAM_GB
is only a ceiling, and main() never takes argv as a cap, so --auto-tier
that set RAM_GB and COLI_PLAN_CAP left the 8 GB default in place.
GLM-5.3 reads GLM53_EXPERT_GB and never RAM_GB.

Export each family's cache knob from expert_cache_bytes so coli plan,
coli doctor, --auto-tier and coli tune apply the budget the engine
actually reads.
@JustVugg
JustVugg merged commit 72e49d2 into JustVugg:dev Sep 23, 2026
29 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants