fix(planner): export K3_EXPERT_GB and GLM53_EXPERT_GB from the plan - #1713
Merged
Merged
Conversation
Kimi K3 sizes the expert LRU from K3_EXPERT_GB (default 8 GB). RAM_GB is only a ceiling, and main() never takes argv as a cap, so --auto-tier that set RAM_GB and COLI_PLAN_CAP left the 8 GB default in place. GLM-5.3 reads GLM53_EXPERT_GB and never RAM_GB. Export each family's cache knob from expert_cache_bytes so coli plan, coli doctor, --auto-tier and coli tune apply the budget the engine actually reads.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Symptom
coli plan,coli doctorand--auto-tiersize a RAM expert cache, then exportRAM_GBandCOLI_PLAN_CAP. Two engines do not read those as the cache size:K3_EXPERT_GB(default 8 GB).RAM_GBis only a ceiling.main()never takes argv as a cap, so the gateway'sCOLI_PLAN_CAPis ignored.GLM53_EXPERT_GBand never readsRAM_GB.Unfixed
dev, a Kimi plan with tens or hundreds of GB of warm experts still started with the 8 GB default.--auto-tierlooked applied. The cache was not.Root cause
environment_for_planalways setRAM_GBfrom the whole-process budget. That is the right variable for colibri, OLMoE, V4. It is the wrong variable for Kimi and GLM-5.3. #1585 stopped advising colibri.c-only knobs on those engines; it did not start advising the knobs they actually read.Fix
After the colibri-only auto-tune block, export the family cache knob from
expert_cache_bytes:K3_EXPERT_GBGLM53_EXPERT_GBcoli plan/coli doctorprint it.--auto-tierapplies it (setdefault, so an explicit value still wins).coli tunethen starts its Kimi resource sweep from the planned cache instead of 8 GB.Tests
In
c/tests/test_resource_plan.py:test_kimi_plan_exports_the_expert_cache_knob_the_engine_reads: a two-layer Kimi fixture getsK3_EXPERT_GBintune, informat_plan, and inenvironment_for_plan; an explicitK3_EXPERT_GB=3.5is kept.test_glm53_plan_exports_the_expert_cache_knob_the_engine_reads: the same forGLM53_EXPERT_GB, and no Kimi key leaks onto GLM-5.3.Fail-before on unfixed
dev:The existing OLMoE test still requires
tune == {}(no colibri knobs, and no *EXPERT_GB either: OLMoE readsRAM_GB).Validation
dev, pass afterpython -m unittest tests.test_resource_plan tests.test_doctor tests.test_autotune tests.test_env_defaults tests.test_doctor_engine_arch tests.test_v4_cli-- 200 OKmake -C c check-- Python-only change; the modules above are the planner/doctor/tune surfacemake -C c cuda-test-- not applicableCompatibility
AI-assisted (Grok)