Repository navigation
Conversation
Collaborator
|
Thanks for this contribution. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes #6912. Both the installer selector and dashboard oracle applied a 3 GB minimum to CPU model memory, even when 35% of detected system RAM was smaller. A 4 GB Tier 0 Linux host therefore received a Qwen 3.5 2B recommendation estimated to need 3 GB, leaving only 1 GB for Linux and ODS services. Remove the floor in both selectors; small hosts now receive no local-model recommendation when no installable model fits the CPU budget. The existing installer fail-closed path tells the operator to choose cloud mode instead of recommending an unsafe model.
Regression coverage
--backend cpu --ram-gb 4 --tier 0 --installable-only, exits with “no installable model fits”. This failed on the unchanged base, which recommended Qwen 3.5 2B at 64K context.ods/tests/test_model_selector_cpu_memory_budget.pyand dashboard selector/parity coverage.Validation
Ran in capped Linux Docker containers with the source mounted read-only:
python /ods/tests/test_model_selector_cpu_memory_budget.py— 2 passed.python /ods/tests/test-pixel-model-selector.py— 15 passed.python /ods/tests/test_model_selector_failure.py— 6 passed.python /ods/tests/test_model_library_apple_context.py— 3 passed.git diff --check— passed.The local selector suite ran in
python:3.11-slimwith 256 MiB, 1 CPU, and no network. The dashboard oracle check ran in the existing API image with 256 MiB and no network. I did not start a full ODS stack or load a model under a 4 GB cgroup. The selector/install fail-closed behavior is tested; actual model-start OOM behavior is not claimed as reproduced. The full API suite is running in GitHub CI after the dashboard parity correction.Overlap check
main; this PR targets pre-install CPU model selection onpublic-beta.main; this PR changes CPU system-RAM selection in both the installer and dashboard, not unified-memory budgeting.