Skip to content

fix(models): reserve host RAM in CPU model selection - #6913

Closed
patil2001 wants to merge 2 commits into
Osmantic:public-betafrom
patil2001:codex/tier0-model-memory-budget
Closed

patil2001 wants to merge 2 commits into
Osmantic:public-betafrom
patil2001:codex/tier0-model-memory-budget

Conversation

@patil2001

@patil2001 patil2001 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Fixes #6912. Both the installer selector and dashboard oracle applied a 3 GB minimum to CPU model memory, even when 35% of detected system RAM was smaller. A 4 GB Tier 0 Linux host therefore received a Qwen 3.5 2B recommendation estimated to need 3 GB, leaving only 1 GB for Linux and ODS services. Remove the floor in both selectors; small hosts now receive no local-model recommendation when no installable model fits the CPU budget. The existing installer fail-closed path tells the operator to choose cloud mode instead of recommending an unsafe model.

Regression coverage

  • The real selector CLI, --backend cpu --ram-gb 4 --tier 0 --installable-only, exits with “no installable model fits”. This failed on the unchanged base, which recommended Qwen 3.5 2B at 64K context.
  • An 8 GB CPU host still selects Qwen 3.5 2B at 64K.
  • The dashboard oracle returns no 3 GB model for a 4 GB CPU host.
  • Added ods/tests/test_model_selector_cpu_memory_budget.py and dashboard selector/parity coverage.

Validation

Ran in capped Linux Docker containers with the source mounted read-only:

  • python /ods/tests/test_model_selector_cpu_memory_budget.py — 2 passed.
  • python /ods/tests/test-pixel-model-selector.py — 15 passed.
  • python /ods/tests/test_model_selector_failure.py — 6 passed.
  • python /ods/tests/test_model_library_apple_context.py — 3 passed.
  • Dashboard oracle 4 GB CPU boundary check against the production ranker passed.
  • git diff --check — passed.

The local selector suite ran in python:3.11-slim with 256 MiB, 1 CPU, and no network. The dashboard oracle check ran in the existing API image with 256 MiB and no network. I did not start a full ODS stack or load a model under a 4 GB cgroup. The selector/install fail-closed behavior is tested; actual model-start OOM behavior is not claimed as reproduced. The full API suite is running in GitHub CI after the dashboard parity correction.

Overlap check

@Lightheartdevs

Copy link
Copy Markdown
Collaborator

Thanks for this contribution. public-beta was promoted into main on 2026-09-24 and no longer receives changes, so we're closing pull requests that target it. This isn't a judgment on the change itself. If it's still needed, please rebase onto main and open a focused PR. See #7253 for details and the contribution policy.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants