Skip to content

feat(local-ai): update RTX Spark 48GB recipe (b11320, Q4_K_S, MTP n=2) - #1595

Open
joelagnel wants to merge 6 commits into
openclaw:mainfrom
joelagnel:feat/spark-recipes-sept30
Open

joelagnel wants to merge 6 commits into
openclaw:mainfrom
joelagnel:feat/spark-recipes-sept30

Conversation

@joelagnel

@joelagnel joelagnel commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Related: #1571

What Problem This Solves

NVIDIA published a revised RTX Spark recipe set on 2026-09-30. The 48GB SKU recipe changed quantization and speculative-decoding depth, and the new set requires a newer llama.cpp build than the one currently pinned.

User Impact

RTX Spark 48GB users are offered the recipe NVIDIA validated for that SKU: Qwen3.6-35B-A3B UD-Q4_K_S with MTP n=2 on llama.cpp b11320.

No user action is required. Non-Spark machines are unaffected. Existing installations keep their recorded runtime and model without re-downloading either artifact, and retained UD-IQ4_XS receipts continue to offer Start and use, Use, or Repair during onboarding.

Implementation

  1. Move the managed runtime pin to b11320, the newest build carrying all four Windows CUDA 13.4 artifacts and satisfying every recipe's minimum build.
  2. Move the 48GB SKU recipe from UD-IQ4_XS to UD-Q4_K_S, retaining the previous model as a retired entry for existing receipts.
  3. Align the offered Qwen3.6-35B recipe with the published MTP depth of 2.
  4. Keep fresh selection limited to offered models while evaluating an existing installation receipt through the installed-model catalog.

The previously shipped b11026 runtime remains in the retired catalog with its own file manifest so existing receipts continue to inspect and launch correctly.

Evidence

Artifact sizes and digests were verified by downloading each archive and recomputing SHA-256. The CUDA runtime archives are byte-identical to those already pinned at b11026, so their entries are unchanged.

The b11320 per-file manifests were generated from the release archives: 27 x64 files and 14 arm64 files, matching the file-selection policy already used for b11026.

Change Type

  • Bug fix
  • Feature
  • Refactor
  • Docs or instructions
  • Tests or validation
  • Security hardening
  • Chore or infrastructure

Scope

  • Tray or WinUI UX
  • Windows node capability
  • Local MCP or winnode
  • Gateway, connection, or pairing
  • Setup or onboarding
  • Permissions, privacy, or security
  • Tests, CI, or docs

Required proof pools

  • windows-wsl-dgx-blackwell: the recipe is selected from NVIDIA GPU detection and ends in fixed-prompt inference.
  • windows-clean-installer-upgrade: the runtime pin moves, so installations recorded against b11026 must remain usable.
  • windows-11-arm64: the pin ships x64 and arm64 artifacts and manifests.

Validation

  • Focused current-head tests on 5f2eee47:
    • installed retired Spark model eligibility plus unchanged fresh-selection rejection: 2 passed, 0 failed.
    • retained receipt onboarding states (Start and use, Use, Repair), repair preflight, and pinned-review ownership: 6 passed, 0 failed.
    • WinUI setup control Windows-target compilation reached the Windows-only XAML compiler; the macOS host cannot execute XamlCompiler.exe.
  • Current-head CI passed on 5f2eee47: fast validation, Core and CLI, Tray/setup/integration, UI/functional/accessibility, Proof-pool contracts, Setup/connect E2E, Revocation recovery E2E, Network recovery E2E, and CI Gate.
  • Full Shared and Setup suites are not valid macOS closeout paths because Windows path, process, WSL, and WinUI tests require Windows semantics; focused cross-platform tests passed and Windows CI is authoritative.

Real Behavior Proof

  • Environment: RTX Spark 48GB (arm64) and Dell Precision RTX 5090 (x64).
  • Recipe/catalog commits under hardware proof were rebased unchanged as f7d17c12, 19e0f9a3, and a3eb7425.
  • Current-head deterministic onboarding proof: Observation_RetainedSparkReceiptPreservesOnboardingActions constructs an installed IQ4_XS receipt and verifies the production SetupLocalAiHost.ObserveAsync path returns Start and use, Use, and Repair for stopped, healthy, and failed runtimes. InstalledRetiredSpark48GbModel_RemainsEligibleWithoutRestoringFreshSelection verifies the retired receipt is eligible while the same ID remains UnknownModel for fresh selection. Preflight_RetainedSparkReceiptUsesInstalledModelCatalogDuringRepair carries the exact receipt-proven model through Repair preflight, while the paired no-receipt test keeps the retired ID unavailable for fresh setup.
  • Hardware steps:
    1. Installed a Dev MSIX on the RTX Spark 48GB host, completed Local AI setup, and sent a chat message.
    2. Installed the published alpha at 76ab8399 on the Dell and completed Local AI setup to create a real b11026 receipt.
    3. Installed the updated build over it with Add-AppxPackage -ForceApplicationShutdown, restarted cold, and sent a chat message.
    4. On the RTX Spark 48GB host, completed the b11026 plus IQ4_XS install from a main-based build (f879af8a), then installed this head over it with Add-AppxPackage -ForceApplicationShutdown and without purging Local AI state.
    5. Completed a real WSL Gateway setup on the RTX Spark 48GB host at this head (app-owned OpenClawGateway-Dev, loopback only, Tailscale off), leaving the retained b11026 plus IQ4_XS installation in place, then opened onboarding.
    6. Ran the shipped CudaHostHardwareProbe and LocalInferenceSelector from OpenClaw.Shared directly against the RTX Spark 48GB adapter at this head, so SKU routing is observed rather than inferred from catalog tests.
  • Result:
    • Spark 48GB served qwen3.6-35b-a3b-mtp-ud-q4-k-s, displayed as Qwen3.6 35B-A3B (UD-Q4_K_S).
    • The Dell retained its b11026 runtime receipt, model size, mtime, and installation timestamp. The model was not downloaded again and setup did not rerun.
    • The non-Spark path remained unchanged and continued to recognize the RTX 5090.
    • The Spark 48GB host retained its IQ4_XS receipt across the upgrade to this head (package 2026.9.5.11): state.json SHA-256 unchanged, model file unchanged at 18,209,036,576 bytes with mtime 2026-10-02T17:32:52, installedAtUtc unchanged, and b11026 still the only runtime on disk. The app resolved that retained receipt and launched the router from it: /v1/models reported qwen3.6-35b-a3b-mtp-ud-iq4-xs, and a fixed prompt returned a complete answer (finish_reason: stop, 135 tokens in 2s). The running worker kept the recorded recipe, --spec-type draft-mtp --spec-draft-n-max 2 --ctx-size 98304 on engines\llama-server\b11026\win-arm64\llama-server.exe.
    • Onboarding at this head resolved the retained retired receipt and offered Repair instead of a fresh install, naming the model directly: LocalAiDescription read ... NVIDIA RTX Spark N1X (6144-core Blackwell RTX GPU) - Qwen3.6 35B-A3B (UD-IQ4_XS) with LocalAiActionText = Repair Local AI. Screenshot and exact automation values are in the follow-up comment. The receipt was unchanged throughout the run.
    • Measured selection on the Spark 48GB adapter: IsRtxSpark=True, HasCompleteFacts=True, GpuVisibleMemoryBytes=48719466496, CudaMajorVersion=13. Select() returned qwen3.6-35b-a3b-mtp-ud-q4-k-s on profile ctx-98304-f16 with origin=Default and the plan bound to that adapter's stable ID, and Evaluate() reported CanInstall=True at 28,971,620,640 bytes required against 48,719,466,496 detected. The reading sits below the 48-to-64 geometric midpoint (55.43e9), so the unit classifies as the 48GB SKU and takes the pinned recipe rather than a largest-that-fits profile. Full output is in the follow-up comment.
  • Screenshot links in the follow-up comment were verified and cropped to the application window.
  • Not verified / blocked:
    • Completing the Repair run itself, and receipt recovery after a failed Repair, are not yet captured on hardware.
    • An in-product transition from b11026 to b11320 was not observed end-to-end on hardware. The retained-receipt run above deliberately keeps b11026 so the upgrade-preservation path could be proven, and the b11320 arm64 runtime was exercised separately by the Spark 48GB install that served qwen3.6-35b-a3b-mtp-ud-q4-k-s. Catalog and digest tests cover the pin itself.

Security Impact

  • New permissions or capabilities? No
  • Secrets or tokens handling changed? No
  • New or changed network calls? Yes
  • Command or tool execution surface changed? No
  • Data access scope changed? No
  • Risk and mitigation: downloads move to b11320 release URLs and new Hugging Face model files. Every artifact is pinned by exact size and SHA-256, and every extracted runtime file is pinned by digest.

Compatibility and Migration

  • Backward compatible? Yes
  • Config or environment changes? No
  • Migration needed? No
  • Existing installations continue using their recorded runtime and model. Retired runtime and model entries remain receipt-readable but are not offered for fresh installs.

Review Conversations

  • I replied to or resolved every bot review conversation addressed by this PR.
  • I left unresolved only conversations that still need maintainer judgment.

🤖 Generated with Claude Code

@clawsweeper

clawsweeper Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@joelagnel

Copy link
Copy Markdown
Contributor Author

@clawsweeper please review.

@clawsweeper

clawsweeper Bot commented Oct 2, 2026

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Command router queued. I will update this comment with the next step.

@clawsweeper clawsweeper Bot added P2 Normal priority bug or improvement with limited blast radius. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. proof: sufficient Contributor real behavior proof is sufficient. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. labels Oct 2, 2026
@clawsweeper

clawsweeper Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex review: needs real behavior proof before merge. Reviewed October 4, 2026, 11:45 AM ET / 15:45 UTC (Revision 20).

ClawSweeper review

What this changes

Updates Windows Local AI to llama.cpp b11320 and the revised RTX Spark 48GB model recipe while preserving older installations through receipt-aware onboarding and Repair.

Merge readiness

⛔ Blocked before merge - 3 items remain

This PR remains useful: current main still offers the older Spark recipe. No blocking code defect was found, but the previously identified Repair proof gap remains unresolved.

Priority: P2
Reviewed head: 5f2eee477cf7d7049a7d351e273419c7b17f7c0b

Review scores

Measure Result What it means
Overall readiness 🦐 gold shrimp (3/6) The patch is focused and has substantial positive hardware evidence, but the remaining Repair coverage gap prevents merge readiness.
Proof confidence 🦐 gold shrimp (3/6) Needs stronger real behavior proof before merge: Inspected screenshots and hardware output establish fresh Q4_K_S inference, current-head GPU selection, retained b11026 inference, and onboarding observation. SetupWindow's changed receipt handoff through actual Repair completion and failed-Repair recovery remains explicitly unverified on hardware. No stored-data contract changes were introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Needs proof Needs stronger real behavior proof before merge: Inspected screenshots and hardware output establish fresh Q4_K_S inference, current-head GPU selection, retained b11026 inference, and onboarding observation. SetupWindow's changed receipt handoff through actual Repair completion and failed-Repair recovery remains explicitly unverified on hardware. No stored-data contract changes were introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.
Evidence reviewed 9 items Applicable repository policy: Read the full root AGENTS.md and the applicable openclaw-proof-validation skill. The policy requires direct evidence of changed onboarding behavior. Tracked nested AGENTS.md files do not govern the changed paths; no maintainer-notes directory was present.
Pinned change ownership: Reviewed the pinned merge-base-to-head changes across all 17 introduced files. The checkout and original PR head are 5f2eee4. The supplied stale test merge was not used to infer removal of current-main behavior.
Still necessary on current main: Fetched main still pins b11026 and routes the 48GB Spark tier to IQ4_XS. It does not implement this PR's b11320 and Q4_K_S update. GitHub independently reports this PR open and unmerged with the pinned head.
Findings None None.
Security None None.

How this fits together

Windows Local AI uses detected GPU capabilities and installation receipts to select and manage a local inference runtime. Setup and onboarding expose installation, launch, and repair actions, and the resulting endpoint serves the user's agent.

flowchart TD
  A[Detected GPU capabilities] --> C[Model and runtime selection]
  B[Existing installation receipt] --> C
  C --> D[Onboarding actions]
  D --> E[Install or repair pipeline]
  E --> F[Local inference endpoint]
  B --> F
  F --> G[Agent responses]
Loading

Before merge

  • Add real behavior proof - Needs stronger real behavior proof before merge: Inspected screenshots and hardware output establish fresh Q4_K_S inference, current-head GPU selection, retained b11026 inference, and onboarding observation. SetupWindow's changed receipt handoff through actual Repair completion and failed-Repair recovery remains explicitly unverified on hardware. No stored-data contract changes were introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.
  • Resolve merge risk (P1) - Retained IQ4_XS installations now enter a changed Repair handoff, but successful Repair and recovery after a failed Repair have not been demonstrated on hardware; preservation through that mutation path remains unverified.
  • Complete next step (P2) - Provide current-head real evidence of retained-IQ4_XS Repair completion and failed-Repair recovery, checking receipt recovery and preservation of model files. Screenshots or recordings are preferred for visible behavior; terminal output and redacted logs also count. Redact private endpoints and credentials, then update the PR body to trigger re-review; if needed, ask a maintainer to comment @clawsweeper re-review.
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Production and test LOC production +251/-38 (net +213); tests +178/-7 (net +171) Production growth chiefly supports pinned runtime manifests, retained catalog identities, and receipt-aware recovery.

Merge-risk options

Maintainer options:

  1. Complete retained-install Repair proof (recommended)
    Provide current-head hardware evidence of successful IQ4_XS Repair and failed-Repair recovery, including retained receipt and model-file checks.

Technical review

Best possible solution:

Offer the revised Spark recipe while keeping retained installations launchable and safely repairable, with demonstrated recovery that preserves existing model files.

Do we have a high-confidence way to reproduce the issue?

Not applicable as a bug reproduction: this updates an offered hardware recipe. Real hardware output establishes fresh selection and inference plus retained-install launch, but does not establish Repair completion.

Is this the best way to solve the issue?

Yes, separating fresh selection from receipt-aware eligibility is a focused way to preserve existing installations without reoffering retired models; the remaining uncertainty is runtime proof of Repair.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning medium; reviewed against 7127c16d537c.

Labels

Label changes:

No label changes.

Label justifications:

  • P2: This is a bounded Local AI recipe improvement with installation compatibility work and no demonstrated urgent regression.
  • merge-risk: 🚨 compatibility: The changed retained-install Repair path lacks real execution and failure-recovery evidence, leaving existing-install preservation uncertain.
  • rating: 🦐 gold shrimp: Overall readiness is 🦐 gold shrimp; proof is 🦐 gold shrimp and patch quality is 🐚 platinum hermit.
  • status: 📣 needs proof: The PR needs real behavior proof before ClawSweeper can clear the contributor ask. Needs stronger real behavior proof before merge: Inspected screenshots and hardware output establish fresh Q4_K_S inference, current-head GPU selection, retained b11026 inference, and onboarding observation. SetupWindow's changed receipt handoff through actual Repair completion and failed-Repair recovery remains explicitly unverified on hardware. No stored-data contract changes were introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.

Evidence

What I checked:

Likely related people:

  • Joel Fernandes: Raw commit 6576752 adds src/OpenClaw.Shared/Inference/Catalog/RtxSparkInferenceSelector.cs:38 relative to its recorded parents. This identifies author metadata, not feature responsibility or a PR merger. (role: source-line author; confidence: high; commits: 6576752a85c1; files: src/OpenClaw.Shared/Inference/Catalog/RtxSparkInferenceSelector.cs)
  • RomneyDa: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • shanselman: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rank-up moves

Optional improvements that raise the rating; they are not merge blockers.

  • Complete current-head retained-IQ4_XS Repair and demonstrate failed-Repair receipt recovery on hardware while preserving existing model files.

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (19 earlier review cycles; latest 8 shown)
  • reviewed 2026-10-02T23:09:34.591Z sha 5f2eee4 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-03T00:21:38.644Z sha 5f2eee4 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-03T20:54:47.983Z sha 5f2eee4 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-03T20:58:39.584Z sha 5f2eee4 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-04T00:13:28.990Z sha 5f2eee4 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-04T09:42:05.769Z sha 5f2eee4 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-04T13:19:13.437Z sha 5f2eee4 :: needs real behavior proof before merge. :: none
  • reviewed 2026-10-04T15:40:22.060Z sha 5f2eee4 :: needs real behavior proof before merge. :: none

@joelagnel
joelagnel force-pushed the feat/spark-recipes-sept30 branch from fcaf024 to 0525726 Compare October 2, 2026 05:51
@joelagnel

Copy link
Copy Markdown
Contributor Author

Hardware validation

Both runs are a clean flow on a wiped machine — uninstall, purge OpenClawTray-Dev,
reinstall the MSIX, first-run wizard, WSL gateway, Local AI install — against this
branch's tip.

RTX Spark N1X, 48GB SKU (arm64) — the recipe this series changes

Spark 48GB chat round trip

The response line carries the proof: qwen3.6-35b-a3b-mtp-ud-q4-k-s · 15.3K/98.3K, i.e.
the new Q4_K_S quantization running at the 48GB tier's 98,304-token context, with the
composer showing Qwen3.6 35B-A3B (UD-Q4_K_S).

Live worker argv on that host:

...\engines\llama-server\b11320\win-arm64\llama-server.exe --port 53676
  --model ...Qwen3.6-35B-A3B-UD-Q4_K_S.gguf
  --spec-type draft-mtp --spec-draft-n-max 2 --ctx-size 98304

RTX 5090 (x64) — generic dGPU path, unchanged by this series

5090 verified model

Verified model: llamacpp/qwen3.8-27b-mtp-ud-q4-k-m, confirming the dGPU path keeps its
own recipe and is not pulled onto the Spark table:

...\engines\llama-server\b11320\win-x64\llama-server.exe --port 61565
  --model ...Qwen3.8-27B-UD-Q4_K_M.gguf
  --spec-type draft-mtp --spec-draft-n-max 3 --ctx-size 196608

n=3 / 196,608 on the dGPU versus n=2 / 98,304 on the Spark SKU.

@clawsweeper clawsweeper Bot added the proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. label Oct 2, 2026
@joelagnel
joelagnel force-pushed the feat/spark-recipes-sept30 branch from 0525726 to 32e839b Compare October 2, 2026 08:58
@joelagnel

Copy link
Copy Markdown
Contributor Author

Thanks — the published-install finding is correct and I had the premise wrong. Addressed at 32e839b1.

P1 / P1 — retain b11026 identities and the IQ4_XS model. I verified the claim rather than taking it on trust: v2026.9.5-alpha.78 is published (not a draft), 2026-10-01T21:28:05Z, with OpenClaw-x64.msix, OpenClaw-arm64.msix and OpenClaw.msixbundle attached. Its tag resolves to 76ab8399, and that tree carries X64RuntimeId = "b11026-cuda13-x64", Arm64RuntimeId = "b11026-cuda13-arm64" and Qwen35B_IQ4XSModelId = "qwen3.6-35b-a3b-mtp-ud-iq4-xs" as the active pins. So installations recording those identities exist, and dropping the entries makes FindInstalled throw during router construction and pushes reconcile past its non-destructive upgrade branch.

The commit that removed them has been dropped from the series. Both the b11026 runtime variants and the IQ4_XS model are back as installed-only entries with their original release tags, artifact pins, and the 98,304-token Spark profile. The branch is the seven commits it was before that change.

P2 — repository root discovery. Fixed, and the finding is right about the mechanism: in a linked worktree .git is a file, so a Directory.Exists probe walks past the real root. It now prefers OPENCLAW_REPO_ROOT and otherwise anchors on src/OpenClaw.SetupEngine/default-config.json, matching the existing override-aware pattern used elsewhere in this test project.

On the VCLibs minimum. The floor comes from openclaw-windows-packaging#145 rather than from this series, but it is checkable: scanning the import strings of the pinned b11320 binaries on real hardware, llama-server.exe, llama-server-impl.dll and ggml-cuda.dll import vcruntime140.dll and msvcp140.dll only — no vcruntime140_1.dll. That is the VS2015-era CRT, which 14.0.24217.0 carries. I have not tested a machine that has only the baseline framework package installed, so I would not call it proven on-device; happy to drop the floor back to 14.0.33728.0 if you would rather not carry that risk.

On the earlier "no installed base" rationale. That was my error. I reasoned from "unlaunched product" without checking the releases list, and the published alpha contradicts it. The compatibility surface is real and the entries stay.

@joelagnel

Copy link
Copy Markdown
Contributor Author

Dropped the VCLibs commit from the series, which removes one of the merge-risk items rather than arguing it.

It should not have been here. It came from openclaw-windows-packaging#145, which I applied on the understanding that it fixed the llama-server --version timed out failure. It did not, and the evidence is direct: on the x64 host where that failure reproduced, Microsoft.VCLibs.140.00.UWPDesktop 14.0.33728.0 X64 was already installed — the higher floor — so the dependency was satisfied and the floor could not have been the cause. The real cause was a runtime directory holding the launcher stub without llama-server-impl.dll, which 73978b4c fixes.

The commit also carried a rationale I never reproduced ("llama-server.exe then fails to start inside the packaged session" on a high floor). That was inherited from the upstream PR and stated as established; it was not.

Lowering the minimum is a real compatibility widening, but it belongs to the packaging repo and wants its own on-device proof: it converts a visible install-time dependency failure into an invisible load-time one, which is the failure shape this series exists to remove. The manifest is back at 14.0.33728.0 with the docs and export scripts matching, and git diff upstream/main..HEAD over those files is empty.

Series is now six commits: three recipe changes and three install-reliability fixes, all demonstrated on hardware.

@joelagnel
joelagnel force-pushed the feat/spark-recipes-sept30 branch from 32e839b to 9e643be Compare October 2, 2026 11:40
@clawsweeper clawsweeper Bot added status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. and removed status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. labels Oct 2, 2026
@joelagnel
joelagnel force-pushed the feat/spark-recipes-sept30 branch from 9e643be to f0862eb Compare October 2, 2026 12:54
@clawsweeper clawsweeper Bot added status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. and removed status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. proof: sufficient Contributor real behavior proof is sufficient. labels Oct 2, 2026
@joelagnel
joelagnel force-pushed the feat/spark-recipes-sept30 branch from f0862eb to a0a2be7 Compare October 2, 2026 13:28
@joelagnel

joelagnel commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor Author

Published-install compatibility: launch path verified on hardware

Correction (edited). This comment originally claimed published-install
compatibility outright. That overstated the evidence. What is proven below is
the launch path: a published install still starts and serves inference on
its retained runtime. The observation path used by onboarding is a
different code path and is genuinely broken, exactly as ClawSweeper's Revision 6
P1 describes — WindowsLlamaRuntimeInspector.ValidateVersionOutput compares
--version output against the active catalog pin rather than the receipt's
recorded runtime, so a retained b11026 install fails inspection. I had seen that
symptom on the test box and wrongly attributed it to a duplicate app instance.
Fix in progress; the evidence below stands for the launch path only.

This closes the P1 raised in the last review. Rather than synthesise a manifest,
I built the published alpha's exact commit (76ab8399, the tag behind
v2026.9.5-alpha.78) as a Dev MSIX, installed it on the Dell RTX 5090, and let
it run a real Local AI install. That produced a genuine published receipt:

"runtimeId":      "b11026-cuda13-x64",
"engineVersion":  "b11026",
"modelCatalogId": "qwen3.8-27b-mtp-ud-q4-k-m",
"contextLength":  196608

with the model on disk at 16464440224 bytes, mtime 2026-09-24T10:07:05.

I then built this branch as 2026.9.5.8 — deliberately one revision above
the alpha, so this is a genuine upgrade and not a forced downgrade — and
installed it over the top with Add-AppxPackage -ForceApplicationShutdown. No
purge, no Remove-AppxPackage.

Assertion Result
Install data survives state.json intact, installedAtUtc unchanged
Model re-downloaded? No — still 16464440224 bytes, mtime still 2026-09-24T10:07:05
Router launches on retained runtime Yes — ...\engines\llama-server\b11026\win-x64\llama-server.exe
Endpoint serves the recorded model Yes — /v1/models returned qwen3.8-27b-mtp-ud-q4-k-m
Survives a cold restart Yes — killed every process, relaunched, llama-server auto-started on b11026
Forced back into setup? No — launched straight to the workspace

The strongest single piece of evidence is that the branch build served real
inference
on the retained install, not merely started a process:

published install upgrade

Assistant · qwen3.8-27b-mtp-ud-q4-k-m · 0/196.6K — and 196.6K matches the
retained receipt's "contextLength": 196608.

The non-Spark path is unchanged on the same box: "Your NVIDIA GeForce RTX 5090
can run Local AI on this PC."

New regression test

PublishedAlphaReceipt_StillResolvesAfterTheRuntimeBump pins the exact receipt
pair measured above. It is load-bearing: removing the retired-runtime lookup from
LlamaRuntimeCatalog.FindInstalled makes it fail with Assert.NotNull() Failure,
restoring it makes it pass. It is squashed into the runtime-bump commit, so the
series is still six commits.

What I did not verify

Onboarding's observation path for a retained install — this is the defect noted
in the correction above, not a gap in the evidence.

The in-product runtime transition b11026 → b11320 was not observed
end-to-end. Both UI entry points were blocked on this box: the setup wizard
required permanently deleting a WSL distro that was out of scope for this test,
and the openclaw://hub/local-ai deep link spawned a second app instance rather
than routing into the running one. That transition is therefore covered at the
catalog/logic level by the new test only. Rollback-after-failed-upgrade is
likewise unverified.

Test counts at a0a2be74 (Windows)

Suite Failed Passed Skipped Total
OpenClaw.Shared.Tests 0 4231 33 4264
OpenClaw.Connection.Tests 3 1496 1 1500
OpenClaw.SetupEngine.Tests 3 2206 1 2210
OpenClaw.Tray.Tests 0 3862 0 3862

.\build.ps1 -Msix Dev exited 0. The 6 failures are pre-existing and live in
files this series never touches, confirmed with git diff --name-only.

The PR description has been rewritten to the repository template with proof pool
declarations and the full validation detail.

@joelagnel
joelagnel force-pushed the feat/spark-recipes-sept30 branch 2 times, most recently from ba9b5e0 to 00d3e73 Compare October 2, 2026 13:44
joelagnel added a commit to joelagnel/openclaw-windows-node that referenced this pull request Oct 2, 2026
@joelagnel
joelagnel force-pushed the feat/spark-recipes-sept30 branch from 00d3e73 to b57027f Compare October 2, 2026 14:10
@joelagnel

Copy link
Copy Markdown
Contributor Author

Scope reduced to the recipe set, and rebased onto current main

This PR is now three commits and touches only the recipe catalog:

Commit
13294415 build(local-ai): bump the managed runtime pin to b11320
ff90ee78 feat(local-ai): move the RTX Spark 48GB recipe to UD-Q4_K_S
b57027f6 feat(local-ai): align the offered Qwen3.6-35B recipe with MTP n=2

The three Local AI reliability fixes that were previously here have been
dropped. Rebasing onto current main showed that area has been reworked
upstream, and keeping them would have put this PR in the middle of that work:

  • Inspection no longer launches llama-server --version at all, so the probe
    budget those commits added is dead code on top of main.
  • ServerImplementationLibraryName and the per-file digest manifests landed
    upstream, which supersedes the missing-implementation-library check.

Rebase conflicts are resolved. One consequence worth calling out: the new
inspection fails closed when a variant has no file manifest, so moving the pin to
b11320 required generating a manifest for it — 27 x64 and 14 arm64 entries,
produced from the release archives themselves using the same file selection the
repository already applies to b11026. The retired b11026 variants now carry their
own manifest too, so an installation recorded against the previous pin still
inspects cleanly instead of being reported as unverifiable.

The CUDA runtime archives are byte-identical to those already pinned at b11026
(same sizes, same digests), which is an independent check that the generation
procedure matches the existing entries.

On the two findings from the last review

"Validate installed runtimes against their recorded pin" — resolved, though
not by me. ValidateVersionOutput no longer exists on main; inspection now
takes the LlamaRuntimeVariant as a parameter, which is exactly the receipt-aware
shape the finding asked for. What remained was that retired variants carried no
file manifest, so they would fail the new inspection — that is fixed here.

"Keep the installed Spark model eligible for onboarding" — not changed, and I
would like a maintainer's view before touching it. The suggested fix is for
LocalInferenceSelector to fall back to FindInstalled when an explicitly
requested model id is no longer offered. I implemented that, and it breaks
Evaluate_Removed16GiBModelIdIsUnknown, which asserts the opposite: that a
retired id passed as an explicit request reports UnknownModel. That test
predates this PR and encodes the same treatment for the already-retired
qwen3.5-9b-mtp-q4-k-m, which sits in the installed-only catalog exactly as
...-iq4-xs now does.

So the behaviour the finding describes is the repository's existing, deliberate
treatment of retired models rather than something this PR introduces. Changing it
would be a product decision affecting every retired model, which seems out of
scope here. I have reverted my attempt and left the invariant intact.

@joelagnel

Copy link
Copy Markdown
Contributor Author

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🦞👀
Exact review queued.

Re-review progress:

@joelagnel

Copy link
Copy Markdown
Contributor Author

Current-head proof: 48GB Spark recipe on b11320, end to end

Re-validated on RTX Spark 48GB hardware (arm64) against the rebased head
b57027f6, so this is proof of the PR as it now stands rather than of the
earlier, larger branch.

The managed runtime directory was purged before installing, which forces
b11320 to be re-acquired and every pinned runtime file to be verified by digest.
The install completed, so the new per-file manifest this PR adds is correct
against the real release archives.

Receipt written by the install:

engineVersion = b11320
runtimeId     = b11320-cuda13-arm64
model         = qwen3.6-35b-a3b-mtp-ud-q4-k-s

Setup routed the SKU to the new recipe before installing anything:

Compatible PC. Review setup before installing. · NVIDIA RTX Spark N1X
(6144-core Blackwell RTX GPU) · Qwen3.6 35B-A3B (UD-Q4_K_S)

and sized it as 27 GiB required. 45.4 GiB CUDA-visible, matching the memory
envelope pinned in Spark48GbRecipe_RequiredMemoryStaysWithinTheSku.

Prompt and response

Prompt: "In one sentence, explain why the sky appears blue."

Sunlight scatters off nitrogen and oxygen molecules in Earth's atmosphere, and
because blue light has a shorter wavelength it scatters more readily than other
colors — a phenomenon called Rayleigh scattering.

Spark 48GB recipe on b11320

The assistant line names the model that actually served it —
qwen3.6-35b-a3b-mtp-ud-q4-k-s · 15.3K/98.3K (15%) — and the composer reads
Qwen3.6 35B-A3B (UD-Q4_K_S).

CI

All checks green on this head, including Core and CLI tests, Tray/setup/
integration, UI/functional/accessibility, Proof-pool contracts, and the three
E2E jobs.

@RomneyDa RomneyDa added the status: 🚢 actively landing A maintainer or agent is actively driving this item through implementation, validation, or merge. label Oct 2, 2026
@RomneyDa RomneyDa changed the title feat(local-ai): RTX Spark 48GB recipe update (b11320, Q4_K_S, MTP n=2) and install reliability fixes feat(local-ai): update RTX Spark 48GB recipe (b11320, Q4_K_S, MTP n=2) Oct 2, 2026
@RomneyDa

RomneyDa commented Oct 2, 2026

Copy link
Copy Markdown
Member

Addressed the actionable ClawSweeper finding in c8c0d826.

  • SetupLocalAiHost.ObserveAsync now uses receipt-aware eligibility when an install receipt exists.
  • The installed path resolves retired catalog entries; fresh selection still returns UnknownModel for UD-IQ4_XS.
  • Added production-path coverage for a retained Spark receipt in stopped, healthy, and failed runtime states, yielding Start and use, Use, and Repair respectively.
  • Focused tests: 5 passed, 0 failed.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🦞👀
Exact review queued.

Re-review progress:

@RomneyDa

RomneyDa commented Oct 2, 2026

Copy link
Copy Markdown
Member

Addressed the Repair-handoff finding in f8289486.

  • Pinned recovery review evaluates the receipt model through EvaluateInstalled.
  • SetupWindow carries a runtime-only, non-serialized InstalledReceiptModelId into preflight.
  • Preflight permits retired-catalog lookup only when the selected ID exactly matches that receipt-proven ID; fresh setup still rejects it.
  • Added focused coverage for retained Repair preflight, missing-receipt rejection, review ownership, and onboarding projection. Focused result: 6 passed, 0 failed.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🦞👀
Exact review queued.

Re-review progress:

joelagnel and others added 6 commits October 2, 2026 16:03
The RTX Spark recipe set published on 2026-09-30 raises its minimum
llama.cpp build: the Qwen3.6-35B-A3B recipes require b11236, the
Qwen3.8-27B DFlash recipes require b11229, and Qwen3.8 Flash-Next
requires b11256. A single managed runtime is pinned for every recipe,
so the pin moves to b11320, the newest build carrying all four Windows
CUDA 13.4 artifacts, which clears every one of those minimums.

Artifact sizes and digests are taken from the release and verified by
downloading each llama-server archive and recomputing SHA-256. The
CUDA runtime archives are byte-identical to the ones already pinned at
b11026, so their digests are unchanged.

b11026 moves into the retired set alongside b10655. It is a shipped
runtime once this release goes out, so an installation recorded
against it has to keep resolving its own receipt and stay launchable
until setup upgrades it.
…tion

The 2026-09-30 recipe set replaces the 48GB-SKU Qwen3.6-35B-A3B
quantization: UD-IQ4_XS becomes UD-Q4_K_S. Both files already exist at
the Hugging Face revision this catalog pins, so only the artifact
changes; the revision, run recipe, and 98,304-token context tier are
untouched.

The previous quantization is retired rather than rewritten in place.
It ships as the 48GB default, so an existing receipt has to keep
resolving its own pinned artifact and profile until setup upgrades it.
Retired entries normally expose only the pre-profile native/F16
profile, which this model was never installed under, so the retired
entry keeps the 48GB SKU's fixed tier instead. A regression covers
that and fails without it.
The 2026-09-30 recipe set runs every Qwen3.6-35B-A3B configuration with
MTP n=2, including the 28GB and 30GB tiers the 64GB and 128GB SKUs
offer. The already-offered UD-Q4_K_M entry inherited the catalog-wide
default of 3 instead, so it launched with a draft depth no published
recipe uses.

Its pinned artifact already matches the new recipes byte for byte --
same revision, same file, same digest -- so only the draft depth
changes, and the model stays reachable exactly as before through the
existing explicit-alternative path.

The sampling block published alongside these recipes is not adopted:
every recipe in the set carries an identical string, including
"speculative draft backend sampling: ON" on MTP-only entries that have
no draft model, so it reads as shared boilerplate rather than per-model
tuning. The 35B entries keep the temperature they ship with today.
@RomneyDa
RomneyDa force-pushed the feat/spark-recipes-sept30 branch from f828948 to 5f2eee4 Compare October 2, 2026 23:05
@RomneyDa

RomneyDa commented Oct 2, 2026

Copy link
Copy Markdown
Member

Rebased onto current main (f879af8a) and preserved the complete contributor commit chain. Current head: 5f2eee47.

Cleanup review narrowed the installed-receipt selector and transient recovery marker from public to internal visibility; no behavior or serialized contract changed. Focused retained-receipt tests remain green.

The prior Setup/connect E2E cancellation was isolated to the stale pre-rebase run; current main CI is green. Fresh exact-head CI is running.

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

🦞👀
Exact review queued.

Re-review progress:

@RomneyDa RomneyDa removed the status: 🚢 actively landing A maintainer or agent is actively driving this item through implementation, validation, or merge. label Oct 3, 2026
@joelagnel

Copy link
Copy Markdown
Contributor Author

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

🦞👀
Exact review queued.

Re-review progress:

@joelagnel

Copy link
Copy Markdown
Contributor Author

Retained-receipt proof on current head (5f2eee47), RTX Spark 48GB, arm64.

Baseline was a main-based build (f879af8a, pins b11026) with a completed IQ4_XS install. This head was then installed over it with Add-AppxPackage -ForceApplicationShutdown, without purging Local AI state, taking the package from 2026.9.5.10 to 2026.9.5.11.

pre-upgrade post-upgrade after app launch
modelCatalogId qwen3.6-35b-a3b-mtp-ud-iq4-xs same same
runtimeId b11026-cuda13-arm64 same same
state.json SHA-256 BC2BA758...6DC1A68B BC2BA758...6DC1A68B endpoint port only
model file 18,209,036,576 B, mtime 2026-10-02T17:32:52 same same
engines on disk b11026 b11026 b11026

installedAtUtc stayed 2026-10-03T00:33:32, so setup did not rerun and the model was not fetched again. The retired model ID resolved from the receipt rather than failing fresh selection, and the router was launched from it:

GET /v1/models  ->  qwen3.6-35b-a3b-mtp-ud-iq4-xs
POST /v1/chat/completions  ->  finish_reason: stop, 135 tokens in 2s

The worker process kept the recorded recipe, unchanged by the newer pin:

...\LocalAI\engines\llama-server\b11026\win-arm64\llama-server.exe
  --alias qwen3.6-35b-a3b-mtp-ud-iq4-xs
  --spec-type draft-mtp  --spec-draft-n-max 2
  --ctx-size 98304  --cache-type-k f16  --cache-type-v f16  --flash-attn on
  --model ...\Qwen3.6-35B-A3B-UD-IQ4_XS.gguf

The b11320 arm64 runtime is covered separately by the Spark 48GB install that served qwen3.6-35b-a3b-mtp-ud-q4-k-s; this run intentionally keeps b11026 so the upgrade-preservation path is what gets proven.

@clawsweeper clawsweeper Bot removed the proof: 📸 screenshot Contributor real behavior proof includes screenshot evidence. label Oct 4, 2026
@joelagnel

Copy link
Copy Markdown
Contributor Author

Retained IQ4_XS onboarding on current head (5f2eee47), RTX Spark 48GB, arm64, package 2026.9.5.11.

This run completed a real WSL Gateway setup on the host (app-owned OpenClawGateway-Dev, loopback only, Tailscale off), with the existing b11026 + IQ4_XS installation left in place. Onboarding then resolved that retained receipt and offered Repair rather than a fresh install:

Retained IQ4_XS onboarding

Exact UI automation values read from the live window:

LocalAiTitle       = Local AI on this PC
LocalAiDescription = The managed installation needs attention. Review repairs
                     without reinstalling the Gateway.
                     · NVIDIA RTX Spark N1X (6144-core Blackwell RTX GPU)
                     · Qwen3.6 35B-A3B (UD-IQ4_XS)
LocalAiActionText  = Repair Local AI

The retired model is resolved and named by the current-head onboarding surface. Under the previous fresh-selection path that ID returns UnknownModel, so the card could not have named it.

The receipt was untouched throughout: state.json SHA-256 stayed 187E362E...76E63B06 before and after, installedAtUtc stayed 2026-10-03T00:33:32, the GGUF stayed 18,209,036,576 bytes at mtime 2026-10-02T17:32:52, and b11026 remained the only runtime on disk.

@joelagnel

Copy link
Copy Markdown
Contributor Author

Recipe selection measured directly on RTX Spark 48GB hardware at current head (5f2eee47), arm64.

This runs the shipped CudaHostHardwareProbe and LocalInferenceSelector from OpenClaw.Shared against the real adapter, so the SKU routing is observed rather than inferred from catalog tests.

CpuArchitecture          : Arm64
HasNvidiaGpu             : True

--- GPU 'NVIDIA RTX Spark N1X (6144-core Blackwell RTX GPU)' ---
  Vendor                   : Nvidia
  GpuVisibleMemoryBytes    : 48719466496
  FreeGpuVisibleMemoryBytes: 48509751296
  CudaMajorVersion         : 13
  StableId                 : 'GPU-d67ee4df-...'
  IsRtxSpark               : True
  HasCompleteFacts         : True

Select(no request): Status=Selected Failure=None
  model   : qwen3.6-35b-a3b-mtp-ud-q4-k-s  (Qwen3.6 35B-A3B (UD-Q4_K_S))
  profile : ctx-98304-f16
  origin  : Default  boundGpu='GPU-d67ee4df-...'

Evaluate(): CanInstall=True Status=Eligible Failure=None
  plan model   : qwen3.6-35b-a3b-mtp-ud-q4-k-s (Qwen3.6 35B-A3B (UD-Q4_K_S))
  plan profile : ctx-98304-f16
  requiredTotal: 28971620640  detected: 48719466496

Points worth noting against this PR's intent:

  • 48719466496 bytes falls below the 48-to-64 geometric midpoint (55.43e9), so the unit classifies as the 48GB SKU and takes the Q4_K_S recipe.
  • The profile is the SKU table's pinned ctx-98304-f16, not a generic largest-that-fits pick.
  • boundGpuStableId is set, so the plan is bound to the adapter that produced it.
  • origin: Default confirms this is the SKU table's own recommendation rather than an explicit request.

Headroom is 28.97 GB required against 48.72 GB visible.

@joelagnel

Copy link
Copy Markdown
Contributor Author

@clawsweeper re-review

@clawsweeper

clawsweeper Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

🦞👀
Exact review queued.

Re-review progress:

@joelagnel

Copy link
Copy Markdown
Contributor Author

Retained IQ4_XS install observed through the tray's Local AI surface on current head (5f2eee47), RTX Spark 48GB, arm64, package 2026.9.5.11.

This is after the real WSL Gateway setup completed on the host, so the Hub reaches SetupLocalAiHost.ObserveAsync through the normal Recovery route rather than a constructed receipt.

Settings > Local AI with the retained IQ4_XS install

Automation values read from the live page:

LocalAiEngineStatus    = Running
LocalAiEngineOwnership = Managed by OpenClaw
EngineVersionText      = b11026
EndpointText           = http://127.0.0.1:52375/v1
ProcessIdText          = 11528
LocalAiModelStatus     = Downloaded and verified
LocalAiModelName       = Qwen3.6 35B-A3B (UD-IQ4_XS)
ModelRecipeText        = 96K context - F16 target + MTP draft KV - Loads on first request
LocalAiGatewayStatus   = Connected
GatewayDetailText      = Local (OpenClawGateway-Dev)

Downloaded and verified is the per-file SHA-256 inspection passing against the retained b11026 variant's manifest, and the engine is serving on the recorded runtime rather than the newer pin. The recipe line reflects the receipt's own 98304 context and MTP draft KV.

Startup for this session recorded the same, with the session resolving to the retained model:

Local AI router startup state: Healthy
[SESSION] model changed '' -> 'qwen3.6-35b-a3b-mtp-ud-iq4-xs'

Receipt was unchanged across the run: model qwen3.6-35b-a3b-mtp-ud-iq4-xs, runtime b11026-cuda13-arm64, installedAtUtc 2026-10-03T00:33:32, GGUF 18,209,036,576 bytes at mtime 2026-10-02T17:32:52.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants