fix(qwen3.8): wire-format detect nvfp4 artifact profile - #107
Conversation
Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile (row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer layout, 18.3 GB artifact). Resolve the weights profile from the artifact's text/token_embedding descriptor instead of always selecting the FP8 profile. The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer (https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs significantly faster on an RTX 5090 while leaving more VRAM, so it supports a larger context (262144 tokens at full speed) than the 21 GB upstream FP8 artifact, which cannot reach those speeds on the same hardware. Existing artifacts are unaffected: the upstream FP8 model still resolves to the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8; only W8-embedded community artifacts are routed to Qwen36Nvfp4.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b03557e82a
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
|
Confirming this from the other side — I've been carrying a local patch doing the same thing, and #107's predicate matches it on both artifacts I have here. Read straight from the container headers on an RTX 5090:
The second is worth noting: it's a third-party abliterated build that happens to use the official object layout, declares the same Without this the Ostfralla artifact dies inside materialisation with a byte-count mismatch naming neither the artifact nor the profile. It reads as a corrupt download, which cost me a re-download before I worked out what it was. One suggestion: the current code falls through to Running the Ostfralla profile at 262,144 context with vision on the 5090 as a daily driver; happy to test a revision. |
|
i've been also patching my own fork/running |
Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile (row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer layout, 18.3 GB artifact). Resolve the weights profile from the artifact's text/token_embedding descriptor instead of always selecting the FP8 profile.
The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer (https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs significantly faster on an RTX 5090 while leaving more VRAM, so it supports a larger context (262144 tokens at full speed) than the 21 GB upstream FP8 artifact, which cannot reach those speeds on the same hardware.
Existing artifacts are unaffected: the upstream FP8 model still resolves to the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8; only W8-embedded community artifacts are routed to Qwen36Nvfp4.