Skip to content

fix(qwen3.8): wire-format detect nvfp4 artifact profile - #107

Open
koloved wants to merge 1 commit into
Neroued:masterfrom
koloved:master
Open

fix(qwen3.8): wire-format detect nvfp4 artifact profile#107
koloved wants to merge 1 commit into
Neroued:masterfrom
koloved:master

Conversation

@koloved

@koloved koloved commented Aug 28, 2026

Copy link
Copy Markdown

Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile (row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer layout, 18.3 GB artifact). Resolve the weights profile from the artifact's text/token_embedding descriptor instead of always selecting the FP8 profile.

The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer (https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs significantly faster on an RTX 5090 while leaving more VRAM, so it supports a larger context (262144 tokens at full speed) than the 21 GB upstream FP8 artifact, which cannot reach those speeds on the same hardware.

Existing artifacts are unaffected: the upstream FP8 model still resolves to the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8; only W8-embedded community artifacts are routed to Qwen36Nvfp4.

Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile
(row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the
community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer
layout, 18.3 GB artifact). Resolve the weights profile from the artifact's
text/token_embedding descriptor instead of always selecting the FP8 profile.

The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer
(https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs
significantly faster on an RTX 5090 while leaving more VRAM, so it supports
a larger context (262144 tokens at full speed) than the 21 GB upstream
FP8 artifact, which cannot reach those speeds on the same hardware.

Existing artifacts are unaffected: the upstream FP8 model still resolves to
the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8;
only W8-embedded community artifacts are routed to Qwen36Nvfp4.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b03557e82a

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/targets/qwen3_6_27b/impl/package.cpp
@stormgraser-ux

Copy link
Copy Markdown

Confirming this from the other side — I've been carrying a local patch doing the same thing, and #107's predicate matches it on both artifacts I have here. Read straight from the container headers on an RTX 5090:

artifact bytes text/token_embedding resolves to
Ostfralla/Qwen3.8-27B-NVFP4-NInfer 18,324,067,840 W8G32_F16S (row-split-k128-v1) Qwen36Nvfp4
huihui-ai/Qwen3.8-27B-abliterated-NVFP4 21,492,695,040 FP8_E4M3FN_ROW_BF16S (row-scale-v1) Qwen38Nvfp4

The second is worth noting: it's a third-party abliterated build that happens to use the official object layout, declares the same qwen3.8-27b/nvfp4 identity, and still resolves correctly — so the descriptor is discriminating layout rather than provenance.

Without this the Ostfralla artifact dies inside materialisation with a byte-count mismatch naming neither the artifact nor the profile. It reads as a corrupt download, which cost me a re-download before I worked out what it was.

One suggestion: the current code falls through to Qwen38Nvfp4 when text/token_embedding is absent, isn't a tensor, or carries a third format — which reinstates that same unnamed failure for any future artifact. Throwing there, naming both expected formats, makes it a one-line diagnosis.

Running the Ostfralla profile at 262,144 context with vision on the 5090 as a daily driver; happy to test a revision.

@knoopx

knoopx commented Aug 31, 2026

Copy link
Copy Markdown

i've been also patching my own fork/running Ostfralla/Qwen3.8-27B-NVFP4-NInfer for some time, if i recall properly prefill speeds are +10% with this layout over official one and i also verified the conversion carries over and improves ThinkingCap-Qwen3.6-27B, there was an open issue regarding this but it was closed #38

Gevil added a commit to Gevil/ninfer that referenced this pull request Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants