-
-
Notifications
You must be signed in to change notification settings - Fork 297
Pull requests: Neroued/ninfer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Remove the empty graph nodes of the MTP draft phase: +0.33% decode on Qwen3.6-35B-A3B
#226
opened Sep 9, 2026 by
MichaelDementii
Contributor
Loading…
Fold the sigmoid gate into the causal-attention reduce epilogue: operator -8.3% where it folds, +0.31% decode
#225
opened Sep 9, 2026 by
MichaelDementii
Contributor
Loading…
Extend rmsnorm_rope to the text profile: one graph node instead of three, +0.60% decode on Qwen3.6-35B-A3B
#222
opened Sep 9, 2026 by
MichaelDementii
Contributor
Loading…
MTP graph profiles carry no topology class, so one executable serves two attention routes and any draft window past five fails at startup
#221
opened Sep 9, 2026 by
MichaelDementii
Contributor
Loading…
perf(ops): choose the predicated W8 GEMM cache policy instead of inheriting it
#201
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(sparse_moe): widen the Q5 routed-down Rows2 window to its measured crossover
#200
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
perf(sparse_moe): keep two Q4 group quads in flight in the routed gate/up dot product
#199
opened Sep 7, 2026 by
MichaelDementii
Contributor
Loading…
Fall back to a preset of the same weights format when no (model, weights) row matches: prefill-cost prediction goes from 3.1x to 1.15x on a registered artifact
#195
opened Sep 6, 2026 by
MichaelDementii
Contributor
Loading…
perf(nvfp4): drop the guarded expf slow path from the fused SwiGLU epilogue
#194
opened Sep 6, 2026 by
MichaelDementii
Contributor
Loading…
feat(serve): override the frontend chat template via --chat-template FILE
#183
opened Sep 5, 2026 by
wojciak
Loading…
feat(kv): rk2v4-e8 compressed-KV (E8-root, 208 B/head-token) on the paged-KV engine
#173
opened Sep 4, 2026 by
danielfparkernz
Loading…
The fp8 A8 GEMM stages its operands through TMA: prefill +2.4% at chunk 1024 and +4.9% at 4096
#167
opened Sep 3, 2026 by
MichaelDementii
Contributor
Loading…
feat(serve): expose llama.cpp-compatible model metadata on /v1/models
#162
opened Sep 2, 2026 by
hecrj
Loading…
The NVFP4 TMA route reads activation scales as one tile: operator 0.81x to 0.87x, prefill +4.1%, bitwise identical
#160
opened Sep 2, 2026 by
MichaelDementii
Contributor
Loading…
feat(serve): automatic shared-prefix write at the system/developer frontier
#152
opened Sep 1, 2026 by
Astrangemaninhere
Loading…
OpenAI Responses API: compatible with reasoning summary and encryption
#148
opened Sep 1, 2026 by
Sha1rholder
Loading…
fix(qwen3.8): wire-format detect nvfp4 artifact profile
#107
opened Aug 28, 2026 by
koloved
Loading…
build: cache native compilation and incremental image builds
#97
opened Aug 26, 2026 by
DuncanBetts
•
Draft
feat(platform): native Windows (MSVC + CUDA) build for ninfer-serve
#84
opened Aug 22, 2026 by
devan-carlin
Loading…
feat(serve): bound each image with a Vision-token budget
#61
opened Aug 20, 2026 by
Sociopacific
•
Draft
fix(serve): name the exception that terminates the process
#54
opened Aug 19, 2026 by
Sociopacific
Loading…
ProTip!
Mix and match filters to narrow down what you’re looking for.