-
Notifications
You must be signed in to change notification settings - Fork 65
Pull requests: gittensor-ai-lab/sparkinfer
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
perf(runtime): requant MoE expert down Q5_K/Q6_K → Q4_K at load
#699
opened Aug 7, 2026 by
kurosawareiji7007-hub
Loading…
1 task done
perf(dflash): widen the verify window past the compact graph's four-row limit
eval-dflash:REJECT
sparkinfer DFlash vs-main speed tier: REJECT
#693
opened Aug 7, 2026 by
coderbench
Loading…
1 task done
perf(dflash): optimize Qwen3.6 target shared-expert decode
dflash-needs-rebase
verified DFlash speedup but not this round's dflash-merge-first
eval-dflash:REJECT
sparkinfer DFlash vs-main speed tier: REJECT
#687
opened Aug 7, 2026 by
flashatten
Loading…
1 task done
server: emit GPU ttft/generation/decode_tps in stream usage
area:runtime
subsystem (emission weight 0.26)
hold
Maintainer override: never auto-merge this PR
#570
opened Jul 21, 2026 by
ai-hpc
Contributor
Loading…
1 of 4 tasks
ProTip!
Find all pull requests that aren't related to any open issues with -linked:issue.