-
Notifications
You must be signed in to change notification settings - Fork 187
Pull requests: 1CatAI/1Cat-vLLM
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Bump the minor-update group across 1 directory with 138 updates
dependencies
Pull requests that update a dependency file
#575
opened Sep 8, 2026 by
dependabot
Bot
Loading…
[Bugfix][Spec Decode] Trim the optimistic spec-decode tokens on every pipeline rank
#574
opened Sep 8, 2026 by
Peuqui
Loading…
[Bugfix][Qwen4Exp] Keep the MTP drafter stage-local under pipeline parallelism
#573
opened Sep 8, 2026 by
Peuqui
Loading…
[Bugfix][Perf][SM70] Make Turing (sm75) boot, compute correctly and keep its Inductor fusions
#572
opened Sep 8, 2026 by
Peuqui
Loading…
[Kernel] Generalize H3 FP16 preparation and residual sharding
#571
opened Sep 8, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Kernel] Optimize TP2 attention and packed GDN verification
#566
opened Sep 8, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Bugfix][Core] Preserve singleton prefill semantics in speculative GDN
#563
opened Sep 8, 2026 by
zhaochengggg
Loading…
3 of 4 tasks
[Kernel] Reduce DFlash2 weight and scale memory on SM70
#561
opened Sep 8, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Perf] Gate DFlash2 layout and QPN2 verification optimizations
#556
opened Sep 7, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Kernel][WIP] Screen SM70 FlashNext EP4 and DeepEP
#547
opened Sep 7, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Kernel][WIP] Integrate native FlashInfer SM70 batch GDN and QSA probes
#523
opened Sep 6, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Kernel][SM70] Add native-g32 AWQ QPN M1 operator
#519
opened Sep 6, 2026 by
Leonccaa
Contributor
Loading…
[Kernel][WIP] Port FlashInfer GDN and HC layer fusion to SM70
#515
opened Sep 5, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Kernel][SM70] Port actual FlashInfer CUDA decode to sparse QSA (prototype)
#513
opened Sep 5, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Kernel][WIP] Follow up FlashNext batch HC fusion and QSA route audit
#504
opened Sep 5, 2026 by
yangzhuxinyzx
Contributor
•
Draft
build: restore SM70 (Volta) compilation — 3 csrc fixes (fabric symbols, bf16 helpers, arch-conditional activation kernels)
#435
opened Aug 31, 2026 by
SabaTech-dev
Loading…
platform: disable custom all-reduce on SM70 (Volta) — TP2 was hitting non-converging capture + Xid 62
#431
opened Aug 31, 2026 by
SabaTech-dev
Loading…
[Build][SM70] Make TokenSpeed MLA optional for V100 builds
ready
#409
opened Aug 28, 2026 by
ga-it
Contributor
Loading…
4 tasks done
[Kernel][SM70] Reduce DFlash2 FP8 verification below 20 ms
#405
opened Aug 28, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Bugfix][SM70] Stabilize PP2 TP4 FP8 tactics
#349
opened Aug 27, 2026 by
yangzhuxinyzx
Contributor
•
Draft
[Kernel][SM70] Revisit exact fused V4 auxiliary GEMVs
#348
opened Aug 27, 2026 by
yangzhuxinyzx
Contributor
•
Draft
Previous Next
ProTip!
What’s not been updated in a month: updated:<2026-08-08.