Skip to content

[Kernel][SM70] Integrate copy-only GDN projection split from #504 - #549

Merged
yangzhuxinyzx merged 1 commit into
mainfrom
codex/v100-sep07-gdn-copy-audit-20260907-100553
Sep 7, 2026
Merged

[Kernel][SM70] Integrate copy-only GDN projection split from #504#549
yangzhuxinyzx merged 1 commit into
mainfrom
codex/v100-sep07-gdn-copy-audit-20260907-100553

Conversation

@yangzhuxinyzx

Copy link
Copy Markdown
Contributor

AI-assisted integration requested by the maintainer. Narrow extraction from #504; the parent experimental HC/QSA bundle stays open. Base: 8d9c351.

Purpose: fuse only the four existing batched GDN output copies without changing either GEMM or latest-main role-specific plans. Copy flag defaults on; unrelated arithmetic flags unchanged.

Test Plan/Result: 61 targeted CPU/GPU tests passed on owned V100; changing-input CUDA Graph replays and exhaustive 65536 half-payload checks passed. Actual layer weights, synthetic activations, paired complete-projection graph latency improves 11–19% at M2/4/8/16/32/64. No E2E suite or model-throughput claim. Full evidence in docs/design/sm70_gdn_projection_copy_audit.md. Local scoped pre-commit passed. Draft pending hosted checks and final SHA review.

Extract the independently verified copy kernel from #504; preserve main GEMV roles and unrelated defaults.

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant