-
Notifications
You must be signed in to change notification settings - Fork 258
All issues
Issue creation is restricted in this repository
Issues
is:issue state:open
is:issue state:open
Search results
[Bug]: [CI] oss:rel:cutlass_dsl_4.8: [Blackwell, shard1] grouped_gemm_wgrad fp4 Tensor-likes mismatch
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.cat-ciCI failures, test flakiness, workflow breakage, or automation issues.CI failures, test flakiness, workflow breakage, or automation issues.orig-qaReported or owned by quality assurance, validation, or testing.Reported or owned by quality assurance, validation, or testing.Status: Open.#727 In NVIDIA/cudnn-frontend;[Bug]: [CI][FROST] frost:cutlass-dsl-4.8:sdpa:sm107 test_sdpa_fp8_bwd_L0 fails on Rubin with fp8 bwd Tensor-likes mismatch
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.cat-ciCI failures, test flakiness, workflow breakage, or automation issues.CI failures, test flakiness, workflow breakage, or automation issues.orig-qaReported or owned by quality assurance, validation, or testing.Reported or owned by quality assurance, validation, or testing.Status: Open.#725 In NVIDIA/cudnn-frontend;frost(sdpa): f16 THD K/V descriptors lack the FP8 packed-total clamp — uninitialized capacity tail poisons O via 0*NaN
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#718 In NVIDIA/cudnn-frontend;- Status: Open.#713 In NVIDIA/cudnn-frontend;
frost(sdpa): pygraph.validate() lowers every SDPA graph to C++ and runs backend validation, even when a FROST engine will serve it
mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#704 In NVIDIA/cudnn-frontend;Validation ordering: layout/shape checks on inferable output tensors must not run in pre_validate_node — audit the support surface
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.mod-frontendcuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.cuDNN frontend APIs, operation graph construction, plans, and user-facing wrappers.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#703 In NVIDIA/cudnn-frontend;frost(sdpa): latent dead-row uninitialized-TMEM NaN hazard in the d192 fp8/mxfp8 and f16 SM100 forward epilogues
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#702 In NVIDIA/cudnn-frontend;[Bug]: DSA sparse_attention_backward (SM90/SM100): NaN gradients on invalid/empty/all-sink top-k rows, GPU deadlock on empty rows, wrong d_sink, and accuracy loss under large logits
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.Status: Open.#676 In NVIDIA/cudnn-frontend;[Feature]: MXFP8 SDPA backward performance and efficient variable-length support
cat-featureRequests for new functionality, APIs, examples, or behavior improvements.Requests for new functionality, APIs, examples, or behavior improvements.Status: Open.frost(sdpa): SM107 fp8 sibling missing the #585 LPT_L2 scheduler port — all causal fp8 declines on Rubin
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.frost(sdpa): THD zero-host-read execute loads uninitialized KV capacity rows — NaN poisons P@V on the f16/fp8 SM100/SM120 rows
cat-bugReports of incorrect behavior, crashes, regressions, or unexpected results.Reports of incorrect behavior, crashes, regressions, or unexpected results.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#624 In NVIDIA/cudnn-frontend;frost(sdpa): evaluate TMA store for the SM120 O epilogue (share the SM100-style setup)
mod-cutedslCuTeDSL kernels, generated kernels, examples, or related integration work.CuTeDSL kernels, generated kernels, examples, or related integration work.orig-nv-engReported or requested by NVIDIA engineering.Reported or requested by NVIDIA engineering.Status: Open.#620 In NVIDIA/cudnn-frontend;