Skip to content

[Core][Performance] Avoid singleton hybrid KV cache group sizing - #343

Merged
yangzhuxinyzx merged 1 commit into
1CatAI:mainfrom
Leonccaa:perf/hybrid-kv-singleton-group-size
Aug 27, 2026
Merged

yangzhuxinyzx merged 1 commit into
1CatAI:mainfrom
Leonccaa:perf/hybrid-kv-singleton-group-size

Conversation

@Leonccaa

@Leonccaa Leonccaa commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Extract hybrid KV cache group-size selection into a focused helper.
  • Prevent a singleton auxiliary state cache from forcing every repeated cache type into groups of one.
  • Select the largest shared divisor whose total padding stays within 20%.
  • Preserve the existing layout heuristic for all non-singleton configurations.

Problem

While validating Qwen3.8 Flash Next with PLE offload, the hybrid cache layer counts were [1, 12, 12, 12, 36]. The current minimum-count rule selected group_size=1, producing 73 cache groups and repeating slot-mapping and QSA metadata work on every decode step.

The singleton is an auxiliary PLE state cache, not the repeating pattern that should define the main hybrid layout.

Solution

When the minimum layer count is exactly one, compute the greatest common divisor of the repeated cache types and choose the largest divisor that keeps aggregate padding at or below 20%. If no safe divisor exists, retain group_size=1.

This is intentionally narrow: configurations without a singleton keep the current behavior.

Validation

  • Added six selector cases covering the Flash Next layout, common divisors, the padding cap, no-common-divisor fallback, and unchanged existing heuristics.
  • Targeted selector tests: 6 passed.
  • All pre-commit hooks for both changed files pass, including Ruff, mypy, SPDX, forbidden imports, and configuration checks.
  • TP4 on 4x Tesla V100 32 GB, 64-token prompt and 64 measured decode intervals, W4A16 Marlin over ModelOpt packed weights, PLE CPU offload, FULL decode graph:
    • cache groups: 73 -> 7
    • slot-mapping kernels per token: 73 -> 7
    • QSA metadata builds per step: 24 -> 2
    • decode mean: 25.809 -> 34.411 tok/s (+33.33%)
    • deterministic output token IDs remained identical across all measured runs

Observed during the Flash Next work in #338, but this change is generic and does not depend on that model implementation.

CI note

The remote pre-commit workflow runs with --all-files and currently fails on unchanged repository-wide Ruff format, typos, clang-format, markdownlint, SPDX, forbidden-import, and torch.cuda baselines. Neither file changed by this PR appears in the failure diff. Both changed files pass the full local pre-commit suite, and the remote Python 3.10, 3.11, 3.12, and 3.13 mypy hooks all pass.

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
@Leonccaa
Leonccaa force-pushed the perf/hybrid-kv-singleton-group-size branch from 06c4583 to 854a4e0 Compare August 27, 2026 14:38
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor

Maintainer audit on latest public main (06724d0fd4): source review confirms the singleton-only branch preserves all existing non-singleton grouping behavior; selected divisors preserve repeated-layer grouping and enforce the documented aggregate padding cap. The six focused cases plus an AST-isolated exhaustive sweep of 262,144 3-count combinations passed, git diff --check passed, DCO is present, and remote pre-run/pre-commit are green. The recorded matched V100 result is +33.33% decode with identical deterministic token IDs. Accepted for merge.

@yangzhuxinyzx
yangzhuxinyzx marked this pull request as ready for review August 27, 2026 15:35
@yangzhuxinyzx
yangzhuxinyzx merged commit 3ce6e9e into 1CatAI:main Aug 27, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants