Skip to content

deepseek41 : add compressed attention runtime - #3

Merged
ajaxdude merged 6 commits into
jeromecoste-microsoft-deepseek-v41-supportfrom
jeromecoste-microsoft-deepseek-v41-attention
Sep 14, 2026
Merged

ajaxdude merged 6 commits into
jeromecoste-microsoft-deepseek-v41-supportfrom
jeromecoste-microsoft-deepseek-v41-attention

Conversation

@ajaxdude

@ajaxdude ajaxdude commented Sep 12, 2026 •

Copy link
Copy Markdown
Owner

Summary

Stacked on the schema layer in halo-box#49 at 65fdc19831d30b12f06b981c3ca74dd18731f90e.

  • load and strictly validate the published deepseek41.* metadata contract, including the 40-entry emitted compression layout and Engram header metadata
  • add distinct DeepSeek V4.1 tensor ownership for KV sources [2,8,14,20] and index sources [2,8,14,20,24,28,32,36]
  • add model-free V4.1 cache/state planning for ratio-0 raw attention, ratio-1 direct compression, ratio-2 two-token pooling, the 128-token raw ring, and source-owned compressed/index state
  • add candidate block selection and propagation for layer 20 with block size 8, top 2048 blocks, causal filtering, stable ties, and forced retention of the final visible block
  • add GGML graph primitives for ratio pooling, one shared raw-plus-compressed softmax, and final four-stream HC collapse through the BF16/RMSNorm/output path without output_hc_*
  • add parameterized raw/compressed/index/carry/candidate/position/workspace memory accounting
  • register a distinct DeepSeek V4.1 model and tensor contract using ATTN_KV_A_NORM, omitting V4-only output HC, compressor APE, indexer-compressor, TID2EID, and nextn families

Design: halo-box#48 and its refinement comment.

Runtime gate

The published model still requires disk-backed Engram decoding and routed-expert streaming. This layer validates tensor metadata and the four published Engram tensor names (engram_embd, engram_q_norm, engram_k_norm, engram_kv), then fails before payload mapping with an explicit dependency error. It does not claim that the published 365 GB GGUF can execute yet.

Validation

  • test-deepseek41-runtime
  • test-deepseek41-schema
  • test-llama-archs
  • git diff --check

No model payload was loaded or run. No performance claim is made.

Stack note

This PR targets the fork-only base branch jeromecoste-microsoft-deepseek-v41-support. It must be recreated or retargeted to halo-box/strix-llama.cpp after halo-box#49 merges.

Written by GPT-5.6 Sol. The implementation was validated only with local synthetic/model-free tests on macOS; no Strix Halo hardware or published-model inference was used.

Assisted-by: GPT-5.6 Sol

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Jerome Coste and others added 2 commits September 12, 2026 17:26
Assisted-by: GPT-5.6 Sol

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Assisted-by: GPT-5.6 Sol

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@ajaxdude

Copy link
Copy Markdown
Owner Author

Stack repair update: merged the current schema base 6a473b7e942a2cc54f1e5bb2b2d5deacd247b888 normally. New attention head: 580986ee4a86ad2ec5a73d367c8f220e06377da8. Focused runtime and schema tests pass; test-llama-archs -a deepseek41 passes, and diff checks pass. The unfiltered local test-llama-archs run traps in unrelated Metal coverage after many non-DeepSeek architectures; the DeepSeek41 case remains intentionally skipped behind the dependency gate.

Assisted-by: GPT-5.6 Sol

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@ajaxdude
ajaxdude marked this pull request as ready for review September 14, 2026 07:27
Assisted-by: GPT-5.6 Sol

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@ajaxdude
ajaxdude merged commit 0fa6b96 into jeromecoste-microsoft-deepseek-v41-support Sep 14, 2026
3 checks passed
@ajaxdude
ajaxdude deleted the jeromecoste-microsoft-deepseek-v41-attention branch September 30, 2026 00:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant