Skip to content

deepseek41 : add core DeepSeek V4.1 runtime (arch, loader, Engram, expert store, memory admission, watchdog) - #70

Open
ajaxdude wants to merge 8 commits into
masterfrom
deepseek-v41-core-runtime
Open

ajaxdude wants to merge 8 commits into
masterfrom
deepseek-v41-core-runtime

Conversation

@ajaxdude

@ajaxdude ajaxdude commented Sep 17, 2026 •

Copy link
Copy Markdown
Member

This replaces #68 (same 7 commits). #68's commits carried a real name + machine hostname as git
author/committer (auto-configured by git because user.name/user.email were never set), and its branch
name embedded a corporate username. That is the same category of issue that got the
original #49/#59/#60 closed ("leaking an incredible amount of personal configuration and pathing") --
just at the commit-metadata/branch-name level instead of in file content. Fixed here: all 7 commits were
rebased with --exec "git commit --amend --no-edit --reset-author" (author/committer -> ajaxdude,
content verified byte-identical via git diff against #68's tip), and the branch was pushed under a
new, leak-free name (deepseek-v41-core-runtime). #68 is closed in favor of this PR; see its closing
comment for the pointer back. The branch has since been rebased onto current master and gained a
typing-only commit 8, so it is no longer byte-identical to #68 (see "Rebase onto current master" and
"Typing-only follow-up (commit 8)" below).

Overview

This is Layer 1 (PR A) of a two-PR stack reviving closed PR #59 ("trace : recover DeepSeek V4.1 correctness harness"), which was closed alongside #49 and #60 for two reasons: (1) all three PRs targeted master directly instead of stacking, so each diff showed the combined content of its unmerged predecessors; (2) leaked personal paths. This PR succeeds #49 with the same substantive content, properly restacked with clean history and no personal paths.

This layer adds the core DeepSeek V4.1 runtime, scoped as the model-execution half of the correctness-harness effort tracked in #48:

  • LLM_ARCH_DEEPSEEK41 registration and GGUF schema (gguf-py/gguf/constants.py, src/llama-arch.{cpp,h})
  • Model graph and loader wiring (src/models/deepseek41.cpp, src/llama-model.cpp, src/llama-model-loader.cpp, src/llama-graph.cpp, src/llama-context.cpp, include/llama.h, common/arg.cpp)
  • Engram state and expert store (src/llama-engram.*, src/llama-expert-store.*, src/llama-dsv41-engram.*, src/llama-dsv41-expert.*, src/llama-dsv41.*)
  • On-disk expert memory admission (src/llama-memory-dsv41.*, src/llama-bounded-file.*, src/llama-ple-disk.cpp integration)
  • A standalone host-memory watchdog script and docs (scripts/strix_memory_watchdog.py, docs/strix-memory-watchdog.md)
  • Unit tests for all of the above
  • A POSIX_FADV_WILLNEED guard in src/llama-ple-disk.cpp (commit 7) that fixes a pre-existing macOS/Apple build error on master (see "Notes" below)

PR B (#71, the correctness trace harness under tools/deepseek-v41-trace/) is stacked on top of this branch. This layer deliberately does not include tools/deepseek-v41-trace/, tools/server/ integration, or the PR60 host-memory-guard-in-loader scope; those are out of scope here.

AI usage disclosure: AGENT-AUTHORED. I (Claude Sonnet 5, running as GitHub Copilot) rebuilt this layer from scratch: reconciled the 12 files that drifted between the old closed branch and current master (see below), re-applied the validated DeepSeek V4.1 content on top of current master, and wrote fresh commits. The later rebase onto current master, its conflict resolution, and the typing-only commit 8 (below) were done by Claude Opus 5.5, also running as GitHub Copilot.

Rebase onto current master

Rebased onto master at 03895887abe6b46473e0a2e07d0a82fe44484894 (previous base: 0636c9aee4a2ab1d38bb7e8c038227c6634c3857, 99 commits behind). Commits 1-3 and 5-7 replayed unchanged (git range-diff reports them as =). All conflicts were in commit 4 (deepseek41 : wire model graph and loader):

  • --ngram-on-disk and lazymode merge (#29) removed the ple_on_disk, ple_direct_io, ple_io_threads and ple_cache_mb params and the --ngram-io-threads/--ngram-cache/--ngram-direct-io options (replaced by --lazy-mode on-direct). This PR had added its expert-cache fields next to those lines. Master's removal is kept; only this PR's additions are re-anchored: --expert-cache-slots/--expert-cache-mib now follow --ngram-on-disk (before --cpu-moe) in common/arg.cpp, the common_params fields and common_model_params_to_llama assignments follow no_host, and expert_cache_bytes/expert_cache_slots in llama_model_params follow main_gpu, in the same order in llama_model_default_params.
  • load_tensors: init_mappings keeps master's lazy-mode-direct skip plus this PR's external-tensor guard, i.e. !(params.lazy_mode == LLAMA_LAZY_MODE_DIRECT || ml.external.any()); params.ple_on_disk is gone.
  • Master added the MAPLE arch. The swiglu-clamp arch check in build_moe_ffn keeps LLM_ARCH_MAPLE and adds LLM_ARCH_DEEPSEEK41; tests/test-llama-archs.cpp keeps master's MAPLE block next to the DeepSeek V4.1 block.

Against the pre-rebase patch, the only changed +/- lines of the whole PR diff are the init_mappings condition and the MAPLE term in the swiglu-clamp condition; the diffstat is unchanged (54 files, +16750/-185). Commit 4's message lost a stale ple_on_disk reference and gained an Assisted-by: Claude Opus 5.5 trailer. All 8 commits, including commit 8 below, are authored and committed as ajaxdude.

Typing-only follow-up (commit 8)

Commit 8 (1568e7903cc07fc83fee1894df491dfcd2f7ba9c, strix : fix ty findings in memory watchdog on Linux) changes only scripts/strix_memory_watchdog.py (+4/-2). ty reported 4 errors in the guardian launch path, but only with --python-platform linux: on macOS that branch is unreachable (use_guardian requires sys.platform.startswith("linux")), so the earlier ty check missed them.

  • ProcessHandle declared pid: int as a writable protocol attribute, but GuardianProcess exposes pid as a read-only property. So ty rejected the assignment to child: ProcessHandle | None, and the two later uses of child fell back to ProcessHandle | None. The protocol now declares pid as a read-only property. subprocess.Popen and the test fakes still satisfy it.
  • _launch_guardian annotated launch_mask as set[signal.Signals], but it gets the result of signal.pthread_sigmask, which is typed set[int]. The annotation is now set[int].

No runtime change: with annotations erased, the normalized AST of the script differs from the previous revision only in the ProcessHandle protocol body. ProcessHandle is used only in annotations (no subclassing, instantiation, or isinstance). The PR diffstat is now 54 files, +16752/-185. Commit 8 is now the watchdog revision that #71 pins (WATCHDOG_REVISION), replacing 77f77e18 (commit 5).

Twelve-file reconciliation

When this layer was first built, twelve files had changed on master since the old branch's merge-base and had to be reconciled (3-way merge of base/master/old-branch content) rather than blindly overwritten:

  • gguf-py/gguf/constants.py, src/llama-arch.cpp, src/llama-arch.h, common/arg.cpp, include/llama.h, src/llama-context.cpp, src/llama-graph.cpp, src/llama-model-loader.cpp, src/llama-model.h, tests/CMakeLists.txt -- merged with zero conflicts (purely additive on both sides).
  • src/llama-model.cpp -- one conflict: both master (LLAMA_LAZY_MODE_DIRECT skip) and this branch (!ml.external.any() skip) touched the same ml.init_mappings(...) call; resolved by combining both conditions.
  • src/llama-ple-disk.cpp -- the significant one: master had independently reworked the row-gather algorithm (old uniq-sorted scheme -> new order/runs/job_kind scheme with calling-thread work-stealing) since the merge-base. This branch's llama_bounded_file abstraction (exception-based I/O error handling, replacing raw pread/O_DIRECT) was hand-integrated onto master's newer scheduling algorithm, including extending exception propagation to cover the calling thread's own work-stealing pass, and keeping a dedicated raw fd for prefetch()'s posix_fadvise hint (a master-only addition postdating the merge-base that llama_bounded_file doesn't expose).

Scope decision made during implementation

tests/test-deepseek41-runtime.cpp (checked out into this layer, per plan) includes tools/deepseek-v41-trace/trace-components.h to validate the trace-tensor naming contract round-trip. Since that header is explicitly PR B's content, its CTest registration is deferred to that layer (the file exists in this tree but is not yet wired into tests/CMakeLists.txt; PR B will add that registration when it adds the header).

Measurements

Not applicable. This layer only adds new architecture registration, loader plumbing, and CPU-side runtime infrastructure; nothing in this PR is reachable without a DeepSeek V4.1 GGUF, and no such model was loaded or run as part of this work (standing constraint: no inference, no GGUF loaded). There is no throughput or latency claim to make, and no baseline to compare against, until PR B's trace harness can actually drive a model through this code. This section will be populated as needed by future performance work once end-to-end correctness is established.

Correctness:

Model-free build and targeted CTest on the current head (1568e790, commit 8), run on a development machine (macOS/arm64, Apple clang 21, not Strix hardware):

cmake -S . -B build -DLLAMA_BUILD_TESTS=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j 10
ctest --test-dir build -R 'deepseek41|test-engram|test-expert-store|test-llama-archs|test-arg-parser|test-ple-disk' --output-on-failure
100% tests passed out of 9
    test-ple-disk             Passed
    test-deepseek41-schema    Passed
    test-deepseek41-engram    Passed
    test-deepseek41-expert    Passed
    test-expert-store         Passed
    test-deepseek41-memory    Passed
    test-engram               Passed
    test-deepseek41-no-alloc  Passed
    test-arg-parser           Passed

Also on the current head:

  • test-arg-parser and test-ple-disk are in the CTest run because common/arg.cpp and src/llama-ple-disk.cpp are touched here. test-deepseek41-no-alloc is the test-llama-archs binary run with -a deepseek41 -s 1.
  • The build has no compiler warnings in files this PR touches (commit 8 changes no C/C++ source). The only warning under src/ is master's missing prototype for qwen4exp_ple_prefetch (src/models/qwen4exp.cpp:1777).
  • python3 -m pytest -q gguf-py/tests/test_deepseek41_schema.py tests/test_strix_memory_watchdog.py: schema 4 passed; watchdog 36 tests, 28 passed, 8 skipped (Linux-only guardian lifecycle cases), 22 subtests passed. The watchdog test is only wired into CTest on Linux.
  • flake8 4.0.1 + flake8-no-print 0.1.1 on Python 3.11 with the repo .flake8 (same versions the flake8 Lint job resolves): clean on the PR's Python files. A whole-repo run flags only conversion/qwen4exp.py (2x F401), which this PR does not touch.
  • ty 0.0.78 check --exit-zero-on-warning on the whole repo with the repo ty.toml, run separately with --python-platform linux and --python-platform darwin, in a Python 3.11 environment with no third-party packages: 0 diagnostics in the PR's Python files (scripts/strix_memory_watchdog.py, tests/test_strix_memory_watchdog.py, gguf-py/gguf/constants.py, gguf-py/tests/test_deepseek41_schema.py). The rest of the output is identical, line for line, to master on the same platform (only unresolved-import errors from the missing packages, plus unused-ignore warnings). Before commit 8, the linux run had the 4 extra errors described above.
  • editorconfig-checker v3.0.3 (whole repo), the Code Style Checker model-naming snippet (OK: 151 mappings validated.), and git diff --check origin/master...HEAD: clean.

Notes:

  • The POSIX_FADV_WILLNEED guard is part of this PR as commit 7 (ple-disk : guard POSIX_FADV_WILLNEED for non-Linux hosts). POSIX_FADV_WILLNEED does not exist on Apple platforms, and the unguarded call at src/llama-ple-disk.cpp:386 currently fails CI (apple) on master itself (use of undeclared identifier 'POSIX_FADV_WILLNEED' in the iOS/tvOS/visionOS jobs). The guard skips the readahead hint when the macro is missing; Linux behaviour is unchanged.
  • CI (apple) on master also fails in the macOS arm64/x64 jobs on the -Werror missing prototype for qwen4exp_ple_prefetch (src/models/qwen4exp.cpp:1777). That is unrelated to this PR and not fixed here. CI (apple) does not run on pull requests in this repository.

Additional information

Related: #48 (tracking issue), #49 (predecessor, closed for stacking/leak hygiene), #59 (predecessor, closed for the same reason), #60 (companion PR covering the host-memory-guard-in-loader integration, deliberately excluded from this layer's scope).

Requirements

  • I have read and agree with the contributing guidelines
  • This change is Strix Halo specific: DeepSeek V4.1's Engram/expert-store design targets Strix Halo's unified-memory constraints (bounded on-disk expert paging, host-memory watchdog for degraded-swap conditions). No performance claim is made in this layer; that is PR B's job once correctness is established end-to-end.
  • AI usage disclosure: AGENT-AUTHORED (Claude Sonnet 5 / GitHub Copilot; rebase onto current master and commit 8 by Claude Opus 5.5 / GitHub Copilot). See above for what was reconciled vs. directly reused.
  • What was NOT verified: no GGUF loaded, no inference run, no ROCm/Vulkan backend exercised, no Strix hardware run. Cross-runtime and hardware-backed correctness are explicitly out of scope for this layer and are PR B's responsibility. Not run on the current head either: the generated-model CTest cases (test-generate-models and the tests that load those models), a Linux/GCC or Windows build, the Linux-only watchdog cases (commit 8 was checked with ty --python-platform linux, not run on Linux), and ty with CI's requirements-all packages installed.

ajaxdude and others added 7 commits September 25, 2026 21:43
Add the DEEPSEEK41 model architecture enum, its GGUF metadata keys
(hidden size, MoE routing, index attention, Engram config) and tensor
names, plus the DeepSeek-V4.1-specific hparams fields (kv/index source
layer maps, engram layer bitset). Reconciled onto the current gguf-py
constants and llama-arch content.

Assisted-by: Claude Sonnet 5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
… infra

Add llama_bounded_file, a direct-IO-aware file reader abstraction with
exception-based error handling, and integrate it into llama_ple_disk's
row-gather path (replacing the old raw pread/O_DIRECT logic) while
keeping master's current gather/dequant job scheduling and read-size
accounting intact.

Add llama_engram (fixed-capacity KV/index recall state) and
llama_expert_store (host-resident MoE expert weight cache) as new
standalone modules, plus their unit tests. Extend llama_mmap,
llama_kv_cache and llama_memory with the small hooks these modules
need.

Assisted-by: Claude Sonnet 5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add llama-dsv41, the DeepSeek V4.1-specific glue that binds Engram
recall state and the expert store to the model runtime hparams, and
llama-memory-dsv41, the llama_memory implementation that admits and
schedules Engram/expert state per sequence. Add unit tests for the
engram runtime, expert runtime, and memory admission paths.

Assisted-by: Claude Sonnet 5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add llama_model_deepseek41's arch graph, hparam and tensor loading in
src/models/deepseek41.cpp, and register it in the model factory and
llama_model::create_memory. Reconcile llama-model.cpp, llama-model.h,
llama-model-loader.cpp, llama-context.cpp and llama-graph.cpp against
master's own drift since the merge-base this content is sourced from:

- llama-model-loader gains register_external_tensor and an
  external_read tracker so expert-store tensors can be read straight
  from disk instead of being mapped or fully loaded; load_all_data and
  init_mappings take explicit load_from_mmap/discard_file_cache
  parameters instead of assuming the loader-wide mmap setting.
- llama-model's load_tensors keeps mmap unsafe for any context that
  overlaps an external tensor range, still applying master's own
  lazy-mode-direct mapping skip alongside it.
- llama-context adds a memory-context commit/rollback transaction
  guard around graph compute, and model runtime acquire/release hooks
  used by the dsv41 engram/expert runtime.
- llama-graph's build_moe_ffn takes a separate expert lookup tensor
  from the routing tensor, needed because dsv41's expert store lets
  the physical MoE weight slot differ from the logical expert id.

Add --expert-cache-slots/--expert-cache-mib CLI options and extend
test-llama-archs.cpp with a no-allocation DeepSeek V4.1 graph-shape
smoke test (its full published geometry is too large for the compact
generated-model fixture used by the rest of that test).

Assisted-by: Claude Sonnet 5
Assisted-by: Claude Opus 5.5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add scripts/strix_memory_watchdog.py, a standalone process-group
supervisor that requires zero active swap and stops a wrapped command
before host-wide memory reaches a validation ceiling. This is a plain
external monitoring script: it does not touch or link against llama.cpp
runtime code, and is unrelated to the (deliberately excluded from this
layer) in-process host-memory-guard integration.

Add its unit tests, usage docs, and a README table entry. Use a full
git commit hash instead of the short form for GGML_BUILD_COMMIT so a
build can be traced back unambiguously during a watchdog-flagged run.

Assisted-by: Claude Sonnet 5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Register the new DeepSeek V4.1 unit tests (schema, engram, expert
store, memory admission, runtime) and the test-deepseek41-no-alloc
allocation-guard case for the deepseek41 arch. Register the Python
strix-memory-watchdog unit test alongside the other python-labeled
tests.

This intentionally omits the trace-harness and native-containment
test registrations (test-server-child-protocol.cpp, test-deepseek41-
trace.py, LLAMA_DEEPSEEK_V41_NATIVE_CONTAINMENT_TESTS); those depend
on tools/deepseek-v41-trace/ and tools/server/ content that is out of
scope for this layer and will be added by the next layer.

Assisted-by: Claude Sonnet 5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
POSIX_FADV_WILLNEED is not declared on macOS, so the unconditional
posix_fadvise() call in prefetch() failed to compile there. Guard it
the same way discard_cache() already guards POSIX_FADV_DONTNEED in
llama-mmap.cpp: skip the readahead hint when the macro is unavailable,
Linux behaviour is unchanged.

Assisted-by: Claude Sonnet 5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@ajaxdude
ajaxdude force-pushed the deepseek-v41-core-runtime branch from 4fa1eb4 to 0541e3e Compare September 26, 2026 04:54
ty reports 4 errors in scripts/strix_memory_watchdog.py when it checks
for Linux. The guardian launch path in run_watchdog() only runs on
Linux, so a darwin check does not reach it.

- ProcessHandle declared pid as a writable attribute, but
  GuardianProcess has pid as a read-only property, so a GuardianProcess
  was not a ProcessHandle. Declare pid as a read-only property in the
  protocol. subprocess.Popen and plain pid attributes still match it.
- _launch_guardian() typed the saved signal mask as set[Signals], but
  signal.pthread_sigmask() returns plain ints for signal numbers with
  no Signals member, and ty types its result as set[int]. Type the
  mask as set[int].

The other 2 errors (child is ProcessHandle | None at
LeaseManager.start() and _monitor_child()) came from the first one and
go away with it.

Typing only, no runtime change: the file uses
"from __future__ import annotations", and nothing subclasses,
instantiates or runs isinstance() on ProcessHandle.

Assisted-by: Claude Opus 5.5
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@ajaxdude
ajaxdude requested review from dzannotti and removed request for dzannotti September 26, 2026 14:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

build conversion documentation Improvements or additions to documentation ggml model testing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant