Conversation
ajaxdude
added this pull request to stack #72
September 17, 2026 02:50
This was referenced Sep 17, 2026
Add the DEEPSEEK41 model architecture enum, its GGUF metadata keys (hidden size, MoE routing, index attention, Engram config) and tensor names, plus the DeepSeek-V4.1-specific hparams fields (kv/index source layer maps, engram layer bitset). Reconciled onto the current gguf-py constants and llama-arch content. Assisted-by: Claude Sonnet 5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
… infra Add llama_bounded_file, a direct-IO-aware file reader abstraction with exception-based error handling, and integrate it into llama_ple_disk's row-gather path (replacing the old raw pread/O_DIRECT logic) while keeping master's current gather/dequant job scheduling and read-size accounting intact. Add llama_engram (fixed-capacity KV/index recall state) and llama_expert_store (host-resident MoE expert weight cache) as new standalone modules, plus their unit tests. Extend llama_mmap, llama_kv_cache and llama_memory with the small hooks these modules need. Assisted-by: Claude Sonnet 5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add llama-dsv41, the DeepSeek V4.1-specific glue that binds Engram recall state and the expert store to the model runtime hparams, and llama-memory-dsv41, the llama_memory implementation that admits and schedules Engram/expert state per sequence. Add unit tests for the engram runtime, expert runtime, and memory admission paths. Assisted-by: Claude Sonnet 5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add llama_model_deepseek41's arch graph, hparam and tensor loading in src/models/deepseek41.cpp, and register it in the model factory and llama_model::create_memory. Reconcile llama-model.cpp, llama-model.h, llama-model-loader.cpp, llama-context.cpp and llama-graph.cpp against master's own drift since the merge-base this content is sourced from: - llama-model-loader gains register_external_tensor and an external_read tracker so expert-store tensors can be read straight from disk instead of being mapped or fully loaded; load_all_data and init_mappings take explicit load_from_mmap/discard_file_cache parameters instead of assuming the loader-wide mmap setting. - llama-model's load_tensors keeps mmap unsafe for any context that overlaps an external tensor range, still applying master's own lazy-mode-direct mapping skip alongside it. - llama-context adds a memory-context commit/rollback transaction guard around graph compute, and model runtime acquire/release hooks used by the dsv41 engram/expert runtime. - llama-graph's build_moe_ffn takes a separate expert lookup tensor from the routing tensor, needed because dsv41's expert store lets the physical MoE weight slot differ from the logical expert id. Add --expert-cache-slots/--expert-cache-mib CLI options and extend test-llama-archs.cpp with a no-allocation DeepSeek V4.1 graph-shape smoke test (its full published geometry is too large for the compact generated-model fixture used by the rest of that test). Assisted-by: Claude Sonnet 5 Assisted-by: Claude Opus 5.5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Add scripts/strix_memory_watchdog.py, a standalone process-group supervisor that requires zero active swap and stops a wrapped command before host-wide memory reaches a validation ceiling. This is a plain external monitoring script: it does not touch or link against llama.cpp runtime code, and is unrelated to the (deliberately excluded from this layer) in-process host-memory-guard integration. Add its unit tests, usage docs, and a README table entry. Use a full git commit hash instead of the short form for GGML_BUILD_COMMIT so a build can be traced back unambiguously during a watchdog-flagged run. Assisted-by: Claude Sonnet 5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Register the new DeepSeek V4.1 unit tests (schema, engram, expert store, memory admission, runtime) and the test-deepseek41-no-alloc allocation-guard case for the deepseek41 arch. Register the Python strix-memory-watchdog unit test alongside the other python-labeled tests. This intentionally omits the trace-harness and native-containment test registrations (test-server-child-protocol.cpp, test-deepseek41- trace.py, LLAMA_DEEPSEEK_V41_NATIVE_CONTAINMENT_TESTS); those depend on tools/deepseek-v41-trace/ and tools/server/ content that is out of scope for this layer and will be added by the next layer. Assisted-by: Claude Sonnet 5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
POSIX_FADV_WILLNEED is not declared on macOS, so the unconditional posix_fadvise() call in prefetch() failed to compile there. Guard it the same way discard_cache() already guards POSIX_FADV_DONTNEED in llama-mmap.cpp: skip the readahead hint when the macro is unavailable, Linux behaviour is unchanged. Assisted-by: Claude Sonnet 5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
ajaxdude
force-pushed
the
deepseek-v41-core-runtime
branch
from
September 26, 2026 04:54
4fa1eb4 to
0541e3e
Compare
ty reports 4 errors in scripts/strix_memory_watchdog.py when it checks for Linux. The guardian launch path in run_watchdog() only runs on Linux, so a darwin check does not reach it. - ProcessHandle declared pid as a writable attribute, but GuardianProcess has pid as a read-only property, so a GuardianProcess was not a ProcessHandle. Declare pid as a read-only property in the protocol. subprocess.Popen and plain pid attributes still match it. - _launch_guardian() typed the saved signal mask as set[Signals], but signal.pthread_sigmask() returns plain ints for signal numbers with no Signals member, and ty types its result as set[int]. Type the mask as set[int]. The other 2 errors (child is ProcessHandle | None at LeaseManager.start() and _monitor_child()) came from the first one and go away with it. Typing only, no runtime change: the file uses "from __future__ import annotations", and nothing subclasses, instantiates or runs isinstance() on ProcessHandle. Assisted-by: Claude Opus 5.5 Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
This is Layer 1 (PR A) of a two-PR stack reviving closed PR #59 ("trace : recover DeepSeek V4.1 correctness harness"), which was closed alongside #49 and #60 for two reasons: (1) all three PRs targeted
masterdirectly instead of stacking, so each diff showed the combined content of its unmerged predecessors; (2) leaked personal paths. This PR succeeds #49 with the same substantive content, properly restacked with clean history and no personal paths.This layer adds the core DeepSeek V4.1 runtime, scoped as the model-execution half of the correctness-harness effort tracked in #48:
LLM_ARCH_DEEPSEEK41registration and GGUF schema (gguf-py/gguf/constants.py,src/llama-arch.{cpp,h})src/models/deepseek41.cpp,src/llama-model.cpp,src/llama-model-loader.cpp,src/llama-graph.cpp,src/llama-context.cpp,include/llama.h,common/arg.cpp)src/llama-engram.*,src/llama-expert-store.*,src/llama-dsv41-engram.*,src/llama-dsv41-expert.*,src/llama-dsv41.*)src/llama-memory-dsv41.*,src/llama-bounded-file.*,src/llama-ple-disk.cppintegration)scripts/strix_memory_watchdog.py,docs/strix-memory-watchdog.md)POSIX_FADV_WILLNEEDguard insrc/llama-ple-disk.cpp(commit 7) that fixes a pre-existing macOS/Apple build error onmaster(see "Notes" below)PR B (#71, the correctness trace harness under
tools/deepseek-v41-trace/) is stacked on top of this branch. This layer deliberately does not includetools/deepseek-v41-trace/,tools/server/integration, or the PR60 host-memory-guard-in-loader scope; those are out of scope here.AI usage disclosure: AGENT-AUTHORED. I (Claude Sonnet 5, running as GitHub Copilot) rebuilt this layer from scratch: reconciled the 12 files that drifted between the old closed branch and current
master(see below), re-applied the validated DeepSeek V4.1 content on top of currentmaster, and wrote fresh commits. The later rebase onto currentmaster, its conflict resolution, and the typing-only commit 8 (below) were done by Claude Opus 5.5, also running as GitHub Copilot.Rebase onto current
masterRebased onto
masterat03895887abe6b46473e0a2e07d0a82fe44484894(previous base:0636c9aee4a2ab1d38bb7e8c038227c6634c3857, 99 commits behind). Commits 1-3 and 5-7 replayed unchanged (git range-diffreports them as=). All conflicts were in commit 4 (deepseek41 : wire model graph and loader):--ngram-on-disk and lazymode merge (#29)removed theple_on_disk,ple_direct_io,ple_io_threadsandple_cache_mbparams and the--ngram-io-threads/--ngram-cache/--ngram-direct-iooptions (replaced by--lazy-mode on-direct). This PR had added its expert-cache fields next to those lines. Master's removal is kept; only this PR's additions are re-anchored:--expert-cache-slots/--expert-cache-mibnow follow--ngram-on-disk(before--cpu-moe) incommon/arg.cpp, thecommon_paramsfields andcommon_model_params_to_llamaassignments followno_host, andexpert_cache_bytes/expert_cache_slotsinllama_model_paramsfollowmain_gpu, in the same order inllama_model_default_params.load_tensors:init_mappingskeeps master's lazy-mode-direct skip plus this PR's external-tensor guard, i.e.!(params.lazy_mode == LLAMA_LAZY_MODE_DIRECT || ml.external.any());params.ple_on_diskis gone.build_moe_ffnkeepsLLM_ARCH_MAPLEand addsLLM_ARCH_DEEPSEEK41;tests/test-llama-archs.cppkeeps master's MAPLE block next to the DeepSeek V4.1 block.Against the pre-rebase patch, the only changed
+/-lines of the whole PR diff are theinit_mappingscondition and the MAPLE term in the swiglu-clamp condition; the diffstat is unchanged (54 files, +16750/-185). Commit 4's message lost a staleple_on_diskreference and gained anAssisted-by: Claude Opus 5.5trailer. All 8 commits, including commit 8 below, are authored and committed asajaxdude.Typing-only follow-up (commit 8)
Commit 8 (
1568e7903cc07fc83fee1894df491dfcd2f7ba9c,strix : fix ty findings in memory watchdog on Linux) changes onlyscripts/strix_memory_watchdog.py(+4/-2).tyreported 4 errors in the guardian launch path, but only with--python-platform linux: on macOS that branch is unreachable (use_guardianrequiressys.platform.startswith("linux")), so the earliertycheck missed them.ProcessHandledeclaredpid: intas a writable protocol attribute, butGuardianProcessexposespidas a read-only property. Sotyrejected the assignment tochild: ProcessHandle | None, and the two later uses ofchildfell back toProcessHandle | None. The protocol now declarespidas a read-only property.subprocess.Popenand the test fakes still satisfy it._launch_guardianannotatedlaunch_maskasset[signal.Signals], but it gets the result ofsignal.pthread_sigmask, which is typedset[int]. The annotation is nowset[int].No runtime change: with annotations erased, the normalized AST of the script differs from the previous revision only in the
ProcessHandleprotocol body.ProcessHandleis used only in annotations (no subclassing, instantiation, orisinstance). The PR diffstat is now 54 files, +16752/-185. Commit 8 is now the watchdog revision that #71 pins (WATCHDOG_REVISION), replacing77f77e18(commit 5).Twelve-file reconciliation
When this layer was first built, twelve files had changed on
mastersince the old branch's merge-base and had to be reconciled (3-way merge of base/master/old-branch content) rather than blindly overwritten:gguf-py/gguf/constants.py,src/llama-arch.cpp,src/llama-arch.h,common/arg.cpp,include/llama.h,src/llama-context.cpp,src/llama-graph.cpp,src/llama-model-loader.cpp,src/llama-model.h,tests/CMakeLists.txt-- merged with zero conflicts (purely additive on both sides).src/llama-model.cpp-- one conflict: both master (LLAMA_LAZY_MODE_DIRECTskip) and this branch (!ml.external.any()skip) touched the sameml.init_mappings(...)call; resolved by combining both conditions.src/llama-ple-disk.cpp-- the significant one: master had independently reworked the row-gather algorithm (olduniq-sorted scheme -> neworder/runs/job_kindscheme with calling-thread work-stealing) since the merge-base. This branch'sllama_bounded_fileabstraction (exception-based I/O error handling, replacing rawpread/O_DIRECT) was hand-integrated onto master's newer scheduling algorithm, including extending exception propagation to cover the calling thread's own work-stealing pass, and keeping a dedicated raw fd forprefetch()'sposix_fadvisehint (a master-only addition postdating the merge-base thatllama_bounded_filedoesn't expose).Scope decision made during implementation
tests/test-deepseek41-runtime.cpp(checked out into this layer, per plan) includestools/deepseek-v41-trace/trace-components.hto validate the trace-tensor naming contract round-trip. Since that header is explicitly PR B's content, its CTest registration is deferred to that layer (the file exists in this tree but is not yet wired intotests/CMakeLists.txt; PR B will add that registration when it adds the header).Measurements
Not applicable. This layer only adds new architecture registration, loader plumbing, and CPU-side runtime infrastructure; nothing in this PR is reachable without a DeepSeek V4.1 GGUF, and no such model was loaded or run as part of this work (standing constraint: no inference, no GGUF loaded). There is no throughput or latency claim to make, and no baseline to compare against, until PR B's trace harness can actually drive a model through this code. This section will be populated as needed by future performance work once end-to-end correctness is established.
Correctness:
Model-free build and targeted CTest on the current head (
1568e790, commit 8), run on a development machine (macOS/arm64, Apple clang 21, not Strix hardware):Also on the current head:
test-arg-parserandtest-ple-diskare in the CTest run becausecommon/arg.cppandsrc/llama-ple-disk.cppare touched here.test-deepseek41-no-allocis thetest-llama-archsbinary run with-a deepseek41 -s 1.src/is master's missing prototype forqwen4exp_ple_prefetch(src/models/qwen4exp.cpp:1777).python3 -m pytest -q gguf-py/tests/test_deepseek41_schema.py tests/test_strix_memory_watchdog.py: schema 4 passed; watchdog 36 tests, 28 passed, 8 skipped (Linux-only guardian lifecycle cases), 22 subtests passed. The watchdog test is only wired into CTest on Linux.flake8-no-print0.1.1 on Python 3.11 with the repo.flake8(same versions theflake8 Lintjob resolves): clean on the PR's Python files. A whole-repo run flags onlyconversion/qwen4exp.py(2x F401), which this PR does not touch.ty 0.0.78 check --exit-zero-on-warningon the whole repo with the repoty.toml, run separately with--python-platform linuxand--python-platform darwin, in a Python 3.11 environment with no third-party packages: 0 diagnostics in the PR's Python files (scripts/strix_memory_watchdog.py,tests/test_strix_memory_watchdog.py,gguf-py/gguf/constants.py,gguf-py/tests/test_deepseek41_schema.py). The rest of the output is identical, line for line, tomasteron the same platform (only unresolved-import errors from the missing packages, plus unused-ignore warnings). Before commit 8, the linux run had the 4 extra errors described above.Code Style Checkermodel-naming snippet (OK: 151 mappings validated.), andgit diff --check origin/master...HEAD: clean.Notes:
POSIX_FADV_WILLNEEDguard is part of this PR as commit 7 (ple-disk : guard POSIX_FADV_WILLNEED for non-Linux hosts).POSIX_FADV_WILLNEEDdoes not exist on Apple platforms, and the unguarded call atsrc/llama-ple-disk.cpp:386currently failsCI (apple)onmasteritself (use of undeclared identifier 'POSIX_FADV_WILLNEED'in the iOS/tvOS/visionOS jobs). The guard skips the readahead hint when the macro is missing; Linux behaviour is unchanged.CI (apple)onmasteralso fails in the macOS arm64/x64 jobs on the-Werrormissing prototype forqwen4exp_ple_prefetch(src/models/qwen4exp.cpp:1777). That is unrelated to this PR and not fixed here.CI (apple)does not run on pull requests in this repository.Additional information
Related: #48 (tracking issue), #49 (predecessor, closed for stacking/leak hygiene), #59 (predecessor, closed for the same reason), #60 (companion PR covering the host-memory-guard-in-loader integration, deliberately excluded from this layer's scope).
Requirements
masterand commit 8 by Claude Opus 5.5 / GitHub Copilot). See above for what was reconciled vs. directly reused.test-generate-modelsand the tests that load those models), a Linux/GCC or Windows build, the Linux-only watchdog cases (commit 8 was checked withty --python-platform linux, not run on Linux), andtywith CI'srequirements-allpackages installed.