Repository navigation
Conversation
added 2 commits
October 3, 2026 11:34
27B loads stop at a varying `loading layer N/64` under host-memory pressure. It is not an allocation failure: at the stall the daemon's main thread burns 100% of one core (399 jiffies / 4 s, no syscalls) while rocm-smi reports GPU 14% / 0% — a lost KFD completion, not a long kernel. Root cause. Under ROCm's defaults the host pages the GPU reads are KFD userptr BOs, in two classes: (1) every hipHostMalloc block (Qwen4 routed experts, partial-GPU-offload spilled layers, clr staging); (2) the source of any large pageable copy — clr pins the mapped HFQ file's page-cache pages in place for every multi-MB weight upload. Reclaiming either class evicts all of the process's GPU queues until KFD's restore worker faults the pages back, and under sustained pressure HIP's signal wait spins a core with no error (ROCm/rocm-systems#12528; already described in keep_host_memory_out_of_reclaim's comment, which records it for host-mapped experts AND for "weight uploads out of the mapped file"). Class (2) is on every HFQ load, on every architecture. The switches were once global (f5731a5, 2026-09-30 15:48) and were narrowed to host-maps_qwen4_experts() by 5df32fe (16:16), on the premise that Qwen4 experts were the only host-mapped class — which was already false for the dense partial-offload path added that morning (e3066e7, 04:24), and left class (2) uncovered for every other model too. The offload work did not introduce the bug; the narrowing did. Fix. One switch, armed for every load: - GPU_PINNED_MIN_XFER_SIZE=100000 in every Linux process (unless set). It routes each pageable copy through clr's pinned staging buffer, so the copy source is never registered as reclaimable userptr — class (2), closed generically for every model and every placement. - HSA_USERPTR_FOR_PAGED_MEM=0 is unchanged from master: still scoped to host_maps_qwen4_experts(). It is the half that costs, and it moves host-mapped memory under TTM pages_limit; the ~20% in 5df32fe's message is this switch, not the one added here. No new tier is introduced. - One `[hipfire] host-memory reclaim mitigation armed: ...` line names what was applied. Neither switch name is a published interface, and the docs now say so: GPU_PINNED_MIN_XFER_SIZE is absent from HIP's environment-variable reference and from AMD's env-var reference (checked 2026-10-03, ROCm 7.2) and a verbatim docs search returns nothing, so docs/env-vars.md and the code comment record it as an internal clr knob, read once at runtime init, default unpublished, that can change without notice. HSA_USERPTR_FOR_PAGED_MEM is a libhsakmt knob of the same kind. f5731a5 relied on both first, for Qwen4. The switch is set in HipRuntime::load(), before any model is chosen, so it applies to every architecture and model by construction — there is no per-model allowlist to keep in sync (the previous gate was one). Cross-arch check on the branch binary. One binary covers both arms: `pre` pre-seeds GPU_PINNED_MIN_XFER_SIZE=0, which is master's effective environment for the copy-source class (the gate never overrides an operator value), and that switch is the only behavioural delta. `ship` lets the gate arm it. model load pre/ship generation (256 tok) sweep unarmed -> armed (3 warm reps) lfm2.5:350m (LFM2 dense) yes/yes "Paris" loader prints no sweep line lfm2.5:1.2b (LFM2 dense) yes/yes "Paris" loader prints no sweep line lfm2.5:8b-a1b (LFM2 MoE) yes/yes no readable text loader prints no sweep line qwen3.5:2b (Qwen35 dense) yes/yes "Paris" 128-131 -> 131-133 ms qwen3.5:9b (Qwen35 dense) yes/yes "Paris" 365-371 -> 364-373 ms mimo-qwen (32 layers) yes/yes "Paris" 367-368 -> 364-387 ms qwen3.8-27b (offload) yes/yes - see the cost block below (under 8 GiB ballast only the armed arm loads) Nothing is slower: every per-model delta is inside that model's own rep spread. The 8b-a1b emits tokens but no readable answer to this prompt. muse-glimmer could not be covered. Its HF SKUs are 16.26 GB (mq4r) / 18.61 GB (mq4), and a local muse-glimmer-30b.mq2-xt (13.43 GB) fails identically armed and unarmed, before any weight copy: load failed: glimmer: unsupported quant_type 50 for model.language_model.layers.0.mlp.gate_proj.weight. That is a glimmer/MQ2-XT support gap (qt=50), not this change; the card also reports 13218 MB free against a 13431 MB file. Cost, measured on gfx1201 (warm cache, arms alternated so drift cancels, the first rep of each arm dropped as a cold-cache outlier; `weight sweep` as the loader prints it; median of the remaining reps in brackets): 27B unarmed 1550 / 1278 / 1151 / 2349 ms [1414] armed 1154 / 1139 / 1454 / 1239 ms [1197] 9B unarmed 381 / 384 / 389 / 384 ms [ 384] armed 380 / 372 / 383 / 381 / 378 ms [379] No measurable penalty. The unarmed arm's own spread (1151-2349 ms) is wider than the gap between arms, so this says "not slower", not "faster": the staged copy is not worse than the page-pinning it replaces. Effectiveness, 8 GiB host-memory ballast, same model: arm partial offload fully resident unarmed stall (layers 9, 45, 59) stall 5/5 (layers 49/35/31/33/42) HSA only load 1/1 stall 4/4 GPU_PINNED only load 3/3 load 4/4 both load 1/1 - HSA alone cannot cover the resident case (no hipHostMalloc weight block to move); GPU_PINNED alone covers both and is free. The fully-resident 27B is reachable on this 17.1 GB card at max_seq=4096 (a larger max_seq is refused by admission before any load, which is why the resident arm needs the smaller context to reproduce). Follow-up, not fixed here: serve exits through process::exit on the second SIGTERM and on pre-warm failure, skipping EngineInner::drop, so the daemon child (with its VRAM and the GPU flock) is orphaned. Observed live — daemon 45609 outlived serve 45608 holding 10 GB — and it explains the operator's `hipfire daemon already running` / GPU-lock FATALs. Separate change. Signed-off-by: Avery Drouillard <avery@averynet.xyz>
…LOAD_TRACE Extends the existing HIPFIRE_LOAD_TRACE knob (previously the packed-expert read/upload timing in qwen35/load.rs) with the allocation-level view: - rdna-compute/src/dispatch.rs: one `[load-trace]` line per upload_raw / upload_raw_host / alloc_host_mapped_tensor / free_tensor park / release_tensor_immediate, each carrying free VRAM and the pool counters (pool_new, pool_reused, pool_alloc_bytes, pool_parked_bytes); - rdna-compute/src/pool.rs: GpuPool::freelist_bytes(), surfaced as Gpu::pool_freelist_bytes() — bytes parked in the free lists, which a failed load's rollback leaves unreachable to upload_raw until drain_pool; - hipfire-runtime/src/weight_backend.rs: proj / norm / raw_f32 markers, so the last line before a stall names the tensor being read; - hipfire-arch-qwen35/src/qwen35/load.rs: one per-layer summary with free VRAM and the pool counters. Gated and a strict no-op unless the knob is set (read through hipfire_config::developer_var_os, so it works as an env var or a [developer] TOML key). Use it to localize, not to benchmark: every line calls hipMemGetInfo. This is the diagnostic that localized the 27B load stall to a spin in the weight-upload path; the fix it produced is the previous commit. It shares one knob with the existing packed-expert line and uses the same `[load-trace]` prefix. docs/env-vars.md: HIPFIRE_LOAD_TRACE's source list gains the four new files. Signed-off-by: Avery Drouillard <avery@averynet.xyz>
CI's `gates (ratchets, layering, registers)` job failed on "Crate maps match
the tree": hip-bridge, rdna-compute, hipfire-runtime and hipfire-arch-qwen35
all gained lines, and each crate's `map.md` is generated from the tree.
Regenerated with:
python3 scripts/check-crate-maps.py hip-bridge hipfire-arch-qwen35 hipfire-runtime rdna-compute
Line-count deltas only; `--check` now reports "45 map(s) match the tree".
Signed-off-by: Avery Drouillard <avery@averynet.xyz>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Fixes HFQ loads that stall forever at a varying
loading layer N/64under host-memory pressure, by arming clr'spageable-copy switch (
GPU_PINNED_MIN_XFER_SIZE=100000) in every Linux process instead of only in processeshost-mapping Qwen4 experts. Behaviour-neutral where it was already set; the
HSA_USERPTR_FOR_PAGED_MEMarm isunchanged from master.
The stall is not an allocation failure: at the hang the daemon's main thread burns one core (399 jiffies / 4 s, no
syscalls) while
rocm-smireports the GPU idle (14 % / 0 %) — HIP's signal wait spinning after KFD evicts theprocess's queues because the kernel reclaimed KFD-userptr pages the GPU was reading from. That class
(ROCm/rocm-systems#12528) is already described in
keep_host_memory_out_of_reclaim's comment, which records it forboth host-mapped experts and "weight uploads out of the mapped file";
5df32fe89(2026-09-30 16:16) narrowed themitigation to Qwen4 experts only, which left the pageable-copy class uncovered for every other model — and for the
dense partial-offload path that had landed 12 h earlier (
e3066e739).Branch:
fix/host-memory-reclaim-stall(forkaldrouil/hipfire), based onmaster@a89ed0a8e— the currentwarpfront/hipfireHEAD, no rebase needed.Which surface(s) does this touch?
crates/hip-bridge(the switch),crates/rdna-compute(load tracing)weight_backend), archload.rshipfire-arch-qwen35(load tracing only)crates/hipfire-quantize/ quant formatsTest plan
./scripts/no-gpu-ci.shpasses — ran 2026-10-03:cargo check --workspace --examples, the per-crate lib tests,pytest (179 + 14 + 6 tests OK),
env-docs: references covered,check-lifecyclecleancargo build --releasecleancargo test --lib --workspace --lockedpasses — not green on this box, for two pre-existing environmentalreasons, neither touched by this branch (both reproduced on mainline
a89ed0a8e): (1)hipfire-loader::admission::tests::head_overlay::head_on_non_maple_refusesfails when this machine's~/.hipfire/config.tomlis visible and passes with a cleanHOME; (2)peacemaker-ir::tests::opcode_examples_match_pinned_llvm_mc_for_every_declared_formfails for want of the pinnedllvm-mcon this host. With a clean HOME the run is 2840 passed / 1 failed (thepeacemaker-irone). CI's jobruns
--lockedon a clean runner and is not expected to hit either.python3 scripts/serve_harness.py --model ~/.hipfire/models/qwen3.8-27b.mq3-xt --tag qwen3.8:27b-mq3-xt --mode battery --kv q8 --max-seq 16384 --max-tokens 4096 --max-think-tokens 0 --thinking-effort noneon hardware — 5/5 turns,runaway=0 empty=0 attractor=0 retrieval_miss=0,finish=stopon all five, avg decode 46.5 tok/s,kv_backend effective: vmm (log-marker), 42 s wall. JSON below.kv_backend,kv_backend_legacy,kv_backend_reason); theLFM2 runs used an automatic legacy KV backend (
carrier lfm2moe has no VMM KV owner) — token isautomatic,not explicit, and it is disclosed in "Notes for the reviewer". No command set
HIPFIRE_KV_BACKEND=legacy.selects clr's staging path for pageable copies. Measured load cost on this box (warm cache, arms alternated,
first rep dropped, median in brackets):
27B weight sweepunarmed 1550/1278/1151/2349 ms[1414]vs armed 1154/1139/1454/1239 ms[1197];9Bunarmed 381/384/389/384 ms[384]vs armed 380/372/383/381/378 ms[379].The unarmed arm's own spread exceeds the gap, so this says "not slower", not "faster".
leanup-thresholds.txtceiling moved)What was verified on hardware
Both arms, one binary (
prepre-seedsGPU_PINNED_MIN_XFER_SIZE=0= master's effective env for this class; thegate never overrides an operator value):
swift-qwen3.8-27b.mq3-pro(partial offload, configgpu_layer_budget=55)qwen3.5:9b,qwen3.5:2b"Paris"lfm2.5:350m,lfm2.5:1.2b(LFM2 dense)"Paris"mimo-qwen(32 layers)"Paris"Stall reproduction/resolution, 8 GiB host-memory ballast held by another process:
HSAonlyGPU_PINNEDonlyArtifacts used, per block:
serve_harnessbatteryqwen3.8-27b.mq3-xt(qwen3.8:27b-mq3-xt), size 117776168963e04fc8db80bda557b965ec60ac876cf2500fced7f340624f3fcbeae134af5c5swift-qwen3.8-27b.mq3-pro, size 136028764161bcd36128f8bb58677d7444234a75ebad82ad8b110843ebcdf0e1262c9369744lfm2.5:350m/:1.2b/:8b-a1b(sha256 verified byhipfire pull)7cd38949…,a7aeba87…,07da9942…hipfire/daemonb980aef35c023fe31b5a23ddef3a0012/a633358a27de5dc003b515b318366d54Neither pinned dense-trunk fixture was used —
qwen3.8:27b-mq4-xts(AGENTS.md§5) andqwen3.8:27b-mq4-xt(
scripts/hw-gate/fixtures.json) are both not installed on this box — and no evidence above depends on trunk identity:the claim is load-path behaviour, measured as a relative before/after on one artifact. The 15.66 GB
qwen3.8-27b.mq4on disk (sha256
5bb556a6…) was used only for the harness-fit diagnosis described in "Reproducing the battery".muse-glimmercould not be covered: its registry SKUs are 16.26 GB (mq4r) / 18.61 GB (mq4) against ~13.2 GB freeon the 17.1 GB card, host-RAM offload is not wired for it, and a locally quantized
kvist-glimmer-30b.mq4(arch 14)fails at load in both arms, before any copy —
glimmer: unsupported quant_type 44 for model.language_model.layers.0.mlp.down_proj.weight— a glimmer/MQ4V2 load-support gap, not this change.local serve_harness battery.json (load / serve / kernel changes)
[ { "request_id": "chatcmpl-231767-1", "ctx": 44, "cached": 0, "gen": 205, "finish": "stop", "think_words": 0, "ans_words": 99, "prefill_ms": 4011.8, "prefill_tok_s": 11.0, "decode_tok_s": 41.7, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 4.649, "wall_s": 8.928, "attractor": false, "empty": false, "runaway": false, "ans_preview": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted lists into a single sorted list.", "assistant_content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted lists into a single sorted list.\n \n Uses a two-pointer approach to combine the lists in O(n + m) time.\n Does not use sorted() or any built-in sorting function.\n \"\"\"\n merged = []\n i = 0\n j = 0\n len_a = len(a)\n len_b = len(b)\n\n while i < len_a and j < len_b:\n if a[i] <= b[j]:\n merged.append(a[i])\n i += 1\n else:\n merged.append(b[j])\n j += 1\n\n # Append any remaining elements from either list\n while i < len_a:\n merged.append(a[i])\n i += 1\n while j < len_b:\n merged.append(b[j])\n j += 1\n\n return merged\n```", "content": "```python\ndef merge_sorted(a, b):\n \"\"\"Merge two sorted lists into a single sorted list.\n \n Uses a two-pointer approach to combine the lists in O(n + m) time.\n Does not use sorted() or any built-in sorting function.\n \"\"\"\n merged = []\n i = 0\n j = 0\n len_a = len(a)\n len_b = len(b)\n\n while i < len_a and j < len_b:\n if a[i] <= b[j]:\n merged.append(a[i])\n i += 1\n else:\n merged.append(b[j])\n j += 1\n\n # Append any remaining elements from either list\n while i < len_a:\n merged.append(a[i])\n i += 1\n while j < len_b:\n merged.append(b[j])\n j += 1\n\n return merged\n```", "reasoning_content": "", "tool_calls": [], "request_md5": "e900b2d56ac907bb93c348f1e90e1da8", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "43ca0d15712d3dfb777b51ae76d8fd5f", "expected_substrings": [], "retrieval_missing": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_warning": null, "kv_backend_reason": null, "kv_mode": null }, { "request_id": "chatcmpl-231767-3", "ctx": 55, "cached": 0, "gen": 332, "finish": "stop", "think_words": 0, "ans_words": 157, "prefill_ms": 101.0, "prefill_tok_s": 544.5, "decode_tok_s": 47.7, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.125, "wall_s": 7.059, "attractor": false, "empty": false, "runaway": false, "ans_preview": "To find the total distance traveled by the train, we need to calculate the distance for ea", "assistant_content": "To find the total distance traveled by the train, we need to calculate the distance for each segment of the trip separately and then sum them up.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first segment.**\n* Speed ($v_1$) = 60 mph\n* Time ($t_1$) = 2.5 hours\n\n$$ d_1 = 60 \\, \\text{mph} \\times 2.5 \\, \\text{hours} = 150 \\, \\text{miles} $$\n\n**Step 2: Calculate the distance for the second segment.**\n* Speed ($v_2$) = 40 mph\n* Time ($t_2$) = 1.5 hours\n\n$$ d_2 = 40 \\, \\text{mph} \\times 1.5 \\, \\text{hours} = 60 \\, \\text{miles} $$\n\n**Step 3: Add the two distances together to find the total distance.**\n\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 \\, \\text{miles} + 60 \\, \\text{miles} = 210 \\, \\text{miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210 miles**.", "content": "To find the total distance traveled by the train, we need to calculate the distance for each segment of the trip separately and then sum them up.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first segment.**\n* Speed ($v_1$) = 60 mph\n* Time ($t_1$) = 2.5 hours\n\n$$ d_1 = 60 \\, \\text{mph} \\times 2.5 \\, \\text{hours} = 150 \\, \\text{miles} $$\n\n**Step 2: Calculate the distance for the second segment.**\n* Speed ($v_2$) = 40 mph\n* Time ($t_2$) = 1.5 hours\n\n$$ d_2 = 40 \\, \\text{mph} \\times 1.5 \\, \\text{hours} = 60 \\, \\text{miles} $$\n\n**Step 3: Add the two distances together to find the total distance.**\n\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 \\, \\text{miles} + 60 \\, \\text{miles} = 210 \\, \\text{miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210 miles**.", "reasoning_content": "", "tool_calls": [], "request_md5": "755023dedf8bc8b2ed308c65ed47551a", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "640e0fd4f55996cb175a422f0a12cef5", "expected_substrings": [], "retrieval_missing": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_warning": null, "kv_backend_reason": null, "kv_mode": null }, { "request_id": "chatcmpl-231767-5", "ctx": 25, "cached": 0, "gen": 81, "finish": "stop", "think_words": 0, "ans_words": 66, "prefill_ms": 56.0, "prefill_tok_s": 446.4, "decode_tok_s": 47.8, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.079, "wall_s": 1.752, "attractor": false, "empty": false, "runaway": false, "ans_preview": "The seasons on Earth are primarily caused by the planet's axial tilt, which is approximate", "assistant_content": "The seasons on Earth are primarily caused by the planet's axial tilt, which is approximately 23.5 degrees relative to its orbital plane. As Earth orbits the Sun, this tilt causes different hemispheres to receive varying amounts of sunlight and different angles of insolation throughout the year. These changes in the intensity and duration of sunlight lead to the cyclical variation of temperature that defines each season.", "content": "The seasons on Earth are primarily caused by the planet's axial tilt, which is approximately 23.5 degrees relative to its orbital plane. As Earth orbits the Sun, this tilt causes different hemispheres to receive varying amounts of sunlight and different angles of insolation throughout the year. These changes in the intensity and duration of sunlight lead to the cyclical variation of temperature that defines each season.", "reasoning_content": "", "tool_calls": [], "request_md5": "956958f04974611b1491658313d2224c", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "8f66b4c97988825bd8e7840aaf44357e", "expected_substrings": [], "retrieval_missing": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_warning": null, "kv_backend_reason": null, "kv_mode": null }, { "request_id": "chatcmpl-231767-7", "ctx": 33, "cached": 0, "gen": 132, "finish": "stop", "think_words": 0, "ans_words": 103, "prefill_ms": 75.0, "prefill_tok_s": 439.9, "decode_tok_s": 47.7, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.098, "wall_s": 2.843, "attractor": false, "empty": false, "runaway": false, "ans_preview": "Elias wiped the salt spray from his goggles and scanned the jagged shoreline, expecting no", "assistant_content": "Elias wiped the salt spray from his goggles and scanned the jagged shoreline, expecting nothing more than the usual driftwood and broken shells. To his surprise, a brass astrolabe, tarnished but remarkably intact, lay nestled in a crevice between the basalt rocks, gleaming faintly under the moonlight. He descended the cliff path with trembling hands, his heart hammering against his ribs, driven by the sudden, intense need to understand how such a delicate instrument had survived the storm. As he brushed the sand from its intricate gears, he felt a profound, inexplicable connection to a time long before the lighthouse was ever built.", "content": "Elias wiped the salt spray from his goggles and scanned the jagged shoreline, expecting nothing more than the usual driftwood and broken shells. To his surprise, a brass astrolabe, tarnished but remarkably intact, lay nestled in a crevice between the basalt rocks, gleaming faintly under the moonlight. He descended the cliff path with trembling hands, his heart hammering against his ribs, driven by the sudden, intense need to understand how such a delicate instrument had survived the storm. As he brushed the sand from its intricate gears, he felt a profound, inexplicable connection to a time long before the lighthouse was ever built.", "reasoning_content": "", "tool_calls": [], "request_md5": "7f37e02b446c34bfa12810a92ce0e6f9", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "8fe0ad36f61bcf4992cc9df81cdf3817", "expected_substrings": [], "retrieval_missing": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_warning": null, "kv_backend_reason": null, "kv_mode": null }, { "request_id": "chatcmpl-231767-9", "ctx": 31, "cached": 0, "gen": 105, "finish": "stop", "think_words": 0, "ans_words": 77, "prefill_ms": 57.4, "prefill_tok_s": 540.5, "decode_tok_s": 47.8, "decode_estimated": false, "tau": null, "cycles": null, "dflash": null, "mtp": null, "mtp_ngram": null, "ngram_mod_windows": null, "ngram_mod_drafts": null, "ngram_mod_accepted": null, "ngram_mod_accept_rate": null, "mtp_windows": null, "ar_windows": null, "mtp_retired": null, "mtp_window_timings": null, "ttft_s": 0.081, "wall_s": 2.257, "attractor": false, "empty": false, "runaway": false, "ans_preview": "1. Prioritize simplicity and single responsibility to ensure each function does exactly on", "assistant_content": "1. Prioritize simplicity and single responsibility to ensure each function does exactly one thing.\n2. Use descriptive, self-explanatory names for variables, functions, and classes to reduce the need for comments.\n3. Keep functions short and focused, aiming for under ten to fifteen lines of logic where possible.\n4. Avoid over-optimizing; focus on clarity and consistency unless performance is a proven bottleneck.\n5. Write tests in parallel with your code to establish a safety net that encourages refactoring.", "content": "1. Prioritize simplicity and single responsibility to ensure each function does exactly one thing.\n2. Use descriptive, self-explanatory names for variables, functions, and classes to reduce the need for comments.\n3. Keep functions short and focused, aiming for under ten to fifteen lines of logic where possible.\n4. Avoid over-optimizing; focus on clarity and consistency unless performance is a proven bottleneck.\n5. Write tests in parallel with your code to establish a safety net that encourages refactoring.", "reasoning_content": "", "tool_calls": [], "request_md5": "a299a1f17034e10c71993553bba3a432", "atem_leak": false, "terminal_count": 1, "terminal_reasons": [ "stop" ], "post_terminal_bytes": 0, "saw_done": true, "stream_error": null, "prompt_md5": "8bed8e2d056dc1d47dccae9d32dbecf4", "expected_substrings": [], "retrieval_missing": [], "kv_backend": "vmm", "kv_backend_legacy": false, "kv_backend_warning": null, "kv_backend_reason": null, "kv_mode": null } ]Hardware validation request (optional)
{ "routes": [ {"mode": "battery", "tag": "qwen3.8:27b"} ], "claim": "loads the dense Qwen3.5/3.8-family text artifacts on the affected path; no regression on the load path" }Notes for the reviewer
hip-bridge) and the gatedHIPFIRE_LOAD_TRACEallocationinstrumentation that localized it. Happy to split them into two PRs if that is preferred — the diagnostic is the
second commit only, so dropping it leaves the fix intact.
GPU_PINNED_MIN_XFER_SIZEis an undocumented clr knob. It is absent from HIP's environment-variable reference(
rocm-systems:projects/hip/docs/data/env_variables_hip.rst) and AMD's env-var page;docs/env-vars.mdand thecode comment now record that provenance, that its default is unpublished, and that it can change without notice.
hipfire already depended on it (
f5731a506, for Qwen4 host-mapped experts). The non-workaround alternative would befor the loader to stage weights through its own pinned buffer instead of DMing from a pageable source — same effect,
no undocumented env var, but an owned copy and the loss of the mmap zero-copy path. Out of scope here.
scripts/serve_harness.pyruns the model in its own HOME(
~/.cache/serve_harness_home), so it does not see the operator'smemory.gpu_layer_budget. With the 15.66 GBqwen3.8-27b.mq4trunk that means a fully-resident load and the daemon refuses it —Qwen VMM load refused: projected weights, minimum prefill scratch and KV do not fit free VRAM— so the battery wasrun on the 11.78 GB
qwen3.8-27b.mq3-xtwith--max-seq 16384. Also:--max-tokensmust exceed the thinking cap(
--max-think-tokens 0here) or the harness fails pre-flight by design. And kill any previous harnessservebeforea rerun — it supervises its daemon, so the first attempts died with
FATAL: GPU … already reserved by holder PID …. That orphan behaviour is the pre-existing defect noted in the fix'sfollow-up, not this change.
logged
WARNING HIPFIRE_KV_BACKEND=legacy: automatic VMM selection unavailable (carrier lfm2moe has no VMM KV owner); using legacy KV storage— that is automatic, not an explicitHIPFIRE_KV_BACKEND=legacytoken, and thereason is the carrier, not this change (identical warning in both arms). No command in this PR set
kv_backend.The
serve_harnessbattery (qwen3.8:27b-mq3-xt) ran the q8 VMM path; itskv_backendfields are in the JSON below.