Skip to content

hip-bridge: keep pageable weight copies out of reclaim on every load - #814

Open
aldrouil wants to merge 3 commits into
warpfront:masterfrom
aldrouil:fix/host-memory-reclaim-stall
Open

aldrouil wants to merge 3 commits into
warpfront:masterfrom
aldrouil:fix/host-memory-reclaim-stall

Conversation

@aldrouil

@aldrouil aldrouil commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

Summary

Fixes HFQ loads that stall forever at a varying loading layer N/64 under host-memory pressure, by arming clr's
pageable-copy switch (GPU_PINNED_MIN_XFER_SIZE=100000) in every Linux process instead of only in processes
host-mapping Qwen4 experts. Behaviour-neutral where it was already set; the HSA_USERPTR_FOR_PAGED_MEM arm is
unchanged from master.

The stall is not an allocation failure: at the hang the daemon's main thread burns one core (399 jiffies / 4 s, no
syscalls) while rocm-smi reports the GPU idle (14 % / 0 %) — HIP's signal wait spinning after KFD evicts the
process's queues because the kernel reclaimed KFD-userptr pages the GPU was reading from. That class
(ROCm/rocm-systems#12528) is already described in keep_host_memory_out_of_reclaim's comment, which records it for
both host-mapped experts and "weight uploads out of the mapped file"; 5df32fe89 (2026-09-30 16:16) narrowed the
mitigation to Qwen4 experts only, which left the pageable-copy class uncovered for every other model — and for the
dense partial-offload path that had landed 12 h earlier (e3066e739).

Branch: fix/host-memory-reclaim-stall (fork aldrouil/hipfire), based on master @ a89ed0a8e — the current
warpfront/hipfire HEAD, no rebase needed.

Which surface(s) does this touch?

  • kernel — crates/hip-bridge (the switch), crates/rdna-compute (load tracing)
  • load — runtime load path (weight_backend), arch load.rs
  • arch crate(s): hipfire-arch-qwen35 (load tracing only)
  • crates/hipfire-quantize / quant formats
  • control plane
  • policy files

Test plan

  • ./scripts/no-gpu-ci.sh passes — ran 2026-10-03: cargo check --workspace --examples, the per-crate lib tests,
    pytest (179 + 14 + 6 tests OK), env-docs: references covered, check-lifecycle clean
  • cargo build --release clean
  • cargo test --lib --workspace --locked passes — not green on this box, for two pre-existing environmental
    reasons, neither touched by this branch
    (both reproduced on mainline a89ed0a8e): (1)
    hipfire-loader::admission::tests::head_overlay::head_on_non_maple_refuses fails when this machine's
    ~/.hipfire/config.toml is visible and passes with a clean HOME; (2)
    peacemaker-ir::tests::opcode_examples_match_pinned_llvm_mc_for_every_declared_form fails for want of the pinned
    llvm-mc on this host. With a clean HOME the run is 2840 passed / 1 failed (the peacemaker-ir one). CI's job
    runs --locked on a clean runner and is not expected to hit either.
  • load / serve / kernel changes: python3 scripts/serve_harness.py --model ~/.hipfire/models/qwen3.8-27b.mq3-xt --tag qwen3.8:27b-mq3-xt --mode battery --kv q8 --max-seq 16384 --max-tokens 4096 --max-think-tokens 0 --thinking-effort none on hardware — 5/5 turns, runaway=0 empty=0 attractor=0 retrieval_miss=0, finish=stop on all five, avg decode 46.5 tok/s, kv_backend effective: vmm (log-marker), 42 s wall. JSON below.
  • I inspected loaded/bench/harness KV backend fields (kv_backend, kv_backend_legacy, kv_backend_reason); the
    LFM2 runs used an automatic legacy KV backend (carrier lfm2moe has no VMM KV owner) — token is automatic,
    not explicit, and it is disclosed in "Notes for the reviewer". No command set HIPFIRE_KV_BACKEND=legacy.
  • Perf: not perf-relevant in the decode sense; the change is confined to an env var read once at HIP init that
    selects clr's staging path for pageable copies. Measured load cost on this box (warm cache, arms alternated,
    first rep dropped, median in brackets):
    27B weight sweep unarmed 1550/1278/1151/2349 ms [1414] vs armed 1154/1139/1454/1239 ms [1197];
    9B unarmed 381/384/389/384 ms [384] vs armed 380/372/383/381/378 ms [379].
    The unarmed arm's own spread exceeds the gap, so this says "not slower", not "faster".
  • ratchet-raise — not applicable (no leanup-thresholds.txt ceiling moved)

What was verified on hardware

Both arms, one binary (pre pre-seeds GPU_PINNED_MIN_XFER_SIZE=0 = master's effective env for this class; the
gate never overrides an operator value):

model load pre/ship generation (256 tok)
swift-qwen3.8-27b.mq3-pro (partial offload, config gpu_layer_budget=55) yes / yes (under the 8 GiB ballast only the armed arm loads) —
qwen3.5:9b, qwen3.5:2b yes / yes "Paris"
lfm2.5:350m, lfm2.5:1.2b (LFM2 dense) yes / yes "Paris"
mimo-qwen (32 layers) yes / yes "Paris"

Stall reproduction/resolution, 8 GiB host-memory ballast held by another process:

arm partial offload fully resident
unarmed stall (layers 9, 45, 59 across runs) stall 5/5 (layers 49/35/31/33/42)
HSA only load 1/1 stall 4/4
GPU_PINNED only load 3/3 load 4/4

Artifacts used, per block:

evidence artifact sha256 / md5
serve_harness battery qwen3.8-27b.mq3-xt (qwen3.8:27b-mq3-xt), size 11777616896 3e04fc8db80bda557b965ec60ac876cf2500fced7f340624f3fcbeae134af5c5
stall matrix / both-arm load checks swift-qwen3.8-27b.mq3-pro, size 13602876416 1bcd36128f8bb58677d7444234a75ebad82ad8b110843ebcdf0e1262c9369744
cross-arch pulls lfm2.5:350m / :1.2b / :8b-a1b (sha256 verified by hipfire pull) 7cd38949…, a7aeba87…, 07da9942…
binaries (branch build) hipfire / daemon md5 b980aef35c023fe31b5a23ddef3a0012 / a633358a27de5dc003b515b318366d54

Neither pinned dense-trunk fixture was used — qwen3.8:27b-mq4-xts (AGENTS.md §5) and qwen3.8:27b-mq4-xt
(scripts/hw-gate/fixtures.json) are both not installed on this box — and no evidence above depends on trunk identity:
the claim is load-path behaviour, measured as a relative before/after on one artifact. The 15.66 GB qwen3.8-27b.mq4
on disk (sha256 5bb556a6…) was used only for the harness-fit diagnosis described in "Reproducing the battery".

muse-glimmer could not be covered: its registry SKUs are 16.26 GB (mq4r) / 18.61 GB (mq4) against ~13.2 GB free
on the 17.1 GB card, host-RAM offload is not wired for it, and a locally quantized kvist-glimmer-30b.mq4 (arch 14)
fails at load in both arms, before any copy — glimmer: unsupported quant_type 44 for model.language_model.layers.0.mlp.down_proj.weight — a glimmer/MQ4V2 load-support gap, not this change.

local serve_harness battery.json (load / serve / kernel changes)
[
 {
  "request_id": "chatcmpl-231767-1",
  "ctx": 44,
  "cached": 0,
  "gen": 205,
  "finish": "stop",
  "think_words": 0,
  "ans_words": 99,
  "prefill_ms": 4011.8,
  "prefill_tok_s": 11.0,
  "decode_tok_s": 41.7,
  "decode_estimated": false,
  "tau": null,
  "cycles": null,
  "dflash": null,
  "mtp": null,
  "mtp_ngram": null,
  "ngram_mod_windows": null,
  "ngram_mod_drafts": null,
  "ngram_mod_accepted": null,
  "ngram_mod_accept_rate": null,
  "mtp_windows": null,
  "ar_windows": null,
  "mtp_retired": null,
  "mtp_window_timings": null,
  "ttft_s": 4.649,
  "wall_s": 8.928,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "ans_preview": "```python\ndef merge_sorted(a, b):\n    \"\"\"Merge two sorted lists into a single sorted list.",
  "assistant_content": "```python\ndef merge_sorted(a, b):\n    \"\"\"Merge two sorted lists into a single sorted list.\n    \n    Uses a two-pointer approach to combine the lists in O(n + m) time.\n    Does not use sorted() or any built-in sorting function.\n    \"\"\"\n    merged = []\n    i = 0\n    j = 0\n    len_a = len(a)\n    len_b = len(b)\n\n    while i < len_a and j < len_b:\n        if a[i] <= b[j]:\n            merged.append(a[i])\n            i += 1\n        else:\n            merged.append(b[j])\n            j += 1\n\n    # Append any remaining elements from either list\n    while i < len_a:\n        merged.append(a[i])\n        i += 1\n    while j < len_b:\n        merged.append(b[j])\n        j += 1\n\n    return merged\n```",
  "content": "```python\ndef merge_sorted(a, b):\n    \"\"\"Merge two sorted lists into a single sorted list.\n    \n    Uses a two-pointer approach to combine the lists in O(n + m) time.\n    Does not use sorted() or any built-in sorting function.\n    \"\"\"\n    merged = []\n    i = 0\n    j = 0\n    len_a = len(a)\n    len_b = len(b)\n\n    while i < len_a and j < len_b:\n        if a[i] <= b[j]:\n            merged.append(a[i])\n            i += 1\n        else:\n            merged.append(b[j])\n            j += 1\n\n    # Append any remaining elements from either list\n    while i < len_a:\n        merged.append(a[i])\n        i += 1\n    while j < len_b:\n        merged.append(b[j])\n        j += 1\n\n    return merged\n```",
  "reasoning_content": "",
  "tool_calls": [],
  "request_md5": "e900b2d56ac907bb93c348f1e90e1da8",
  "atem_leak": false,
  "terminal_count": 1,
  "terminal_reasons": [
   "stop"
  ],
  "post_terminal_bytes": 0,
  "saw_done": true,
  "stream_error": null,
  "prompt_md5": "43ca0d15712d3dfb777b51ae76d8fd5f",
  "expected_substrings": [],
  "retrieval_missing": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_warning": null,
  "kv_backend_reason": null,
  "kv_mode": null
 },
 {
  "request_id": "chatcmpl-231767-3",
  "ctx": 55,
  "cached": 0,
  "gen": 332,
  "finish": "stop",
  "think_words": 0,
  "ans_words": 157,
  "prefill_ms": 101.0,
  "prefill_tok_s": 544.5,
  "decode_tok_s": 47.7,
  "decode_estimated": false,
  "tau": null,
  "cycles": null,
  "dflash": null,
  "mtp": null,
  "mtp_ngram": null,
  "ngram_mod_windows": null,
  "ngram_mod_drafts": null,
  "ngram_mod_accepted": null,
  "ngram_mod_accept_rate": null,
  "mtp_windows": null,
  "ar_windows": null,
  "mtp_retired": null,
  "mtp_window_timings": null,
  "ttft_s": 0.125,
  "wall_s": 7.059,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "ans_preview": "To find the total distance traveled by the train, we need to calculate the distance for ea",
  "assistant_content": "To find the total distance traveled by the train, we need to calculate the distance for each segment of the trip separately and then sum them up.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first segment.**\n*   Speed ($v_1$) = 60 mph\n*   Time ($t_1$) = 2.5 hours\n\n$$ d_1 = 60 \\, \\text{mph} \\times 2.5 \\, \\text{hours} = 150 \\, \\text{miles} $$\n\n**Step 2: Calculate the distance for the second segment.**\n*   Speed ($v_2$) = 40 mph\n*   Time ($t_2$) = 1.5 hours\n\n$$ d_2 = 40 \\, \\text{mph} \\times 1.5 \\, \\text{hours} = 60 \\, \\text{miles} $$\n\n**Step 3: Add the two distances together to find the total distance.**\n\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 \\, \\text{miles} + 60 \\, \\text{miles} = 210 \\, \\text{miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210 miles**.",
  "content": "To find the total distance traveled by the train, we need to calculate the distance for each segment of the trip separately and then sum them up.\n\nThe formula for distance is:\n$$ \\text{Distance} = \\text{Speed} \\times \\text{Time} $$\n\n**Step 1: Calculate the distance for the first segment.**\n*   Speed ($v_1$) = 60 mph\n*   Time ($t_1$) = 2.5 hours\n\n$$ d_1 = 60 \\, \\text{mph} \\times 2.5 \\, \\text{hours} = 150 \\, \\text{miles} $$\n\n**Step 2: Calculate the distance for the second segment.**\n*   Speed ($v_2$) = 40 mph\n*   Time ($t_2$) = 1.5 hours\n\n$$ d_2 = 40 \\, \\text{mph} \\times 1.5 \\, \\text{hours} = 60 \\, \\text{miles} $$\n\n**Step 3: Add the two distances together to find the total distance.**\n\n$$ \\text{Total Distance} = d_1 + d_2 $$\n$$ \\text{Total Distance} = 150 \\, \\text{miles} + 60 \\, \\text{miles} = 210 \\, \\text{miles} $$\n\n**Final Answer:**\nThe train traveled a total of **210 miles**.",
  "reasoning_content": "",
  "tool_calls": [],
  "request_md5": "755023dedf8bc8b2ed308c65ed47551a",
  "atem_leak": false,
  "terminal_count": 1,
  "terminal_reasons": [
   "stop"
  ],
  "post_terminal_bytes": 0,
  "saw_done": true,
  "stream_error": null,
  "prompt_md5": "640e0fd4f55996cb175a422f0a12cef5",
  "expected_substrings": [],
  "retrieval_missing": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_warning": null,
  "kv_backend_reason": null,
  "kv_mode": null
 },
 {
  "request_id": "chatcmpl-231767-5",
  "ctx": 25,
  "cached": 0,
  "gen": 81,
  "finish": "stop",
  "think_words": 0,
  "ans_words": 66,
  "prefill_ms": 56.0,
  "prefill_tok_s": 446.4,
  "decode_tok_s": 47.8,
  "decode_estimated": false,
  "tau": null,
  "cycles": null,
  "dflash": null,
  "mtp": null,
  "mtp_ngram": null,
  "ngram_mod_windows": null,
  "ngram_mod_drafts": null,
  "ngram_mod_accepted": null,
  "ngram_mod_accept_rate": null,
  "mtp_windows": null,
  "ar_windows": null,
  "mtp_retired": null,
  "mtp_window_timings": null,
  "ttft_s": 0.079,
  "wall_s": 1.752,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "ans_preview": "The seasons on Earth are primarily caused by the planet's axial tilt, which is approximate",
  "assistant_content": "The seasons on Earth are primarily caused by the planet's axial tilt, which is approximately 23.5 degrees relative to its orbital plane. As Earth orbits the Sun, this tilt causes different hemispheres to receive varying amounts of sunlight and different angles of insolation throughout the year. These changes in the intensity and duration of sunlight lead to the cyclical variation of temperature that defines each season.",
  "content": "The seasons on Earth are primarily caused by the planet's axial tilt, which is approximately 23.5 degrees relative to its orbital plane. As Earth orbits the Sun, this tilt causes different hemispheres to receive varying amounts of sunlight and different angles of insolation throughout the year. These changes in the intensity and duration of sunlight lead to the cyclical variation of temperature that defines each season.",
  "reasoning_content": "",
  "tool_calls": [],
  "request_md5": "956958f04974611b1491658313d2224c",
  "atem_leak": false,
  "terminal_count": 1,
  "terminal_reasons": [
   "stop"
  ],
  "post_terminal_bytes": 0,
  "saw_done": true,
  "stream_error": null,
  "prompt_md5": "8f66b4c97988825bd8e7840aaf44357e",
  "expected_substrings": [],
  "retrieval_missing": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_warning": null,
  "kv_backend_reason": null,
  "kv_mode": null
 },
 {
  "request_id": "chatcmpl-231767-7",
  "ctx": 33,
  "cached": 0,
  "gen": 132,
  "finish": "stop",
  "think_words": 0,
  "ans_words": 103,
  "prefill_ms": 75.0,
  "prefill_tok_s": 439.9,
  "decode_tok_s": 47.7,
  "decode_estimated": false,
  "tau": null,
  "cycles": null,
  "dflash": null,
  "mtp": null,
  "mtp_ngram": null,
  "ngram_mod_windows": null,
  "ngram_mod_drafts": null,
  "ngram_mod_accepted": null,
  "ngram_mod_accept_rate": null,
  "mtp_windows": null,
  "ar_windows": null,
  "mtp_retired": null,
  "mtp_window_timings": null,
  "ttft_s": 0.098,
  "wall_s": 2.843,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "ans_preview": "Elias wiped the salt spray from his goggles and scanned the jagged shoreline, expecting no",
  "assistant_content": "Elias wiped the salt spray from his goggles and scanned the jagged shoreline, expecting nothing more than the usual driftwood and broken shells. To his surprise, a brass astrolabe, tarnished but remarkably intact, lay nestled in a crevice between the basalt rocks, gleaming faintly under the moonlight. He descended the cliff path with trembling hands, his heart hammering against his ribs, driven by the sudden, intense need to understand how such a delicate instrument had survived the storm. As he brushed the sand from its intricate gears, he felt a profound, inexplicable connection to a time long before the lighthouse was ever built.",
  "content": "Elias wiped the salt spray from his goggles and scanned the jagged shoreline, expecting nothing more than the usual driftwood and broken shells. To his surprise, a brass astrolabe, tarnished but remarkably intact, lay nestled in a crevice between the basalt rocks, gleaming faintly under the moonlight. He descended the cliff path with trembling hands, his heart hammering against his ribs, driven by the sudden, intense need to understand how such a delicate instrument had survived the storm. As he brushed the sand from its intricate gears, he felt a profound, inexplicable connection to a time long before the lighthouse was ever built.",
  "reasoning_content": "",
  "tool_calls": [],
  "request_md5": "7f37e02b446c34bfa12810a92ce0e6f9",
  "atem_leak": false,
  "terminal_count": 1,
  "terminal_reasons": [
   "stop"
  ],
  "post_terminal_bytes": 0,
  "saw_done": true,
  "stream_error": null,
  "prompt_md5": "8fe0ad36f61bcf4992cc9df81cdf3817",
  "expected_substrings": [],
  "retrieval_missing": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_warning": null,
  "kv_backend_reason": null,
  "kv_mode": null
 },
 {
  "request_id": "chatcmpl-231767-9",
  "ctx": 31,
  "cached": 0,
  "gen": 105,
  "finish": "stop",
  "think_words": 0,
  "ans_words": 77,
  "prefill_ms": 57.4,
  "prefill_tok_s": 540.5,
  "decode_tok_s": 47.8,
  "decode_estimated": false,
  "tau": null,
  "cycles": null,
  "dflash": null,
  "mtp": null,
  "mtp_ngram": null,
  "ngram_mod_windows": null,
  "ngram_mod_drafts": null,
  "ngram_mod_accepted": null,
  "ngram_mod_accept_rate": null,
  "mtp_windows": null,
  "ar_windows": null,
  "mtp_retired": null,
  "mtp_window_timings": null,
  "ttft_s": 0.081,
  "wall_s": 2.257,
  "attractor": false,
  "empty": false,
  "runaway": false,
  "ans_preview": "1. Prioritize simplicity and single responsibility to ensure each function does exactly on",
  "assistant_content": "1. Prioritize simplicity and single responsibility to ensure each function does exactly one thing.\n2. Use descriptive, self-explanatory names for variables, functions, and classes to reduce the need for comments.\n3. Keep functions short and focused, aiming for under ten to fifteen lines of logic where possible.\n4. Avoid over-optimizing; focus on clarity and consistency unless performance is a proven bottleneck.\n5. Write tests in parallel with your code to establish a safety net that encourages refactoring.",
  "content": "1. Prioritize simplicity and single responsibility to ensure each function does exactly one thing.\n2. Use descriptive, self-explanatory names for variables, functions, and classes to reduce the need for comments.\n3. Keep functions short and focused, aiming for under ten to fifteen lines of logic where possible.\n4. Avoid over-optimizing; focus on clarity and consistency unless performance is a proven bottleneck.\n5. Write tests in parallel with your code to establish a safety net that encourages refactoring.",
  "reasoning_content": "",
  "tool_calls": [],
  "request_md5": "a299a1f17034e10c71993553bba3a432",
  "atem_leak": false,
  "terminal_count": 1,
  "terminal_reasons": [
   "stop"
  ],
  "post_terminal_bytes": 0,
  "saw_done": true,
  "stream_error": null,
  "prompt_md5": "8bed8e2d056dc1d47dccae9d32dbecf4",
  "expected_substrings": [],
  "retrieval_missing": [],
  "kv_backend": "vmm",
  "kv_backend_legacy": false,
  "kv_backend_warning": null,
  "kv_backend_reason": null,
  "kv_mode": null
 }
]

Hardware validation request (optional)

{
  "routes": [
    {"mode": "battery", "tag": "qwen3.8:27b"}
  ],
  "claim": "loads the dense Qwen3.5/3.8-family text artifacts on the affected path; no regression on the load path"
}

Notes for the reviewer

  • Two commits, one logical change each: the fix (hip-bridge) and the gated HIPFIRE_LOAD_TRACE allocation
    instrumentation that localized it. Happy to split them into two PRs if that is preferred — the diagnostic is the
    second commit only, so dropping it leaves the fix intact.
  • GPU_PINNED_MIN_XFER_SIZE is an undocumented clr knob. It is absent from HIP's environment-variable reference
    (rocm-systems:projects/hip/docs/data/env_variables_hip.rst) and AMD's env-var page; docs/env-vars.md and the
    code comment now record that provenance, that its default is unpublished, and that it can change without notice.
    hipfire already depended on it (f5731a506, for Qwen4 host-mapped experts). The non-workaround alternative would be
    for the loader to stage weights through its own pinned buffer instead of DMing from a pageable source — same effect,
    no undocumented env var, but an owned copy and the loss of the mmap zero-copy path. Out of scope here.
  • Reproducing the battery. scripts/serve_harness.py runs the model in its own HOME
    (~/.cache/serve_harness_home), so it does not see the operator's memory.gpu_layer_budget. With the 15.66 GB
    qwen3.8-27b.mq4 trunk that means a fully-resident load and the daemon refuses it —
    Qwen VMM load refused: projected weights, minimum prefill scratch and KV do not fit free VRAM — so the battery was
    run on the 11.78 GB qwen3.8-27b.mq3-xt with --max-seq 16384. Also: --max-tokens must exceed the thinking cap
    (--max-think-tokens 0 here) or the harness fails pre-flight by design. And kill any previous harness serve before
    a rerun — it supervises its daemon, so the first attempts died with
    FATAL: GPU … already reserved by holder PID …. That orphan behaviour is the pre-existing defect noted in the fix's
    follow-up, not this change.
  • KV backend disclosure. The Qwen runs above used the q8 VMM path. The three LFM2 runs (350m, 1.2b, 8b-a1b) each
    logged WARNING HIPFIRE_KV_BACKEND=legacy: automatic VMM selection unavailable (carrier lfm2moe has no VMM KV owner); using legacy KV storage — that is automatic, not an explicit HIPFIRE_KV_BACKEND=legacy token, and the
    reason is the carrier, not this change (identical warning in both arms). No command in this PR set kv_backend.
    The serve_harness battery (qwen3.8:27b-mq3-xt) ran the q8 VMM path; its kv_backend fields are in the JSON below.

Avery Drouillard added 2 commits October 3, 2026 11:34
27B loads stop at a varying `loading layer N/64` under host-memory pressure.
It is not an allocation failure: at the stall the daemon's main thread burns
100% of one core (399 jiffies / 4 s, no syscalls) while rocm-smi reports
GPU 14% / 0% — a lost KFD completion, not a long kernel.

Root cause. Under ROCm's defaults the host pages the GPU reads are KFD userptr
BOs, in two classes: (1) every hipHostMalloc block (Qwen4 routed experts,
partial-GPU-offload spilled layers, clr staging); (2) the source of any large
pageable copy — clr pins the mapped HFQ file's page-cache pages in place for
every multi-MB weight upload. Reclaiming either class evicts all of the
process's GPU queues until KFD's restore worker faults the pages back, and
under sustained pressure HIP's signal wait spins a core with no error
(ROCm/rocm-systems#12528; already described in keep_host_memory_out_of_reclaim's
comment, which records it for host-mapped experts AND for "weight uploads out
of the mapped file").

Class (2) is on every HFQ load, on every architecture. The switches were once
global (f5731a5, 2026-09-30 15:48) and were narrowed to host-maps_qwen4_experts()
by 5df32fe (16:16), on the premise that Qwen4 experts were the only
host-mapped class — which was already false for the dense partial-offload path
added that morning (e3066e7, 04:24), and left class (2) uncovered for every
other model too. The offload work did not introduce the bug; the narrowing did.

Fix. One switch, armed for every load:

- GPU_PINNED_MIN_XFER_SIZE=100000 in every Linux process (unless set). It
  routes each pageable copy through clr's pinned staging buffer, so the copy
  source is never registered as reclaimable userptr — class (2), closed
  generically for every model and every placement.
- HSA_USERPTR_FOR_PAGED_MEM=0 is unchanged from master: still scoped to
  host_maps_qwen4_experts(). It is the half that costs, and it moves
  host-mapped memory under TTM pages_limit; the ~20% in 5df32fe's message is
  this switch, not the one added here. No new tier is introduced.
- One `[hipfire] host-memory reclaim mitigation armed: ...` line names what
  was applied.

Neither switch name is a published interface, and the docs now say so:
GPU_PINNED_MIN_XFER_SIZE is absent from HIP's environment-variable reference
and from AMD's env-var reference (checked 2026-10-03, ROCm 7.2) and a verbatim
docs search returns nothing, so docs/env-vars.md and the code comment record it
as an internal clr knob, read once at runtime init, default unpublished, that
can change without notice. HSA_USERPTR_FOR_PAGED_MEM is a libhsakmt knob of the
same kind. f5731a5 relied on both first, for Qwen4.

The switch is set in HipRuntime::load(), before any model is chosen, so it
applies to every architecture and model by construction — there is no per-model
allowlist to keep in sync (the previous gate was one).

Cross-arch check on the branch binary. One binary covers both arms: `pre`
pre-seeds GPU_PINNED_MIN_XFER_SIZE=0, which is master's effective environment
for the copy-source class (the gate never overrides an operator value), and
that switch is the only behavioural delta. `ship` lets the gate arm it.

  model                        load pre/ship  generation (256 tok)  sweep unarmed -> armed (3 warm reps)
  lfm2.5:350m  (LFM2 dense)    yes/yes        "Paris"               loader prints no sweep line
  lfm2.5:1.2b  (LFM2 dense)    yes/yes        "Paris"               loader prints no sweep line
  lfm2.5:8b-a1b (LFM2 MoE)     yes/yes        no readable text      loader prints no sweep line
  qwen3.5:2b   (Qwen35 dense)  yes/yes        "Paris"               128-131 -> 131-133 ms
  qwen3.5:9b   (Qwen35 dense)  yes/yes        "Paris"               365-371 -> 364-373 ms
  mimo-qwen    (32 layers)     yes/yes        "Paris"               367-368 -> 364-387 ms
  qwen3.8-27b  (offload)       yes/yes        -                     see the cost block below
                                             (under 8 GiB ballast only the armed arm loads)

Nothing is slower: every per-model delta is inside that model's own rep spread.
The 8b-a1b emits tokens but no readable answer to this prompt.

muse-glimmer could not be covered. Its HF SKUs are 16.26 GB (mq4r) / 18.61 GB
(mq4), and a local muse-glimmer-30b.mq2-xt (13.43 GB) fails identically armed
and unarmed, before any weight copy:

  load failed: glimmer: unsupported quant_type 50 for
    model.language_model.layers.0.mlp.gate_proj.weight.

That is a glimmer/MQ2-XT support gap (qt=50), not this change; the card also
reports 13218 MB free against a 13431 MB file.

Cost, measured on gfx1201 (warm cache, arms alternated so drift cancels, the
first rep of each arm dropped as a cold-cache outlier; `weight sweep` as the
loader prints it; median of the remaining reps in brackets):

  27B   unarmed  1550 / 1278 / 1151 / 2349 ms [1414]   armed  1154 / 1139 / 1454 / 1239 ms [1197]
  9B    unarmed   381 /  384 /  389 /  384 ms [ 384]   armed   380 /  372 /  383 /  381 /  378 ms [379]

No measurable penalty. The unarmed arm's own spread (1151-2349 ms) is wider
than the gap between arms, so this says "not slower", not "faster": the staged
copy is not worse than the page-pinning it replaces.

Effectiveness, 8 GiB host-memory ballast, same model:

  arm              partial offload           fully resident
  unarmed          stall (layers 9, 45, 59)  stall 5/5 (layers 49/35/31/33/42)
  HSA only         load 1/1                  stall 4/4
  GPU_PINNED only  load 3/3                  load 4/4
  both             load 1/1                  -

HSA alone cannot cover the resident case (no hipHostMalloc weight block to
move); GPU_PINNED alone covers both and is free. The fully-resident 27B is
reachable on this 17.1 GB card at max_seq=4096 (a larger max_seq is refused by
admission before any load, which is why the resident arm needs the smaller
context to reproduce).

Follow-up, not fixed here: serve exits through process::exit on the second
SIGTERM and on pre-warm failure, skipping EngineInner::drop, so the daemon
child (with its VRAM and the GPU flock) is orphaned. Observed live — daemon
45609 outlived serve 45608 holding 10 GB — and it explains the operator's
`hipfire daemon already running` / GPU-lock FATALs. Separate change.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
…LOAD_TRACE

Extends the existing HIPFIRE_LOAD_TRACE knob (previously the packed-expert
read/upload timing in qwen35/load.rs) with the allocation-level view:

- rdna-compute/src/dispatch.rs: one `[load-trace]` line per upload_raw /
  upload_raw_host / alloc_host_mapped_tensor / free_tensor park /
  release_tensor_immediate, each carrying free VRAM and the pool counters
  (pool_new, pool_reused, pool_alloc_bytes, pool_parked_bytes);
- rdna-compute/src/pool.rs: GpuPool::freelist_bytes(), surfaced as
  Gpu::pool_freelist_bytes() — bytes parked in the free lists, which a failed
  load's rollback leaves unreachable to upload_raw until drain_pool;
- hipfire-runtime/src/weight_backend.rs: proj / norm / raw_f32 markers, so the
  last line before a stall names the tensor being read;
- hipfire-arch-qwen35/src/qwen35/load.rs: one per-layer summary with free VRAM
  and the pool counters.

Gated and a strict no-op unless the knob is set (read through
hipfire_config::developer_var_os, so it works as an env var or a [developer]
TOML key). Use it to localize, not to benchmark: every line calls hipMemGetInfo.

This is the diagnostic that localized the 27B load stall to a spin in the
weight-upload path; the fix it produced is the previous commit. It shares one
knob with the existing packed-expert line and uses the same `[load-trace]`
prefix.

docs/env-vars.md: HIPFIRE_LOAD_TRACE's source list gains the four new files.

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
@aldrouil
aldrouil requested a review from Kaden-Schutt as a code owner October 3, 2026 16:46
CI's `gates (ratchets, layering, registers)` job failed on "Crate maps match
the tree": hip-bridge, rdna-compute, hipfire-runtime and hipfire-arch-qwen35
all gained lines, and each crate's `map.md` is generated from the tree.

Regenerated with:

    python3 scripts/check-crate-maps.py hip-bridge hipfire-arch-qwen35 hipfire-runtime rdna-compute

Line-count deltas only; `--check` now reports "45 map(s) match the tree".

Signed-off-by: Avery Drouillard <avery@averynet.xyz>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant