Skip to content

fix(atlas): EAN ggml_sqr contiguity + POST /atlas/reset?scope=ean|all (hydra_vortex#806) - #158

Merged
ddvnguyen merged 1 commit into
feat/flash-next-colibrifrom
fix/ean-contiguity-atlas-reset
Sep 27, 2026
Merged

ddvnguyen merged 1 commit into
feat/flash-next-colibrifrom
fix/ean-contiguity-atlas-reset

Conversation

@ddvnguyen

@ddvnguyen ddvnguyen commented Sep 27, 2026 •

Copy link
Copy Markdown
Owner

What

Closes ddvnguyen/hydra_vortex#806 (RULING 1 + RULING 3, docs/design-moe-demand-admission-phase15.md:1850-1868).

  1. T3 — EAN contiguity fix (src/llama-graph.cpp, inside the HYDRA_EAN_STATS gate): the per-slot reduction ran ggml_sqr on a strided ggml_view_2d → ggml_cuda_op_sqr contiguity assert aborted warmup. Now slot_c = ggml_is_contiguous(slot) ? slot : ggml_cont(ctx0, slot) + build-time GGML_ASSERT. Conditional single copy: decode graphs (n_tokens==1) add zero ops; env-unset graphs are byte-identical (OFF-parity structural).
  2. T4 — POST /atlas/reset?scope=ean|all (tools/server/server.cpp + server-atlas.{cpp,h}): new hydra_atlas::reset_scope(); scope=ean zeroes g_ean_gxn/g_ean_nsel in place under lock (sizes kept → /experts shows ean cells=0 until the next decode refills), scope=all delegates to reset(), 400 on unknown scope. Deliberately separate from the zero-caller per-probe reset().
  3. Docs: stale in-tree pinfile-parser citations in tools/expert-atlas/{README.md,export_pinfile.py} reworded (parser absent; format contract only).

Evidence (A/B smoke, RTX 5060 Ti, Ornith-1.5-35B-A3B Compact, --n-cpu-moe 12)

Arm Result
pre-fix binary, HYDRA_EAN_STATS=1 abort at warmup 6.7 s, unary.cu:144 GGML_ASSERT(ggml_is_contiguous(src0)) — exact ggml-org#806 defect
post-fix, HYDRA_EAN_STATS=1 HYDRA_EXPERT_META=1 listening in 39 s, /health 200, completion 200, /experts ean accumulates 1185 cells, POST scope=ean 200 (cells 0 → 702 refill next completion), scope=all 200, bad scope 400
post-fix, env unset (OFF) normal start, byte-identical greedy completion, ean key absent

Build: build-colibri-parity llama-server clean (CUDA 13.2.1, sm_120).

Refs ddvnguyen/hydra_vortex#806

… (hydra_vortex#806)

RULING 1 (docs/design-moe-demand-admission-phase15.md:1850-1854): the EAN
per-slot reduction ran ggml_sqr on a strided ggml_view_2d, tripping
ggml_cuda_op_sqr's contiguity assert at load warmup whenever
HYDRA_EAN_STATS=1 (reproduced pre-fix: unary.cu:144 abort in 6.7 s).
Materialize one contiguous copy (ggml_is_contiguous ? slot : ggml_cont)
plus a build-time GGML_ASSERT before the sqr; the whole change stays
inside the getenv gate, so env-unset graphs stay byte-identical
(OFF-parity structural). The conditional means decode graphs
(n_tokens==1, view passes ggml_is_contiguous) add zero ops.

RULING 3 (:1864-1868): new POST /atlas/reset?scope=ean|all backed by
hydra_atlas::reset_scope() — scope=ean zeroes g_ean_gxn/g_ean_nsel in
place under g_mtx (sizes kept, /experts shows ean cells=0 until the next
decode refills), scope=all delegates to the existing reset();
deliberately separate from the zero-caller per-probe reset(). 400 on an
unknown scope.

Also reword the stale in-tree pinfile-parser citations in
tools/expert-atlas/{README.md,export_pinfile.py} (hydra_cpu_init /
ggml-cuda.cu parsers are not present in this tree; format contract only).

Tests: A/B smoke on RTX 5060 Ti, model Ornith-1.5-35B-A3B Compact,
--n-cpu-moe 12: pre-fix binary aborts at warmup in 6.7 s (unary.cu:144,
the exact defect); post-fix with HYDRA_EAN_STATS=1 HYDRA_EXPERT_META=1
reaches listening in 39 s, /health 200, completion 200, /experts ean
accumulates 1185 cells, POST scope=ean 200 (cells 0 -> 702 refill after
the next completion), scope=all 200, bad scope 400. Env-unset run:
normal start, byte-identical greedy completion, ean key absent.
build-colibri-parity llama-server builds clean.

Refs ddvnguyen/hydra_vortex#806
Co-Authored-By: Oh My Pi (opencode-go/mimo-v2.6-flash)
@ddvnguyen
ddvnguyen merged commit 6387773 into feat/flash-next-colibri Sep 27, 2026
7 of 30 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant