Skip to content

F1-F3: NextN geometry on colibri line + F2 current-turn hits + F3 bit-length heat - #157

Merged
ddvnguyen merged 5 commits into
feat/flash-next-colibrifrom
feat/flash-next-colibri-f1f3
Sep 26, 2026
Merged

ddvnguyen merged 5 commits into
feat/flash-next-colibrifrom
feat/flash-next-colibri-f1f3

Conversation

@ddvnguyen

@ddvnguyen ddvnguyen commented Sep 26, 2026 •

Copy link
Copy Markdown
Owner

Architect handoff package (hydra_vortex track t-34d7bdf5d6, decision d-9981fa1092), delivered on the rebased feat/flash-next-colibri line (base 6b7dc6a).

Commits

  1. 380d7c9 — artifact-routes feature (0f766227d) rebased onto the colibri line. 0abcf9287 + 3840b2de3 are already patch-equivalent on the base (verified via git cherry); only the artifact-routes commit was missing.
  2. df421e4 — F1: NextN/MTP block classified as nextn_rows, not a trunk row (re-applied 6dbf575). Merged hunk verified semantically: vector sized block_count + nextn, trunk = block_count - nextn, moe_rows 0..trunk-1, nextn_rows from trunk. EAN mtp_excluded:true already present on this line.
  3. b2e496d — F2 + F3.

F2 — sticky hits fix

  • New hydra_atlas::begin_step() zeroes g_hits_step under lock; decode hook (server-context.cpp) calls it once per step BEFORE the per-row accumulate loop.
  • /experts now serves the CURRENT TURN window (g_turn_hits, folded per step in end_step, cleared at record_turn) — never the per-step bitmap. NO clear-on-read.
  • end_step also clears the step bitmap (belt-and-braces: a skipped begin_step can never resurrect stale bits).

F3 — heat encoding: bit-length (upstream parity)

  • encode_cell: verbatim port of upstream c/telemetry.h emap_emit: while(u){heat++;u>>=1;} if(heat>63) heat=63; (was linear-clamp-63).
  • Precondition verified before the change: sweep_analyze_790.py and colibri_gates.py never decode EMAP bytes (heat computed from raw probe counts); tools/expert-atlas/sweep.sh computes deltas over the monotone byte (encoding-agnostic; the saturates-at-63 note remains true).

Verification (local, CPU-only)

  • Configure: -DGGML_CUDA=OFF -DLLAMA_CURL=OFF, 20-core -j16.
  • test-hydra-atlas-file builds clean; 31/31 checks ok, exit 0 — including 10 new F2/F3 checks: window hits = current step bits; OR-accumulation within a turn; record_turn clears the served window (sticky-bug regression); per-row popcount <= k*n_out; heat = bit-length at counts 1/3/4/70 (emap_emit parity; pre-F3 linear would say 1,3,4,63).

Not in this PR

Correction (architect re-review d-7045376b02, fixed in e8f271b)

The F3 commit message's claim that sweep.sh deltas are "a monotone byte (encoding-agnostic)" was wrong: analyze.py sums and validate.py shares treat the recorded "selections" as linear counts, so EMAP-byte deltas under bit-length heat are not counts (hot experts saturate → delta 0 → vanish; cold dominate) and would invert the affinity ranking. Superseded: sweep.sh now reads the probe turn's exact per-turn routing counts from GET /turns/<seq> (source: "turns-routing", turn_seq provenance, MTP rows excluded) and never decodes EMAP bytes. Also includes N1: clean return false on negative trunk in load_impl. Verified: bash -n OK, test-hydra-atlas-file 31/31 exit 0.

ddvnguyen and others added 5 commits September 26, 2026 23:01
…ifact routes

Stage-C/D owner ruling: the fork ships the two atlas files; atlas-web proxies
GET {engine}/experts.json and propagates the upstream status verbatim
(server.ts:250-320), so a missing/refused artifact must surface as 404 —
never a misleading 200, never another model's atlas.

- server-atlas: atlas_file()/atlas_json() — resolution HYDRA_EXPERT_ATLAS
  (dir override) -> model file's own directory; engine-id refusal gate on
  provenance.engine_id (exact 16-hex FNV geometry id or the
  "$arch:$basename[:$size]" short form); 0/404/500 contract (404 missing or
  refused, 500 corrupt/unreadable). Geometry gains model_path/total_size;
  load_impl scratch-state reset on partial init keeps dense-model results
  fully cleared (byte-identity argued; single-shard MoE re-verified live).
- server.cpp: routes /experts.json + /v1/experts.json (observability tier)
  and /expert-ranks.json + /v1/expert-ranks.json (ranking tier twins).
- tests/test-hydra-atlas-file: synthetic 2-row MoE GGUF; 22 checks covering
  happy paths (both id spellings), 404 missing/wrong-id/no-provenance,
  500 corrupt JSON, tier independence, unknown kind.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
With NextN models (qwen35moe, nextn_predict_layers=1), llama-model counts
the nextn layer inside block_count: n_layer_all=41, trunk n_layer()=40.
The MTP tensors live in blk.40 (blk.40.nextn.* alongside blk.40.ffn_*_exps),
so the old geometry init (trunk = 0..block_count-1) classified blk.40 as a
trunk MoE row. The topk hook then reported it dead forever ("atlas: 1/41
MoE rows lack topk tensors (dead layers: 40)") and the EMAP row read as a
permanently unrouted VRAM tier row — misleading on the atlas UI.

Fix: trunk = block_count - nextn; blocks >= trunk with exps tensors are
classified as nextn_rows. Geometry now reports moe_rows 0..39 and
nextn_rows [40]; consumers (UI /turns labels, EMAP grid) can render the
MTP draft layer honestly.

Found via track t-34d7bdf5d6 (owner spotted layer 40 never routing).

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
…m parity)

F2 (architect package d-9981fa1092): the per-step hits bitmap never
zeroed between decode steps, so g_hits_step OR-accumulated lifetime bits
and /experts served a sticky cumulative hits picture. Add
hydra_atlas::begin_step() (zeroes g_hits_step under lock), call it from
the decode hook before the per-row accumulate loop (server-context),
serve /experts hits from the current TURN window (g_turn_hits, folded in
end_step, cleared at record_turn) with a size-guard, and clear the step
bitmap in end_step as belt-and-braces. NO clear-on-read semantics.

F3: EMAP heat encoding switches from linear-clamp-63 to bit-length,
verbatim port of upstream c/telemetry.h emap_emit
(`while(u){heat++;u>>=1;} if(heat>63) heat=63;`). Precondition checked:
sweep_analyze_790.py and colibri_gates.py never decode EMAP bytes
(heat from raw probe counts); sweep.sh deltas a monotone byte (encoding-
agnostic; saturates-at-63 comment still true).

Tests: test-hydra-atlas-file extended (+10 checks) — window hits = step
bits, OR within turn, record_turn clears served window (sticky-bug
regression), per-row popcount <= k*n_out, and heat = bit-length at
counts 1/3/4/70 (emap_emit parity). CPU build: test target compiles
clean; ./test-hydra-atlas-file: 31/31 ok, exit 0.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
…as) + N1 trunk guard

Architect re-review d-7045376b02 on PR #157:

B1 (blocking, script-only): sweep.sh computed probe selections as
after-before deltas on EMAP bytes (map & 63). With F3 bit-length heat
those deltas are not selection counts — hot experts saturate (127 stays
heat 7, delta 0, drops out) while cold experts dominate, inverting the
affinity/lift ranking downstream (analyze.py sums, validate.py shares
treat "selections" as linear counts). The sweep now reads the probe
turn's exact routing counts from the newest /turns/<seq> record (same
decode hook as the counters; {row, expert, count} mapped through
geometry.moe_rows/nextn_rows), records source:"turns-routing" +
turn_seq for provenance, and falls back to telemetry:false when the ring
did not advance. The before.json snapshot is kept only for its seq and
telemetry_enabled gate.

N1 (nit): load_impl returns false cleanly on negative trunk
(malformed-GGUF insurance; llama-model asserts n_layer_nextn <=
n_layer_all before the atlas loads, but the atlas parses independently).

Verification: bash -n sweep.sh OK; test-hydra-atlas-file rebuilt clean,
31/31 "all ok" exit 0. PR description corrected: the F3 commit's
"sweep.sh deltas a monotone byte (encoding-agnostic)" claim was WRONG
(consumers treat deltas as linear counts) — superseded by this fix.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
…d 2)

Architect round-2 review on PR #157 (e8f271b):

B2 (blocking): the stats-writer argv indexes were off by one — argv[3]
got the category, int(argv[4]) crashed on it ("invalid literal for
int(): 'math'"), and the output opened argv[2] (a file named after the
turn seq). set -uo pipefail without -e made every probe fail silently.
Now: category argv[4], idx int(argv[5]), output open(argv[3]).

B3 (blocking): BEFORE_SEQ read /experts .seq = g_seq (decode-step
counter), while TURN_SEQ is g_turn_seq (turn counter) — the comparison
was always false after any decoding, so every probe after the first
would record telemetry:false. Now the before-value comes from
GET /turns .seq, snapshotted BEFORE generation; /experts is polled only
for the telemetry_enabled gate.

Should-fixes: exact-one-turn gate (TURN_SEQ == BEFORE+1) — concurrent
client turns no longer silently counted as the probe's selections;
mismatch writes telemetry:false with reason (no-new-turn |
ring-advanced-by-N | telemetry-disabled | turns-unavailable). Row
filter corrected to row < len(moe_rows) (comment now matches code).

Hermetic smoke (architect-required, stub engine with /experts.seq ~900K
>> /turns.seq ~50 to exercise B3): 4 probes → 4 sidecars telemetry:true,
correct category/idx, source:turns-routing, linear counts, MTP row
excluded, turn_seq advancing; analyze.py exits 0 (3 experts, 2
categories). bash -n clean.

🤖 Generated with Codebuff
Co-Authored-By: Codebuff <noreply@codebuff.com>
@ddvnguyen
ddvnguyen merged commit 4176202 into feat/flash-next-colibri Sep 26, 2026
3 of 23 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant