Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
87 changes: 87 additions & 0 deletions analysis/ghidra/benchmarks/corpus/coldprobe.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,87 @@
#!/usr/bin/env bash
# coldprobe.sh -- #3023: did llm-worker contention move phase-2 scores, or only
# wall clock?
#
# Phase 2 ran without the cold-slot protocol's worker-stop step, so 36 of its 52
# models were measured while hp-llm-worker issued a competing /api/chat every
# ~4 minutes against the same OLLAMA_MAX_LOADED_MODELS=1 slot. Phase 2 shows 12
# escalations; phase 4's cold cells were +-0 across three runs on four
# architectures. That is suggestive, not proof -- so measure it instead of
# arguing about it.
#
# Two models, both ALREADY measured and both still local (no pull, no delete):
#
# DeepHat-V1-7B:Q4_K_M 4.7 GB, fully GPU-resident stored A=[63,63]
# -> the control. A resident model has no CPU-offload nondeterminism, so if
# its cold score differs from its contended one, contention moved scores
# and phase 2 needs re-running.
#
# ravenx-cyberagent-35b:Q4_K_M 21 GB, spills to CPU stored B=[64,63,64]
# -> the escalated cell. If cold gives three identical values the escalation
# was contention; if it still disagrees the nondeterminism is intrinsic to
# spilling (float reduction order across threads), which would exonerate
# phase 2 rather than condemn it.
#
# Results go to a SEPARATE directory so no roster row is polluted, and the
# stored phase-2 files are never touched.
#
# Preconditions this asserts rather than assumes: hp-llm-worker down,
# sweep_extra.sh not running, no record_baseline in flight.
set -u
REPO=/mnt-1/benchmarks/APIARY
OUT=/mnt-1/benchmarks/coldprobe
CACHE=/mnt-1/benchmarks/tierb-cache
mkdir -p "$OUT/logs"

die() { echo "ABORT: $*" >&2; exit 1; }

pgrep -f "sweep_extra.sh" >/dev/null && die "sweep_extra.sh is running -- stop it between models first"
pgrep -f "record_baseline.py" >/dev/null && die "a record_baseline run is in flight"
docker ps --format '{{.Names}}' | grep -qx hp-llm-worker && die "hp-llm-worker is up -- this probe measures its absence"
[ -d "$CACHE" ] || die "tierb-cache missing"
head=$(git -C "$REPO" rev-parse --short HEAD)
[ "$head" = "a99e765" ] || die "repo head is $head, not a99e765 -- wrong scoring vintage"

cd "$REPO" || die "no repo"

run() { # tier tag slug n
local tier="$1" tag="$2" slug="$3" n="$4"
local out="$OUT/cold_tier${tier}_${slug}_run${n}.json"
[ -f "$out" ] && { echo "skip $tier $slug run$n (present)"; return 0; }
local extra=""; [ "$tier" = "B" ] && extra="--ghidra-cache $CACHE"
docker exec ghidra-ollama-1 ollama stop "$tag" >/dev/null 2>&1 # cold every run
sleep 5
echo "$(date -u +%H:%M:%S) start $tier $slug run$n (load $(cut -d' ' -f1 /proc/loadavg))"
timeout 10800 python3 analysis/ghidra/benchmarks/corpus/record_baseline.py \
--tier "$tier" $extra --model "$tag" \
--operator coldprobe-3023 --provenance synthetic \
--output "$out" > "$OUT/logs/cold_tier${tier}_${slug}_run${n}.log" 2>&1
local rc=$?
if [ $rc -eq 0 ] && [ -f "$out" ]; then
echo "$(date -u +%H:%M:%S) done $tier $slug run$n score=$(python3 -c "import json;print(json.load(open('$out'))['total_score'])")"
else
echo "$(date -u +%H:%M:%S) FAIL $tier $slug run$n rc=$rc"
rm -f "$out"
fi
}

echo "=== $(date -u +%FT%TZ) COLDPROBE_START head=$head ==="

# control: fully resident, stored Tier A = [63, 63]
run A 'hf.co/mradermacher/DeepHat-V1-7B-GGUF:Q4_K_M' 'deephat_7b' 1
run A 'hf.co/mradermacher/DeepHat-V1-7B-GGUF:Q4_K_M' 'deephat_7b' 2

# the escalated cell: spills to CPU, stored Tier B = [64, 63, 64]
run B 'ravenx-cyberagent-35b:Q4_K_M' 'ravenx35b' 1
run B 'ravenx-cyberagent-35b:Q4_K_M' 'ravenx35b' 2
run B 'ravenx-cyberagent-35b:Q4_K_M' 'ravenx35b' 3

echo
echo "=== verdict inputs ==="
printf 'deephat_7b TierA contended=[63,63] cold=['
for n in 1 2; do f="$OUT/cold_tierA_deephat_7b_run${n}.json"; [ -f "$f" ] && printf '%s,' "$(python3 -c "import json;print(json.load(open('$f'))['total_score'])")"; done
printf ']\n'
printf 'ravenx35b TierB contended=[64,63,64] cold=['
for n in 1 2 3; do f="$OUT/cold_tierB_ravenx35b_run${n}.json"; [ -f "$f" ] && printf '%s,' "$(python3 -c "import json;print(json.load(open('$f'))['total_score'])")"; done
printf ']\n'
echo "=== $(date -u +%FT%TZ) COLDPROBE_COMPLETE ==="
69 changes: 69 additions & 0 deletions analysis/ghidra/benchmarks/corpus/keep_and_sample.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,69 @@
#!/usr/bin/env bash
# keep_and_sample.sh -- runs alongside sweep_extra.sh, changes nothing about it.
#
# Two jobs, both read-mostly:
#
# 1. RE-CREATE THE LOST SAMPLER (#2245). /mnt-1/benchmarks/vram_samples.tsv was
# the empirical record of served size and CPU/GPU split per model -- the input
# #2245's category-1 "which models actually spill" list was supposed to rest
# on. It was wiped with the rest of the work area (#2971) and is in no mirror,
# and the result JSONs record no VRAM at all, so it cannot be re-derived after
# the fact. Sample `ollama ps` while each model is loaded or it is lost again.
#
# 2. STOP LOSING THE REQUANT INPUTS. sweep_extra.sh:153 does `ollama rm "$TAG"`
# on every tag it pulled -- correct when /var was at 92%, wrong now that it
# has 6.5T free, and actively destructive to #2245: all ten sources named in
# requant_plan.txt have already been deleted this way. `ollama cp` makes a
# second manifest over the SAME blobs (verified: cp+rm of the copy left the
# blob count unchanged), so a keep-alias costs no disk and survives the
# sweep's rm of the roster tag.
#
# Deliberately does NOT edit sweep_extra.sh: bash reads a running script lazily
# by byte offset, and the keep-aliases use names no roster entry can match, so
# the sweep's own "already local -> will not delete" test at line 119 is
# unaffected either way.
#
# Stop with: pkill -f keep_and_sample.sh (leaves every alias in place)
set -u
OLLAMA=ghidra-ollama-1
SAMPLES=/mnt-1/benchmarks/vram_samples.tsv
KEPT=/mnt-1/benchmarks/kept-aliases.tsv
INTERVAL=60

oll() { docker exec "$OLLAMA" ollama "$@" 2>/dev/null; }

[ -f "$SAMPLES" ] || printf 'ts\tname\tsize\tprocessor\tcontext\n' > "$SAMPLES"
[ -f "$KEPT" ] || printf 'ts\troster_tag\tkeep_alias\n' > "$KEPT"

echo "$(date -u +%FT%TZ) KEEPSAMPLE_START interval=${INTERVAL}s"

while true; do
ts=$(date -u +%FT%TZ)

# --- 1. sample whatever is resident right now -------------------------------
# `ollama ps` columns, whitespace-separated:
# $1 NAME $2 ID $3+$4 SIZE ("57 GB") $5+$6 PROCESSOR ("65%/35% CPU/GPU")
# $7 CONTEXT $8.. UNTIL
# SIZE is the *served* footprint (weights + KV at the served context), which
# is the number #2245 needs -- not the on-disk GGUF size. PROCESSOR is the
# spill: anything not "100% GPU" is a category-1 requantization candidate.
oll ps | tail -n +2 | awk -v ts="$ts" 'NF>=7 {
printf "%s\t%s\t%s %s\t%s %s\t%s\n", ts, $1, $3, $4, $5, $6, $7
}' >> "$SAMPLES"

# --- 2. keep-alias anything new the sweep pulled ----------------------------
oll list | tail -n +2 | awk '{print $1}' | while IFS= read -r tag; do
[ -z "$tag" ] && continue
case "$tag" in keep/*) continue;; esac
slug=$(printf '%s' "$tag" | tr ':/' '__' | tr '[:upper:]' '[:lower:]')
alias="keep/${slug}:src"
if ! oll list | awk '{print $1}' | grep -qixF "$alias"; then
if oll cp "$tag" "$alias" >/dev/null 2>&1; then
printf '%s\t%s\t%s\n' "$ts" "$tag" "$alias" >> "$KEPT"
echo "$ts KEPT $tag -> $alias"
fi
fi
done

sleep "$INTERVAL"
done
54 changes: 54 additions & 0 deletions analysis/ghidra/benchmarks/corpus/probe_at_gap.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,54 @@
#!/usr/bin/env bash
# probe_at_gap.sh -- run coldprobe.sh in the next natural gap of the phase-2
# sweep, then put the sweep back.
#
# The gap matters: the probe must measure a genuinely uncontended slot, so it
# cannot run alongside sweep_extra.sh. But sweep_extra must not be killed
# mid-run either -- do_run's own `rm -f "$out"` cleanup is skipped when the
# parent dies, which is how the 2026-08-31 abort left a valid-looking partial
# result carrying 11 of 14 cases (see STATE-2026-08-31-cpu-swap.md).
#
# So: wait for the next MODEL_DONE, stop the sweep's driver only, let any
# in-flight run finish on its own, probe, then relaunch. sweep_extra skips
# every model that already has both tier run1 files, so the relaunch resumes
# exactly where it stopped.
set -u
BASE=/mnt-1/benchmarks
LOG=$BASE/probe_at_gap.log

say() { echo "$(date -u +%FT%TZ) $*" | tee -a "$LOG"; }

baseline=$(grep -c MODEL_DONE "$BASE/extra.log")
say "ARMED baseline_model_done=$baseline"

# 1. wait for the sweep to finish its current model (up to 6h)
for _ in $(seq 1 720); do
[ "$(grep -c MODEL_DONE "$BASE/extra.log")" -gt "$baseline" ] && break
sleep 30
done
[ "$(grep -c MODEL_DONE "$BASE/extra.log")" -gt "$baseline" ] || { say "TIMEOUT waiting for MODEL_DONE -- doing nothing"; exit 1; }
say "GAP model finished: $(grep MODEL_DONE "$BASE/extra.log" | tail -1)"

# 2. stop the driver only, never an in-flight record_baseline
pkill -f "$BASE/sweep_extra.sh" 2>/dev/null
pkill -f "bash sweep_extra.sh" 2>/dev/null
sleep 3
say "driver stopped (sweep_extra procs now: $(pgrep -cf sweep_extra.sh))"

# 3. let any run that was already started finish by itself (up to 3h)
for _ in $(seq 1 360); do
pgrep -f record_baseline.py >/dev/null || break
sleep 30
done
pgrep -f record_baseline.py >/dev/null && { say "ABORT: a record_baseline is still in flight after 3h"; exit 1; }
say "slot is idle -- starting cold probe"

# 4. probe
bash "$BASE/coldprobe.sh" 2>&1 | tee -a "$LOG"
say "COLDPROBE finished rc=$?"

# 5. put the sweep back
cd "$BASE" || exit 1
setsid nohup bash "$BASE/sweep_extra.sh" >> "$BASE/extra.log" 2>&1 < /dev/null &
sleep 5
say "sweep relaunched (procs: $(pgrep -cf sweep_extra.sh))"
35 changes: 35 additions & 0 deletions docs/CI-CD.md
Original file line number Diff line number Diff line change
Expand Up @@ -809,6 +809,41 @@ runner, that cache survives between job runs on this same machine, so the
second and every later run skips the download entirely. This is most of
where the actual speed win comes from, not raw CPU.

#### How many instances, and why it is 2 during a benchmark window

The box ran **seven** `honeypot-ci` instances (`supermicro`,
`supermicro-ci-2` .. `-7`) plus the separate `honeypot-home` deploy runner.
Reduced to **two** (`supermicro` + `supermicro-ci-2`) on 2026-09-05.

This is a shared box, not a CI box. The homeserver is also the GPU host for
the #1947 model benchmark, and CI is that benchmark's largest source of
measurement noise: identical work on the same model spread **62%** (78.2 vs
126.7 min/run) purely from runner contention, and load average sat at
**15-27** with seven instances live. #1947's whole premise is that scores
separated by less than the error bar are not results, so a benchmark run
taken against an unknown, CI-driven load is not a measurement -- see
`docs/benchmarks/plans/2026-09-05-1947-resume-plan.md` step P5, which cannot
be answered at all without this.

Seven was never a measured optimum; it was the ceiling #2572 allowed once
one instance stopped serialising the matrix. The CPU swap (Xeon Silver 4110
8c/16t -> Xeon Gold 5220R 24c/48t) also means two instances today have more
cores behind them than four did before.

To restore the extra instances once the benchmark window closes:

```bash
ssh homeserver 'for n in 3 4 5 6 7; do
sudo systemctl enable --now actions.runner.Xore-APIARY.supermicro-ci-$n.service
done'
```

They are only `systemctl disable --now`, not deregistered -- the registration,
`_work` dir and tool cache all survive, so re-enabling costs nothing and needs
no token. **Stopping a busy instance fails its in-flight job**, so cancel the
run first or wait for idle; the 2026-09-05 reduction cancelled a run and cost
16 jobs a re-run.

### Buildx layer cache: `type=local` on the homeserver (#2822)

`containers.yml` used to export every image's layer cache to `type=gha`,
Expand Down
Loading
Loading