Skip to content

ax on Agent Substrate across the fleet (nas control plane, coordinator gVisor harness pool) - #464

Closed
mecattaf wants to merge 32 commits into
chore/llm-agents-bump-drop-harvest-hookfrom
ax/fleet-bringup
Closed

mecattaf wants to merge 32 commits into
chore/llm-agents-bump-drop-harvest-hookfrom
ax/fleet-bringup

Conversation

@mecattaf

Copy link
Copy Markdown
Owner

Closes #463. Supersedes #446, #447 and #454 (left open for you to close).

Draft: round 4 of review found one high-severity item that is not fixed (below). Merge and switch are yours.

What it declares

  • nas (myAxFleet.role = control): k3s server, Agent Substrate (server at upstream d277088b; ax keeps its vendored client pin), Postgres, RustFS, registry, Redis, ax-server and ax-controller. Bulk state lives on the data pool, not the eMMC root.
  • coordinator (harness): k3s agent, gVisor worker pool (2 workers), ax client and proxy, and a deny-by-default LAN guard. The CPU weight is set so the cluster does not crowd the desk.
  • worker (inference): no cluster code. Sandboxes reach Halogen :8731 through the NAS egress gateway.
  • Kill switch: one per host.
  • ax: v0.3.0 carrying sandbox-class.patch and p1-completion.patch.

Evidence (all at this head, f1bba47e)

  • nix flake check --no-build: all checks pass.
  • All three toplevels build.
  • checks.ax-fleet-topology passes.
  • checks.x86_64-linux.ax-fleet: the 4-VM test (nas, coordinator, worker, tailnet peer) passes 50 subtests. nix eval of its outPath at this head is /nix/store/6d586q2qzm4y3fm97ci467g4q9q42379-vm-test-run-ax-fleet, the green round-3 run.
  • In that test:
    • An ax Task ran under gVisor on the coordinator, reached a Halogen stand-in, and stayed Completed.
    • Four Tasks ran on the 2-worker pool.
    • A LAN flap and both rollbacks passed.
  • Three fix rounds are logged in ~/today/evals-2026-09-23/ax-fleet/REVIEW-LOG.md.

Switch runbook

Base. Merge this onto chore/llm-agents-bump-drop-harvest-hook (#461), which carries the live coordinator revision 3a658991. Basing it on main would roll back #461 and #460.

Order. nas, then worker, then coordinator.

Verify after the switch:

  1. kubectl get nodes shows nas and coordinator Ready.
  2. Substrate is healthy.
  3. ax-fleet-smoke probe: the Task shows a gVisor /proc/version and a Halogen reply.
  4. ax-fleet-smoke floor: 4 Tasks with no ResourceExhausted.

Note. The ax API proxy now listens on :8099, not :8080. Use :8099 in any manual check (see open items).

Rollback. Switch back to the previous generation, then run ax-fleet-teardown. There is also a per-host kill switch.

Open items found in review round 4 (not fixed here)

  • high: ax egress fails open. A Task with no gateway, or naming a gateway its atespace lacks, gets allow-all egress. Nothing on the fleet enforces the day-one gateway. Round 4 drafted a fix as a new ax patch, but you prefer not to patch ax, so that work sits on the unreviewed branch ax/fleet-r4-wip (e82cb885, not pushed). Proposed default: refuse dispatch when the gateway is missing, in the link or CONWIP layer, rather than in ax.
  • medium:
    • On the NAS, local services can reach the unauthenticated ax-server API and the password-less Redis.
    • The kubelet --eviction-hard replaces k3s's disk-pressure defaults.
    • Rollback teardown drops the coordinator's flannel.1 guard and pod-input rules.
    • The runbook's 127.0.0.1:8080 check should be :8099.
    • The VM base generation already carries the guards, so the first-time switch is not fully exercised.

Planned follow-ups (decided after this branch was built)

  • Drop sandbox-class.patch: stock ax already writes gVisor for every Task. MEASURED: 4-VM pass on branch probe/ax-fleet-nosc.
  • Drop p1-completion.patch: the job reports to the floor, then the link deletes the Task (the Buildkite pattern).
    • One green 4-VM run on stock ax with zero patches (probe/ax-fleet-zeropatch).
    • One delete hung in 4 runs. The link needs a re-delete and escalate step first; see ~/today/evals-2026-09-23/zero-patch/DELETE-HANG-2026-09-23.md.
  • Credentials: mount credentials into sandboxes by path (Sandbox runtimes: mount Claude and cloud credentials from the coordinator by path (full profile) #462).

Unknowns and proposed defaults

  • Behaviour on real hardware versus the VM test (wifi on the coordinator, eMMC on the NAS). Proposed default: switch the NAS first and watch it for a day.
  • Round-4 medium items. Proposed default: fix them in a follow-up PR, before running unattended jobs.

🤖 Generated with Claude Code

mecattaf and others added 30 commits September 17, 2026 06:25
`systemctl status | head -5` SIGPIPEd systemctl once head had its lines,
and under writeShellApplication's pipefail every successful switch exited
141 (seen on the worker 2026-09-17). Print the unit's state with
`systemctl show` instead. Verified on the coordinator: qwen38-27b and off
both exit 0.

Refs #397

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
halogen-switch: exit 0 on a successful switch
Tom, 2026-09-20: "having local kubernetes s3 or redis or postgres or whatever
it needs on the NAS". This is that sentence for the three Substrate actually
reads. State on the appliance, machines on the Strix boxes -- the NAS has 8
threads and 22 GiB and runs the house's DNS, so it holds the control plane's
state and schedules nothing (Appendix J section 5).

A SECOND DATABASE, not a second PostgreSQL. hosts/nas/media.nix already
brought the instance up and pinned its dataDir to the NVMe; this rides it, so
there is one instance and one backup story. ensureDatabases/ensureUsers with
ensureDBOwnership, because ateapi runs goose migrations at startup and creates
its own tables. enableTCPIP is forced by the deployment and not by taste:
ateapi is a pod on a Strix box, the k3s server here sets disableAgent, so
there is no unix socket to peer-authenticate over. scram-sha-256 from the LAN
and from the pod CIDR, never trust.

THE OBJECT STORE IS A HAND-WRITTEN UNIT ON PURPOSE. This host rides
nixpkgs-stable (nixos-26.05) and stable has no rustfs at all -- no module, no
package. The main pin has both; the package comes across the
hosts/nas/unstable-pkgs.nix seam that attic-server and Immich already use, and
the unit is modelled line for line on the unstable module's own serviceConfig.
Delete it for services.rustfs when the NAS next rides a stable that ships it.

Every unit carries RequiresMountsFor on /mnt/fast. hosts/nas/attic.nix learned
the second edition of the signing-key trap: /mnt/fast is nofail, so a failed
mount otherwise yields an empty state tree and a service that starts happily
against it. For an object store that is silently lost snapshots. Refusing to
start is a Tuesday.

Secrets are runbook-placed root-owned files, not agenix, per attic.nix's
no-agenix doctrine -- the rustfs key pair and the Substrate role's password
are consumed by this host alone and a runbook can place them once. The
password oneshot reads its file through LoadCredential and binds it as a psql
value, so it lands in no argv and no journal line.

Firewall: one block, three ports, iifname "enp1s0" only. Not openFirewall,
which would publish all three on the tailnet as well. All three are on
hosts/nas/cloudflared.nix's never-routed list.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Tom, 2026-09-20: "kubernetes is world-class for that ... they each stay in
their lane. Effects ts handles the ultracode-level json dag specification and
kubernetes schedules it on the right machines."

One shared modules/k3s-fleet.nix imported by all three hosts, because the
numbers have to agree and a second copy of them is how they stop agreeing.
Server on the NAS with disableAgent, so the appliance holds the control plane
and schedules nothing; agents on the Strix boxes, which have the cores, the
memory and /dev/kvm.

THE POD CIDR MOVES, AND THE ASSERTION IS OUTSIDE THE GATE. k3s defaults to
10.42.0.0/16 for pods and the house LAN is 10.42.0.0/24, which contains the
NAS, the coordinator, the worker, the printer and the router. Tom's 2026-09-09
research already picked the replacement: pods 10.200.0.0/16, services
10.201.0.0/16. Those are kept exactly, and the overlap check is real
arithmetic in a real assertion placed OUTSIDE the enable gate, so it
evaluates on every host whether k3s is on or off. Verified against six cases,
including that k3s's own default IS caught.

TRACK K AND TRACK S ARE THE SAME FILE, and that is the point. Design K is
plain k3s plus Cilium plus the gvisor RuntimeClass; Design S adds Substrate on
top of exactly this. The only thing S needs from the bottom layer that K does
not is four feature-gate settings, and a gate nothing asks for is inert: no
scheduling change, no new controller, no memory. So the gates are in
unconditionally and there is one PR instead of two that drift.

The gate names are upstream's own, MEASURED from ~/Downloads/substrate,
hack/create-kind-cluster.sh:104-112: ClusterTrustBundle,
ClusterTrustBundleProjection, PodCertificateRequest, plus runtimeConfig
certificates.k8s.io/v1beta1. The kubelet is handed only the two that are its
own, because an unrecognised gate name is fatal to kubelet and
ClusterTrustBundle is apiserver-side. Whether k3s 1.35.6+k3s1 accepts them at
all is U9 and is not claimed here.

Cilium 1.18.14 through autoDeployCharts, so the CNI lives in the generation
rather than in a `cilium install` somebody has to remember. The chart hash is
measured, not fakeHash: built with fakeHash, took the reported hash, rebuilt
clean. The gvisor RuntimeClass is a server manifest and its containerd runtime
is an agent containerdConfigTemplate that keeps `{{ template "base" . }}` --
one line, load-bearing, and dropping it costs the node its CNI, its
snapshotter and its registry mirrors at once. kata is written out and
commented, pointing at tonight's U14 report.

The kube API is LAN-only: 6443 on the NAS's enp1s0 and nothing else, plus the
never-routed doctrine in the cloudflared PR, because a tunnel ingress bypasses
nftables entirely and "firewalled" is therefore a separate guarantee from
"not in the tunnel".

All three hosts land with the gate OFF. secrets/k3s-token.age does not exist;
minting it needs Tom's admin age key.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A5a stood the pinned k3s up inside a throwaway guest built from this exact
nixpkgs revision the same night and measured three things this file had
wrong. Report: ~/today/review/2026-09-20/sandbox-spike/U9-U13-U15-K3S.md.

1. THE THREE GATES ALONE SERVE NOTHING. Boot 5 held the gates on and dropped
   the runtime-config flag: all three kubernetes_feature_enabled metrics read
   1 while /apis/certificates.k8s.io listed only v1 and api-resources returned
   NO_RESOURCES. A check that read only the metric would have reported a false
   pass. --kube-apiserver-arg=runtime-config=certificates.k8s.io/v1beta1=true
   is what turns the group version on, and podcertcontroller has nothing to
   talk to without it. It now sits beside the gates, and the flag values
   inside a -arg= carry no leading dashes of their own, which is A5a's
   measured working form.

   U9 is therefore ANSWERED YES on the pin: v1beta1 serves
   clustertrustbundles and podcertificaterequests. Appendix J section 9's
   fallback does not have to be taken and no newer k3s is needed. kubelet also
   accepts all three gate names, so the cautious subset this file used to hand
   it is gone.

2. THE NIXPKGS EXAMPLE'S CONTAINERD KEY PATH IS WRONG FOR THIS k3s. The
   bundled containerd is 2.2.5-k3s2 and the config it generates starts
   `version = 3` with runtimes under
   [plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.runc]. A
   grpc.v1.cri block would be parsed, accepted and silently ignored, which is
   the worst of the three outcomes. A5a's exact working table is used here.

   And runtime_path, NOT options.BinaryName. The BinaryName shape this file
   used is a trap that looks right: a pod under it reaches Running and stays
   1/1 Running in kubelet's view, produces no logs at all, and kubectl exec
   fails with "in state stopped". The generic runc shim starts runsc but
   carries neither its stdio nor its state. containerd-shim-runsc-v1 is what
   goes here, and nixpkgs' gvisor builds it.

3. k3s DOES NOT FIND runsc ON ITS OWN. Not by auto-detection (with gvisor in
   systemPackages the generated config still carried only runc and
   runhcs-wcow-process) and not from the unit PATH, which the nixpkgs rancher
   module leaves empty but for zfs. Without the path line every sandbox dies
   at creation with `exec: "runsc": executable file not found in $PATH`. The
   line was already here; it now carries why, and A5a's caution that the
   assignment replaces rather than extends.

Also corrected: this branch's earlier claim that identical gate-off
derivation paths proved the module inert. They do not prove it. The flake
embeds the configuration revision, so the toplevel tracks the commit. The
sound check is that with the gate off, services.k3s.enable is false, no 6443
rule is emitted and no registries.yaml exists on any of the three hosts,
which was measured directly.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pkgs/ax builds google/ax v0.3.0 pinned by commit d8ed0fe (tag v0.3.0),
wired through overlays/default.nix and the flake's packages.<system>
list like every sibling package. modules/ax-client.nix puts kubectl and
ax into environment.systemPackages on coordinator, worker and client
when myAxClient.enable is true; it lands false on all three. Refs #453.

The one judgement call: ax's go.mod opens `go 1.27.1` and Go refuses to
build a module whose go directive is newer than the running toolchain
(`go: go.mod requires go >= 1.27.1 (running go 1.27.0; GOTOOLCHAIN=local)`),
with no network in the sandbox to fetch one. MEASURED 2026-09-23: nixpkgs
has go_1_27 = 1.27rc2 and nixpkgs-fresh has 1.27.0, so neither pin already
in this flake can build it. Adds one input, nixpkgs-go, pinned BY REVISION
and supplying exactly one attribute to exactly one package - the same shape
as the existing nixpkgs-paperless. Bumping nixpkgs-fresh instead was
rejected on purpose: that input also carries the twins' linux 7.2 kernel.

checks.ax-client-topology asserts the option exists on the three hosts
that import the module, that it is false on all three, that kubectl and
ax are absent from all three systemPackages, and that hosts/nas carries
no such option at all. It goes red on the flip by design.

No switch, no rebuild, no gate flipped, no cluster, no Substrate, no
patches - this builds pristine v0.3.0.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds `string sandbox_class = 11` to TaskSpec, regenerates ax.pb.go with the
same protoc-gen-go v1.36.11 upstream used, refuses unknown values in
ValidateTask, and threads the value from the reconciler through
BuildActorTemplate, replacing the SandboxClass_SANDBOX_CLASS_GVISOR hardcode
at internal/substrate/client.go:273 (the only SANDBOX_CLASS occurrence in the
tree). Empty means gvisor, so existing manifests behave exactly as before.

The generated Go is in the patch on purpose: ax bridges YAML through protojson
with unknown fields rejected, so a .proto-only edit would make every manifest
naming sandboxClass fail strict decode.

This does NOT deliver a workerd sandbox. Agent Substrate's SandboxClass enum
has three members (UNSPECIFIED, GVISOR, MICROVM), so "gvisor" and "microvm"
are the only values that can reach a real substrate. A workerd class needs an
upstream Agent Substrate change that does not exist.

vendorHash is unchanged: the patch touches no go.mod or go.sum line.

Measured on the patched scratch tree with go 1.27.1: `go build ./...` rc 0,
`go test ./...` rc 0. In this clone: `nix build --no-link .#ax` rc 0, with the
patch applying to all six files and checkPhase passing.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
home/ax-conwip.nix declares `myAxConwip`, a home-manager module that WOULD run
the CONWIP scheduler at /home/tom/mecattaf/ax-conwip as one long-running user
service against a declared ax server, with the rewrite's seat meter directory as
a read-only input. It is imported from home/home.nix and gated OFF, so it defines
no unit on any host today.

Seven options, every one of them data in this repository rather than a command
line somebody types: enable (false), serverUrl (127.0.0.1:8080, the loopback the
mock stack listens on, so an accidental enable reaches nothing real), metersDir
(the REWRITE's meters, ~/.local/state/tally-rewrite/meters, never branch (a)'s,
read only, no tmpfiles rule over it), wipCap (1, stricter than the program's own
default of 2), sourceDir, recordsDir (a path that does not exist, so an
accidental enable fails legibly instead of passing over an empty directory) and
stateDir.

Four things deliberately not done, and the module header says all four out loud:

  - no flake input, because the CONWIP repository has NO REMOTE. A fetchGit of a
    path on one box is not a declaration, it is a machine-local accident that
    would break every other host's eval;
  - no package, because packaging follows the input. The unit runs the checkout
    in place through `pnpm exec tsx`. That is a development shape, not a
    delivered one, and it is the biggest single reason the gate stays off;
  - no timer: the release signal is a WatchTask stream, not a poll;
  - nothing enabled and nothing armed. `enable` is set nowhere, and even with
    the gate flipped the unit carries NO Install section, so a rebuild declares
    it and a `systemctl --user start` is a second, separate act.

checks.x86_64-linux.ax-conwip-topology asserts the option EXISTS on all three
home-manager hosts (a dropped import makes it undefined, not false, which is the
mutation hint), that it is false on all three, that no ax-conwip service, timer
or socket is rendered anywhere, that there is no system-bus twin on any of the
four hosts, and that no tmpfiles rule naming ax-conwip exists while the gate is
off. The `enable == false` line goes red on the flip, deliberately.

MEASURED in this tree:
  nix build --no-link .#checks.x86_64-linux.ax-conwip-topology -> rc 0, so every
    assertion above ran and passed
  nix eval .#nixosConfigurations.{coordinator,worker,client}.config
    .home-manager.users.tom.systemd.user.services --apply builtins.attrNames
    -> rc 0, no ax-conwip on any of the three
  the same eval under extendModules with the gate flipped renders the unit with
    no Install section, so the OFF branch is not a dead branch

docs/ax-conwip.md carries what the CONWIP is, what the unit would run, the
nine-step sequence to enable it, and the line that says the repository has no
remote so nothing is packaged.

Nothing here is enabled, nothing was started and nothing was switched. The flip
is a separate act and it is Tom's.

Refs #455 (the gate's own pre-existing bug: nix flake check does not reach a
verdict on this repository).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ers directory

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…/fleet-bringup

# Conflicts:
#	hosts/nas/default.nix
Stock ax v0.3.0 never learns that a Task's command exited, so the Task stays
Running and its actor keeps a worker. The 2026-09-23 Substrate probe measured
the consequence: finished Tasks hold every worker of a small pool and the next
Task fails ResourceExhausted.

p1-completion.patch (on top of sandbox-class.patch, both now under
pkgs/ax/patches/):
- runner serves /metadata/v1alpha1/ax/{exit,result,usage}; result capped 1 MiB
- controller: terminal guard, exit read after resume, --running-resync (15s),
  Completed / Failed ExitCode=N write-back, result and usage copy, then
  SuspendActor; a CRASHED actor with no exit report is Failed ActorCrashed
- API: TaskStatus.command, UsageStats.tool_calls, rpc GetTaskResult and
  `ax result task <name>`; generated Go regenerated

Tests run in checkPhase (doCheck stays true), including a four-Task floor
test on a 2-worker mock pool and the no-exit-report contrast. vendorHash and
the vendored Substrate client pin 672533541dbf are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…asset

buildGoModule with the nixpkgs-go toolchain (go 1.27.1), vendored tree,
CGO off, version ldflags d277088b: ate-setup plus ateapi, atecontroller,
atelet, atenet, podcertcontroller and ateom-gvisor.

images.nix: the six component images as OCI layouts with a digest file,
the eight third-party images the kind install pulls as digest-preserving
fixed-output copies of their linux/amd64 manifests, and the gVisor
nightly tarball SandboxConfig gvisor-default names (sha256 d547d814).

Patches touch manifests only: 0001 points pauseImage at the NAS registry
through atelet's localhost rewrite; 0002 pins third-party images to their
linux/amd64 child digests so the NAS store holds about 0.7 GB instead of
every platform.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mage

OCI layouts with a `digest` file (pkgs/ax/oci-layout.nix), the shape
myAxFleet.registrySeed takes. The digest is read at build time; nothing
imports from a derivation.

- pkgs/ax/images.nix: the patched ax's server and controller (uid 65532,
  cacert) and nixpkgs redis for ax-redis.
- pkgs/ax-agent-image: ax-task-runner at /usr/local/bin (PID 1), pi from
  llm-agents, and the ax-agent adapter with modes halogen-smoke, pi
  (the probe's adapter-pi.sh and validate.py), fetch URL and exit N.
  pi-models.json names Halogen only, no secret; the endpoint comes from
  $HALOGEN_URL at run time so one digest serves every host. No claude-code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…modules/ax-fleet)

One module, one kill switch per host (myAxFleet.enable), three roles:
control (nas: k3s server with its kubelet, untainted; registry; seed and
bootstrap runners), harness (coordinator: k3s agent tainted
ate.dev/sandboxClass=gvisor:NoSchedule, labelled with the Substrate
version at registration; guard chain; NetworkManager conf.d drop-in and a
config reload, never a restart), inference (worker: one assertion).

Folds #447's CIDRs, assertions, feature gates and runtime-config; drops
its Cilium, containerd template and runsc RuntimeClass (Substrate runs its
own runsc). Keeps #446's registry only, on the data pool. flannel VXLAN and
kube-proxy bound to the LAN leg. k3s state bind-mounted from /mnt/fast,
PersistentVolumes under /mnt/nas/services/ax-fleet/local-path.

pkgs/ax-fleet-teardown wraps the pinned k3s-killall.sh, removes the guard
chain and restores the sysctls snapshotted before k3s first ran.

substrate.nix and ax.nix are empty for their tracks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…447 gates

nas = control, coordinator = harness (myAxClient on by mkDefault),
worker = inference (renders one assertion; not switched in this motion).
Deletes modules/k3s-fleet.nix (with its RuntimeClass manifest) and
hosts/nas/state-services.nix, superseded by modules/ax-fleet. The
k3s-token comment in secrets.nix no longer claims the ciphertext is missing.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eet-teardown

tests/ax-fleet: nas, coordinator, worker (Halogen stub), peer (tailnet
stand-in); base configs with ax off and an ax-on specialisation switched
live, NAS first. Phases 10-cluster and 90-rollback are the cluster track's;
20-substrate and 30-ax are empty for their tracks. Receipt in
$out/receipt.json.

ax-fleet-topology asserts the real hosts' rendered flags, firewall scope,
the NetworkManager.conf no-restart property, the kill switch (mkForce
false renders nothing) and parity with the VM nodes. ax-client-topology
now expects myAxClient ON on the coordinator only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… test

modules/ax-fleet/ax.nix writes only the myAxFleet extension points and the
coordinator's ax scripts:
- manifests.ax-fleet-40-ax: Namespace ax-system; ax-redis (AOF on a local-path
  PVC, no password, nothing on argv); ax-server on ClusterIP 10.201.0.80,
  never a NodePort; ax-controller with upstream's projected token and
  ClusterTrustBundle, --running-resync (P1), ATENET_ROUTER_ADDR and
  AX_SNAPSHOTS_BUCKET=gs://ate-snapshots/ax/. Every image by digest, the
  digests substituted at build time and checked (no IFD).
- registrySeed: ax-server, ax-controller, ax-redis, ax-agent. Images come from
  the flake's own nixpkgs and `ax`, never the host's pkgs, so the NAS seeds
  the digest the coordinator names.
- harness: ax-fleet-image-ref, ax-fleet-smoke (halogen, pi, exit N,
  egress-deny, floor N; JSON receipts) and AX_SERVER=http://127.0.0.1:8080.
- options myAxFleet.ax.{runningResyncSeconds,atespace}.

tests/ax-fleet/phases/30-ax.py: T1 to T5 and the resilience checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…with ate-setup

Fills only the myAxFleet extension points, on the control role:
- registrySeed: six component images (substrate/<name>:d277088b) and the
  eight third-party images under their upstream repositories;
- bootstrap 20-registry-svc: Namespace ate-system and the kind-registry
  Service plus EndpointSlice to the NAS registry, which atelet's
  --localhost-registry-replacement names;
- 30-substrate: upstream ate-setup --kind deploy ate-system with
  --image-repo <registry>/substrate --image-tag d277088b, once per
  (installer, images) closure, stamped in kube-system/ax-fleet-substrate;
- 40-gvisor-asset: bucket gvisor in the in-cluster RustFS holds the pinned
  runsc tarball at atelet's fallback key, sha256-verified; the upstream kind
  credential is read from the Deployment and kept off argv;
- 50-workerpool: WorkerPool ateom-gvisor, replicas and memory limit from the
  options, harness nodeSelector and taint toleration, by digest.

tests/ax-fleet/phases/20-substrate.py: the Substrate subtests of the 4-VM
check. images.nix: drop an unused argument (deadnix).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…losure

The bootstrap referenced the full substrate output (seven binaries, 377 MB)
and the patched source (172 MB) at run time. It now runs passthru.ate-setup
(53 MB) from passthru.installTree (go.mod, manifests/, hack/: 1.4 MB), which
is everything deploy ate-system reads. The component binaries reach the NAS
only inside the images. Saves about 0.5 GB on the NAS's eMMC root.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ax-fleet-kubeconfig: quote the ssh host (shellcheck SC2086 failed the
  coordinator toplevel build).
- tests: wait out k3s's transient node taints and the flannel.1 link
  instead of asserting once; A-record DNS query (the stand-in has no
  upstream, AAAA is REFUSED); a cross-node pod-to-pod subtest over VXLAN;
  typed receipt (test-driver type check); records also printed to the log.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts:
#	modules/ax-fleet/substrate.nix
#	tests/ax-fleet/phases/20-substrate.py
# Conflicts:
#	modules/ax-fleet/ax.nix
#	tests/ax-fleet/phases/30-ax.py
…, prove it in the VM test

Merges ax/fleet-cluster (cc39eac), ax/fleet-substrate (1165567) and
ax/fleet-ax (71ce913); the four add/add placeholder conflicts take each
track's own file. The gates stay ON in hosts/{nas,coordinator,worker}
(control, harness, inference).

- pkgs/ax-agent-image: pi's bun binary lists PT_LOAD out of p_vaddr order.
  Linux runs it; gVisor refuses it (exit 126 inside the sandbox, MEASURED in
  the VM test). elf-sort-load.py reorders the program header table only, so
  the Task image's pi runs in gVisor (T2 schema-valid, MEASURED).
- pkgs/ax-agent-image: a `probe` mode (claude --version when the image has
  claude-code, a GET of Halogen's /v1/models, /proc/version); pi mode records
  its stderr tail; test-only `extraPaths`/`variant` arguments.
- ax-fleet-smoke: --image, --sandbox-class and the `probe` case; an exited
  command with protobuf-omitted exitCode reads as 0.
- tests/ax-fleet: phase 25-harness-probe (one sandboxClass gvisor Task on the
  coordinator in a test-only image variant with claude-code and no
  credential: Completed, sampled Completed for 60 s); the NAS test node seeds
  that variant; test nodes import modules/ax-client.nix as the hosts do; the
  Halogen stub answers a JSON Schema prompt with {"answer": 42}; 20-substrate
  checks the pod-certificate controller in its own namespace and stops
  shadowing the driver's `step` and `log`.
- flake: packages ax-agent-image and substrate-images.

checks.x86_64-linux.ax-fleet (4 VMs): rc 0, 37 subtests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ngup

The coordinator runs configurationRevision 3a65899, which is not in
origin/main. Stacking on it keeps claude-code 2.1.280, the closed 8731
wifi door (#460), the theme switcher and the harvest-hook removal, so a
switch from this branch adds only the ax-fleet delta.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…(DF-6)

a775d58 (remove the SessionEnd harvest hook) also carried the
home-profiles asserts from f1d689b (#448, parakeet socket activation),
which this branch does not contain. The coordinator home here still has
parakeet-service with Install.WantedBy = [ "default.target" ] and no
socket, so checks.home-profiles failed with
`attribute 'parakeet-service' missing` at flake.nix:2036.

Restore main's asserts for this check. The socket asserts belong in #448
and return when it merges. No module or host change; only the check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sure, panic tunables, RBAC, RustFS credential, real flap)

- Stacked on the live coordinator revision 3a65899 (merge commit before
  this one) plus f7b34239 (DF-6, check-only), so a coordinator switch
  adds only the ax-fleet delta: same home-manager generation as live, no
  tcp 8731 rule, no harness version rollback.
- ax-fleet-teardown: inheritPath = false, refuse if tailscale is on PATH,
  so k3s-killall.sh never runs `tailscale set --advertise-routes=` (the
  NAS's 10.42.0.0/24 route). Also drops the NAS guard table.
- NAS: docker-registry waits up to 30 s for 10.42.0.1 and restarts on
  failure; the seed wants (not requires) it and restarts on failure;
  k3s waits for the LAN address on both nodes. The boot test now adds
  the address 40 s late and proves the registry recovers.
- NAS: inet ax-fleet-guard prerouting at priority raw drops routed LAN
  traffic to the pod and Service ranges before kube-proxy DNAT.
- Coordinator guard: flannel.1 payloads must come from the pod CIDR; pod
  traffic out the LAN leg is dropped for every private range.
- kubelet's kernel.panic, panic_on_oops, overcommit_memory are put back
  after each kubelet start (myAxFleet.kubelet.keepHostKernelTunables,
  default true, Tom's ruling pending).
- ax-controller: no ClusterRole on Secrets, no automounted API token.
- Substrate patch 0003: RustFS, ate-api-server, atelet and the bucket-init
  Job read a per-cluster credential from Secret ax-fleet-rustfs, generated
  once on the NAS by bootstrap step 25-rustfs-secret.
- k3s bind mounts unmount lazily, so the rollback switch no longer fails
  on a busy /var/lib/kubelet.
- VM test: the flap waits for Ready=Unknown and the unreachable taint;
  new subtests for the routed-LAN paths, the Freebox-style foreign /24,
  the tunables restore, the RBAC and credential surface, and the
  teardown never calling tailscale.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mecattaf and others added 2 commits September 23, 2026 12:41
…en tunables, conntrack, teardown on PATH, read-only registry, split k3s token, NAS parity, boot snapshot, receipt)

- harness: refuse NEW connections from cni0/flannel.1 at the head of nixos-fw
  (v4 and v6); user.slice and system.slice CPUWeight 10000 (deskCpuWeight).
- k3s: kernel tunables watcher bound to k3s (no 900 s bound); boot snapshot
  takes non-ax declared values (NIXOS_ACTION unset at boot); kube-proxy
  conntrack args 0; conntrack keys snapshotted; teardown on PATH for the
  control and harness roles whatever the switch says; server and agent
  tokens split (agentTokenFile on the NAS).
- control: registry read-only (maintenance.readonly, delete off); the seed
  pushes through a loopback writer that lives only for the seed run.
- substrate: images built from inputs.nixpkgs, one digest set everywhere.
- secrets: k3s-token.age re-minted for editors ++ nasOnly, new
  k3s-agent-token.age for editors ++ coordinatorOnly ++ nasOnly.
- tests: ax-fleet-boot on nixpkgs-stable with the systemd initrd and the
  NAS kernel; VM coordinator on the desk kernel with sshd; new assertions;
  every phase subtest recorded in receipt.json.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pod egress, role-scoped guards through the kill switch, proxy ARP by interface, API owner match, proxy off :8080, bounded tracing)

- harness guard: pod egress policed by destination on any interface, nothing
  enters cni0/flannel.1 but the NAS host; enp191s0 is a LAN leg
  (lan.extraInterfaces) with NetworkManager route metric 700
- control: forward chain in inet ax-fleet-guard; NAS pods reach only
  Halogen's host:port among private addresses (the Gateway port is ignored
  upstream), never the tailnet or ve-*
- guards, pod-input refusal and API owner match render for the role
  whatever enable says; the teardown leaves declared guards
- proxy_arp on cni0, flannel.1, veth* by name; conf.default untouched
- ax-server-proxy on 127.0.0.1:8099 as a static user; OUTPUT owner match
  admits only root, apiUsers and the proxy to it and the cluster ranges
- substrate patch 0004: jaeger max-traces and limits, collector
  memory_limiter, limits, debug basic
- VM: eth3 DHCP leg, worker :2222, T4b, proxy ARP, API owner, rollback
  pod-to-host probes; topology asserts

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@mecattaf

Copy link
Copy Markdown
Owner Author

Superseded by #467 (ax/fleet-zero 9cfc57dd, base main): same fleet bring-up with zero ax patches (P1 and sandbox-class dropped, no p2), the egress fail-open closed without an ax patch, the substrate-link module vendored (gate OFF), round-1 test fixes, and main (#466, #461) merged in. Branch ax/fleet-bringup (f1bba47e) is kept.

@mecattaf mecattaf closed this Sep 23, 2026
mecattaf added a commit that referenced this pull request Sep 23, 2026
ax on Agent Substrate across the fleet, zero ax patches (supersedes #464)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant