ax on Agent Substrate across the fleet, zero ax patches (supersedes #464) - #467
Merged
Merged
Conversation
Tom, 2026-09-20: "having local kubernetes s3 or redis or postgres or whatever it needs on the NAS". This is that sentence for the three Substrate actually reads. State on the appliance, machines on the Strix boxes -- the NAS has 8 threads and 22 GiB and runs the house's DNS, so it holds the control plane's state and schedules nothing (Appendix J section 5). A SECOND DATABASE, not a second PostgreSQL. hosts/nas/media.nix already brought the instance up and pinned its dataDir to the NVMe; this rides it, so there is one instance and one backup story. ensureDatabases/ensureUsers with ensureDBOwnership, because ateapi runs goose migrations at startup and creates its own tables. enableTCPIP is forced by the deployment and not by taste: ateapi is a pod on a Strix box, the k3s server here sets disableAgent, so there is no unix socket to peer-authenticate over. scram-sha-256 from the LAN and from the pod CIDR, never trust. THE OBJECT STORE IS A HAND-WRITTEN UNIT ON PURPOSE. This host rides nixpkgs-stable (nixos-26.05) and stable has no rustfs at all -- no module, no package. The main pin has both; the package comes across the hosts/nas/unstable-pkgs.nix seam that attic-server and Immich already use, and the unit is modelled line for line on the unstable module's own serviceConfig. Delete it for services.rustfs when the NAS next rides a stable that ships it. Every unit carries RequiresMountsFor on /mnt/fast. hosts/nas/attic.nix learned the second edition of the signing-key trap: /mnt/fast is nofail, so a failed mount otherwise yields an empty state tree and a service that starts happily against it. For an object store that is silently lost snapshots. Refusing to start is a Tuesday. Secrets are runbook-placed root-owned files, not agenix, per attic.nix's no-agenix doctrine -- the rustfs key pair and the Substrate role's password are consumed by this host alone and a runbook can place them once. The password oneshot reads its file through LoadCredential and binds it as a psql value, so it lands in no argv and no journal line. Firewall: one block, three ports, iifname "enp1s0" only. Not openFirewall, which would publish all three on the tailnet as well. All three are on hosts/nas/cloudflared.nix's never-routed list. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Tom, 2026-09-20: "kubernetes is world-class for that ... they each stay in
their lane. Effects ts handles the ultracode-level json dag specification and
kubernetes schedules it on the right machines."
One shared modules/k3s-fleet.nix imported by all three hosts, because the
numbers have to agree and a second copy of them is how they stop agreeing.
Server on the NAS with disableAgent, so the appliance holds the control plane
and schedules nothing; agents on the Strix boxes, which have the cores, the
memory and /dev/kvm.
THE POD CIDR MOVES, AND THE ASSERTION IS OUTSIDE THE GATE. k3s defaults to
10.42.0.0/16 for pods and the house LAN is 10.42.0.0/24, which contains the
NAS, the coordinator, the worker, the printer and the router. Tom's 2026-09-09
research already picked the replacement: pods 10.200.0.0/16, services
10.201.0.0/16. Those are kept exactly, and the overlap check is real
arithmetic in a real assertion placed OUTSIDE the enable gate, so it
evaluates on every host whether k3s is on or off. Verified against six cases,
including that k3s's own default IS caught.
TRACK K AND TRACK S ARE THE SAME FILE, and that is the point. Design K is
plain k3s plus Cilium plus the gvisor RuntimeClass; Design S adds Substrate on
top of exactly this. The only thing S needs from the bottom layer that K does
not is four feature-gate settings, and a gate nothing asks for is inert: no
scheduling change, no new controller, no memory. So the gates are in
unconditionally and there is one PR instead of two that drift.
The gate names are upstream's own, MEASURED from ~/Downloads/substrate,
hack/create-kind-cluster.sh:104-112: ClusterTrustBundle,
ClusterTrustBundleProjection, PodCertificateRequest, plus runtimeConfig
certificates.k8s.io/v1beta1. The kubelet is handed only the two that are its
own, because an unrecognised gate name is fatal to kubelet and
ClusterTrustBundle is apiserver-side. Whether k3s 1.35.6+k3s1 accepts them at
all is U9 and is not claimed here.
Cilium 1.18.14 through autoDeployCharts, so the CNI lives in the generation
rather than in a `cilium install` somebody has to remember. The chart hash is
measured, not fakeHash: built with fakeHash, took the reported hash, rebuilt
clean. The gvisor RuntimeClass is a server manifest and its containerd runtime
is an agent containerdConfigTemplate that keeps `{{ template "base" . }}` --
one line, load-bearing, and dropping it costs the node its CNI, its
snapshotter and its registry mirrors at once. kata is written out and
commented, pointing at tonight's U14 report.
The kube API is LAN-only: 6443 on the NAS's enp1s0 and nothing else, plus the
never-routed doctrine in the cloudflared PR, because a tunnel ingress bypasses
nftables entirely and "firewalled" is therefore a separate guarantee from
"not in the tunnel".
All three hosts land with the gate OFF. secrets/k3s-token.age does not exist;
minting it needs Tom's admin age key.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A5a stood the pinned k3s up inside a throwaway guest built from this exact nixpkgs revision the same night and measured three things this file had wrong. Report: ~/today/review/2026-09-20/sandbox-spike/U9-U13-U15-K3S.md. 1. THE THREE GATES ALONE SERVE NOTHING. Boot 5 held the gates on and dropped the runtime-config flag: all three kubernetes_feature_enabled metrics read 1 while /apis/certificates.k8s.io listed only v1 and api-resources returned NO_RESOURCES. A check that read only the metric would have reported a false pass. --kube-apiserver-arg=runtime-config=certificates.k8s.io/v1beta1=true is what turns the group version on, and podcertcontroller has nothing to talk to without it. It now sits beside the gates, and the flag values inside a -arg= carry no leading dashes of their own, which is A5a's measured working form. U9 is therefore ANSWERED YES on the pin: v1beta1 serves clustertrustbundles and podcertificaterequests. Appendix J section 9's fallback does not have to be taken and no newer k3s is needed. kubelet also accepts all three gate names, so the cautious subset this file used to hand it is gone. 2. THE NIXPKGS EXAMPLE'S CONTAINERD KEY PATH IS WRONG FOR THIS k3s. The bundled containerd is 2.2.5-k3s2 and the config it generates starts `version = 3` with runtimes under [plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.runc]. A grpc.v1.cri block would be parsed, accepted and silently ignored, which is the worst of the three outcomes. A5a's exact working table is used here. And runtime_path, NOT options.BinaryName. The BinaryName shape this file used is a trap that looks right: a pod under it reaches Running and stays 1/1 Running in kubelet's view, produces no logs at all, and kubectl exec fails with "in state stopped". The generic runc shim starts runsc but carries neither its stdio nor its state. containerd-shim-runsc-v1 is what goes here, and nixpkgs' gvisor builds it. 3. k3s DOES NOT FIND runsc ON ITS OWN. Not by auto-detection (with gvisor in systemPackages the generated config still carried only runc and runhcs-wcow-process) and not from the unit PATH, which the nixpkgs rancher module leaves empty but for zfs. Without the path line every sandbox dies at creation with `exec: "runsc": executable file not found in $PATH`. The line was already here; it now carries why, and A5a's caution that the assignment replaces rather than extends. Also corrected: this branch's earlier claim that identical gate-off derivation paths proved the module inert. They do not prove it. The flake embeds the configuration revision, so the toplevel tracks the commit. The sound check is that with the gate off, services.k3s.enable is false, no 6443 rule is emitted and no registries.yaml exists on any of the three hosts, which was measured directly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pkgs/ax builds google/ax v0.3.0 pinned by commit d8ed0fe (tag v0.3.0), wired through overlays/default.nix and the flake's packages.<system> list like every sibling package. modules/ax-client.nix puts kubectl and ax into environment.systemPackages on coordinator, worker and client when myAxClient.enable is true; it lands false on all three. Refs #453. The one judgement call: ax's go.mod opens `go 1.27.1` and Go refuses to build a module whose go directive is newer than the running toolchain (`go: go.mod requires go >= 1.27.1 (running go 1.27.0; GOTOOLCHAIN=local)`), with no network in the sandbox to fetch one. MEASURED 2026-09-23: nixpkgs has go_1_27 = 1.27rc2 and nixpkgs-fresh has 1.27.0, so neither pin already in this flake can build it. Adds one input, nixpkgs-go, pinned BY REVISION and supplying exactly one attribute to exactly one package - the same shape as the existing nixpkgs-paperless. Bumping nixpkgs-fresh instead was rejected on purpose: that input also carries the twins' linux 7.2 kernel. checks.ax-client-topology asserts the option exists on the three hosts that import the module, that it is false on all three, that kubectl and ax are absent from all three systemPackages, and that hosts/nas carries no such option at all. It goes red on the flip by design. No switch, no rebuild, no gate flipped, no cluster, no Substrate, no patches - this builds pristine v0.3.0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds `string sandbox_class = 11` to TaskSpec, regenerates ax.pb.go with the same protoc-gen-go v1.36.11 upstream used, refuses unknown values in ValidateTask, and threads the value from the reconciler through BuildActorTemplate, replacing the SandboxClass_SANDBOX_CLASS_GVISOR hardcode at internal/substrate/client.go:273 (the only SANDBOX_CLASS occurrence in the tree). Empty means gvisor, so existing manifests behave exactly as before. The generated Go is in the patch on purpose: ax bridges YAML through protojson with unknown fields rejected, so a .proto-only edit would make every manifest naming sandboxClass fail strict decode. This does NOT deliver a workerd sandbox. Agent Substrate's SandboxClass enum has three members (UNSPECIFIED, GVISOR, MICROVM), so "gvisor" and "microvm" are the only values that can reach a real substrate. A workerd class needs an upstream Agent Substrate change that does not exist. vendorHash is unchanged: the patch touches no go.mod or go.sum line. Measured on the patched scratch tree with go 1.27.1: `go build ./...` rc 0, `go test ./...` rc 0. In this clone: `nix build --no-link .#ax` rc 0, with the patch applying to all six files and checkPhase passing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
home/ax-conwip.nix declares `myAxConwip`, a home-manager module that WOULD run
the CONWIP scheduler at /home/tom/mecattaf/ax-conwip as one long-running user
service against a declared ax server, with the rewrite's seat meter directory as
a read-only input. It is imported from home/home.nix and gated OFF, so it defines
no unit on any host today.
Seven options, every one of them data in this repository rather than a command
line somebody types: enable (false), serverUrl (127.0.0.1:8080, the loopback the
mock stack listens on, so an accidental enable reaches nothing real), metersDir
(the REWRITE's meters, ~/.local/state/tally-rewrite/meters, never branch (a)'s,
read only, no tmpfiles rule over it), wipCap (1, stricter than the program's own
default of 2), sourceDir, recordsDir (a path that does not exist, so an
accidental enable fails legibly instead of passing over an empty directory) and
stateDir.
Four things deliberately not done, and the module header says all four out loud:
- no flake input, because the CONWIP repository has NO REMOTE. A fetchGit of a
path on one box is not a declaration, it is a machine-local accident that
would break every other host's eval;
- no package, because packaging follows the input. The unit runs the checkout
in place through `pnpm exec tsx`. That is a development shape, not a
delivered one, and it is the biggest single reason the gate stays off;
- no timer: the release signal is a WatchTask stream, not a poll;
- nothing enabled and nothing armed. `enable` is set nowhere, and even with
the gate flipped the unit carries NO Install section, so a rebuild declares
it and a `systemctl --user start` is a second, separate act.
checks.x86_64-linux.ax-conwip-topology asserts the option EXISTS on all three
home-manager hosts (a dropped import makes it undefined, not false, which is the
mutation hint), that it is false on all three, that no ax-conwip service, timer
or socket is rendered anywhere, that there is no system-bus twin on any of the
four hosts, and that no tmpfiles rule naming ax-conwip exists while the gate is
off. The `enable == false` line goes red on the flip, deliberately.
MEASURED in this tree:
nix build --no-link .#checks.x86_64-linux.ax-conwip-topology -> rc 0, so every
assertion above ran and passed
nix eval .#nixosConfigurations.{coordinator,worker,client}.config
.home-manager.users.tom.systemd.user.services --apply builtins.attrNames
-> rc 0, no ax-conwip on any of the three
the same eval under extendModules with the gate flipped renders the unit with
no Install section, so the OFF branch is not a dead branch
docs/ax-conwip.md carries what the CONWIP is, what the unit would run, the
nine-step sequence to enable it, and the line that says the repository has no
remote so nothing is packaged.
Nothing here is enabled, nothing was started and nothing was switched. The flip
is a separate act and it is Tom's.
Refs #455 (the gate's own pre-existing bug: nix flake check does not reach a
verdict on this repository).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ers directory Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…' into ax/fleet-bringup
…/fleet-bringup # Conflicts: # hosts/nas/default.nix
# Conflicts: # hosts/worker/default.nix
Stock ax v0.3.0 never learns that a Task's command exited, so the Task stays
Running and its actor keeps a worker. The 2026-09-23 Substrate probe measured
the consequence: finished Tasks hold every worker of a small pool and the next
Task fails ResourceExhausted.
p1-completion.patch (on top of sandbox-class.patch, both now under
pkgs/ax/patches/):
- runner serves /metadata/v1alpha1/ax/{exit,result,usage}; result capped 1 MiB
- controller: terminal guard, exit read after resume, --running-resync (15s),
Completed / Failed ExitCode=N write-back, result and usage copy, then
SuspendActor; a CRASHED actor with no exit report is Failed ActorCrashed
- API: TaskStatus.command, UsageStats.tool_calls, rpc GetTaskResult and
`ax result task <name>`; generated Go regenerated
Tests run in checkPhase (doCheck stays true), including a four-Task floor
test on a 2-worker mock pool and the no-exit-report contrast. vendorHash and
the vendored Substrate client pin 672533541dbf are unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…asset buildGoModule with the nixpkgs-go toolchain (go 1.27.1), vendored tree, CGO off, version ldflags d277088b: ate-setup plus ateapi, atecontroller, atelet, atenet, podcertcontroller and ateom-gvisor. images.nix: the six component images as OCI layouts with a digest file, the eight third-party images the kind install pulls as digest-preserving fixed-output copies of their linux/amd64 manifests, and the gVisor nightly tarball SandboxConfig gvisor-default names (sha256 d547d814). Patches touch manifests only: 0001 points pauseImage at the NAS registry through atelet's localhost rewrite; 0002 pins third-party images to their linux/amd64 child digests so the NAS store holds about 0.7 GB instead of every platform. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mage OCI layouts with a `digest` file (pkgs/ax/oci-layout.nix), the shape myAxFleet.registrySeed takes. The digest is read at build time; nothing imports from a derivation. - pkgs/ax/images.nix: the patched ax's server and controller (uid 65532, cacert) and nixpkgs redis for ax-redis. - pkgs/ax-agent-image: ax-task-runner at /usr/local/bin (PID 1), pi from llm-agents, and the ax-agent adapter with modes halogen-smoke, pi (the probe's adapter-pi.sh and validate.py), fetch URL and exit N. pi-models.json names Halogen only, no secret; the endpoint comes from $HALOGEN_URL at run time so one digest serves every host. No claude-code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…modules/ax-fleet) One module, one kill switch per host (myAxFleet.enable), three roles: control (nas: k3s server with its kubelet, untainted; registry; seed and bootstrap runners), harness (coordinator: k3s agent tainted ate.dev/sandboxClass=gvisor:NoSchedule, labelled with the Substrate version at registration; guard chain; NetworkManager conf.d drop-in and a config reload, never a restart), inference (worker: one assertion). Folds #447's CIDRs, assertions, feature gates and runtime-config; drops its Cilium, containerd template and runsc RuntimeClass (Substrate runs its own runsc). Keeps #446's registry only, on the data pool. flannel VXLAN and kube-proxy bound to the LAN leg. k3s state bind-mounted from /mnt/fast, PersistentVolumes under /mnt/nas/services/ax-fleet/local-path. pkgs/ax-fleet-teardown wraps the pinned k3s-killall.sh, removes the guard chain and restores the sysctls snapshotted before k3s first ran. substrate.nix and ax.nix are empty for their tracks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…447 gates nas = control, coordinator = harness (myAxClient on by mkDefault), worker = inference (renders one assertion; not switched in this motion). Deletes modules/k3s-fleet.nix (with its RuntimeClass manifest) and hosts/nas/state-services.nix, superseded by modules/ax-fleet. The k3s-token comment in secrets.nix no longer claims the ciphertext is missing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eet-teardown tests/ax-fleet: nas, coordinator, worker (Halogen stub), peer (tailnet stand-in); base configs with ax off and an ax-on specialisation switched live, NAS first. Phases 10-cluster and 90-rollback are the cluster track's; 20-substrate and 30-ax are empty for their tracks. Receipt in $out/receipt.json. ax-fleet-topology asserts the real hosts' rendered flags, firewall scope, the NetworkManager.conf no-restart property, the kill switch (mkForce false renders nothing) and parity with the VM nodes. ax-client-topology now expects myAxClient ON on the coordinator only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… test modules/ax-fleet/ax.nix writes only the myAxFleet extension points and the coordinator's ax scripts: - manifests.ax-fleet-40-ax: Namespace ax-system; ax-redis (AOF on a local-path PVC, no password, nothing on argv); ax-server on ClusterIP 10.201.0.80, never a NodePort; ax-controller with upstream's projected token and ClusterTrustBundle, --running-resync (P1), ATENET_ROUTER_ADDR and AX_SNAPSHOTS_BUCKET=gs://ate-snapshots/ax/. Every image by digest, the digests substituted at build time and checked (no IFD). - registrySeed: ax-server, ax-controller, ax-redis, ax-agent. Images come from the flake's own nixpkgs and `ax`, never the host's pkgs, so the NAS seeds the digest the coordinator names. - harness: ax-fleet-image-ref, ax-fleet-smoke (halogen, pi, exit N, egress-deny, floor N; JSON receipts) and AX_SERVER=http://127.0.0.1:8080. - options myAxFleet.ax.{runningResyncSeconds,atespace}. tests/ax-fleet/phases/30-ax.py: T1 to T5 and the resilience checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…with ate-setup Fills only the myAxFleet extension points, on the control role: - registrySeed: six component images (substrate/<name>:d277088b) and the eight third-party images under their upstream repositories; - bootstrap 20-registry-svc: Namespace ate-system and the kind-registry Service plus EndpointSlice to the NAS registry, which atelet's --localhost-registry-replacement names; - 30-substrate: upstream ate-setup --kind deploy ate-system with --image-repo <registry>/substrate --image-tag d277088b, once per (installer, images) closure, stamped in kube-system/ax-fleet-substrate; - 40-gvisor-asset: bucket gvisor in the in-cluster RustFS holds the pinned runsc tarball at atelet's fallback key, sha256-verified; the upstream kind credential is read from the Deployment and kept off argv; - 50-workerpool: WorkerPool ateom-gvisor, replicas and memory limit from the options, harness nodeSelector and taint toleration, by digest. tests/ax-fleet/phases/20-substrate.py: the Substrate subtests of the 4-VM check. images.nix: drop an unused argument (deadnix). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…losure The bootstrap referenced the full substrate output (seven binaries, 377 MB) and the patched source (172 MB) at run time. It now runs passthru.ate-setup (53 MB) from passthru.installTree (go.mod, manifests/, hack/: 1.4 MB), which is everything deploy ate-system reads. The component binaries reach the NAS only inside the images. Saves about 0.5 GB on the NAS's eMMC root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ax-fleet-kubeconfig: quote the ssh host (shellcheck SC2086 failed the coordinator toplevel build). - tests: wait out k3s's transient node taints and the flannel.1 link instead of asserting once; A-record DNS query (the stand-in has no upstream, AAAA is REFUSED); a cross-node pod-to-pod subtest over VXLAN; typed receipt (test-driver type check); records also printed to the log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # modules/ax-fleet/substrate.nix # tests/ax-fleet/phases/20-substrate.py
# Conflicts: # modules/ax-fleet/ax.nix # tests/ax-fleet/phases/30-ax.py
…, prove it in the VM test Merges ax/fleet-cluster (cc39eac), ax/fleet-substrate (1165567) and ax/fleet-ax (71ce913); the four add/add placeholder conflicts take each track's own file. The gates stay ON in hosts/{nas,coordinator,worker} (control, harness, inference). - pkgs/ax-agent-image: pi's bun binary lists PT_LOAD out of p_vaddr order. Linux runs it; gVisor refuses it (exit 126 inside the sandbox, MEASURED in the VM test). elf-sort-load.py reorders the program header table only, so the Task image's pi runs in gVisor (T2 schema-valid, MEASURED). - pkgs/ax-agent-image: a `probe` mode (claude --version when the image has claude-code, a GET of Halogen's /v1/models, /proc/version); pi mode records its stderr tail; test-only `extraPaths`/`variant` arguments. - ax-fleet-smoke: --image, --sandbox-class and the `probe` case; an exited command with protobuf-omitted exitCode reads as 0. - tests/ax-fleet: phase 25-harness-probe (one sandboxClass gvisor Task on the coordinator in a test-only image variant with claude-code and no credential: Completed, sampled Completed for 60 s); the NAS test node seeds that variant; test nodes import modules/ax-client.nix as the hosts do; the Halogen stub answers a JSON Schema prompt with {"answer": 42}; 20-substrate checks the pod-certificate controller in its own namespace and stops shadowing the driver's `step` and `log`. - flake: packages ax-agent-image and substrate-images. checks.x86_64-linux.ax-fleet (4 VMs): rc 0, 37 subtests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ngup The coordinator runs configurationRevision 3a65899, which is not in origin/main. Stacking on it keeps claude-code 2.1.280, the closed 8731 wifi door (#460), the theme switcher and the harvest-hook removal, so a switch from this branch adds only the ax-fleet delta. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…(DF-6) a775d58 (remove the SessionEnd harvest hook) also carried the home-profiles asserts from f1d689b (#448, parakeet socket activation), which this branch does not contain. The coordinator home here still has parakeet-service with Install.WantedBy = [ "default.target" ] and no socket, so checks.home-profiles failed with `attribute 'parakeet-service' missing` at flake.nix:2036. Restore main's asserts for this check. The socket asserts belong in #448 and return when it merges. No module or host change; only the check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r deletes the Task Probe branch, not for merge. pkgs/ax builds v0.3.0 with sandbox-class.patch only (p1-completion.patch removed) and ax-controller runs without --running-resync. The VM test keeps phases 10 and 20 and replaces the rest with 30-nop1: each Task's command POSTs its result to a floor stand-in on the worker stub, and the driver (standing in for the link) deletes the Task through the stock ax client. Test-only kubectl-ate on the nas node for actor, template and worker listings. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… check Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…refuses a multi-line command in AX_TASK_YAML) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s requeue); ateapi diagnostics Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nosc); port stock-ax sandboxClass control into 30-nop1; assert gVisor kernel per report Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sure, panic tunables, RBAC, RustFS credential, real flap) - Stacked on the live coordinator revision 3a65899 (merge commit before this one) plus f7b34239 (DF-6, check-only), so a coordinator switch adds only the ax-fleet delta: same home-manager generation as live, no tcp 8731 rule, no harness version rollback. - ax-fleet-teardown: inheritPath = false, refuse if tailscale is on PATH, so k3s-killall.sh never runs `tailscale set --advertise-routes=` (the NAS's 10.42.0.0/24 route). Also drops the NAS guard table. - NAS: docker-registry waits up to 30 s for 10.42.0.1 and restarts on failure; the seed wants (not requires) it and restarts on failure; k3s waits for the LAN address on both nodes. The boot test now adds the address 40 s late and proves the registry recovers. - NAS: inet ax-fleet-guard prerouting at priority raw drops routed LAN traffic to the pod and Service ranges before kube-proxy DNAT. - Coordinator guard: flannel.1 payloads must come from the pod CIDR; pod traffic out the LAN leg is dropped for every private range. - kubelet's kernel.panic, panic_on_oops, overcommit_memory are put back after each kubelet start (myAxFleet.kubelet.keepHostKernelTunables, default true, Tom's ruling pending). - ax-controller: no ClusterRole on Secrets, no automounted API token. - Substrate patch 0003: RustFS, ate-api-server, atelet and the bucket-init Job read a per-cluster credential from Secret ax-fleet-rustfs, generated once on the NAS by bootstrap step 25-rustfs-secret. - k3s bind mounts unmount lazily, so the rollback switch no longer fails on a busy /var/lib/kubelet. - VM test: the flap waits for Ready=Unknown and the unreachable taint; new subtests for the routed-LAN paths, the Freebox-style foreign /24, the tunables restore, the RBAC and credential surface, and the teardown never calling tailscale. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…stream state when a Task delete does not finish; fix ateapi deploy name in diagnose Run 2 failed with nop1-ctl-2 stuck in Terminating and none of these logs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…en tunables, conntrack, teardown on PATH, read-only registry, split k3s token, NAS parity, boot snapshot, receipt) - harness: refuse NEW connections from cni0/flannel.1 at the head of nixos-fw (v4 and v6); user.slice and system.slice CPUWeight 10000 (deskCpuWeight). - k3s: kernel tunables watcher bound to k3s (no 900 s bound); boot snapshot takes non-ax declared values (NIXOS_ACTION unset at boot); kube-proxy conntrack args 0; conntrack keys snapshotted; teardown on PATH for the control and harness roles whatever the switch says; server and agent tokens split (agentTokenFile on the NAS). - control: registry read-only (maintenance.readonly, delete off); the seed pushes through a loopback writer that lives only for the seed run. - substrate: images built from inputs.nixpkgs, one digest set everywhere. - secrets: k3s-token.age re-minted for editors ++ nasOnly, new k3s-agent-token.age for editors ++ coordinatorOnly ++ nasOnly. - tests: ax-fleet-boot on nixpkgs-stable with the systemd initrd and the NAS kernel; VM coordinator on the desk kernel with sshd; new assertions; every phase subtest recorded in receipt.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pod egress, role-scoped guards through the kill switch, proxy ARP by interface, API owner match, proxy off :8080, bounded tracing) - harness guard: pod egress policed by destination on any interface, nothing enters cni0/flannel.1 but the NAS host; enp191s0 is a LAN leg (lan.extraInterfaces) with NetworkManager route metric 700 - control: forward chain in inet ax-fleet-guard; NAS pods reach only Halogen's host:port among private addresses (the Gateway port is ignored upstream), never the tailnet or ve-* - guards, pod-input refusal and API owner match render for the role whatever enable says; the teardown leaves declared guards - proxy_arp on cni0, flannel.1, veth* by name; conf.default untouched - ax-server-proxy on 127.0.0.1:8099 as a static user; OUTPUT owner match admits only root, apiUsers and the proxy to it and the cluster ranges - substrate patch 0004: jaeger max-traces and limits, collector memory_limiter, limits, debug basic - VM: eth3 DHCP leg, worker :2222, T4b, proxy ARP, API owner, rollback pod-to-host probes; topology asserts Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ed patches Ports the zero-patch probe onto ax/fleet-bringup f1bba47, keeping the bring-up's lan-guard, rollback, DF-6 and fix rounds 1 to 3. - pkgs/ax: patches = [ ]; sandbox-class.patch and p1-completion.patch removed. - ax-controller: no --running-resync; the runningResyncSeconds option and AX_FLEET_RESYNC_SECONDS are gone. - ax-fleet-smoke: the nosc change (the probe sends no spec.sandboxClass). - tests: 30-nop1 adopted (floor reports, the link's delete, the no-delete control, the stock-ax sandboxClass refusal, the delete-hang controller log capture), pointed at the :8099 proxy that fix round 3 introduced. 25-harness-probe and 30-ax asserted P1 (Completed, ax result); their non-P1 cases (claude --version probe, T4, T4b, API owner match, ax-redis and ax-controller restarts, the LAN flap) move to 32-fleet in the floor shape. 35-lan-guard and 90-rollback stay as they were. Not run here: the 4-VM test (next step, one at a time machine-wide). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…b885) Takes round 4's non-patch items and leaves p2-egress-deny.patch out (Tom, 09-23: no ax patches): - NAS: chain output in inet ax-fleet-guard; only root and clusterClientUids open NEW connections from the NAS host to the pod and Service CIDRs and the seed registry (ax-server API, ax-redis). - kubelet evictionHard restores the nodefs and imagefs signals; WorkerPool pods get ephemeral-storage 1Gi request, 32Gi limit. - ax-fleet-guard-apply, re-run by ax-fleet-teardown after k3s-killall.sh strips flannel's rules; the pod-input rules are idempotent. - ax-fleet-smoke --gateway NAME|none. - VM tests for the above (10-cluster, 20-substrate, 35-lan-guard, 90-rollback, ax-fleet-boot) and the TEST-NET-2 public stand-in on the worker. 30-ax's T4c (which asserted p2's EgressDenied) is not taken; the gateway-default phase replaces it. Round 4 was unreviewed and its VM runs 1 and 2 failed (REPORTED evals-2026-09-23/ax-fleet/receipts/fix-r4); not re-run in this step. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Stock ax v0.3.0 gives a Task with no gateway, or naming a missing one, an allow-all EgressPolicy (reconciler.go "Default to allow all egress"; MEASURED from the pinned source). Instead of p2-egress-deny.patch: - modules/ax-fleet/gateways.nix: myAxFleet.ax.atespaces declares every atespace the fleet creates (fleet and ax's CLI default "default"), each with its Gateways and a default one (Halogen only). Assertions: the default Gateway is declared for every atespace, its allowlist is not empty and allows no "*", 0.0.0.0/0 or ::/0; the smoke atespace is declared. - bootstrap step 80-ax-gateways applies every declared Gateway. - ax-fleet-gateway-default (control role, root, poll 2s): points any Task in a declared atespace without a usable gateway at the default through ax's own UpdateTask. pkgs/ax-gateway-default is a separate client binary built from stock ax's module tree (same vendorHash); ax is unchanged. - tests/ax-fleet/phases/33-gateway-default.py: the default exists in both atespaces; a Task with no gateway and one naming a missing gateway are pointed at the default, report to the floor, and cannot reach the TEST-NET-2 public stand-in; a Gateway allowlisting it gets 200 (positive control). Records the repoint window. The link's capacity 0 on a missing gateway (B10) stays the second guard. Not run in this step. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…and package it pkgs/substrate-link/src is `git archive a021003` of package.json, pnpm-lock.yaml, pnpm-workspace.yaml, tsconfig.json and apps/link from ax-conwip eval/2026-09-23-link (SYNC.md names the sha and the resync recipe). The package is the conwip-link derivation from evals-2026-09-23/link, renamed, with src = ./src and a real pnpmDeps hash (MEASURED: fakeHash build, then "got:"). Built: esbuild bundle of 2,103,690 bytes; run with no config it logs token-missing and exits 78. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…eet-80658978.patch) Applied by hand: the patch no longer applies at f1bba47 (secrets.nix:98). - hosts/nas/substrate-link.nix: myNas.substrateLink, enable = false by default, LoadCredential floor-link-token, RestartPreventExitStatus 78, outbound only, no port. package defaults to pkgs/substrate-link (the "package is null" assertion is gone with the vendored source). - completion defaults to "guest": ax carries no P1, so "p1" and "auto" would lease nothing (B6); guest needs a fleet-internal Complete relay that does not exist yet, so arming fails its assertion until then. - hosts/nas/default.nix imports it after modules/ax-fleet. - secrets.nix: secrets/floor-link-token.age for editors ++ nasOnly (the .age file is not minted; Tom's hand). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
The phases are concatenated into one test script; the test driver's type check rejected rebinding gw (dict[str, Any] in 30-nop1) to an optional gateway name. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…le connection r1 run 2 (MEASURED): after the LAN flap, after-flap-a1 failed ActorResumeFailed (Unavailable, stale ateapi->atelet connection). fleet_run deleted it; Substrate's atelet Terminate then failed on every try with 'failed to read sandbox record during terminate', because the aborted resume never wrote the record (cmd/atelet/main.go:1281). The actor stayed in ACTOR_STATE_DELETING and ax delete timed out after 300 s. Stock ax reconciles on every SaveTask and re-runs ResumeActor with no phase guard, so the recovery with no ax patch is to apply the same Task again. fleet_run and phase 33 now re-apply up to REAPPLY_RESUME times before falling back to delete-and-recreate. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
r1 run 3 (MEASURED): after the in-ax-on teardown and the k3s restart, kubectl wait Ready passed in 0.24 s on the pre-teardown status, and the test read status.podIP once (10.200.1.5, the old sandbox) and curled it for 300 s (curl rc 7). The teardown recreated every sandbox (four new veths on cni0). The wait now re-reads the IP on each try, records the stale and the final reads, and on a timeout records the pod path (neighbours, forwarding, flannel FDB, listener) before re-raising. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Conflicts in flake.nix (checks), hosts/coordinator and hosts/worker: both sides are additive (ax-fleet checks and myAxFleet vs DF-5/G1 checks and myGvisor.enable = false); kept both. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Sep 23, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ax on Agent Substrate across the fleet (nas control plane, coordinator gVisor harness pool), now with zero ax patches. Supersedes #464 (
ax/fleet-bringupf1bba47e), which is closed with a pointer here. Issue #463.What changed since #464 (
f1bba47e)0c3467c1mergeprobe/ax-fleet-zeropatch(6ada9d2e): stock ax v0.3.0.sandbox-class.patchandp1-completion.patchare gone;nix eval --json .#ax.patches=[](MEASURED on9cfc57dd). Substrate's four patches (0001to0004) stay unchanged.a3009eadfix round 4 fromax/fleet-r4-wipe82cb885withoutp2-egress-deny.patch: NASchain outputguard (only root andclusterClientUidsmay open new connections to pod/Service CIDRs and the seed registry), kubeletevictionHardnodefs/imagefs, WorkerPool ephemeral-storage 1Gi/32Gi,ax-fleet-guard-applyre-run after teardown, VM phases.e948734degress fail-open closed without an ax patch. Stock ax gives a Task with no gateway (or an unknown one){Host: "*"}.modules/ax-fleet/gateways.nixdeclares every atespace the fleet creates (fleet,default) with adefaultGateway that allows Halogen only (assertions forbid*,0.0.0.0/0,::/0and empty lists), bootstrap step80-ax-gatewaysapplies them, and theax-fleet-gateway-defaultenforcer points gateway-less Tasks at it. VM phase33-gateway-defaultproves it.5b90c6a7,716dfe5elink module vendored:pkgs/substrate-link(sourceax-conwipa021003apps/link,SYNC.mdnames the sha) and the NAS modulemyNas.substrateLinkfromconwip-link-nixos-on-ax-fleet-80658978.patch, gate OFF,completion = "guest"default (fails closed until a guest Complete relay exists),LoadCredential floor-link-token.aaf7433d,a6965fed,40a2ec87test fixes from 4-VM round 1: phase 33 variable rename (driver type check); re-apply (up to 3) instead of delete after a transientActorResumeFailedon a stale ateapi to atelet connection; rollback (a) re-reads the pod IP on each try.9cfc57ddmerge ofmainafter nas-topology: 8731 follows autoStart (Tom's 09-23 choice: close), plus the green fixes #466 and llm-agents 2026-09-23 (claude-code 2.1.280), drop the SessionEnd harvest hook, shut the coordinator's :8731 wifi door (#460) #461 (conflicts inflake.nixchecks andhosts/{coordinator,worker}were additive; both sides kept: ax-fleet checks andmyAxFleetplus DF-5/G1 checks andmyGvisor.enable = false).Evidence (paths under
/home/tom/today/wednesday-prep-2026-09-23/fleet/)checks.ax-fleet(4-VM) run 4 rc 0, 58 of 58 subtests on40a2ec87(MEASURED,receipts/r1/ax-fleet-4.log,receipts/r1/ax-fleet-4.receipt.json).checks.ax-fleet-bootrc 0 (receipts/r1/ax-fleet-boot.log). VM tests ran one at a time under a machine-wide lock.checks.ax-fleet=w4jfrm7b...-vm-test-run-ax-fleet.drvandchecks.ax-fleet-boot=nki40x0j...-vm-test-run-ax-fleet-boot.drvon40a2ec87and on9cfc57dd; both outputs present in the store (MEASURED,logs/land/04-drvpaths.txt,logs/land/05-topology-builds.log). So the run-4 result applies to this head.9cfc57dd:nix flake check --no-buildrc 0 "all checks passed!" (logs/land/03-landzero-9cfc57dd-flake-check.log);ax-fleet-topology,nas-topology,ax-client-topology,gvisor-module,home-profilesbuild rc 0 (logs/land/05-topology-builds.log).40a2ec87(receipts/r1/final-toplevel-*.log).PORT.md; round notes:STATE.mdhandoffs 1 and 2.Not in this PR / known
ax-fleet-smokestill waits forCompleted, which stock ax never writes; it must move to the floor-report shape before the switch-phase verify block relies on it.UpdateTask) before delete, like the test does (INFERRED from round 1, not yet inpkgs/substrate-link).Unknowns and proposed defaults
feat/substrate-wrap) targetsax/fleet-zero. Default: it is retargeted to main by its own lane after this merges; theax/fleet-zerobranch is kept, not deleted.🤖 Generated with Claude Code