Conversation
`systemctl status | head -5` SIGPIPEd systemctl once head had its lines, and under writeShellApplication's pipefail every successful switch exited 141 (seen on the worker 2026-09-17). Print the unit's state with `systemctl show` instead. Verified on the coordinator: qwen38-27b and off both exit 0. Refs #397 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
halogen-switch: exit 0 on a successful switch
Tom, 2026-09-20: "having local kubernetes s3 or redis or postgres or whatever it needs on the NAS". This is that sentence for the three Substrate actually reads. State on the appliance, machines on the Strix boxes -- the NAS has 8 threads and 22 GiB and runs the house's DNS, so it holds the control plane's state and schedules nothing (Appendix J section 5). A SECOND DATABASE, not a second PostgreSQL. hosts/nas/media.nix already brought the instance up and pinned its dataDir to the NVMe; this rides it, so there is one instance and one backup story. ensureDatabases/ensureUsers with ensureDBOwnership, because ateapi runs goose migrations at startup and creates its own tables. enableTCPIP is forced by the deployment and not by taste: ateapi is a pod on a Strix box, the k3s server here sets disableAgent, so there is no unix socket to peer-authenticate over. scram-sha-256 from the LAN and from the pod CIDR, never trust. THE OBJECT STORE IS A HAND-WRITTEN UNIT ON PURPOSE. This host rides nixpkgs-stable (nixos-26.05) and stable has no rustfs at all -- no module, no package. The main pin has both; the package comes across the hosts/nas/unstable-pkgs.nix seam that attic-server and Immich already use, and the unit is modelled line for line on the unstable module's own serviceConfig. Delete it for services.rustfs when the NAS next rides a stable that ships it. Every unit carries RequiresMountsFor on /mnt/fast. hosts/nas/attic.nix learned the second edition of the signing-key trap: /mnt/fast is nofail, so a failed mount otherwise yields an empty state tree and a service that starts happily against it. For an object store that is silently lost snapshots. Refusing to start is a Tuesday. Secrets are runbook-placed root-owned files, not agenix, per attic.nix's no-agenix doctrine -- the rustfs key pair and the Substrate role's password are consumed by this host alone and a runbook can place them once. The password oneshot reads its file through LoadCredential and binds it as a psql value, so it lands in no argv and no journal line. Firewall: one block, three ports, iifname "enp1s0" only. Not openFirewall, which would publish all three on the tailnet as well. All three are on hosts/nas/cloudflared.nix's never-routed list. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Tom, 2026-09-20: "kubernetes is world-class for that ... they each stay in
their lane. Effects ts handles the ultracode-level json dag specification and
kubernetes schedules it on the right machines."
One shared modules/k3s-fleet.nix imported by all three hosts, because the
numbers have to agree and a second copy of them is how they stop agreeing.
Server on the NAS with disableAgent, so the appliance holds the control plane
and schedules nothing; agents on the Strix boxes, which have the cores, the
memory and /dev/kvm.
THE POD CIDR MOVES, AND THE ASSERTION IS OUTSIDE THE GATE. k3s defaults to
10.42.0.0/16 for pods and the house LAN is 10.42.0.0/24, which contains the
NAS, the coordinator, the worker, the printer and the router. Tom's 2026-09-09
research already picked the replacement: pods 10.200.0.0/16, services
10.201.0.0/16. Those are kept exactly, and the overlap check is real
arithmetic in a real assertion placed OUTSIDE the enable gate, so it
evaluates on every host whether k3s is on or off. Verified against six cases,
including that k3s's own default IS caught.
TRACK K AND TRACK S ARE THE SAME FILE, and that is the point. Design K is
plain k3s plus Cilium plus the gvisor RuntimeClass; Design S adds Substrate on
top of exactly this. The only thing S needs from the bottom layer that K does
not is four feature-gate settings, and a gate nothing asks for is inert: no
scheduling change, no new controller, no memory. So the gates are in
unconditionally and there is one PR instead of two that drift.
The gate names are upstream's own, MEASURED from ~/Downloads/substrate,
hack/create-kind-cluster.sh:104-112: ClusterTrustBundle,
ClusterTrustBundleProjection, PodCertificateRequest, plus runtimeConfig
certificates.k8s.io/v1beta1. The kubelet is handed only the two that are its
own, because an unrecognised gate name is fatal to kubelet and
ClusterTrustBundle is apiserver-side. Whether k3s 1.35.6+k3s1 accepts them at
all is U9 and is not claimed here.
Cilium 1.18.14 through autoDeployCharts, so the CNI lives in the generation
rather than in a `cilium install` somebody has to remember. The chart hash is
measured, not fakeHash: built with fakeHash, took the reported hash, rebuilt
clean. The gvisor RuntimeClass is a server manifest and its containerd runtime
is an agent containerdConfigTemplate that keeps `{{ template "base" . }}` --
one line, load-bearing, and dropping it costs the node its CNI, its
snapshotter and its registry mirrors at once. kata is written out and
commented, pointing at tonight's U14 report.
The kube API is LAN-only: 6443 on the NAS's enp1s0 and nothing else, plus the
never-routed doctrine in the cloudflared PR, because a tunnel ingress bypasses
nftables entirely and "firewalled" is therefore a separate guarantee from
"not in the tunnel".
All three hosts land with the gate OFF. secrets/k3s-token.age does not exist;
minting it needs Tom's admin age key.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
A5a stood the pinned k3s up inside a throwaway guest built from this exact nixpkgs revision the same night and measured three things this file had wrong. Report: ~/today/review/2026-09-20/sandbox-spike/U9-U13-U15-K3S.md. 1. THE THREE GATES ALONE SERVE NOTHING. Boot 5 held the gates on and dropped the runtime-config flag: all three kubernetes_feature_enabled metrics read 1 while /apis/certificates.k8s.io listed only v1 and api-resources returned NO_RESOURCES. A check that read only the metric would have reported a false pass. --kube-apiserver-arg=runtime-config=certificates.k8s.io/v1beta1=true is what turns the group version on, and podcertcontroller has nothing to talk to without it. It now sits beside the gates, and the flag values inside a -arg= carry no leading dashes of their own, which is A5a's measured working form. U9 is therefore ANSWERED YES on the pin: v1beta1 serves clustertrustbundles and podcertificaterequests. Appendix J section 9's fallback does not have to be taken and no newer k3s is needed. kubelet also accepts all three gate names, so the cautious subset this file used to hand it is gone. 2. THE NIXPKGS EXAMPLE'S CONTAINERD KEY PATH IS WRONG FOR THIS k3s. The bundled containerd is 2.2.5-k3s2 and the config it generates starts `version = 3` with runtimes under [plugins.'io.containerd.cri.v1.runtime'.containerd.runtimes.runc]. A grpc.v1.cri block would be parsed, accepted and silently ignored, which is the worst of the three outcomes. A5a's exact working table is used here. And runtime_path, NOT options.BinaryName. The BinaryName shape this file used is a trap that looks right: a pod under it reaches Running and stays 1/1 Running in kubelet's view, produces no logs at all, and kubectl exec fails with "in state stopped". The generic runc shim starts runsc but carries neither its stdio nor its state. containerd-shim-runsc-v1 is what goes here, and nixpkgs' gvisor builds it. 3. k3s DOES NOT FIND runsc ON ITS OWN. Not by auto-detection (with gvisor in systemPackages the generated config still carried only runc and runhcs-wcow-process) and not from the unit PATH, which the nixpkgs rancher module leaves empty but for zfs. Without the path line every sandbox dies at creation with `exec: "runsc": executable file not found in $PATH`. The line was already here; it now carries why, and A5a's caution that the assignment replaces rather than extends. Also corrected: this branch's earlier claim that identical gate-off derivation paths proved the module inert. They do not prove it. The flake embeds the configuration revision, so the toplevel tracks the commit. The sound check is that with the gate off, services.k3s.enable is false, no 6443 rule is emitted and no registries.yaml exists on any of the three hosts, which was measured directly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
pkgs/ax builds google/ax v0.3.0 pinned by commit d8ed0fe (tag v0.3.0), wired through overlays/default.nix and the flake's packages.<system> list like every sibling package. modules/ax-client.nix puts kubectl and ax into environment.systemPackages on coordinator, worker and client when myAxClient.enable is true; it lands false on all three. Refs #453. The one judgement call: ax's go.mod opens `go 1.27.1` and Go refuses to build a module whose go directive is newer than the running toolchain (`go: go.mod requires go >= 1.27.1 (running go 1.27.0; GOTOOLCHAIN=local)`), with no network in the sandbox to fetch one. MEASURED 2026-09-23: nixpkgs has go_1_27 = 1.27rc2 and nixpkgs-fresh has 1.27.0, so neither pin already in this flake can build it. Adds one input, nixpkgs-go, pinned BY REVISION and supplying exactly one attribute to exactly one package - the same shape as the existing nixpkgs-paperless. Bumping nixpkgs-fresh instead was rejected on purpose: that input also carries the twins' linux 7.2 kernel. checks.ax-client-topology asserts the option exists on the three hosts that import the module, that it is false on all three, that kubectl and ax are absent from all three systemPackages, and that hosts/nas carries no such option at all. It goes red on the flip by design. No switch, no rebuild, no gate flipped, no cluster, no Substrate, no patches - this builds pristine v0.3.0. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adds `string sandbox_class = 11` to TaskSpec, regenerates ax.pb.go with the same protoc-gen-go v1.36.11 upstream used, refuses unknown values in ValidateTask, and threads the value from the reconciler through BuildActorTemplate, replacing the SandboxClass_SANDBOX_CLASS_GVISOR hardcode at internal/substrate/client.go:273 (the only SANDBOX_CLASS occurrence in the tree). Empty means gvisor, so existing manifests behave exactly as before. The generated Go is in the patch on purpose: ax bridges YAML through protojson with unknown fields rejected, so a .proto-only edit would make every manifest naming sandboxClass fail strict decode. This does NOT deliver a workerd sandbox. Agent Substrate's SandboxClass enum has three members (UNSPECIFIED, GVISOR, MICROVM), so "gvisor" and "microvm" are the only values that can reach a real substrate. A workerd class needs an upstream Agent Substrate change that does not exist. vendorHash is unchanged: the patch touches no go.mod or go.sum line. Measured on the patched scratch tree with go 1.27.1: `go build ./...` rc 0, `go test ./...` rc 0. In this clone: `nix build --no-link .#ax` rc 0, with the patch applying to all six files and checkPhase passing. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
home/ax-conwip.nix declares `myAxConwip`, a home-manager module that WOULD run
the CONWIP scheduler at /home/tom/mecattaf/ax-conwip as one long-running user
service against a declared ax server, with the rewrite's seat meter directory as
a read-only input. It is imported from home/home.nix and gated OFF, so it defines
no unit on any host today.
Seven options, every one of them data in this repository rather than a command
line somebody types: enable (false), serverUrl (127.0.0.1:8080, the loopback the
mock stack listens on, so an accidental enable reaches nothing real), metersDir
(the REWRITE's meters, ~/.local/state/tally-rewrite/meters, never branch (a)'s,
read only, no tmpfiles rule over it), wipCap (1, stricter than the program's own
default of 2), sourceDir, recordsDir (a path that does not exist, so an
accidental enable fails legibly instead of passing over an empty directory) and
stateDir.
Four things deliberately not done, and the module header says all four out loud:
- no flake input, because the CONWIP repository has NO REMOTE. A fetchGit of a
path on one box is not a declaration, it is a machine-local accident that
would break every other host's eval;
- no package, because packaging follows the input. The unit runs the checkout
in place through `pnpm exec tsx`. That is a development shape, not a
delivered one, and it is the biggest single reason the gate stays off;
- no timer: the release signal is a WatchTask stream, not a poll;
- nothing enabled and nothing armed. `enable` is set nowhere, and even with
the gate flipped the unit carries NO Install section, so a rebuild declares
it and a `systemctl --user start` is a second, separate act.
checks.x86_64-linux.ax-conwip-topology asserts the option EXISTS on all three
home-manager hosts (a dropped import makes it undefined, not false, which is the
mutation hint), that it is false on all three, that no ax-conwip service, timer
or socket is rendered anywhere, that there is no system-bus twin on any of the
four hosts, and that no tmpfiles rule naming ax-conwip exists while the gate is
off. The `enable == false` line goes red on the flip, deliberately.
MEASURED in this tree:
nix build --no-link .#checks.x86_64-linux.ax-conwip-topology -> rc 0, so every
assertion above ran and passed
nix eval .#nixosConfigurations.{coordinator,worker,client}.config
.home-manager.users.tom.systemd.user.services --apply builtins.attrNames
-> rc 0, no ax-conwip on any of the three
the same eval under extendModules with the gate flipped renders the unit with
no Install section, so the OFF branch is not a dead branch
docs/ax-conwip.md carries what the CONWIP is, what the unit would run, the
nine-step sequence to enable it, and the line that says the repository has no
remote so nothing is packaged.
Nothing here is enabled, nothing was started and nothing was switched. The flip
is a separate act and it is Tom's.
Refs #455 (the gate's own pre-existing bug: nix flake check does not reach a
verdict on this repository).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ers directory Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…' into ax/fleet-bringup
…/fleet-bringup # Conflicts: # hosts/nas/default.nix
# Conflicts: # hosts/worker/default.nix
Stock ax v0.3.0 never learns that a Task's command exited, so the Task stays
Running and its actor keeps a worker. The 2026-09-23 Substrate probe measured
the consequence: finished Tasks hold every worker of a small pool and the next
Task fails ResourceExhausted.
p1-completion.patch (on top of sandbox-class.patch, both now under
pkgs/ax/patches/):
- runner serves /metadata/v1alpha1/ax/{exit,result,usage}; result capped 1 MiB
- controller: terminal guard, exit read after resume, --running-resync (15s),
Completed / Failed ExitCode=N write-back, result and usage copy, then
SuspendActor; a CRASHED actor with no exit report is Failed ActorCrashed
- API: TaskStatus.command, UsageStats.tool_calls, rpc GetTaskResult and
`ax result task <name>`; generated Go regenerated
Tests run in checkPhase (doCheck stays true), including a four-Task floor
test on a 2-worker mock pool and the no-exit-report contrast. vendorHash and
the vendored Substrate client pin 672533541dbf are unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…asset buildGoModule with the nixpkgs-go toolchain (go 1.27.1), vendored tree, CGO off, version ldflags d277088b: ate-setup plus ateapi, atecontroller, atelet, atenet, podcertcontroller and ateom-gvisor. images.nix: the six component images as OCI layouts with a digest file, the eight third-party images the kind install pulls as digest-preserving fixed-output copies of their linux/amd64 manifests, and the gVisor nightly tarball SandboxConfig gvisor-default names (sha256 d547d814). Patches touch manifests only: 0001 points pauseImage at the NAS registry through atelet's localhost rewrite; 0002 pins third-party images to their linux/amd64 child digests so the NAS store holds about 0.7 GB instead of every platform. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…mage OCI layouts with a `digest` file (pkgs/ax/oci-layout.nix), the shape myAxFleet.registrySeed takes. The digest is read at build time; nothing imports from a derivation. - pkgs/ax/images.nix: the patched ax's server and controller (uid 65532, cacert) and nixpkgs redis for ax-redis. - pkgs/ax-agent-image: ax-task-runner at /usr/local/bin (PID 1), pi from llm-agents, and the ax-agent adapter with modes halogen-smoke, pi (the probe's adapter-pi.sh and validate.py), fetch URL and exit N. pi-models.json names Halogen only, no secret; the endpoint comes from $HALOGEN_URL at run time so one digest serves every host. No claude-code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…modules/ax-fleet) One module, one kill switch per host (myAxFleet.enable), three roles: control (nas: k3s server with its kubelet, untainted; registry; seed and bootstrap runners), harness (coordinator: k3s agent tainted ate.dev/sandboxClass=gvisor:NoSchedule, labelled with the Substrate version at registration; guard chain; NetworkManager conf.d drop-in and a config reload, never a restart), inference (worker: one assertion). Folds #447's CIDRs, assertions, feature gates and runtime-config; drops its Cilium, containerd template and runsc RuntimeClass (Substrate runs its own runsc). Keeps #446's registry only, on the data pool. flannel VXLAN and kube-proxy bound to the LAN leg. k3s state bind-mounted from /mnt/fast, PersistentVolumes under /mnt/nas/services/ax-fleet/local-path. pkgs/ax-fleet-teardown wraps the pinned k3s-killall.sh, removes the guard chain and restores the sysctls snapshotted before k3s first ran. substrate.nix and ax.nix are empty for their tracks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…447 gates nas = control, coordinator = harness (myAxClient on by mkDefault), worker = inference (renders one assertion; not switched in this motion). Deletes modules/k3s-fleet.nix (with its RuntimeClass manifest) and hosts/nas/state-services.nix, superseded by modules/ax-fleet. The k3s-token comment in secrets.nix no longer claims the ciphertext is missing. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eet-teardown tests/ax-fleet: nas, coordinator, worker (Halogen stub), peer (tailnet stand-in); base configs with ax off and an ax-on specialisation switched live, NAS first. Phases 10-cluster and 90-rollback are the cluster track's; 20-substrate and 30-ax are empty for their tracks. Receipt in $out/receipt.json. ax-fleet-topology asserts the real hosts' rendered flags, firewall scope, the NetworkManager.conf no-restart property, the kill switch (mkForce false renders nothing) and parity with the VM nodes. ax-client-topology now expects myAxClient ON on the coordinator only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… test modules/ax-fleet/ax.nix writes only the myAxFleet extension points and the coordinator's ax scripts: - manifests.ax-fleet-40-ax: Namespace ax-system; ax-redis (AOF on a local-path PVC, no password, nothing on argv); ax-server on ClusterIP 10.201.0.80, never a NodePort; ax-controller with upstream's projected token and ClusterTrustBundle, --running-resync (P1), ATENET_ROUTER_ADDR and AX_SNAPSHOTS_BUCKET=gs://ate-snapshots/ax/. Every image by digest, the digests substituted at build time and checked (no IFD). - registrySeed: ax-server, ax-controller, ax-redis, ax-agent. Images come from the flake's own nixpkgs and `ax`, never the host's pkgs, so the NAS seeds the digest the coordinator names. - harness: ax-fleet-image-ref, ax-fleet-smoke (halogen, pi, exit N, egress-deny, floor N; JSON receipts) and AX_SERVER=http://127.0.0.1:8080. - options myAxFleet.ax.{runningResyncSeconds,atespace}. tests/ax-fleet/phases/30-ax.py: T1 to T5 and the resilience checks. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…with ate-setup Fills only the myAxFleet extension points, on the control role: - registrySeed: six component images (substrate/<name>:d277088b) and the eight third-party images under their upstream repositories; - bootstrap 20-registry-svc: Namespace ate-system and the kind-registry Service plus EndpointSlice to the NAS registry, which atelet's --localhost-registry-replacement names; - 30-substrate: upstream ate-setup --kind deploy ate-system with --image-repo <registry>/substrate --image-tag d277088b, once per (installer, images) closure, stamped in kube-system/ax-fleet-substrate; - 40-gvisor-asset: bucket gvisor in the in-cluster RustFS holds the pinned runsc tarball at atelet's fallback key, sha256-verified; the upstream kind credential is read from the Deployment and kept off argv; - 50-workerpool: WorkerPool ateom-gvisor, replicas and memory limit from the options, harness nodeSelector and taint toleration, by digest. tests/ax-fleet/phases/20-substrate.py: the Substrate subtests of the 4-VM check. images.nix: drop an unused argument (deadnix). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…losure The bootstrap referenced the full substrate output (seven binaries, 377 MB) and the patched source (172 MB) at run time. It now runs passthru.ate-setup (53 MB) from passthru.installTree (go.mod, manifests/, hack/: 1.4 MB), which is everything deploy ate-system reads. The component binaries reach the NAS only inside the images. Saves about 0.5 GB on the NAS's eMMC root. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- ax-fleet-kubeconfig: quote the ssh host (shellcheck SC2086 failed the coordinator toplevel build). - tests: wait out k3s's transient node taints and the flannel.1 link instead of asserting once; A-record DNS query (the stand-in has no upstream, AAAA is REFUSED); a cross-node pod-to-pod subtest over VXLAN; typed receipt (test-driver type check); records also printed to the log. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # modules/ax-fleet/substrate.nix # tests/ax-fleet/phases/20-substrate.py
# Conflicts: # modules/ax-fleet/ax.nix # tests/ax-fleet/phases/30-ax.py
…, prove it in the VM test Merges ax/fleet-cluster (cc39eac), ax/fleet-substrate (1165567) and ax/fleet-ax (71ce913); the four add/add placeholder conflicts take each track's own file. The gates stay ON in hosts/{nas,coordinator,worker} (control, harness, inference). - pkgs/ax-agent-image: pi's bun binary lists PT_LOAD out of p_vaddr order. Linux runs it; gVisor refuses it (exit 126 inside the sandbox, MEASURED in the VM test). elf-sort-load.py reorders the program header table only, so the Task image's pi runs in gVisor (T2 schema-valid, MEASURED). - pkgs/ax-agent-image: a `probe` mode (claude --version when the image has claude-code, a GET of Halogen's /v1/models, /proc/version); pi mode records its stderr tail; test-only `extraPaths`/`variant` arguments. - ax-fleet-smoke: --image, --sandbox-class and the `probe` case; an exited command with protobuf-omitted exitCode reads as 0. - tests/ax-fleet: phase 25-harness-probe (one sandboxClass gvisor Task on the coordinator in a test-only image variant with claude-code and no credential: Completed, sampled Completed for 60 s); the NAS test node seeds that variant; test nodes import modules/ax-client.nix as the hosts do; the Halogen stub answers a JSON Schema prompt with {"answer": 42}; 20-substrate checks the pod-certificate controller in its own namespace and stops shadowing the driver's `step` and `log`. - flake: packages ax-agent-image and substrate-images. checks.x86_64-linux.ax-fleet (4 VMs): rc 0, 37 subtests. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ngup The coordinator runs configurationRevision 3a65899, which is not in origin/main. Stacking on it keeps claude-code 2.1.280, the closed 8731 wifi door (#460), the theme switcher and the harvest-hook removal, so a switch from this branch adds only the ax-fleet delta. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…(DF-6) a775d58 (remove the SessionEnd harvest hook) also carried the home-profiles asserts from f1d689b (#448, parakeet socket activation), which this branch does not contain. The coordinator home here still has parakeet-service with Install.WantedBy = [ "default.target" ] and no socket, so checks.home-profiles failed with `attribute 'parakeet-service' missing` at flake.nix:2036. Restore main's asserts for this check. The socket asserts belong in #448 and return when it merges. No module or host change; only the check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…sure, panic tunables, RBAC, RustFS credential, real flap) - Stacked on the live coordinator revision 3a65899 (merge commit before this one) plus f7b34239 (DF-6, check-only), so a coordinator switch adds only the ax-fleet delta: same home-manager generation as live, no tcp 8731 rule, no harness version rollback. - ax-fleet-teardown: inheritPath = false, refuse if tailscale is on PATH, so k3s-killall.sh never runs `tailscale set --advertise-routes=` (the NAS's 10.42.0.0/24 route). Also drops the NAS guard table. - NAS: docker-registry waits up to 30 s for 10.42.0.1 and restarts on failure; the seed wants (not requires) it and restarts on failure; k3s waits for the LAN address on both nodes. The boot test now adds the address 40 s late and proves the registry recovers. - NAS: inet ax-fleet-guard prerouting at priority raw drops routed LAN traffic to the pod and Service ranges before kube-proxy DNAT. - Coordinator guard: flannel.1 payloads must come from the pod CIDR; pod traffic out the LAN leg is dropped for every private range. - kubelet's kernel.panic, panic_on_oops, overcommit_memory are put back after each kubelet start (myAxFleet.kubelet.keepHostKernelTunables, default true, Tom's ruling pending). - ax-controller: no ClusterRole on Secrets, no automounted API token. - Substrate patch 0003: RustFS, ate-api-server, atelet and the bucket-init Job read a per-cluster credential from Secret ax-fleet-rustfs, generated once on the NAS by bootstrap step 25-rustfs-secret. - k3s bind mounts unmount lazily, so the rollback switch no longer fails on a busy /var/lib/kubelet. - VM test: the flap waits for Ready=Unknown and the unreachable taint; new subtests for the routed-LAN paths, the Freebox-style foreign /24, the tunables restore, the RBAC and credential surface, and the teardown never calling tailscale. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…en tunables, conntrack, teardown on PATH, read-only registry, split k3s token, NAS parity, boot snapshot, receipt) - harness: refuse NEW connections from cni0/flannel.1 at the head of nixos-fw (v4 and v6); user.slice and system.slice CPUWeight 10000 (deskCpuWeight). - k3s: kernel tunables watcher bound to k3s (no 900 s bound); boot snapshot takes non-ax declared values (NIXOS_ACTION unset at boot); kube-proxy conntrack args 0; conntrack keys snapshotted; teardown on PATH for the control and harness roles whatever the switch says; server and agent tokens split (agentTokenFile on the NAS). - control: registry read-only (maintenance.readonly, delete off); the seed pushes through a loopback writer that lives only for the seed run. - substrate: images built from inputs.nixpkgs, one digest set everywhere. - secrets: k3s-token.age re-minted for editors ++ nasOnly, new k3s-agent-token.age for editors ++ coordinatorOnly ++ nasOnly. - tests: ax-fleet-boot on nixpkgs-stable with the systemd initrd and the NAS kernel; VM coordinator on the desk kernel with sshd; new assertions; every phase subtest recorded in receipt.json. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…pod egress, role-scoped guards through the kill switch, proxy ARP by interface, API owner match, proxy off :8080, bounded tracing) - harness guard: pod egress policed by destination on any interface, nothing enters cni0/flannel.1 but the NAS host; enp191s0 is a LAN leg (lan.extraInterfaces) with NetworkManager route metric 700 - control: forward chain in inet ax-fleet-guard; NAS pods reach only Halogen's host:port among private addresses (the Gateway port is ignored upstream), never the tailnet or ve-* - guards, pod-input refusal and API owner match render for the role whatever enable says; the teardown leaves declared guards - proxy_arp on cni0, flannel.1, veth* by name; conf.default untouched - ax-server-proxy on 127.0.0.1:8099 as a static user; OUTPUT owner match admits only root, apiUsers and the proxy to it and the cluster ranges - substrate patch 0004: jaeger max-traces and limits, collector memory_limiter, limits, debug basic - VM: eth3 DHCP leg, worker :2222, T4b, proxy ARP, API owner, rollback pod-to-host probes; topology asserts Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Owner
Author
|
Superseded by #467 ( |
mecattaf
added a commit
that referenced
this pull request
Sep 23, 2026
ax on Agent Substrate across the fleet, zero ax patches (supersedes #464)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #463. Supersedes #446, #447 and #454 (left open for you to close).
Draft: round 4 of review found one high-severity item that is not fixed (below). Merge and switch are yours.
What it declares
myAxFleet.role = control): k3s server, Agent Substrate (server at upstreamd277088b; ax keeps its vendored client pin), Postgres, RustFS, registry, Redis, ax-server and ax-controller. Bulk state lives on the data pool, not the eMMC root.harness): k3s agent, gVisor worker pool (2 workers), ax client and proxy, and a deny-by-default LAN guard. The CPU weight is set so the cluster does not crowd the desk.inference): no cluster code. Sandboxes reach Halogen:8731through the NAS egress gateway.sandbox-class.patchandp1-completion.patch.Evidence (all at this head,
f1bba47e)nix flake check --no-build: all checks pass.checks.ax-fleet-topologypasses.checks.x86_64-linux.ax-fleet: the 4-VM test (nas, coordinator, worker, tailnet peer) passes 50 subtests.nix evalof its outPath at this head is/nix/store/6d586q2qzm4y3fm97ci467g4q9q42379-vm-test-run-ax-fleet, the green round-3 run.~/today/evals-2026-09-23/ax-fleet/REVIEW-LOG.md.Switch runbook
Base. Merge this onto
chore/llm-agents-bump-drop-harvest-hook(#461), which carries the live coordinator revision3a658991. Basing it onmainwould roll back #461 and #460.Order. nas, then worker, then coordinator.
Verify after the switch:
kubectl get nodesshows nas and coordinator Ready.ax-fleet-smoke probe: the Task shows a gVisor/proc/versionand a Halogen reply.ax-fleet-smoke floor: 4 Tasks with no ResourceExhausted.Note. The ax API proxy now listens on
:8099, not:8080. Use:8099in any manual check (see open items).Rollback. Switch back to the previous generation, then run
ax-fleet-teardown. There is also a per-host kill switch.Open items found in review round 4 (not fixed here)
ax/fleet-r4-wip(e82cb885, not pushed). Proposed default: refuse dispatch when the gateway is missing, in the link or CONWIP layer, rather than in ax.--eviction-hardreplaces k3s's disk-pressure defaults.flannel.1guard and pod-input rules.127.0.0.1:8080check should be:8099.Planned follow-ups (decided after this branch was built)
sandbox-class.patch: stock ax already writes gVisor for every Task. MEASURED: 4-VM pass on branchprobe/ax-fleet-nosc.p1-completion.patch: the job reports to the floor, then the link deletes the Task (the Buildkite pattern).probe/ax-fleet-zeropatch).~/today/evals-2026-09-23/zero-patch/DELETE-HANG-2026-09-23.md.Unknowns and proposed defaults
🤖 Generated with Claude Code