Skip to content

ax on Agent Substrate across the fleet: one switch #463

Description

@mecattaf

Tom's ruling (2026-09-23, verbatim)

this is a dotfiles task to do on my nixos fleet. the decisions there were already made: hypervisor on NAS, agent harnesses on coordinator, halogen inference mainly on worker (can also run on coordinator if we need redundancy or a second parallel halogen task). i am certain that this can be knocked down in one issue>pr>nixos switch motion.

Placement

  • nas: k3s server, Agent Substrate (server at upstream d277088b), Postgres, RustFS object store, registry, Redis, ax-server and ax-controller.
  • coordinator: a k3s agent with the gVisor worker pool, where ax Tasks with coding agents land.
  • worker: Halogen at :8731, reached from sandboxes through the NAS egress gateway. The worker is not a cluster node in this motion.

Acceptance (post-switch)

  • kubectl get nodes shows nas and coordinator Ready.
  • Substrate is healthy.
  • An ax Task placed on the coordinator runs in gVisor (/proc/version shows 4.19.0-gvisor), reaches Halogen, and completes.
  • Four Tasks on the 2-worker pool never hit ResourceExhausted.
  • Rollback returns the coordinator to its prior generation with tailnet, Caddy, herdr and Halogen unaffected.

Evidence and design: ~/today/evals-2026-09-23/ax-fleet/ (DESIGN.md, INTEGRATE.md, REVIEW-LOG.md) and ~/today/evals-2026-09-23/substrate/.

🤖 Generated with Claude Code

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions