Repository navigation
substrate: coordinator pusher and puller as user units, ON; pkgs/substrate-apps at fc2f8bd; tokens sealed - #465
Merged
Merged
Conversation
…ubstrate dc7cd1d, one pnpm package set pkgs/substrate-apps: the box-side programs of the Cloudflare Substrate as one package set (link, pusher; puller pending, absent upstream at this sha). Source taken with git archive dc7cd1d05b1e3938dace9d3d3cddc1a22d98d6cc (package.json, pnpm-lock.yaml, pnpm-workspace.yaml, tsconfig*.json, apps/link, apps/pusher), unchanged; SYNC.md records the sha and what is not vendored; sync.sh re-vendors from a sha and picks up apps/puller once it exists. link is bundled with esbuild exactly as pkgs/substrate-link (a021003) is; apps/link/src is identical between the two shas (MEASURED git diff: tests, tsconfig and vitest config only). pnpmDeps is fixed-output, fetcherVersion 3 (this nixpkgs pin removed 2), hash from a failing build (MEASURED). pusher needs no node_modules (its imports are node: builtins and its own src), so it is installed as files and its checkPhase runs one --dry-run --once tick against the vendored seat-capacity/1 fixture. Exposed as .#substrate-pusher and .#substrate-apps-link. pkgs/substrate (Agent Substrate, Go) is a different project and stays where it is, which is why this set is not at pkgs/substrate/. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…declared OFF
modules/substrate.nix: the coordinator side of the Cloudflare Substrate.
Two systemd user units for Tom's manager (ConditionUser), both gated off:
pusher substrate-pusher --config <store pusher.json> --token-file
%d/floor-token. seats --json --no-spend under the gentle policy,
posted to POST /capacity/snapshots. Seats cc, cc2, codex,
pi-qwencloud, halogen; gpu-worker published as halogen; cc3 not
dispatchable; owners.codex = tom on the coordinator (E6).
puller the interpreter host. PENDING its package (apps/puller is not
upstream yet), so package defaults to null and arming asserts.
The unit renders runtimes.toml (host, herdr, gvisor with nix
runsc and pasta, ssh:worker as pi on halogen; default host;
[seats] claude = cc2) into AX_CONWIP_RUNTIMES, a substrate.json
into its RuntimeDirectory with capacityFloorTokenFile pointing at
the credential, and CLAUDE_CONFIG_DIR=/home/tom/.claude-work
(cc2, by path). An optional runtime-test wrapper around ExecStart.
The bearer is an agenix secret NAME (substrate-floor-token, owner tom,
0400) handed to each unit by LoadCredential; no rendered file or unit line
carries a value. hosts/coordinator imports the module and sets floorUrl
https://substrate.mecattaf.dev with both enables false.
tests/substrate-modules (checks.substrate-modules): eval-time assertions
that the gates are OFF and no unit or secret is rendered on the real host,
then the same host armed through extendModules with fixtures, whose
pusher.json and runtimes.toml the check parses (TOML decodes, runtime
types known, no token-shaped key).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…er unit on the real [puller] config
Re-vendored from agency-agency/substrate origin/main b1051790f376c103ba4e901619ba11efeb78cef6
(017893d added apps/puller, the coordinator-side interpreter host; b105179 is
docs). The vendored set now carries src/, proto/, packages/* and deploy/ as
well, because the puller runs through tsx's ESM loader over the workspace
(substrate root, @substrate/{api,link,interpreter,runners}); apps/floor,
apps/cli, apps/mcp, test/, tools/ and docs/ stay out. sync.sh's path list
grew accordingly. The pnpm hash did not change (MEASURED).
pkgs/substrate-apps.puller (`.#substrate-puller`): pnpm install --prod from
the shared pnpmDeps, the tree with its production node_modules copied to
$out/lib/substrate-apps, node wrapped over apps/puller/bin/substrate-puller.mjs
(86M, closure 327 MB, MEASURED). checkPhase starts it with no config and
requires exit 78 with a config-invalid line, so every import resolved and
main() ran.
modules/substrate.nix puller: package defaults to the real puller; its
configuration is the [puller] table of a client config.toml
(packages/api/src/config.ts): holder coordinator, link_token_file, state_dir,
pidfile, max_runs 1, node_dispatch local, runtimes (the rendered TOML),
seat cc2, cap 2, default_model, node_runs_on, optional ax_server. The unit
renders config.toml and substrate.json into its RuntimeDirectory at start
with the two credential paths (floor-token, link-token: a second agenix
secret name, substrate-link-token-coordinator) and exports
SUBSTRATE_CLIENT_CONFIG, SUBSTRATE_CONFIG, AX_CONWIP_RUNTIMES. The INFERRED
command-line flags of the previous commit are gone.
tests/substrate-modules: asserts both real packages on the declared host
(evaluated, not built), the two LoadCredential entries, and parses the
rendered config.toml (placeholders in place, seat cc2, holder coordinator,
no /run/agenix or /run/user path in any rendered file).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…tor for the coordinator (editors ++ coordinatorOnly)
…odexSandbox, peerCacheDir); packages build Vendored src re-archived from agency-agency/substrate fc2f8bd (6f681d3, the checkout the live coordinator puller and pusher run from, plus two floor-only commits). pnpmDeps.hash is unchanged (measured with lib.fakeHash: the lockfile only gains the apps/evaluator importer). The puller check still sees exit 78 with a config-invalid line; its log tail is widened to 40 lines so an import failure is readable. SYNC.md records what changed for the box: seat routing, [credentials.seats]/seat_dirs, codexSandbox, peerCacheDir, the new [puller] keys demand_dir, capacity_wait_s, drain_timeout_s, health_addr, and that new upstream files must be git-added before the flake sees them. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…e 2026-09-24 proven deployment (opus/halogen/codex/codex-rw, credentials.seats, peerCacheDir); secrets declared - modules/substrate.nix: runtimes default = opus/halogen/codex/codex-rw host runtimes (herdr, gvisor, ssh:worker no longer declared by default), [credentials.seats] cc/cc2 (RG-1); [puller] gains demand_dir, drain_timeout_s, capacity_wait_s and an optional health_addr; pusher gains peerCacheDir (default "inherit") and gpu-coordinator in not_dispatchable; puller RestartPreventExitStatus 3/75/78, TimeoutStopSec above the drain. Header documents the ON gate, the no-runtime-test choice and the pidfile cutover step. - hosts/coordinator: pusher.enable and puller.enable = true. - secrets.nix: substrate-floor-token and substrate-link-token-coordinator, editors ++ coordinatorOnly. - tests/substrate-modules: asserts the declared host ON with both age files named, and the proven shape of all three rendered files. - docs/ax-conwip.md: dated note pointing at the substrate units. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
NixOS renders a unit's `path` as Environment=PATH=..., which replaces the user manager's PATH instead of extending it. The built coordinator closure showed both units with store-only PATHs, so the puller's host runtimes (bare `claude`, `pi`, `codex` in packages/runners/src/harness.ts) and the pusher's oracle (`seats` spawns bare `codex` for the codex seat) would fail with ENOENT. Both units now append ~/.local/bin, /etc/profiles/per-user/<user>/bin and /run/current-system/sw/bin, in the manager's order. The substrate-modules check asserts this for the armed and declared units. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…it cutover order The puller counts an EPERM pid as alive, so a pidfile on persistent disk left by an unclean exit (or the namespace pid 2 left by the hand-started run) could make every start exit 3, and RestartPreventExitStatus=3 left the unit dead. The pidfile now defaults to $RUNTIME_DIRECTORY/puller.pid (placeholder filled at start), which systemd clears on stop and at boot, and 3 is retried. The cutover comments now say: SIGTERM both nohup processes by HOST pid, rm -f puller.pid, pusher.pid, puller.host.pid and puller.supervisor.pid, then switch (a unit pusher treats pid 2 as stale and would otherwise run beside the nohup pusher). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…ectory in the substrate checkout, MemoryHigh/Max 2G/4G, session sockets and CREDENTIALS_DIRECTORY dropped from agents' environment (review 2026-09-24 medium/low)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Substrate's coordinator side in dotfiles: pkgs/substrate-apps, modules/substrate.nix (both gates OFF)
Charter: PLAN A2 and ruling E1 (2026-09-23): the NAS link, the coordinator puller and the capacity pusher live in dotfiles as packages and modules, vendored from agency-agency/substrate at a pinned sha. Touches nothing under
pkgs/ax,modules/ax-fleet, orhosts/beyond the coordinator's import and an OFF declaration.Target and base
ax/fleet-zero(branched ata6965fed; Lane A's head is now40a2ec87, one commit later,tests/ax-fleet/phases/90-rollback.pyonly). Test merge of this head onto40a2ec87: clean (MEASUREDgit merge-tree --write-treerc 0, 2026-09-23T21:38Z).ax/fleet-zerois not inmain(MEASUREDgit merge-base --is-ancestorfalse;main=7d4704db, after Lane A's nas-topology: 8731 follows autoStart (Tom's 09-23 choice: close), plus the green fixes #466 merge). This PR retargets tomainafter the fleet stack (Lane A) lands. Lane A owns the merge and the switch tonight; this lane does not merge.main7d4704dbconflicts inflake.nix,hosts/coordinator/default.nix,hosts/worker/default.nix;ax/fleet-zero40a2ec87alone conflicts withmainin the same three files (MEASUREDgit merge-tree --name-only, 21:39Z), so the conflicts are the fleet stack against nas-topology: 8731 follows autoStart (Tom's 09-23 choice: close), plus the green fixes #466, not this PR's commits. Once the fleet stack is inmain, this PR should merge clean.What is vendored, from which sha
pkgs/substrate-apps/:apps/link,apps/pusher,apps/pullerand the workspace they import (src/,proto/,packages/*,deploy/, rootpackage.json,pnpm-lock.yaml,pnpm-workspace.yaml,tsconfig*.json) taken unchanged withgit archivefrom agency-agency/substrateb1051790f376c103ba4e901619ba11efeb78cef6("docs: DEPLOY.md ..."; the puller landed one commit earlier in 017893d). 303 files, 3.2 MB.pkgs/substrate-apps/SYNC.mdrecords this;pkgs/substrate-apps/sync.sh <checkout> <sha>re-vendors. One fixed-outputpnpmDeps(fetcherVersion 3). First vendored at dc7cd1d (link and pusher); resynced the same night whenapps/pullerlanded.test/,tools/,docs/,apps/floor(Cloudflare Worker, deployed from the substrate repo),apps/cli,apps/mcp.0aaf6a2("capacity: ask the floor seat-wide for a model that names no family, so codex nodes are admitted"): it changessrc/capacity/gate.ts(the floor's admission gate, run in the Worker) and adds a test. The three vendored coordinator and NAS programs do not execute that path, so no resync tonight; a follow-upsync.shpicks it up (MEASUREDgit show --stat).pkgs/substrate/because that is Agent Substrate (Go).pkgs/substrate-link(a021003) andhosts/nas/substrate-link.nixare untouched (the NAS side;apps/linkat b105179 differs from a021003 only in tests and tsconfig, MEASURED).The two services (
modules/substrate.nix), both gated OFFservices.substrate.pusher: systemd user unitsubstrate-pusherin Tom's manager, runssubstrate-pusheron a timer-less loop againstfloorUrl; bearersubstrate-floor-tokenviaLoadCredential, renderedpusher.jsoncarries paths only.services.substrate.puller: systemd user unitsubstrate-puller; renders the[puller]table of a clientconfig.toml(holdercoordinator, seatcc2, node_dispatch local, max_runs 1, cap 2),runtimes.toml(host, herdr, gvisor with nix runsc and pasta, ssh:worker as pi on halogen; cc2 by path~/.claude-work) andsubstrate.json; two credentials,substrate-floor-tokenandsubstrate-link-token-coordinator(Lease session), filled at unit start into the RuntimeDirectory.hosts/coordinator/default.nix: imports the module withfloorUrl = https://substrate.mecattaf.dev,pusher.enable = false,puller.enable = false,owners.codex = tom(E6). The built coordinator toplevel contains 0 substrate units and no substrate secret (MEASURED).tests/substrate-modules(checks.substrate-modules): eval-time proof that both gates are OFF on the real host and that the armed rendering parses.How to enable (Tom's hand, morning; not tonight)
https://substrate.mecattaf.dev(Lane B);wrangler secret put FLOOR_TOKEN; seal the same value assecrets/substrate-floor-token.age(owner tom, editors ++ coordinator key insecrets.nix).services.substrate.pusher.enable = trueinhosts/coordinator/default.nix; switch;systemctl --user status substrate-pusher;GET /capacityshows 5 seats.LINK_TOKENSbound to holdercoordinator, labelsruntime:interpreter,seat:coordinator(substratedocs/DEPLOY.md); seal assecrets/substrate-link-token-coordinator.age;services.substrate.puller.enable = true; switch. Check the unit's PATH sees claude, pi, herdr (extraPathif the profile is not seen).Arm the pusher first, the puller second; keep both OFF until the two age files exist.
Proofs (MEASURED on head
2d97f40c; the same rows were green on0861cebabefore the resync)nix build .#substrate-pusher --no-linkrc 0 (dry tick against the vendored fixture)nix build .#substrate-apps-link --no-linkrc 0 (bundle 2,103,690 B)nix build .#substrate-puller --no-linkrc 0 (starts, exits 78config-invalidwith no config; 86M, closure 327 MB)nix build .#checks.x86_64-linux.substrate-modules --no-linkrc 0 ("pusher.json, runtimes.toml and config.toml render as expected")nix build .#nixosConfigurations.coordinator.config.system.build.toplevel --no-linkrc 0:/nix/store/gabhx87g4a8c0ninvar6jhjmwm132wkj-nixos-system-coordinator-26.11.20260723.e2587canix flake check --no-buildrc 0, "all checks passed!"Evidence paths (coordinator, Tom's home)
/home/tom/today/wednesday-prep-2026-09-23/wrap/RECEIPT.md(every push, build and this PR, with sha)/home/tom/today/wednesday-prep-2026-09-23/wrap/STATE.md(handoffs)/home/tom/today/wednesday-prep-2026-09-23/wrap/VENDOR.md(layout, packaging decisions, unknowns)/home/tom/today/wednesday-prep-2026-09-23/wrap/logs/flake-check-no-build-2d97f40c.log/home/tom/today/wednesday-prep-2026-09-23/wrap/logs/coordinator-toplevel-build-2d97f40c.log/home/tom/today/wednesday-prep-2026-09-23/wrap/logs/flake-check-no-build.log,coordinator-toplevel-build.log(on0861ceba)/home/tom/mecattaf/dotfiles-wt-substrate-wrap; upstream checkout/home/tom/mecattaf/substrateUnknowns and proposed defaults
ax/fleet-zerobefore the main merge or retargets it tomainafterwards: Lane A's call. Default: retarget tomainonceax/fleet-zerois inmain; the paths are disjoint exceptflake.nixpackages and checks blocks and the coordinator import list.0aaf6a2(floor gate) is not vendored. Default: nextsync.shrun picks it up with the next upstream change that touchesapps/*orpackages/*.🤖 Generated with Claude Code
Update 2026-09-24 (retargeted to main, units ON)
pkgs/substrate-appsresynced b105179 → fc2f8bd (seat routing byagent()option,[credentials.seats],codexSandbox,peerCacheDir); pusher, puller and link build;nix flake checkandchecks.substrate-modulespass;nixos-rebuild build --flake .#coordinatorsucceeds at cbe157e (fleet-revision.jsondirty=false).modules/substrate.nixrenders the configuration proven live since 2026-09-23 (runtimes opus/halogen/codex, plus codex-rw withworkspace-write;[credentials.seats]cc/cc2;[puller]demand_dir, drain_timeout_s, capacity_wait_s; pusherpeerCacheDir = "inherit"). Units carry the user profile on PATH (systemdpathreplaces the manager's PATH), the puller's pidfile lives in its RuntimeDirectory, KillMode=mixed, OOMPolicy=continue, WorkingDirectory = the substrate checkout, MemoryHigh/Max 2G/4G, session sockets and CREDENTIALS_DIRECTORY dropped from agents' environment.hosts/coordinator:services.substrate.pusher.enable = true,puller.enable = true.secrets/substrate-floor-token.age,secrets/substrate-link-token-coordinator.agesealed for editors ++ coordinatorOnly (same recipient stanzas as tailscale-authkey-coordinator).hosts/nas/substrate-link.nixstays OFF; needs the floor-side LINK_TOKENS entry, a guest Complete relay and Tom's ack).~/.local/state/substrate/, switch,systemctl --user start substrate-pusher substrate-puller, then the checks: accepted push in the pusher journal, lease loop in the puller journal, coordinator live on/holders,substrate submit prove-three.js --waitsucceeds.codex, a call routed only byseat:codexis refused as ambiguous; workflows nameruntime: "codex"or"codex-rw".🤖 Generated with Claude Code