Repository navigation
Embedded Nebula overlay host, off by default - #3
Merged
Merged
Conversation
The agent treated NATS as a construction-time dependency: nats.Connect without RetryOnFailedConnect, a fatal JetStream AccountInfo probe, and agent.New returning an error that ended the process. That was a reasonable design while the agent owned only itself — the service manager restarted it and nothing else was affected. It stops being reasonable the moment the agent owns a dataplane. An agent whose NATS endpoint lives on a Nebula overlay could never start: the connect fails, the process exits, and the overlay that would have made the connect succeed never gets the chance to come up. A deadlock by structure rather than by timing. Even on the underlay, a NATS outage would bounce the mesh on a restart timer. So: retry the connect, and check JetStream on every connect rather than once, fatally, at startup. The check still turns "JetStream is not enabled for this account" into a loud error instead of telemetry that silently goes nowhere — it is now reported through cmd.health (new nats.jetstream field, and connected-but-no-JetStream reports degraded) rather than by refusing to run. Configuration errors are unchanged and still fail fast: an unknown auth type or unreadable TLS material is a config error, and unreachability is not. The part that was not obvious: retrying silently removed the revoked-credential recovery path. cmd/agent/main.go documents that the agent exits rather than limps when NATS rejects it, because being restarted is what re-runs the platform credential sync that heals it. That exit came from the constructor failure being removed here. nats.go abandons a connection after two consecutive authorization failures — exactly what a revoked credential looks like — which would have left the agent alive against a permanently closed connection: scheduler firing, every publish failing, nothing ever reconnecting. A ClosedHandler now distinguishes a connection the agent closed from one nats.go gave up on, and the latter exits for restart. internal/nats had no tests. Adds three covering the behaviour that changed: construction succeeds against an unreachable server, still fails on a bad auth type, and Lost does not fire on a deliberate close. Also adds docs/nebula-design.md, which is why this change exists. It is deliberately not linked from the README's documentation index, which lists shipped behaviour only. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The platform already manages a Nebula mesh: pb-nebula runs the CA, signs host certificates, and generates each host's complete config. What it cannot do is make a device adopt the config it just generated, which until now was a person with curl. That gap is load-bearing. Nebula has no CRL — revoking a certificate means listing its fingerprint in pki.blocklist in every other host's config — so a mesh only converges if its members re-read their configs on their own. CA rotation's three-step interlock and certificate renewal both rest on the same assumption. This closes all three loops. The feature is convergence; connectivity is a side effect. Nothing was needed on the platform side. things.nebula_host already sits beside things.nats_user, and the nebula_hosts rules already carry an explicit branch letting a thing read the host assigned to it. internal/nebula runs Nebula in-process via nebula.Main and keeps its config current. Applying is a transaction: reload, and if the overlay does not come back, restart, and if that fails, roll back to the last config known to have worked. The agent deliberately does NOT classify which config keys can hot-reload — Nebula's firewall, lighthouse, hostmap, connection manager and tun device each register their own reload callback and decide for themselves — so the ladder is expressed in outcomes instead, and there is no table to drift on a Nebula upgrade. Verification waits for a lighthouse handshake only when the config names a lighthouse to reach: a false positive here causes the outage it is meant to prevent. Fetching reuses the platform client's session rather than duplicating its authentication: a thing has one platform identity. Same discipline as the credential path — refresh without expanding, probe ?fields=updated, and pull the body only when the revision moved. A Nebula config embeds the host private key inline, so the routine path must not fetch it just to learn nothing changed. Both syncs now read-modify-write one session file, so Client grows a mutex; Sync splits into a locking wrapper and syncLocked, because EnsureCredentials calls it and Go mutexes are not reentrant. cmd.nebula takes sync and restart. It answers "accepted" and then acts: when NATS rides the overlay, both actions interrupt the tunnel the request arrived through, so a reply sent afterwards never lands and the caller sees a timeout on an operation that succeeded. The outcome shows up in cmd.health instead, which is why mesh state belongs there — it is the completion channel, not just a dashboard. There is no stop: it would sever the only channel that could deliver start, and "turn Nebula off" is nebula.enabled: false, which survives a restart. nebula.sync_interval is the revocation latency for the device, so it is bounded at 1m-1h rather than left open. Costs: the binary goes from 19 MB to 30.5 MB, and Nebula v1.11 raises the module to Go 1.26 — unavoidable, since every Nebula new enough to issue v2 certificates needs at least 1.25. CI and goreleaser read go-version-file: go.mod and follow on their own. Bringing Nebula up needs a real TUN device, so the apply ladder is verified on a host; the unit tests cover the state machine around it — sources, the unchanged path, unparseable configs, and the cache. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The overlay shipped with no user-facing documentation and no way to turn it on from a stock install. Closing that. docs/nebula.md is the guide: prerequisites, the config block, what happens at startup and when a config changes, the cmd.nebula verbs and why their reply does not carry the result, gateway requirements, and troubleshooting keyed to the errors the agent actually emits. It leads with sync_interval being the revocation latency rather than a tuning knob, because that is the setting most likely to be treated as performance. nebula-design.md stays what it is — a design record, not a guide — and is now labelled as such where the docs are listed. The three example configs ship in the release archive and are what a new install copies, so they carry the block with the same warnings. The Windows one caught a real bug: its paths are double-quoted YAML, where a lone backslash is an invalid escape, and the first version of the block stopped the file parsing. TestExampleConfigsParse loads all three through the real loader so the next one is caught by CI rather than by somebody's first boot. On Windows, Nebula loads wintun.dll from a fixed path relative to the executable, and the agent does not ship that DLL. What surfaces today is LoadDLL's "The system cannot find the path specified", naming neither the DLL nor where it was looked for. windowsTunHint appends the path to a start failure when the file is genuinely absent. That is a hint on failure rather than a preflight check, deliberately: a preflight would hard-code a path out of Nebula's internals, and if Nebula ever moved it we would refuse to start a host that works. A hint cannot fail that way. The cost is that it is appended to Windows start failures unrelated to Wintun, which is why it is worded as a condition. CLAUDE.md was stale after the last two commits — it still described a fail-fast JetStream probe and had no idea internal/nebula existed. It now carries the package, the config block, cmd.nebula and its answer-before-acting rule, the session mutex and why exported platform methods delegate to Locked variants, and the Nebula dependency's effect on the Go floor and binary size. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adding a longer field name moved the struct's alignment column and Expand was left behind. gofmt catches it; I did not, because core.autocrlf=true makes `gofmt -l` list every file in the working copy and the real complaint was invisible in the noise. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Makes the agent a Nebula host on its organization's mesh, opt-in via
nebula.enabled. Design record and rejected alternatives indocs/nebula-design.md; user guide indocs/nebula.md.Why
The platform already manages a Nebula mesh —
pb-nebularuns the CA, signs certificates, generates each host's config. What it could not do is make a device adopt the config it just generated. That was a person withcurl.That gap is load-bearing. Nebula has no CRL: revoking a certificate means listing its fingerprint in
pki.blocklistin every other host's config, so a revoked device stays on the mesh until each peer re-reads its own config. CA rotation's three-step interlock and certificate renewal rest on the same assumption. This closes all three loops.The feature is convergence. Connectivity is a side effect.
Nothing was needed platform-side:
things.nebula_hostalready sits besidethings.nats_user, and thenebula_hostsrules already carry an explicit branch letting a thing read the host assigned to it.Commits
refactor:— retry the NATS connection instead of refusing to startfeat:— the overlay itselfdocs:— user guide, example configs, Wintun hintCommit 1 stands alone: the agent treated NATS as a construction-time dependency, so an agent whose NATS lives on the overlay could never start — connect fails, process exits, and the overlay that would have made it succeed never comes up. A deadlock by structure.
Decisions worth reviewing
The apply path is a ladder of outcomes, not a table of config keys. Reload → if the overlay does not come back, restart → if still not, roll back to the cached config. An earlier draft had the agent classify which keys hot-reload; Nebula already does that internally (firewall, lighthouse, hostmap, connection manager and each tun device register their own callbacks), so a table here would drift on every upgrade. The one key that genuinely cannot reload is
listen.port, andpb-nebulaissues port 0 to anything that is not a lighthouse or relay.cmd.nebulaanswers before it acts.syncandrestartreplyacceptedand work asynchronously, because both can interrupt the tunnel the request arrived through — a reply sent afterwards never lands, and the caller sees a timeout on an operation that succeeded. Outcome is read fromcmd.health. There is nostop: it would sever the only channel that could deliverstart.Verification waits for a handshake only when the config names a lighthouse to reach. A lighthouse has no peer to hand shake with, and rolling back a good config because peers were offline causes the outage the rollback prevents.
nebula.sync_intervalis bounded above (1m–1h), not just below. It is this device's revocation latency.Two bugs found while building
platform.Clientgot a mutex — which deadlocked startup, becauseEnsureCredentialscallsSyncand Go mutexes are not reentrant. Existing platform tests caught it by hanging 600s. Split into a locking wrapper plussyncLocked.TestExampleConfigsParsenow loads all three examples through the real loader.Costs
The Go bump is unavoidable — every Nebula new enough to issue v2 certificates needs ≥1.25, and v1.11 needs 1.26. Pinned to the same version
pb-nebulauses so the library generating configs and the one reading them cannot disagree. CI and goreleaser readgo-version-file: go.modand follow on their own.Testing
internal/nebulais new (16 tests),internal/natshad none before (3 added), plus Nebula config validation and the example-config parse test. All packages pass;go vetclean; cross-compiles for linux/freebsd/windows.Not covered: bringing Nebula up needs a real TUN device and root, so the reload/restart/rollback ladder is unverified by
go test— the unit tests cover the state machine around it. This is called out in the test file and the design doc. A real-host smoke test is wanted before this is relied on.internal/tasksfails three tests, identically onmain— Windows service-manager access and path-allowlist tests, unrelated to this branch.Not included
wintun.dllis not shipped (third-party binary; documented instead, with the agent naming the path it looked in). IP-forwarding detection was cut — no gateway is deployed yet, andunsafe_networksfrom the live certificate already identifies a certified gateway.leaf-syncgaining the same feature is step 4, in the platform repo.🤖 Generated with Claude Code