Skip to content

Embedded Nebula overlay host, off by default - #3

Merged
skeeeon merged 4 commits into
mainfrom
feat/nebula-support
Sep 13, 2026
Merged

skeeeon merged 4 commits into
mainfrom
feat/nebula-support

Conversation

@skeeeon

@skeeeon skeeeon commented Sep 13, 2026

Copy link
Copy Markdown
Contributor

Makes the agent a Nebula host on its organization's mesh, opt-in via nebula.enabled. Design record and rejected alternatives in docs/nebula-design.md; user guide in docs/nebula.md.

Why

The platform already manages a Nebula mesh — pb-nebula runs the CA, signs certificates, generates each host's config. What it could not do is make a device adopt the config it just generated. That was a person with curl.

That gap is load-bearing. Nebula has no CRL: revoking a certificate means listing its fingerprint in pki.blocklist in every other host's config, so a revoked device stays on the mesh until each peer re-reads its own config. CA rotation's three-step interlock and certificate renewal rest on the same assumption. This closes all three loops.

The feature is convergence. Connectivity is a side effect.

Nothing was needed platform-side: things.nebula_host already sits beside things.nats_user, and the nebula_hosts rules already carry an explicit branch letting a thing read the host assigned to it.

Commits

  1. refactor: — retry the NATS connection instead of refusing to start
  2. feat: — the overlay itself
  3. docs: — user guide, example configs, Wintun hint

Commit 1 stands alone: the agent treated NATS as a construction-time dependency, so an agent whose NATS lives on the overlay could never start — connect fails, process exits, and the overlay that would have made it succeed never comes up. A deadlock by structure.

Decisions worth reviewing

The apply path is a ladder of outcomes, not a table of config keys. Reload → if the overlay does not come back, restart → if still not, roll back to the cached config. An earlier draft had the agent classify which keys hot-reload; Nebula already does that internally (firewall, lighthouse, hostmap, connection manager and each tun device register their own callbacks), so a table here would drift on every upgrade. The one key that genuinely cannot reload is listen.port, and pb-nebula issues port 0 to anything that is not a lighthouse or relay.

cmd.nebula answers before it acts. sync and restart reply accepted and work asynchronously, because both can interrupt the tunnel the request arrived through — a reply sent afterwards never lands, and the caller sees a timeout on an operation that succeeded. Outcome is read from cmd.health. There is no stop: it would sever the only channel that could deliver start.

Verification waits for a handshake only when the config names a lighthouse to reach. A lighthouse has no peer to hand shake with, and rolling back a good config because peers were offline causes the outage the rollback prevents.

nebula.sync_interval is bounded above (1m–1h), not just below. It is this device's revocation latency.

Two bugs found while building

  • A self-deadlock I introduced. The session file now has two writers (creds sync, nebula sync), so platform.Client got a mutex — which deadlocked startup, because EnsureCredentials calls Sync and Go mutexes are not reentrant. Existing platform tests caught it by hanging 600s. Split into a locking wrapper plus syncLocked.
  • The Windows example config stopped parsing. Its paths are double-quoted YAML where a lone backslash is an invalid escape. TestExampleConfigsParse now loads all three examples through the real loader.

Costs

before after
binary 19.2 MB 30.5 MB
Go floor 1.24 1.26

The Go bump is unavoidable — every Nebula new enough to issue v2 certificates needs ≥1.25, and v1.11 needs 1.26. Pinned to the same version pb-nebula uses so the library generating configs and the one reading them cannot disagree. CI and goreleaser read go-version-file: go.mod and follow on their own.

Testing

internal/nebula is new (16 tests), internal/nats had none before (3 added), plus Nebula config validation and the example-config parse test. All packages pass; go vet clean; cross-compiles for linux/freebsd/windows.

Not covered: bringing Nebula up needs a real TUN device and root, so the reload/restart/rollback ladder is unverified by go test — the unit tests cover the state machine around it. This is called out in the test file and the design doc. A real-host smoke test is wanted before this is relied on.

internal/tasks fails three tests, identically on main — Windows service-manager access and path-allowlist tests, unrelated to this branch.

Not included

wintun.dll is not shipped (third-party binary; documented instead, with the agent naming the path it looked in). IP-forwarding detection was cut — no gateway is deployed yet, and unsafe_networks from the live certificate already identifies a certified gateway. leaf-sync gaining the same feature is step 4, in the platform repo.

🤖 Generated with Claude Code

skeeeon and others added 4 commits September 12, 2026 22:05
The agent treated NATS as a construction-time dependency: nats.Connect
without RetryOnFailedConnect, a fatal JetStream AccountInfo probe, and
agent.New returning an error that ended the process. That was a
reasonable design while the agent owned only itself — the service
manager restarted it and nothing else was affected.

It stops being reasonable the moment the agent owns a dataplane. An
agent whose NATS endpoint lives on a Nebula overlay could never start:
the connect fails, the process exits, and the overlay that would have
made the connect succeed never gets the chance to come up. A deadlock
by structure rather than by timing. Even on the underlay, a NATS outage
would bounce the mesh on a restart timer.

So: retry the connect, and check JetStream on every connect rather than
once, fatally, at startup. The check still turns "JetStream is not
enabled for this account" into a loud error instead of telemetry that
silently goes nowhere — it is now reported through cmd.health (new
nats.jetstream field, and connected-but-no-JetStream reports degraded)
rather than by refusing to run. Configuration errors are unchanged and
still fail fast: an unknown auth type or unreadable TLS material is a
config error, and unreachability is not.

The part that was not obvious: retrying silently removed the
revoked-credential recovery path. cmd/agent/main.go documents that the
agent exits rather than limps when NATS rejects it, because being
restarted is what re-runs the platform credential sync that heals it.
That exit came from the constructor failure being removed here. nats.go
abandons a connection after two consecutive authorization failures —
exactly what a revoked credential looks like — which would have left
the agent alive against a permanently closed connection: scheduler
firing, every publish failing, nothing ever reconnecting. A
ClosedHandler now distinguishes a connection the agent closed from one
nats.go gave up on, and the latter exits for restart.

internal/nats had no tests. Adds three covering the behaviour that
changed: construction succeeds against an unreachable server, still
fails on a bad auth type, and Lost does not fire on a deliberate close.

Also adds docs/nebula-design.md, which is why this change exists. It is
deliberately not linked from the README's documentation index, which
lists shipped behaviour only.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The platform already manages a Nebula mesh: pb-nebula runs the CA, signs
host certificates, and generates each host's complete config. What it
cannot do is make a device adopt the config it just generated, which
until now was a person with curl.

That gap is load-bearing. Nebula has no CRL — revoking a certificate
means listing its fingerprint in pki.blocklist in every other host's
config — so a mesh only converges if its members re-read their configs
on their own. CA rotation's three-step interlock and certificate renewal
both rest on the same assumption. This closes all three loops. The
feature is convergence; connectivity is a side effect.

Nothing was needed on the platform side. things.nebula_host already sits
beside things.nats_user, and the nebula_hosts rules already carry an
explicit branch letting a thing read the host assigned to it.

internal/nebula runs Nebula in-process via nebula.Main and keeps its
config current. Applying is a transaction: reload, and if the overlay
does not come back, restart, and if that fails, roll back to the last
config known to have worked. The agent deliberately does NOT classify
which config keys can hot-reload — Nebula's firewall, lighthouse,
hostmap, connection manager and tun device each register their own
reload callback and decide for themselves — so the ladder is expressed
in outcomes instead, and there is no table to drift on a Nebula upgrade.
Verification waits for a lighthouse handshake only when the config names
a lighthouse to reach: a false positive here causes the outage it is
meant to prevent.

Fetching reuses the platform client's session rather than duplicating
its authentication: a thing has one platform identity. Same discipline
as the credential path — refresh without expanding, probe ?fields=updated,
and pull the body only when the revision moved. A Nebula config embeds
the host private key inline, so the routine path must not fetch it just
to learn nothing changed. Both syncs now read-modify-write one session
file, so Client grows a mutex; Sync splits into a locking wrapper and
syncLocked, because EnsureCredentials calls it and Go mutexes are not
reentrant.

cmd.nebula takes sync and restart. It answers "accepted" and then acts:
when NATS rides the overlay, both actions interrupt the tunnel the
request arrived through, so a reply sent afterwards never lands and the
caller sees a timeout on an operation that succeeded. The outcome shows
up in cmd.health instead, which is why mesh state belongs there — it is
the completion channel, not just a dashboard. There is no stop: it would
sever the only channel that could deliver start, and "turn Nebula off"
is nebula.enabled: false, which survives a restart.

nebula.sync_interval is the revocation latency for the device, so it is
bounded at 1m-1h rather than left open.

Costs: the binary goes from 19 MB to 30.5 MB, and Nebula v1.11 raises
the module to Go 1.26 — unavoidable, since every Nebula new enough to
issue v2 certificates needs at least 1.25. CI and goreleaser read
go-version-file: go.mod and follow on their own.

Bringing Nebula up needs a real TUN device, so the apply ladder is
verified on a host; the unit tests cover the state machine around it —
sources, the unchanged path, unparseable configs, and the cache.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The overlay shipped with no user-facing documentation and no way to turn
it on from a stock install. Closing that.

docs/nebula.md is the guide: prerequisites, the config block, what
happens at startup and when a config changes, the cmd.nebula verbs and
why their reply does not carry the result, gateway requirements, and
troubleshooting keyed to the errors the agent actually emits. It leads
with sync_interval being the revocation latency rather than a tuning
knob, because that is the setting most likely to be treated as
performance. nebula-design.md stays what it is — a design record, not a
guide — and is now labelled as such where the docs are listed.

The three example configs ship in the release archive and are what a new
install copies, so they carry the block with the same warnings. The
Windows one caught a real bug: its paths are double-quoted YAML, where a
lone backslash is an invalid escape, and the first version of the block
stopped the file parsing. TestExampleConfigsParse loads all three
through the real loader so the next one is caught by CI rather than by
somebody's first boot.

On Windows, Nebula loads wintun.dll from a fixed path relative to the
executable, and the agent does not ship that DLL. What surfaces today is
LoadDLL's "The system cannot find the path specified", naming neither
the DLL nor where it was looked for. windowsTunHint appends the path to
a start failure when the file is genuinely absent.

That is a hint on failure rather than a preflight check, deliberately: a
preflight would hard-code a path out of Nebula's internals, and if
Nebula ever moved it we would refuse to start a host that works. A hint
cannot fail that way. The cost is that it is appended to Windows start
failures unrelated to Wintun, which is why it is worded as a condition.

CLAUDE.md was stale after the last two commits — it still described a
fail-fast JetStream probe and had no idea internal/nebula existed. It
now carries the package, the config block, cmd.nebula and its
answer-before-acting rule, the session mutex and why exported platform
methods delegate to Locked variants, and the Nebula dependency's effect
on the Go floor and binary size.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Adding a longer field name moved the struct's alignment column and
Expand was left behind. gofmt catches it; I did not, because
core.autocrlf=true makes `gofmt -l` list every file in the working copy
and the real complaint was invisible in the noise.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@skeeeon
skeeeon merged commit ec67cb3 into main Sep 13, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant