Gameplane's dashboard is deliberately internet-exposed — that's the whole point. Assume:
- the login page is enumerable by any scanner,
- cluster-internal attackers may land a pod in
gameplane-gamesvia a compromised game image (Minecraft plugins, Valheim mods, etc.), - the game pods themselves should be treated as low-trust,
- a dashboard user with only
servers:readmust not be able to read game passwords (RCON, admin or server passwords) out of a GameServer, - a usage-telemetry receiver is reachable by anyone on the internet, so it must stay safe when fed hostile or invented reports, and must never reveal figures to an unauthenticated viewer (see Telemetry).
Operator-managed agent and capture-sidecar control traffic uses mTLS with TLS 1.2 or newer and a verified
client certificate. Both servers reload their serving certificate, key and
client CA on each new handshake, so operator Secret renewal takes effect
once Kubernetes updates the mounted projection. Each server reads one pinned
..data generation from the Secret rather than mixing files during a
projection swap. Invalid rotated material rejects new handshakes until
repaired; it never falls back to old trust. A decoded client-CA
CERTIFICATE block with invalid DER rejects the entire bundle, even
when it also contains valid CAs. Valid multi-CA bundles remain supported.
TLS session tickets are
disabled so new connections cannot resume authentication against a
removed CA. Existing connections retain their established TLS state.
Two modes, configurable independently:
- Local accounts — argon2id (64 MiB, t=3, p=2) password hashing.
Session cookies are HttpOnly, Secure, SameSite=Lax. The database stores only a SHA-256 digest of the session cookie value, so a leaked database or backup does not yield usable session cookies; the digest itself is rejected if presented as a cookie. CSRF protection
via a double-submit
X-Gameplane-CSRFheader on mutating requests. - OIDC — Keycloak, Google, GitHub, any RFC-7519 compliant IdP.
State validated through a short-lived cookie;
id_tokensignature verified against the provider's JWKS.
On first OIDC login for a subject, Gameplane creates a user row with
role viewer. Admins must promote new OIDC users manually.
OIDC providers can be added at runtime under Admin Settings →
Authentication: issuer + client id live in the auth config row (they
are public OAuth identifiers), and the client secret is stored as an
API-managed Secret gameplane-auth-<name> in the control-plane
namespace. Two labels bound the API's reach: it only reads Secrets
labelled gameplane.local/auth-provider=true, and only deletes ones
additionally labelled gameplane.local/managed-by=gameplane-api — so a
config:manage user can neither exfiltrate arbitrary control-plane
Secrets through a provider's configRef nor delete kubectl-/GitOps-
created ones over HTTP. Provider changes apply on save, no restart: the
registry re-reads the config row per auth request and rebuilds OIDC
clients lazily (issuer discovery cached, failures back off).
A provider configured through Helm flags (api.oidc.*) appears as the
read-only helm provider; it is owned by values.yaml and cannot be
edited, disabled, or deleted from the dashboard.
Saving an auth config with zero enabled providers is rejected (the Helm provider counts as always-enabled). If you still lock yourself out — local login disabled while the only OIDC provider is broken — run the break-glass inside the API pod:
kubectl -n gameplane-system exec deploy/gameplane-api -- \
/api bootstrap-admin --enable-local-loginIt force-enables the local provider in the auth config row, keeping
everything else in that row as it was (the other providers and the
helmOverride role-mapping overlay), and takes effect on the next login
attempt.
bootstrap-admin --username <name> --force resets that account's password,
promotes it to admin, and ends every existing session of the account, the
same way a dashboard password reset does.
The API records a client IP for every request. Login rate limiting
(per-IP caps) and audit records key on it. The IP comes from the TCP peer
and, only when that peer is a trusted proxy, from the X-Forwarded-For
header. Trusted proxies are set by api.trustedProxies (default:
loopback and private ranges 127.0.0.0/8, 10.0.0.0/8, 172.16.0.0/12,
192.168.0.0/16, 169.254.0.0/16, ::1/128, fc00::/7, fe80::/10).
The API applies these rules in order (api/internal/auth/clientip.go):
- A TCP peer outside
api.trustedProxiesis the client. AnyX-Forwarded-Forheader on its request is ignored. - When the peer is a trusted proxy, the API reads
X-Forwarded-Forfrom right to left and takes the first address outsideapi.trustedProxiesas the client. - When every address in the chain is inside
api.trustedProxies, the leftmost address is the client. - An entry that isn't an IP address ends the walk at the last address
already reached. With no
X-Forwarded-Forheader that is the peer. - With
api.trustedProxiesempty, the peer is always the client.
IPv4-mapped IPv6 prefixes. Prefixes must be standard IPv4 (e.g., 10.42.0.0/16)
or IPv6 (e.g., fc00::/7). IPv4-mapped IPv6 prefixes (e.g., ::ffff:10.42.0.0/112)
are automatically normalized: they are converted to their unmapped IPv4 form and must
be at least /96 bits to be valid (::ffff:10.42.0.0/96 becomes 10.42.0.0/0,
::ffff:10.42.0.0/112 becomes 10.42.0.0/16). Prefixes shorter than /96 are
rejected at startup.
In a normal Kubernetes install (ingress controller, then the web front end, then the API), the default works out of the box: the proxies run in pod and node ranges the default covers, and a client on the public internet is recorded by its own address.
Clients on a private network. The default treats every private-range
address as a possible proxy, so for clients that connect from a private
range the recorded IP depends on the forwarded chain those addresses
present (rule 3). If dashboard users reach Gameplane from a private
network and you rely on per-client limits for them, narrow
api.trustedProxies to the ranges your proxies actually run in, usually
the cluster's pod CIDR plus any load balancer in front of the ingress:
api:
trustedProxies: "10.42.0.0/16" # k3s default pod CIDR; use your cluster'sWhen the API is directly exposed (no proxy), set
api.trustedProxies to "": the TCP peer is then always the client. If
you place a proxy in front of the API, list that proxy's addresses so the
API reads X-Forwarded-For from it. A proxy outside the list is recorded
as the client itself, so every user behind it shares one rate-limit
bucket.
Example for a single proxy at 203.0.113.1:
api:
trustedProxies: "203.0.113.1/32"The client IP is used for login rate limiting (per-IP caps) and audit records, so misconfiguring this can either record a proxy instead of the client in audit logs or group legitimate users behind one proxy address in a single rate-limit bucket.
RBAC is permission-based. A permission is a fixed resource:action
string from the server-defined catalog (api/internal/rbac/catalog.go,
e.g. servers:write, backups:restore, users:manage). A role is a
named set of permissions, and a user is bound to roles per namespace.
- Roles live in the API database (
roles/role_permissions). The built-inadmin,operator, andviewerroles are seeded so their cluster-wide grants reproduce the historical role matrix exactly.adminholds the*wildcard and is immutable;operator/viewerare editable templates; custom roles can be created with any subset of the catalog (the*wildcard is never grantable through the API). Built-in roles and roles still assigned to a user cannot be deleted. - Bindings (
user_role_bindings) grant a role in a namespace;*means cluster-wide. A user's primary role (PATCH /users/{id}) is their cluster-wide binding; additional per-namespace grants are managed via…/users/{id}/bindings. Allowed namespaces are theGAMEPLANE_EXTRA_NAMESPACESallow-list plus the defaultgameplane-games. - Enforcement (
api/internal/rbac/rbac.go): each route maps to one required permission; the middleware resolves the request's target namespace and checks the caller's resolved permission set. A namespaced permission is granted by a cluster-wide binding or a binding in the target namespace; a cluster-scoped permission requires a cluster-wide binding — the same Role vs ClusterRole split Kubernetes uses. Unmatched routes fail closed. - Lockout guards. The API refuses to demote or delete the last user who
can manage users, and refuses self-demotion below
users:manage. Role edits follow the same rules: a change that removesusers:managefrom the caller's own primary role, or from the role every user manager holds, is refused. - Event stream.
GET /eventscarries only the resource kinds the caller may read in the resolved cluster and namespace, using the same read permission as each kind's list route. - Account removal.
DELETE /users/{id}removes the account and every row tied to it in one transaction (SSO links, preferences, sessions, API tokens, role bindings) and revokes the share links the account created. The API does this itself rather than relying on foreign-key cascades, which the shipped SQLite DSN leaves off. An SSO user who is deleted and signs in again is provisioned as a new user.
In addition to namespace-based RBAC, GameServers support ownership and
collaboration: the owner (who created the server) and any collaborators
(managed via PUT /servers/{name}:collaborators) gain operational control over
that specific server, regardless of their namespace role. This is purely additive
— it does not override namespace bindings. Collaborators retain: read, console,
WebSocket access, start/stop/restart/clone operations, and files/players/config
subroutes. Destructive operations are owner-only: delete, wipe-data, ownership
transfer, and collaborator list edits. Only the server's owner or an admin (a
role holding * in the server's cluster and namespace) can perform owner-only
operations. The namespace servers:write permission alone does not grant them,
and a server with no recorded owner (for example one created with kubectl or
GitOps) can be transferred, wiped or deleted only by an admin. The transfer,
collaborator-edit and wipe patches are conditional on the server's
resourceVersion as read by the ownership check: if the server changes in
between (for example its ownership is transferred), the API re-reads it and
repeats the check, so a caller who is no longer the owner is refused, and
after three conflicting attempts the request fails with 409. Backups,
restore jobs, schedules, and events remain namespace-gated in this release.
Generic GameServer create/update requests cannot set or remove annotations in
the gameplane.local domain or its subdomains, except the user-editable
description and unused legacy grace-period-seconds hints. The API strips supplied values
on creation before assigning the owner, and preserves live values on updates.
Wipe requests, controller acknowledgements and lifecycle guards therefore cannot
be injected or erased through ordinary settings writes. Use the authorized
lifecycle and ownership endpoints for those operations.
GameServer.spec.config holds wizard values as stored in Kubernetes, including
the values of type: password fields of the template's configSchema. The API
never returns those in clear: every response that carries a GameServer
(/servers, /servers/{name} including create/update replies, :clone,
/users/me/servers, /fleet/servers and the /events stream) replaces each
non-empty password value with the marker __gameplane_redacted__, through one
helper (api/internal/handlers/config_redact.go). If the GameTemplate cannot be
read, every config value is redacted (fail closed).
These responses also omit the kubectl.kubernetes.io/last-applied-configuration
annotation, which can contain passwords from an older manifest even after they
are removed from current config. This applies even to servers without config;
the annotation remains stored in Kubernetes.
On PUT /servers/{name} a password value equal to the marker keeps the stored
value, a different non-empty value replaces it, and an empty string clears it
only when the field is optional (a required field keeps its value); a required
password omitted from the body is restored. Share links, capture, mods, tunnel
credentials, notifications, WebSockets and the audit log (which never records
request bodies) do not carry spec.config.
Residual exposure: anyone who can read the GameServer object through the
Kubernetes API (kubectl, GitOps tooling, the optional read-only mcp-server)
still sees the stored values; restrict that with cluster RBAC.
Share links (api/internal/db/shares.go, schema in api/internal/db/migrations/sqlite/006_share_links.sql, Postgres twin in migrations/postgres/) grant unauthenticated, token-bearing access to a single GameServer's status and connection address, optionally with permission to wake it. Because they are unauthenticated, their security rests entirely on the token being both hard to guess and hard to recover if the database or a backup leaks.
Storage. Only a SHA-256 hash of the token is persisted (token_hash, indexed for O(1) lookup); the raw 32-byte random token is generated at creation, returned exactly once in the create response, and never stored, logged, or recoverable afterwards.
Why a fast hash is the right choice here, not a weakness. SHA-256 is not a password-hashing function (no salt, no work factor), and that is deliberate:
- The input is a 32-byte cryptographically random value, not a low-entropy human-chosen secret. There is no dictionary or brute-force search over a keyspace of 2^256 that a fast hash makes newly feasible — the security margin comes entirely from the token's entropy, not from how expensive the hash is to compute.
- A database or backup leak exposes only hashes. Recovering a usable token from a hash would require reversing SHA-256 or brute-forcing a 256-bit random space; both are computationally infeasible regardless of the hash function's speed. A slow KDF (bcrypt/argon2/scrypt) would add no meaningful protection here, since the "weak secret" scenario those algorithms defend against does not apply.
- The token is looked up on every public request to
GET /shares/{token}andPOST /shares/{token}/start. A fast, indexed equality lookup keeps that path cheap; a deliberately slow KDF would turn every anonymous share-page load into an expensive hashing operation — the opposite of what a slow KDF is for (limiting an attacker's guesses per second), since here the attacker's limiting factor is token entropy, not hash speed.
Indistinguishability. An unknown token, an expired token, and a revoked token all produce the identical response (ErrShareLinkInvalid internally; a uniform 404 at the HTTP layer). This prevents an attacker from learning, by probing, whether a guessed token ever existed, expired, or was deliberately revoked.
Token placement. The token travels as a path segment (/shares/{token}), not a query parameter, keeping it out of query-string capture by proxies, access logs and analytics, and letting audit logging redact it from recorded request paths with a simple prefix rewrite (redactShareToken in api/internal/audit/audit.go). A path segment is still part of the URL, so it does not by itself keep the token out of Referer headers; API responses set Referrer-Policy: no-referrer (secureHeaders in api/cmd/main.go) for that.
Expiry. Every share link either has an expiry timestamp or is explicitly created with no expiry (owner's choice; see specs/done_017-share-link-expiry/). There is no platform-enforced maximum lifetime: a non-expiring or long-lived link is exactly as hard to guess on any given day as a short-lived one, because guessing difficulty comes from the token's entropy, not from its age. The tradeoff of a long-lived or non-expiring link is operational — a forgotten link stays live until the owner revokes it — not cryptographic, which is why the create-link UI warns the owner explicitly ("This link works until you revoke it." for no expiry; a long-lived-token warning for a custom date a year or more out) rather than the system silently capping the choice.
Revocation. Revocation sets revoked_at (never a delete, preserving the audit trail) and is checked independently of, and prior to, any expiry check, so it applies uniformly regardless of whether the link expires, expires far in the future, or never expires. DELETE /servers/{name}/shares/{id} revokes a link only when it belongs to that server (cluster, namespace and name); any other id answers 404. Deleting a user revokes every share link that user created.
Rate limiting. The public resolve/start endpoints are rate-limited (auth.ShareLimiter) specifically because tokens are guessable-by-brute-force in principle (just computationally infeasible in practice); the rate limit is defense in depth against automated probing, not a substitute for token entropy.
mTLS. The Helm chart provisions a self-signed CA via a post-install
hook (or takes an existing gameplane-agent-ca Secret). The operator
uses the CA to sign per-pod server certs; the API uses a single client
cert. Agent refuses plain-HTTP traffic when TLS material is present.
Fallback: a shared-secret bearer token via --api-token-file. Only
intended for local kind development where mTLS is overkill.
Console-injection guard. Free text folded into an RCON/stdin command (module action string params and player kick/ban reasons) passes through gameaction.CheckText, which rejects ASCII control characters and the metacharacters ; & | $ ` \ " '. The agent applies it independently of the API.
When networkPolicies.enabled=true (default) the chart applies:
default-deny-ingress— denies all ingress to all pods in the games namespace. A Kubernetes Service does not create any NetworkPolicy allowance; ingress traffic is fully isolated unless an allow-policy explicitly permits it.default-deny-egress— denies all egress from all pods in the games namespace except DNS (UDP/TCP port 53 tokube-system). This is the most restrictive policy; outbound downloads (binaries, assets, mods) are gated by theallow-game-public-egresspolicy below, and apiserver access byallow-agent-to-apiserver.allow-agent-to-apiserver— allows game pods to reach the kube-apiserver for status heartbeat (GameServer status patches). By default permits TCP 443 and 6443 to all addresses in RFC1918 and link-local ranges (10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 169.254.0.0/16), not limited to the apiserver endpoint. Customize vianetworkPolicies.apiServerCIDRsto narrow to specific API server addresses.allow-api-to-agent— allows the API and operator pods (in the control-plane namespace) to reach every game pod's agent on TCP port 8090. The API proxies console/files/logs/players; the operator calls/quiescebefore backups.allow-api-to-capture— rendered only whencapture.enabled=true; allows TCP 9091 to game pods from this release's operator and enabled API pods in the control-plane namespace. Both the namespace and pod selectors must match. This admits operator capture control and local API file access without opening other ports or depending on the kubelet probe allowance. Remote file access has its separate gateway policy.<gameserver-name>-game-ingress— created by the operator for each GameServer, allows external traffic to reach only the advertised ports declared in the GameTemplate. Selects traffic fromnetworkPolicies.gameIngress.fromCIDRs(default0.0.0.0/0, the internet) to only the container ports markedAdvertise: trueat their declared protocol (UDP preserved). This policy does not itself open RCON or the agent's port 8090; however, those ports remain reachable fromkubeletCIDRsvia theallow-kubelet-probespolicy (see below), which has no port restrictions by default. To close that hole, narrownetworkPolicies.probePorts— previously unsafe to do, but now safe since this policy protects advertised player traffic separately. Operators may also narrowfromCIDRsto a private range (e.g. a LAN, a VPN CIDR) to gate player access. If a template declares no advertised ports ornetworkPolicies.gameIngress.enabled: false, the operator ensures this policy does not exist. The policy is owned by the GameServer and cascade-deletes with it.allow-kubelet-probes— allows kubelet to reach game pods for liveness/readiness probes. By default targets RFC1918 + link-local ranges, or customizable vianetworkPolicies.kubeletCIDRs; probe ports vianetworkPolicies.probePorts.allow-prometheus-to-agent— opt-in (rendered only whenserviceMonitors.scrapeNamespaceSelectoris set) policy allowing Prometheus pods in the selected namespace to reach the agent's plain-HTTP metrics port (TCP 9090). That port is unauthenticated HTTP, not mTLS, and is otherwise already reachable fromkubeletCIDRsviaallow-kubelet-probesabove unlessnetworkPolicies.probePortsis narrowed to exclude it.allow-game-public-egress— (enabled by default, gated bynetworkPolicies.gameEgress.enabled) allows game pods to reach the public internet for binary/asset/mod downloads. SetnetworkPolicies.gameEgress.enabled: falseto withhold public egress. Permits egress to 0.0.0.0/0 except private ranges defined innetworkPolicies.gameEgress.privateCIDRs(by default: 10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16, 169.254.0.0/16 to block in-cluster and cloud-metadata access). Ports customizable vianetworkPolicies.gameEgress.ports.
Note: default-deny-egress applies to all pods in the games namespace
(podSelector: {}), while allow-agent-to-apiserver and
allow-game-public-egress only select game pods labelled
app.kubernetes.io/name: gameplane-game. Any pod in the games namespace
without that label receives DNS-only egress (connecting to kube-dns on port
53). Do not place helper or debug pods in the games namespace expecting them
to route traffic—place them in a different namespace instead.
Every Gameplane-managed pod (operator, api, agent, and the optional audit-syslog-bridge) runs as:
runAsNonRoot: true(uid 65532)readOnlyRootFilesystem: trueseccompProfile.type: RuntimeDefaultcapabilities.drop: [ALL]allowPrivilegeEscalation: false
The agent sidecar in each game pod gets the same fixed hardening. The game
container itself does not: it runs as the template's
spec.security.runAsUser/runAsGroup when the template sets them (see
module authoring), and otherwise as
the image's own default user, which may be root. The operator sets no
runAsNonRoot, allowPrivilegeEscalation, readOnlyRootFilesystem,
seccompProfile or capability drop on the game container, so it keeps the
container runtime's defaults. This keeps arbitrary third-party game images
working unchanged.
Game pods are shaped per-template. For a hostile game module, enable
Pod Security Standards restricted on the games namespace via
podSecurity.enforceRestricted=true.
The optional network capture feature [optional] (see
docs/roadmap.md) adds an ephemeral sidecar container to game pods when
capture is enabled. This sidecar requires allowPrivilegeEscalation: true in
its securityContext, which violates the Pod Security Standards restricted
profile.
Why the exception is necessary: The sidecar acquires CAP_NET_RAW capability
via file capabilities (setcap cap_net_raw+ep), applied at container image build
time. This mechanism is necessary because Kubernetes does not set ambient
capabilities on a container's process by default: declaring
securityContext.capabilities.add: ["NET_RAW"] alone on a non-root user grants
nothing — the kernel clears the effective capability set when the entrypoint
binary is executed (via execve), leaving the non-root process with no
capabilities. File capabilities survive this exec because they are consulted
independently by the kernel at exec time, independent of ambient state.
The container's securityContext.capabilities still lists
Drop: ["ALL"], Add: ["NET_RAW"], but that Add is not what grants the
capability — it exists only because Drop: ["ALL"] on its own would also empty
the process's bounding capability set, and the kernel refuses to grant a file
capability at execve that isn't in the bounding set (EPERM). Re-adding NET_RAW
keeps it available in the bounding set so the file capability grant can proceed;
the process's own effective set is still empty at start, and the actual grant
comes from the setcap'd binary at exec, not from capabilities.add. Separately,
file capabilities are ignored by the kernel when no_new_privs is set, and
Kubernetes sets no_new_privs whenever allowPrivilegeEscalation: false.
Therefore, the container must also set allowPrivilegeEscalation: true for the
file capability to function.
The game container does not share this exception: the operator never sets
allowPrivilegeEscalation: true on it or adds NET_RAW to it, and the capture
sidecar is the only container Gameplane grants CAP_NET_RAW for capture. The
game container keeps the posture described in Pod security:
the template's uid/gid or the image's default user, and the container
runtime's default capability set. Which capabilities the game container holds
therefore depends on that runtime default and on the user the image runs as,
not on the capture feature.
Trade-off with PodSecurity restricted: A cluster enforcing the restricted
Pod Security Standards profile on the games namespace will reject any pod with
allowPrivilegeEscalation: true. If you require restricted admission on your
games namespace, you have three options:
- Disable capture — leave the cluster's capture feature disabled via Helm
value
capture.enabled: false(default is false). Captures are not required for normal operation; this is the safest option if you cannot or prefer not to relax therestrictedprofile. - Exempt the games namespace — remove or relax the Pod Security Standards
restrictedenforcement for the games namespace, usingbaselineorprivilegedinstead. The games namespace remains an untrusted environment (game code can run arbitrary containers), but the admission level permits the capture sidecar to be injected when needed. - Leave the games-namespace label off — the chart only adds the
pod-security.kubernetes.io/enforce: restrictedlabel to the games namespace whenpodSecurity.enforceRestricted=true; that value defaults tofalse. A default install therefore already leaves the label off, and captures work without any change. If you previously setpodSecurity.enforceRestricted=trueand accept the operational trade-off, set it back tofalsein the Helm values to drop the label. This setting only affects the gamesNamespaceobject; it has no cluster-wide effect and there is no per-pod opt-in.
Data sensitivity: Captures contain binary game protocols, player IP addresses,
and may include sensitive data like in-game chat or credentials. An admin with
captures:manage permission can start a capture and read all traffic reaching
that pod, including other players' packets. This access is admin-only and fully
audited (every capture operation is logged). Default capture retention is 24 hours,
reducing the exposure window for captured data.
A GameTemplate materialized from a module chooses the container image,
command, and config a game pod runs — so a module source is a trust
boundary. Three controls protect it:
- Fetch SSRF guard. The operator's
git/httpsource fetchers (netguard.IsAllowed) refuse link-local, cloud-metadata (169.254.169.254), unspecified, and multicast destinations, at dial time (so a DNS name rebinding to one is caught). This blocks a source from being aimed at the instance-metadata endpoint to steal the operator's IAM credentials. Private/loopback addresses stay reachable for self-hosted registries.ModuleSourcemutation is admin-only, so this is defense-in-depth. - Signature verification.
ModuleSource.spec.verify(OCI sources) makes the operator refuse any bundle without a valid cosign signature — keyed (a public key) or keyless (a pinned Fulcio issuer + identity). Use it for any source you don't fully control. The officialmodules/*bundles are keyed-signed by the release pipeline (and now also recorded in the public Sigstore Rekor log), and verify offline (no Rekor/Fulcio reachability needed). The operator's verification is intentionally offline/keyed, keeping air-gapped and self-hosted clusters functional. Opt-in enforcement of transparency log inclusion is future work. Signing is an OCI concept, so switch the default source totype: ociand enabledefaultModuleSource.oci.verify.enabled. - Digest pinning.
Module.spec.digestpins exact bundle content; a moved tag fails the install withDigestMismatch.
Verification and pinning are opt-in. A source with neither is trusted to
serve a GameTemplate whose image/command runs in your cluster — only point
Gameplane at module sources you trust, and prefer signed, pinned installs for
third-party games. Authoring details: module-authoring.md.
The agent's /files/* routes (list, read, download, write, upload, mkdir,
delete) operate on the game data volume and are reachable by anyone who can
use the dashboard file browser, while the game container (a different uid)
can create symlinks on the same volume. The agent therefore treats the data
directory contents as untrusted and does not trust a path it validated
earlier:
- The requested path is validated lexically (no
..escape, no dot-prefixed component) and every operation then runs relative to a directory descriptor opened on the data root, opening one component at a time withopenat(O_NOFOLLOW). A symlink at any component is refused (HTTP 400); there is no "follow it if it stays inside the root" case. - Reads, writes, creates and deletes use the descriptors they obtained
(
fstatat,mkdirat,unlinkat,renameat, serving from the opened file) and never re-open an absolute path, so swapping a checked directory or file for a symlink between validation and use cannot redirect the operation out of the root. Writes are atomic: temp file in the parent descriptor,fsync,renameatin the same directory. - Recursive delete walks descriptors and unlinks links as links; it refuses trees that contain dot-prefixed entries (agent state such as the mods manifest).
- The data-root path itself is operator-configured and trusted (it may be a symlink); only components below it are checked.
The mods package confines its own paths separately (ConfinePath/ConfineRelPath)
and does not use this descriptor walk.
Separately from the module supply chain above, a running server can install
mods/plugins at runtime if its template declares
capabilities.mods.install — a user-supplied URL the agent downloads into
the server's data volume (see
module-authoring.md). This is a distinct trust
boundary: the target is whatever host the logged-in user types, not an
admin-configured ModuleSource, so it is guarded more strictly
(netguard.IsPublic): only globally routable addresses are allowed —
loopback, private/ULA, link-local, and CGNAT/reserved ranges are all
refused, not just the cloud-metadata range. An allowedHosts allow-list is
also required before installs are enabled at all, and redirects are
re-checked against both the host allow-list and the address guard. This
guard and the operator's fetch guard above share their dial-time enforcement
machinery in the netguard/ module; its package doc explains why the two
policies (IsAllowed vs IsPublic) stay separately selectable rather than
being collapsed into one.
Notification sinks (docs) are a third outbound-dial
surface, between the two above in trust: the URLs are configured at runtime
through the dashboard (unlike the deploy-time audit webhook flag), but only
by users holding config:manage — admin-tier, the same trust class as the
operator's ModuleSources. They get the same guard: every sink dial
(HTTP and SMTP) goes through netguard.IsAllowed, so LAN/in-cluster
receivers (ntfy, a syslog bridge, an SMTP smarthost) keep working while
link-local (cloud metadata), unspecified, multicast, and NAT64/6to4
destinations are refused at dial time — DNS rebinding can't slip past.
Two further containments: sink credentials resolve only from Secrets
labelled gameplane.local/notification-sink=true in the control-plane
namespace (so a sink configRef can't be aimed at an arbitrary Secret),
and delivery errors are sanitized to never echo the sink URL, whose path
often embeds a capability token.
The API server can resolve Steam IDs in player lists to public display names
through the Steam Web API (ISteamUser/GetPlayerSummaries/v2, batched up to
100 IDs per call). It is an optional outbound dependency, off unless
api.steam.apiKeySecretRef.name is set:
- Egress origin. Requests leave only from the API pod in the control-plane
namespace, to
api.steampowered.comover HTTPS. Game pods in the games namespace stay underdefault-deny-egress; they get no path to Steam and never see the key. - Dial guard. The HTTP client is built on
netguardwith the strictIsPublicpolicy, so loopback, private, link-local and reserved destinations are refused at dial time, including after DNS resolution. - Credential. The key is a user-supplied Secret, injected as the
environment variable
GAMEPLANE_STEAM_API_KEY(never a flag, so it does not appear in the pod spec orpsoutput). It is never logged, and failure messages from the resolver are key-free. It is never returned in an API response or sent to the browser. - Cache. Resolved names are kept in a bounded in-process LRU cache keyed on Steam ID, with a positive TTL (default 12h) and a shorter negative TTL (default 15m) for IDs Steam does not return. The cache holds no credential, is never written to disk or the database, and is discarded on restart. Failed calls are not cached.
- Identity. A resolved name is display-only. It is never an identifier for moderation: kick, ban and unban key on the Steam ID alone, so a name change or a spoofed display name cannot change who is acted on.
- Failure mode. If the key is unset, Steam is unreachable, or a call times out (the lookup has a per-request budget of about 1.5 seconds), the player list falls back to raw Steam IDs. The player list is never blocked or failed by the lookup.
audit_events is a hash chain (migration 005_audit_chain.sql): every row
inserted after that migration stores hash = SHA-256(prev_hash || canonical(row)), and GET /admin/audit/verify re-walks the chain to report
the first broken link. Two config-table entries bound the walk: a
Prune-written checkpoint anchors the oldest surviving row after a
retention sweep, and a per-insert head anchors the newest row, so a
DELETE FROM audit_events WHERE id > N — truncating only the tail, which
would otherwise leave every surviving link internally consistent — is
detected too.
Be precise about what this catches. The chain is unkeyed: it recomputes
hashes from row content and two config-table entries, and config is
writable by anyone with the same database access an attacker would need to
tamper with audit_events in the first place. This mechanism reliably
detects:
- naive in-DB tampering —
UPDATE/DELETE(including tail truncation) againstaudit_eventsalone, without also touchingconfig; and - accidental corruption (a bad migration, a restore from an inconsistent backup, etc).
It does not detect a sophisticated attacker who has DB write access and also recomputes and rewrites the checkpoint and head to match — that attacker can forge an internally-consistent chain from any starting point. Nothing server-side can close that gap while the verification data lives in the same database the attacker can already write to.
The real append-only record of last resort is the external sinks — stdout (cluster log aggregation), the audit webhook, and the S3 batch sink (see Secrets below for how their credentials are contained). Because delivery is push-based and decoupled from the request path, an attacker who compromises the database after the fact cannot retroactively alter what was already shipped to those destinations. Treat the hash chain as tamper-evidence for common-case tampering and corruption, and the external sinks as the actual tamper-proof trail.
A documented future hardening is HMAC-keyed chaining (hash = HMAC-SHA256(key, prev_hash || canonical(row)) with the key held outside the
database — e.g. a K8s Secret the API process reads but never writes back),
which would raise the bar to compromising that external key as well. Not
implemented today; tracked in roadmap.md.
Gameplane's release binaries and container images are built via GitHub Actions. The CI pipeline is a trust boundary: compromised build steps can inject malicious code into what users deploy. Several controls harden the pipeline:
External GitHub Actions in .github/workflows/ and .github/actions/ are pinned
to a 40-character commit SHA with an inline # vX.Y.Z semver comment, never to
a mutable tag like @v4. The threat this closes: a compromised or malicious
action maintainer can repoint the tag at new code, giving arbitrary code execution
inside CI with that workflow's GITHUB_TOKEN privileges — a supply-chain attack
on the entire user base.
Pinning to an immutable commit SHA is enforced mechanically by the zizmor linter
in the workflow-lint job (.github/zizmor.yml config). Dependabot's
github-actions ecosystem entry maintains pins through reviewable pull requests,
so the operational trade-off — pins grow stale and miss security updates — is
managed rather than ignored: every update lands as a visible, auditable PR before
it ships.
The threat: a compromised or buggy build step can use the full scope of the
GITHUB_TOKEN. An over-broad top-level permissions grants those scopes to
every job and step, multiplying the blast radius.
The defense: .github/workflows/ci.yaml and release.yaml now grant
permissions: {contents: read} at the top level — the minimum viable scope —
and elevate scopes only on the specific job(s) that need them, in an explicit
per-job permissions block. Concrete reductions from earlier config:
- ci.yaml:
statuses: write(needed only for thewebjob to mark PR checks) was inherited by every job; it now lives onwebalone. - release.yaml: top-level
permissionswas{contents: write}, granting write access to every release job; it is now{contents: read}at the top, withcontents: writeelevated only to thegithub-releasejob.
Every job in every workflow has an explicit timeout-minutes to bound the
duration a compromised job can run (see below).
The threat: attacker-controlled PR text — github.event.pull_request.title,
.body, github.head_ref — can be interpolated directly into a run: shell
body without quoting, giving arbitrary shell execution. A malicious PR title
like "; rm -rf /; #" then becomes a command.
The safe pattern: pass attacker-controlled values through the env: block and
reference them as quoted shell variables ("$VAR", not $VAR), so they are
treated as data, not code. Example:
env:
PR_TITLE: ${{ github.event.pull_request.title }}
run: echo "Title is: \"$PR_TITLE\""The actionlint linter (run in the workflow-lint job) detects the unsafe form
— direct interpolation of github.event.* or github.head_ref into run: —
and rejects it. This repo does not use pull_request_target (the trigger that
runs workflows on untrusted fork code with the base repo's GITHUB_TOKEN),
closing the attack surface entirely: pull_request_target was designed for
workflows that must access repo secrets (e.g., automated releases), but it
introduces risk if any step trusts PR body content as code.
Every job carries an explicit timeout-minutes to bound job duration. The threat
is twofold: a compromised job could hang indefinitely (DoS on CI capacity and
cost), and a build failure could leave a job in a partially-modified state if the
termination is not clean.
Default timeout budget is ≤30 minutes. Documented exceptions with inline justification comments:
- The five e2e jobs in ci.yaml run at 60 minutes (the
e2e-gojob uses ajob_timeoutmatrix value;e2e-multicluster,e2e-upgrade,e2e-web-live,e2e-game-botset it directly). E2E test suites on kind clusters are inherently slow; 60 minutes is measured from prior runs. - publish-edge.yaml: the
imagesjob runs at 35 minutes (measured 31-minute historical max for full image build and push across all components).
The threat: CI failures on a public repository produce world-readable artifacts
(downloaded test logs, pod state dumps, $GITHUB_STEP_SUMMARY markdown). Any
unredacted secret — a pod env var, a log line, a manifest dump — becomes public
and compromised immediately. Multi-cluster environments and complex setups make
this harder to spot: a cluster dump is hundreds of lines and secrets can hide in
labels, annotation values, or environment variable lists.
The defense: .github/actions/dump-cluster-state/action.yml applies a redact()
filter at every emit boundary — before any data reaches $GITHUB_STEP_SUMMARY,
before any artifact is uploaded, before logs are written. The filter is
conservative: it redacts values (preserving keys so dumps stay debuggable)
by matching known patterns:
- Labelled key/value pairs: any value for keys containing
password,passwd,token,secret,api,key,bearer, orauthorization(case-insensitive, with optional dashes/underscores). - Bare tokens: JWT-shaped values (
eyJ...), PEM private-key blocks.
An important limitation: redaction is pattern-based and therefore best-effort.
A credential in an unrecognised shape — a long hex string, a custom token format,
or a value buried in a JSON log without a recognisable key — may not be caught.
No Kubernetes Secret object is ever collected into dumps (for obj in deployments statefulsets daemonsets jobs configmaps), so high-entropy database passwords and
OIDC secrets bound to the pod via Secrets are outside the dump scope entirely.
Operators should: treat CI artifacts as sensitive (not suitable for sharing with untrusted parties without review), verify no live credentials appear in failures, and rotate any that do immediately. This is not a substitute for not logging credentials in the first place.
Optional feature: the repository uses the CodeRabbit GitHub App for automated
code review (configured in .coderabbit.yaml). The app is structurally safer than
a self-hosted API-key reviewer:
- No repository secret in untrusted job: A self-hosted API-key reviewer would
require placing a repository secret (e.g.,
ANTHROPIC_API_KEY) inside a job that checks out untrusted fork code. If the fork is malicious, it can exfiltrate the secret from$GITHUB_TOKEN, env vars, or the runner's filesystem. The CodeRabbit app is a GitHub App, not a personal API key — it integrates via GitHub's OAuth flow and never places a repository secret alongside untrusted code. - PR content treated as data, not instructions: Review is advisory (cannot
block a merge) and does not execute instructions from PR body, title, or branch
name. However, the app's configuration (
.coderabbit.yaml) is read from the PR's HEAD branch — meaning a pull request, including one from a fork, can modifypath_instructionsandlabeling_instructionsand thereby change how the review is conducted. This is a real limitation of the GitHub App model: the bot's advice is weaker or differently-oriented on that PR than intended. It does not grant the PR repository permission, exfiltrate a secret (no repo secret is in the review path at all), or block or force a merge — the review remains advisory. The mitigation is the same as for any PR that edits CI workflows: human review should scrutinize changes to.coderabbit.yaml. - No implicit privilege escalation: The app cannot request scopes it was not granted at install time, and it cannot modify its own permissions.
The review is advisory — a human reviewer must still validate changes before merge — and the tool can be disabled at any time via GitHub's app management UI.
Usage telemetry has two ends, each with its own trust boundary. For what is sent and how to turn it off, see install.md. For running a receiver, see telemetry-provider.md. The project's data-handling statement is at https://gameplane.net/telemetry/.
A receiver's /ingest is open to the internet by design: installs have no
project-issued credential, so it cannot require one. (A self-hosted
receiver can require AUTH_TOKEN; the project's default receiver cannot.)
Anyone can therefore post reports. Defences:
- strict decoding: one JSON object, no duplicate or unknown keys, a 16 KiB body cap, and every string matched against a pattern or a fixed list, so hostile text can't create unbounded categories or reach the dashboard
- a per-source daily limit on accepted reports (
INGEST_SOURCE_DAILY_LIMIT, default20, tolerant of shared NAT) and a rate limit on the public summary, both held in memory only - optional proof-of-work (
INGEST_POW, off by default; see below), which raises the cost of bulk reports in proportion to the request rate - only daily aggregates and expiring activity records are stored; raw reports and source addresses are never written to disk
GET /v1/summary is off unless PUBLIC_SUMMARY=true. It returns five
fixed counts and nothing else: no breakdown, no per-day series, no install
ID. It is cached and rate-limited so public polling can't degrade ingest.
What this does not stop: someone can invent installs and send fabricated basic reports, within the per-source limits. Figures are approximate and self-reported, and the dashboard labels them that way.
When a receiver sets INGEST_POW=true, /ingest requires a solved
challenge from GET /v1/challenge in the Gameplane-Telemetry-PoW header.
Challenges are HMAC-signed with a key held only in the receiver's memory,
expire after 15 minutes and work once. The difficulty is INGEST_POW_MIN_BITS (default zero) while at most
than INGEST_POW_TARGET_PER_MIN challenges a minute are issued, then rises
with the rate (up to INGEST_POW_MAX_BITS, at most 26) and falls one bit per
five minutes. The check runs before the body is read, so a refused report
costs the receiver one HMAC and one hash, and it answers 428
(pow_required or pow_invalid) or, when its used-challenge set (1,000,000
entries) is full, 503 pow_busy. Challenge requests have their own
per-source limit.
What it stops: cheap bulk fabrication. The work a flood must do grows faster than its rate, while an install sending one report a day pays nothing under normal load and a few seconds during a flood. It also needs no per-install secret, so it works against an attacker who holds many source addresses, which the per-source limits do not.
What it does not stop:
- a determined or well-resourced sender: SHA-256 is cheap on GPUs and dedicated hardware, so it deters volume, not a patient attacker
- fabricated installs: solving a challenge proves work was done, not that an install exists, so invented installs remain possible and figures stay approximate
- many sources raising the difficulty for everyone: the difficulty is
global and capped by
INGEST_POW_MAX_BITS; the per-source challenge limit only stops a single source doing it alone - anything on the dashboard login: it has no proof-of-work (the dashboard uses no JavaScript); its defence is the credential minimum below
Proof-of-work does not replace signing or the per-source daily limit; all three apply. A restart drops the challenge key and the used set, so outstanding challenges fail once and installs retry with a new one.
The extended tier carries a random UUID. It is pseudonymous, not anonymous:
it links one install's reports over time. It is never derived from the
cluster, host, network or users; it is kept only on the install; an admin
can reset it at any time (the install then looks new to the receiver, and
the old ID lapses); and it is deleted when extended is turned off. The
receiver stores only HMAC-SHA256(pepper, installID) and never stores an
extended attribute (environment, games, features) next to it; those exist
only as daily aggregates.
Every extended report is signed (Ed25519) with a key derived from a random secret kept in the install's database combined with the install ID. The secret and private key never leave the install. The receiver binds an install ID to the first valid key that uses it (the claim) and then refuses, changing no stored data:
- a bad or missing signature (
403 bad_signature) - a send time outside
[now - 36h, now + 1h](403 stale) - a different key for a claimed ID (
409 id_claimed) - a send time not later than the last accepted one (
403 replay)
Signing stops impersonation of an existing install and replayed reports. It does not stop fabricated installs: anyone can generate a new ID and key pair, so a flood of invented installs is limited only by the per-source limits. A claim lasts as long as the activity record (90 days without a report by default), then the ID can be claimed again. If the receiver says an ID is claimed by another key, the install replaces its ID and resends once; any other refusal never rotates the ID.
The receiver's dashboard runs on its own listener (:8081), started only
when DASHBOARD_TOKEN is set; with no token there is no dashboard. Keep
that port off the public internet (the chart's NetworkPolicy admits only
configured peers). It uses one operator token from a Secret, which must be
at least 32 characters (generate it from 32 random bytes); the receiver
refuses to start with a shorter one. Per-source login limits can be
outrun by anyone holding many IPv6 /64s, so the token's length, not the
limiter, is what makes guessing infeasible:
- browser login compares the token in constant time and sets an
HttpOnly; Secure; SameSite=Strictsession cookie valid 12 hours; login and logout also require a same-originOrigin; logins are limited to 5 per minute per source - scripts can send
Authorization: Bearer <token>;/metricsaccepts only the Bearer token - the cookie key is derived from the token, so replacing the token invalidates every session and loses no data
- every dashboard response carries a strict
Content-Security-Policy, with no inline script or style, andCache-Control: no-store
Refusal invariant: an unauthenticated response never contains a figure,
a date other than page chrome, a version, category or module name, or
anything that varies with whether data exists. A wrong token gets the same
"Invalid credentials" login page whether or not the receiver is empty, and
an unauthenticated JSON request gets 401 {"error":"unauthorized"}. CI
compares the unauthenticated response against the login page of an empty
receiver.
The receiver uses the TCP peer address as the source for limits. It trusts
X-Forwarded-For only when the peer is inside TRUSTED_PROXY_CIDRS
(empty by default), so a client can't pick its own source by setting the
header. IPv4 addresses and IPv6 /64 prefixes are treated as one source
each. Sources live in memory only (capped at 100,000 per limiter,
least-recently-used evicted), are never logged to the database, and
counters reset at UTC midnight. If you run behind a proxy and leave
TRUSTED_PROXY_CIDRS empty, every client shares the proxy's address and
the per-source limits throttle everyone together.
ID_PEPPER keys the HMAC that turns an install ID into the stored
activity-record key, so a stored value can't be matched to an ID taken
from an install without the pepper. If unset, the receiver generates one
on first start and keeps it in its database, so it is as exposed as the
database file. Set it from a Secret to keep it out of backups of the
database. Changing the pepper makes every existing record unmatchable: all
installs look new for one day and their activity history is lost; no
install is locked out. See
telemetry-provider.md.
The API's <release>-api-read ClusterRole grants list on Pods in every
namespace, so the Cluster page can show per-node pod usage
(countNodePods in api/internal/handlers/cluster.go). A pod list returns
full pod specs, including plain-text environment values, for every workload
sharing the cluster — not just gameplane-games. A compromised API could
read them. The maintainer accepted this tradeoff on 2026-10-05 for the pod
meter. Values held in Secrets are not exposed: the API has no cluster-wide
Secret read. Without the grant, the API omits "used" and the dashboard shows
"—".
The optional MCP server [optional] (mcpServer.enabled, see mcp-server/README.md)
is strictly read-only — no tool it exposes can create, update, patch,
delete, or apply anything, enforced structurally (its tool handlers only
ever hold a client whose exported methods are List/Get-shaped) and by RBAC
(a ClusterRole granting only get/list/watch, plus get on
pods/log).
That RBAC grant is cluster-wide, not scoped to gameplane-games or any
other single namespace: the server can list/read Pods, Events, and pod logs
in every namespace, including kube-system and any other workload's
namespace sharing the cluster. Pod logs in particular can surface secrets
an application logs at startup or during errors (API keys, connection
strings, stack traces) — Kubernetes has no mechanism to redact those.
Combined with write-freedom and opt-in, admin-only installation
(mcpServer.enabled plus whatever gates kubectl exec access to the
gameplane-mcp-server pod), this is an accepted tradeoff, not an oversight
— but install it knowing that anyone who can reach a serve session gets
read access to cluster-wide pod state and logs, not just Gameplane-managed
namespaces. If that blast radius is wider than acceptable for a given
cluster, don't enable mcpServer there.
Secrets Gameplane reads or creates, by convention:
gameplane-<gameserver>-rcon— per-game RCON password, created by operatorgameplane-agent-ca— CA bundle the API trustsgameplane-agent-client— API's client cert/keygameplane-oidc— OIDC client secret (user-supplied)gameplane-backup-repo— restic repo URL + password (user-supplied)- audit-webhook auth — any Secret you reference via
api.audit.webhook.authSecretRef(user-supplied). The token is injected as an env var, never a flag, so it does not appear in the pod spec orpsoutput. - audit S3 credentials — any Secret you reference via
api.audit.s3.credentialsSecretRef(user-supplied). The access key and secret key are injected as env vars (GAMEPLANE_AUDIT_S3_ACCESS_KEY,GAMEPLANE_AUDIT_S3_SECRET_KEY), never flags, so they do not appear in the pod spec orpsoutput. - notification sinks — any Secret labelled
gameplane.local/notification-sink=truein the control-plane namespace (user-supplied; referenced by name from Admin Settings → Notifications, read by the API at delivery time — see notifications). - telemetry receiver credentials —
api.telemetry.receiver.dashboard.tokenSecretRef(dashboard and/metricstoken, at least 32 characters) andapi.telemetry.receiver.pepperSecretRef(install-ID pepper), both user-supplied and injected as env vars (DASHBOARD_TOKEN,ID_PEPPER), never flags. See telemetry-provider.md.
Rotation: deleting the -rcon secret triggers a reconciliation and
generates a fresh password on the next pod restart.
In a multi-cluster setup, each target cluster is referenced by a Secret containing its kubeconfig. Access to cluster credentials is protected by several layers:
- Embedded credentials only. The API and operator reject token files,
client certificate/key files, CA files,
execauthentication, andauth-providerplugins before creating a client. This applies to every entry, including unused contexts. Remote kubeconfigs must carry tokens or certificate/key/CA data directly; they cannot read control-plane files or run local authentication commands. - Label guard. The API only reads Secrets labelled
gameplane.local/cluster-kubeconfig=truewhen registering a cluster via the dashboard or API. This prevents a user from pointing at an arbitrary control-plane Secret (e.g., the OIDC client secret or backup credentials) and using it as a kubeconfig. - Delete guard.
DELETE /clusters/{name}drops the cluster's client at once and deletes the referenced Secret only when it is the one POST generates for that cluster (cluster--kubeconfig) and carriesgameplane.local/cluster-kubeconfig=true(Secrets created before the managed-by label was added are also cleaned up). Any other Secret, including one named for a different cluster or one without the kubeconfig label, is left in place. A kubeconfig Secret you create with kubectl or GitOps under another name is never deleted over HTTP. - Never logged or returned. The kubeconfig is never logged by the API, never echoed in responses, never visible in audit trails. It exists only to bootstrap the Kubernetes client for that cluster.
- Permission gating. Only users holding the
cluster:managepermission (admin-only) can register, list, or delete clusters via the API. Dashboard access to/clustersis similarly gated. - No implicit RBAC. Registering a cluster does not grant any user access to resources on that cluster. Access is determined by role bindings created independently on the target cluster, not by federation. See install.md.
When OIDC authentication is configured at install time via Helm values
(api.oidc.groupsClaim, api.oidc.roleMappings, api.oidc.defaultRole) (unreleased; ships in the next release),
Gameplane automatically assigns roles to users based on their OIDC provider's
group/role claims on every login. This eliminates the need for a bootstrap-admin
account in OIDC-only deployments — an operator can configure group mappings at
install time and the first user to log in receives the correct role immediately.
The security model consists of:
Core risk: Gameplane trusts the IdP's group claim unconditionally. Whoever controls IdP group membership effectively controls Gameplane role assignments. If an attacker compromises the IdP or its group directory (LDAP, Active Directory, cloud identity service), they can add themselves to a mapped group and gain that group's Gameplane role on their next login — up to and including admin access.
This is an accepted architectural boundary: the IdP is a trust root. If the IdP is compromised, Gameplane cannot defend against that. Mitigation is at the IdP level: strong authentication to the IdP, audit logging of group membership changes, and monitoring for suspicious group additions.
Install-time values (api.oidc.roleMappings.*) seed the role-mapping policy in
the database when the API starts. An admin can then override one or more roles'
group lists through the dashboard (PUT /admin/config/auth with helmOverride)
at runtime, without restarting the API or re-running Helm. Each role's effective
mapping is determined independently:
- Database override present (even if an empty list
[]): That list is the effective mapping for that role, used on every login. An empty list means "nobody maps to this role from any group" — a valid and meaningful override distinct from "no override set." - Database override absent: The Helm-seeded value is the effective mapping for that role.
When an operator runs helm upgrade and changes a Helm value (api.oidc.roleMappings.*),
the new value updates the seed — but it does NOT overwrite a database override that
has already been set for that role. The override persists until explicitly reset via
the dashboard (DELETE /admin/config/auth/role-mappings/{role}). This is deliberate:
Helm upgrades should not silently undo admin customizations made through the UI.
Consequences for operators: the effective role mappings in the dashboard may not
match what is in values.yaml after one or more roles have been dashboard-overridden.
To audit what is actually configured, consult the dashboard's /admin/config view
(installTimeSettings.oidcHelmProvider shows the Helm seed, auth.helmOverride
shows any database overrides) rather than relying on Helm values alone.
When resolving a user's role, Gameplane matches the user's groups against every
role's effective mapping (seed + overrides merged) and assigns the highest
privilege match: admin > operator > viewer. This is applied after the
per-role merge of Helm seed and database override, so a user matching both an
overridden (database-managed) viewer group and a Helm-seeded admin group still
resolves to admin.
Overriding a lower role does not revoke a higher one. Example: if the database
overrides the viewer list to [] (nobody maps to viewer), but the Helm-seeded
admin list is ["admins"], a user in the "admins" group still resolves to admin
on the next login — the admin mapping was never overridden.
On each OIDC login, Gameplane:
- Extracts the user's group membership from the OIDC token's group claim (configured
via
api.oidc.groupsClaim; defaults to"groups"). - Reads the effective role mappings: the Helm-seeded values, merged with any database overrides for each role independently.
- Matches the user's groups against the effective mappings to compute their role.
- If no role matches, assigns the default role (configured via
api.oidc.defaultRole; defaults toviewer; can be set todenyto reject login).
This re-evaluation runs whenever the effective role mappings exist: Helm-seeded
api.oidc.roleMappings (at least one non-empty role array), a dashboard
helmOverride.roleMappings overlay (which counts even when the Helm values set
no mappings), or the mappings of a dashboard-managed provider. If none is
configured, new OIDC users receive the fixed viewer role and existing users'
roles are never re-evaluated.
Two guards prevent lockout during re-evaluation:
- No lockout rule at login: If re-evaluating a user's role would remove the
last user able to manage users (hold the
users:managepermission), the re-evaluation is not applied — that user retains their old role and can still manage other users. This prevents an operator from accidentally creating an unrecoverable lockout via an override change. - Break-glass mechanism: If role mappings are misconfigured such that nobody
can reach admin, an operator can run the
bootstrap-adminbreak-glass command to create a local admin account and fix the mappings. Bootstrap-admin and OIDC-mapped admin accounts coexist peacefully.
Gameplane always warns when an operator configures or changes a role mapping that includes a group. The warning text states: "Be aware that an OIDC group may include a large number of users, and assigning it to the admin role grants admin access to all members of that group." This warning is unconditional — it appears on every such configuration change, not just when the operator first enables role mappings.
Why unconditional: Gameplane cannot enumerate OIDC group membership — it cannot tell whether a group name refers to a 3-person team or a 3000-person organization. An attacker with dashboard access who knows the group structure could configure a mapping for an unexpectedly large group to gain access. Operators must be aware of this risk at every configuration step. The warning does not prevent the change, but it ensures operators cannot claim they were not aware of the risk.
Every role assignment driven by OIDC group mappings (on initial login or re-evaluation)
is recorded in audit_events with the matched group name and role transition:
- Action: OIDC login with role assignment
- Target: the user (subject of the OIDC token)
- Details recorded:
- Which OIDC provider performed the assignment (
"helm"for the Helm-seeded provider, otherwise the dashboard-managed provider's name) - Which group matched a mapping rule (or
"none"if no mapping matched) - The user's old role (
"new_user"on first login, or the previous role) - The assigned role (
"viewer","operator","admin", or"denied"if rejected)
- Which OIDC provider performed the assignment (
Examples of what gets logged:
- First login, matched admin group: user created with admin role from group membership
- Re-evaluation, no mapping match: role re-evaluated to default role on next login
- Subsequent login, role upgraded: user's role changed from viewer to operator based on new group membership
Every dashboard change to a role's override — writing a new group list via
PUT /admin/config/auth or resetting it via DELETE /admin/config/auth/role-mappings/{role}
— is also audited as a configuration change:
- Action: Role mapping override write or reset
- Target: the affected role (
"admin","operator", or"viewer") - Details recorded:
- Which admin made the change (from session)
- The new or reset value (the group list, if changed)
This allows operators to track who changed what mappings and when, and to correlate unexpected role assignments with dashboard configuration changes.
api.oidc.groupsClaim (the claim name from the OIDC token that holds group
information) and api.oidc.defaultRole (the fallback role when no group matches)
are configured via Helm values only — there is no dashboard write path for either
in v1.
Why this matters for security: An attacker with dashboard admin access can
override individual role mappings but cannot repoint the group claim to a
different OIDC token field (e.g., changing from "groups" to a field they control),
nor can they change the default role fallback to "deny" to lock everyone out. These
two settings remain under the operator's full control, via Helm values only, and
require a helm upgrade to change — an out-of-band action that can be audited and
gated by access controls on the cluster itself (e.g., who can run Helm in production).
No internal infrastructure metrics are displayed on the login page or
any other unauthenticated surface. This is a hard requirement — see
web/src/routes/Login.tsx for the enforcement.
The API's Prometheus metrics follow the same rule. They are served on a
dedicated listener (--metrics-addr, chart value api.metricsPort,
default 9090), never on the public API port that the Ingress and the web
front end route to, so /metrics on the dashboard host answers 404. The
chart's ServiceMonitor scrapes the metrics port from inside the cluster.