This document outlines the security architecture integrated into the operator-go SDK. It adopts a defense-in-depth approach, split into two primary layers:
- Application Security: Focused on safely injecting sensitive data (Secrets, Keys) into workloads.
- Infrastructure Security: Focused on securing the Kubernetes execution environment (RBAC, Service Accounts, Pod Constraints).
The core design philosophy is "Zero-Touch Security". The Product Operator does not directly handle sensitive data; it delegates provisioning to a specialized secret-operator.
SecretClass is a resource managed by secret-operator. It defines "how" to obtain security artifacts, while the workload (Pod) simply declares "what" it needs by referencing a SecretClass by name. The CRD itself — its scope and schema — is owned by the secret-operator, not by this SDK; operator-go only emits the secrets.kubedoop.dev/class: <name> annotation and never reads the object.
This mechanism is implemented using the Kubernetes CSI (Container Storage Interface). The secret-operator provides a CSI driver that intercepts volume mount requests, generates or retrieves the required secrets on-the-fly, and injects them into the container file system as files.
- Definition: Admin creates a
SecretClasscontaining the policy (e.g., "Issue certificates using ClusterIssuer 'my-ca'"). - Reference: The Product CR (e.g., HdfsCluster) specifies
secretClass: "hdfs-secret-class". - Declaration: The Product Operator registers the need on a
security.SecretProvisionerand appends it toRoleGroupBuildContext.VolumeProviders(or callsAutoInject(stsBuilder)). The SDK then emits, on the Pod template, a generic ephemeral volume whosevolumeClaimTemplatecarries the secret-operator annotations (secrets.kubedoop.dev/class, and…/scope,…/format, … when set) and thesecrets.kubedoop.devStorageClass. Kubernetes' ephemeral-volume controller materializes one Pod-owned PVC per Pod, so the operator needs no PVC create RBAC and the PVC lifecycle is Pod-bound. This step is opt-in: nothing is mounted unless the product registers a volume. - Injection: When a Pod starts, the CSI driver calls the backend to generate artifacts (TLS
certs, Keytabs) and mounts them read-only under the SDK's canonical base,
/kubedoop/mount/<volumeName>(constant.KubedoopMountDir, overridable per provisioner withSecretProvisioner.WithMountBasePath).
Never hardcode the mount path. Ask the provisioner:
provisioner.Path("server-tls")(orMustPath) returns/kubedoop/mount/server-tls, and stays correct when the base path is overridden. Some helpers deliberately use a different base — e.g.s3.ConnectionInfo.CredentialsProvisionermounts underconstant.KubedoopSecretDir(/kubedoop/secret/<volume>) — which is exactly why the path must come from the API.
security ships one constructor per common need. Each requests a 10Mi PVC and sets the scope shown
below (CredentialsVolume deliberately sets none — credential secrets are usually class-resolved;
add one with WithScope, e.g. from ScopeString):
| Constructor | Format | Default scope |
|---|---|---|
TLS(volumeName, secretClass) |
tls-p12 |
pod,node |
TLSPEMFormat(volumeName, secretClass) |
tls-pem |
pod,node |
ServiceTLS(volumeName, secretClass, serviceName) |
tls-p12 |
pod,node,service=<serviceName> |
KerberosVolume(volumeName, secretClass, serviceName, …) |
kerberos |
pod,node |
ListenerVolume(volumeName, secretClass, listenerVolumeName, format) |
caller-supplied | listener-volume=<listenerVolumeName> |
Custom(volumeName, secretClass, format) |
caller-supplied | pod,node |
CredentialsVolume(volumeName, secretClass) |
none | none |
Named scopes must carry a name. The secret-operator parses service and listener-volume
entries as key=<value> and unconditionally reads the value, so a bare service or
listener-volume entry is unresolvable. The SDK refuses to build one:
ListenerVolumerequireslistenerVolumeName(the name of the Pod volume that mounts the listener, which the secret-operator resolves to that listener's addresses) and panics on an empty string. The emitted scope islistener-volume=<listenerVolumeName>.SecretVolumeRegistration.WithScopeandSecretProvisioner.Registerpanic on a scope whoseservice/listener-volumeentry has no value, and on an empty comma entry. Unknown scope keys pass — the scope vocabulary belongs to the secret-operator and may grow ahead of this SDK.security.ScopeString(*commonsv1alpha1.CredentialsScope)renders a CRD-declared scope (node,pod,services,listenerVolumes) into that annotation value, emitting thekey=prefix for the named entries. It returns""for a nil/empty scope, in which case the annotation is omitted.
A scope name may contain neither , nor =. The annotation is one comma-delimited string of
key=value entries with no escaping, so a name carrying either character does not quote itself —
it adds scopes. services: ["mysvc,node"] renders service=mysvc,node, which the
secret-operator reads as a service scope and a node scope: the CR author silently receives a
certificate covering the node's hostname and IP, and a reviewer reading the CR sees nothing
unusual. Two layers stop that:
CredentialsScope.Servicesand.ListenerVolumescarry+kubebuilder:validation:items:Pattern(^[^,=]+$) anditems:MinLength=1, so the API server rejects it atkubectl apply— where the user sees the error and can fix it. This is the real defence.ScopeStringdrops any entry it cannot render as itself, covering what admission cannot: a CR stored before those markers existed, and a scope built in Go. Dropping is the safe direction — splicing grants a broader credential than requested, invisibly, while dropping withholds a requested scope, which surfaces as the application rejecting the certificate.
The default PKCS12 password (changeit) is stored as a PVC template annotation and is therefore
readable by anyone with get on PVCs. Use WithPassword or WithNoPassword.
Certificate rotation is configured on the registration (WithCertLifetime, WithCertJitter,
WithCertBuffer), which the secret-operator honours when issuing the artifact.
Listener volumes themselves are declared through listener.NewVolume(name, class) plus optional
.WithListenerName(name) on a listener.ListenerProvisioner. pkg/listener has no scope API —
there is no WithScope, no ListenerScope type and no listeners.kubedoop.dev/scope annotation on
listener PVC templates. Scoping a secret to a listener is done on the secret side, with
ListenerVolume above.
The backends below are implemented by the secret-operator; the SDK's part is declaring the volume
and its annotations (§2.1.2) and resolving the mount path. Backend selection is a property of the
SecretClass the admin creates, not of the SDK call.
The mechanisms described in §2.2.1–§2.2.3 are
secret-operatorbehavior, documented here so product authors know what the platform provides. They are not implemented inoperator-goand cannot be verified against this repository — consult thesecret-operatordocumentation for the authoritative contract. Only thesecurity.*API names and mount paths in §2.1 are SDK behavior.
Calculates and issues TLS certificates for components.
- Scenario: Internal mTLS communication (e.g., DataNode <-> NameNode) or external HTTPS access.
- Mechanism:
- Automatically generates SANs (Subject Alternative Names) based on Pod DNS names (e.g.,
*.hdfs.svc.cluster.local). - Solves the comprehensive trust problem: Components from different products (e.g., Flink connecting to HDFS) can trust each other if they use
SecretClassessigned by the same Root CA.
- Automatically generates SANs (Subject Alternative Names) based on Pod DNS names (e.g.,
Automates Kerberos integration for Hadoop/Big Data ecosystems.
- Scenario: Secure clusters requiring Kerberos authentication.
- Mechanism:
- Dynamic Principal: Supports generating principals based on the Pod's specific hostname (e.g.,
nn/hdfs-namenode-0.hdfs.svc@REALM). This is critical for K8s StatefulSets where Pod names are deterministic but distinct. - Keytab Injection: Generates the keytab on the KDC and securely mounts it to the container.
- Dynamic Principal: Supports generating principals based on the Pod's specific hostname (e.g.,
Searches and injects existing Kubernetes Secrets or ConfigMaps.
- Scenario: Legacy applications or reusing existing static secrets.
- Mechanism:
- Searches for resources matching specific labels or names in the cluster.
- Security Benefit: The Product Operator does not need
LIST/WATCH/GET Secretpermissions for the entire namespace. Only the privilegedsecret-operatoraccesses the data, minimizing the attack surface.
Unlike §2.2, OIDC is not a secret-operator CSI backend in this SDK. It is delivered by the
sidecar.OAuth2ProxySidecarProvider, which terminates the login flow in an oauth2-proxy sidecar in
front of the product's HTTP port.
- Scenario: Workloads requiring modern authentication (e.g., a product Web UI behind an IdP).
- Mechanism:
- Configuration source: an
AuthenticationClasswith anOIDCProvider(pkg/apis/authentication/v1alpha1). The product resolves the class and passes the provider tosidecar.NewOAuth2ProxySidecarProvider(oidcProvider, clientCredentialsSecret, upstreamPort, opts...). - Credential injection: a plain Kubernetes Secret named by
clientCredentialsSecret, carrying the keysCLIENT_ID,CLIENT_SECRETandCOOKIE_SECRET(sidecar.OIDCClientIDKey/OIDCClientSecretKey/OIDCCookieSecretKey). All three reach the sidecar asOAUTH2_PROXY_*env vars viasecretKeyRef— they are not mounted through the secret-operator CSI driver, and they are never written into a config file or inlined into the PodSpec.Validatefails the reconcile before the StatefulSet is applied if the Secret is missing, or if it does not carry the cookie-secret key. - Configuration of the proxy: the provider sets the
OAUTH2_PROXY_*env vars (issuer URL, scopes, provider hint, upstream, listen address, PKCES256). The SDK does not inject JVM system properties or configure the product application's own OIDC module — a product that needs in-application OIDC generates that configuration itself. - Cookie secret: read from the Secret above (relocate it with
WithOAuth2ProxyCookieSecretRef). It is never derived from the CR and never inlined: this value signs every session cookie the proxy subsequently trusts, so anything readable through the API — a UID, a name, an envvaluein the PodSpec — would let a reader forge a session and bypass authentication entirely. Generate it once withsidecar.GenerateCookieSecret()and store it; the value must stay stable, or each reconcile would roll the pods and log every user out. - Authorization: authenticating against an IdP is not authorization for this cluster. On a
shared realm, every account the IdP can issue a token for would otherwise reach the product, so
the provider requires an explicit policy —
WithOAuth2ProxyEmailDomains(...)to restrict logins, orWithOAuth2ProxyAllowAllEmails()to admit everyone deliberately. Exactly one: building a proxy with neither fails, and so does building one with both, because "allow all" would otherwise win and silently discard the domain list. Post-login redirects are likewise closed by default (the proxy's own host only);WithOAuth2ProxyWhitelistDomains(...)widens that, one domain at a time. - Probes: the sidecar carries a readiness probe and a liveness probe on
/ping— never/ready, which oauth2-proxy documents as a deep health check and which would make a runtime IdP outage evict the pod from every Service. Readiness is the right gate here, and this is an availability property with a security edge: this container terminates client traffic, but pod readiness is otherwise decided by the main container's probe on the product's own port, so without it the pod joined its Services — and received requests — while the proxy was still starting. Requests arriving then are refused rather than authenticated, which on a rollout looks to clients like an outage and to an operator like a working auth layer. There is deliberately no startup probe: as a native sidecar, a proxy that could never satisfy it would stop the product's own container from ever starting.
- Configuration source: an
This layer covers both directions: how the Operator constructs the Kubernetes Pods and Resources to minimize the attack surface (§3.1, §3.2, §3.4), and what the Operator process itself must be granted in order to do so (§3.3).
A Product Cluster managed by the SDK can operate with its own distinct identity — distinct both from
the namespace's default account and from the ServiceAccount the operator process itself runs
as, whose permissions are a separate axis entirely (§3.3). This section gives the workload an
identity; §3.2 gives that identity its permissions.
-
Unconditional Provisioning: every cluster the framework reconciles gets a workload
ServiceAccount. The reconciler creates (or updates) it in the CR's namespace with the CR as controller owner at step 0 of the reconcile — before any extension hook or role is processed — and propagates the name throughRoleGroupBuildContext.ServiceAccountNameinto the Pod template. There is no switch to turn it off: a ServiceAccount is pure metadata, and the opt-in it replaced mostly produced clusters running as the namespacedefaultby accident, which is the opposite of what this section promises. -
The name is DERIVED, not configured:
reconciler.ServiceAccountResourceName(kind, cluster)yields"<lowercased kind>-<cluster>"— anHdfsClusternamedprodruns ashdfscluster-prod. The Kind is part of it because a CR name alone is not unique in a namespace: anHdfsClusterand aTrinoClusterboth namedprodwould otherwise select one ServiceAccount and the second controller could never take ownership of it.This follows the framework's general rule that it owns the name of every slot it owns the lifecycle of (
docs/architecture.md§3): the framework creates the object, controller-owns it and garbage-collects it with the CR, so nothing needs to address it by a product-chosen name. It also removes a whole failure class rather than detecting it — a configurable static name was a constant, so every CR of a product in one namespace resolved to the SAME ServiceAccount: the second cluster failed withAlreadyOwnedErrorforever, and deleting the first garbage-collected the SA out from under the second's running pods. -
Granularity is per CR, not per RoleGroup. One identity per cluster; there is no per-role or per-role-group ServiceAccount.
-
Scope:
BaseRoleGroupHandlerbinds this ServiceAccount to the Pod template (frombuildCtx.ServiceAccountName), so pods run as it and Kubernetes audit logs reflect the specific application identity rather than a genericdefaultaccount. A product implementingRoleGroupHandlerdirectly must bind it itself — the framework creates and owns the ServiceAccount but does not verify that the built StatefulSet uses it. -
Predictability is part of the contract:
ServiceAccountResourceNameis exported so that anything outside the operator that must name this ServiceAccount — a hand-written RoleBinding, an admission policy, an audit query — derives it from the formula rather than from the operator's source. Call the exported function rather than re-implementing the concatenation: past 253 bytes (the DNS-subdomain limit) it truncates and appends a deterministic 8-character hash, so a hand-derived name for a long CR name would refer to a ServiceAccount that does not exist. -
Not configurable — but reachable through
podOverrides: the common CRD types (GenericClusterSpec,RoleGroupConfigSpec) carry noserviceAccountNamefield, so there is no way to configure a different identity.podOverridesis a different matter: it is strategic-merged onto the assembled pod template last, sopodOverrides.spec.serviceAccountNamereplaces the derived name outright. What the framework guarantees is this identity's existence, name and lifecycle — not that the pods use it. Treat write access to the CR as write access to the workload's identity, and if that matters in your threat model, constrain it at admission rather than expecting the SDK to.
Workloads often need to interact with the Kubernetes API (e.g., a Flink JobManager creating Jobs, a Spark driver creating executor pods, a member electing a leader through a Lease).
This is the other half of §3.1. That section gives the workload an identity; this one gives that identity its permissions. The framework maintains both at the same point in the same way — once per CR, at cluster level, before any role is built — and the role group merely consumes the result.
-
The product declares rules; the framework owns everything else. Set
GenericReconcilerConfig.WorkloadRBACRules func(cr CR) []rbacv1.PolicyRule. The framework maintains a namespacedRoleandRoleBinding, named after the derived ServiceAccount, in the CR's namespace, with the CR as controller owner. The split is deliberate: only the product knows what its pods call, and only the framework can guarantee the naming, the ownership, the idempotent apply and the revocation.WorkloadRBACRules: func(cr *v1alpha1.NifiCluster) []rbacv1.PolicyRule { return []rbacv1.PolicyRule{{ APIGroups: []string{"coordination.k8s.io"}, Resources: []string{"leases"}, Verbs: []string{"get", "list", "watch", "create", "update"}, }} },
-
The framework passes the ServiceAccount name it derived itself, so the identity and its permissions cannot drift apart. That is why this is a config field rather than only the exported
reconciler.EnsureWorkloadRBAChelper: the helper takes the name as a parameter, so a caller can pass one that is not the workload's, producing a Role that grants to a ServiceAccount no pod uses — two objects that both exist, pods that start fine, and a 403 on the first API call. The helper stays exported for a product driving the reconciler's pieces itself. -
Namespaced only, by construction. There is no
ClusterRolepath: a namespaced CR cannot controller-own a cluster-scoped object, so the framework would have no lifecycle for one — no garbage collection with the cluster, and no ownership gate on the reclaim. A product that genuinely needs cluster-scoped permissions maintains those objects itself, including their cleanup. -
An empty rule set REVOKES. Returning
nilor an empty slice deletes both objects (gated on this CR controller-owning them). Like every other optional slot in the framework — a nilMetricsServicereclaims the Service — nil is an instruction, not "leave it alone". Without that reading, narrowing a rule set to nothing would be the one change that silently did nothing while the pods kept permissions the product had stopped granting. A nil hook is different from a hook returning empty: nil means the product never opted in and no RBAC object is touched at all. -
A pre-existing RoleBinding is never re-pointed.
roleRefis immutable, so a RoleBinding already at this name pointing elsewhere cannot be converged. The framework fails with a*reconciler.ValidationErrornaming both refs and the command that fixes it, rather than rewriting only the subject — which would hand this cluster's pods whatever the old ref allows, and report success.One already pointing at this cluster's Role is the other case, and it is adopted — that is the intended migration path off a hand-maintained or Helm-installed binding, and a hand-written one usually does match, since both the binding's name and the Role's are the derived ServiceAccount name. Adoption is not gentle, though: the subjects are replaced with the single derived ServiceAccount, the labels are overwritten, and the CR takes the controller reference, so the object is thereafter garbage-collected with the cluster. Anything else that binding was granting — a second subject, a CI account — disappears on the next reconcile.
-
The operator must hold what it grants. Kubernetes refuses to let a subject grant permissions it does not itself hold, so the operator's own ClusterRole must be a superset of every rule passed here. This is invisible at compile time and invisible in any test running as cluster-admin (envtest does), so it needs two kinds of marker:
// the rules being granted — Kubernetes forbids granting what the granter lacks // +kubebuilder:rbac:groups=coordination.k8s.io,resources=leases,verbs=get;list;watch;create;update // write access to the RBAC API itself, without which nothing here can be created at all // +kubebuilder:rbac:groups=rbac.authorization.k8s.io,resources=roles;rolebindings,verbs=get;list;watch;create;update;patch;delete
A 403 has those two distinct causes needing opposite fixes, so the SDK re-explains the API server's own message rather than pre-checking: that message has already computed rule covering against the operator's real effective permissions (wildcards,
resourceNames, aggregated ClusterRoles) and names the missing rule.The escalation check runs twice, on two writes, with two different bypass verbs. The Role create/update is checked against
escalateonroles; the RoleBinding create is checked separately, and its bypass verb isbindon the referenced role. Soescalatealone is not an escape hatch — it is the worst of both: the Role is created, the RoleBinding is refused, andensureWorkloadRBACthen fails at step 0b on every pass, before any hook or role runs, leaving a Role bound to nothing. If you take the hatch at all it isverbs: ["bind", "escalate"]; the supported answer remains the superset. -
Watches are registered only when the hook is set. An unconditional
Owns(&rbacv1.Role{})would force every operator built on this SDK to grant itself cluster-widelist;watchon roles and rolebindings — and a forbidden informer failsWaitForCacheSyncfor all sources, somanager.Startreturns and the process exits. Operators that never use the feature pay nothing. -
The RBAC builders remain (
builder.RoleBuilder,RoleBindingBuilder,ClusterRoleBuilder,ClusterRoleBindingBuilder) for products assembling RBAC objects the framework does not maintain. They write nothing to the API server.
§3.1 and §3.2 are about the workload's identity and permissions. This section is the other axis:
the permissions the operator process must hold, because GenericReconciler performs API writes
on the operator's own ServiceAccount on behalf of every CR it reconciles.
The SDK cannot declare these for you. controller-gen only walks the packages named in its paths
argument and never descends into a dependency, so a +kubebuilder:rbac marker inside pkg/ would
generate nothing anywhere. The framework consumes the permissions; the adopting operator declares
them. That split is the entire reason this section exists — before it, the only way to derive this
set was to read examples/trino-operator's markers and hope they were complete.
The two axes differ in provenance, not just in subject, and that is the reliable way to tell them
apart. The objects in §3.1 and §3.2 are never authored by anyone: the reconcile loop computes them
per CR, creates them on the next pass, revokes them when the rule set empties, and lets garbage
collection reclaim them with the cluster. Nothing about them is in a YAML file in your repository.
The permissions in this section are the mirror image — they are authored, as +kubebuilder:rbac
markers in your own controller package; make manifests renders them into config/rbac/role.yaml;
your kustomize overlay or helm chart deploys that as a ClusterRole bound to the operator's own
ServiceAccount, once per install. It exists before any CR does, outlives every CR, and changes only
when you regenerate and redeploy.
So an adopter who adds a conditional grant from §3.3.2 and does not run make manifests and
redeploy gets exactly the 403 this section exists to prevent. And if your repository also ships a
helm chart, its templates/clusterrole.yaml is a second copy that nothing regenerates — keep it in
step deliberately, or the deployed operator runs on rules that no longer match its markers.
Derived from the framework's own call sites, not from the example. It describes what the code does today; when the framework's API usage changes, this table is what must be updated with it.
// The CR itself. The framework Gets it at the top of every pass and watches it via For().
// It never writes the CR BODY — only the status subresource. See "granting the write back" below
// if your operator registers a finalizer.
// +kubebuilder:rbac:groups=<your.group>,resources=<yourplurals>,verbs=get;list;watch
// +kubebuilder:rbac:groups=<your.group>,resources=<yourplurals>/status,verbs=get;update;patch
// +kubebuilder:rbac:groups=<your.group>,resources=<yourplurals>/finalizers,verbs=update
// The workload the framework builds and reclaims.
// +kubebuilder:rbac:groups=apps,resources=statefulsets,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=core,resources=services;configmaps,verbs=get;list;watch;create;update;patch;delete
// +kubebuilder:rbac:groups=policy,resources=poddisruptionbudgets,verbs=get;list;watch;create;update;patch;delete
// The workload identity (§3.1). Every cluster gets one and NOTHING ever deletes it.
// +kubebuilder:rbac:groups=core,resources=serviceaccounts,verbs=get;list;watch;create;update;patch
// Health evaluation: one label-selected List per pass. `get` is there for the exported
// util.ExecUtil.PodIsReady, which the framework itself never calls. NOT self-announcing — §3.3.3.
// +kubebuilder:rbac:groups=core,resources=pods,verbs=get;list;watch
// Orphaned PVCs, when the delete-pvcs annotation is set on the CR at runtime (§3.3.2).
// +kubebuilder:rbac:groups=core,resources=persistentvolumeclaims,verbs=get;list;watch;delete
// Events. NOT self-announcing — see §3.3.3.
// +kubebuilder:rbac:groups=core,resources=events,verbs=create;patchTwo verbs are deliberately absent, and both absences remove a capability the operator genuinely does not need:
- No
deleteonserviceaccounts. The SDK has no code path that deletes one — it carries a controller owner reference and is reclaimed by garbage collection with the CR.deleteis not reachable through any other verb, so withholding it is a real reduction. Collapsingservices;configmaps;serviceaccountsonto one marker line is how it gets granted by accident. - No
update/patchon the CR body. The framework Gets the CR and writes onlyStatus().Update. An operator that can rewrite its users'specis a materially different trust proposition from one that cannot, so this is worth withholding.
Everything else matches what kubebuilder scaffolds, including patch on kinds the framework
only ever Updates, and that is deliberate rather than sloppy. In RBAC, patch alongside update
grants no additional capability: anything achievable with a PATCH is achievable with a
read-modify-write PUT. So omitting it would reduce privilege by exactly zero while costing three
real things — util.K8sUtil.Patch, a helper this SDK exports for product code, would 403; any future move to server-side apply would 403; and every adopter's
diff against the scaffold would grow noise for no benefit.
The rule this section follows, therefore: omit a verb only when omitting it actually removes a capability. "The framework does not call it" is a fact worth recording — and §3.3.3 depends on knowing exactly which calls happen — but it is not by itself a reason to withhold a grant.
Several rows grant list;watch on kinds the framework never Lists — core/secrets and
s3.kubedoop.dev/* in §3.3.2 are read exclusively through Get (Dependencies validation,
FetchSecret, the oauth2-proxy Validate, the S3 reference resolvers). That is not sloppiness, and
tightening it to get is the one "obvious" correction that breaks an operator.
The framework reads through mgr.GetClient(), which is cache-backed: a read is served from an
informer, and a read of a kind that has no informer yet lazily creates one — which LISTs and then
WATCHes. So a code path that only ever calls Get still requires get;list;watch, and it fails at
the first read, mid-reconcile, rather than at startup.
The consequence worth weighing before you grant it: an informer is cluster-wide and unfiltered by
default, so core/secrets does not mean "the operator can read the one Secret it needs" — it means
the operator LISTs and caches every Secret in every namespace in its own memory. That is a memory
cost on a large cluster and a much larger blast radius if the operator is compromised. Scope it in
your manager options if that matters:
Cache: cache.Options{ByObject: map[client.Object]cache.ByObject{
&corev1.Secret{}: {Namespaces: map[string]cache.Config{"my-namespace": {}}},
}}The same is true of any ExtraResources kind you cache, and it is why the s3 rows are worth
skipping entirely for a product whose users always write inline:.
Granting the CR-body write back. The SDK registers no finalizer, but it explicitly contemplates
products that do (a product finalizer is one of the two paths that leave a CR readable with a
deletionTimestamp). Adding or removing a finalizer writes metadata.finalizers, which is the CR
body — so an operator that registers one needs
+kubebuilder:rbac:groups=<your.group>,resources=<yourplurals>,verbs=get;list;watch;update;patch.
That is a conscious re-grant, not scaffolding.
<yourplurals>/finalizers is required for a reason its name does not suggest — the SDK registers
no finalizer anywhere. controllerutil.SetControllerReference stamps blockOwnerDeletion: true
on every owner reference the framework writes, and the OwnerReferencesPermissionEnforcement
admission plugin gates that field on update of the owner's finalizers subresource. It is
therefore required on clusters running that plugin (OpenShift enables it) and inert everywhere else.
Keep it, and keep this sentence with it, or the next reader will delete it as dead.
Grant these only if you use the feature. Each row names its exact trigger, because the point of publishing a baseline is that adopters can stop copying permissions they do not need.
| Grant | Needed when |
|---|---|
rbac.authorization.k8s.io/roles;rolebindings — get;list;watch;create;update;patch;delete |
WorkloadRBACRules is set (§3.2) — plus every rule your hook returns, since Kubernetes forbids granting what the granter lacks. That second half cannot be tabulated here, because it is whatever your product passes; without it the operator 403s at step 0b on every pass, before any hook or role runs. A nil hook registers neither the watches nor any write. |
core/secrets — get;list;watch |
Dependencies returns a DependencySecret, the oauth2-proxy sidecar is registered, or a handler calls FetchSecret. |
core/secrets — get;list;watch;create;update;patch |
A product calls EnsureGeneratedSecret (§4.9.4 in architecture.md) — use this row instead of the one above. It is effectively mandatory with oauth2-proxy, whose Validate fails when the cookie key is missing. |
core/persistentvolumeclaims — get;list;watch;delete |
Listed in the baseline above because of the trap below, not because every operator reclaims PVCs. |
core/pods/exec — create |
A product builds util.NewExecUtil (e.g. an in-container ServiceHealthCheck). This is arbitrary command execution in the product's pods; it is deliberately not in the baseline. |
s3.kubedoop.dev/s3connections;s3buckets — get;list;watch |
A product resolves S3 through pkg/s3 and users write reference: rather than inline: — the inline branch performs no I/O. |
your ExtraResources kinds — get;list;watch;create;update;patch;delete |
A handler ships RoleGroupResources.ExtraResources. The list;watch half is load-bearing at startup, not only for cleanup: these kinds are registered through SetupWithManagerOptions.ExtraOwns. |
Both write rows carry patch for the same reason the baseline does — these paths are
controllerutil.CreateOrUpdate like every other, so next to update the verb grants nothing
extra, while omitting it would 403 the exported util.K8sUtil.Patch. The read-only rows and
persistentvolumeclaims deliberately do not get it: the PVC row has no update either, so
there patch would genuinely add the ability to modify a claim.
The PVC trap. operator.zncdata.dev/delete-pvcs is set by whoever operates the cluster, on the
CR, at runtime — not by the operator's author at build time. "We do not use that feature" is
therefore not a decision you get to make: without the grant, a user who sets the annotation gets
silent non-reclamation. Cluster deletion never touches PVCs (the SDK sets no finalizer and no
retention policy), so this is strictly the orphan path.
Two more the framework never calls, listed because they are part of a working operator and no
call-site scan of pkg/ reveals them — both are wired in your main.go:
coordination.k8s.io/leases (get;create;update) under --leader-elect, and
authentication.k8s.io/tokenreviews + authorization.k8s.io/subjectaccessreviews (create) when
the metrics endpoint is protected.
Most grants above fail loudly, which is why those get fixed on first deployment. Two mechanisms do that, and which one applies depends on where the watch came from:
- a kind registered with
Owns()—statefulsets,services,configmaps,serviceaccounts,poddisruptionbudgetsand yourExtraOwnskinds — starts its informer at boot. A forbidden one failsWaitForCacheSyncfor all sources, somanager.Startreturns and the process exits (§3.2 relies on the same mechanism). You cannot miss it. - a forbidden create/update returns an error that fails the role group and sets
Degradedwith the API server's own message.
A lazily-created informer is the quieter middle case: the kinds read only through Get
(§3.3.1's "Why a Get needs list;watch") have no Owns() line, so nothing fails at boot and the
403 surfaces on the first read instead — during a reconcile, on whichever cluster happened to
trigger it.
Three things are quieter still, and the third is a whole code path rather than a grant:
Three things are quiet, and the third is a whole code path rather than a grant:
-
the cleanup path.
RoleGroupCleaner's errors are logged and swallowed — only a 429 aborts the pass — so a 403 on any teardown delete sets no condition, emits no event, and the reconcile still reportsReconcileComplete=True. In this baseline that is thepersistentvolumeclaimsgrant anddeleteon the workload kinds, which is the same "silent non-reclamation" §3.3.2's PVC trap warns about, reached from the other direction. -
podsandevents, below. Neither reaches even the first-read failure above: one is read through a path that swallows the error, the other is written fire-and-forget. -
events— client-go classifies a 403 on an event as permanent: it logsServer rejected event (will not retry!)and discards the event. Emission is fire-and-forget onto a broadcaster channel, so the reconcile never sees an error, the status is untouched, and the pass reports success. What you lose is everyWarninginarchitecture.md§4.14's vocabulary — includingImmutableFieldIgnored, which is the only one with no paired log line and no status condition, i.e. the only framework warning whose information exists nowhere else. Its own comment records why it was added: doing it silently is what let a storage resize be accepted, reported asReconcileComplete=True, and never applied. -
pods— the health pass Lists the cluster's pods once per reconcile, through the manager's cache, so a 403 does not arrive as a clean error: the lazily-created pods informer never syncs. The visible symptom is client-go loggingFailed to watch *v1.Pod … is forbiddenon a backoff loop while the cluster's conditions stop moving. Where the read does return an error — an uncached reader — the framework deliberately does not treat it as the cluster's fault: it is logged and "the verdict falls back to 'no failures observed'". Either way the grant's absence does not report a fault, it removes the ability to report one, andDegraded=Falsecomes to mean "not checked" rather than "healthy". (Failing pods are one of three inputs toDegraded; an unreadable StatefulSet and a failingServiceHealthCheckdo not depend on this grant.)
The project's runtime gates cannot catch any of this: make test runs envtest as cluster-admin, so
a missing grant is invisible, and make verify-generate only diffs a product's own markers against
its own generated YAML — a framework requirement appears in neither.
What does catch it is a static check over the generated ClusterRole, and this repository ships one
worth copying: examples/trino-operator/internal/controller/rbac_test.go reads
config/rbac/role.yaml off disk, flattens it into group/resource:verb triples and asserts
equality with the set published here. No cluster, no envtest. Equality rather than coverage is
the load-bearing part — it fails when a grant the framework needs goes missing and when the
published minimum quietly grows, and a wildcard rule is recorded verbatim so it fails the comparison
instead of satisfying it.
The SDK generates PodSpecs that adhere to modern container security best practices. The base
role-group handler applies a single, canonical default pod/container SecurityContext with no
opt-in required. This default hardcodes the kubedoop org-standard identity, because all kubedoop
product images run as uid 1001.
The pod-level context lands on .spec.template.spec.securityContext (so it covers every container
in the Pod); the container-level context is set on the primary container the base handler
builds. Sidecar containers carry whatever SidecarConfig.SecurityContext their provider was given.
Pod-level (spec.securityContext):
| Field | Value | Rationale |
|---|---|---|
runAsUser |
1001 |
kubedoop images run as uid 1001 |
runAsGroup |
0 |
OpenShift-compatible: OpenShift assigns an arbitrary uid but keeps gid 0, and kubedoop images are group-0 readable/writable |
fsGroup |
1001 |
mounted volumes are chowned so the non-root process can write to them |
fsGroupChangePolicy |
OnRootMismatch |
apply that chown once, not on every pod start — see below |
runAsNonRoot |
true |
refuse to start as root |
seccompProfile.type |
RuntimeDefault |
apply the runtime's default seccomp profile |
fsGroupChangePolicy is paired with fsGroup deliberately. Kubernetes documents the field as
"Valid values are OnRootMismatch and Always. If not specified, Always is used" — and Always
means the kubelet walks the entire volume, chown'ing and chmod'ing every file, before the
container starts. On a data PVC holding millions of files (an HDFS DataNode, a Kafka broker) that
is tens of minutes to hours, repeated on every restart, every rollout and every node drain, while
the pod sits in ContainerCreating with nothing in its events explaining why.
OnRootMismatch performs the same recursion only when the volume's root directory does not already
have the expected owner and mode — true exactly once, the first time a freshly provisioned volume is
mounted. The trade-off is deliberate: ownership that drifts inside a volume whose root is still
correct will not be repaired. That is a repair the framework never promised, and paying for it on
every start of every stateful pod is the wrong price. A product that wants it back sets
fsGroupChangePolicy: Always through PodOverrides (§3.4.2).
The policy has no effect on ephemeral volume types (secret, configMap, emptyDir), so the config mount and the shared log volume behave identically either way; the data PVC is what it is about.
Container-level (container.securityContext):
| Field | Value | Rationale |
|---|---|---|
runAsUser |
1001 |
kubedoop images run as uid 1001 |
runAsGroup |
0 |
OpenShift-compatible group 0 |
runAsNonRoot |
true |
refuse to start as root |
allowPrivilegeEscalation |
false |
block privilege escalation |
capabilities.drop |
[ALL] |
drop all Linux capabilities |
seccompProfile.type |
RuntimeDefault |
apply the runtime's default seccomp profile |
Products customize the security context through MergedConfig.PodOverrides, which is applied
as a Kubernetes Strategic Merge Patch (the merge strategy docs/architecture.md §2.5
documents for the whole pod template). Security contexts deep-merge per field: an override
stating only runAsUser keeps the framework-hardened remainder (runAsNonRoot,
capabilities.drop, seccompProfile, …).
A field the override does not mention keeps its default, so an override must explicitly
restate any default it wants to change — e.g. an image that must run as root sets both
runAsUser: 0 and runAsNonRoot: false; setting only runAsUser: 0 inherits
runAsNonRoot: true and the kubelet refuses to start the container.
Two handler-wide escape hatches sit alongside the per-role-group overrides:
BaseRoleGroupHandler.WithSecurityContext(containerCtx, podCtx) replaces the defaults for every
role group, and WithoutDefaultSecurityContext() disables them entirely (the StatefulSet is then
built with no SecurityContext unless PodOverrides supplies one).
-
Access Isolation: Product Operators operate with minimal RBAC privileges, reducing the blast radius if an operator is compromised.
-
Consistency: Standardizes security configurations across all data products (HDFS, Hive, Trino, etc.).
-
Lifecycle Management: certificate issuance and renewal are the
secret-operator's job; the SDK's part is the rotation annotations on the registration (WithCertLifetime,WithCertJitter,WithCertBuffer). Propagating a renewal into a running Pod is handled by the separate restarter component, whose contractpkg/constant/restarter.godefines (restarter.kubedoop.dev/enable=trueon the workload;secret.restarter.kubedoop.dev/*andconfigmap.restarter.kubedoop.dev/*annotations;restarter.kubedoop.dev/expires-at.*for expiry-driven Pod restarts).Not applied by the SDK.
pkg/constantonly declares these names — no builder or reconciler setsrestarter.kubedoop.dev/enableon the StatefulSet. Enabling it is a deployment decision: label the cluster CR, and the reconciler propagates the CR's labels into every resource it builds, including theStatefulSet.metadata.labelsthe restarter watches. (Its watch predicate and itsMatchingLabelslist both read object metadata, so a pod-template-only label — everythingpodOverridescan reach — enables nothing.) Without the label and a deployed restarter, a rotated certificate reaches the container only if the application hot-reloads it, or on the next manual rolling restart.