From 3e1eeb566452e8a22a674720c21234d2db63c29b Mon Sep 17 00:00:00 2001 From: Simon Lauger Date: Fri, 4 Sep 2026 00:19:26 +0200 Subject: [PATCH 1/7] docs: correct manifest examples that the API server rejects Three classes of error that make a copied example fail rather than mislead slowly. spec.code is a list of sources, not a single one. Ten examples showed it as an object; the API server answers 'unknown field spec.code.image' and refuses the manifest. Verified against the CRD with kubectl apply --dry-run=server: the list form is accepted, the object form is not. The image tags 8.12.1 and 8.13.0 appeared at ten places and cannot exist. images/openvox-versions.yaml states that the image name encodes the OpenVox major while the tag carries the operator release version; the published tags are 0.12.0, develop and latest. Three label selectors in the troubleshooting and CA import guides do not exist in labels.go: app.kubernetes.io/instance is never set, app.kubernetes.io/name is openvox rather than openvox-server, and the CA label is certificateauthority without a hyphen. A selector that matches nothing is worse than no command at all, because it reads like an answer. Also corrects two resource names: the Config ConfigMap is {name}-config, and the CA setup Job is {ca}-ca-setup. The remaining names in the reference tables were checked against the code and are correct. --- docs/concepts/certificate-signing.md | 2 +- docs/concepts/code-deployment.md | 24 +++++++++---------- docs/concepts/database.md | 2 +- docs/concepts/external-node-classification.md | 2 +- docs/concepts/report-processing.md | 2 +- docs/examples/index.md | 10 ++++---- docs/getting-started/quickstart.md | 2 +- docs/guides/ca-import.md | 2 +- docs/reference/certificateauthority.md | 2 +- docs/reference/config.md | 4 ++-- docs/troubleshooting.md | 4 ++-- 11 files changed, 28 insertions(+), 28 deletions(-) diff --git a/docs/concepts/certificate-signing.md b/docs/concepts/certificate-signing.md index d2b21941..47d87df6 100644 --- a/docs/concepts/certificate-signing.md +++ b/docs/concepts/certificate-signing.md @@ -29,7 +29,7 @@ sequenceDiagram Operator->>K8s: Create PVC ({ca}-data) Operator->>K8s: Create Service ({ca}-internal) Operator->>K8s: Create ServiceAccount + RBAC - Operator->>K8s: Create Job ({ca}-setup) + Operator->>K8s: Create Job ({ca}-ca-setup) Job->>PVC: Run puppetserver ca setup Job->>K8s: Create Secret {ca}-ca (public cert) Job->>K8s: Create Secret {ca}-ca-key (private key) diff --git a/docs/concepts/code-deployment.md b/docs/concepts/code-deployment.md index 0ea0d947..3f5c5925 100644 --- a/docs/concepts/code-deployment.md +++ b/docs/concepts/code-deployment.md @@ -58,9 +58,9 @@ metadata: spec: image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" code: - image: ghcr.io/example/puppet-code:v1.0.0 + - image: ghcr.io/example/puppet-code:v1.0.0 ``` ### Pull Policy @@ -70,8 +70,8 @@ Control when the image is pulled via `imagePullPolicy`. Defaults to `IfNotPresen ```yaml spec: code: - image: ghcr.io/example/puppet-code:v1.0.0 - imagePullPolicy: Always + - image: ghcr.io/example/puppet-code:v1.0.0 + imagePullPolicy: Always ``` Supported values: `Always`, `IfNotPresent`, `Never`. @@ -83,7 +83,7 @@ For immutable, reproducible deployments you can reference images by digest inste ```yaml spec: code: - image: ghcr.io/example/puppet-code@sha256:45b23dee08af5e43a7fea6c4cf9c25ccf269ee113168c19722f87876677c5cb2 + - image: ghcr.io/example/puppet-code@sha256:45b23dee08af5e43a7fea6c4cf9c25ccf269ee113168c19722f87876677c5cb2 ``` A tag+digest combination also works: @@ -91,7 +91,7 @@ A tag+digest combination also works: ```yaml spec: code: - image: ghcr.io/example/puppet-code:v1.0.0@sha256:45b23dee08af5e43a7fea6c4cf9c25ccf269ee113168c19722f87876677c5cb2 + - image: ghcr.io/example/puppet-code:v1.0.0@sha256:45b23dee08af5e43a7fea6c4cf9c25ccf269ee113168c19722f87876677c5cb2 ``` ### Rolling Out Code Changes @@ -101,7 +101,7 @@ Update the image reference to deploy new code. The operator detects the change a ```yaml spec: code: - image: ghcr.io/example/puppet-code:v1.1.0 + - image: ghcr.io/example/puppet-code:v1.1.0 ``` ### Private Registries @@ -111,8 +111,8 @@ For private registries, create a pull secret and reference it: ```yaml spec: code: - image: registry.example.com/puppet-code:v1.0.0 - imagePullSecret: registry-credentials + - image: registry.example.com/puppet-code:v1.0.0 + imagePullSecret: registry-credentials ``` ### Rollout Visibility @@ -163,7 +163,7 @@ spec: configRef: production certificateRef: canary-cert code: - image: ghcr.io/example/puppet-code:v2.0.0-rc1 + - image: ghcr.io/example/puppet-code:v2.0.0-rc1 ``` ## PVC @@ -181,9 +181,9 @@ metadata: spec: image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" code: - claimName: puppet-code + - claimName: puppet-code ``` Like the image volume, the PVC is mounted at the configured `environmentPath` (default `/etc/puppetlabs/code/environments`), so its root must contain the environment directories directly (`production/`, `staging/`, ...). diff --git a/docs/concepts/database.md b/docs/concepts/database.md index 6f9604eb..971c5792 100644 --- a/docs/concepts/database.md +++ b/docs/concepts/database.md @@ -71,7 +71,7 @@ spec: databaseRef: production-db # operator reads Database.status.url image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" ``` The operator reads `Database.status.url` (e.g. `https://production-db.namespace.svc.cluster.local:8081`) and renders it into `puppetdb.conf`. When the Database is not yet `Running`, the Config controller waits. diff --git a/docs/concepts/external-node-classification.md b/docs/concepts/external-node-classification.md index 6f5d19fd..c12b0914 100644 --- a/docs/concepts/external-node-classification.md +++ b/docs/concepts/external-node-classification.md @@ -136,7 +136,7 @@ spec: nodeClassifierRef: pe-classifier image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" ``` This generates the following puppet.conf entries in the `[server]` section: diff --git a/docs/concepts/report-processing.md b/docs/concepts/report-processing.md index 8fa5858f..a77190bd 100644 --- a/docs/concepts/report-processing.md +++ b/docs/concepts/report-processing.md @@ -107,7 +107,7 @@ spec: authorityRef: production-ca image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" ``` ```yaml diff --git a/docs/examples/index.md b/docs/examples/index.md index f23a52a8..90419ac5 100644 --- a/docs/examples/index.md +++ b/docs/examples/index.md @@ -13,7 +13,7 @@ spec: authorityRef: lab-ca image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" --- apiVersion: openvox.voxpupuli.org/v1alpha1 kind: CertificateAuthority @@ -74,7 +74,7 @@ spec: databaseRef: production-db image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" puppet: environmentTimeout: unlimited storeconfigs: true @@ -232,7 +232,7 @@ spec: replicas: 3 maxActiveInstances: 2 code: - claimName: puppet-code + - claimName: puppet-code resources: requests: cpu: "1" @@ -250,10 +250,10 @@ spec: certificateRef: canary-cert poolRefs: [puppet] image: - tag: "8.13.0" + tag: "latest" replicas: 1 code: - claimName: puppet-code + - claimName: puppet-code resources: requests: cpu: "1" diff --git a/docs/getting-started/quickstart.md b/docs/getting-started/quickstart.md index 0b37421d..fb03c9d8 100644 --- a/docs/getting-started/quickstart.md +++ b/docs/getting-started/quickstart.md @@ -75,7 +75,7 @@ This guide sets up an OpenVox Server deployment. Choose between the Helm chart ( authorityRef: lab-ca image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" --- apiVersion: openvox.voxpupuli.org/v1alpha1 kind: CertificateAuthority diff --git a/docs/guides/ca-import.md b/docs/guides/ca-import.md index f92e972f..a5d4eeb8 100644 --- a/docs/guides/ca-import.md +++ b/docs/guides/ca-import.md @@ -28,7 +28,7 @@ If you have an existing CA and want the operator to manage it going forward, you ```bash # Find the PVC - kubectl get pvc -l openvox.voxpupuli.org/certificate-authority=production-ca + kubectl get pvc -l openvox.voxpupuli.org/certificateauthority=production-ca # Create a temporary pod to copy data kubectl run ca-import --image=busybox --restart=Never \ diff --git a/docs/reference/certificateauthority.md b/docs/reference/certificateauthority.md index 51a69a7d..6818a351 100644 --- a/docs/reference/certificateauthority.md +++ b/docs/reference/certificateauthority.md @@ -182,7 +182,7 @@ The failed Job is left in place; its logs are the only record of what went wrong: ```bash -kubectl logs -n job/-setup +kubectl logs -n job/-ca-setup ``` The attempt counter lives in the `openvox.voxpupuli.org/setup-attempts` diff --git a/docs/reference/config.md b/docs/reference/config.md index 790e6083..191c55a2 100644 --- a/docs/reference/config.md +++ b/docs/reference/config.md @@ -13,7 +13,7 @@ spec: authorityRef: production-ca image: repository: ghcr.io/slauger/openvox-server-8 - tag: "8.12.1" + tag: "latest" puppet: environmentTimeout: "0" storeconfigs: true @@ -223,6 +223,6 @@ credentials along. Secrets for code images are always added on top. | Resource | Name | Description | |---|---|---| -| ConfigMap | `{name}` | puppet.conf, puppetserver.conf, auth.conf, webserver.conf, `routes.yaml` (facts terminus, when PuppetDB is the active backend), etc. | +| ConfigMap | `{name}-config` | puppet.conf, puppetserver.conf, auth.conf, webserver.conf, `routes.yaml` (facts terminus, when PuppetDB is the active backend), etc. | | Secret | `{name}-enc` | ENC config for openvox-enc binary (only when `nodeClassifierRef` is set) | | ServiceAccount | `{name}-server` | Shared ServiceAccount for all Server pods (`automountServiceAccountToken: false`) | diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index 062bd9f6..a8fc52ad 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -68,7 +68,7 @@ condition names it. ```bash kubectl describe deployment -n -kubectl describe pod -l app.kubernetes.io/instance= -n +kubectl describe pod -l openvox.voxpupuli.org/server= -n kubectl get events -n --sort-by='.lastTimestamp' ``` @@ -117,7 +117,7 @@ kubectl logs -n --previous 1. Verify the Pool Service exists: ```bash - kubectl get svc -n -l app.kubernetes.io/name=openvox-server + kubectl get svc -n -l app.kubernetes.io/name=openvox ``` 2. Check endpoints are populated: From 7397a4238cc89123a532c6931e47317b57c0e44b Mon Sep 17 00:00:00 2001 From: Simon Lauger Date: Fri, 4 Sep 2026 00:21:10 +0200 Subject: [PATCH 2/7] docs: correct behaviour the controller does not have The Error phase is documented for Config, Server and Database as 'reconciliation failed'. It exists in the CRD enum but no code path assigns it: a failing reconcile leaves the phase where it was and reports the reason in a condition. Someone waiting for phase Error waits forever, which is the worst kind of wrong - it looks like a working diagnosis. Replaced with the condition to watch, and the jsonpath was checked against a live object. The Gateway API concept showed serverRef on a Pool. That field exists neither on the CRD nor in the chart; a Server names its pools through poolRefs and the Pool selects those pods. Removed, and the direction of the relationship spelled out, since getting it backwards is the natural guess. Added what the code does but nothing described: renewal reuses the existing private key deliberately, because the CA renews for the same public key. So renewal extends validity without rotating the key, and replacing a key needs a new Certificate. Checked and left alone: the Renewing phase description is already accurate, and no page claims renewal generates a new key. The review report asserted both; neither held up. --- docs/concepts/gateway-api.md | 5 ++++- docs/reference/certificate.md | 11 +++++++++++ docs/reference/config.md | 12 +++++++++++- docs/reference/database.md | 12 +++++++++++- docs/reference/server.md | 12 +++++++++++- 5 files changed, 48 insertions(+), 4 deletions(-) diff --git a/docs/concepts/gateway-api.md b/docs/concepts/gateway-api.md index 3bbacb6b..515ec044 100644 --- a/docs/concepts/gateway-api.md +++ b/docs/concepts/gateway-api.md @@ -107,6 +107,10 @@ spec: The `openvox-stack` chart provides a `gateway` section for shared Gateway settings: +A Pool does not name its servers. The relationship runs the other way: each +entry under `servers` lists the pools it joins via `poolRefs`, and the Pool +selects those pods through its Service. + ```yaml gateway: name: puppet-gateway @@ -114,7 +118,6 @@ gateway: pools: - name: puppet - serverRef: ca service: type: ClusterIP port: 8140 diff --git a/docs/reference/certificate.md b/docs/reference/certificate.md index 881383e2..9ce87a1e 100644 --- a/docs/reference/certificate.md +++ b/docs/reference/certificate.md @@ -89,6 +89,17 @@ Certificates issued before this field existed carry an empty hash. The controller adopts the current spec as the baseline for them rather than re-signing every certificate after an operator upgrade. +### Renewal reuses the private key + +A renewal submits a CSR for the key the certificate already has +(`certificate_signing.go`: the existing `key.pem` is read from the TLS Secret +and reused). The CA renews for the same public key, so the key material must +not change between the old and the new certificate. + +The consequence is worth stating: renewal extends validity, it does not rotate +the key. A key that must be replaced needs a new Certificate under a different +name, since `certname` is immutable and the CA keeps one entry per name. + ### CSR Poll Backoff When the CA does not immediately sign the CSR (e.g. autosigning is disabled), the controller enters `WaitingForSigning` after 10 unsuccessful poll attempts and retries with exponential backoff: diff --git a/docs/reference/config.md b/docs/reference/config.md index 191c55a2..9facd5b0 100644 --- a/docs/reference/config.md +++ b/docs/reference/config.md @@ -208,7 +208,17 @@ Controls Puppet Server metrics.conf settings. |---|---| | `Pending` | Config created, waiting for reconciliation | | `Running` | ConfigMap created, ready for use | -| `Error` | Reconciliation failed | +| `Error` | Defined in the API, but never set by the controller (see below) | + +`Error` is part of the API but the controller never assigns it. A failing +reconcile leaves the phase at its previous value, reports the reason in the +`ConfigReady` condition and emits a warning event. Watch the condition rather than +the phase: + +```bash +kubectl get config -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}: {.message}{"\n"}{end}' +``` + ### Image resolution diff --git a/docs/reference/database.md b/docs/reference/database.md index 8d8cf676..122bf9a8 100644 --- a/docs/reference/database.md +++ b/docs/reference/database.md @@ -99,7 +99,17 @@ When enabled, the default policy allows TCP/8081 only from pods with `app.kubern | `Pending` | Database created, resolving references | | `WaitingForCert` | Certificate not yet `Signed` | | `Running` | Deployment created and running | -| `Error` | Reconciliation failed | +| `Error` | Defined in the API, but never set by the controller (see below) | + +`Error` is part of the API but the controller never assigns it. A failing +reconcile leaves the phase at its previous value, reports the reason in the +`DatabaseReady` condition and emits a warning event. Watch the condition rather than +the phase: + +```bash +kubectl get database -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}: {.message}{"\n"}{end}' +``` + ## Pod Anatomy diff --git a/docs/reference/server.md b/docs/reference/server.md index 18cbb515..bbc7ad06 100644 --- a/docs/reference/server.md +++ b/docs/reference/server.md @@ -131,7 +131,17 @@ When enabled, the default policy allows TCP/8140 from all sources (agents may co | `Pending` | Server created, resolving references | | `WaitingForCert` | Certificate not yet `Signed` | | `Running` | Deployment created and running | -| `Error` | Reconciliation failed | +| `Error` | Defined in the API, but never set by the controller (see below) | + +`Error` is part of the API but the controller never assigns it. A failing +reconcile leaves the phase at its previous value, reports the reason in the +`Ready` condition and emits a warning event. Watch the condition rather than +the phase: + +```bash +kubectl get server -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}: {.message}{"\n"}{end}' +``` + ## Deployment Strategy From aab8ce77946f8f3a53ecc9063022a9f11314e325 Mon Sep 17 00:00:00 2001 From: Simon Lauger Date: Fri, 4 Sep 2026 00:23:04 +0200 Subject: [PATCH 3/7] docs: explain that no SigningPolicy means deny-all The operator writes autosign = /usr/local/bin/openvox-autosign as soon as a CertificateAuthority exists, regardless of whether any SigningPolicy does, and evaluatePolicies returns false for an empty list. An empty policy list is therefore deny-all rather than off, which is the state every fresh install starts in. Nothing surfaces it: the servers run, the Config reports Running, and every agent sits in --waitforcert until it gives up. Documented in the quickstart and in the agent section of the troubleshooting guide, with the manual puppetserver ca sign path as the immediate way out. Also separates two cases the troubleshooting guide conflated. A Certificate resource stuck in Pending is almost never a policy problem: for an internal CA the operator signs its own Certificates over the CA API with the operator signing certificate, so autosign is not involved. That entry now lists the causes that do apply - CA not ready, signing certificate not yet available, certname already claimed, external CA - each with a command that reads the condition rather than the phase. Policies govern agents, and that is where the note now lives. --- docs/getting-started/quickstart.md | 12 ++++++ docs/troubleshooting.md | 63 +++++++++++++++++++++++++++--- 2 files changed, 69 insertions(+), 6 deletions(-) diff --git a/docs/getting-started/quickstart.md b/docs/getting-started/quickstart.md index fb03c9d8..a22d5da6 100644 --- a/docs/getting-started/quickstart.md +++ b/docs/getting-started/quickstart.md @@ -162,6 +162,18 @@ NAME TYPE ENDPOINTS AGE pool.openvox.voxpupuli.org/puppet ClusterIP 1 2m ``` +!!! warning "Without a SigningPolicy nothing gets signed" + + The operator points `autosign` at its own binary as soon as a + CertificateAuthority exists, and that binary denies every CSR it has no + matching policy for. An empty policy list therefore means deny-all, not + off. + + Nothing surfaces this. The servers come up, the Config reports `Running`, + and every agent sits in `puppet agent --waitforcert` until it gives up. + Check with `kubectl get signingpolicy -n `; if the list is + empty, see [SigningPolicy](../reference/signingpolicy.md). + ## Next Steps See the [Examples](../examples/index.md) section for production setups with separate CA, server pools, canary deployments, and code deployment via OCI image volumes. diff --git a/docs/troubleshooting.md b/docs/troubleshooting.md index a8fc52ad..2cebf701 100644 --- a/docs/troubleshooting.md +++ b/docs/troubleshooting.md @@ -43,16 +43,45 @@ condition names it. **Symptoms:** Certificate never reaches `Signed` phase. +A SigningPolicy is usually *not* the cause here. For an internal CA the +operator signs its own Certificate resources over the CA API, authenticated +with the operator signing certificate, so autosign is not involved. Policies +govern agents, not Certificate resources. + **Possible causes:** -1. **CA not ready:** The CertificateAuthority must be in `Ready` phase. -2. **No matching SigningPolicy:** No policy exists that would sign this certificate. +1. **CA not ready.** The CertificateAuthority must report `CAReady`. + + ```bash + kubectl get certificateauthority -n \ + -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}{"\n"}{end}' + ``` + +2. **Operator signing certificate not available yet.** Without it the operator + cannot sign and falls back to polling, which only succeeds if something else + signs the CSR. + + ```bash + kubectl get certificateauthority -n \ + -o jsonpath='{.status.signingSecretName}{"\n"}' + ``` + + An empty value means the `{ca}-operator-signing` Certificate is not signed + yet. During bootstrap this resolves on its own. + +3. **Certname already claimed.** Two Certificates cannot share a certname + against the same CA. The condition names the holder: ```bash - kubectl get signingpolicy -n + kubectl get certificate -n \ + -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}: {.message}{"\n"}{end}' ``` -**Solution:** Check CA status and verify a SigningPolicy with matching criteria exists. +4. **External CA.** With `spec.external` the operator has no admin access and + cannot sign. The CSR must be signed on the external CA. + +**Agents** stuck in `--waitforcert` are the case where SigningPolicy matters - +see [Agents cannot connect to server](#agents-cannot-connect-to-server). ### Server pods not starting @@ -110,9 +139,31 @@ kubectl logs -n --previous ### Agents cannot connect to server -**Symptoms:** Puppet agents fail to connect to the server endpoint. +**Symptoms:** Puppet agents fail to connect to the server endpoint, or hang in +`puppet agent --waitforcert`. -**Debugging steps:** +An agent that reaches the server but hangs waiting for its certificate has a +signing problem, not a connectivity one. The operator points `autosign` at its +own binary as soon as a CertificateAuthority exists, and that binary denies +every CSR no policy matches - so **no SigningPolicy means deny-all, not off**. +The servers run normally and the Config reports `Running` either way. + +```bash +kubectl get signingpolicy -n +``` + +An empty list is the common cause on a fresh install. See +[SigningPolicy](reference/signingpolicy.md) for the available match rules, or +sign by hand: + +```bash +kubectl exec -n deploy/ -- \ + puppetserver ca list --all +kubectl exec -n deploy/ -- \ + puppetserver ca sign --certname +``` + +**Debugging steps for connectivity:** 1. Verify the Pool Service exists: From ab6c28e5fe8fc98adaf56b10908927c58af180e0 Mon Sep 17 00:00:00 2001 From: Simon Lauger Date: Fri, 4 Sep 2026 00:39:21 +0200 Subject: [PATCH 4/7] docs: close the gaps between the reference pages and the CRDs Compared every documented field against the generated CRDs, in both directions. observedGeneration exists on all nine status types since #549 but appeared on two reference pages. Added to the remaining seven, with the sentence that makes it useful: a value below metadata.generation means the rest of the status has not caught up yet. Certificate gained signedSpecHash and effectiveDNSAltNames in the status table. Both drive behaviour a reader needs to predict - the first decides when a certificate is re-signed, the second says which alt names it is actually issued for once Pools contribute. Server gained readOnlyRootFilesystem, the per-Server override added in #575. The ImageSpec table still listed defaults that #549 and #576 removed: repository ghcr.io/slauger/openvox-server-8, tag latest, pullPolicy IfNotPresent. It contradicted the prose directly beneath it. Corrected, and the reason spelled out - a nested default is applied whether or not the parent object was given, so a defaulted field can never mean inherit - plus the pullSecrets rule that a Server list replaces rather than extends. Checked and found correct: every other documented default matches its CRD, and no reference page names a field that does not exist. --- docs/reference/certificate.md | 3 +++ docs/reference/certificateauthority.md | 1 + docs/reference/config.md | 1 + docs/reference/database.md | 1 + docs/reference/index.md | 27 ++++++++++++++++++-------- docs/reference/nodeclassifier.md | 1 + docs/reference/reportprocessor.md | 1 + docs/reference/server.md | 1 + docs/reference/signingpolicy.md | 1 + 9 files changed, 29 insertions(+), 8 deletions(-) diff --git a/docs/reference/certificate.md b/docs/reference/certificate.md index 9ce87a1e..90e6a10f 100644 --- a/docs/reference/certificate.md +++ b/docs/reference/certificate.md @@ -31,9 +31,12 @@ spec: | Field | Type | Description | |---|---|---| +| `observedGeneration` | int64 | The `.metadata.generation` the status was last derived from. A value below `.metadata.generation` means the rest of this status has not caught up with the current spec yet | | `phase` | string | Current lifecycle phase | | `secretName` | string | Name of the Secret containing `cert.pem` and `key.pem` | | `notAfter` | time | Expiry time of the signed certificate | +| `signedSpecHash` | string | Digest of what the current certificate was issued for: certname, effective alt names and CSR extensions. A mismatch triggers re-signing; empty means the hash was never recorded and is adopted rather than triggering one | +| `effectiveDNSAltNames` | []string | The alt names the certificate is actually issued for: `spec.dnsAltNames` plus the route hostname of every Pool with `injectDNSAltName` that a Server using this Certificate joins | | `conditions` | []Condition | `CertSigned` | ## Deletion diff --git a/docs/reference/certificateauthority.md b/docs/reference/certificateauthority.md index 6818a351..b1a8c831 100644 --- a/docs/reference/certificateauthority.md +++ b/docs/reference/certificateauthority.md @@ -81,6 +81,7 @@ spec: | Field | Type | Description | |---|---|---| +| `observedGeneration` | int64 | The `.metadata.generation` the status was last derived from. A value below `.metadata.generation` means the rest of this status has not caught up with the current spec yet | | `phase` | string | Current lifecycle phase | | `caSecretName` | string | Name of the Secret containing `ca_crt.pem` (public CA certificate) | | `serviceName` | string | Name of the internal ClusterIP Service for operator communication. Empty when `spec.external` is set. | diff --git a/docs/reference/config.md b/docs/reference/config.md index 9facd5b0..afc3e97a 100644 --- a/docs/reference/config.md +++ b/docs/reference/config.md @@ -199,6 +199,7 @@ Controls Puppet Server metrics.conf settings. | Field | Type | Description | |---|---|---| +| `observedGeneration` | int64 | The `.metadata.generation` the status was last derived from. A value below `.metadata.generation` means the rest of this status has not caught up with the current spec yet | | `phase` | string | Current lifecycle phase | | `conditions` | []Condition | `ConfigReady` | diff --git a/docs/reference/database.md b/docs/reference/database.md index 122bf9a8..d7f7e83e 100644 --- a/docs/reference/database.md +++ b/docs/reference/database.md @@ -86,6 +86,7 @@ When enabled, the default policy allows TCP/8081 only from pods with `app.kubern | Field | Type | Description | |---|---|---| +| `observedGeneration` | int64 | The `.metadata.generation` the status was last derived from. A value below `.metadata.generation` means the rest of this status has not caught up with the current spec yet | | `phase` | string | Current lifecycle phase | | `url` | string | HTTPS endpoint of the Database Service (e.g. `https://production-db:8081`) | | `ready` | int32 | Number of ready replicas | diff --git a/docs/reference/index.md b/docs/reference/index.md index 517a0015..c4654701 100644 --- a/docs/reference/index.md +++ b/docs/reference/index.md @@ -52,17 +52,28 @@ These types are reused across multiple CRDs. | Field | Type | Default | Description | |---|---|---|---| -| `repository` | string | `ghcr.io/slauger/openvox-server-8` | Container image repository | -| `tag` | string | `latest` | Container image tag | -| `pullPolicy` | string | `IfNotPresent` | Image pull policy | +| `repository` | string | - | Container image repository | +| `tag` | string | - | Container image tag | +| `pullPolicy` | string | - | Image pull policy | | `pullSecrets` | []LocalObjectReference | - | Image pull secrets | -`repository` and `tag` carry no API-level default. They are required on `Config` -and `Database`; on `Server` both are optional and fall back to the referenced -`Config`, which is what lets one Config drive a whole set of Servers. +No field here carries an API-level default. That is deliberate: a nested +default is applied whether or not the parent object was given, so a defaulted +field can never express "inherit from the Config" - it is simply never empty. -The defaults live in the Helm charts (`config.image.*`, `database.image.*`), -where changing the registry is a values change rather than a CRD update. +`repository` and `tag` are required on `Config` and `Database`. On `Server` +both are optional and fall back to the referenced `Config`, which is what lets +one Config drive a whole set of Servers. + +`pullPolicy` follows the same rule, falling back to the Config and then to +`IfNotPresent`. `pullSecrets` behaves differently: a non-empty list on a +`Server` *replaces* the Config's rather than extending it, so a Server pulling +from another registry does not carry the Config's credentials along. Secrets +for code images are always added on top. + +The registry defaults live in the Helm charts (`config.image.*`, +`database.image.*`), where changing them is a values change rather than a CRD +update. ### StorageSpec diff --git a/docs/reference/nodeclassifier.md b/docs/reference/nodeclassifier.md index ff1c19d7..cd5c612c 100644 --- a/docs/reference/nodeclassifier.md +++ b/docs/reference/nodeclassifier.md @@ -166,6 +166,7 @@ At most one authentication method may be configured. | Field | Type | Description | |---|---|---| +| `observedGeneration` | int64 | The `.metadata.generation` the status was last derived from. A value below `.metadata.generation` means the rest of this status has not caught up with the current spec yet | | `phase` | string | Current lifecycle phase | | `conditions` | []Condition | `Ready` | diff --git a/docs/reference/reportprocessor.md b/docs/reference/reportprocessor.md index 628fb85f..e757d7d3 100644 --- a/docs/reference/reportprocessor.md +++ b/docs/reference/reportprocessor.md @@ -175,6 +175,7 @@ Either `value` or `valueFrom` may be set, not both. | Field | Type | Description | |---|---|---| +| `observedGeneration` | int64 | The `.metadata.generation` the status was last derived from. A value below `.metadata.generation` means the rest of this status has not caught up with the current spec yet | | `phase` | string | Current lifecycle phase | | `conditions` | []Condition | `Ready` | diff --git a/docs/reference/server.md b/docs/reference/server.md index bbc7ad06..fa1cd6b6 100644 --- a/docs/reference/server.md +++ b/docs/reference/server.md @@ -45,6 +45,7 @@ spec: | `envFrom` | []EnvFromSource | - | ConfigMap/Secret sources to populate environment variables from | | `extraVolumes` | []Volume | - | Extra volumes added to the Server pods | | `extraVolumeMounts` | []VolumeMount | - | Extra volume mounts for the `openvox-server` container | +| `readOnlyRootFilesystem` | *bool | *(inherits from Config)* | Overrides the Config's setting for this Server. One Config backs several Servers with different roles, so the CA and the compilers can differ. Unset inherits | | `securityContext` | [PodSecurityContextSpec](index.md#podsecuritycontextspec) | - | Override pod-level security context (runAsUser/runAsGroup/fsGroup) | ### Extra Environment and Volumes diff --git a/docs/reference/signingpolicy.md b/docs/reference/signingpolicy.md index 4b4c0690..4277ce64 100644 --- a/docs/reference/signingpolicy.md +++ b/docs/reference/signingpolicy.md @@ -127,6 +127,7 @@ Either `value` or `valueFrom` must be set. | Field | Type | Description | |---|---|---| +| `observedGeneration` | int64 | The `.metadata.generation` the status was last derived from. A value below `.metadata.generation` means the rest of this status has not caught up with the current spec yet | | `phase` | string | Current lifecycle phase | | `conditions` | []Condition | `Ready` | From a61c13595ca42089208884c2b2fc6740034c078c Mon Sep 17 00:00:00 2001 From: Simon Lauger Date: Fri, 4 Sep 2026 00:41:31 +0200 Subject: [PATCH 5/7] docs: add a guide for connecting agents The user documentation never showed an agent connecting. puppet agent, --waitforcert, ca_server and puppetserver ca sign appeared nowhere, so the moment the product exists for was the one step a reader had to work out alone. The guide covers what an agent needs, where to point it, running one inside the cluster and reaching one from outside, signing by hand when no policy matches, revoking, and how to tell from the server side that it worked. Two things it states that are easy to get wrong. ca_server is only needed once the CA and the compilers are in separate pools - with the default poolRefs [ca, server] one address serves both, which is why the setting appears nowhere in this repository. And puppet agent --test returns 2 when it applied changes, so a Job that treats non-zero as failure reports a successful run as broken. It also names what the agent image is: openvox-agent- is built by CI and not by the release workflow, so it carries only the develop tag - no latest, no version. Documented as the test artifact it is, rather than implying it is a supported way to run agents. Every claim was checked: the Service default is ClusterIP, the Deployment carries the Server name, crlRefreshInterval defaults to 5m, and the image tags are what ghcr actually holds. --- docs/guides/connecting-agents.md | 163 +++++++++++++++++++++++++++++++ mkdocs.yml | 1 + 2 files changed, 164 insertions(+) create mode 100644 docs/guides/connecting-agents.md diff --git a/docs/guides/connecting-agents.md b/docs/guides/connecting-agents.md new file mode 100644 index 00000000..6fbd0b18 --- /dev/null +++ b/docs/guides/connecting-agents.md @@ -0,0 +1,163 @@ +# Connecting Agents + +The operator manages the server side. Agents are ordinary Puppet agents: they +are not Kubernetes resources, and nothing in the operator creates or tracks +them. This page covers what they need from a stack deployed by the operator. + +## What an agent needs + +Three things, and they are the usual ones: + +| | | +|---|---| +| **A reachable server address** | the Service of a Pool the server joins, on port 8140 | +| **A certname** | the identity the certificate is issued for | +| **A signed certificate** | either autosigned by a [SigningPolicy](../reference/signingpolicy.md) or signed by hand | + +!!! warning "Without a SigningPolicy nothing is signed" + + The operator points `autosign` at its own binary as soon as a + CertificateAuthority exists, and that binary denies every CSR no policy + matches. An empty policy list is deny-all, not off. Agents then sit in + `--waitforcert` while the servers look perfectly healthy. + +## Where to point the agent + +Agents talk to the Service of a Pool. With the default layout the CA server +joins both pools (`poolRefs: [ca, server]`), so one address serves catalog +requests and the CA: + +```bash +puppet agent --test \ + --server -server \ + --certname web01.example.com \ + --waitforcert 30 +``` + +If you separate the roles - the CA in the `ca` pool only, compilers in +`server` - the agent needs both addresses, because catalog and CA no longer +share one: + +```ini +[main] +server = -server +ca_server = -ca +``` + +## From inside the cluster + +Any Puppet agent works. This project also builds `openvox-agent-`, but +be aware of what it is: a **test artifact**. It is built by CI and not by the +release workflow, so it carries only the `develop` tag - no `latest`, no +version. For anything but a throwaway check, use your own agent image and pin +it. + +The shape below is what the e2e suite runs: + +```yaml +apiVersion: batch/v1 +kind: Job +metadata: + name: puppet-agent +spec: + backoffLimit: 0 + template: + spec: + restartPolicy: Never + containers: + - name: puppet-agent + image: ghcr.io/slauger/openvox-agent-8:develop # test artifact, see above + command: ["sh", "-c"] + args: + - | + puppet agent --test \ + --server -server \ + --certname agent-01 \ + --waitforcert 30 + # 0 = no changes, 2 = changes applied; both are success + EXIT=$? + if [ $EXIT -eq 0 ] || [ $EXIT -eq 2 ]; then exit 0; else exit $EXIT; fi +``` + +The exit code check matters: `puppet agent --test` returns **2** when it +applied changes, which is a success and not a failure. + +## From outside the cluster + +A Pool Service defaults to `ClusterIP`, which is unreachable from outside. +Choose one: + +| Option | How | +|---|---| +| **LoadBalancer** | `pools[].service.type: LoadBalancer` | +| **NodePort** | `pools[].service.type: NodePort`, optionally `nodePort` | +| **Gateway API** | `pools[].route`, see [Gateway API](../concepts/gateway-api.md) | + +Whichever you pick, the name agents connect through must be in the server +certificate. The chart derives the Service names of every Pool a server joins +into `dnsAltNames` automatically; an external name such as a load balancer +address has to be added explicitly: + +```yaml +servers: + - name: ca + certificate: + certname: puppet + dnsAltNames: + - puppet.example.com +``` + +A missing name shows up as a TLS error on the agent, not as an operator +problem: + +``` +Server hostname 'puppet.example.com' did not match server certificate +``` + +## Signing by hand + +Without a matching policy the CSR waits. List and sign on the CA pod: + +```bash +kubectl exec -n deploy/ -- \ + puppetserver ca list --all + +kubectl exec -n deploy/ -- \ + puppetserver ca sign --certname web01.example.com +``` + +## Removing an agent + +Revoking is a CA operation, and the operator does not do it for you - it only +manages the certificates that belong to its own resources: + +```bash +kubectl exec -n deploy/ -- \ + puppetserver ca clean --certname web01.example.com +``` + +The revocation reaches agents with the next CRL refresh, which runs on +`spec.crlRefreshInterval` (default `5m`). Until then the revoked certificate is +still accepted. + +## Verifying it worked + +The agent reports success itself, but two things are worth checking on the +server side. + +**Did the catalog compile?** + +```bash +kubectl logs -n deploy/ | grep -i "Compiled catalog" +``` + +**Did the facts reach PuppetDB?** Only when a Database is wired up: + +```bash +kubectl exec -n deploy/ -- \ + curl -sf "http://127.0.0.1:8080/pdb/query/v4/nodes" +``` + +An empty result with a healthy agent run usually means `routes.yaml` is not +routing the facts terminus to PuppetDB - see +[Database](../reference/database.md). diff --git a/mkdocs.yml b/mkdocs.yml index d4e6d0c7..156cba3e 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -66,6 +66,7 @@ nav: - Installation: getting-started/installation.md - Quick Start: getting-started/quickstart.md - Guides: + - Connecting Agents: guides/connecting-agents.md - CA Import & External CA: guides/ca-import.md - Monitoring: guides/monitoring.md - Pausing Reconciliation: guides/pausing-reconciliation.md From 93d4757892342e44cde60935609733b970950a19 Mon Sep 17 00:00:00 2001 From: Simon Lauger Date: Fri, 4 Sep 2026 00:44:31 +0200 Subject: [PATCH 6/7] docs: fix the remaining broken example and two claims about behaviour Extracted all 42 manifests from the documentation and ran them against a real API server with kubectl apply --dry-run=server. One was rejected for a reason a reader could not guess: the NodeClassifier page abbreviated the image block as 'image: ...', and the resulting error talks about a type mismatch rather than the ellipsis. Replaced with the real block; all 42 are accepted now. The Database concept claimed the Config controller waits when the Database is not Running. It does not: renderPuppetDBConf returns a puppetdb.conf without server_urls and the servers start regardless. Together with soft_write_failure = true, which the operator sets in every path, a server in that state compiles catalogs normally while reports and exported resources go nowhere and nothing is raised. That resolves itself during bring-up and only bites when the Database never becomes ready, so the symptom is an empty PuppetDB rather than an error. Monitoring now states that CertificateAuthority and Certificate write the same expiry metric with the same labels and nothing distinguishing them, so the .*-ca matcher in the CA alert is a naming convention and not a guarantee. Added an absence rule for the CRL metric: a staleness alert cannot fire on a series that was never created, which is exactly the case when the operator never got far enough to refresh one. --- docs/concepts/database.md | 13 ++++++++++++- docs/guides/monitoring.md | 30 ++++++++++++++++++++++++++++++ docs/reference/nodeclassifier.md | 4 +++- 3 files changed, 45 insertions(+), 2 deletions(-) diff --git a/docs/concepts/database.md b/docs/concepts/database.md index 971c5792..51e4f9a8 100644 --- a/docs/concepts/database.md +++ b/docs/concepts/database.md @@ -74,7 +74,18 @@ spec: tag: "latest" ``` -The operator reads `Database.status.url` (e.g. `https://production-db.namespace.svc.cluster.local:8081`) and renders it into `puppetdb.conf`. When the Database is not yet `Running`, the Config controller waits. +The operator reads `Database.status.url` (e.g. `https://production-db.namespace.svc.cluster.local:8081`) and renders it into `puppetdb.conf`. + +When the Database has no URL yet, the Config controller does **not** wait. It +renders a `puppetdb.conf` without `server_urls` and carries on, so the servers +start regardless. Combined with `soft_write_failure = true`, which the operator +always sets, a server in that state compiles catalogs normally while reports +and exported resources go nowhere and no error is raised. + +The Config is re-reconciled when the Database status changes, so this resolves +by itself during bring-up. It becomes a problem only if the Database never +reaches `Running`: the symptom is an empty PuppetDB with healthy-looking +servers, not a failure. ### Via static `puppetdb.serverUrls` diff --git a/docs/guides/monitoring.md b/docs/guides/monitoring.md index 376a080a..8ef40ea3 100644 --- a/docs/guides/monitoring.md +++ b/docs/guides/monitoring.md @@ -114,6 +114,24 @@ Alert when a certificate expires within 30 days: description: "{{ $labels.name }} in {{ $labels.namespace }} expires in {{ $value | humanizeDuration }}" ``` +### Missing CRL series + +A stale CRL is alertable only while the series exists. If the operator never +refreshed the CRL - it crashed early, or the CA never became ready - there is +no series at all and the staleness rule below stays silent. Alert on the +absence separately: + +```yaml + - alert: OpenVoxCRLMetricMissing + expr: absent(openvox_crl_last_refresh_timestamp_seconds) + for: 15m + labels: + severity: warning + annotations: + summary: "No CRL refresh has been recorded" + description: "The operator has not refreshed any CRL since it started. Revocations are not reaching agents." +``` + ### Stale CRL A CRL that is no longer refreshed means revoked agents keep being accepted. @@ -132,6 +150,18 @@ Alert when the last successful refresh is more than a day old: ### CA Expiring +!!! note "CAs and certificates share one metric" + + `openvox_certificate_expiry_timestamp_seconds` is written by both the + Certificate and the CertificateAuthority controller, with the same + `name`/`namespace` labels and nothing that says which kind a series belongs + to. Telling them apart in a query means matching on the name. + + The rule below uses `.*-ca`, which is a convention rather than a guarantee: + it also catches a Certificate that happens to end in `-ca`, and it misses a + CertificateAuthority named otherwise. Replace the matcher with your actual + CA names if you rely on the distinction. + Alert when a CA certificate expires within 90 days: ```yaml diff --git a/docs/reference/nodeclassifier.md b/docs/reference/nodeclassifier.md index cd5c612c..db7b1c84 100644 --- a/docs/reference/nodeclassifier.md +++ b/docs/reference/nodeclassifier.md @@ -80,8 +80,10 @@ kind: Config metadata: name: production spec: - image: ... authorityRef: production-ca + image: + repository: ghcr.io/slauger/openvox-server-8 + tag: "latest" nodeClassifierRef: foreman ``` From baa6b554b34d0aa815d0ac26db06b86dbcd23073 Mon Sep 17 00:00:00 2001 From: Simon Lauger Date: Fri, 4 Sep 2026 00:46:05 +0200 Subject: [PATCH 7/7] docs: stop advertising JVM auto-tuning that does not happen README and the feature list promise 'Heap size calculated from memory limits (90%) - no manual -Xmx tuning needed'. The controller does contain that calculation, but it is unreachable: ServerSpec.JavaArgs carries a CRD default, so the field is never empty and the first branch of resolveJavaArgs always wins. Every Server runs with -Xms512m -Xmx1024m no matter how much memory it is given. Replaced the claim with what actually applies, and noted it on the field in the Server reference with a link to #592, where the fix is tracked. Once the default is removed the derivation works and the wording can go back. Third instance of the same pattern after #550 and #576: a CRD default takes away the empty state a fallback depends on. --- README.md | 2 +- docs/_snippets/features.md | 2 +- docs/reference/server.md | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index adb3473a..be8ed47a 100644 --- a/README.md +++ b/README.md @@ -17,7 +17,7 @@ A Kubernetes Operator that maps [OpenVox Server](https://github.com/OpenVoxProje - 🔄 **Multi-Version Deployments** - Run different server versions side by side - canary deployments, rolling upgrades - 🔒 **Rootless & OpenShift Ready** - Random UID compatible, no root, no ezbake, no privilege escalation - ðŸŠķ **Minimal Image** - UBI9-based, no agent Ruby, no ezbake packaging - smaller footprint, fewer updates -- 🧠 **Auto-tuned JVM** - Heap size calculated from memory limits (90%) - no manual `-Xmx` tuning needed +- 🧠 **JVM sizing** - Set `javaArgs` per Server or Database; see [Server](docs/reference/server.md) for the current default - ðŸ“Ķ **OCI Image Volumes** - Package Puppet code as OCI images, deploy immutably with automatic rollout (K8s 1.35+) - 🌐 **Gateway API** - SNI-based TLSRoute support - share a single LoadBalancer across environments via TLS passthrough - 🗄ïļ **Managed OpenVox DB** - Deploy OpenVox DB (PuppetDB) with external PostgreSQL - TLS, config, and credentials managed by the operator diff --git a/docs/_snippets/features.md b/docs/_snippets/features.md index 39bb1998..72473a9b 100644 --- a/docs/_snippets/features.md +++ b/docs/_snippets/features.md @@ -6,7 +6,7 @@ - 🔄 **Multi-Version Deployments** - Run different server versions side by side - canary deployments, rolling upgrades - 🔒 **Rootless & OpenShift Ready** - Random UID compatible, no root, no ezbake, no privilege escalation - ðŸŠķ **Minimal Image** - UBI9-based, no agent Ruby, no ezbake packaging - smaller footprint, fewer updates -- 🧠 **Auto-tuned JVM** - Heap size calculated from memory limits (90%) - no manual `-Xmx` tuning needed +- 🧠 **JVM sizing** - Set `javaArgs` per Server or Database; see [Server](reference/server.md) for the current default - ðŸ“Ķ **OCI Image Volumes** - Package Puppet code as OCI images, deploy immutably with automatic rollout (K8s 1.35+) - 🌐 **Gateway API** - SNI-based TLSRoute support - share a single LoadBalancer across environments via TLS passthrough - 🗄ïļ **Managed OpenVox DB** - Deploy OpenVox DB (PuppetDB) with external PostgreSQL - TLS, config, and credentials managed by the operator diff --git a/docs/reference/server.md b/docs/reference/server.md index fa1cd6b6..0995066c 100644 --- a/docs/reference/server.md +++ b/docs/reference/server.md @@ -33,7 +33,7 @@ spec: | `replicas` | int32 | `1` | Number of pod replicas | | `autoscaling` | [AutoscalingSpec](#autoscalingspec) | - | HPA configuration | | `resources` | ResourceRequirements | - | CPU/memory requests and limits | -| `javaArgs` | string | `-Xms512m -Xmx1024m` | JVM arguments | +| `javaArgs` | string | `-Xms512m -Xmx1024m` | JVM arguments. The controller can derive the heap from the memory limit, but the CRD default makes that path unreachable today - see [#592](https://github.com/slauger/openvox-operator/issues/592). Set this explicitly to size the heap | | `maxActiveInstances` | int32 | `1` | Number of JRuby instances per pod | | `code` | [[]CodeSpec](index.md#codespec) | - | Override the Config's code sources (replace, not merge). A list; see [CodeSpec](index.md#codespec) | | `topologySpreadConstraints` | []TopologySpreadConstraint | - | Pod spread constraints across topology domains |