Skip to content

fix: drop dead app-chart values and set qr-code-generator replicas via deployment.replicas - #1500

Merged
venkatamutyala merged 1 commit into
feat/k8s-1.35-component-upgradesfrom
fix/app-chart-dead-keys
Sep 17, 2026
Merged

venkatamutyala merged 1 commit into
feat/k8s-1.35-component-upgradesfrom
fix/app-chart-dead-keys

Conversation

@venkatamutyala

Copy link
Copy Markdown
Contributor

Follow-up to #1498 (targets its branch). Reviewing the app chart 0.13.0 -> 0.14.1 bump turned up values keys the chart has never read (verified against the templates at 0.13.0 and 0.14.1):

  • image.pullPolicy: the chart reads deployment.imagePullPolicy (_podTemplate.tpl), which these Applications already set.
  • top-level replicaCount: the chart reads deployment.replicas (deployment.yaml).

What

Application Change Rendered effect
cluster-info-page, go-healthz, pull-request-bot remove image.pullPolicy none, byte-identical
qr-code-generator remove replicaCount: '2' and image.pullPolicy; add deployment.replicas: 2 and deployment.imagePullPolicy: IfNotPresent replicas: 1 -> 2, and the chart's multi-replica strategy applies (maxSurge: 50%, maxUnavailable: 1 instead of 100%/0)

⚠️ The qr-code-generator change is a real behaviour change: the file has said 2 replicas since #1168, but the chart silently applied its default of 1. If 1 replica was actually fine, drop the replicas: 2 line and only the dead keys go away.

Not touched: application-loki-alert-group-controller.yaml also carries image.pullPolicy, but that file is on chart 0.9.0 and is reserved for #1486.

Verification

  • helm lint . -f ci/values.yaml: clean.
  • For each of the four Applications, the exact helm.values from the platform render before/after were templated through app chart 0.14.1 and diffed: three identical, qr-code-generator differs only in replicas and the two strategy fields.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

…a deployment.replicas

The GlueOps app chart never reads image.pullPolicy (it reads
deployment.imagePullPolicy, which these Applications already set) or a
top-level replicaCount (it reads deployment.replicas). Verified against the
chart templates at 0.13.0 and 0.14.1.

- cluster-info-page, go-healthz, pull-request-bot: remove image.pullPolicy.
  Rendered manifests are byte-identical.
- qr-code-generator: replace the ignored replicaCount: '2' and
  image.pullPolicy with deployment.replicas: 2 and
  deployment.imagePullPolicy: IfNotPresent. This changes the rendered
  Deployment from 1 replica (the chart default that was silently applied) to
  the 2 replicas the file has intended since #1168, and the chart's
  multi-replica rollout strategy (maxSurge 50%, maxUnavailable 1) follows.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k
@venkatamutyala
venkatamutyala merged commit ba2fadf into feat/k8s-1.35-component-upgrades Sep 17, 2026
2 checks passed
@venkatamutyala
venkatamutyala deleted the fix/app-chart-dead-keys branch September 17, 2026 12:15
hamzabouissi added a commit that referenced this pull request Sep 18, 2026
* feat!: upgrade platform components for Kubernetes 1.35

Bump charts and images that are unsupported or untested on Kubernetes 1.35,
plus low-risk bumps that were already pending.

- cert-manager v1.18.2 -> v1.21.2
- ingress-nginx 4.13.3 / v1.13.3 -> 4.15.1 / v1.15.1
- external-secrets 0.19.2 / v0.16.2 -> 2.10.0 / v2.10.0
- metacontroller v4.12.5 -> v4.17.2
- traefik 39.0.0 / v3.6.7 -> 41.5.0 / v3.7.13 (logs -> log/accessLog)
- openbao 0.19.3 / 2.4.4 -> 0.29.4 / 2.6.2
- external-dns 1.20.0 / v0.20.0 -> 1.22.0 / v0.22.0 (pin annotationPrefix
  to external-dns.alpha.kubernetes.io/)
- reflector 10.0.65, keda 2.20.2, goldilocks 11.1.0 / v4.16.1
- dex v2.45.1, oauth2-proxy v7.15.4, backup-tools v2.17.0,
  network_exporter 1.8.0, curl 8.22.0
- app chart 0.13.0/0.8.1 -> 0.14.1 (non-monitoring apps)

BREAKING CHANGE: requires platform-crds with matching CRDs (cert-manager
v1.21.2, external-secrets v2.10.0, metacontroller v4.17.2, traefik chart
41.5.0, keda v2.20.2) to be applied before this release syncs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SW6t6HkN3k4wYusAkQ55eS

* fix: remove etcd secret for metrics in prometheus

* fix: pin the reflector image and pull it through the dockerhub mirror (#1499)

reflector was the only component whose image was neither routed through a
*.repo.gpkg.io mirror nor pinned under container_images; the chart pulled
docker.io/emberstack/kubernetes-reflector at its appVersion with no digest.

Add container_images.app_reflector with the 10.0.65 multi-arch index digest
(resolved from Docker Hub and confirmed served by dockerhub.repo.gpkg.io) and
pass it to the chart's image.repository/image.tag. Rendered image:
dockerhub.repo.gpkg.io/emberstack/kubernetes-reflector:10.0.65@sha256:51dbd58...

README regenerated with helm-docs 1.14.2.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: drop dead app-chart values and set qr-code-generator replicas via deployment.replicas (#1500)

The GlueOps app chart never reads image.pullPolicy (it reads
deployment.imagePullPolicy, which these Applications already set) or a
top-level replicaCount (it reads deployment.replicas). Verified against the
chart templates at 0.13.0 and 0.14.1.

- cluster-info-page, go-healthz, pull-request-bot: remove image.pullPolicy.
  Rendered manifests are byte-identical.
- qr-code-generator: replace the ignored replicaCount: '2' and
  image.pullPolicy with deployment.replicas: 2 and
  deployment.imagePullPolicy: IfNotPresent. This changes the rendered
  Deployment from 1 replica (the chart default that was silently applied) to
  the 2 replicas the file has intended since #1168, and the chart's
  multi-replica rollout strategy (maxSurge 50%, maxUnavailable 1) follows.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat!: bump VPA to chart 5.1.0 / 1.7.1 to match goldilocks 4.16.1 (#1501)

VPA 1.7.x supports Kubernetes 1.28+ (upstream docs/installation.md at
vertical-pod-autoscaler-1.7.1); only the alpha InPlace/InPlaceOrRecreate/
CPUStartupBoost features need 1.33+. Nothing requires 1.35, so there is no
reason to hold VPA back from the 1.35 component upgrade. goldilocks 4.16.1
(already in this branch) links the VPA 1.7.1 library and its chart pairs
with vpa 5.0.*.

- templates/application-vpa.yaml: chart 4.12.3 -> 5.1.0. Chart 5.x only
  changes appVersion, kubeVersion (>= 1.28.0-0) and the test image; with our
  values the render differs only in the recommender image and labels.
- values.yaml: recommender/updater/admission-controller 1.6.0 -> 1.7.1 with
  multi-arch index digests (identical on registry.k8s.io and k8s.repo.gpkg.io).
  Only the recommender is enabled.

BREAKING CHANGE: platform-crds must ship the VPA 1.7.1 CRD
(GlueOps/platform-crds#75) before this syncs. The CRD change is additive
(InPlace update mode, startupBoost, evictAfterOOMSeconds,
memoryAggregationIntervalSeconds, status.observedGeneration).

README regenerated with helm-docs 1.14.2.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: pin the KEDA images by digest under container_images (#1502)

* fix: pin the KEDA images by digest under container_images

KEDA was the only component whose images had neither a pinned tag nor a
digest: the Application only set global.image.registry, so the three images
followed the chart appVersion as mutable tags.

Add container_images.app_keda (operator, metrics-apiserver,
admission-webhooks) with the 2.20.2 multi-arch index digests, resolved from
ghcr.io and confirmed identical on ghcr.repo.gpkg.io, and pass them to the
chart's per-component image.{keda,metricsApiServer,webhooks} values. The
global.image.registry override is dropped because the chart lets it take
precedence over the per-image registry, which would silently defeat the pin.

README regenerated with helm-docs 1.14.2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

* style: match the 4-space indentation used by the rest of the keda values block

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

* fix: keep global.image.registry so base_registries always decides the KEDA registry

Per-image registry values are dropped from the Application because the chart
gives global.image.registry precedence; only repository and tag come from
container_images.app_keda.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: enforce the cert-manager image digests via image.tag parameters (#1503)

* fix: enforce the cert-manager image digests via image.tag parameters

container_images.app_cert_manager.cert_manager.image.tag carried a digest that
no template referenced: only registry and repository were passed to the chart,
so all five cert-manager images rendered as bare :v1.21.2 tags following the
chart appVersion.

Add webhook, cainjector, acmesolver and startupapicheck entries under
container_images.app_cert_manager (each with registry, repository and
tag@digest resolved from quay.io and confirmed identical on quay.repo.gpkg.io)
and pass every component's repository and tag as chart parameters. The chart
renders image:tag@digest for each, including the --acme-http01-solver-image
argument.

README regenerated with helm-docs 1.14.2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

* fix: drive the cert-manager registry from base_registries via the chart's imageRegistry

Replace the five per-image repository parameters with the chart's global
imageRegistry (from base_registries.quay_io). The chart composes
imageRegistry/imageNamespace/<image.name> for every component and falls back
to quay.io/jetstack when imageRegistry is absent. Per-image tag@digest
parameters stay. Rendered images are unchanged on a real cluster.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat: update backup-tools to v2.18.0 (#1504)

Bundles the OpenBao CLI 2.6.2, matching the OpenBao server this branch
deploys, and switches to the renamed upstream release asset. Also aws-cli
2.36.25, gh 2.97.0, loki logcli 3.7.6, refreshed ubuntu base. The
backup-vault script the platform runs is unchanged.

Digest is the manifest digest served identically by ghcr.io and
ghcr.repo.gpkg.io.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: keep the external-secrets webhook readiness port in the 450xx host-port range (#1506)

external-secrets 2.10.0 declares the webhook readiness port as a containerPort
(upstream v1.3.0). With webhook.hostNetwork=true that becomes a host port, and
the chart default 8081 sits outside the platform's reserved 450xx range.

- values.yaml: host_network.external_secrets.webhook_readiness_port: 45012
- application-external-secrets.yaml: pass it as webhook.readinessProbe.port,
  which drives both --healthz-addr and the containerPort.
- Remove the dead inline webhook.port: 10751 (the webhook.port parameter
  overrides it) and the CRD ignoreDifferences block (this Application has not
  rendered CRDs since #1480).

Chart render diff: --healthz-addr=:8081 -> :45012 and containerPort 8081 ->
45012; nothing else changes.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat: update goldilocks to chart 11.1.1 / v4.16.2 (#1508)

v4.16.2 (2026-09-15) is a vulnerability-fix release (goldilocks #889) plus
CI/test changes; chart 11.1.1 only bumps appVersion. Rendered diff with our
values: the two container images. Image confirmed on gcp.repo.gpkg.io.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat: update vault-backup-validator to v2.19.0 (#1510)

Ships GlueOps/vault-backup-validator#255: bundled OpenBao 2.6.2 and the
renamed release asset name. With backup-tools v2.18.0 (#1504) both sides now
run OpenBao 2.6.2, matching the server, so the validator no longer
re-downloads a CLI on every backup run and restores snapshots on the same
version that produced them.

Digest is the manifest digest served identically by ghcr.io and the mirror.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: set openbao autopilot min_quorum to the actual voter count (#1507)

min_quorum is autopilot's floor for cleanup_dead_servers: it never prunes a
peer if doing so would take the cluster below this number. With 3 replicas
and min_quorum = 5 the cluster is always below the floor, so dead peers left
behind by an unplanned member replacement (pod recreated without its PVC)
were never removed and raft quorum would silently count them. Set it to 3,
matching server.ha.replicas, as upstream documents.

The StatefulSet pod template is unchanged (chart has no config checksum), so
this does not roll pods on its own; the value takes effect on each pod's next
restart, i.e. the 2.6.2 rollout in this branch.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: traefik chart 41.6.0 and move service.type under service.spec (#1509)

Chart 40.0.0 moved the Service spec fields under service.spec (#1686); a
top-level service.type has been ignored since, and LoadBalancer only rendered
because it is the chart default. Move it next to externalTrafficPolicy in all
three Applications so the value is honoured again. Verified: with the key in
its old place spec.type=ClusterIP renders LoadBalancer; in the new place it
renders ClusterIP.

Also bump 41.5.0 -> 41.6.0 (2026-09-16): PDB apiVersion chosen by kubeVersion,
Hub transparency-log values. Proxy stays v3.7.13 and the chart's crds/ are
byte-identical to 41.5.0, so the platform-crds pin is unaffected. All three
instances render identically apart from the chart label.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* chore: drop unused kubeadm.kube_etcd.serviceMonitor values

The etcd scrape moved to the plaintext metrics listener on 2381 in 0043e56,
which removed the only template references to these caFile/certFile/keyFile
values and to the etcd-client-certs secret mount. Chart render is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Venkat <venkata@venkatamutyala.com>
Co-authored-by: venkatamutyala <venkata.mutyala@glueops.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant