Skip to content

feat!: bump VPA to chart 5.1.0 / 1.7.1 to match goldilocks 4.16.1 - #1501

Merged
venkatamutyala merged 1 commit into
feat/k8s-1.35-component-upgradesfrom
feat/vpa-1.7.1
Sep 17, 2026
Merged

venkatamutyala merged 1 commit into
feat/k8s-1.35-component-upgradesfrom
feat/vpa-1.7.1

Conversation

@venkatamutyala

Copy link
Copy Markdown
Contributor

Follow-up to #1498 (targets its branch). The PR body says VPA 1.7 must wait because "1.7 is listed only from 1.35". That is not what upstream documents, so this brings VPA along with the rest of the 1.35 bumps.

Evidence

  • Upstream compatibility table at tag vertical-pod-autoscaler-1.7.1 (docs/installation.md): 1.7.x -> Kubernetes 1.28+, with InPlace, InPlaceOrRecreate and CPUStartupBoost requiring 1.33+. Same row shape as 1.6.x. Nothing mentions 1.35.
  • Fairwinds vpa chart 5.x README ("BREAKING Upgrading to >=5.0.0"): "VPA 1.7 requires Kubernetes 1.28 or later"; kubeVersion: >= 1.28.0-0.
  • goldilocks v4.16.0 (breaking: the dependency kube-prometheus-stack has been updated to a new major version (62.7.0), which may include breaking changes. #major #879) moved its k8s.io/autoscaler/vertical-pod-autoscaler dependency 1.4.1 -> 1.7.1, and goldilocks chart 11.1.0 (already in feat!: upgrade platform components for Kubernetes 1.35 #1498) bundles vpa 5.0.*. Running goldilocks 4.16.1 against VPA 1.6.0 works (it only uses autoscaling.k8s.io/v1), but 1.7.1 is what upstream pairs it with.
  • VPA 1.7.0 release notes: InPlace update mode, startupBoost, per-VPA evictAfterOOMSeconds / memoryAggregationIntervalSeconds, status.observedGeneration, memory reduction in informers, k8s deps 1.36. 1.7.1: recommender checkpoint-loading regression fix, Prometheus history flag fix. All alpha features stay off by default, and we run only the recommender.

What

Before After
chart vpa 4.12.3 5.1.0 (latest; 5.0.1 -> 5.1.0 is only the metrics-server subchart, which we disable)
recommender / updater / admission-controller images 1.6.0 (no digest) 1.7.1 with index digests, identical on registry.k8s.io and k8s.repo.gpkg.io

README regenerated with helm-docs 1.14.2. Supersedes Renovate #1479 and #1496.

⚠️ Rollout prerequisite

platform-crds must ship the VPA 1.7.1 CRD first (GlueOps/platform-crds#75, open; the 1.35 CRD bundle PR #83 excludes VPA). The chart README is explicit that CRDs go first. The CRD diff 1.6.0 -> 1.7.1 is purely additive (no removed lines apart from the controller-gen annotation): InPlace in the updateMode enum, startupBoost, evictAfterOOMSeconds, memoryAggregationIntervalSeconds, status.observedGeneration. v1 stays served+storage. Without the new CRD the 1.7.1 recommender still runs (it JSON-patches status, and the API server prunes the unknown observedGeneration), but do it in the documented order.

Verification

  • helm lint . -f ci/values.yaml: clean.
  • Platform render diff vs the PR branch: only targetRevision and the three image tags in the vpa Application.
  • vpa chart 4.12.3 (old values) vs 5.1.0 (new values), templated with the exact Application values, --skip-crds: 23 objects both sides; the only diffs are helm.sh/chart / app.kubernetes.io/version labels, the recommender image, and the alpine/kubectl image in the chart's helm.sh/hook: test pods (ignored by Argo CD). RBAC, args, probes, scheduling unchanged. Updater and admission controller stay disabled.

🤖 Generated with Claude Code

https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

VPA 1.7.x supports Kubernetes 1.28+ (upstream docs/installation.md at
vertical-pod-autoscaler-1.7.1); only the alpha InPlace/InPlaceOrRecreate/
CPUStartupBoost features need 1.33+. Nothing requires 1.35, so there is no
reason to hold VPA back from the 1.35 component upgrade. goldilocks 4.16.1
(already in this branch) links the VPA 1.7.1 library and its chart pairs
with vpa 5.0.*.

- templates/application-vpa.yaml: chart 4.12.3 -> 5.1.0. Chart 5.x only
  changes appVersion, kubeVersion (>= 1.28.0-0) and the test image; with our
  values the render differs only in the recommender image and labels.
- values.yaml: recommender/updater/admission-controller 1.6.0 -> 1.7.1 with
  multi-arch index digests (identical on registry.k8s.io and k8s.repo.gpkg.io).
  Only the recommender is enabled.

BREAKING CHANGE: platform-crds must ship the VPA 1.7.1 CRD
(GlueOps/platform-crds#75) before this syncs. The CRD change is additive
(InPlace update mode, startupBoost, evictAfterOOMSeconds,
memoryAggregationIntervalSeconds, status.observedGeneration).

README regenerated with helm-docs 1.14.2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k
@venkatamutyala
venkatamutyala merged commit e508077 into feat/k8s-1.35-component-upgrades Sep 17, 2026
2 checks passed
@venkatamutyala
venkatamutyala deleted the feat/vpa-1.7.1 branch September 17, 2026 12:24
hamzabouissi added a commit that referenced this pull request Sep 18, 2026
* feat!: upgrade platform components for Kubernetes 1.35

Bump charts and images that are unsupported or untested on Kubernetes 1.35,
plus low-risk bumps that were already pending.

- cert-manager v1.18.2 -> v1.21.2
- ingress-nginx 4.13.3 / v1.13.3 -> 4.15.1 / v1.15.1
- external-secrets 0.19.2 / v0.16.2 -> 2.10.0 / v2.10.0
- metacontroller v4.12.5 -> v4.17.2
- traefik 39.0.0 / v3.6.7 -> 41.5.0 / v3.7.13 (logs -> log/accessLog)
- openbao 0.19.3 / 2.4.4 -> 0.29.4 / 2.6.2
- external-dns 1.20.0 / v0.20.0 -> 1.22.0 / v0.22.0 (pin annotationPrefix
  to external-dns.alpha.kubernetes.io/)
- reflector 10.0.65, keda 2.20.2, goldilocks 11.1.0 / v4.16.1
- dex v2.45.1, oauth2-proxy v7.15.4, backup-tools v2.17.0,
  network_exporter 1.8.0, curl 8.22.0
- app chart 0.13.0/0.8.1 -> 0.14.1 (non-monitoring apps)

BREAKING CHANGE: requires platform-crds with matching CRDs (cert-manager
v1.21.2, external-secrets v2.10.0, metacontroller v4.17.2, traefik chart
41.5.0, keda v2.20.2) to be applied before this release syncs.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01SW6t6HkN3k4wYusAkQ55eS

* fix: remove etcd secret for metrics in prometheus

* fix: pin the reflector image and pull it through the dockerhub mirror (#1499)

reflector was the only component whose image was neither routed through a
*.repo.gpkg.io mirror nor pinned under container_images; the chart pulled
docker.io/emberstack/kubernetes-reflector at its appVersion with no digest.

Add container_images.app_reflector with the 10.0.65 multi-arch index digest
(resolved from Docker Hub and confirmed served by dockerhub.repo.gpkg.io) and
pass it to the chart's image.repository/image.tag. Rendered image:
dockerhub.repo.gpkg.io/emberstack/kubernetes-reflector:10.0.65@sha256:51dbd58...

README regenerated with helm-docs 1.14.2.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: drop dead app-chart values and set qr-code-generator replicas via deployment.replicas (#1500)

The GlueOps app chart never reads image.pullPolicy (it reads
deployment.imagePullPolicy, which these Applications already set) or a
top-level replicaCount (it reads deployment.replicas). Verified against the
chart templates at 0.13.0 and 0.14.1.

- cluster-info-page, go-healthz, pull-request-bot: remove image.pullPolicy.
  Rendered manifests are byte-identical.
- qr-code-generator: replace the ignored replicaCount: '2' and
  image.pullPolicy with deployment.replicas: 2 and
  deployment.imagePullPolicy: IfNotPresent. This changes the rendered
  Deployment from 1 replica (the chart default that was silently applied) to
  the 2 replicas the file has intended since #1168, and the chart's
  multi-replica rollout strategy (maxSurge 50%, maxUnavailable 1) follows.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat!: bump VPA to chart 5.1.0 / 1.7.1 to match goldilocks 4.16.1 (#1501)

VPA 1.7.x supports Kubernetes 1.28+ (upstream docs/installation.md at
vertical-pod-autoscaler-1.7.1); only the alpha InPlace/InPlaceOrRecreate/
CPUStartupBoost features need 1.33+. Nothing requires 1.35, so there is no
reason to hold VPA back from the 1.35 component upgrade. goldilocks 4.16.1
(already in this branch) links the VPA 1.7.1 library and its chart pairs
with vpa 5.0.*.

- templates/application-vpa.yaml: chart 4.12.3 -> 5.1.0. Chart 5.x only
  changes appVersion, kubeVersion (>= 1.28.0-0) and the test image; with our
  values the render differs only in the recommender image and labels.
- values.yaml: recommender/updater/admission-controller 1.6.0 -> 1.7.1 with
  multi-arch index digests (identical on registry.k8s.io and k8s.repo.gpkg.io).
  Only the recommender is enabled.

BREAKING CHANGE: platform-crds must ship the VPA 1.7.1 CRD
(GlueOps/platform-crds#75) before this syncs. The CRD change is additive
(InPlace update mode, startupBoost, evictAfterOOMSeconds,
memoryAggregationIntervalSeconds, status.observedGeneration).

README regenerated with helm-docs 1.14.2.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: pin the KEDA images by digest under container_images (#1502)

* fix: pin the KEDA images by digest under container_images

KEDA was the only component whose images had neither a pinned tag nor a
digest: the Application only set global.image.registry, so the three images
followed the chart appVersion as mutable tags.

Add container_images.app_keda (operator, metrics-apiserver,
admission-webhooks) with the 2.20.2 multi-arch index digests, resolved from
ghcr.io and confirmed identical on ghcr.repo.gpkg.io, and pass them to the
chart's per-component image.{keda,metricsApiServer,webhooks} values. The
global.image.registry override is dropped because the chart lets it take
precedence over the per-image registry, which would silently defeat the pin.

README regenerated with helm-docs 1.14.2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

* style: match the 4-space indentation used by the rest of the keda values block

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

* fix: keep global.image.registry so base_registries always decides the KEDA registry

Per-image registry values are dropped from the Application because the chart
gives global.image.registry precedence; only repository and tag come from
container_images.app_keda.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: enforce the cert-manager image digests via image.tag parameters (#1503)

* fix: enforce the cert-manager image digests via image.tag parameters

container_images.app_cert_manager.cert_manager.image.tag carried a digest that
no template referenced: only registry and repository were passed to the chart,
so all five cert-manager images rendered as bare :v1.21.2 tags following the
chart appVersion.

Add webhook, cainjector, acmesolver and startupapicheck entries under
container_images.app_cert_manager (each with registry, repository and
tag@digest resolved from quay.io and confirmed identical on quay.repo.gpkg.io)
and pass every component's repository and tag as chart parameters. The chart
renders image:tag@digest for each, including the --acme-http01-solver-image
argument.

README regenerated with helm-docs 1.14.2.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

* fix: drive the cert-manager registry from base_registries via the chart's imageRegistry

Replace the five per-image repository parameters with the chart's global
imageRegistry (from base_registries.quay_io). The chart composes
imageRegistry/imageNamespace/<image.name> for every component and falls back
to quay.io/jetstack when imageRegistry is absent. Per-image tag@digest
parameters stay. Rendered images are unchanged on a real cluster.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat: update backup-tools to v2.18.0 (#1504)

Bundles the OpenBao CLI 2.6.2, matching the OpenBao server this branch
deploys, and switches to the renamed upstream release asset. Also aws-cli
2.36.25, gh 2.97.0, loki logcli 3.7.6, refreshed ubuntu base. The
backup-vault script the platform runs is unchanged.

Digest is the manifest digest served identically by ghcr.io and
ghcr.repo.gpkg.io.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: keep the external-secrets webhook readiness port in the 450xx host-port range (#1506)

external-secrets 2.10.0 declares the webhook readiness port as a containerPort
(upstream v1.3.0). With webhook.hostNetwork=true that becomes a host port, and
the chart default 8081 sits outside the platform's reserved 450xx range.

- values.yaml: host_network.external_secrets.webhook_readiness_port: 45012
- application-external-secrets.yaml: pass it as webhook.readinessProbe.port,
  which drives both --healthz-addr and the containerPort.
- Remove the dead inline webhook.port: 10751 (the webhook.port parameter
  overrides it) and the CRD ignoreDifferences block (this Application has not
  rendered CRDs since #1480).

Chart render diff: --healthz-addr=:8081 -> :45012 and containerPort 8081 ->
45012; nothing else changes.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat: update goldilocks to chart 11.1.1 / v4.16.2 (#1508)

v4.16.2 (2026-09-15) is a vulnerability-fix release (goldilocks #889) plus
CI/test changes; chart 11.1.1 only bumps appVersion. Rendered diff with our
values: the two container images. Image confirmed on gcp.repo.gpkg.io.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat: update vault-backup-validator to v2.19.0 (#1510)

Ships GlueOps/vault-backup-validator#255: bundled OpenBao 2.6.2 and the
renamed release asset name. With backup-tools v2.18.0 (#1504) both sides now
run OpenBao 2.6.2, matching the server, so the validator no longer
re-downloads a CLI on every backup run and restores snapshots on the same
version that produced them.

Digest is the manifest digest served identically by ghcr.io and the mirror.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: set openbao autopilot min_quorum to the actual voter count (#1507)

min_quorum is autopilot's floor for cleanup_dead_servers: it never prunes a
peer if doing so would take the cluster below this number. With 3 replicas
and min_quorum = 5 the cluster is always below the floor, so dead peers left
behind by an unplanned member replacement (pod recreated without its PVC)
were never removed and raft quorum would silently count them. Set it to 3,
matching server.ha.replicas, as upstream documents.

The StatefulSet pod template is unchanged (chart has no config checksum), so
this does not roll pods on its own; the value takes effect on each pod's next
restart, i.e. the 2.6.2 rollout in this branch.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix: traefik chart 41.6.0 and move service.type under service.spec (#1509)

Chart 40.0.0 moved the Service spec fields under service.spec (#1686); a
top-level service.type has been ignored since, and LoadBalancer only rendered
because it is the chart default. Move it next to externalTrafficPolicy in all
three Applications so the value is honoured again. Verified: with the key in
its old place spec.type=ClusterIP renders LoadBalancer; in the new place it
renders ClusterIP.

Also bump 41.5.0 -> 41.6.0 (2026-09-16): PDB apiVersion chosen by kubeVersion,
Hub transparency-log values. Proxy stays v3.7.13 and the chart's crds/ are
byte-identical to 41.5.0, so the platform-crds pin is unaffected. All three
instances render identically apart from the chart label.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* chore: drop unused kubeadm.kube_etcd.serviceMonitor values

The etcd scrape moved to the plaintext metrics listener on 2381 in 0043e56,
which removed the only template references to these caFile/certFile/keyFile
values and to the etcd-client-certs secret mount. Chart render is unchanged.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Venkat <venkata@venkatamutyala.com>
Co-authored-by: venkatamutyala <venkata.mutyala@glueops.dev>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant