Skip to content

fix: disable argo-helm 10.x default NetworkPolicies - #69

Merged
venkatamutyala merged 1 commit into
mainfrom
fix/disable-default-networkpolicies
Sep 18, 2026
Merged

venkatamutyala merged 1 commit into
mainfrom
fix/disable-default-networkpolicies

Conversation

@venkatamutyala

Copy link
Copy Markdown
Contributor

Problem

argo-helm chart 10.0.0 (released 2026-06-26) flipped global.networkPolicy.create from false to true. Its entire changelog entry for this is:

artifacthub.io/changes: |
  - kind: fixed
    description: Enable network policies for all components by default

Filed as kind: fixed, despite flipping every Argo CD component to default-deny ingress.

v0.23.0 (#67) moved these values to chart 10.2.2 and set networkPolicy.create: true explicitly — which matches the new chart default, so the policies ship either way.

Each generated policy sets policyTypes: [Ingress], which makes the selected pod default-deny and re-opens only what it lists. The metrics allow-rules are written as:

- from:
  - namespaceSelector: {}
  ports:
  - port: metrics      # by NAME, and no protocol

On a GlueOps cluster those rules do not take effect.

Impact, measured on a live cluster

Prometheus went from 75 up / 7 down, with every failure an Argo CD metrics endpoint:

target port result
argocd-application-controller-metrics 8082 connection refused
argocd-repo-server-metrics 8084 connection refused (×4 pods)
argocd-applicationset-controller-metrics 8080 connection refused
argocd-notifications-controller-metrics 9001 connection refused

argocd-server-metrics:8083 stayed up — its policy renders as ingress: [{}], a blanket allow, so it restricts nothing.

Worth noting the failure mode is connection refused, not a timeout, which reads like "nothing is listening" and sent the first round of diagnosis down the wrong path. This CNI rejects with RST rather than dropping.

Verification

helm upgrade --reuse-values --set global.networkPolicy.create=false on a live cluster. Dry-run diff first: the only change is the removal of the 5 NetworkPolicy objects (152 lines), nothing else.

Result: 82 up / 0 down. Every Argo CD target recovered on the same pod IPs that were refusing connections, and no pod restarted — so the metrics servers had been listening the whole time and the policies were blocking them. All Applications stayed Synced/Healthy.

Why false rather than fixing the rule

The rule is hardcoded in the chart template — no value reshapes it — and it is byte-identical in the latest chart, 10.9.2, so upgrading does not help. The per-component toggles are or'd with the global one (if or .Values.X.networkPolicy.create .Values.global.networkPolicy.create), so a component-level false cannot override a global true.

Chart 9.3.7 defaulted to false, so this preserves the behaviour every existing cluster already runs rather than changing it. It also decouples two unrelated risks currently bundled together: upgrading Argo CD 3.2.12 → 3.4.9, and introducing default-deny NetworkPolicies across every component — the latter arriving silently via a chart default, during a Kubernetes 1.35 platform upgrade, while removing the alerting you would use to catch problems.

Follow-up (not this PR)

Re-enabling is worth doing deliberately. repo-server is the policy with real value: it restricts port 8081 — which holds repository credentials and executes Helm/Kustomize rendering — to the four Argo CD components. The shape to aim for is create: true plus a supplemental allow-policy via extraObjects for the metrics ports by number, with scrape targets verified before rollout.

Also worth knowing: these policies are Ingress-only. There are no egress restrictions on Argo CD, which is one of the highest-privilege workloads in the cluster.

Related

Upstream has a recurring pattern of these policies breaking metrics and webhooks: argoproj/argo-helm#3431 → #3501, #3678 → #3681, #3915 → #3930, #3948 → #3950, #3569.

Blocks/relates to GlueOps/terraform-module-cloud-multy-prerequisites#741, which pins docs-argocd v0.23.0.

🤖 Generated with Claude Code

argo-helm chart 10.0.0 flipped global.networkPolicy.create from false to true.
Each generated policy sets policyTypes: [Ingress], making the pod default-deny,
and the metrics allow-rules use an empty namespaceSelector with the port named
rather than numbered. On a live GlueOps cluster those rules do not take effect:
kube-prometheus-stack lost the application-controller, repo-server, applicationset
and notifications scrape targets (7 down). Removing the policies recovered all 7
with no pod restart, confirming the policies were the cause.

Chart 9.3.7 defaulted to false, so this preserves the behaviour existing clusters
already run instead of bundling a default-deny change into the 3.2.12 -> 3.4.9
upgrade. Re-enabling should be a separate change with a supplemental allow-policy
for the metrics ports by number.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@venkatamutyala
venkatamutyala merged commit 2033176 into main Sep 18, 2026
2 checks passed
@venkatamutyala
venkatamutyala deleted the fix/disable-default-networkpolicies branch September 18, 2026 14:06
venkatamutyala added a commit to GlueOps/terraform-module-cloud-multy-prerequisites that referenced this pull request Sep 18, 2026
Picks up GlueOps/docs-argocd#69, which sets global.networkPolicy.create: false.

argo-helm chart 10.0.0 flipped that default from false to true, adding a
NetworkPolicy per Argo CD component. Each sets policyTypes: [Ingress], making
the pod default-deny, and the metrics allow-rules use an empty namespaceSelector
with the port referenced by name. On a GlueOps cluster those rules do not take
effect: kube-prometheus-stack lost the application-controller (8082), repo-server
(8084), applicationset (8080) and notifications (9001) scrape targets.

Verified on a live cluster: 75 up / 7 down before, 82 up / 0 down after, with the
targets recovering on the same pod IPs and no pod restart.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
hamzabouissi added a commit to GlueOps/terraform-module-cloud-multy-prerequisites that referenced this pull request Sep 18, 2026
* feat: bump versions for argocd, platform. terraform_module

* fix: bump platform version

* fix: pin Argo CD v3.4.9 and leave the released CHANGELOG entry alone (#742)

- argocd_app_version v3.4.6 -> v3.4.9: newest 3.4 patch (2026-09-14) and the
  version platform-crds pins (glueops.dev/pin.argo-cd). Chart 10.2.2 ships
  v3.4.6 and no 10.x chart ships 3.4.7+, so the image tag override carries
  the patch, as it does today (9.3.7 ships v3.2.6, we run v3.2.12).
- CHANGELOG.md restored to main: release-please owns it, and the edited
  0.92.3 entry described the v0.79.2 bump that actually happened.


Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* feat: take the released docs-argocd v0.23.0 and codespaces v0.162.0 (#743)

* feat: take the released docs-argocd v0.23.0 and codespaces v0.162.0

- docs-argocd module ref v0.22.0 -> v0.23.0 (docs-argocd#67): values for
  Argo CD v3.4.9 on argo-helm 10.2.2, explicit NetworkPolicies, extension
  installer fail-open, extension fetched through the raw-github proxy.
  Required with argocd_helm_chart_version 10.2.2.
- codespace_version v0.161.1 -> v0.162.0 (codespaces#602): argocd CLI 3.4.9
  matching the server, kubectl 1.34.11, opentofu 1.11.14 and other patch
  bumps.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

* docs: automated update of terraform docs

* fix: platform_crds_version will be v0.1.5, not v0.2.0

platform-crds' release-please sets bump-patch-for-minor-pre-major, and
platform-crds#83 is a plain feat: PR, so the release carrying the 1.35 CRD
bundle will be v0.1.5.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01FSau4e6nS8Y7eM1nQHwq6k

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>

* feat: update AWS EKS versions in aws.yaml (#751)

* fix: bump docs-argocd to v0.23.1 (#752)

Picks up GlueOps/docs-argocd#69, which sets global.networkPolicy.create: false.

argo-helm chart 10.0.0 flipped that default from false to true, adding a
NetworkPolicy per Argo CD component. Each sets policyTypes: [Ingress], making
the pod default-deny, and the metrics allow-rules use an empty namespaceSelector
with the port referenced by name. On a GlueOps cluster those rules do not take
effect: kube-prometheus-stack lost the application-controller (8082), repo-server
(8084), applicationset (8080) and notifications (9001) scrape targets.

Verified on a live cluster: 75 up / 7 down before, 82 up / 0 down after, with the
targets recovering on the same pod IPs and no pod restart.

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Venkat <venkata@venkatamutyala.com>
Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: github-actions[bot] <github-actions[bot]@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants