Skip to content

fix: Remove a deleted ingress sidecar pod's map rows and advertisements - #785

Merged
privateip merged 2 commits into
mainfrom
fix/issue-773
Oct 7, 2026
Merged

privateip merged 2 commits into
mainfrom
fix/issue-773

Conversation

@privateip

@privateip privateip commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A deleted ingress sidecar pod leaves its rows in the node's shared eBPF maps and its gateway advertisements in the API, and nothing removes either.

A replacement sidecar now removes its predecessor's rows at startup. The host deletes advertisements whose pod interface is gone, and removes the node's sidecar rows once every advertisement has stayed dead for two sweeps.

Either step can mistake a live sidecar for a gone one, so every sidecar now reads its own rows back each sweep and rewrites any that are missing, which bounds that mistake to one 5 second sweep.

Merge order

This touches the same files as #634, #613 and #598 and merges independently of them. Whichever lands second rebases.

Test plan

  • Deleting a sidecar pod with no replacement removes its rows and advertisements within about a minute
  • Replacing a sidecar pod removes the old pod's rows and advertisements while the new pod keeps serving
  • Rows removed from under a live sidecar come back within one sweep
  • Build, lint, and unit tests pass

Fixes #773

🤖 Generated with Claude Code

@privateip privateip self-assigned this Oct 7, 2026
@privateip privateip changed the title fix: Remove what a deleted ingress sidecar pod leaves on its node fix: Remove a deleted ingress sidecar pod's map rows and advertisements Oct 7, 2026
privateip and others added 2 commits October 7, 2026 14:11
The ingress sidecar leaves its rows in the node's shared eBPF maps when it exits, so a restarted container keeps serving. A deleted pod takes its VRFs and veths with it but leaves the rows, and the host's sweeps never touch the sidecar's share of the maps: GC keeps every vrf_table row under the sidecar's Block, nothing sweeps ifindex_vrf_table, and the egress route refresh skips sidecar tables. Its gateway BGPAdvertisements stay too, since they carry no netns annotation and the orphaned-CRD GC counts that as live. The installer then fails to install a return path for each one on every tick.

Two things now remove them:

- A new sidecar, after its startup inventory, removes every sidecar row its own VRFs do not account for. That covers a replaced pod, including VPCs the replacement no longer serves.
- The installer's sidecar return sweep deletes a sidecar advertisement whose recorded host-side interface is gone or is no longer a veth, with a UID and resourceVersion precondition so a republished one survives. When every sidecar advertisement on the node is dead on two sweeps in a row, with the same versions, it first removes every sidecar row from the maps. The second sweep covers a replacement pod that has written its rows but not yet republished, which a single sweep would strip. A failed interface lookup counts as unknown and is logged at debug level.

Rows are picked only by the sidecar's reserved key ranges (Block, ifindex base, table ID base), shared through a new sidecarmap package, so host rows are never touched. A node with no sidecar advertisements gives no evidence and keeps its rows.

Fixes #773

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two paths remove ingress sidecar rows from the node's shared eBPF maps, and either can mistake a live sidecar for a deleted one. A new sidecar's startup prune removes rows it does not track, which includes another live pod's when two sidecar pods overlap on a node. The installer's reaper removes every sidecar row once all sidecar advertisements look dead on two sweeps, which a replacement pod that fails to publish for that long cannot prevent. In both cases the live sidecar only rewrote its rows on a datapath reload, so its traffic stayed broken.

Each sweep, the sidecar now reads its own vrf_table and ifindex_vrf_table rows back and runs the same reapply a reload triggers when any are missing. That bounds either mistake to one sweep interval (5s by default). An unreadable map counts as present, matching how an unreadable generation is treated.

Related to #773

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@privateip
privateip marked this pull request as ready for review October 7, 2026 18:36
@privateip
privateip requested a review from a team as a code owner October 7, 2026 18:36
@privateip
privateip enabled auto-merge October 7, 2026 18:36
@privateip

Copy link
Copy Markdown
Collaborator Author

Second-pass review on head 1034ede.

pr-rereviewer: VERDICT merge. All applied findings were confirmed fixed, the final tree is byte-identical to the prior head, and the first commit builds and tests on its own.

pr-conventions-reviewer: VERDICT merge. The only note is a nit: two commit subjects run 52 and 51 characters against a target of about 50, under the repository's 72 limit. It was not acted on.

Both reviewers' RAN: lists are in their reports; no findings beyond that nit remain open, and no code, commits, title or body changed in this phase.

CI: every check on the head passed (Build, Lint, Unit Tests, Unit Tests (root), E2E Tests, license/cla, all seven image publish jobs and the kustomize bundle publish). Earlier failures were transient (module proxy stream errors, a Lint run cancelled during apt-get) and passed on re-run.

Auto-merge is enabled with the merge-commit method, the only method the ruleset allows. The ruleset requires one approving code-owner review on the latest push, so the PR waits for a human approval.

@privateip
privateip merged commit 7292b26 into main Oct 7, 2026
24 of 27 checks passed
@privateip
privateip deleted the fix/issue-773 branch October 7, 2026 18:54
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Deleted sidecar pods leave rows in the shared eBPF maps

2 participants