Summary
k8s-tests/operator-agent/sigterm_grace/ asserts that progress holds exactly one start and one end after the package pod is deleted mid-step. That is a transient state, and the assertion passes in CI only because it is evaluated within a few seconds of the first end appearing. The settled state has two start/end pairs.
Why
Deleting the package pod makes the Job controller count that attempt as failed regardless of the agent's exit code (a pod deleted mid-init never reaches Succeeded), so it creates a replacement pod, which runs the apply stage again. The shellscript package's steps declare "idempotence": true, which is Idempotence.Disabled in both agents: the completion flag is ignored and the step runs regardless. So apply.sh runs twice by design. Observed on a real cluster (#685 release testing):
start 1790793529
end 1790793604 <- first attempt finished 75s later: gracefulShutdown held
start 1790793608 <- replacement pod, 4s after the first end
end 1790793683
The scenario's comments also claim "the retry pod finds apply.sh's completion flag, skips it", which is not what happens for a Disabled step.
What the scenario should assert
What gracefulShutdown guarantees is that the attempt which received SIGTERM finished: its start is followed by its own end about 60 seconds later. A killed attempt is the only thing that leaves a start without an end. So, after the package converges: every start has an end (equal counts), and the first end minus the first start is at least the step's sleep. Fix in the linked PR, together with the corrected comments.
Summary
k8s-tests/operator-agent/sigterm_grace/asserts thatprogressholds exactly onestartand oneendafter the package pod is deleted mid-step. That is a transient state, and the assertion passes in CI only because it is evaluated within a few seconds of the firstendappearing. The settled state has two start/end pairs.Why
Deleting the package pod makes the Job controller count that attempt as failed regardless of the agent's exit code (a pod deleted mid-init never reaches
Succeeded), so it creates a replacement pod, which runs the apply stage again. The shellscript package's steps declare"idempotence": true, which isIdempotence.Disabledin both agents: the completion flag is ignored and the step runs regardless. Soapply.shruns twice by design. Observed on a real cluster (#685 release testing):The scenario's comments also claim "the retry pod finds apply.sh's completion flag, skips it", which is not what happens for a
Disabledstep.What the scenario should assert
What
gracefulShutdownguarantees is that the attempt which received SIGTERM finished: itsstartis followed by its ownendabout 60 seconds later. A killed attempt is the only thing that leaves astartwithout anend. So, after the package converges: everystarthas anend(equal counts), and the firstendminus the firststartis at least the step's sleep. Fix in the linked PR, together with the corrected comments.