Skyhook Version
Operator v0.18.0; agent v6.4.2 in the original reproduction.
Installation method
Helm (managed through Flux).
Kubernetes Version
Kubelet v1.34.2 in the captured affected-node logs. The API-server version was not captured.
Component
Operator (controller-manager), Agent (package executor)
Describe the bug
A native reboot interrupt successfully reboots the host, but its container exits 143 during shutdown. After the node returns Ready, the rollout remains stopped and the node remains cordoned. We observed package state stage=interrupt, state=complete alongside stale node-level erroring status, with post-interrupt work not progressing.
Recovery requires updating the affected node's status annotation and label to waiting, then resetting the NodeWright's compartment batch state. Resetting only the NodeWright batch state is insufficient. I reconfirmed this on v0.18.0.
The tested recovery changes both the annotation and label. I have not isolated whether either metadata change alone is sufficient; please do not reduce the documented recovery to the status patch alone.
Expected behavior
After NodeWright verifies the requested reboot completed, it should reconcile node status with package progress, run post-interrupt checks, and release its cordon when the lifecycle completes. Shutdown-related container termination should not permanently strand a successfully recovered node.
Real reboot or post-interrupt failures must still respect the deployment policy. This report does not request treating every exit 143 as success or disabling the failure threshold.
If an explicit reset is required by policy, the supported recovery operation should reconcile node metadata and batch state together rather than requiring manual edits to multiple objects.
Steps to reproduce
Run a package with native interrupt: {type: reboot} on the environment below, using fixed one-node batches, a 100% batch success threshold, and failureThreshold: 1. The observed shutdown ordering is relevant; this is an observed reproduction, not a claim that every reboot triggers it.
- NodeWright config completes and the native reboot interrupt starts.
- The host begins rebooting. During shutdown, systemd stops the interrupt container's CRI-O scope, and that container exits 143.
- The node successfully reboots and returns Ready.
- The package reaches
interrupt/complete, but node-level status remains erroring and the deployment batch is stopped.
- Post-interrupt checks do not progress and the node remains cordoned.
- Patching the NodeWright batch state alone does not recover the rollout.
- Setting the node status annotation and label to
waiting, followed by the batch-state patch below, permits recovery.
Relevant log output
The following is a timestamped summary of the original host journal, on September 10, 2026, showed this ordering:
22:16:32.822 UTC logind announces reboot
22:16:32.853 UTC systemd begins stopping the interrupt container's scope
22:16:32.871 UTC conmon reports that container exited 143
22:16:33.193 UTC kubelet observes exit 143
22:16:33.204 UTC CRI-O receives a request to recreate the container during shutdown
22:16:44.832 UTC systemd begins stopping kubelet
22:17:02.972 UTC systemd begins stopping CRI-O
This supports shutdown-related termination of the initiating container. It does not establish the complete controller-side cause of the stale status and stopped batch.
Environment details
- NodeWright operator: v0.18.0
- Agent in the original reproduction: v6.4.2
- Runtime: CRI-O
- Infrastructure: OCI bare-metal GPU nodes; observed on A100 and H200 nodes
- Example NodeWright/package:
rdma-netns-exclusive, using native interrupt: {type: reboot} at the time of reproduction
- Deployment policy: fixed one-node batches, 100% batch success threshold,
failureThreshold: 1
An OKE systemd ordering cycle involving multi-user.target, oci-oke-node-client.service, kubelet.service, and kubelet-monitor.service was also observed. We have not established that this cycle caused the shutdown race or the NodeWright stall.
Additional context
Confirmed recovery sequence
Only use this for the observed recovered hosts whose package state is already interrupt/complete. Replace NODE_NAME with an affected node. Run the annotation and label commands for every affected node before resetting the batch.
kubectl annotate node NODE_NAME \
nodewright.nvidia.com/status_rdma-netns-exclusive=waiting --overwrite
kubectl label node NODE_NAME \
nodewright.nvidia.com/status_rdma-netns-exclusive=waiting --overwrite
kubectl patch nodewright rdma-netns-exclusive --subresource=status --type=json -p '[{"op":"replace","path":"/status/compartmentStatuses/__default__/batchState","value":{"currentBatch":1,"consecutiveFailures":0,"completedNodes":0,"failedNodes":0,"shouldStop":false,"lastBatchSize":0,"lastBatchFailed":false}}]'
These commands preserve package progress; they do not erase the nodeState annotation or force another reboot. For other NodeWrights or compartments, substitute the corresponding status key, resource name, and compartment path.
Source observations in operator/v0.18.0
These paths provide a plausible mechanism for stale node error status to defeat a batch-only reset. The exact reconcile ordering still needs a controller regression test or trace.
Relationship to existing issues
Please cross-reference these issues rather than assuming either existing fix covers this reproduction.
Regression coverage requested
Exercise native reboot with the above policy, including termination of the interrupt container during shutdown, successful host return, and interrupt/complete package state. Verify post-interrupt completion and uncordoning without manual metadata edits or an extra reboot. Also test recovery from the existing stale erroring plus stopped-batch state, and verify that actual failures still stop the rollout.
We have used interrupt: noop with guarded package-managed reboots as a workaround, but would prefer reliable native reboot handling.
Code of Conduct
Skyhook Version
Operator v0.18.0; agent v6.4.2 in the original reproduction.
Installation method
Helm (managed through Flux).
Kubernetes Version
Kubelet v1.34.2 in the captured affected-node logs. The API-server version was not captured.
Component
Operator (controller-manager), Agent (package executor)
Describe the bug
A native reboot interrupt successfully reboots the host, but its container exits 143 during shutdown. After the node returns Ready, the rollout remains stopped and the node remains cordoned. We observed package state
stage=interrupt, state=completealongside stale node-levelerroringstatus, with post-interrupt work not progressing.Recovery requires updating the affected node's status annotation and label to
waiting, then resetting the NodeWright's compartment batch state. Resetting only the NodeWright batch state is insufficient. I reconfirmed this on v0.18.0.The tested recovery changes both the annotation and label. I have not isolated whether either metadata change alone is sufficient; please do not reduce the documented recovery to the status patch alone.
Expected behavior
After NodeWright verifies the requested reboot completed, it should reconcile node status with package progress, run post-interrupt checks, and release its cordon when the lifecycle completes. Shutdown-related container termination should not permanently strand a successfully recovered node.
Real reboot or post-interrupt failures must still respect the deployment policy. This report does not request treating every exit 143 as success or disabling the failure threshold.
If an explicit reset is required by policy, the supported recovery operation should reconcile node metadata and batch state together rather than requiring manual edits to multiple objects.
Steps to reproduce
Run a package with native
interrupt: {type: reboot}on the environment below, using fixed one-node batches, a 100% batch success threshold, andfailureThreshold: 1. The observed shutdown ordering is relevant; this is an observed reproduction, not a claim that every reboot triggers it.interrupt/complete, but node-level status remainserroringand the deployment batch is stopped.waiting, followed by the batch-state patch below, permits recovery.Relevant log output
The following is a timestamped summary of the original host journal, on September 10, 2026, showed this ordering:
This supports shutdown-related termination of the initiating container. It does not establish the complete controller-side cause of the stale status and stopped batch.
Environment details
rdma-netns-exclusive, using nativeinterrupt: {type: reboot}at the time of reproductionfailureThreshold: 1An OKE systemd ordering cycle involving
multi-user.target,oci-oke-node-client.service,kubelet.service, andkubelet-monitor.servicewas also observed. We have not established that this cycle caused the shutdown race or the NodeWright stall.Additional context
Confirmed recovery sequence
Only use this for the observed recovered hosts whose package state is already
interrupt/complete. ReplaceNODE_NAMEwith an affected node. Run the annotation and label commands for every affected node before resetting the batch.kubectl annotate node NODE_NAME \ nodewright.nvidia.com/status_rdma-netns-exclusive=waiting --overwrite kubectl label node NODE_NAME \ nodewright.nvidia.com/status_rdma-netns-exclusive=waiting --overwrite kubectl patch nodewright rdma-netns-exclusive --subresource=status --type=json -p '[{"op":"replace","path":"/status/compartmentStatuses/__default__/batchState","value":{"currentBatch":1,"consecutiveFailures":0,"completedNodes":0,"failedNodes":0,"shouldStop":false,"lastBatchSize":0,"lastBatchFailed":false}}]'These commands preserve package progress; they do not erase the nodeState annotation or force another reboot. For other NodeWrights or compartments, substitute the corresponding status key, resource name, and compartment path.
Source observations in operator/v0.18.0
Status()reads the node status annotation;SetStatus()writes the annotation and label. The annotation is functional controller state, not merely display metadata.SetStatus()also gates those metadata writes on the annotation value, so it does not necessarily repair a differing label when the annotation already matches.EvaluateCurrentBatch()counts incomplete nodes with statuserroringas failures.GetNodesForNextBatch()returns no nodes when the strategy hasShouldStopset.These paths provide a plausible mechanism for stale node error status to defeat a batch-only reset. The exact reconcile ordering still needs a controller regression test or trace.
Relationship to existing issues
InProgress. It is related, but does not cover the shutdown-exit/stale-error-status recovery sequence reported here. Our original policy used one-node batches.Please cross-reference these issues rather than assuming either existing fix covers this reproduction.
Regression coverage requested
Exercise native reboot with the above policy, including termination of the interrupt container during shutdown, successful host return, and
interrupt/completepackage state. Verify post-interrupt completion and uncordoning without manual metadata edits or an extra reboot. Also test recovery from the existing staleerroringplus stopped-batch state, and verify that actual failures still stop the rollout.We have used
interrupt: noopwith guarded package-managed reboots as a workaround, but would prefer reliable native reboot handling.Code of Conduct