Skip to content

[BUG]: Successful reboot leaves nodes erroring and rollout stopped; batch reset alone does not recover (v0.18.0) #633

Description

@crookedstorm

Skyhook Version

Operator v0.18.0; agent v6.4.2 in the original reproduction.

Installation method

Helm (managed through Flux).

Kubernetes Version

Kubelet v1.34.2 in the captured affected-node logs. The API-server version was not captured.

Component

Operator (controller-manager), Agent (package executor)

Describe the bug

A native reboot interrupt successfully reboots the host, but its container exits 143 during shutdown. After the node returns Ready, the rollout remains stopped and the node remains cordoned. We observed package state stage=interrupt, state=complete alongside stale node-level erroring status, with post-interrupt work not progressing.

Recovery requires updating the affected node's status annotation and label to waiting, then resetting the NodeWright's compartment batch state. Resetting only the NodeWright batch state is insufficient. I reconfirmed this on v0.18.0.

The tested recovery changes both the annotation and label. I have not isolated whether either metadata change alone is sufficient; please do not reduce the documented recovery to the status patch alone.

Expected behavior

After NodeWright verifies the requested reboot completed, it should reconcile node status with package progress, run post-interrupt checks, and release its cordon when the lifecycle completes. Shutdown-related container termination should not permanently strand a successfully recovered node.

Real reboot or post-interrupt failures must still respect the deployment policy. This report does not request treating every exit 143 as success or disabling the failure threshold.

If an explicit reset is required by policy, the supported recovery operation should reconcile node metadata and batch state together rather than requiring manual edits to multiple objects.

Steps to reproduce

Run a package with native interrupt: {type: reboot} on the environment below, using fixed one-node batches, a 100% batch success threshold, and failureThreshold: 1. The observed shutdown ordering is relevant; this is an observed reproduction, not a claim that every reboot triggers it.

  1. NodeWright config completes and the native reboot interrupt starts.
  2. The host begins rebooting. During shutdown, systemd stops the interrupt container's CRI-O scope, and that container exits 143.
  3. The node successfully reboots and returns Ready.
  4. The package reaches interrupt/complete, but node-level status remains erroring and the deployment batch is stopped.
  5. Post-interrupt checks do not progress and the node remains cordoned.
  6. Patching the NodeWright batch state alone does not recover the rollout.
  7. Setting the node status annotation and label to waiting, followed by the batch-state patch below, permits recovery.

Relevant log output

The following is a timestamped summary of the original host journal, on September 10, 2026, showed this ordering:

22:16:32.822 UTC  logind announces reboot
22:16:32.853 UTC  systemd begins stopping the interrupt container's scope
22:16:32.871 UTC  conmon reports that container exited 143
22:16:33.193 UTC  kubelet observes exit 143
22:16:33.204 UTC  CRI-O receives a request to recreate the container during shutdown
22:16:44.832 UTC  systemd begins stopping kubelet
22:17:02.972 UTC  systemd begins stopping CRI-O

This supports shutdown-related termination of the initiating container. It does not establish the complete controller-side cause of the stale status and stopped batch.

Environment details

  • NodeWright operator: v0.18.0
  • Agent in the original reproduction: v6.4.2
  • Runtime: CRI-O
  • Infrastructure: OCI bare-metal GPU nodes; observed on A100 and H200 nodes
  • Example NodeWright/package: rdma-netns-exclusive, using native interrupt: {type: reboot} at the time of reproduction
  • Deployment policy: fixed one-node batches, 100% batch success threshold, failureThreshold: 1

An OKE systemd ordering cycle involving multi-user.target, oci-oke-node-client.service, kubelet.service, and kubelet-monitor.service was also observed. We have not established that this cycle caused the shutdown race or the NodeWright stall.

Additional context

Confirmed recovery sequence

Only use this for the observed recovered hosts whose package state is already interrupt/complete. Replace NODE_NAME with an affected node. Run the annotation and label commands for every affected node before resetting the batch.

kubectl annotate node NODE_NAME \
  nodewright.nvidia.com/status_rdma-netns-exclusive=waiting --overwrite

kubectl label node NODE_NAME \
  nodewright.nvidia.com/status_rdma-netns-exclusive=waiting --overwrite

kubectl patch nodewright rdma-netns-exclusive --subresource=status --type=json -p '[{"op":"replace","path":"/status/compartmentStatuses/__default__/batchState","value":{"currentBatch":1,"consecutiveFailures":0,"completedNodes":0,"failedNodes":0,"shouldStop":false,"lastBatchSize":0,"lastBatchFailed":false}}]'

These commands preserve package progress; they do not erase the nodeState annotation or force another reboot. For other NodeWrights or compartments, substitute the corresponding status key, resource name, and compartment path.

Source observations in operator/v0.18.0

These paths provide a plausible mechanism for stale node error status to defeat a batch-only reset. The exact reconcile ordering still needs a controller regression test or trace.

Relationship to existing issues

Please cross-reference these issues rather than assuming either existing fix covers this reproduction.

Regression coverage requested

Exercise native reboot with the above policy, including termination of the interrupt container during shutdown, successful host return, and interrupt/complete package state. Verify post-interrupt completion and uncordoning without manual metadata edits or an extra reboot. Also test recovery from the existing stale erroring plus stopped-batch state, and verify that actual failures still stop the rollout.

We have used interrupt: noop with guarded package-managed reboots as a workaround, but would prefer reliable native reboot handling.

Code of Conduct

  • I agree to follow Skyhook's Code of Conduct
  • I have searched the open bugs and have found no duplicates for this bug report

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

component/agentSkyhook agent (package executor)component/operatorSkyhook operator (controller-manager)

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions