Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
47 changes: 47 additions & 0 deletions docs/architecture/interrupt-flow.md
Original file line number Diff line number Diff line change
Expand Up @@ -191,13 +191,60 @@ a pod that cannot finish terminating — a stuck finalizer, an unresponsive
kubelet — blocks the node's interrupt until it is cleared. Set
`spec.drainConfig.timeout` to bound that wait; there is no default.

### DrainBlocked Condition

A `DrainBlocked` condition is set on the NodeWright whenever one or more selected
nodes have a drain that cannot currently make progress. It names the blocked
nodes and, where the blocker is a PodDisruptionBudget, includes the apiserver's
own message verbatim:

```yaml
- type: DrainBlocked
status: "True"
reason: PodDisruptionBudget
message: "2/5 nodes blocked draining (node-a, node-c); default/web-0 on node-a:
The disruption budget web-pdb needs 3 healthy pods and has 3 currently"
```

`reason` is one of `PodDisruptionBudget`, `UnmanagedPod`, `EmptyDirData`, or
`MultipleCauses` when more than one kind of blocker is present across the
affected nodes. The condition clears automatically once every previously
blocked node has drained — no action is required beyond removing the
underlying blocker.

`DrainBlocked` is independent of the `Blocked` condition (which is reserved for
an uninstalled dependency): a NodeWright can be both dependency-blocked and
drain-blocked at the same time, so the two conditions never share a type.

A PodDisruptionBudget rejection is treated as a self-resolving wait state, not
a reconcile error: it no longer aborts the reconcile pass for the remaining
nodes. However, this also means it no longer puts the NodeWright into
controller-runtime's exponential backoff. With a PDB at zero allowed
disruptions and `spec.drainConfig.timeout` unset, the operator currently
retries the eviction every 2 seconds indefinitely, where it previously backed
off toward roughly 1000 seconds — a real increase in eviction-API load.
Per-node throttling for the drain-blocked case specifically is tracked in
[#632](https://github.com/NVIDIA/nodewright/issues/632) and not yet shipped.

Unmanaged pods (`force: false`) and `emptyDir` pods (`deleteEmptyDirData:
false`) are also reported in `DrainBlocked`, though — unlike a PDB rejection —
these were already wait states before this condition existed; `DrainBlocked`
only makes them visible without reading operator logs.

### Recovering From a Drain Timeout

When `spec.drainConfig.timeout` expires, the operator records a `DrainTimeout`
warning event, marks the node and NodeWright `erroring`, and leaves the node
cordoned. The operator stops issuing further evict/delete actions while the
blocking condition remains, so package stages do not proceed on that node.

`DrainTimeout` and `DrainBlocked` answer different questions: `DrainBlocked`
names what is currently preventing progress and clears itself once the
blocker is gone, with no configured limit on how long that may take.
`DrainTimeout` only fires once `spec.drainConfig.timeout` is set and elapses —
which, for a PDB blocker, is now the only path that turns a stuck drain into
an `erroring` node; the PDB rejection itself no longer does so directly.

To recover, remove the underlying blocker first, such as a PDB with zero allowed
disruptions, an unmanaged pod when `force: false`, or an `emptyDir` pod when
`deleteEmptyDirData: false`. Then reset the failed rollout metadata:
Expand Down
1 change: 1 addition & 0 deletions docs/architecture/operator-status.md
Original file line number Diff line number Diff line change
Expand Up @@ -90,6 +90,7 @@ The operator also sets additional condition types that may be useful for trouble
- `NodesIgnored`: selected nodes are skipped because they have the ignore label set
- `ApplyPackage`: the controller is applying a package to a node
- `DeploymentPolicyNotFound`: the referenced `DeploymentPolicy` is missing at reconcile time
- `DrainBlocked`: one or more selected nodes have a drain that cannot currently make progress. Reasons: `PodDisruptionBudget`, `UnmanagedPod`, `EmptyDirData`, or `MultipleCauses` when more than one kind of blocker is present. Independent of `Blocked` — a NodeWright can be both dependency-blocked and drain-blocked at once.

These conditions complement, rather than replace, `.status.status` and `Ready`.

Expand Down
77 changes: 77 additions & 0 deletions operator/internal/controller/cluster_state_v2.go
Original file line number Diff line number Diff line change
Expand Up @@ -19,6 +19,7 @@
package controller

import (
"context"
"fmt"
"sort"
"strings"
Expand Down Expand Up @@ -402,6 +403,7 @@ type SkyhookNodes interface {
IsPaused() bool
HasUninstallWork() (bool, error)
UpdateBlockedCondition() error
UpdateDrainBlockedCondition(ctx context.Context, logger logr.Logger)
UpdateUninstallConditions() error
UpdateNodeStateMalformedCondition()
NodeCount() int
Expand Down Expand Up @@ -626,6 +628,81 @@ func (s *skyhookNodes) UpdateBlockedCondition() error {
return nil
}

// UpdateDrainBlockedCondition rebuilds the DrainBlocked condition from each node's
// persisted drain-blocker annotation (SkyhookNode.DrainBlocked), rather than from a
// single pass's live findings. This makes the condition level-triggered, matching
// UpdateBlockedCondition: it is correct from persisted state alone regardless of
// whether this reconcile pass actually ran RunSkyhookPackages for this Skyhook (paused,
// disabled, complete Skyhooks skip it) or returned from it early (an error, or
// spec.serial stopping after the first node) — those nodes simply keep whatever was
// last recorded for them, rather than being wrongly treated as unblocked.
//
// A node whose annotation fails to parse is skipped for this computation, the same
// tolerance UpdateBlockedCondition applies to unreadable nodeState — this is a
// NodeStateMalformed-adjacent concern, but has no dedicated user-visible signal of
// its own; a node dropping out of this aggregate silently is an accepted limitation
// rather than a deliberate design, tracked for follow-up.
func (s *skyhookNodes) UpdateDrainBlockedCondition(ctx context.Context, logger logr.Logger) {
blocks := make([]wrapper.DrainBlockedNode, 0, len(s.nodes))
for _, node := range s.nodes {
// A node with no runnable, interrupt-requiring package this pass will never
// reach EnsureNodeIsReadyForInterrupt, so nothing else clears its persisted
// drain-blocker annotation. Clear it here instead: this function is the one
// thing that runs on every pass regardless of paused/disabled/complete state
// or an error elsewhere in the reconcile, which is what keeps a removed or
// finished drain from reporting a blocker that no longer exists.
if !nodeNeedsInterruptDrain(ctx, node) {
if err := node.SetDrainBlocked(nil); err != nil {
logger.Error(err, "clearing stale drain blocked state", "node", node.GetNode().Name)
}
}

blocked, err := node.DrainBlocked()
if err != nil {
continue
}
if len(blocked) == 0 {
continue
}
blocks = append(blocks, wrapper.DrainBlockedNode{
NodeName: node.GetNode().Name,
Blocked: blocked,
})
}

if len(blocks) == 0 {
wrapper.RemoveSkyhookConditionTypes(s.skyhook, wrapper.SkyhookConditionDrainBlocked)
return
}

// The message builder truncates detail lines past drainBlockedDetailLineLimit and
// points the reader at controller logs for the rest, mirroring
// updateTaintToleranceCondition's log-before-truncate pattern — so log the same
// persisted set the message is built from here, not a caller-local subset.
if totalBlockedPods := countBlockedPods(blocks); totalBlockedPods > wrapper.ReadyConditionNodeListLimit {
logger.Info("DrainBlocked condition message truncated; full blocked set", "nodewright", s.skyhook.Name, "drainBlocks", blocks)
}

wrapper.AddSkyhookCondition(s.skyhook, metav1.Condition{
Type: wrapper.SkyhookConditionDrainBlocked,
Status: metav1.ConditionTrue,
ObservedGeneration: s.skyhook.Generation,
LastTransitionTime: metav1.Now(),
Reason: wrapper.DrainBlockedConditionReason(blocks),
Message: wrapper.DrainBlockedConditionMessage(blocks, len(s.nodes)),
})
}

// countBlockedPods sums Blocked across every node, for deciding whether the DrainBlocked
// message will be truncated and the full set needs logging.
func countBlockedPods(blocks []wrapper.DrainBlockedNode) int {
total := 0
for _, b := range blocks {
total += len(b.Blocked)
}
return total
}

// isPackageCompleteOnAllNodes reports whether the package has reached its
// terminal-complete stage (per node.IsPackageComplete semantics) on every node
// this Skyhook selects. Returns false when there are no selected nodes: with
Expand Down
75 changes: 66 additions & 9 deletions operator/internal/controller/mock/SkyhookNodes.go

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

Loading
Loading