Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
129 changes: 129 additions & 0 deletions .github/actions/dump-operator-agent-diagnostics/action.yml
Original file line number Diff line number Diff line change
@@ -0,0 +1,129 @@
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0
#
#
# Licensed under the Apache License, Version 2.0 (the "License");
# you may not use this file except in compliance with the License.
# You may obtain a copy of the License at
#
# http://www.apache.org/licenses/LICENSE-2.0
#
# Unless required by applicable law or agreed to in writing, software
# distributed under the License is distributed on an "AS IS" BASIS,
# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
# See the License for the specific language governing permissions and
# limitations under the License.

name: Dump operator-agent diagnostics
description: >
Dump agent pod logs and the agent's on-node state after an operator-agent
chainsaw failure. Shared by the Python and Go suites so a parity failure is
diagnosed from the same evidence on both sides.

inputs:
node:
description: Node whose agent state is dumped, via the debugger pod setup.sh creates.
required: false
default: kind-worker

runs:
using: composite
steps:
# Every command is `|| true`-guarded and the step never fails: this runs
# under `if: failure()`, so a missing path here would replace the real
# chainsaw failure with a confusing one from the diagnostics themselves.
- name: Dump cluster state
shell: bash
env:
NODE: ${{ inputs.node }}
run: |
echo "::group::Nodes"
kubectl get nodes -o wide || true
echo "::endgroup::"

echo "::group::Pods (all namespaces)"
kubectl get pods -A -o wide || true
echo "::endgroup::"

echo "::group::NodeWright custom resources"
kubectl get nodewrights.nodewright.nvidia.com -o yaml || true
echo "::endgroup::"

echo "::group::Node annotations (nodeState lives here, not on the CR)"
kubectl get node "$NODE" -o jsonpath='{.metadata.annotations}' | tr ',' '\n' || true
echo "::endgroup::"

echo "::group::Recent events"
kubectl get events -A --sort-by=.lastTimestamp | tail -100 || true
echo "::endgroup::"

- name: Dump agent pod logs
shell: bash
run: |
# Package pods are created per node/package; dump every container of each,
# including terminated ones, since the failing stage has usually exited.
#
# Cluster-infrastructure namespaces are skipped. They contribute thousands
# of unrelated reflector errors on a kind cluster, and a dump nobody can
# read is the same as no dump: on the first real failure this action
# caught, the chainsaw error was buried under kube-system scheduler noise.
for ns in $(kubectl get ns -o jsonpath='{.items[*].metadata.name}' || true); do
case "$ns" in
kube-system|kube-public|kube-node-lease|local-path-storage) continue ;;
esac
for pod in $(kubectl get pods -n "$ns" -o jsonpath='{.items[*].metadata.name}' 2>/dev/null || true); do
case "$pod" in
*-debugger) continue ;;
esac
echo "::group::logs ${ns}/${pod}"
kubectl logs -n "$ns" "$pod" --all-containers --prefix --timestamps --tail=-1 || true
echo "--- previous ---"
kubectl logs -n "$ns" "$pod" --all-containers --prefix --timestamps --previous --tail=-1 || true
echo "::endgroup::"
done
done

- name: Dump agent on-node state
shell: bash
env:
NODE: ${{ inputs.node }}
run: |
# setup.sh leaves a privileged debugger pod on the node with the host
# root bind-mounted at /host, which is the only way to read the agent's
# flag, history and log files from inside the job.
#
# /var/lib/skyhook, not the agent's standalone /etc/skyhook default: under
# the operator SKYHOOK_ROOT_DIR is <CopyDirRoot>/<skyhook-name>
# (skyhook_controller.go getAgentConfigEnvVars, CopyDirRoot default
# /var/lib/skyhook), and that is where flags/, check_results,
# *_ALL_CHECKED, history/ and interrupts/flags/ actually land.
DEBUGGER="${NODE}-debugger"
if ! kubectl get pod -n default "$DEBUGGER" >/dev/null 2>&1; then
echo "No ${DEBUGGER} pod; skipping on-node dump."
exit 0
fi

for dir in /host/var/lib/skyhook /host/var/log/skyhook; do
echo "::group::tree ${dir#/host}"
kubectl exec -n default "$DEBUGGER" -- ls -laR "$dir" || true
echo "::endgroup::"
done

echo "::group::file contents under /var/lib/skyhook (flags + history)"
kubectl exec -n default "$DEBUGGER" -- \
find /host/var/lib/skyhook -type f -exec sh -c 'echo "===== $1 ====="; cat "$1"' _ {} \; || true
echo "::endgroup::"

echo "::group::agent log files under /var/log/skyhook"
kubectl exec -n default "$DEBUGGER" -- \
find /host/var/log/skyhook -type f -exec sh -c 'echo "===== $1 ====="; cat "$1"' _ {} \; || true
echo "::endgroup::"

- name: Dump operator logs
shell: bash
run: |
# The suite runs the operator via `make run` as a background process on
# the runner, not as a pod, so its log is a file rather than kubectl logs.
echo "::group::operator manager stdout"
cat operator/reporting/int/std.out || true
echo "::endgroup::"
4 changes: 4 additions & 0 deletions .github/workflows/agent-ci.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -392,6 +392,10 @@ jobs:
make build-cli
make setup-kind-cluster operator-agent-tests

- name: Dump diagnostics on failure
if: failure()
uses: ./.github/actions/dump-operator-agent-diagnostics

# See operator-ci.yaml's ci-gate for rationale. operator-ci, agent-ci,
# and lint-ci all publish a check named `ci-gate` — GitHub composes
# same-named required checks, so every ci-gate that posts must pass.
Expand Down
Loading
Loading