diff --git a/.claude/skills/nodewright-managing-openvex/SKILL.md b/.claude/skills/nodewright-managing-openvex/SKILL.md new file mode 100644 index 000000000..e0dea3e14 --- /dev/null +++ b/.claude/skills/nodewright-managing-openvex/SKILL.md @@ -0,0 +1,332 @@ +--- +name: nodewright-managing-openvex +description: | + Triage findings from the Image Vulnerability Scan workflow and maintain + `.openvex.json`, the repository's OpenVEX v0.2.0 document. Use when asked to + suppress a CVE, mark a finding not affected, write or edit a VEX statement, + review Grype code scanning alerts on the operator or agent image, or work out + why a suppression did not take effect. Suppression happens in `.openvex.json` + and never by dismissing a code scanning alert, because #607 will publish this + document as signed evidence attested onto every released image. +user-invocable: true +version: 0.1.0 +--- + +# Managing `.openvex.json` + +Operating instructions for triaging container image vulnerability findings and +maintaining the repository's OpenVEX document. + +## Files this skill owns + +| Path | Role | +| --- | --- | +| `.openvex.json` | The OpenVEX v0.2.0 document. The **only** suppression mechanism. Currently `statements: []`. | +| `.grype.yaml` | Scan config: path excludes, plus an `ignore:` list that is **not** for suppression (see below). | +| `.github/workflows/vuln-scan-images.yaml` | Weekly Thursday 09:00 UTC scan of the published `:latest` operator and agent images; uploads SARIF under categories `grype-operator` and `grype-agent`. Never fails a build. | +| `.claude/skills/nodewright-managing-openvex/vex-state.yaml` | Working notes and open verifications for this skill. | + +## Step 1 (always first): read `vex-state.yaml` + +Read `.claude/skills/nodewright-managing-openvex/vex-state.yaml` before doing +anything else. It records checks that could not finish in an earlier session, +negative results nobody should repeat, and decisions not to act. Doing the +triage without it means redoing work that was already done, or silently +reversing a decision someone made deliberately. + +## Step 2: suppress in `.openvex.json`, never in the Security tab + +Do **not** dismiss a code scanning alert to make a finding go away. + +Grype applies `.openvex.json` at scan time through the workflow's `vex:` input, +so a suppressed finding never becomes an alert in the first place. A UI +dismissal is invisible to the document, so the two drift apart. That matters +beyond tidiness: #607 will publish `.openvex.json` as *signed* evidence attested +onto every released image. That has not shipped yet, so nothing consumers can +fetch carries this document today. Once it does, a suppression living only in +the GitHub UI makes the signed artifact the wrong one, and the signed artifact +is the one consumers read. + +`.grype.yaml`'s `ignore:` list is not the alternative. It is reserved for the +opposite case: a finding that **is** reachable in our code and has no fixed +version upstream yet, which is accepted risk rather than "not affected". Its own +header states the four things every entry must record. If you are reaching for +it, reread that header first. + +## Step 3: check `main` before writing any statement + +The workflow scans the moving `:latest` tag on purpose, because the subject of +the scan is the artifact users actually pull. That artifact lags `main`. + +So a finding can mean either of two very different things: + +1. **Unfixed.** The dependency is still vulnerable on `main`. Fix it, or write a + VEX statement if we are genuinely not affected. +2. **Already fixed on `main`, not yet released.** Nothing is wrong with the code. + The remedy is cutting a release. + +Issue #629 is exactly case 2: the released operator image reports 2 HIGH that +`operator/go.mod` already fixed. Writing a VEX statement for that would be +false. The document would assert we are not affected when we simply have not +shipped the fix. + +Check before writing: + +Which file answers the question depends on where the package came from, and for +these images most findings are **not** direct dependencies. Check the artifact +type in the scan output first (`.matches[].artifact.type`) and follow it: + +```bash +# Go modules (artifact.type == "go-module"): operator, and the Go agent. +grep -n '' operator/go.mod agent/go/go.mod +git log --oneline -5 -- operator/go.mod agent/go/go.mod + +# Python packages (artifact.type == "python") in the agent venv. +grep -rn '' agent/skyhook-agent/pyproject.toml agent/vendor/ + +# deb packages and the CPython binary (artifact.type == "deb" or "binary") come +# from the base image, not from our source. Nothing in this repo pins their +# versions: `scripts/latest-distroless.sh` resolves the newest base at build +# time, so "is it fixed on main?" means "does a newer base carry the fix?". +grep -n 'DISTROLESS_VERSION\|FROM nvcr.io' containers/agent.Dockerfile containers/operator.Dockerfile +``` + +For a base-image finding, compare the base directly rather than guessing. Scan +the base tag the released image was built on (its +`org.opencontainers.image.base.name` label) against the newest available tag. If +the newer base clears it, the remedy is a rebuild and release, not a statement. +That is #628. + +If `main` already carries the fix, stop. Record the finding on the release +issue, not in `.openvex.json`. + +## Step 4: get the vulnerability ID right + +`vulnerability.name` in a statement must equal grype's **primary** ID, the +`.matches[].vulnerability.id` field. For ecosystem advisories (Go, PyPI, npm) +that is usually the **GHSA**, with the CVE appearing only under +`relatedVulnerabilities`. + +OpenVEX matches by exact string, across `@id`, `name`, and `aliases`. A statement +naming only a CVE will not match a finding whose primary ID is a GHSA: it will +look correct and suppress nothing. Read the alias warning below before reaching +for `aliases` to fix that. + +Pull the primary ID and its aliases out of a scan: + +```bash +GRYPE_DB_VALIDATE_AGE=false GRYPE_CHECK_FOR_APP_UPDATE=false \ + grype ghcr.io/nvidia/nodewright/operator:latest -o json \ + | jq -r '.matches[] + | select(.vulnerability.severity | ascii_downcase | IN("high","critical")) + | "\(.vulnerability.id)\t\(.vulnerability.severity)\t\(.artifact.name)@\(.artifact.version)\taliases=\([.relatedVulnerabilities[]?.id] | join(","))"' +``` + +The first column is what goes in `vulnerability.name`. + +**`aliases` is also a match key, so do not treat it as searchable decoration.** +go-vex matches a statement against a finding by checking `@id`, `name`, and +every entry in `aliases`. Adding the CVE alongside a GHSA therefore widens the +statement to any *other* finding in the image whose primary ID is that CVE, in +a package you never reasoned about. That is live here: in the agent image most +HIGH+ findings are Debian packages whose primary ID is a CVE, while the Go and +PyPI ones are GHSA-primary with CVE aliases. List an alias only when you mean +the statement to cover it. + +Narrow the blast radius with `subcomponents` when a statement is about one +package rather than the whole image. Grype passes the matched package's purl as +the subcomponent identifier, so a product with no `subcomponents` matches the +entire image for that vulnerability ID. + +`GRYPE_DB_VALIDATE_AGE=false GRYPE_CHECK_FOR_APP_UPDATE=false` is a **local-only +workaround**: this sandbox blocks grype's database and update hosts, so without +it grype refuses a stale local DB. It must never appear in a workflow. CI has +network access and a stale database there would silently under-report. + +## Step 5: get the product PURL right + +Grype derives the OCI product PURL from the **registry repository basename**, +not from the `org.opencontainers.image.title` label. For these images that makes +the expected values: + +- `pkg:oci/operator` +- `pkg:oci/agent` + +`pkg:oci/operator` is **verified**: a test statement carrying it dropped the +image's HIGH+ count from 2 to 1. `pkg:oci/agent` is inferred rather than +measured, but follows by the same code path: grype derives +`pkg:oci/@sha256:...?repository_url=...` from the image's +`RepoDigests`, and a version-less, qualifier-less purl matches it. + +A mismatched PURL silently suppresses nothing while looking perfectly correct in +review, which is the worst failure mode this document has. + +Carry the PURL in **both** `@id` and `identifiers.purl` on every product, with +the same value. OpenVEX v0.2.0 has no `products[].purl` field: a product is a +Component carrying `@id`, `identifiers`, and `hashes`, so the purl lives at +`identifiers.purl`. Grype will in practice also match a `@id` that begins with +`pkg:`, so writing both is belt and braces rather than strictly required, but it +is what the sibling project's working document does on every statement and it +costs nothing. + +**Every statement must be verified empirically before it is merged, not just +the first one.** This is not a one-time bootstrap. A mistyped `status` is the +clearest reason why: `"not-affected"` with a hyphen instead of an underscore +makes grype exit 0 with no warning and suppress nothing. So do a wrong purl, a +CVE where the primary ID is a GHSA, and a typo'd package name. None of them +produce an error; all of them produce a document that reads correctly and does +nothing. + +The measurement is the only thing that distinguishes a working statement from a +decorative one. Do not verify by reading; verify by measuring the HIGH+ count +with and without `--vex` against the same image: + +```bash +IMG=ghcr.io/nvidia/nodewright/operator:latest +COUNT='[.matches[] | select(.vulnerability.severity | ascii_downcase | IN("high","critical"))] | length' + +before=$(GRYPE_DB_VALIDATE_AGE=false GRYPE_CHECK_FOR_APP_UPDATE=false \ + grype "$IMG" -o json | jq "$COUNT") +after=$(GRYPE_DB_VALIDATE_AGE=false GRYPE_CHECK_FOR_APP_UPDATE=false \ + grype "$IMG" --vex .openvex.json -o json | jq "$COUNT") + +echo "before=$before after=$after" +``` + +`after` must be lower than `before` by exactly the number of findings the new +statement covers. If the counts are identical, the statement matched nothing: +suspect the PURL first, then the vulnerability ID. Do not merge a statement that +has not moved the count. + +If the count does not move, the statement is decorative. Do not merge it. + +## Step 6: write the statement to the v0.2.0 contract + +`status` must be one of exactly: + +- `not_affected` +- `affected` +- `fixed` +- `under_investigation` + +Two hard requirements: + +- A `not_affected` statement **must** carry a `justification` **or** an + `impact_statement`. +- An `affected` statement **must** carry an `action_statement`. + +If a `justification` is present it must be one of exactly these five. There is +no free-text justification: + +- `component_not_present` +- `vulnerable_code_not_present` +- `vulnerable_code_not_in_execute_path` +- `vulnerable_code_cannot_be_controlled_by_adversary` +- `inline_mitigations_already_exist` + +If none of the five is honestly true, use an `impact_statement` and say why in +prose, or do not write the statement at all. Picking the nearest-sounding +justification is how a signed document ends up asserting something false. + +Shape: + +```json +{ + "vulnerability": { + "name": "GHSA-xxxx-xxxx-xxxx", + "aliases": ["CVE-2026-NNNNN"] + }, + "products": [ + { + "@id": "pkg:oci/operator", + "identifiers": { "purl": "pkg:oci/operator" } + } + ], + "status": "not_affected", + "justification": "vulnerable_code_not_in_execute_path", + "impact_statement": "One or two sentences naming the specific call path we do not take." +} +``` + +## Step 7: keep document-level fields as identifiers, not prose + +`@id`, `author`, `role`, `timestamp`, `version`, and `tooling` are metadata +fields. They identify the document. They are not a place to explain reasoning. + +`tooling` in particular. A sibling NVIDIA project let that field grow into an +8,010-byte changelog, which then shipped, signed, on seven release images before +anyone noticed (NVIDIA/aicr#2706). Ours reads +`manual curation (.claude/skills/nodewright-managing-openvex)` and should stay +that length. + +Working notes, open questions, and rejected options go in `vex-state.yaml`. +Reasoning a reader needs goes in the statement's `impact_statement`, which is +the field designed for it. Nothing narrative goes in a document-level field. + +On every change to the document: + +- Bump `version` (integer, monotonic). +- Update `timestamp` to the current UTC time in RFC 3339. +- Leave `@id`, `author`, `role`, and `tooling` alone. + +## Step 8: triage by severity, not by alert count + +The uploaded SARIF contains **every finding with an available fix**, not only +HIGH and above. `severity-cutoff: high` in the workflow neither filters the +report nor grades it: the SARIF level is a pure function of the finding's own +severity (critical and high become `error`, medium becomes `warning`, anything +lower becomes `note`). All `severity-cutoff` does is set grype's `--fail-on`, +which is inert while `fail-build` is false. It is the knob #630 will flip, and +until then changing it has no observable effect. + +Grype sets a `security-severity` (CVSS) property on each rule, so GitHub grades +alerts correctly and the Security tab can filter by severity. Use that filter. + +**Triage HIGH and CRITICAL.** Low and Medium findings do not need a statement. +Writing one for every Low alert is how the document stops being reviewable. + +Filter existing alerts by category and severity: + +```bash +gh api -X GET repos/NVIDIA/nodewright/code-scanning/alerts \ + -f state=open -f tool_name=grype --paginate \ + | jq -r '.[] + | select(.rule.security_severity_level | IN("high","critical")) + | "\(.most_recent_instance.category)\t\(.rule.security_severity_level)\t\(.rule.id)"' \ + | sort | uniq -c +``` + +## Step 9 (always last): update `vex-state.yaml` + +Before finishing, write back to +`.claude/skills/nodewright-managing-openvex/vex-state.yaml`: + +- **Delete** any `deferred_verifications` entry whose check has now been run + **and passed**. A failed check is not a resolution; it is the finding you are + now triaging. +- **Delete** any `known_negatives` or `deliberate_exclusions` entry whose + `invalidated_by` condition has fired. Do not rewrite it as history. +- **Add** an entry for anything this session could not finish, already ruled out, + or deliberately chose not to do. + +Every entry must state what would invalidate it. An entry that cannot state that +is not state; it is prose, and it belongs in a PR description or nowhere. The +file's header explains the lifecycle rules in full. + +## Quick reference: local scan + +```bash +GRYPE_DB_VALIDATE_AGE=false GRYPE_CHECK_FOR_APP_UPDATE=false \ + grype ghcr.io/nvidia/nodewright/agent:latest \ + --config .grype.yaml \ + --vex .openvex.json \ + --only-fixed \ + -o table +``` + +Run the workflow on demand rather than waiting for Thursday: + +```bash +gh workflow run "Image Vulnerability Scan" --repo NVIDIA/nodewright +gh run list --workflow "Image Vulnerability Scan" --repo NVIDIA/nodewright --limit 3 +``` diff --git a/.claude/skills/nodewright-managing-openvex/vex-state.yaml b/.claude/skills/nodewright-managing-openvex/vex-state.yaml new file mode 100644 index 000000000..9281710fd --- /dev/null +++ b/.claude/skills/nodewright-managing-openvex/vex-state.yaml @@ -0,0 +1,79 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Working state for the `nodewright-managing-openvex` skill. Read at the start +# of every invocation, updated at the end of every invocation. +# +# THE LIFECYCLE IS THE ENTIRE POINT OF THIS FILE. Entries are written so they +# can be deleted, and deleting them is the job, not an optional tidy-up. State +# that is never removed is exactly what turned a sibling project's document-level +# `tooling` field into 8,010 bytes of accumulated prose that then shipped, +# signed, on seven release images (NVIDIA/aicr#2706). This file exists so that +# never has to live in `.openvex.json`; it only works if entries leave. +# +# THE RULE: every entry must say what would invalidate it. An entry that cannot +# state that is not state. It is prose, and it belongs in a PR description or +# nowhere. +# +# Three kinds of entry, each with its own deletion rule: +# +# deferred_verifications +# A check that could not complete in the session that wrote it because it +# needs a later CI run or a later release. `resolves_when` must name the +# OUTCOME, never the event that merely makes the check runnable. "The +# release published" is not a result; "the HIGH+ count dropped by the +# number of findings the statement covers" is. +# DELETE once the check has been run AND passed. A failed check is not a +# resolution; it is the finding you are now triaging, and it belongs in an +# issue rather than here. +# +# known_negatives +# Work already done that came back negative, recorded so nobody repeats it. +# Carries `invalidated_by`: the upstream event that makes the negative worth +# rechecking. +# DELETE when `invalidated_by` fires. Do not rewrite it as history; the +# recheck's result is the new entry, if one is needed at all. +# +# deliberate_exclusions +# A decision not to act, recorded so the next maintainer does not silently +# reverse it. Carries `invalidated_by`: the condition under which the +# decision should be revisited. +# DELETE when that condition changes. The decision is then live again and +# gets made afresh, not inherited. + +deferred_verifications: + - id: confirm-first-scheduled-run + opened: 2026-09-16 + what: >- + `.github/workflows/vuln-scan-images.yaml` has never executed. Two things + cannot be confirmed by reading it and can only be confirmed by a real run: + that a job-level `security-events: write` is sufficient for a non-CodeQL + workflow to upload SARIF through `github/codeql-action/upload-sarif`, and + that the two upload categories (`grype-operator`, `grype-agent`) produce + two independent alert sets rather than one overwriting the other. + resolves_when: >- + A completed run of the workflow is followed by code scanning alerts + present under BOTH categories, confirmed via the code-scanning alerts API + grouped by `most_recent_instance.category`. A run that uploads without + error but yields alerts under only one category is a failure of the second + half, not a resolution. + blocks: >- + Nothing, but until it resolves, an empty Security tab means "unverified + pipeline" and not "no findings". Do not read absence of alerts as a + clean result. + +known_negatives: [] + +deliberate_exclusions: [] diff --git a/.github/workflows/vuln-scan-images.yaml b/.github/workflows/vuln-scan-images.yaml new file mode 100644 index 000000000..8c4cd1f0f --- /dev/null +++ b/.github/workflows/vuln-scan-images.yaml @@ -0,0 +1,168 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +# Scans the published container images for known vulnerabilities and publishes +# the results as code scanning alerts. +# +# This workflow never fails on a finding. Gating a release on the results is +# #630, deliberately deferred: at the time this landed the released agent image +# carried 58 HIGH+ findings (#628) and the operator image 2 (#629), so a gate +# would have blocked every release on day one. +# +# Suppression belongs in `.openvex.json` and nowhere else. Grype applies it at +# scan time via the `vex:` input below, so a suppressed finding never becomes an +# alert. Dismissing an alert in the Security tab instead would be invisible to +# that document, which #607 will publish as signed evidence on every released +# image. #607 has not shipped, so nothing attests this document yet; the rule is +# stated now so the habit is in place before the signed artifact exists. +# +# The images are scanned by their moving `latest` tag on purpose: the subject of +# this scan is the artifact users actually pull, which lags `main`. A finding +# here can therefore mean "already fixed on main, not yet released" (#629 is +# exactly that) rather than "unfixed". Check `main` before writing a VEX +# statement. +# +# `severity-cutoff` neither filters nor grades the report, which is easy to +# misread. Grype's SARIF level is a pure function of the finding's own severity +# (critical/high -> error, medium -> warning, lower -> note); `severity-cutoff` +# only sets `--fail-on`, which is inert while `fail-build` is false. It is kept +# because it is the knob #630 flips, not because it does anything today. +# +# So every finding with an available fix is uploaded, graded by the +# `security-severity` (CVSS) property grype sets on each rule. The Security tab +# carries Medium and Low alerts alongside the HIGH+ ones and can filter between +# them. #630 must therefore filter its gate by severity: keying a gate on "any +# open Grype alert" would block releases on Low findings. +# +# `only-fixed` is a deliberate blind spot: a HIGH or CRITICAL with no upstream +# patch is not reported here at all. That keeps the alert list actionable, and +# it means an empty Security tab is not the same as "no exposure". + +name: Image Vulnerability Scan + +on: + schedule: + # Thursday 09:00 UTC. Kept clear of CodeQL (Monday 05:00) and the daily + # 06:00/07:00 housekeeping workflows. + - cron: '0 9 * * 4' + workflow_dispatch: {} + # The scan itself never runs on a pull request, but its two input files are + # parsed by nothing else in CI. A dropped comma in `.openvex.json` would merge + # green and then hard-fail the scan the following Thursday, where the only + # signal is a red cron nobody watches. Validate the inputs on any PR that + # touches them. + pull_request: + paths: + - '.openvex.json' + - '.grype.yaml' + - '.github/workflows/vuln-scan-images.yaml' + +permissions: + contents: read + +concurrency: + group: ${{ github.workflow }}-${{ github.ref }} + cancel-in-progress: true + +jobs: + validate-inputs: + name: Validate scan inputs + runs-on: ubuntu-latest + timeout-minutes: 5 + steps: + - name: Checkout repository + # SHA-pinned rather than @v7: this job carries security-events: write, + # so a retargeted tag could alter code scanning results. codeql.yaml and + # renovate.yaml, the repo's other security-events workflows, pin for the + # same reason. + uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + + - name: Validate .openvex.json + run: | + set -euo pipefail + jq -e ' + (.["@context"] | type == "string") + and (.["@id"] | type == "string") + and (.author | type == "string") + and (.timestamp | type == "string") + and (.version | type == "number") + and (.statements | type == "array") + ' .openvex.json > /dev/null + echo "✅ .openvex.json is valid JSON and carries the required OpenVEX fields" + + - name: Validate .grype.yaml + run: | + set -euo pipefail + # ruby rather than python3: PyYAML is not guaranteed on the runner + # image, ruby ships YAML in its stdlib and is preinstalled. + ruby -ryaml -e 'YAML.load_file(".grype.yaml")' + echo "✅ .grype.yaml parses" + + scan: + name: Scan (${{ matrix.name }}) + runs-on: ubuntu-latest + needs: [validate-inputs] + # Never scan on a pull request: the job pulls published images and writes + # code scanning alerts, neither of which belongs in PR feedback. The + # repository guard keeps a fork with Actions enabled from scanning NVIDIA's + # images weekly and filing alerts about them in the fork's Security tab, + # matching lock-threads.yaml and inactive-pr-reminder.yaml. + if: github.event_name != 'pull_request' && github.repository == 'NVIDIA/nodewright' + timeout-minutes: 20 + permissions: + contents: read + security-events: write + strategy: + fail-fast: false + matrix: + include: + - name: operator + image: ghcr.io/nvidia/nodewright/operator + - name: agent + image: ghcr.io/nvidia/nodewright/agent + steps: + - name: Checkout repository + # SHA-pinned rather than @v7: this job carries security-events: write, + # so a retargeted tag could alter code scanning results. codeql.yaml and + # renovate.yaml, the repo's other security-events workflows, pin for the + # same reason. + uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1 + with: + persist-credentials: false + + - name: Scan image + id: scan + uses: anchore/scan-action@27805bf3b4e84b4a5c980df22ed233c00390a439 # v7.4.2 + with: + image: '${{ matrix.image }}:latest' + severity-cutoff: 'high' + only-fixed: true + # Never fail the run. See #630. + fail-build: false + output-format: sarif + config: .grype.yaml + vex: .openvex.json + + # One category per image. Code scanning keys an analysis on the category, + # so uploads sharing one can replace each other rather than accumulate; + # `vex-state.yaml` carries this as an unverified check until a real run + # shows alerts under both. + - name: Upload SARIF to code scanning + uses: github/codeql-action/upload-sarif@cdf488f595d80d6e07e03d4674febd5ab45fa938 # v4.37.9 + with: + sarif_file: ${{ steps.scan.outputs.sarif }} + category: grype-${{ matrix.name }} diff --git a/.grype.yaml b/.grype.yaml new file mode 100644 index 000000000..28a55eecb --- /dev/null +++ b/.grype.yaml @@ -0,0 +1,41 @@ +# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved. +# SPDX-License-Identifier: Apache-2.0 +# +# Licensed under the Apache License, Version 2.0 (the "License"); +# you may not use this file except in compliance with the License. +# You may obtain a copy of the License at +# +# http://www.apache.org/licenses/LICENSE-2.0 +# +# Unless required by applicable law or agreed to in writing, software +# distributed under the License is distributed on an "AS IS" BASIS, +# WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +# See the License for the specific language governing permissions and +# limitations under the License. + +exclude: + # Coding-agent worktrees are full checkouts at older commits. Their stale + # dependency manifests resurface already-fixed findings. Local-only: CI + # scans a fresh checkout, so this path never exists there. + - './.claude/worktrees/**' + - './operator/bin/**' + - './operator/vendor/**' + +# `ignore` is NOT the suppression mechanism for "this CVE does not apply to us". +# That is `.openvex.json`, which is structured, publishable, and (from #607) +# signed and attested onto every released image. +# +# This list is for the opposite case: a finding that IS reachable in our code +# and has no fixed version upstream yet, which is accepted risk rather than +# "not affected". Every entry MUST record the reachable path, the no-fix +# status, the target fixed version, and a review date, and MUST be deleted the +# moment a fixed version is available. An entry with no expiry condition is how +# a suppression becomes permanent. +# +# Note the tension with the scan workflow, which passes `only-fixed: true`: a +# finding with no available fix is never reported there, so an entry of this +# kind has nothing to suppress in CI today. It would still apply to a local +# scan run without `--only-fixed`. Do not add one speculatively. +# +# Currently empty. Keep it that way unless an entry can state all four. +ignore: [] diff --git a/.openvex.json b/.openvex.json new file mode 100644 index 000000000..32608f792 --- /dev/null +++ b/.openvex.json @@ -0,0 +1,10 @@ +{ + "@context": "https://openvex.dev/ns/v0.2.0", + "@id": "https://github.com/NVIDIA/nodewright/.openvex.json", + "author": "NVIDIA NodeWright maintainers", + "role": "document creator", + "timestamp": "2026-09-16T00:00:00Z", + "version": 1, + "tooling": "manual curation (.claude/skills/nodewright-managing-openvex)", + "statements": [] +} diff --git a/docs/contributing/release-process.md b/docs/contributing/release-process.md index 750739df5..f309e0485 100644 --- a/docs/contributing/release-process.md +++ b/docs/contributing/release-process.md @@ -482,6 +482,41 @@ Use the same command pattern for each released artifact: | GHCR agent image | `ghcr.io/nvidia/nodewright/agent@sha256:` | | GHCR Helm chart | `ghcr.io/nvidia/nodewright/charts/nodewright@sha256:` | +## Vulnerability Scanning + +The `Image Vulnerability Scan` workflow (`.github/workflows/vuln-scan-images.yaml`) runs [Grype](https://github.com/anchore/grype) against the published container images on a weekly cron, Thursday 09:00 UTC, plus `workflow_dispatch`. + +It scans two images, both by their moving `latest` tag: `ghcr.io/nvidia/nodewright/operator:latest` and `ghcr.io/nvidia/nodewright/agent:latest`. Not every released version, only whatever `latest` currently points at. + +Findings are uploaded as SARIF and show up as GitHub code scanning alerts in the repository Security tab, under the categories `grype-operator` and `grype-agent`. One category per image is required, not cosmetic: uploads sharing a category replace one another. Alerts dedupe across runs and close themselves once a later run stops reporting them. + +**The scan never fails a build.** No part of the release path is gated on a finding today; gating is deferred to #630. + +Run it on demand rather than waiting for Thursday: + +```bash +gh workflow run "Image Vulnerability Scan" --repo NVIDIA/nodewright +gh run list --workflow "Image Vulnerability Scan" --repo NVIDIA/nodewright --limit 3 +``` + +### The report is not limited to HIGH+ + +The uploaded SARIF contains every finding that has an available fix, not only HIGH and above. `severity-cutoff: high` neither filters the report nor grades it: the SARIF level is a pure function of the finding's own severity (critical and high become `error`, medium becomes `warning`, anything lower becomes `note`), and `severity-cutoff` only sets grype's `--fail-on`, which is inert because `fail-build` is false. It is kept because it is the knob #630 flips, not because it changes anything today. Grype sets a `security-severity` (CVSS) property on each rule, so GitHub grades the alerts correctly and the Security tab can filter by severity. Expect Medium and Low alerts alongside the HIGH+ ones, and filter by severity rather than reading a raw alert count. + +### Suppressing a finding + +Suppression goes in `.openvex.json`, the repository's OpenVEX document, and nowhere else. **Never dismiss a code scanning alert in the Security tab.** Grype applies the VEX document at scan time, so a correctly suppressed finding never becomes an alert at all, and #607 will publish that document as signed evidence attested onto every released image. #607 has not shipped yet, so nothing attests the document today; the rule is stated now so the habit precedes the signed artifact. Once it lands, a UI dismissal would leave the signed artifact inaccurate while looking resolved in the UI. + +The `ignore:` list in `.grype.yaml` is not a second suppression mechanism. It is reserved for the opposite case: a finding that is reachable in our code and has no upstream fix, accepted as risk. Read that file's header before adding an entry. + +Triage is driven by the [`nodewright-managing-openvex`](../../.claude/skills/nodewright-managing-openvex/SKILL.md) skill, which covers matching on the primary vulnerability ID, product PURLs, the OpenVEX v0.2.0 statement contract, and proving that a new statement actually suppresses something. + +### Known caveats + +- **A finding on `:latest` may already be fixed on `main`.** The scan deliberately targets the artifact users pull, which lags `main`. #629 is exactly that: the released operator image reports 2 HIGH that `operator/go.mod` already fixed. The remedy there is cutting a release, not writing a suppression. Check `main` before writing any VEX statement. +- **A release-candidate tag republishes `:latest`.** During an RC cycle `:latest` can point at a prerelease, so a scan in that window measures the RC instead of the shipped release. Tracked as #631. +- **Findings with no upstream fix are not reported at all.** The scan passes `only-fixed: true`, so a HIGH or CRITICAL with no patch available never becomes an alert. That keeps the alert list actionable, and it means an empty Security tab is not the same as no exposure. + ## Common Commands ```bash