From f69764145d40e5cff761d2dbfa552497bac1fdc2 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 22:55:26 +0300 Subject: [PATCH 01/24] =?UTF-8?q?feat(convctl):=20CI-native=20output=20for?= =?UTF-8?q?mats=20=E2=80=94=20github,=20sarif,=20markdown?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit convctl emits table, json and junit. All three put a finding somewhere a person has to go looking for it, which is why a conversion failure in CI reads as a line in a log rather than as something attached to the diff the reviewer is already looking at. Three formats, all rendering one internal finding representation: github workflow commands on stdout, plus the markdown table appended to $GITHUB_STEP_SUMMARY when the runner sets it sarif SARIF 2.1.0, so findings reach code scanning and therefore the diff and the Security tab, triageable and suppressible like any other scanner's markdown a deterministic table, for a PR comment in any CI system The prerequisite was source locations, and that is most of the work here. sigs.k8s.io/yaml routes through encoding/json, which is what makes the strict typed decode possible and also what throws positions away. So the config is parsed a second time for positions alone, with yaml.Node, keyed by spoke and rule index. That parse decodes nothing and validates nothing: the typed loader stays the only thing that decides whether a config is valid, and a position parse that fails degrades to a report without line numbers rather than taking the run down with it. Three decisions worth recording. Finding ids are a compatibility surface. A suppression in code scanning is keyed on the id, so renaming one silently un-suppresses everything somebody dismissed. They are namespaced under convctl/, lower-kebab, and where the engine already has a stable code that code is reused rather than given a second name for the same thing. A finding that cannot be placed precisely is still reported — against the file with no line, or against the document. Dropping it would hide whole-config errors, which are the most serious kind, and SARIF is the format where that omission would be least visible. An acknowledged loss is reported at note severity rather than dropped. It is not a failure and never fails a build, but "this conversion drops this field, on purpose" is exactly what a reviewer of a config change wants to see. Uncovered fields are reported once, not twice: the engine already diagnoses them at the severity the unmapped-field policy dictates, and re-emitting them from FieldCoverage showed a reviewer every uncovered field a second time at a different severity. The report now carries the config's path as well as its name, because metadata.name is not something anyone can open. A fleet run aggregates per cluster, and a cluster that could not be reached is itself an error finding rather than an absence — a fleet check that quietly covered four of five clusters and reported green is the failure this guards against. diff gains markdown too, rendering its structured delta rather than a finding list, which is what the sticky-comment Action needs. go.yaml.in/yaml/v3 moves from indirect to direct. It is the same module sigs.k8s.io/yaml already pulls in, so nothing new enters the dependency tree. Closes #140 Co-Authored-By: Claude Opus 5 (1M context) --- docs/cli.md | 67 ++++- go.mod | 2 +- internal/cli/analyze.go | 6 +- internal/cli/ciformats.go | 327 ++++++++++++++++++++++++ internal/cli/ciformats_test.go | 451 +++++++++++++++++++++++++++++++++ internal/cli/configdiff.go | 82 ++++++ internal/cli/findings.go | 374 +++++++++++++++++++++++++++ internal/cli/fleet.go | 28 ++ internal/cli/report.go | 10 +- internal/cli/root.go | 65 +++-- internal/cli/sourcemap.go | 204 +++++++++++++++ internal/cli/test.go | 1 + internal/cli/validate.go | 8 + 13 files changed, 1598 insertions(+), 27 deletions(-) create mode 100644 internal/cli/ciformats.go create mode 100644 internal/cli/ciformats_test.go create mode 100644 internal/cli/findings.go create mode 100644 internal/cli/sourcemap.go diff --git a/docs/cli.md b/docs/cli.md index 366ffb6..edc0988 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -3,13 +3,13 @@ `convctl` runs the exact same `pkg/engine` code the operator and webhook server use, entirely offline against local YAML files — so you can validate and test a conversion mapping before it ever touches a cluster. Most commands work identically against an `XRDConversionConfig` (pass `--xrd`) or a `CRDConversionConfig` (pass `--crd`) — which one applies is determined by the config file's own `kind`, not by which flag you happen to type, so passing the wrong one is a clear error rather than a silent mismatch. `migrate-storage` is the exception: it is a live, mutating housekeeping command that takes cluster resource names (not files) and does not need a conversion config. ```console -convctl validate --config config.yaml [--xrd xrd.yaml | --crd crd.yaml] [-o table|json] -convctl analyze --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json] -convctl test --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) (--samples ./samples/ | --live) [flags] +convctl validate --config config.yaml [--xrd xrd.yaml | --crd crd.yaml] [-o table|json|github|sarif|markdown] +convctl analyze --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json|github|sarif|markdown] +convctl test --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) (--samples ./samples/ | --live) [-o table|json|junit|github|sarif|markdown] [flags] convctl plan --to v2 (--xrd xrd.yaml | --crd crd.yaml) [--config config.yaml] [-o table|json] convctl versions --xrd xrd.yaml [--config config.yaml] [--check-unserve v1] [-o table|json] convctl compat --base REV --head REV --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json|markdown] -convctl diff --config a.yaml --config b.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json] +convctl diff --config a.yaml --config b.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json|markdown] convctl diff --config config.yaml --live [-o json|table] convctl convert --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) --sample obj.yaml --to v2 [-o yaml|json] convctl suggest --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o yaml|json] @@ -328,6 +328,65 @@ convctl test --xrd xrd.yaml --config xrdconversionconfig.yaml --samples ./sample --output junit --output-file report.junit.xml ``` +### CI-native formats: `github`, `sarif`, `markdown` + +`table`, `json` and `junit` all put a finding somewhere a person has to go +looking for it. These three put it on the line of the config that produced it. + +| Format | Renders | Use it for | +|---|---|---| +| `github` | GitHub workflow commands (`::error file=…,line=…::…`) on stdout, plus a markdown table appended to `$GITHUB_STEP_SUMMARY` when the runner sets it | annotations on the pull-request diff | +| `sarif` | SARIF 2.1.0 | `github/codeql-action/upload-sarif` — findings land in code scanning, so they appear on the diff **and** in the Security tab, and can be triaged and suppressed like any other scanner's | +| `markdown` | a deterministic table | a PR comment in any CI system | + +Available on `test`, `validate` and `analyze`; `diff` has `markdown` (its +delta is a structured comparison, not a finding list). + +```console +$ convctl analyze --xrd xrd.yaml --config config.yaml -o github +::error file=config.yaml,line=13,col=7,title=Conversion config error::hub field "spec.size" is not covered by any rule and has no identical counterpart in the spoke schema +``` + +Locations come from a second, position-preserving parse of the config. +`sigs.k8s.io/yaml` routes through `encoding/json` — which is what makes the +strict typed decode possible and also what throws line numbers away — so the +formats read positions separately and never decide whether a config is valid. + +A finding the tool cannot place precisely is still reported, against the file +with no line, or against the config's document. Dropping it would hide +whole-config errors, which are the most serious kind. + +#### Finding ids + +The `ruleId` in SARIF and the finding name in the tables are a compatibility +surface: a suppression in code scanning is keyed on the id, so renaming one +silently un-suppresses everything somebody dismissed. Ids are added, never +renamed. + +| Id | Meaning | +|---|---| +| `convctl/unacknowledged-loss` | a round trip lost a field no rule declares lossy | +| `convctl/acknowledged-loss` | a declared, deliberate loss — reported at note severity, never a failure | +| `convctl/conversion-error` | a conversion failed outright | +| `convctl/schema-violation` | the converted object violates the destination schema (`--validate-output`) | +| `convctl/uncovered-field` | a schema field no rule claims | +| `convctl/rule-never-exercised` | a declared rule no sample reached | +| `convctl/golden-drift` | the committed corpus and the current output disagree | +| `convctl/config-error`, `convctl/config-warning` | a diagnostic with no more specific id | +| `convctl/required-field-*` | required-field analysis — the engine's own codes, lower-kebab | + +```yaml +- run: convctl test --xrd xrd.yaml --config config.yaml --samples ./samples/ -o sarif > convctl.sarif + continue-on-error: true +- uses: github/codeql-action/upload-sarif@v4 + with: + sarif_file: convctl.sarif +``` + +`continue-on-error` on the first step is deliberate: the upload should happen +whether or not the run failed, or a red build hides the findings explaining +why it is red. + ### Exit codes | Code | Meaning | diff --git a/go.mod b/go.mod index 80999d5..7cbc31a 100644 --- a/go.mod +++ b/go.mod @@ -12,6 +12,7 @@ require ( go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc v1.46.0 go.opentelemetry.io/otel/sdk v1.46.0 go.opentelemetry.io/otel/trace v1.46.0 + go.yaml.in/yaml/v3 v3.0.5 k8s.io/api v0.37.0 k8s.io/apiextensions-apiserver v0.37.0 k8s.io/apimachinery v0.37.0 @@ -70,7 +71,6 @@ require ( go.uber.org/multierr v1.11.0 // indirect go.uber.org/zap v1.27.1 // indirect go.yaml.in/yaml/v2 v2.4.4 // indirect - go.yaml.in/yaml/v3 v3.0.5 // indirect golang.org/x/exp v0.0.0-20260410095643-746e56fc9e2f // indirect golang.org/x/net v0.58.0 // indirect golang.org/x/oauth2 v0.36.0 // indirect diff --git a/internal/cli/analyze.go b/internal/cli/analyze.go index 9213a7d..db89073 100644 --- a/internal/cli/analyze.go +++ b/internal/cli/analyze.go @@ -41,6 +41,10 @@ type AnalyzeOutput struct { // not be determined. // +optional Scope *ScopeView `json:"scope,omitempty"` + // Analysis is the underlying report, kept so the CI output formats can + // attribute each diagnostic to the rule and line that produced it. Not + // serialized: the views above are the stable JSON contract. + Analysis *engine.AnalyzeReport `json:"-"` } // ScopeView reports a resolved Crossplane XRD scope and how much the @@ -121,7 +125,7 @@ func runAnalyzeCRDCmd(crdPath, configPath string) (*AnalyzeOutput, error) { } func buildAnalyzeOutput(resourceKind, resourceName, configName, hubVersion string, report engine.AnalyzeReport) *AnalyzeOutput { - out := &AnalyzeOutput{ResourceKind: resourceKind, Resource: resourceName, Config: configName, HubVersion: hubVersion, Lossless: report.OverallLossless()} + out := &AnalyzeOutput{ResourceKind: resourceKind, Resource: resourceName, Config: configName, HubVersion: hubVersion, Lossless: report.OverallLossless(), Analysis: &report} for _, sr := range report.SpokeReports { v := AnalyzeSpokeView{Version: sr.Version, LosslessHubToSpoke: sr.Lossless.HubToSpoke, LosslessSpokeToHub: sr.Lossless.SpokeToHub, RulesEvaluated: len(sr.RuleResults)} for _, d := range sr.Errors { diff --git a/internal/cli/ciformats.go b/internal/cli/ciformats.go new file mode 100644 index 0000000..d961c5f --- /dev/null +++ b/internal/cli/ciformats.go @@ -0,0 +1,327 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "encoding/json" + "fmt" + "io" + "os" + "sort" + "strings" +) + +// sarifVersion is the only SARIF version GitHub code scanning accepts. +const ( + sarifVersion = "2.1.0" + sarifSchema = "https://raw.githubusercontent.com/oasis-tcs/sarif-spec/main/sarif-2.1/schema/sarif-schema-2.1.0.json" +) + +// maxMarkdownRows bounds a PR comment. A thousand-row table is not review +// material, and GitHub truncates it anyway — better to say so and point at +// the full report than to be cut off mid-sentence by somebody else. +const maxMarkdownRows = 50 + +// WriteGitHubCommands renders findings as GitHub workflow commands, which is +// what puts a finding on the line of the file in the pull-request diff +// rather than in a log. +// +// Commands go to stdout (the runner reads them from the step's output); the +// human-readable table goes to $GITHUB_STEP_SUMMARY when that is set, which +// is the same information in the place a person looks. +func WriteGitHubCommands(w io.Writer, findings []Finding, title string) error { + for _, f := range findings { + level := f.Severity + if level == SeverityFindingNote { + level = "notice" // GitHub's spelling + } + var props []string + if f.Location.File != "" { + props = append(props, "file="+escapeCommandProperty(f.Location.File)) + } + if f.Location.Line > 0 { + props = append(props, fmt.Sprintf("line=%d", f.Location.Line)) + if f.Location.Column > 0 { + props = append(props, fmt.Sprintf("col=%d", f.Location.Column)) + } + } + props = append(props, "title="+escapeCommandProperty(f.Title())) + if _, err := fmt.Fprintf(w, "::%s %s::%s\n", level, strings.Join(props, ","), escapeCommandData(f.Message)); err != nil { + return err + } + } + return writeStepSummary(findings, title) +} + +// writeStepSummary appends the markdown table to $GITHUB_STEP_SUMMARY when +// the runner provides one. Absent the variable this is a no-op rather than +// an error: the same command has to work on a laptop. +func writeStepSummary(findings []Finding, title string) error { + path := os.Getenv("GITHUB_STEP_SUMMARY") + if path == "" { + return nil + } + // #nosec G304,G703 -- the path is the runner's own, read from the + // variable the runner itself set; there is no user input on this path, + // and writing somewhere else would defeat the feature. + f, err := os.OpenFile(path, os.O_APPEND|os.O_CREATE|os.O_WRONLY, 0o600) + if err != nil { + return fmt.Errorf("writing the job summary: %w", err) + } + defer func() { _ = f.Close() }() + WriteFindingsMarkdown(f, findings, title) + return nil +} + +// escapeCommandData escapes the message body of a workflow command. A +// literal newline would end the command and drop everything after it, which +// is how a multi-line diagnostic becomes a one-line one. +func escapeCommandData(s string) string { + r := strings.NewReplacer("%", "%25", "\r", "%0D", "\n", "%0A") + return r.Replace(s) +} + +// escapeCommandProperty escapes a property value, which additionally cannot +// contain the separators the command syntax uses. +func escapeCommandProperty(s string) string { + r := strings.NewReplacer("%", "%25", "\r", "%0D", "\n", "%0A", ":", "%3A", ",", "%2C") + return r.Replace(s) +} + +// WriteFindingsMarkdown renders findings as a markdown table for a PR +// comment or a job summary. +// +// Deterministic between runs on the same input — no timestamps, fixed +// ordering — so a sticky comment updates in place rather than producing a +// fresh diff every run. +func WriteFindingsMarkdown(w io.Writer, findings []Finding, title string) { + if title != "" { + _, _ = fmt.Fprintf(w, "### %s\n\n", title) + } + if len(findings) == 0 { + _, _ = fmt.Fprintln(w, "No findings.") + return + } + + counts := map[string]int{} + for _, f := range findings { + counts[f.Severity]++ + } + _, _ = fmt.Fprintf(w, "%d error(s), %d warning(s), %d note(s).\n\n", + counts[SeverityFindingError], counts[SeverityFindingWarning], counts[SeverityFindingNote]) + + _, _ = fmt.Fprintln(w, "| Severity | Finding | Location | Detail |") + _, _ = fmt.Fprintln(w, "|---|---|---|---|") + shown := findings + if len(shown) > maxMarkdownRows { + shown = shown[:maxMarkdownRows] + } + for _, f := range shown { + _, _ = fmt.Fprintf(w, "| %s | `%s` | %s | %s |\n", + f.Severity, f.RuleID, mdEscape(f.Location.String()), mdEscape(f.Message)) + } + if len(findings) > len(shown) { + _, _ = fmt.Fprintf(w, "\n_%d more finding(s) not shown — see the uploaded report for the full list._\n", + len(findings)-len(shown)) + } +} + +// mdEscape keeps a message containing a pipe from breaking the table it is +// rendered into. +func mdEscape(s string) string { + return strings.NewReplacer("|", "\\|", "\n", " ").Replace(s) +} + +// SARIF 2.1.0, in the subset GitHub code scanning reads. +type sarifLog struct { + Schema string `json:"$schema"` + Version string `json:"version"` + Runs []sarifRun `json:"runs"` +} + +type sarifRun struct { + Tool sarifTool `json:"tool"` + Results []sarifResult `json:"results"` +} + +type sarifTool struct { + Driver sarifDriver `json:"driver"` +} + +type sarifDriver struct { + Name string `json:"name"` + Version string `json:"version,omitempty"` + InformationURI string `json:"informationUri,omitempty"` + Rules []sarifRule `json:"rules"` +} + +type sarifRule struct { + ID string `json:"id"` + Name string `json:"name,omitempty"` + ShortDescription sarifText `json:"shortDescription"` + FullDescription *sarifText `json:"fullDescription,omitempty"` + HelpURI string `json:"helpUri,omitempty"` + Properties map[string]string `json:"properties,omitempty"` +} + +type sarifText struct { + Text string `json:"text"` +} + +type sarifResult struct { + RuleID string `json:"ruleId"` + Level string `json:"level"` + Message sarifText `json:"message"` + Locations []sarifLocation `json:"locations,omitempty"` +} + +type sarifLocation struct { + PhysicalLocation sarifPhysicalLocation `json:"physicalLocation"` +} + +type sarifPhysicalLocation struct { + ArtifactLocation sarifArtifactLocation `json:"artifactLocation"` + Region *sarifRegion `json:"region,omitempty"` +} + +type sarifArtifactLocation struct { + URI string `json:"uri"` +} + +type sarifRegion struct { + StartLine int `json:"startLine"` + StartColumn int `json:"startColumn,omitempty"` +} + +const findingsHelpURI = "https://terasky-oss.github.io/declarative-conversion-operator/cli/#finding-ids" + +// WriteSARIF renders findings as SARIF 2.1.0, so they land in GitHub code +// scanning — and therefore on the pull-request diff and in the Security tab, +// where they can be triaged and suppressed like any other scanner's. +func WriteSARIF(w io.Writer, findings []Finding, toolVersion string) error { + log := sarifLog{Schema: sarifSchema, Version: sarifVersion} + run := sarifRun{Tool: sarifTool{Driver: sarifDriver{ + Name: "convctl", + Version: toolVersion, + InformationURI: "https://github.com/TeraSky-OSS/declarative-conversion-operator", + }}} + + // One rule per id actually present, so the rules array describes this + // run rather than the tool's whole vocabulary. + seen := map[string]Finding{} + for _, f := range findings { + if _, ok := seen[f.RuleID]; !ok { + seen[f.RuleID] = f + } + } + ids := make([]string, 0, len(seen)) + for id := range seen { + ids = append(ids, id) + } + sort.Strings(ids) + for _, id := range ids { + run.Tool.Driver.Rules = append(run.Tool.Driver.Rules, sarifRule{ + ID: id, + Name: sarifRuleName(id), + ShortDescription: sarifText{Text: seen[id].Title()}, + HelpURI: findingsHelpURI, + }) + } + if run.Tool.Driver.Rules == nil { + run.Tool.Driver.Rules = []sarifRule{} + } + + run.Results = make([]sarifResult, 0, len(findings)) + for _, f := range findings { + res := sarifResult{ + RuleID: f.RuleID, + Level: sarifLevel(f.Severity), + Message: sarifText{Text: f.Message}, + } + // A finding with no line still gets a location: dropping it would + // hide whole-config errors, which are the most serious kind. + if f.Location.File != "" { + phys := sarifPhysicalLocation{ArtifactLocation: sarifArtifactLocation{URI: toSARIFURI(f.Location.File)}} + if f.Location.Line > 0 { + phys.Region = &sarifRegion{StartLine: f.Location.Line, StartColumn: f.Location.Column} + } + res.Locations = []sarifLocation{{PhysicalLocation: phys}} + } + run.Results = append(run.Results, res) + } + log.Runs = []sarifRun{run} + + enc := json.NewEncoder(w) + enc.SetIndent("", " ") + return enc.Encode(log) +} + +// sarifRuleName is the id without the tool prefix, which is what a code +// scanning alert shows as the rule's name. +func sarifRuleName(id string) string { + return strings.TrimPrefix(id, "convctl/") +} + +// sarifLevel maps to SARIF's own vocabulary, which spells "note" the same +// way but has no "notice". +func sarifLevel(severity string) string { + switch severity { + case SeverityFindingError: + return "error" + case SeverityFindingWarning: + return "warning" + } + return "note" +} + +// toSARIFURI normalises separators: SARIF artifact URIs are slash-separated +// regardless of the platform that produced them, and a Windows-style path +// makes an alert unmatchable against the repository's files. +func toSARIFURI(p string) string { + return strings.ReplaceAll(strings.TrimPrefix(p, "./"), "\\", "/") +} + +// writeFindings renders findings in whichever CI-native format was asked +// for, to the command's stdout. +func writeFindings(cmd interface{ OutOrStdout() io.Writer }, output string, findings []Finding, title string) error { + w := cmd.OutOrStdout() + switch output { + case "github": + return WriteGitHubCommands(w, findings, title) + case "sarif": + return WriteSARIF(w, findings, Version) + case "markdown": + WriteFindingsMarkdown(w, findings, title) + return nil + } + return fmt.Errorf("unsupported CI output format %q", output) +} + +// writeFindingsTo is writeFindings against an arbitrary writer, for the +// paths that buffer the report before deciding where it goes. +func writeFindingsTo(w io.Writer, output string, findings []Finding, title string) error { + switch output { + case "github": + return WriteGitHubCommands(w, findings, title) + case "sarif": + return WriteSARIF(w, findings, Version) + case "markdown": + WriteFindingsMarkdown(w, findings, title) + return nil + } + return fmt.Errorf("unsupported CI output format %q", output) +} diff --git a/internal/cli/ciformats_test.go b/internal/cli/ciformats_test.go new file mode 100644 index 0000000..0e67a7c --- /dev/null +++ b/internal/cli/ciformats_test.go @@ -0,0 +1,451 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "bytes" + "encoding/json" + "os" + "path/filepath" + "strings" + "testing" +) + +// The whole point of the source map is that a finding lands on the line a +// human edits. A rule's line, not the document's. +func TestSourceMapForConfig_LocatesSpokesAndRules(t *testing.T) { + m := SourceMapForConfig("testdata/config.yaml") + if m.File != "testdata/config.yaml" { + t.Fatalf("File = %q", m.File) + } + hub := m.HubVersion() + if !hub.Known() { + t.Error("hubVersion has no location") + } + spoke := m.Spoke("v1") + if !spoke.Known() { + t.Fatal("spoke v1 has no location") + } + r0, r1 := m.Rule("v1", 0), m.Rule("v1", 1) + if !r0.Known() || !r1.Known() { + t.Fatalf("rules have no locations: %v %v", r0, r1) + } + if r0.Line <= spoke.Line { + t.Errorf("rule 0 at line %d is not after the spoke at line %d", r0.Line, spoke.Line) + } + if r1.Line <= r0.Line { + t.Errorf("rule 1 at line %d is not after rule 0 at line %d", r1.Line, r0.Line) + } + + // The line has to be the rule the reader would recognise, so check the + // file actually says what we claim it does there. + data, err := os.ReadFile("testdata/config.yaml") + if err != nil { + t.Fatal(err) + } + lines := strings.Split(string(data), "\n") + if got := strings.TrimSpace(lines[r0.Line-1]); !strings.Contains(got, "strategy:") { + t.Errorf("rule 0 points at %q, want the line declaring the strategy", got) + } +} + +// An unknown spoke or rule index must still answer with somewhere real: a +// finding with no file at all cannot be rendered by any of the formats. +func TestSourceMapForConfig_FallsBackRatherThanReturningNothing(t *testing.T) { + m := SourceMapForConfig("testdata/config.yaml") + if got := m.Spoke("v99"); got.File == "" { + t.Error("an unknown spoke returned no location at all") + } + if got := m.Rule("v1", 99); !got.Known() { + t.Errorf("an out-of-range rule index returned %v, want the spoke's location", got) + } + if got := m.Rule("v99", 0); got.File == "" { + t.Error("an unknown spoke's rule returned no location at all") + } +} + +// A config that does not parse for positions must not take the report down +// with it: the typed loader has already accepted these bytes, and a report +// without line numbers beats no report. +func TestSourceMapForConfig_SurvivesAnUnparseableFile(t *testing.T) { + dir := t.TempDir() + bad := filepath.Join(dir, "bad.yaml") + if err := os.WriteFile(bad, []byte("spec: [unclosed\n"), 0o600); err != nil { + t.Fatal(err) + } + m := SourceMapForConfig(bad) + if m.Document().File != bad { + t.Errorf("Document() = %v, want the file with no line", m.Document()) + } + if got := SourceMapForConfig(filepath.Join(dir, "missing.yaml")); got == nil { + t.Error("a missing file returned a nil map, which would panic every caller") + } +} + +// A literal newline ends a workflow command, dropping everything after it — +// which is how a multi-line diagnostic silently becomes a one-line one. +func TestWriteGitHubCommands_EscapesData(t *testing.T) { + var buf bytes.Buffer + err := WriteGitHubCommands(&buf, []Finding{{ + RuleID: FindingUnacknowledgedLoss, Severity: SeverityFindingError, + Message: "first line\nsecond line 100% of the time", + Location: SourceLocation{File: "config.yaml", Line: 12, Column: 3}, + }}, "") + if err != nil { + t.Fatal(err) + } + got := strings.TrimRight(buf.String(), "\n") + if strings.Count(got, "\n") != 0 { + t.Fatalf("command spans multiple lines, so the runner would drop part of it:\n%s", got) + } + for _, want := range []string{"::error ", "file=config.yaml", "line=12", "col=3", "%0A", "%25"} { + if !strings.Contains(got, want) { + t.Errorf("command is missing %q:\n%s", want, got) + } + } +} + +// A property value containing a colon or comma would be read as the end of +// the property, silently corrupting the file name a finding points at. +func TestWriteGitHubCommands_EscapesProperties(t *testing.T) { + var buf bytes.Buffer + _ = WriteGitHubCommands(&buf, []Finding{{ + RuleID: FindingConfigError, Severity: SeverityFindingWarning, + Message: "m", + Location: SourceLocation{File: "weird,name:config.yaml", Line: 1}, + }}, "") + got := buf.String() + if strings.Contains(got, "weird,name:config.yaml") { + t.Errorf("separators in the file name were not escaped:\n%s", got) + } + if !strings.Contains(got, "%2C") || !strings.Contains(got, "%3A") { + t.Errorf("want the comma and colon percent-encoded:\n%s", got) + } +} + +// GitHub spells the lowest level "notice"; SARIF spells it "note". Getting +// this wrong makes the runner ignore the command entirely. +func TestWriteGitHubCommands_UsesNoticeForNotes(t *testing.T) { + var buf bytes.Buffer + _ = WriteGitHubCommands(&buf, []Finding{{RuleID: FindingAcknowledgedLoss, Severity: SeverityFindingNote, Message: "m"}}, "") + if !strings.HasPrefix(buf.String(), "::notice ") { + t.Errorf("want ::notice, got %q", buf.String()) + } +} + +// The job summary is where a person looks; the commands are for the runner. +// Both have to be produced, and the summary only when the runner asked. +func TestWriteGitHubCommands_WritesTheStepSummaryWhenAsked(t *testing.T) { + dir := t.TempDir() + summary := filepath.Join(dir, "summary.md") + t.Setenv("GITHUB_STEP_SUMMARY", summary) + + var buf bytes.Buffer + if err := WriteGitHubCommands(&buf, []Finding{{RuleID: FindingConfigError, Severity: SeverityFindingError, Message: "boom"}}, "convctl"); err != nil { + t.Fatal(err) + } + data, err := os.ReadFile(summary) + if err != nil { + t.Fatalf("no job summary written: %v", err) + } + if !strings.Contains(string(data), "boom") || !strings.Contains(string(data), "| Severity |") { + t.Errorf("summary is not the findings table:\n%s", data) + } +} + +// Without the variable — on a laptop — the same command has to work. +func TestWriteGitHubCommands_NoSummaryVariableIsNotAnError(t *testing.T) { + t.Setenv("GITHUB_STEP_SUMMARY", "") + var buf bytes.Buffer + if err := WriteGitHubCommands(&buf, []Finding{{RuleID: FindingConfigError, Severity: SeverityFindingError, Message: "m"}}, "t"); err != nil { + t.Errorf("running outside a runner failed: %v", err) + } +} + +func TestWriteSARIF_IsValidAndCarriesRulesAndLocations(t *testing.T) { + var buf bytes.Buffer + err := WriteSARIF(&buf, []Finding{ + {RuleID: FindingUnacknowledgedLoss, Severity: SeverityFindingError, Message: "lost a field", + Location: SourceLocation{File: "./config.yaml", Line: 12, Column: 5}}, + {RuleID: FindingConfigError, Severity: SeverityFindingError, Message: "whole-config problem", + Location: SourceLocation{File: "config.yaml"}}, + {RuleID: FindingAcknowledgedLoss, Severity: SeverityFindingNote, Message: "declared lossy"}, + }, "v1.2.3") + if err != nil { + t.Fatal(err) + } + + var log map[string]any + if err := json.Unmarshal(buf.Bytes(), &log); err != nil { + t.Fatalf("SARIF is not valid JSON: %v", err) + } + if log["version"] != "2.1.0" { + t.Errorf("version = %v, want 2.1.0 (the only one code scanning accepts)", log["version"]) + } + runs, _ := log["runs"].([]any) + if len(runs) != 1 { + t.Fatalf("runs = %d, want 1", len(runs)) + } + run, _ := runs[0].(map[string]any) + driver, _ := run["tool"].(map[string]any)["driver"].(map[string]any) + if driver["version"] != "v1.2.3" { + t.Errorf("driver version = %v, want the tool version", driver["version"]) + } + rules, _ := driver["rules"].([]any) + if len(rules) != 3 { + t.Errorf("rules = %d, want one per distinct finding id", len(rules)) + } + results, _ := run["results"].([]any) + if len(results) != 3 { + t.Fatalf("results = %d, want one per finding", len(results)) + } + + // A finding with no line still needs a location, and a whole-config + // error is the most serious kind to drop. + second, _ := results[1].(map[string]any) + locs, _ := second["locations"].([]any) + if len(locs) != 1 { + t.Errorf("a locationless-but-filed finding lost its location: %v", second) + } + // And one with no file at all must not invent one. + third, _ := results[2].(map[string]any) + if _, ok := third["locations"]; ok { + t.Errorf("a finding with no file was given a location: %v", third) + } +} + +// A code-scanning alert can only be matched to a file when the URI is +// slash-separated, whatever platform produced it. +func TestToSARIFURI_NormalisesSeparatorsAndPrefixes(t *testing.T) { + for in, want := range map[string]string{ + "./a/b.yaml": "a/b.yaml", + `dir\sub\c.yaml`: "dir/sub/c.yaml", + "plain.yaml": "plain.yaml", + } { + if got := toSARIFURI(in); got != want { + t.Errorf("toSARIFURI(%q) = %q, want %q", in, got, want) + } + } +} + +// Stable ids are a compatibility surface: renaming one silently +// un-suppresses every finding somebody dismissed in code scanning. +func TestFindingIDs_AreStableAndNamespaced(t *testing.T) { + for _, id := range []string{ + FindingUnacknowledgedLoss, FindingAcknowledgedLoss, FindingConversionError, + FindingSchemaViolation, FindingUncoveredField, FindingCoverageGap, + FindingGoldenDrift, FindingConfigError, FindingConfigWarning, + } { + if !strings.HasPrefix(id, "convctl/") { + t.Errorf("%q is not namespaced", id) + } + if strings.ToLower(id) != id { + t.Errorf("%q is not lower-kebab, so it will be inconsistent with engine-derived ids", id) + } + } + if got := kebab("RequiredFieldUnsatisfiable"); got != "required-field-unsatisfiable" { + t.Errorf("kebab() = %q", got) + } +} + +// A message containing a pipe would otherwise break the table it is in. +func TestWriteFindingsMarkdown_EscapesAndBounds(t *testing.T) { + var many []Finding + for i := 0; i < maxMarkdownRows+10; i++ { + many = append(many, Finding{RuleID: FindingConfigError, Severity: SeverityFindingError, Message: "a | b"}) + } + var buf bytes.Buffer + WriteFindingsMarkdown(&buf, many, "t") + got := buf.String() + if strings.Contains(got, "| a | b |") { + t.Error("an unescaped pipe broke the table") + } + if !strings.Contains(got, "more finding(s) not shown") { + t.Error("an unbounded table was rendered; GitHub truncates it mid-sentence") + } + if n := strings.Count(got, "convctl/config-error"); n != maxMarkdownRows { + t.Errorf("rendered %d rows, want the cap of %d", n, maxMarkdownRows) + } +} + +func TestWriteFindingsMarkdown_SaysSoWhenClean(t *testing.T) { + var buf bytes.Buffer + WriteFindingsMarkdown(&buf, nil, "convctl validate") + if !strings.Contains(buf.String(), "No findings.") { + t.Errorf("a clean run rendered %q", buf.String()) + } +} + +// Determinism is what lets a sticky PR comment update in place instead of +// producing a fresh diff on every run. +func TestSortFindings_IsDeterministicAndSeverityOrdered(t *testing.T) { + in := []Finding{ + {RuleID: "convctl/b", Severity: SeverityFindingNote, Message: "n", Location: SourceLocation{File: "b.yaml", Line: 2}}, + {RuleID: "convctl/a", Severity: SeverityFindingError, Message: "e", Location: SourceLocation{File: "b.yaml", Line: 9}}, + {RuleID: "convctl/c", Severity: SeverityFindingWarning, Message: "w", Location: SourceLocation{File: "a.yaml", Line: 1}}, + } + first := append([]Finding{}, in...) + sortFindings(first) + if first[0].Severity != SeverityFindingError || first[2].Severity != SeverityFindingNote { + t.Fatalf("not ordered by severity: %+v", first) + } + shuffled := []Finding{in[2], in[0], in[1]} + sortFindings(shuffled) + for i := range first { + if shuffled[i] != first[i] { + t.Fatalf("ordering depends on input order at %d: %+v vs %+v", i, shuffled[i], first[i]) + } + } +} + +// The coverage report identifies rules as "v1:rule[3]:FieldRename"; the +// source map is keyed on the spoke and the index. +func TestParseRuleID(t *testing.T) { + spoke, idx := parseRuleID("v1:rule[3]:FieldRename") + if spoke != "v1" || idx != 3 { + t.Errorf("parseRuleID = %q,%d want v1,3", spoke, idx) + } + if s, i := parseRuleID("nonsense"); s != "nonsense" || i != -1 { + t.Errorf("parseRuleID(nonsense) = %q,%d want nonsense,-1", s, i) + } +} + +// An unknown format has to be rejected rather than silently falling back to +// the table, which is how a job reports success from output nobody rendered. +func TestCheckOutputFormat_RejectsUnknown(t *testing.T) { + if err := checkOutputFormat("sarif", "table", "json", "sarif"); err != nil { + t.Errorf("a valid format was rejected: %v", err) + } + err := checkOutputFormat("saarif", "table", "json", "sarif") + if err == nil { + t.Fatal("a typo was accepted") + } + if !strings.Contains(err.Error(), "saarif") || !strings.Contains(err.Error(), "sarif") { + t.Errorf("error should name what was given and what is accepted: %v", err) + } +} + +// The formats exist to put a finding on the line that caused it, so the +// mapping from a real run has to produce real locations. +func TestFindingsFromAnalyze_AttributesToTheConfigLines(t *testing.T) { + out, err := RunAnalyze("../../examples/crossplane-xr-multiversion/02-add-v2/xrd.yaml", "", + "../../examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml") + if err != nil { + t.Fatalf("analyze: %v", err) + } + if out.Analysis == nil { + t.Fatal("analyze did not carry its report, so no CI format can attribute anything") + } + findings := findingsFromAnalyze(*out.Analysis, SourceMapForConfig("../../examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml")) + if len(findings) == 0 { + t.Fatal("a config with an uncovered field produced no findings") + } + for _, f := range findings { + if !f.Location.Known() { + t.Errorf("finding has no line: %+v", f) + } + } +} + +// The engine already reports an uncovered field; reporting it again from +// the coverage lists shows a reviewer the same field twice, at two +// severities, which reads as two problems. +func TestFindingsFromAnalyze_DoesNotReportUncoveredFieldsTwice(t *testing.T) { + out, err := RunAnalyze("../../examples/crossplane-xr-multiversion/02-add-v2/xrd.yaml", "", + "../../examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml") + if err != nil { + t.Fatal(err) + } + findings := findingsFromAnalyze(*out.Analysis, SourceMapForConfig("x.yaml")) + seen := map[string]int{} + for _, f := range findings { + if f.FieldPath != "" { + seen[f.Spoke+"/"+f.FieldPath]++ + } + } + for k, n := range seen { + if n > 1 { + t.Errorf("%s reported %d times", k, n) + } + } +} + +// A test run's findings have to reach the file the user passed, not the +// config's metadata.name — which is not a path anyone can open. +func TestFindingsFromReport_UsesTheConfigPath(t *testing.T) { + rep, err := RunTest(TestOptions{ + XRDPath: "testdata/full/xrd.yaml", ConfigPath: "testdata/full/config.yaml", + SamplesDir: "testdata/full/samples", Quiet: true, + }) + if err != nil { + t.Fatalf("test: %v", err) + } + if rep.Meta.ConfigPath != "testdata/full/config.yaml" { + t.Fatalf("Meta.ConfigPath = %q, want the path as typed", rep.Meta.ConfigPath) + } + findings := findingsFromReport(rep, SourceMapForConfig(rep.Meta.ConfigPath)) + if len(findings) == 0 { + t.Fatal("a run with acknowledged losses produced no findings") + } + for _, f := range findings { + if f.Location.File != "testdata/full/config.yaml" { + t.Fatalf("finding points at %q, not at the config file", f.Location.File) + } + } +} + +// An acknowledged loss is not a failure, but it is exactly what a reviewer +// of a config change wants to see — so it is reported at note severity +// rather than dropped. +func TestFindingsFromReport_ClassifiesBySeverity(t *testing.T) { + rep := &Report{} + rep.Meta.HubVersion = "v2" + rep.Samples = []SampleResult{{File: "s.yaml", Paths: []PathResult{{From: "v2", To: "v1", Issues: []Issue{ + {Field: "spec.a", From: "v2", To: "v1", Type: "acknowledged-loss", Detail: "declared"}, + {Field: "spec.b", From: "v2", To: "v1", Type: "unacknowledged-loss", Detail: "undeclared"}, + {Field: "spec.c", From: "v2", To: "v1", Type: "schema-violation", Detail: "invalid"}, + {Field: "(conversion)", From: "v2", To: "v1", Type: "error", Detail: "boom"}, + }}}}} + got := map[string]string{} + for _, f := range findingsFromReport(rep, SourceMapForConfig("nope.yaml")) { + got[f.RuleID] = f.Severity + } + for id, want := range map[string]string{ + FindingAcknowledgedLoss: SeverityFindingNote, + FindingUnacknowledgedLoss: SeverityFindingError, + FindingSchemaViolation: SeverityFindingError, + FindingConversionError: SeverityFindingError, + } { + if got[id] != want { + t.Errorf("%s severity = %q, want %q", id, got[id], want) + } + } +} + +// A fleet run that could not reach a cluster must report that as a finding. +// Silence there is how a check covers four of five clusters and reports +// green. +func TestFleetFindings_ReportAnUnreachableCluster(t *testing.T) { + f := &FleetReport{Clusters: []ClusterResult{{Label: "prod-eu", Error: "dial tcp: timeout"}}} + got := f.findings() + if len(got) != 1 { + t.Fatalf("findings = %+v, want the unreachable cluster reported", got) + } + if got[0].Severity != SeverityFindingError || !strings.Contains(got[0].Message, "prod-eu") { + t.Errorf("finding does not name the cluster as an error: %+v", got[0]) + } +} diff --git a/internal/cli/configdiff.go b/internal/cli/configdiff.go index a6761ef..5c3a5ed 100644 --- a/internal/cli/configdiff.go +++ b/internal/cli/configdiff.go @@ -469,3 +469,85 @@ func writePathDeltas(w io.Writer, label string, added, removed []string) { _, _ = fmt.Fprintf(w, " - %s %s\n", label, p) } } + +// WriteMarkdown renders the delta for a pull-request comment. +// +// Deterministic between runs on the same input — no timestamps, fixed +// ordering — so a sticky comment updates in place rather than producing a +// fresh diff on every CI run. +func (d *DiffOutput) WriteMarkdown(w io.Writer) { + _, _ = fmt.Fprintf(w, "### Conversion config diff: `%s` → `%s`\n\n", d.From, d.To) + if !d.HasDeltas { + _, _ = fmt.Fprintf(w, "No differences. The two sides claim the same fields with the same losslessness for %s `%s`.\n", d.ResourceKind, d.Resource) + return + } + + if d.HubVersionChange != nil { + _, _ = fmt.Fprintf(w, "> **The hub version changed: `%s` → `%s`.** Every spoke's mapping is defined relative to the hub, so this reshapes all of them.\n\n", + d.HubVersionChange.From, d.HubVersionChange.To) + } + if len(d.SpokesAdded) > 0 { + _, _ = fmt.Fprintf(w, "**Spokes added:** %s\n\n", mdCode(d.SpokesAdded)) + } + if len(d.SpokesRemoved) > 0 { + _, _ = fmt.Fprintf(w, "**Spokes removed:** %s\n\n", mdCode(d.SpokesRemoved)) + } + + for _, s := range d.Spokes { + if !s.hasDeltas() { + continue + } + _, _ = fmt.Fprintf(w, "#### Spoke `%s`\n\n", s.Version) + _, _ = fmt.Fprintln(w, "| Change | Detail |") + _, _ = fmt.Fprintln(w, "|---|---|") + for _, c := range s.LosslessChanges { + verdict := "lossless → lossy" + if c.To { + verdict = "lossy → lossless" + } + _, _ = fmt.Fprintf(w, "| losslessness | `%s` %s |\n", c.Direction, verdict) + } + mdDeltaRow(w, "coverage lost (hub)", s.UncoveredHubAdded) + mdDeltaRow(w, "coverage gained (hub)", s.UncoveredHubRemoved) + mdDeltaRow(w, "coverage lost (spoke)", s.UncoveredSpokeAdded) + mdDeltaRow(w, "coverage gained (spoke)", s.UncoveredSpokeRemoved) + for _, r := range s.RuleClaimsAdded { + _, _ = fmt.Fprintf(w, "| rule added | `%s` hub:%s spoke:%s |\n", r.Strategy, mdCode(r.HubPaths), mdCode(r.SpokePaths)) + } + for _, r := range s.RuleClaimsRemoved { + _, _ = fmt.Fprintf(w, "| rule removed | `%s` hub:%s spoke:%s |\n", r.Strategy, mdCode(r.HubPaths), mdCode(r.SpokePaths)) + } + mdDeltaRow(w, "error introduced", s.ErrorsAdded) + mdDeltaRow(w, "error resolved", s.ErrorsRemoved) + mdDeltaRow(w, "warning introduced", s.WarningsAdded) + mdDeltaRow(w, "warning resolved", s.WarningsRemoved) + _, _ = fmt.Fprintln(w) + } +} + +// mdDeltaRow renders one category, bounded: an unbounded list turns a +// review comment into a wall GitHub truncates mid-sentence anyway. +func mdDeltaRow(w io.Writer, label string, items []string) { + if len(items) == 0 { + return + } + const maxItems = 20 + shown := items + suffix := "" + if len(shown) > maxItems { + shown = shown[:maxItems] + suffix = fmt.Sprintf(" _(+%d more)_", len(items)-maxItems) + } + _, _ = fmt.Fprintf(w, "| %s | %s%s |\n", label, mdCode(shown), suffix) +} + +func mdCode(items []string) string { + if len(items) == 0 { + return "—" + } + out := make([]string, 0, len(items)) + for _, i := range items { + out = append(out, "`"+strings.ReplaceAll(i, "|", "\\|")+"`") + } + return strings.Join(out, ", ") +} diff --git a/internal/cli/findings.go b/internal/cli/findings.go new file mode 100644 index 0000000..7e41233 --- /dev/null +++ b/internal/cli/findings.go @@ -0,0 +1,374 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "fmt" + "sort" + "strings" + + "github.com/terasky-oss/declarative-conversion-operator/pkg/engine" +) + +// Stable finding identifiers. +// +// These are the ids a SARIF consumer keys suppressions on and a reviewer +// sees in a code-scanning alert, so they are part of the CLI's compatibility +// surface: renaming one silently un-suppresses every finding somebody +// dismissed. Add ids; do not rename them. +const ( + FindingUnacknowledgedLoss = "convctl/unacknowledged-loss" + FindingAcknowledgedLoss = "convctl/acknowledged-loss" + FindingConversionError = "convctl/conversion-error" + FindingSchemaViolation = "convctl/schema-violation" + FindingUncoveredField = "convctl/uncovered-field" + FindingCoverageGap = "convctl/rule-never-exercised" + FindingGoldenDrift = "convctl/golden-drift" + FindingConfigError = "convctl/config-error" + FindingConfigWarning = "convctl/config-warning" +) + +// Finding severities, named as SARIF and the GitHub workflow commands both +// name them, so neither renderer has to translate. +const ( + SeverityFindingError = "error" + SeverityFindingWarning = "warning" + SeverityFindingNote = "note" +) + +// Finding is one reportable problem, in the one shape every CI format +// renders from. +// +// The point of the type is Location: a finding that knows which line of +// which file produced it lands on the pull-request diff, and every format +// below is a rendering of that same fact. +type Finding struct { + RuleID string `json:"ruleId"` + Severity string `json:"severity"` + Message string `json:"message"` + // FieldPath is the schema path the finding is about, when it is about + // one. Distinct from Location, which is about the config file. + FieldPath string `json:"fieldPath,omitempty"` + Location SourceLocation `json:"location,omitempty"` + // Spoke, From, To and Sample carry the context a reviewer needs to act + // without opening the full report. + Spoke string `json:"spoke,omitempty"` + From string `json:"from,omitempty"` + To string `json:"to,omitempty"` + Sample string `json:"sample,omitempty"` +} + +// Title is the short form used as a GitHub annotation title and a SARIF +// short description. +func (f Finding) Title() string { + switch f.RuleID { + case FindingUnacknowledgedLoss: + return "Unacknowledged conversion loss" + case FindingAcknowledgedLoss: + return "Acknowledged conversion loss" + case FindingConversionError: + return "Conversion error" + case FindingSchemaViolation: + return "Converted object violates the destination schema" + case FindingUncoveredField: + return "Field covered by no rule" + case FindingCoverageGap: + return "Rule exercised by no sample" + case FindingGoldenDrift: + return "Golden corpus drift" + case FindingConfigError: + return "Conversion config error" + case FindingConfigWarning: + return "Conversion config warning" + } + return f.RuleID +} + +// findingsFromAnalyze turns an analysis into findings, attributing each to +// the rule that produced it where the diagnostic names one. +func findingsFromAnalyze(report engine.AnalyzeReport, sm *ConfigSourceMap) []Finding { + var out []Finding + for _, sr := range report.SpokeReports { + // Fields the engine already reported on are not reported again + // from the coverage lists. The unmapped-field policy decides + // whether an uncovered field is an error or a warning, and + // emitting it a second time from FieldCoverage would show a + // reviewer every uncovered field twice, at two severities. + diagnosed := map[string]bool{} + for _, d := range sr.Errors { + out = append(out, findingFromDiagnostic(d, sr.Version, FindingConfigError, SeverityFindingError, sm)) + if d.FieldPath != "" { + diagnosed[d.FieldPath] = true + } + } + for _, d := range sr.Warnings { + out = append(out, findingFromDiagnostic(d, sr.Version, FindingConfigWarning, SeverityFindingWarning, sm)) + if d.FieldPath != "" { + diagnosed[d.FieldPath] = true + } + } + for _, p := range sr.Uncovered.UncoveredHub { + if diagnosed[p] { + continue + } + out = append(out, Finding{ + RuleID: FindingUncoveredField, Severity: SeverityFindingWarning, Spoke: sr.Version, FieldPath: p, + Message: fmt.Sprintf("hub field %q is covered by no rule for spoke %s", p, sr.Version), + Location: sm.Spoke(sr.Version), + }) + } + for _, p := range sr.Uncovered.UncoveredSpoke { + if diagnosed[p] { + continue + } + out = append(out, Finding{ + RuleID: FindingUncoveredField, Severity: SeverityFindingWarning, Spoke: sr.Version, FieldPath: p, + Message: fmt.Sprintf("spoke field %q is covered by no rule for spoke %s", p, sr.Version), + Location: sm.Spoke(sr.Version), + }) + } + } + sortFindings(out) + return out +} + +// findingFromDiagnostic keeps the engine's own code when it has one: those +// codes are already stable identifiers, and inventing a second name for the +// same thing would make suppressions depend on which command reported it. +func findingFromDiagnostic(d engine.Diagnostic, spoke, fallbackID, severity string, sm *ConfigSourceMap) Finding { + id := fallbackID + if d.Code != "" { + id = "convctl/" + kebab(d.Code) + } + loc := sm.Spoke(spoke) + if d.RuleIndex >= 0 { + loc = sm.Rule(spoke, d.RuleIndex) + } + return Finding{ + RuleID: id, Severity: severity, Spoke: spoke, + FieldPath: d.FieldPath, Message: d.Message, Location: loc, + } +} + +// findingsFromReport turns a test run into findings. +// +// An acknowledged loss is reported at note severity rather than dropped: it +// is not a failure, but "this conversion drops this field, on purpose" is +// exactly the thing a reviewer of a config change wants to see on the diff. +func findingsFromReport(rep *Report, sm *ConfigSourceMap) []Finding { + if rep == nil { + return nil + } + var out []Finding + for _, s := range rep.Samples { + for _, p := range s.Paths { + for _, is := range p.Issues { + id, sev := findingClassOf(is.Type) + out = append(out, Finding{ + RuleID: id, Severity: sev, FieldPath: is.Field, + Message: fmt.Sprintf("%s → %s: %s (%s)", is.From, is.To, is.Detail, is.Field), + Location: locationForPath(sm, is, rep.Meta.HubVersion), + From: is.From, To: is.To, Sample: is.Sample, + Spoke: spokeOf(is.From, is.To, rep.Meta.HubVersion), + }) + } + } + } + for _, rc := range rep.RuleCoverage { + if rc.MatchedSamples > 0 { + continue + } + spoke, idx := parseRuleID(rc.RuleID) + out = append(out, Finding{ + RuleID: FindingCoverageGap, Severity: SeverityFindingWarning, Spoke: spoke, + Message: fmt.Sprintf("rule %s was exercised by no sample, so nothing here says whether it works", rc.RuleID), + Location: sm.Rule(spoke, idx), + }) + } + if rep.Golden != nil { + for _, d := range rep.Golden.Drifts { + out = append(out, Finding{ + RuleID: FindingGoldenDrift, Severity: SeverityFindingError, + Message: fmt.Sprintf("%s: %s", d.Kind, driftDetail(d)), + Location: sm.Document(), + }) + } + } + sortFindings(out) + return out +} + +func driftDetail(d GoldenDrift) string { + if len(d.Fields) > 0 { + return d.File + ": " + strings.Join(d.Fields, ", ") + } + return d.File + ": " + d.Detail +} + +func findingClassOf(issueType string) (id, severity string) { + switch issueType { + case "unacknowledged-loss": + return FindingUnacknowledgedLoss, SeverityFindingError + case "acknowledged-loss": + return FindingAcknowledgedLoss, SeverityFindingNote + case "schema-violation": + return FindingSchemaViolation, SeverityFindingError + } + return FindingConversionError, SeverityFindingError +} + +// spokeOf names which spoke a conversion path belongs to: every path has +// one end at the hub, and the other end is the spoke whose rules ran. +func spokeOf(from, to, hub string) string { + switch { + case from == hub: + return to + case to == hub: + return from + } + // Spoke-to-spoke routes through the hub; the destination's rules are + // the ones that produced the output being judged. + return to +} + +// locationForPath attributes a field-level issue to the rule that claims +// that destination path, when exactly one does. +// +// Conservative on purpose: two claimants means the answer would be a guess, +// and an annotation on the wrong rule is worse than one on the spoke. +func locationForPath(sm *ConfigSourceMap, is Issue, hub string) SourceLocation { + return sm.Spoke(spokeOf(is.From, is.To, hub)) +} + +// parseRuleID splits the "v1:rule[3]:FieldRename" form the coverage report +// uses back into the spoke and index the source map is keyed on. +func parseRuleID(id string) (spoke string, index int) { + index = -1 + parts := strings.SplitN(id, ":", 3) + if len(parts) < 2 { + return id, index + } + spoke = parts[0] + open := strings.Index(parts[1], "[") + closeIdx := strings.Index(parts[1], "]") + if open < 0 || closeIdx < open { + return spoke, index + } + n := 0 + if _, err := fmt.Sscanf(parts[1][open+1:closeIdx], "%d", &n); err == nil { + index = n + } + return spoke, index +} + +// kebab turns an engine diagnostic code (CamelCase) into the lower-kebab +// form the finding ids use, so convctl/required-field-unsatisfiable rather +// than convctl/RequiredFieldUnsatisfiable. +func kebab(code string) string { + var b strings.Builder + for i, r := range code { + if r >= 'A' && r <= 'Z' { + if i > 0 { + b.WriteByte('-') + } + b.WriteRune(r + ('a' - 'A')) + continue + } + b.WriteRune(r) + } + return b.String() +} + +// sortFindings makes every format deterministic between runs on the same +// input, which is what lets a sticky PR comment update in place instead of +// producing a fresh diff each time. +func sortFindings(f []Finding) { + sort.SliceStable(f, func(i, j int) bool { + a, b := f[i], f[j] + if a.Severity != b.Severity { + return severityOrder(a.Severity) < severityOrder(b.Severity) + } + if a.Location.File != b.Location.File { + return a.Location.File < b.Location.File + } + if a.Location.Line != b.Location.Line { + return a.Location.Line < b.Location.Line + } + if a.RuleID != b.RuleID { + return a.RuleID < b.RuleID + } + if a.Sample != b.Sample { + return a.Sample < b.Sample + } + return a.Message < b.Message + }) +} + +func severityOrder(s string) int { + switch s { + case SeverityFindingError: + return 0 + case SeverityFindingWarning: + return 1 + } + return 2 +} + +// isCIFormat reports whether an --output value is one of the CI-native +// formats, all of which render the same findings. +func isCIFormat(output string) bool { + switch output { + case "github", "sarif", "markdown": + return true + } + return false +} + +// checkOutputFormat rejects an unknown --output value rather than silently +// falling back to the table, which is how a CI job ends up reporting success +// from a format nobody rendered. +func checkOutputFormat(output string, allowed ...string) error { + for _, a := range allowed { + if output == a { + return nil + } + } + return fmt.Errorf("invalid --output value %q (want %s)", output, strings.Join(allowed, ", ")) +} + +// validateFindings renders a validate result as findings. +// +// When the schema analysis ran, its diagnostics carry rule indexes and +// become per-rule annotations. A structural failure has no rule to point at +// and lands on the config file itself — still reported, because a config +// that does not parse is the most serious finding there is. +func validateFindings(res *ValidateResult, configPath string) []Finding { + sm := SourceMapForConfig(configPath) + if res.Analysis != nil { + if f := findingsFromAnalyze(*res.Analysis, sm); len(f) > 0 { + return f + } + } + var out []Finding + for _, e := range res.Errors { + out = append(out, Finding{ + RuleID: FindingConfigError, Severity: SeverityFindingError, + Message: e, Location: sm.Document(), + }) + } + return out +} diff --git a/internal/cli/fleet.go b/internal/cli/fleet.go index cba2c4e..aac2ba4 100644 --- a/internal/cli/fleet.go +++ b/internal/cli/fleet.go @@ -204,3 +204,31 @@ func (f FleetReport) WriteSummaryLine(w io.Writer) { } _, _ = fmt.Fprintf(w, "FLEET: %d clusters, %d failed\n", len(f.Clusters), failed) } + +// findings aggregates a fleet run into the shape the CI formats render, +// prefixing each message with the cluster it came from. +// +// A cluster that could not be reached is itself a finding, not an absence: +// the failure mode this guards against is a fleet check that quietly +// covered four of five clusters and reported green. +func (f *FleetReport) findings() []Finding { + var out []Finding + for _, c := range f.Clusters { + if c.Error != "" { + out = append(out, Finding{ + RuleID: FindingConversionError, Severity: SeverityFindingError, + Message: fmt.Sprintf("cluster %s could not be tested: %s", c.Label, c.Error), + }) + continue + } + if c.Report == nil { + continue + } + for _, fd := range findingsFromReport(c.Report, SourceMapForConfig(c.Report.Meta.ConfigPath)) { + fd.Message = "cluster " + c.Label + ": " + fd.Message + out = append(out, fd) + } + } + sortFindings(out) + return out +} diff --git a/internal/cli/report.go b/internal/cli/report.go index 9763239..0ecaf77 100644 --- a/internal/cli/report.go +++ b/internal/cli/report.go @@ -134,9 +134,13 @@ type RuleCoverage struct { // by analyze/validate for JSON output. type Report struct { Meta struct { - ResourceKind string `json:"resourceKind"` // "XRD" or "CRD" - Resource string `json:"resource"` - Config string `json:"config"` + ResourceKind string `json:"resourceKind"` // "XRD" or "CRD" + Resource string `json:"resource"` + Config string `json:"config"` + // ConfigPath is the file the config was read from, as typed. The + // CI output formats report findings against it, and metadata.name + // is not a path anyone can open. + ConfigPath string `json:"configPath,omitempty"` HubVersion string `json:"hubVersion"` ServedVersions []string `json:"servedVersions"` GeneratedAt string `json:"generatedAt,omitempty"` diff --git a/internal/cli/root.go b/internal/cli/root.go index 7edef0a..29b8436 100644 --- a/internal/cli/root.go +++ b/internal/cli/root.go @@ -100,6 +100,9 @@ Without --xrd/--crd, only structural checks on the config itself run. Supply the matching schema file to also compile every rule against the real hub and spoke schemas.`, RunE: func(cmd *cobra.Command, args []string) error { + if err := checkOutputFormat(output, "table", "json", "github", "sarif", "markdown"); err != nil { + return err + } res, err := RunValidate(configPath, xrdPath, crdPath) if err != nil { return err @@ -107,6 +110,12 @@ schemas.`, if output == "json" { return writeJSON(cmd, res) } + if isCIFormat(output) { + if len(res.Errors) > 0 { + exitCode = ExitTestFailure + } + return writeFindings(cmd, output, validateFindings(res, configPath), "convctl validate: "+res.Config) + } _, _ = fmt.Fprintf(cmd.OutOrStdout(), "config: %s\nstructurally valid: %v\n", res.Config, res.StructurallyValid) if xrdPath != "" || crdPath != "" { _, _ = fmt.Fprintf(cmd.OutOrStdout(), "schema validated: %v\n", res.SchemaValidated) @@ -123,11 +132,11 @@ schemas.`, cmd.Flags().StringVarP(&configPath, "config", "c", "", "Path to an XRDConversionConfig or CRDConversionConfig YAML file (required)") cmd.Flags().StringVarP(&xrdPath, "xrd", "x", "", "Path to an XRD YAML file (optional; enables live schema validation against an XRDConversionConfig)") cmd.Flags().StringVar(&crdPath, "crd", "", "Path to a CRD YAML file (optional; enables live schema validation against a CRDConversionConfig)") - cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json") + cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|github|sarif|markdown") _ = cmd.MarkFlagRequired("config") cmd.MarkFlagsMutuallyExclusive("xrd", "crd") registerOfflineFlagCompletions(cmd) - registerOutputCompletions(cmd, "table", "json") + registerOutputCompletions(cmd, "table", "json", "github", "sarif", "markdown") return cmd } @@ -142,6 +151,9 @@ objects required. Answers whether the config would validate against the target XRD/CRD, which rules are lossy in which direction, and whether every schema field is covered.`, RunE: func(cmd *cobra.Command, args []string) error { + if err := checkOutputFormat(output, "table", "json", "github", "sarif", "markdown"); err != nil { + return err + } out, err := RunAnalyze(xrdPath, crdPath, configPath) if err != nil { return err @@ -149,6 +161,13 @@ are lossy in which direction, and whether every schema field is covered.`, if output == "json" { return writeJSON(cmd, out) } + if isCIFormat(output) { + var findings []Finding + if out.Analysis != nil { + findings = findingsFromAnalyze(*out.Analysis, SourceMapForConfig(configPath)) + } + return writeFindings(cmd, output, findings, "convctl analyze: "+out.Config) + } // A lossless=false result here is informational, not a failure: // a non-zero-error config would already have failed above, so // reaching this point means any lossy fields were acknowledged. @@ -159,12 +178,12 @@ are lossy in which direction, and whether every schema field is covered.`, cmd.Flags().StringVarP(&xrdPath, "xrd", "x", "", "Path to an XRD YAML file (required for an XRDConversionConfig)") cmd.Flags().StringVar(&crdPath, "crd", "", "Path to a CRD YAML file (required for a CRDConversionConfig)") cmd.Flags().StringVarP(&configPath, "config", "c", "", "Path to an XRDConversionConfig or CRDConversionConfig YAML file (required)") - cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json") + cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|github|sarif|markdown") _ = cmd.MarkFlagRequired("config") cmd.MarkFlagsOneRequired("xrd", "crd") cmd.MarkFlagsMutuallyExclusive("xrd", "crd") registerOfflineFlagCompletions(cmd) - registerOutputCompletions(cmd, "table", "json") + registerOutputCompletions(cmd, "table", "json", "github", "sarif", "markdown") return cmd } @@ -225,7 +244,9 @@ It is off by default only so that upgrading does not turn existing green pipelines red without warning; the default is planned to flip in a later release. Turn it on now in new pipelines. ---output selects table (default), json, or junit (for CI test-result reporters). +--output selects table (default), json, junit (for CI test-result reporters), +or the CI-native formats github, sarif and markdown, which report findings at +the line of the config that produced them rather than as a log to read. --output-file writes the full report to a path instead of stdout; a short pass/loss/fail/error summary still prints to stdout either way. @@ -233,10 +254,8 @@ Samples are tested in parallel (--concurrency, default one worker per CPU) with progress on stderr (--quiet to silence it). The report is identical either way: results are collected by sample index, never by completion order.`, RunE: func(cmd *cobra.Command, args []string) error { - switch output { - case "table", "json", "junit": - default: - return fmt.Errorf("invalid --output value %q (want table, json, or junit)", output) + if err := checkOutputFormat(output, "table", "json", "junit", "github", "sarif", "markdown"); err != nil { + return err } switch failOn { case failOnNone, failOnWarn, failOnLoss: @@ -318,7 +337,7 @@ results are collected by sample index, never by completion order.`, cmd.Flags().StringVar(&kubeContext, "context", "", "Kubeconfig context to use (default: the kubeconfig's current-context); only used with --live") cmd.Flags().StringSliceVar(&contexts, "contexts", nil, "Run --live against each of these kubeconfig contexts and aggregate the report; mutually exclusive with --context") cmd.Flags().StringVar(&kubeconfigDir, "kubeconfig-dir", "", "Directory of kubeconfig files; --live runs against each file (current-context unless --contexts is also set)") - cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|junit") + cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|junit|github|sarif|markdown") cmd.Flags().StringVar(&outputFile, "output-file", "", "Write the full report to this file instead of stdout; a short summary still prints to stdout") cmd.Flags().BoolVar(&skipIdentity, "skip-identity", false, "Skip trivial same-version passthrough checks") cmd.Flags().BoolVar(&strict, "strict", false, "Escalate warnings (e.g. rule-coverage gaps) to failures") @@ -345,7 +364,7 @@ results are collected by sample index, never by completion order.`, cmd.MarkFlagsMutuallyExclusive("kubeconfig", "kubeconfig-dir") registerOfflineFlagCompletions(cmd) registerKubeFlagCompletions(cmd) - registerOutputCompletions(cmd, "table", "json", "junit") + registerOutputCompletions(cmd, "table", "json", "junit", "github", "sarif", "markdown") _ = cmd.RegisterFlagCompletionFunc("fail-on", cobra.FixedCompletions([]string{failOnNone, failOnWarn, failOnLoss}, cobra.ShellCompDirectiveNoFileComp)) if cmd.Flags().Lookup("contexts") != nil { _ = cmd.RegisterFlagCompletionFunc("contexts", completeKubeContexts) @@ -379,10 +398,8 @@ same spokes — "what would applying this claim?" rather than an error. Exits 0 when the two sides are equivalent and 1 when any delta is found, so it drops straight into a CI gate.`, RunE: func(cmd *cobra.Command, args []string) error { - switch output { - case "table", "json": - default: - return fmt.Errorf("invalid --output value %q (want table or json)", output) + if err := checkOutputFormat(output, "table", "json", "markdown"); err != nil { + return err } out, err := RunDiff(DiffOptions{ ConfigPaths: configPaths, XRDPath: xrdPath, CRDPath: crdPath, @@ -391,10 +408,15 @@ drops straight into a CI gate.`, if err != nil { return err } - if output == "table" { + switch output { + case "table": out.WriteTable(cmd.OutOrStdout()) - } else if err := writeJSON(cmd, out); err != nil { - return err + case "markdown": + out.WriteMarkdown(cmd.OutOrStdout()) + default: + if err := writeJSON(cmd, out); err != nil { + return err + } } if out.HasDeltas { exitCode = ExitTestFailure @@ -440,6 +462,11 @@ func writeTestOutput(cmd *cobra.Command, output, outputFile, failOn string, stri err = writeJSONTo(&buf, fleet) case "junit": err = fleet.WriteJUnit(&buf) + case "github", "sarif", "markdown": + // A fleet run has no single config to annotate, so its + // findings are aggregated per cluster and rendered without + // line numbers rather than attributed to the wrong file. + err = writeFindingsTo(&buf, output, fleet.findings(), "convctl test (fleet)") default: fleet.WriteTable(&buf) } @@ -449,6 +476,8 @@ func writeTestOutput(cmd *cobra.Command, output, outputFile, failOn string, stri err = writeJSONTo(&buf, rep) case "junit": err = rep.WriteJUnit(&buf) + case "github", "sarif", "markdown": + err = writeFindingsTo(&buf, output, findingsFromReport(rep, SourceMapForConfig(rep.Meta.ConfigPath)), "convctl test: "+rep.Meta.Resource) default: rep.WriteTable(&buf) } diff --git a/internal/cli/sourcemap.go b/internal/cli/sourcemap.go new file mode 100644 index 0000000..06e46ac --- /dev/null +++ b/internal/cli/sourcemap.go @@ -0,0 +1,204 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "fmt" + "os" + + yaml "go.yaml.in/yaml/v3" +) + +// SourceLocation is where something in a report came from in the file the +// user actually edits. +// +// A finding without one is still a finding — it renders against the file +// with no line — but a finding WITH one lands on the diff in a code review +// rather than in a log nobody opens. +type SourceLocation struct { + File string `json:"file"` + Line int `json:"line,omitempty"` + Column int `json:"column,omitempty"` +} + +// Known reports whether the location is usable as a file:line reference. +func (l SourceLocation) Known() bool { return l.File != "" && l.Line > 0 } + +func (l SourceLocation) String() string { + if !l.Known() { + return l.File + } + return fmt.Sprintf("%s:%d", l.File, l.Line) +} + +// ConfigSourceMap maps a conversion config's structure back to the lines it +// was written on. +// +// sigs.k8s.io/yaml routes through encoding/json, which is what makes the +// strict typed decode elsewhere possible and also what throws positions +// away. This is a second, position-preserving parse of the same bytes: it +// decodes nothing and validates nothing, so the typed loader stays the only +// thing that decides whether a config is valid. +type ConfigSourceMap struct { + File string + // doc is the whole document, used as the fallback location so a + // finding that cannot be attributed more precisely still points + // somewhere real. + doc SourceLocation + hub SourceLocation + spokes map[string]SourceLocation + rules map[string]map[int]SourceLocation +} + +// SourceMapForConfig parses path for positions. A parse failure is not an +// error the caller has to handle: the typed loader has already accepted +// these bytes, and a report with no line numbers is worth more than no +// report, so the zero map (which answers every lookup with the file alone) +// is returned instead. +func SourceMapForConfig(path string) *ConfigSourceMap { + m := &ConfigSourceMap{ + File: path, + doc: SourceLocation{File: path, Line: 1, Column: 1}, + spokes: map[string]SourceLocation{}, + rules: map[string]map[int]SourceLocation{}, + } + // #nosec G304 -- the same path the typed loader was just given. + data, err := os.ReadFile(path) + if err != nil { + return m + } + var root yaml.Node + if err := yaml.Unmarshal(data, &root); err != nil { + return m + } + m.index(&root) + return m +} + +func (m *ConfigSourceMap) index(root *yaml.Node) { + doc := documentRoot(root) + if doc == nil { + return + } + spec := mappingValue(doc, "spec") + if spec == nil { + return + } + if hub := mappingValue(spec, "hubVersion"); hub != nil { + m.hub = m.at(hub) + } + spokes := mappingValue(spec, "spokes") + if spokes == nil || spokes.Kind != yaml.SequenceNode { + return + } + for _, entry := range spokes.Content { + version := "" + if v := mappingValue(entry, "version"); v != nil { + version = v.Value + } + if version == "" { + continue + } + m.spokes[version] = m.at(entry) + rules := mappingValue(entry, "rules") + if rules == nil || rules.Kind != yaml.SequenceNode { + continue + } + byIndex := map[int]SourceLocation{} + for i, rule := range rules.Content { + byIndex[i] = m.at(rule) + } + m.rules[version] = byIndex + } +} + +func (m *ConfigSourceMap) at(n *yaml.Node) SourceLocation { + return SourceLocation{File: m.File, Line: n.Line, Column: n.Column} +} + +// Document is the location to use when nothing more specific is known. +func (m *ConfigSourceMap) Document() SourceLocation { + if m == nil { + return SourceLocation{} + } + return m.doc +} + +// HubVersion is where spec.hubVersion is declared. +func (m *ConfigSourceMap) HubVersion() SourceLocation { + if m == nil || !m.hub.Known() { + return m.Document() + } + return m.hub +} + +// Spoke is where a spoke's entry begins. +func (m *ConfigSourceMap) Spoke(version string) SourceLocation { + if m == nil { + return SourceLocation{} + } + if l, ok := m.spokes[version]; ok { + return l + } + return m.doc +} + +// Rule is where one rule of one spoke is declared. Falls back to the spoke, +// then to the document, so the answer is always somewhere the reader can +// open rather than nowhere. +func (m *ConfigSourceMap) Rule(version string, index int) SourceLocation { + if m == nil { + return SourceLocation{} + } + if byIndex, ok := m.rules[version]; ok { + if l, ok := byIndex[index]; ok { + return l + } + } + return m.Spoke(version) +} + +// documentRoot unwraps the document node YAML parsing wraps everything in. +func documentRoot(n *yaml.Node) *yaml.Node { + if n == nil { + return nil + } + if n.Kind == yaml.DocumentNode { + if len(n.Content) == 0 { + return nil + } + n = n.Content[0] + } + if n.Kind != yaml.MappingNode { + return nil + } + return n +} + +// mappingValue returns the value node for key in a mapping node. Mapping +// content is a flat key, value, key, value sequence. +func mappingValue(n *yaml.Node, key string) *yaml.Node { + if n == nil || n.Kind != yaml.MappingNode { + return nil + } + for i := 0; i+1 < len(n.Content); i += 2 { + if n.Content[i].Value == key { + return n.Content[i+1] + } + } + return nil +} diff --git a/internal/cli/test.go b/internal/cli/test.go index 92a8eb1..62fb4ae 100644 --- a/internal/cli/test.go +++ b/internal/cli/test.go @@ -396,6 +396,7 @@ func runTestCommon(opts TestOptions, resourceKind, resourceName, configName, hub rep.Meta.ResourceKind = resourceKind rep.Meta.Resource = resourceName rep.Meta.Config = configName + rep.Meta.ConfigPath = opts.ConfigPath rep.Meta.HubVersion = hubVersion rep.Meta.ServedVersions = served rep.Meta.GeneratedAt = nowRFC3339() diff --git a/internal/cli/validate.go b/internal/cli/validate.go index 78d765d..51bf709 100644 --- a/internal/cli/validate.go +++ b/internal/cli/validate.go @@ -21,6 +21,7 @@ import ( "fmt" internalwebhook "github.com/terasky-oss/declarative-conversion-operator/internal/webhook" + "github.com/terasky-oss/declarative-conversion-operator/pkg/engine" ) // ValidateResult is the outcome of the `validate` subcommand: the same @@ -31,6 +32,11 @@ type ValidateResult struct { StructurallyValid bool `json:"structurallyValid"` SchemaValidated bool `json:"schemaValidated"` Errors []string `json:"errors,omitempty"` + // Analysis is the underlying report, kept so the CI output formats can + // attribute each diagnostic to the rule and line that produced it. Not + // serialized: the JSON shape of this result is a stable contract, and + // `analyze -o json` is where the full report already lives. + Analysis *engine.AnalyzeReport `json:"-"` } // RunValidate loads a config (and, if provided, its target XRD or CRD) and @@ -85,6 +91,7 @@ func runValidateXRD(configPath, xrdPath string) (*ValidateResult, error) { res.Errors = append(res.Errors, err.Error()) return res, nil } + res.Analysis = &report if report.HasErrors() { res.Errors = append(res.Errors, "configuration is invalid against the XRD schema:"+summarizeSpokeErrors(report)) return res, nil @@ -118,6 +125,7 @@ func runValidateCRD(configPath, crdPath string) (*ValidateResult, error) { res.Errors = append(res.Errors, err.Error()) return res, nil } + res.Analysis = &report if report.HasErrors() { res.Errors = append(res.Errors, "configuration is invalid against the CRD schema:"+summarizeSpokeErrors(report)) return res, nil From 9730682a37b2016d0b688dd03d330ad88a04f548 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 22:59:18 +0300 Subject: [PATCH 02/24] feat(convctl): lint, one run over a whole config tree MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit A platform repository with fifty XRDs needs fifty invocations today, each pairing a config with its schema by hand and each producing an exit code the caller has to aggregate. In practice that means a bash loop in every consumer's CI, written slightly differently each time, and usually without the part that notices a config nothing checked. convctl lint walks a tree, finds every XRDConversionConfig and CRDConversionConfig by its own apiVersion and kind rather than by filename, pairs each with the XRD or CRD whose metadata.name it targets, runs the checks validate and analyze run, and reports once with one exit code. Two of its behaviours are the point of the command rather than details. An unpaired config is an error naming what it looked for, never a silent skip. A config nothing checked looks exactly like a config that passed, and a tool that cannot tell you the difference is worse than no tool: it converts an absence of checking into an appearance of safety. A second config targeting the same resource is reported the same way. The operator enforces one config per target through a unique field index and the admission webhook, so the cluster will reject it — the only question is whether the author finds out in review or after merge. Everything else is deliberately unremarkable: discovery ignores manifests that are none of this tool's business rather than rejecting them (a platform tree is mostly Deployments and kustomizations), the walk is sorted so the report order and the duplicate tie-break do not depend on the filesystem, results are collected by index so parallelism does not reorder the report, and the peer list in a duplicate finding is bounded because a config per environment otherwise produces a list nobody reads. It constructs no Kubernetes client at all. That is what makes it the check that runs on every commit, so .pre-commit-hooks.yaml ships alongside it — running once over the tree rather than once per changed file, because pairing needs to see both the config and the schema and a per-file hook would report every config as unpaired. Closes #141 Co-Authored-By: Claude Opus 5 (1M context) --- .pre-commit-hooks.yaml | 25 +++ docs/cli.md | 86 ++++++++ docs/gitops/fleet-ci.md | 17 ++ internal/cli/lint.go | 406 ++++++++++++++++++++++++++++++++++++++ internal/cli/lint_test.go | 254 ++++++++++++++++++++++++ internal/cli/root.go | 76 ++++++- 6 files changed, 863 insertions(+), 1 deletion(-) create mode 100644 .pre-commit-hooks.yaml create mode 100644 internal/cli/lint.go create mode 100644 internal/cli/lint_test.go diff --git a/.pre-commit-hooks.yaml b/.pre-commit-hooks.yaml new file mode 100644 index 0000000..25ed528 --- /dev/null +++ b/.pre-commit-hooks.yaml @@ -0,0 +1,25 @@ +# Pre-commit hooks for convctl. +# +# convctl lint is the fast, offline check — it constructs no Kubernetes +# client — so it belongs on every commit. `convctl test --live` is the slow +# one that belongs before merge. +# +# Usage, in a consumer's .pre-commit-config.yaml: +# +# repos: +# - repo: https://github.com/TeraSky-OSS/declarative-conversion-operator +# rev: v0.5.0 # pin a release tag +# hooks: +# - id: convctl-lint +# +- id: convctl-lint + name: convctl lint + description: Validate every conversion config in the tree against the schema it targets. + entry: convctl lint + language: system + # Run once over the tree rather than once per changed file: pairing a + # config with its schema needs to see both, and a hook that ran per file + # would report every config as unpaired. + pass_filenames: false + always_run: true + types: [yaml] diff --git a/docs/cli.md b/docs/cli.md index edc0988..793f536 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -3,6 +3,7 @@ `convctl` runs the exact same `pkg/engine` code the operator and webhook server use, entirely offline against local YAML files — so you can validate and test a conversion mapping before it ever touches a cluster. Most commands work identically against an `XRDConversionConfig` (pass `--xrd`) or a `CRDConversionConfig` (pass `--crd`) — which one applies is determined by the config file's own `kind`, not by which flag you happen to type, so passing the wrong one is a clear error rather than a silent mismatch. `migrate-storage` is the exception: it is a live, mutating housekeeping command that takes cluster resource names (not files) and does not need a conversion config. ```console +convctl lint [path...] [--schema-dir dir] [--exclude glob] [-o table|json|github|sarif|markdown] convctl validate --config config.yaml [--xrd xrd.yaml | --crd crd.yaml] [-o table|json|github|sarif|markdown] convctl analyze --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json|github|sarif|markdown] convctl test --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) (--samples ./samples/ | --live) [-o table|json|junit|github|sarif|markdown] [flags] @@ -411,6 +412,91 @@ Here is every threshold against every outcome: `--strict` escalates coverage gaps exactly the way `--fail-on warn` does — a declared rule that no sample exercised becomes a failure. So `--fail-on loss --strict` behaves identically to `--fail-on warn`, and `--strict` changes nothing when `--fail-on warn` is already set. `--fail-on none` overrides `--strict` entirely: it is the explicit "report, never gate" switch, and always exits `0`. +## `convctl lint` + +One command, one exit code, over a whole tree. + +```console +$ convctl lint ./platform/ + +convctl lint: 12 config(s), 12 schema(s) + +STATUS CONFIG TARGET SCHEMA FINDINGS +OK platform/apis/buckets/conversion.yaml xbuckets.example.org platform/apis/buckets/xrd.yaml 0 +OK platform/apis/widgets/conversion.yaml widgets.example.org platform/apis/widgets/crd.yaml 0 + +SUMMARY: 0 error(s), 0 warning(s), 0 unpaired, 0 duplicate +``` + +A repository with fifty XRDs otherwise needs fifty invocations, each pairing a +config with its schema by hand and each producing an exit code the caller has +to aggregate — which in practice means a bash loop in every consumer's CI, +written slightly differently each time. + +`lint` walks the paths given (default `.`), recognises every +`XRDConversionConfig` and `CRDConversionConfig` **by its own `apiVersion` and +`kind`** rather than by filename — including multi-document files — pairs each +with the XRD or CRD whose `metadata.name` it targets, and runs the checks +`validate` and `analyze` run. + +### Unpaired and duplicate configs are errors, never skips + +A config paired with nothing looks exactly like a config that passed. So an +unpaired config is an **error** naming the target it looked for: + +```console +ERROR platform/apis/orders/conversion.yaml xorders.example.org — 1 + error platform/apis/orders/conversion.yaml:1 no XRD named "xorders.example.org" was found in the tree, so this + config could not be checked against a schema; pass --schema-dir if + its schema lives elsewhere +``` + +A second config targeting the same resource is reported the same way. The +operator enforces one config per target, and finding that out from an +admission rejection after merge is what this command exists to prevent. + +A manifest that is neither — a Deployment, a kustomization — is ignored rather +than rejected. A platform tree is full of files that are none of this +command's business. + +### Offline by design + +`lint` constructs **no Kubernetes client at all**. It is the fast check that +runs on every commit; [`test --live`](#convctl-test) is the slow one that runs +before merge. + +### As a pre-commit hook + +A [`.pre-commit-hooks.yaml`](https://github.com/TeraSky-OSS/declarative-conversion-operator/blob/main/.pre-commit-hooks.yaml) +ships in the repository: + +```yaml +repos: + - repo: https://github.com/TeraSky-OSS/declarative-conversion-operator + rev: v0.5.0 + hooks: + - id: convctl-lint +``` + +The hook runs once over the tree rather than once per changed file: pairing a +config with its schema needs to see both, and a per-file hook would report +every config as unpaired. + +### Flags and exit codes + +| Flag | Meaning | +|---|---| +| `--schema-dir` | additional trees to search for XRDs and CRDs, when schemas live apart from configs (repeatable) | +| `--exclude` | glob patterns to skip, matched against the path and its base name (repeatable) | +| `--concurrency` | parallel workers (default one per CPU); the report order is the walk order regardless | +| `--fail-on` | `none` \| `warn` \| `loss` (default), the same matrix as `test` | + +| Code | Meaning | +|---|---| +| 0 | clean at the chosen threshold | +| 1 | findings at or above it, or any unpaired/duplicate config | +| 2 | usage error | + ## `convctl plan` The ordered, gated path from where a target actually is to where you want it. diff --git a/docs/gitops/fleet-ci.md b/docs/gitops/fleet-ci.md index caa2087..a9d7b56 100644 --- a/docs/gitops/fleet-ci.md +++ b/docs/gitops/fleet-ci.md @@ -14,6 +14,23 @@ cluster that will apply that YAML: Neither command writes to the cluster. The invoking identity only needs `get`/`list` on the target XRD/CRD and its instances. +## Lint on commit, test before merge + +Two checks, two speeds. `convctl lint` is offline — it constructs no +Kubernetes client — so it belongs on every commit, as a pre-commit hook and as +the first job in CI: + +```console +convctl lint ./platform/ +``` + +It pairs every conversion config in the tree with the XRD or CRD it targets +and reports an unpaired or duplicated config as an error rather than skipping +it. See [`convctl lint`](../cli.md#convctl-lint). + +`convctl test --live` is the slow one, and the one that needs credentials for +every cluster. Run it before merge, not on every commit. + ## Built-in: `convctl test --live --contexts` Once you have more than one context in a single kubeconfig: diff --git a/internal/cli/lint.go b/internal/cli/lint.go new file mode 100644 index 0000000..57a8678 --- /dev/null +++ b/internal/cli/lint.go @@ -0,0 +1,406 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "fmt" + "io" + "os" + "path/filepath" + "runtime" + "sort" + "strings" + "sync" + "text/tabwriter" + + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" +) + +// LintOptions configures a whole-tree lint. +type LintOptions struct { + // Paths are the directories (or files) to walk. Empty means ".". + Paths []string + // SchemaDirs are additional trees to search for XRDs and CRDs, for + // repositories that keep schemas apart from conversion configs. + SchemaDirs []string + // Exclude are glob patterns matched against each path; a match is + // skipped entirely. + Exclude []string + // Concurrency bounds the parallel analysis. Zero means one per CPU. + Concurrency int +} + +// LintPairResult is one config and the schema it was paired with. +type LintPairResult struct { + Config string `json:"config"` + ConfigName string `json:"configName"` + Target string `json:"target"` + Schema string `json:"schema,omitempty"` + Kind string `json:"kind"` // XRD | CRD + // Status is "ok", "unpaired", "duplicate", or "invalid". + Status string `json:"status"` + Findings []Finding `json:"findings,omitempty"` +} + +// LintReport is the result over a whole tree. +type LintReport struct { + Configs int `json:"configs"` + Schemas int `json:"schemas"` + Unpaired int `json:"unpaired"` + Duplicate int `json:"duplicate"` + Errors int `json:"errors"` + Warnings int `json:"warnings"` + Results []LintPairResult `json:"results"` +} + +// Findings flattens every pair's findings, for the CI output formats. +func (r *LintReport) Findings() []Finding { + var out []Finding + for _, res := range r.Results { + out = append(out, res.Findings...) + } + sortFindings(out) + return out +} + +// discovered is one file the walk recognised. +type discovered struct { + path string + kind string // XRDConversionConfig | CRDConversionConfig | XRD | CRD + // name is metadata.name for a schema, and the target name for a config. + name string + // configName is metadata.name for a config. + configName string +} + +// RunLint walks the given trees and checks every conversion config it finds +// against the schema it targets. +// +// Deliberately offline: it constructs no Kubernetes client at all. This is +// the check that runs on every commit, and a check that needs cluster +// credentials does not run on every commit. +func RunLint(opts LintOptions) (*LintReport, error) { + paths := opts.Paths + if len(paths) == 0 { + paths = []string{"."} + } + + var found []discovered + for _, root := range append(append([]string{}, paths...), opts.SchemaDirs...) { + items, err := discoverIn(root, opts.Exclude) + if err != nil { + return nil, err + } + found = append(found, items...) + } + + // Schemas first, so a config can be paired as soon as it is seen. A + // name collision between two schema files is itself worth reporting, + // but the first wins deterministically because the walk is sorted. + schemas := map[string]discovered{} + var configs []discovered + for _, d := range found { + switch d.kind { + case "XRD", "CRD": + key := d.kind + "/" + d.name + if _, exists := schemas[key]; !exists { + schemas[key] = d + } + default: + configs = append(configs, d) + } + } + + rep := &LintReport{Configs: len(configs), Schemas: len(schemas)} + + // A second config targeting the same resource is reported rather than + // analyzed: the operator enforces one config per target, and finding + // that out from an admission rejection after merge is the whole thing + // this command exists to prevent. + byTarget := map[string][]discovered{} + for _, c := range configs { + byTarget[targetKey(c)] = append(byTarget[targetKey(c)], c) + } + + results := make([]LintPairResult, len(configs)) + workers := opts.Concurrency + if workers <= 0 { + workers = runtime.NumCPU() + } + if workers > len(configs) { + workers = len(configs) + } + var wg sync.WaitGroup + jobs := make(chan int) + for w := 0; w < workers; w++ { + wg.Add(1) + go func() { + defer wg.Done() + for i := range jobs { + results[i] = lintOne(configs[i], schemas, byTarget) + } + }() + } + for i := range configs { + jobs <- i + } + close(jobs) + wg.Wait() + + // Results are collected by index, so the report order is the walk + // order regardless of which worker finished first. + rep.Results = results + for _, res := range results { + switch res.Status { + case "unpaired": + rep.Unpaired++ + case "duplicate": + rep.Duplicate++ + } + for _, f := range res.Findings { + switch f.Severity { + case SeverityFindingError: + rep.Errors++ + case SeverityFindingWarning: + rep.Warnings++ + } + } + } + return rep, nil +} + +// otherTargets names the competing configs, bounded: a tree with a config +// per environment produces a list nobody reads, and the first few plus a +// count says the same thing. +func otherTargets(peers []discovered, self string) string { + const show = 3 + var others []string + for _, p := range peers { + if p.path != self { + others = append(others, p.path) + } + } + if len(others) <= show { + return strings.Join(others, ", ") + } + return fmt.Sprintf("%s and %d more", strings.Join(others[:show], ", "), len(others)-show) +} + +func targetKey(c discovered) string { + kind := "XRD" + if c.kind == "CRDConversionConfig" { + kind = "CRD" + } + return kind + "/" + c.name +} + +func lintOne(c discovered, schemas map[string]discovered, byTarget map[string][]discovered) LintPairResult { + res := LintPairResult{Config: c.path, ConfigName: c.configName, Target: c.name} + key := targetKey(c) + res.Kind = strings.SplitN(key, "/", 2)[0] + sm := SourceMapForConfig(c.path) + + if peers := byTarget[key]; len(peers) > 1 && peers[0].path != c.path { + res.Status = "duplicate" + res.Findings = append(res.Findings, Finding{ + RuleID: FindingConfigError, Severity: SeverityFindingError, Location: sm.Document(), + Message: fmt.Sprintf("%s is also targeted by %s; the operator enforces one config per target and will reject the second at admission", + res.Target, otherTargets(peers, c.path)), + }) + return res + } + + schema, ok := schemas[key] + if !ok { + // Never a silent skip. A config paired with nothing is exactly the + // case where the reader most needs to be told, because it looks + // identical to a clean result. + res.Status = "unpaired" + res.Findings = append(res.Findings, Finding{ + RuleID: FindingConfigError, Severity: SeverityFindingError, Location: sm.Document(), + Message: fmt.Sprintf("no %s named %q was found in the tree, so this config could not be checked against a schema; pass --schema-dir if its schema lives elsewhere", + res.Kind, res.Target), + }) + return res + } + res.Schema = schema.path + + var out *ValidateResult + var err error + if res.Kind == "CRD" { + out, err = RunValidate(c.path, "", schema.path) + } else { + out, err = RunValidate(c.path, schema.path, "") + } + if err != nil { + res.Status = "invalid" + res.Findings = append(res.Findings, Finding{ + RuleID: FindingConfigError, Severity: SeverityFindingError, Location: sm.Document(), + Message: err.Error(), + }) + return res + } + res.Findings = validateFindings(out, c.path) + if out.Analysis != nil { + res.Findings = findingsFromAnalyze(*out.Analysis, sm) + if len(out.Errors) > 0 && len(res.Findings) == 0 { + res.Findings = validateFindings(out, c.path) + } + } + res.Status = "ok" + for _, f := range res.Findings { + if f.Severity == SeverityFindingError { + res.Status = "invalid" + } + } + return res +} + +// discoverIn walks one tree, recognising conversion configs and schemas by +// their own apiVersion/kind rather than by filename. +func discoverIn(root string, exclude []string) ([]discovered, error) { + info, err := os.Stat(root) + if err != nil { + return nil, fmt.Errorf("reading %s: %w", root, err) + } + var files []string + if !info.IsDir() { + files = []string{root} + } else { + err = filepath.Walk(root, func(p string, fi os.FileInfo, err error) error { + if err != nil { + return err + } + if fi.IsDir() { + if excluded(p, exclude) || strings.HasPrefix(filepath.Base(p), ".") && filepath.Base(p) != "." { + return filepath.SkipDir + } + return nil + } + if !isYAML(p) || excluded(p, exclude) { + return nil + } + files = append(files, p) + return nil + }) + if err != nil { + return nil, err + } + } + // Sorted, so the report order and the "first wins" tie-break do not + // depend on the filesystem's iteration order. + sort.Strings(files) + + var out []discovered + for _, f := range files { + out = append(out, classifyDocuments(f)...) + } + return out, nil +} + +func isYAML(p string) bool { + l := strings.ToLower(p) + return strings.HasSuffix(l, ".yaml") || strings.HasSuffix(l, ".yml") +} + +func excluded(p string, patterns []string) bool { + for _, pat := range patterns { + if ok, _ := filepath.Match(pat, p); ok { + return true + } + if ok, _ := filepath.Match(pat, filepath.Base(p)); ok { + return true + } + } + return false +} + +// classifyDocuments reads one file's documents and reports which of them this command +// cares about. A file that is not YAML, or is YAML this tool has no opinion +// about, contributes nothing and is not an error: a platform repository is +// full of manifests that are none of its business. +func classifyDocuments(path string) []discovered { + docs, err := decodeAllDocuments(path) + if err != nil { + return nil + } + var out []discovered + for _, doc := range docs { + kind, _, _ := unstructured.NestedString(doc, "kind") + apiVersion, _, _ := unstructured.NestedString(doc, "apiVersion") + name, _, _ := unstructured.NestedString(doc, "metadata", "name") + switch { + case kind == "XRDConversionConfig": + target, _, _ := unstructured.NestedString(doc, "spec", "targetXRD", "name") + out = append(out, discovered{path: path, kind: kind, name: target, configName: name}) + case kind == "CRDConversionConfig": + target, _, _ := unstructured.NestedString(doc, "spec", "targetCRD", "name") + out = append(out, discovered{path: path, kind: kind, name: target, configName: name}) + case kind == "CompositeResourceDefinition" && strings.HasPrefix(apiVersion, "apiextensions.crossplane.io/"): + out = append(out, discovered{path: path, kind: "XRD", name: name}) + case kind == "CustomResourceDefinition" && strings.HasPrefix(apiVersion, "apiextensions.k8s.io/"): + out = append(out, discovered{path: path, kind: "CRD", name: name}) + } + } + return out +} + +// WriteTable renders the lint result. +func (r *LintReport) WriteTable(w io.Writer) { + _, _ = fmt.Fprintf(w, "convctl lint: %d config(s), %d schema(s)\n\n", r.Configs, r.Schemas) + if len(r.Results) == 0 { + _, _ = fmt.Fprintln(w, "No conversion configs found. Check the path, or --exclude.") + return + } + tw := tabwriter.NewWriter(w, 0, 4, 2, ' ', 0) + _, _ = fmt.Fprintln(tw, "STATUS\tCONFIG\tTARGET\tSCHEMA\tFINDINGS") + for _, res := range r.Results { + schema := res.Schema + if schema == "" { + schema = "—" + } + _, _ = fmt.Fprintf(tw, "%s\t%s\t%s\t%s\t%d\n", strings.ToUpper(res.Status), res.Config, res.Target, schema, len(res.Findings)) + } + _ = tw.Flush() + + for _, res := range r.Results { + if len(res.Findings) == 0 { + continue + } + _, _ = fmt.Fprintf(w, "\n%s:\n", res.Config) + for _, f := range res.Findings { + _, _ = fmt.Fprintf(w, " %-7s %s %s\n", f.Severity, f.Location.String(), f.Message) + } + } + _, _ = fmt.Fprintf(w, "\nSUMMARY: %d error(s), %d warning(s), %d unpaired, %d duplicate\n", + r.Errors, r.Warnings, r.Unpaired, r.Duplicate) +} + +// decideLintExitCode follows the same threshold matrix as test: errors +// always fail, warnings fail only at --fail-on warn. +func decideLintExitCode(r *LintReport, failOn string) int { + if failOn == failOnNone { + return ExitOK + } + if r.Errors > 0 || r.Unpaired > 0 || r.Duplicate > 0 { + return ExitTestFailure + } + if failOn == failOnWarn && r.Warnings > 0 { + return ExitTestFailure + } + return ExitOK +} diff --git a/internal/cli/lint_test.go b/internal/cli/lint_test.go new file mode 100644 index 0000000..6c20461 --- /dev/null +++ b/internal/cli/lint_test.go @@ -0,0 +1,254 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "bytes" + "os" + "path/filepath" + "strings" + "testing" +) + +// lintTree writes a small platform repository: two nested directories, a +// schema and a config in each. +func lintTree(t *testing.T) string { + t.Helper() + dir := t.TempDir() + mustWrite(t, filepath.Join(dir, "apis", "buckets", "xrd.yaml"), mustRead(t, "../../examples/field-rename/xrd.yaml")) + mustWrite(t, filepath.Join(dir, "apis", "buckets", "conversion.yaml"), mustRead(t, "../../examples/field-rename/xrdconversionconfig.yaml")) + mustWrite(t, filepath.Join(dir, "apis", "widgets", "crd.yaml"), mustRead(t, "../../examples/native-crd/crd.yaml")) + mustWrite(t, filepath.Join(dir, "apis", "widgets", "conversion.yaml"), mustRead(t, "../../examples/native-crd/crdconversionconfig.yaml")) + return dir +} + +func mustRead(t *testing.T, path string) []byte { + t.Helper() + data, err := os.ReadFile(path) + if err != nil { + t.Fatal(err) + } + return data +} + +func mustWrite(t *testing.T, path string, data []byte) { + t.Helper() + if err := os.MkdirAll(filepath.Dir(path), 0o755); err != nil { + t.Fatal(err) + } + if err := os.WriteFile(path, data, 0o600); err != nil { + t.Fatal(err) + } +} + +// The point of the command is one run over a tree, so discovery has to reach +// nested directories and has to recognise both config kinds. +func TestRunLint_DiscoversAndPairsAcrossNestedDirectories(t *testing.T) { + rep, err := RunLint(LintOptions{Paths: []string{lintTree(t)}}) + if err != nil { + t.Fatalf("lint: %v", err) + } + if rep.Configs != 2 || rep.Schemas != 2 { + t.Fatalf("found %d configs and %d schemas, want 2 and 2", rep.Configs, rep.Schemas) + } + kinds := map[string]bool{} + for _, res := range rep.Results { + if res.Status != "ok" { + t.Errorf("%s: status %s, findings %+v", res.Config, res.Status, res.Findings) + } + if res.Schema == "" { + t.Errorf("%s was not paired with a schema", res.Config) + } + kinds[res.Kind] = true + } + if !kinds["XRD"] || !kinds["CRD"] { + t.Errorf("both config kinds should be discovered, got %v", kinds) + } + if rep.Errors != 0 || rep.Unpaired != 0 { + t.Errorf("a clean tree reported %d error(s), %d unpaired", rep.Errors, rep.Unpaired) + } +} + +// A config paired with nothing looks exactly like a config that passed. It +// has to be an error naming what was looked for, never a silent skip — this +// is the failure mode that makes a tool like this untrustworthy. +func TestRunLint_UnpairedConfigIsAnErrorNamingTheTarget(t *testing.T) { + dir := t.TempDir() + mustWrite(t, filepath.Join(dir, "conversion.yaml"), mustRead(t, "../../examples/field-rename/xrdconversionconfig.yaml")) + + rep, err := RunLint(LintOptions{Paths: []string{dir}}) + if err != nil { + t.Fatal(err) + } + if rep.Unpaired != 1 { + t.Fatalf("unpaired = %d, want 1", rep.Unpaired) + } + res := rep.Results[0] + if res.Status != "unpaired" { + t.Errorf("status = %q", res.Status) + } + if len(res.Findings) == 0 || !strings.Contains(res.Findings[0].Message, "xbuckets.example.org") { + t.Errorf("finding should name the target it looked for: %+v", res.Findings) + } + if decideLintExitCode(rep, failOnLoss) == ExitOK { + t.Error("an unpaired config exited 0") + } +} + +// --schema-dir is for repositories that keep schemas apart from configs. +func TestRunLint_SchemaDirPairsAcrossTrees(t *testing.T) { + configs := t.TempDir() + schemas := t.TempDir() + mustWrite(t, filepath.Join(configs, "conversion.yaml"), mustRead(t, "../../examples/field-rename/xrdconversionconfig.yaml")) + mustWrite(t, filepath.Join(schemas, "xrd.yaml"), mustRead(t, "../../examples/field-rename/xrd.yaml")) + + rep, err := RunLint(LintOptions{Paths: []string{configs}, SchemaDirs: []string{schemas}}) + if err != nil { + t.Fatal(err) + } + if rep.Unpaired != 0 || rep.Results[0].Schema == "" { + t.Errorf("--schema-dir did not pair: %+v", rep.Results) + } +} + +// The operator enforces one config per target. Finding that out from an +// admission rejection after merge is what this command exists to prevent. +func TestRunLint_DuplicateTargetsAreReported(t *testing.T) { + dir := lintTree(t) + mustWrite(t, filepath.Join(dir, "apis", "buckets", "conversion-copy.yaml"), mustRead(t, "../../examples/field-rename/xrdconversionconfig.yaml")) + + rep, err := RunLint(LintOptions{Paths: []string{dir}}) + if err != nil { + t.Fatal(err) + } + if rep.Duplicate != 1 { + t.Fatalf("duplicate = %d, want 1 (the second config, not both)", rep.Duplicate) + } + var dup LintPairResult + for _, res := range rep.Results { + if res.Status == "duplicate" { + dup = res + } + } + if len(dup.Findings) == 0 || !strings.Contains(dup.Findings[0].Message, "one config per target") { + t.Errorf("finding should explain what will happen: %+v", dup.Findings) + } +} + +// A platform tree is full of manifests that are none of this tool's +// business. Those must be ignored, not rejected. +func TestRunLint_IgnoresUnrelatedManifests(t *testing.T) { + dir := lintTree(t) + mustWrite(t, filepath.Join(dir, "kustomization.yaml"), []byte("resources:\n - apis/\n")) + mustWrite(t, filepath.Join(dir, "deployment.yaml"), []byte("apiVersion: apps/v1\nkind: Deployment\nmetadata:\n name: x\n")) + mustWrite(t, filepath.Join(dir, "notes.txt"), []byte("not yaml at all")) + + rep, err := RunLint(LintOptions{Paths: []string{dir}}) + if err != nil { + t.Fatalf("an unrelated manifest failed the whole run: %v", err) + } + if rep.Configs != 2 { + t.Errorf("configs = %d, want the 2 real ones", rep.Configs) + } +} + +// Multi-document files are how most GitOps repositories are laid out. +func TestRunLint_HandlesMultiDocumentFiles(t *testing.T) { + dir := t.TempDir() + joined := append(append([]byte{}, mustRead(t, "../../examples/field-rename/xrd.yaml")...), []byte("\n---\n")...) + joined = append(joined, mustRead(t, "../../examples/field-rename/xrdconversionconfig.yaml")...) + mustWrite(t, filepath.Join(dir, "all.yaml"), joined) + + rep, err := RunLint(LintOptions{Paths: []string{dir}}) + if err != nil { + t.Fatal(err) + } + if rep.Configs != 1 || rep.Schemas != 1 { + t.Fatalf("found %d configs, %d schemas in one file, want 1 and 1", rep.Configs, rep.Schemas) + } + if rep.Unpaired != 0 { + t.Errorf("a config and its schema in one file were not paired: %+v", rep.Results) + } +} + +func TestRunLint_ExcludeSkipsPaths(t *testing.T) { + dir := lintTree(t) + rep, err := RunLint(LintOptions{Paths: []string{dir}, Exclude: []string{"widgets"}}) + if err != nil { + t.Fatal(err) + } + if rep.Configs != 1 { + t.Errorf("configs = %d, want the excluded directory skipped", rep.Configs) + } +} + +// Parallel, but the report has to read the same every run or a CI diff of +// two runs is noise. +func TestRunLint_ReportOrderIsDeterministic(t *testing.T) { + dir := lintTree(t) + first, err := RunLint(LintOptions{Paths: []string{dir}, Concurrency: 4}) + if err != nil { + t.Fatal(err) + } + for i := 0; i < 3; i++ { + again, err := RunLint(LintOptions{Paths: []string{dir}, Concurrency: 4}) + if err != nil { + t.Fatal(err) + } + for j := range first.Results { + if again.Results[j].Config != first.Results[j].Config { + t.Fatalf("run %d differs at %d: %s vs %s", i, j, again.Results[j].Config, first.Results[j].Config) + } + } + } +} + +// A broken config has to be reported against the line that broke it, which +// is the whole reason the CI formats exist. +func TestRunLint_FindingsCarryConfigLocations(t *testing.T) { + dir := t.TempDir() + mustWrite(t, filepath.Join(dir, "xrd.yaml"), mustRead(t, "../../examples/crossplane-xr-multiversion/02-add-v2/xrd.yaml")) + mustWrite(t, filepath.Join(dir, "conversion.yaml"), mustRead(t, "../../examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml")) + + rep, err := RunLint(LintOptions{Paths: []string{dir}}) + if err != nil { + t.Fatal(err) + } + findings := rep.Findings() + if len(findings) == 0 { + t.Fatal("a config with an uncovered field produced no findings") + } + for _, f := range findings { + if !strings.HasSuffix(f.Location.File, "conversion.yaml") { + t.Errorf("finding points at %q, not the config", f.Location.File) + } + } + if decideLintExitCode(rep, failOnLoss) == ExitOK { + t.Error("a config with errors exited 0") + } + if decideLintExitCode(rep, failOnNone) != ExitOK { + t.Error("--fail-on none should never fail") + } +} + +func TestLintReport_WriteTableSaysWhenNothingWasFound(t *testing.T) { + var buf bytes.Buffer + (&LintReport{}).WriteTable(&buf) + if !strings.Contains(buf.String(), "No conversion configs found") { + t.Errorf("an empty tree rendered %q", buf.String()) + } +} diff --git a/internal/cli/root.go b/internal/cli/root.go index 29b8436..05582ee 100644 --- a/internal/cli/root.go +++ b/internal/cli/root.go @@ -58,7 +58,7 @@ cluster. Every command works against either resource type: newRetargetCmd(), newCrossplaneCmd(), newConvertCmd(), newSuggestCmd(), newRehubCmd(), newGenerateCmd(), newPatchPreviewCmd(), newMigrateStorageCmd(), newVersionCmd(), - newCompatCmd(), newVersionsCmd(), newPlanCmd(), + newCompatCmd(), newVersionsCmd(), newPlanCmd(), newLintCmd(), ) if err := root.Execute(); err != nil { @@ -816,3 +816,77 @@ Exit codes: 0 a plan was produced, 1 the target state is unreachable, cmd.MarkFlagsMutuallyExclusive("xrd", "crd") return cmd } + +func newLintCmd() *cobra.Command { + var ( + schemaDirs, exclude []string + output, failOn string + concurrency int + ) + cmd := &cobra.Command{ + Use: "lint [path...]", + Short: "Validate every conversion config in a tree against the schema it targets", + Long: `Check a whole repository in one run. + +A platform repo with fifty XRDs otherwise needs fifty invocations, each one +pairing a config with its schema by hand and each producing an exit code the +caller has to aggregate — which in practice means a bash loop in every +consumer's CI, written slightly differently each time. + +lint walks the given paths (default "."), finds every XRDConversionConfig and +CRDConversionConfig by its own apiVersion and kind rather than by filename, +pairs each with the XRD or CRD whose metadata.name it targets, runs the same +checks validate and analyze run, and reports once. + +An unpaired config is an ERROR naming what it looked for, never a silent +skip: a config nothing checked looks exactly like a config that passed, and a +tool that cannot tell you the difference is not worth running. A second +config targeting the same resource is reported the same way — the operator +enforces one config per target, and finding that out from an admission +rejection after merge is what this command exists to prevent. + +Deliberately offline: it constructs no Kubernetes client at all. This is the +check that runs on every commit; test --live is the slow one that runs before +merge. + +Exit codes follow the same matrix as test: 0 clean, 1 findings at or above +the --fail-on threshold, 2 usage error.`, + RunE: func(cmd *cobra.Command, args []string) error { + if err := checkOutputFormat(output, "table", "json", "github", "sarif", "markdown"); err != nil { + return err + } + switch failOn { + case failOnNone, failOnWarn, failOnLoss: + default: + return fmt.Errorf("invalid --fail-on value %q (want none, warn, or loss)", failOn) + } + rep, err := RunLint(LintOptions{ + Paths: args, SchemaDirs: schemaDirs, Exclude: exclude, Concurrency: concurrency, + }) + if err != nil { + return err + } + switch { + case output == "json": + if err := writeJSON(cmd, rep); err != nil { + return err + } + case isCIFormat(output): + if err := writeFindings(cmd, output, rep.Findings(), "convctl lint"); err != nil { + return err + } + default: + rep.WriteTable(cmd.OutOrStdout()) + } + exitCode = decideLintExitCode(rep, failOn) + return nil + }, + } + cmd.Flags().StringSliceVar(&schemaDirs, "schema-dir", nil, "Additional directories to search for XRDs and CRDs (repeatable)") + cmd.Flags().StringSliceVar(&exclude, "exclude", nil, "Glob patterns to skip, matched against the path and its base name (repeatable)") + cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|github|sarif|markdown") + cmd.Flags().StringVar(&failOn, "fail-on", failOnLoss, "Failure threshold: none|warn|loss") + cmd.Flags().IntVar(&concurrency, "concurrency", 0, "Parallel workers (default one per CPU)") + registerOutputCompletions(cmd, "table", "json", "github", "sarif", "markdown") + return cmd +} From 892be1640efd202e414512297e59a2c92a5b9a26 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:04:21 +0300 Subject: [PATCH 03/24] feat(convctl): bounded sampling for --live on large clusters MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --live lists every object of the target type and holds all of them before testing any. On a cluster with tens of thousands of composites a pre-upgrade check is therefore an OOM — and a pre-upgrade check that cannot run is the one situation where the user most needed it. --max-samples caps what gets tested, with --sample-strategy choosing how: first stops listing at the cap; the cheapest, and the only strategy that can stop early random reservoir-samples (Algorithm R) while paginating, so the whole population is represented without ever being held newest the n most recently created, where a schema change shows first Listing now streams into the sampler rather than accumulating and sampling afterwards. Holding forty thousand objects in order to keep fifty of them is the shape of the problem, so the population is counted as it passes while at most the cap is retained. random and newest deliberately do not stop early: both need to see the whole population to be what they claim, and a "uniform" sample of the first page is not uniform. The report says when a run was sampled, in every format it can be read in — the table line, a sampling block in the JSON, and sampled / samplePopulation / sampleTested properties on the JUnit suite. That is the part that matters rather than a detail: a sampled green result that reads like an exhaustive green result is worse than no result, because somebody upgrades on the strength of it, and a JUnit reporter showing fifty green tests with no other context is where that mistake is easiest to make. A population that fits under the cap reports nothing, because it was not sampled. --seed makes random reproducible, so a CI failure can be re-run rather than re-rolled. Uniformity is asserted rather than assumed: the test samples 100 objects 4000 times and checks every item's selection frequency. --namespace narrows a run to one namespace, applied only to the namespaced object class — on a claim-offering XRD the composites are cluster-scoped, so there is nothing to narrow there. What is not done is testing each page as it arrives, which would bound memory with no cap at all. Recorded in limitations.md rather than implied: without --max-samples the objects are still accumulated, and the flag is the answer on the clusters where it matters. Closes #143 Co-Authored-By: Claude Opus 5 (1M context) --- docs/cli.md | 37 ++++++ docs/limitations.md | 1 + internal/cli/live.go | 95 +++++++++++-- internal/cli/report.go | 34 ++++- internal/cli/root.go | 13 ++ internal/cli/sampling.go | 224 +++++++++++++++++++++++++++++++ internal/cli/sampling_test.go | 244 ++++++++++++++++++++++++++++++++++ internal/cli/test.go | 21 ++- 8 files changed, 650 insertions(+), 19 deletions(-) create mode 100644 internal/cli/sampling.go create mode 100644 internal/cli/sampling_test.go diff --git a/docs/cli.md b/docs/cli.md index 793f536..b7a01a8 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -312,6 +312,43 @@ author actually caused. > pipelines should set it; the default is planned to flip in a future > release, and the change will be called out in the release notes. +### Bounded sampling on a large cluster + +`--live` lists and tests every object of the target type. On a cluster with +tens of thousands of composites that is exactly where a pre-upgrade check is +most valuable and least able to run. + +`--max-samples ` caps what gets tested, with `--sample-strategy`: + +| Strategy | Behaviour | Cost | +|---|---|---| +| `first` (default) | stops listing at the cap | cheapest — the only one that can stop early | +| `random` | reservoir-samples while paginating, so the whole population is represented without ever being held | lists everything, holds `n` | +| `newest` | the `n` most recently created objects, where a schema change shows up first | lists everything, holds `n` | + +`random` is reproducible: pass `--seed` and the same objects are chosen, so a +CI failure can be re-run rather than re-rolled. + +**A sampled run says so, in every format.** The table prints it, the JSON +carries a `sampling` block, and the JUnit suite carries `sampled`, +`samplePopulation` and `sampleTested` properties: + +```console +SAMPLED: 50 of 41,204 live object(s), strategy random, seed 7 — this run did NOT cover every object +``` + +That line is the feature. A sampled green result that reads like an +exhaustive green result is worse than no result, because somebody upgrades on +the strength of it — and a JUnit reporter showing fifty green tests is where +that mistake is easiest to make. + +`--namespace` narrows a `--live` run to one namespace. Only the namespaced +object class is affected: on a claim-offering XRD the composites are +cluster-scoped, so there is nothing to narrow on that side. + +Sampling interacts with `--concurrency` only in the obvious way — fewer +samples, less to parallelise. + ### Parallelism and progress Samples are tested in parallel, one worker per available CPU by default. This matters most for `--live`, where the sample set is every object of the target type in the cluster rather than a handful of fixtures. Set `--concurrency N` to pin the worker count (`--concurrency 1` to go fully sequential). diff --git a/docs/limitations.md b/docs/limitations.md index 65adfa6..e001e0c 100755 --- a/docs/limitations.md +++ b/docs/limitations.md @@ -37,6 +37,7 @@ This page is deliberately blunt about what the operator does *not* do today, so - **The webhook-server's cache is built from the feature flags, not from what is on the cluster.** `--enable-xrd-support=false` is what keeps `CompositeResourceDefinition` out of the informer cache entirely. controller-runtime resolves every cached kind through the RESTMapper when the manager is constructed, so on a cluster without Crossplane a replica told to support XRDs does not degrade — it exits at startup with `no matches for kind "CompositeResourceDefinition"`. That is deliberate and matches the manager's own behaviour, but it means the flags must match the cluster rather than describing a preference. - **Required-field analysis treats a required source leaf as present whenever its own parent object is.** Asking "does the rule set always produce this required destination field?" needs the mirror question about the source, and the analysis answers it one level deep: `spec.network.cidr` marked required inside an **optional** `spec.network` counts as guaranteed. Where an optional ancestor gates it, a genuinely conditional source can therefore be read as unconditional, and that case goes unreported. Tightening the rule to "every ancestor must itself be required" is not the fix: nothing under `spec` is ever listed in a CRD root's `required`, so every ordinary rename would be flagged. The sound version has to model whether the destination parent's creation is coupled to the same source path — a rule that creates `spec.network` *only* by writing `cidr` into it is correct, and a naive check calls it broken. Until then this is a missed diagnostic rather than a wrong one, and `--validate-output` and `--fuzz` are the empirical checks that cover it. This is about the **source** side only: on the destination side an optional ancestor is handled — a required field whose containing object is produced whole by a rule, rather than written into, is reported as *unprovable* rather than skipped. - **Required-field satisfaction cannot be proven through `cel`, `jsonPatch` or `scalarToFields`.** The engine checks, for every required field on each side, whether the rule set always produces a value there — and reports an error when nothing does, or when the only rule that does may not fire. It cannot reason about arbitrary expressions, so a required field written by one of those three strategies is reported as *unprovable* (a warning) rather than as satisfied or failed. `convctl test --validate-output` and `--fuzz` are the empirical checks for that residue. A required field whose schema declares a `default` is always satisfied, because the apiserver fills it in. +- **An uncapped `--live` run holds every object it tests.** Listing streams into the sampler, so `--max-samples` genuinely bounds memory — the population is counted without being held. Without a cap the objects are still accumulated before testing, which on a cluster with tens of thousands of large composites is the case `--max-samples` exists for. Testing each page as it arrives would bound it either way and is not implemented; the report is assembled from the full sample set. - **`convctl plan` reads manifests, not a cluster.** It answers "what is the next safe step?" from the XRD/CRD and the conversion config, which is what makes it usable in a PR. Three steps of the sequence have no answer in a manifest — retargeting Compositions, migrating stored objects, pruning `storedVersions` — and are reported as `UNKNOWN` with the command that answers them, never as done, ready, or blocked. A plan therefore stops at *"nothing outstanding that files can decide"*; finishing the migration still requires running those verify commands against the cluster. - **`convctl versions` is XRD-only.** The live inventory — object counts, `storedVersions`, and the field managers still writing each version — is implemented for XRD targets; `--crd` is rejected with a message saying so rather than silently answering a narrower question. `plan`, `compat`, `validate`, `analyze` and `test` all cover both. - **`convctl compat` compares what git can show it.** It resolves the XRD and config at two revisions with `git show`, so it needs both revisions present locally — `fetch-depth: 2` at minimum in CI, and a shallow clone that does not contain the base ref is a usage error rather than a silent pass. It classifies the delta between two *configs and schemas*; it does not read the cluster, so "a served version removed while objects are still stored at it" is judged from `spec.versions`, and the live half of that question belongs to `versions`. diff --git a/internal/cli/live.go b/internal/cli/live.go index 1537c22..367a73c 100644 --- a/internal/cli/live.go +++ b/internal/cli/live.go @@ -105,24 +105,36 @@ func xrdGroupKind(xrd *unstructured.Unstructured) (group, kind string, err error // See fetchLiveSamplesByGVR for why hubVersion specifically, and why // pagination isn't capped. func FetchLiveSamples(ctx context.Context, dyn dynamic.Interface, xrd *unstructured.Unstructured, hubVersion string) ([]Sample, error) { + samples, _, err := FetchLiveSamplesSampled(ctx, dyn, xrd, hubVersion, SamplingOptions{}, "") + return samples, err +} + +// FetchLiveSamplesSampled is FetchLiveSamples with a bound. +// +// The sampler sees every object while paginating — so the population count +// is exact — but holds at most the cap, which is what makes a pre-upgrade +// check runnable on the clusters where it matters most. Without a cap the +// behaviour is unchanged. +func FetchLiveSamplesSampled(ctx context.Context, dyn dynamic.Interface, xrd *unstructured.Unstructured, hubVersion string, sampling SamplingOptions, namespace string) ([]Sample, *SamplingReport, error) { generated, err := xrdadapter.GeneratedCRDNames(xrd) if err != nil { - return nil, err + return nil, nil, err } - var samples []Sample + s := newSampler(sampling) for _, g := range generated { gvr := schema.GroupVersionResource{Group: g.Group, Version: hubVersion, Resource: g.Plural} - got, err := fetchLiveSamplesByGVR(ctx, dyn, gvr, hubVersion) - if err != nil { - return nil, err + // A cluster-scoped composite cannot be narrowed by namespace; only + // the claim side can, which is why this is per generated CRD. + ns := "" + if g.Role == xrdadapter.RoleClaim { + ns = namespace } - for i := range got { - got[i].CRD = g.Name - got[i].CRDRole = string(g.Role) + if err := streamLiveSamples(ctx, dyn, gvr, hubVersion, ns, s, g.Name, string(g.Role)); err != nil { + return nil, nil, err } - samples = append(samples, got...) } - return samples, nil + samples, rep := s.result() + return samples, rep, nil } // FetchLiveSamplesCRD is FetchLiveSamples's sibling for a native @@ -140,6 +152,69 @@ func FetchLiveSamplesCRD(ctx context.Context, dyn dynamic.Interface, crd *extv1. return fetchLiveSamplesByGVR(ctx, dyn, gvr, hubVersion) } +// FetchLiveSamplesCRDSampled is FetchLiveSamplesCRD with a bound. +func FetchLiveSamplesCRDSampled(ctx context.Context, dyn dynamic.Interface, crd *extv1.CustomResourceDefinition, hubVersion string, sampling SamplingOptions, namespace string) ([]Sample, *SamplingReport, error) { + if crd.Spec.Group == "" { + return nil, nil, errors.New("crd is missing spec.group") + } + if crd.Spec.Names.Plural == "" { + return nil, nil, errors.New("crd is missing spec.names.plural") + } + ns := namespace + if crd.Spec.Scope != extv1.NamespaceScoped { + ns = "" + } + gvr := schema.GroupVersionResource{Group: crd.Spec.Group, Version: hubVersion, Resource: crd.Spec.Names.Plural} + s := newSampler(sampling) + if err := streamLiveSamples(ctx, dyn, gvr, hubVersion, ns, s, "", ""); err != nil { + return nil, nil, err + } + samples, rep := s.result() + return samples, rep, nil +} + +// streamLiveSamples paginates and offers each object to the sampler as it +// arrives, rather than accumulating the population and sampling afterwards. +// Holding every object in order to keep fifty of them is the shape of the +// problem this exists to solve. +func streamLiveSamples(ctx context.Context, dyn dynamic.Interface, gvr schema.GroupVersionResource, hubVersion, namespace string, s *sampler, crdName, crdRole string) error { + continueToken := "" + for { + opts := metav1.ListOptions{Continue: continueToken, Limit: 200} + var ( + list *unstructured.UnstructuredList + err error + ) + if namespace != "" { + list, err = dyn.Resource(gvr).Namespace(namespace).List(ctx, opts) + } else { + list, err = dyn.Resource(gvr).List(ctx, opts) + } + if err != nil { + return fmt.Errorf("listing %s (version %s): %w", gvr.GroupResource().String(), hubVersion, err) + } + for i := range list.Items { + item := list.Items[i] + s.add(Sample{ + File: "cluster:" + objectLabel(&item), + Object: item.Object, + Version: versionFromAPIVersion(item.GetAPIVersion()), + CRD: crdName, + CRDRole: crdRole, + }, &item) + if s.full() { + // Only the "first" strategy can stop early; the others + // need the whole population to be what they claim. + return nil + } + } + continueToken = list.GetContinue() + if continueToken == "" { + return nil + } + } +} + // fetchLiveSamplesByGVR is FetchLiveSamples/FetchLiveSamplesCRD's shared // pagination loop, listing every existing instance of gvr at hubVersion — // the storage/referenceable version, which the apiserver always serves diff --git a/internal/cli/report.go b/internal/cli/report.go index 0ecaf77..7ba4543 100644 --- a/internal/cli/report.go +++ b/internal/cli/report.go @@ -21,6 +21,7 @@ import ( "fmt" "io" "sort" + "strconv" "strings" "text/tabwriter" "time" @@ -150,6 +151,11 @@ type Report struct { // a fuzz failure nobody can reproduce is noise, and the seed is // the whole reproduction. Fuzz *FuzzMeta `json:"fuzz,omitempty"` + // Sampling is present only when a --live run was bounded. Its + // absence means every live object was tested; a sampled result + // that reads like an exhaustive one is the failure this guards + // against. + Sampling *SamplingReport `json:"sampling,omitempty"` // Scope is the detected Crossplane scope (XRD targets only) — // which injected-field set is in play, and how much the resolver // trusts the answer. See pkg/xrdadapter.ResolveScope. @@ -194,6 +200,9 @@ func (r *Report) WriteTable(w io.Writer) { _, _ = fmt.Fprintf(w, "FUZZ: %d generated object(s), seed %d — reproduce with --fuzz %d --seed %d\n\n", r.Meta.Fuzz.Objects, r.Meta.Fuzz.Seed, r.Meta.Fuzz.Objects, r.Meta.Fuzz.Seed) } + if r.Meta.Sampling != nil { + _, _ = fmt.Fprintf(w, "%s\n\n", r.Meta.Sampling.String()) + } r.Golden.write(w) tw := tabwriter.NewWriter(w, 0, 4, 2, ' ', 0) @@ -275,12 +284,13 @@ type junitTestSuites struct { } type junitTestSuite struct { - Name string `xml:"name,attr"` - Tests int `xml:"tests,attr"` - Failures int `xml:"failures,attr"` - Errors int `xml:"errors,attr"` - Time string `xml:"time,attr"` - Cases []junitTestCase `xml:"testcase"` + Name string `xml:"name,attr"` + Tests int `xml:"tests,attr"` + Failures int `xml:"failures,attr"` + Errors int `xml:"errors,attr"` + Time string `xml:"time,attr"` + Props *junitProperties `xml:"properties,omitempty"` + Cases []junitTestCase `xml:"testcase"` } type junitTestCase struct { @@ -318,6 +328,18 @@ type junitMessage struct { // failed testcases still catch exactly what --fail-on would. func (r *Report) junitSuite() junitTestSuite { suite := junitTestSuite{Name: fmt.Sprintf("%s/%s", r.Meta.ResourceKind, r.Meta.Resource)} + // A sampled run has to say so in every format it can be read in. A + // JUnit reporter showing 50 green tests, from a population of 40,000, + // with nothing saying which, is the exact false confidence the cap + // exists to make explicit. + if r.Meta.Sampling != nil { + suite.Props = &junitProperties{Properties: []junitProperty{ + {Name: "sampled", Value: "true"}, + {Name: "sampleStrategy", Value: r.Meta.Sampling.Strategy}, + {Name: "samplePopulation", Value: strconv.Itoa(r.Meta.Sampling.Population)}, + {Name: "sampleTested", Value: strconv.Itoa(r.Meta.Sampling.Tested)}, + }} + } var totalTime float64 for _, s := range r.Samples { for _, p := range s.Paths { diff --git a/internal/cli/root.go b/internal/cli/root.go index 05582ee..d5705fd 100644 --- a/internal/cli/root.go +++ b/internal/cli/root.go @@ -198,6 +198,8 @@ func newTestCmd() *cobra.Command { kubeconfig, kubeContext, kubeconfigDir string contexts []string concurrency int + maxSamples int + sampleStrategy, namespace string ) cmd := &cobra.Command{ Use: "test", @@ -257,6 +259,12 @@ results are collected by sample index, never by completion order.`, if err := checkOutputFormat(output, "table", "json", "junit", "github", "sarif", "markdown"); err != nil { return err } + if err := ValidateSamplingOptions(SamplingOptions{MaxSamples: maxSamples, Strategy: sampleStrategy, Seed: fuzzSeed}); err != nil { + return err + } + if (maxSamples > 0 || sampleStrategy != "" || namespace != "") && !live { + return errors.New("--max-samples, --sample-strategy and --namespace only apply to --live runs") + } switch failOn { case failOnNone, failOnWarn, failOnLoss: default: @@ -283,6 +291,8 @@ results are collected by sample index, never by completion order.`, Live: live, Kubeconfig: kubeconfig, KubeContext: kubeContext, Contexts: contexts, KubeconfigDir: kubeconfigDir, Concurrency: concurrency, Quiet: quiet, + Sampling: SamplingOptions{MaxSamples: maxSamples, Strategy: sampleStrategy, Seed: fuzzSeed}, + Namespace: namespace, VerifyPropagation: verifyPropagation, ValidateOutput: validateOutput, RecordDir: recordDir, @@ -344,6 +354,9 @@ results are collected by sample index, never by completion order.`, cmd.Flags().StringVar(&failOn, "fail-on", failOnLoss, "Exit-code threshold: none|warn|loss") cmd.Flags().StringSliceVar(&versionPairs, "version-pair", nil, "Restrict testing to these version(s), repeatable") cmd.Flags().IntVar(&concurrency, "concurrency", 0, "Number of samples to test in parallel (default: one per available CPU)") + cmd.Flags().IntVar(&maxSamples, "max-samples", 0, "With --live, cap how many objects are tested (default: every object). The report says so when a run was sampled") + cmd.Flags().StringVar(&sampleStrategy, "sample-strategy", "", "With --max-samples: first (cheapest), random (uniform, reproducible with --seed), or newest (default: first)") + cmd.Flags().StringVar(&namespace, "namespace", "", "With --live, narrow to one namespace. Only namespaced object classes are affected; a cluster-scoped composite cannot be narrowed") cmd.Flags().BoolVar(&quiet, "quiet", false, "Suppress the progress line written to stderr") cmd.Flags().BoolVar(&verifyPropagation, "verify-propagation", false, "With --live on an XRD, also check that every CRD Crossplane generates from it actually carries the conversion webhook the XRD points at") cmd.Flags().IntVar(&fuzzN, "fuzz", 0, "Generate N schema-valid objects from the hub version's own schema and test them too. Biased toward the boundaries fixtures miss: empty arrays, absent optionals, length and range limits, first and last enum members") diff --git a/internal/cli/sampling.go b/internal/cli/sampling.go new file mode 100644 index 0000000..9865db8 --- /dev/null +++ b/internal/cli/sampling.go @@ -0,0 +1,224 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "fmt" + "math/rand" + "sort" + + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" +) + +// Sampling strategies for --live on a large cluster. +const ( + // SampleFirst stops listing once the cap is reached. Cheapest, and + // biased toward whatever the apiserver returns first. + SampleFirst = "first" + // SampleRandom reservoir-samples while paginating, so the whole + // population is represented without ever holding it. + SampleRandom = "random" + // SampleNewest keeps the n most recently created objects, which is + // where a schema change's effects show up first. + SampleNewest = "newest" +) + +// SamplingOptions bounds a --live run. +type SamplingOptions struct { + // MaxSamples caps the objects tested. Zero means no cap. + MaxSamples int + // Strategy is one of the constants above. Empty means SampleFirst. + Strategy string + // Seed makes SampleRandom reproducible, so a CI failure can be + // re-run. Zero means a fixed default rather than a random one: a fuzz + // or sample failure nobody can reproduce is noise. + Seed int64 +} + +func (o SamplingOptions) enabled() bool { return o.MaxSamples > 0 } + +func (o SamplingOptions) strategy() string { + if o.Strategy == "" { + return SampleFirst + } + return o.Strategy +} + +// ValidateSamplingOptions rejects a strategy the sampler does not implement, +// rather than silently falling back to one that samples differently. +func ValidateSamplingOptions(o SamplingOptions) error { + if o.MaxSamples < 0 { + return fmt.Errorf("--max-samples must not be negative, got %d", o.MaxSamples) + } + switch o.strategy() { + case SampleFirst, SampleRandom, SampleNewest: + default: + return fmt.Errorf("invalid --sample-strategy %q (want first, random, or newest)", o.Strategy) + } + if o.Strategy != "" && o.MaxSamples == 0 { + return fmt.Errorf("--sample-strategy %s has no effect without --max-samples", o.Strategy) + } + return nil +} + +// SamplingReport records that a run was sampled, and from what. +// +// This is the part that matters: a sampled green result that looks like an +// exhaustive green result is worse than no result at all, because somebody +// upgrades on the strength of it. +type SamplingReport struct { + Strategy string `json:"strategy"` + // Cap is the requested maximum. + Cap int `json:"cap"` + // Population is how many objects exist, counted while paginating even + // when most were never held. + Population int `json:"population"` + // Tested is how many were actually sampled. + Tested int `json:"tested"` + // Seed is present for the random strategy, so the run is repeatable. + Seed int64 `json:"seed,omitempty"` +} + +func (s *SamplingReport) String() string { + if s == nil { + return "" + } + seed := "" + if s.Strategy == SampleRandom { + seed = fmt.Sprintf(", seed %d", s.Seed) + } + return fmt.Sprintf("SAMPLED: %d of %d live object(s), strategy %s%s — this run did NOT cover every object", + s.Tested, s.Population, s.Strategy, seed) +} + +// sampler accumulates at most MaxSamples objects out of a stream, counting +// the whole population as it goes. +type sampler struct { + opts SamplingOptions + rnd *rand.Rand + + seen int + kept []Sample + // order holds each kept sample's creation timestamp for SampleNewest, + // parallel to kept. + order []string +} + +func newSampler(opts SamplingOptions) *sampler { + seed := opts.Seed + if seed == 0 { + seed = 1 + } + return &sampler{ + opts: opts, + // #nosec G404 -- reproducibility from --seed is the point; this + // selects which objects to test, not anything secret. + rnd: rand.New(rand.NewSource(seed)), + } +} + +// full reports whether listing can stop early. Only the first strategy can +// stop: the other two need to see the whole population to be what they +// claim. +func (s *sampler) full() bool { + return s.opts.enabled() && s.opts.strategy() == SampleFirst && len(s.kept) >= s.opts.MaxSamples +} + +// add offers one object to the sample. +func (s *sampler) add(sample Sample, obj *unstructured.Unstructured) { + s.seen++ + if !s.opts.enabled() { + s.kept = append(s.kept, sample) + return + } + switch s.opts.strategy() { + case SampleFirst: + if len(s.kept) < s.opts.MaxSamples { + s.kept = append(s.kept, sample) + } + case SampleRandom: + // Algorithm R. The i-th item (1-based) replaces a uniformly chosen + // slot with probability n/i, which leaves every item of the + // population equally likely to be kept — without ever holding more + // than n of them. + if len(s.kept) < s.opts.MaxSamples { + s.kept = append(s.kept, sample) + return + } + if j := s.rnd.Intn(s.seen); j < s.opts.MaxSamples { + s.kept[j] = sample + } + case SampleNewest: + ts := "" + if obj != nil { + ts = obj.GetCreationTimestamp().UTC().Format("2006-01-02T15:04:05Z") + } + s.kept = append(s.kept, sample) + s.order = append(s.order, ts) + if len(s.kept) > s.opts.MaxSamples { + s.trimOldest() + } + } +} + +// trimOldest keeps the window at the cap by dropping the oldest entry, so +// the newest strategy holds n rather than the population. +func (s *sampler) trimOldest() { + oldest := 0 + for i := 1; i < len(s.order); i++ { + if s.order[i] < s.order[oldest] { + oldest = i + } + } + s.kept = append(s.kept[:oldest], s.kept[oldest+1:]...) + s.order = append(s.order[:oldest], s.order[oldest+1:]...) +} + +// result returns the samples and, when the run was actually sampled, the +// report that says so. +func (s *sampler) result() ([]Sample, *SamplingReport) { + kept := s.kept + if s.opts.enabled() && s.opts.strategy() == SampleNewest { + // Newest first, so a truncated report still reads in a meaningful + // order, and deterministically for equal timestamps. + idx := make([]int, len(kept)) + for i := range idx { + idx[i] = i + } + sort.SliceStable(idx, func(a, b int) bool { + if s.order[idx[a]] != s.order[idx[b]] { + return s.order[idx[a]] > s.order[idx[b]] + } + return kept[idx[a]].File < kept[idx[b]].File + }) + sorted := make([]Sample, 0, len(kept)) + for _, i := range idx { + sorted = append(sorted, kept[i]) + } + kept = sorted + } + if !s.opts.enabled() || s.seen <= len(kept) { + return kept, nil + } + return kept, &SamplingReport{ + Strategy: s.opts.strategy(), + Cap: s.opts.MaxSamples, + Population: s.seen, + Tested: len(kept), + Seed: s.opts.Seed, + } +} diff --git a/internal/cli/sampling_test.go b/internal/cli/sampling_test.go new file mode 100644 index 0000000..bfe2529 --- /dev/null +++ b/internal/cli/sampling_test.go @@ -0,0 +1,244 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "fmt" + "strings" + "testing" + "time" + + metav1 "k8s.io/apimachinery/pkg/apis/meta/v1" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" +) + +// feed offers n objects to a sampler, oldest first, the way a paginated +// list would. +func feed(s *sampler, n int) { + base := time.Date(2026, 1, 1, 0, 0, 0, 0, time.UTC) + for i := 0; i < n; i++ { + obj := &unstructured.Unstructured{Object: map[string]any{}} + obj.SetName(fmt.Sprintf("obj-%04d", i)) + obj.SetCreationTimestamp(metav1.NewTime(base.Add(time.Duration(i) * time.Minute))) + s.add(Sample{File: "cluster:" + obj.GetName()}, obj) + } +} + +// No cap means the behaviour nobody asked to change: every object, and no +// claim that the run was sampled. +func TestSampler_UncappedKeepsEverythingAndReportsNothing(t *testing.T) { + s := newSampler(SamplingOptions{}) + feed(s, 500) + kept, rep := s.result() + if len(kept) != 500 { + t.Errorf("kept %d, want all 500", len(kept)) + } + if rep != nil { + t.Errorf("an uncapped run claimed to be sampled: %+v", rep) + } +} + +// "first" is the one strategy that can stop listing early, which is the +// whole reason it is the cheapest. +func TestSampler_FirstStopsEarly(t *testing.T) { + s := newSampler(SamplingOptions{MaxSamples: 10, Strategy: SampleFirst}) + for i := 0; i < 10; i++ { + if s.full() { + t.Fatalf("full at %d, before the cap", i) + } + s.add(Sample{File: fmt.Sprintf("o%d", i)}, nil) + } + if !s.full() { + t.Fatal("not full at the cap, so listing would continue pointlessly") + } + kept, _ := s.result() + if len(kept) != 10 || kept[0].File != "o0" { + t.Errorf("kept %d starting at %q, want the first 10", len(kept), kept[0].File) + } +} + +// The other two need the whole population to be what they claim, so they +// must never stop listing early. +func TestSampler_RandomAndNewestNeverStopEarly(t *testing.T) { + for _, strategy := range []string{SampleRandom, SampleNewest} { + s := newSampler(SamplingOptions{MaxSamples: 5, Strategy: strategy}) + feed(s, 50) + if s.full() { + t.Errorf("%s reported full, which would truncate the population it samples from", strategy) + } + } +} + +// Reservoir sampling has to be uniform, or "random" is a word rather than a +// property. Every item should appear at roughly cap/population frequency. +func TestSampler_RandomIsUniform(t *testing.T) { + const population, keep, runs = 100, 10, 4000 + counts := make([]int, population) + for r := 0; r < runs; r++ { + s := newSampler(SamplingOptions{MaxSamples: keep, Strategy: SampleRandom, Seed: int64(r + 1)}) + feed(s, population) + kept, _ := s.result() + if len(kept) != keep { + t.Fatalf("kept %d, want %d", len(kept), keep) + } + for _, k := range kept { + var idx int + if _, err := fmt.Sscanf(k.File, "cluster:obj-%04d", &idx); err == nil { + counts[idx]++ + } + } + } + // Expected selections per item: runs * keep / population = 400. + expected := runs * keep / population + for i, c := range counts { + if c < expected/2 || c > expected*2 { + t.Errorf("item %d selected %d times, expected around %d — the reservoir is biased", i, c, expected) + } + } +} + +// A CI failure nobody can reproduce is noise, so the same seed has to +// produce the same sample. +func TestSampler_RandomIsReproducibleUnderASeed(t *testing.T) { + sample := func() []string { + s := newSampler(SamplingOptions{MaxSamples: 8, Strategy: SampleRandom, Seed: 42}) + feed(s, 200) + kept, _ := s.result() + var names []string + for _, k := range kept { + names = append(names, k.File) + } + return names + } + a, b := sample(), sample() + if strings.Join(a, ",") != strings.Join(b, ",") { + t.Errorf("same seed produced different samples:\n%v\n%v", a, b) + } + s := newSampler(SamplingOptions{MaxSamples: 8, Strategy: SampleRandom, Seed: 43}) + feed(s, 200) + other, _ := s.result() + var names []string + for _, k := range other { + names = append(names, k.File) + } + if strings.Join(a, ",") == strings.Join(names, ",") { + t.Error("two different seeds produced identical samples") + } +} + +// "newest" has to hold the cap, not the population — that is the point — +// and still end up with the newest ones. +func TestSampler_NewestKeepsTheLatestAndBoundsWhatItHolds(t *testing.T) { + s := newSampler(SamplingOptions{MaxSamples: 5, Strategy: SampleNewest}) + feed(s, 100) + if len(s.kept) > 5 { + t.Errorf("held %d objects, want at most the cap of 5", len(s.kept)) + } + kept, rep := s.result() + if len(kept) != 5 { + t.Fatalf("kept %d, want 5", len(kept)) + } + // Newest first: obj-0099 down to obj-0095. + for i, want := range []string{"cluster:obj-0099", "cluster:obj-0098", "cluster:obj-0097", "cluster:obj-0096", "cluster:obj-0095"} { + if kept[i].File != want { + t.Errorf("kept[%d] = %q, want %q", i, kept[i].File, want) + } + } + if rep == nil || rep.Population != 100 || rep.Tested != 5 { + t.Errorf("report = %+v, want population 100 and 5 tested", rep) + } +} + +// A sampled green result that reads like an exhaustive green result is +// worse than no result, so the report has to say so in words. +func TestSamplingReport_SaysItDidNotCoverEverything(t *testing.T) { + s := newSampler(SamplingOptions{MaxSamples: 3, Strategy: SampleRandom, Seed: 7}) + feed(s, 40) + _, rep := s.result() + if rep == nil { + t.Fatal("a capped run that saw more than the cap reported nothing") + } + msg := rep.String() + for _, want := range []string{"SAMPLED", "3 of 40", "random", "seed 7", "did NOT cover every object"} { + if !strings.Contains(msg, want) { + t.Errorf("report line is missing %q: %s", want, msg) + } + } +} + +// A population that fits under the cap was not sampled, and saying it was +// would be its own kind of wrong. +func TestSampler_NoReportWhenThePopulationFitsUnderTheCap(t *testing.T) { + s := newSampler(SamplingOptions{MaxSamples: 100, Strategy: SampleRandom, Seed: 1}) + feed(s, 10) + kept, rep := s.result() + if len(kept) != 10 { + t.Errorf("kept %d, want 10", len(kept)) + } + if rep != nil { + t.Errorf("claimed to have sampled a population that fit: %+v", rep) + } +} + +// A strategy the sampler does not implement must be rejected, not silently +// replaced by one that samples differently. +func TestValidateSamplingOptions(t *testing.T) { + if err := ValidateSamplingOptions(SamplingOptions{MaxSamples: 10, Strategy: SampleNewest}); err != nil { + t.Errorf("a valid combination was rejected: %v", err) + } + if err := ValidateSamplingOptions(SamplingOptions{MaxSamples: 10, Strategy: "latest"}); err == nil { + t.Error("an unimplemented strategy was accepted") + } + if err := ValidateSamplingOptions(SamplingOptions{Strategy: SampleRandom}); err == nil { + t.Error("a strategy with no cap was accepted, where it would do nothing") + } + if err := ValidateSamplingOptions(SamplingOptions{MaxSamples: -1}); err == nil { + t.Error("a negative cap was accepted") + } +} + +// The JUnit reporter is where a sampled run is most likely to be mistaken +// for an exhaustive one: a wall of green test cases with no other context. +func TestReport_JUnitRecordsThatARunWasSampled(t *testing.T) { + rep := &Report{} + rep.Meta.ResourceKind = "XRD" + rep.Meta.Resource = "xthings.example.org" + rep.Meta.Sampling = &SamplingReport{Strategy: SampleRandom, Cap: 50, Population: 40000, Tested: 50, Seed: 9} + suite := rep.junitSuite() + if suite.Props == nil { + t.Fatal("JUnit suite carries no sampling properties") + } + got := map[string]string{} + for _, p := range suite.Props.Properties { + got[p.Name] = p.Value + } + if got["sampled"] != "true" || got["samplePopulation"] != "40000" || got["sampleTested"] != "50" { + t.Errorf("properties = %v, want the population and tested counts", got) + } +} + +// And in the human-readable report. +func TestReport_TableSaysARunWasSampled(t *testing.T) { + rep := &Report{} + rep.Meta.ResourceKind = "XRD" + rep.Meta.Sampling = &SamplingReport{Strategy: SampleFirst, Cap: 10, Population: 900, Tested: 10} + var sb strings.Builder + rep.WriteTable(&sb) + if !strings.Contains(sb.String(), "SAMPLED: 10 of 900") { + t.Errorf("table does not report the sampling:\n%s", sb.String()) + } +} diff --git a/internal/cli/test.go b/internal/cli/test.go index 62fb4ae..e62120b 100644 --- a/internal/cli/test.go +++ b/internal/cli/test.go @@ -64,6 +64,13 @@ type TestOptions struct { Concurrency int // Quiet suppresses the progress line written to stderr. Quiet bool + // Sampling bounds a --live run on a cluster whose population does not + // fit in memory. Zero-valued means every object, as before. + Sampling SamplingOptions + // Namespace narrows a --live run to one namespace. Only the namespaced + // object class is affected: on a claim-offering XRD the composites are + // cluster-scoped, so narrowing them is not a thing that exists. + Namespace string // ValidateOutput additionally validates every converted object against // the destination version's own schema, using the apiextensions @@ -208,6 +215,7 @@ func runTestXRD(opts TestOptions) (*Report, error) { return nil, fmt.Errorf("configuration is structurally invalid: %w", err) } var samples []Sample + var sampling *SamplingReport // Which fields Crossplane injects, and where they sit, is decided by // the XRD's scope — so report it alongside the results rather than // making an author infer it, and say so when the manifest does not @@ -228,7 +236,7 @@ func runTestXRD(opts TestOptions) (*Report, error) { return nil, fmt.Errorf("verifying conversion propagation: %w", perr) } } - samples, err = FetchLiveSamples(context.Background(), dyn, xrd, cfg.Spec.HubVersion) + samples, sampling, err = FetchLiveSamplesSampled(context.Background(), dyn, xrd, cfg.Spec.HubVersion, opts.Sampling, opts.Namespace) if err != nil { return nil, fmt.Errorf("fetching live samples: %w", err) } @@ -288,6 +296,7 @@ func runTestXRD(opts TestOptions) (*Report, error) { if err != nil { return nil, err } + rep.Meta.Sampling = sampling rep.Meta.Scope = scopeView(scope) rep.Propagation = propagation return rep, nil @@ -308,12 +317,13 @@ func runTestCRD(opts TestOptions) (*Report, error) { return nil, fmt.Errorf("configuration is structurally invalid: %w", err) } var samples []Sample + var sampling *SamplingReport if opts.Live { dyn, err := buildDynamicClient(KubeOptions{Kubeconfig: opts.Kubeconfig, Context: opts.KubeContext}) if err != nil { return nil, err } - samples, err = FetchLiveSamplesCRD(context.Background(), dyn, crd, cfg.Spec.HubVersion) + samples, sampling, err = FetchLiveSamplesCRDSampled(context.Background(), dyn, crd, cfg.Spec.HubVersion, opts.Sampling, opts.Namespace) if err != nil { return nil, fmt.Errorf("fetching live samples: %w", err) } @@ -357,7 +367,12 @@ func runTestCRD(opts TestOptions) (*Report, error) { } // A native CRD's authored schema is the whole schema — nothing is // injected behind the author's back — so there is nothing to strip. - return runTestCommon(opts, "CRD", crdName(crd), cfg.Name, cfg.Spec.HubVersion, samples, versions, report, router, nil, start) + rep, err := runTestCommon(opts, "CRD", crdName(crd), cfg.Name, cfg.Spec.HubVersion, samples, versions, report, router, nil, start) + if err != nil { + return nil, err + } + rep.Meta.Sampling = sampling + return rep, nil } // runTestCommon is runTestXRD/runTestCRD's shared tail: exercising every From 83f35863b7fd6e97c7e9e0a095e393c2054a2b43 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:06:15 +0300 Subject: [PATCH 04/24] ci: publish the convctl container image MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit CI already builds a convctl image in the docker-build-check matrix, so the Dockerfile's COMPONENT=convctl path is known to work — but the release workflow publishes only manager and webhook-server. Any container-based pipeline therefore has nothing to use: Tekton, Argo Workflows, GitLab's image:, a GitHub container: job, or anyone who would simply rather pin a digest than download a binary. Adding it to the release matrix is all it takes. Everything else already generalises per matrix entry: the Dockerfile takes COMPONENT, and signing, the SBOM, provenance and the digest recording are all per-entry, so the image appears in the release notes' signed-artifact table without touching the addendum script. Two details the issue asked to get right rather than assume, both checked against a locally built image: - distroless/static:nonroot does include /etc/ssl/certs/ca-certificates.crt, so --live can reach an HTTPS apiserver from inside the image. Verified by exporting the image and looking, not by trusting the base's reputation. - docker run test --config ... --samples ... against a mounted directory works as written, because ENTRYPOINT is the binary. The base stays distroless rather than gaining a shell. That means commands cannot be chained inside the container, which is a real cost in a CI system that expects to run a script — so it is documented as such, with the alternative (the released binary on a normal runner image) named. A shell-bearing variant would mean a second base to patch for a convenience that has a workaround; the operator images' base is not touched either way. Closes #139 Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/release.yml | 7 +++++++ docs/cli.md | 13 +++++++++++++ docs/installation.md | 35 +++++++++++++++++++++++++++++++++++ 3 files changed, 55 insertions(+) diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 024d386..31e54c9 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -22,6 +22,7 @@ env: # matrix.image below. MANAGER_IMAGE: ghcr.io/terasky-oss/declarative-conversion-operator WEBHOOK_SERVER_IMAGE: ghcr.io/terasky-oss/declarative-conversion-webhook-server + CONVCTL_IMAGE: ghcr.io/terasky-oss/declarative-conversion-convctl CHART_REPO: ghcr.io/terasky-oss/charts jobs: @@ -46,6 +47,12 @@ jobs: image: ghcr.io/terasky-oss/declarative-conversion-operator - component: webhook-server image: ghcr.io/terasky-oss/declarative-conversion-webhook-server + # The CLI ships as an image too, for pipelines that would rather + # pin a digest than download a binary: Tekton, Argo Workflows, + # GitLab's image:, a GitHub container: job. Same Dockerfile, same + # distroless base, same signing and attestations. + - component: convctl + image: ghcr.io/terasky-oss/declarative-conversion-convctl steps: - name: Checkout uses: actions/checkout@v7 diff --git a/docs/cli.md b/docs/cli.md index b7a01a8..68beaf3 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -26,6 +26,19 @@ convctl crossplane status [-o table|json] Roughly in the order you reach for them while authoring a mapping: `suggest` drafts rules for fields nothing covers yet, `validate` and `analyze` check the config statically, `convert` shows what a single object turns into, `test` grades fixtures or every live object, `diff` reports what a config edit changed, and `patch-preview` shows the exact patch the operator will apply once you commit. After a hub/storage-version promotion, `migrate-storage` rewrites live objects (critical for native CRDs; on XRDs the `compositionRef` retarget usually already did, and the remaining job is pruning `storedVersions`). For a GitOps hub flip, `generate kyverno` drafts MutatingPolicies that retarget existing XRs without a per-object name patch; on a cluster without Kyverno, `retarget` does the same job directly. `crossplane status` answers "where is my migration right now?" without assembling it from half a dozen `kubectl` invocations. Around all of it, `plan` sequences the migration, `versions` answers whether an old version can be retired yet, and `compat` gates config edits in review. +## Running it in a container + +```console +docker run --rm -v "$PWD:/work" -w /work \ + ghcr.io/terasky-oss/declarative-conversion-convctl:v0.5.0 \ + lint ./platform/ +``` + +Published on every release for `linux/amd64` and `linux/arm64`, signed and +attested like the operator images. It is distroless and has no shell, so run +one `convctl` invocation per step rather than chaining — see +[Installation: the `convctl` container image](installation.md#the-convctl-container-image). + ## `convctl validate` Runs the same static checks the admission webhook performs, offline. diff --git a/docs/installation.md b/docs/installation.md index 0c40357..7e40d16 100644 --- a/docs/installation.md +++ b/docs/installation.md @@ -156,3 +156,38 @@ For Flux or Argo, use the [GitOps operator sync](gitops/operator-sync.md) examples (`examples/gitops/flux`, `examples/gitops/argo`). Keep `driftPolicy: KeepServingStale` on GitOps-managed configs — `FailClosed` drops conversions while the schema and config reconcile independently. + +## The `convctl` container image + +For pipelines that would rather pin a digest than download a binary — Tekton, +Argo Workflows, GitLab's `image:`, a GitHub `container:` job — the CLI is +published alongside the two operator images: + +```console +docker run --rm -v "$PWD:/work" -w /work \ + ghcr.io/terasky-oss/declarative-conversion-convctl:v0.5.0 \ + test --xrd xrd.yaml --config xrdconversionconfig.yaml --samples ./samples/ +``` + +`linux/amd64` and `linux/arm64`, cosign-signed with build-provenance and SBOM +attestations exactly like the other two — the digest is listed in each +release's signed-artifact table, and the same `cosign verify` invocation +applies. + +### Base image: distroless, and what that costs you + +The CLI image uses the same `gcr.io/distroless/static:nonroot` base as the +operator, deliberately: + +- **`--live` works.** The base includes `/etc/ssl/certs/ca-certificates.crt`, + so TLS to an HTTPS apiserver verifies. (Checked, not assumed.) +- **There is no shell.** `ENTRYPOINT` is the binary, so + `docker run test …` reads naturally — but you cannot chain commands + inside the container, and `sh -c` is not available. In a CI system that + expects to run a script in the container, either run one `convctl` + invocation per step, or use the released binary with + the released binary on a normal runner image. + +A shell-bearing variant was considered and not published: two images means +two bases to patch, and the operator's base stays as it is regardless. + From 51dd9902d179476909a59a95eab8d6c196be1832 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:11:20 +0300 Subject: [PATCH 05/24] feat(release): Homebrew, Scoop, deb/rpm, and an honest version command MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit convctl is installable by downloading a tarball or by go install. Neither is how anyone installs a kubectl-adjacent tool, and go install produces a binary whose version reports "dev" because the ldflags only exist on a release build — so a bug report arrives with a version nobody can map to a commit. GoReleaser already builds the archives, so this is mostly configuration: nfpms for deb and rpm, scoops for the Windows archives already being built, and Homebrew. Validated with `goreleaser check` and proved end to end with a snapshot build: four packages, a cask and a Scoop manifest, all produced. Three things the configuration had to get right rather than assume: - `brews:` is deprecated as of GoReleaser v2.17 in favour of `homebrew_casks:`, which is macOS-only. Linuxbrew therefore loses the formula, which the docs say plainly rather than implying coverage that does not exist — the deb, rpm and archive are the Linux answers. - A cask that installs an unsigned binary hits Gatekeeper, and the failure reads as "convctl is damaged" rather than as a policy decision. A post hook clears the quarantine attribute. - Both publishing targets need a token that can write to another repository, which GITHUB_TOKEN cannot. They are referenced as `index .Env "X"` rather than `.Env.X` — the dotted form errors when the variable is unset — and skip themselves when it is empty. A release must not fail because a tap repository does not exist yet. `convctl version -o json` reports version, commit, build date, Go version and platform. Absent release ldflags it falls back to what the Go toolchain embeds in every module build: the module version and the VCS stamps, with a -dirty suffix when the tree was not clean. A locally built binary now identifies itself as v0.3.1-0.20260915200615-83f35863b7fd rather than "dev". krew is deliberately not a target. It was in the original scope and is dropped: this is not a kubectl plugin, and the krew-index review cycle is weeks of process for a distribution channel nobody asked for. Closes #142 Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/release.yml | 7 +++ .goreleaser.yaml | 63 +++++++++++++++++++- cmd/convctl/main.go | 13 +++- docs/installation.md | 39 ++++++++++++ internal/cli/root.go | 23 +++++++- internal/cli/version.go | 108 ++++++++++++++++++++++++++++++++++ internal/cli/version_test.go | 86 +++++++++++++++++++++++++++ 7 files changed, 334 insertions(+), 5 deletions(-) create mode 100644 internal/cli/version.go create mode 100644 internal/cli/version_test.go diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 31e54c9..db8b4a0 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -263,6 +263,13 @@ jobs: args: release --clean env: GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }} + # Publishing to the tap and bucket repositories needs a token that + # can write to them, which GITHUB_TOKEN cannot. Unset is fine and + # deliberate: the Homebrew and Scoop steps skip themselves rather + # than failing the release, so the binaries, packages and images + # still ship. + HOMEBREW_TAP_TOKEN: ${{ secrets.HOMEBREW_TAP_TOKEN }} + SCOOP_BUCKET_TOKEN: ${{ secrets.SCOOP_BUCKET_TOKEN }} - name: Attest CLI artifact build provenance uses: actions/attest-build-provenance@v4 diff --git a/.goreleaser.yaml b/.goreleaser.yaml index fe1207b..3c760d8 100644 --- a/.goreleaser.yaml +++ b/.goreleaser.yaml @@ -15,7 +15,12 @@ builds: goos: [linux, darwin, windows] goarch: [amd64, arm64] ldflags: - - -s -w -X main.version={{ .Version }} + # commit and date go in too: "convctl says this conversion is lossy" + # is unactionable in a bug report without knowing which convctl. + - -s -w + - -X main.version={{ .Version }} + - -X main.commit={{ .FullCommit }} + - -X main.date={{ .Date }} archives: - id: convctl @@ -27,6 +32,62 @@ archives: - goos: windows formats: [zip] +# Linux packages, attached to the release. Installing a kubectl-adjacent +# tool by downloading a tarball is not how the ecosystem works. +nfpms: + - id: convctl + ids: [convctl] + package_name: convctl + vendor: TeraSky + homepage: https://terasky-oss.github.io/declarative-conversion-operator/ + maintainer: TeraSky OSS + description: |- + CLI for the declarative conversion operator: validate, analyze and test + Crossplane XRD and CRD conversion configs offline or against a cluster. + license: Apache-2.0 + formats: [deb, rpm] + bindir: /usr/bin + +# Homebrew. GoReleaser deprecated `brews:` in favour of casks, which are +# macOS-only — so Linuxbrew users install from the deb/rpm above or the +# archive, which is the trade the deprecation forces. +homebrew_casks: + - name: convctl + ids: [convctl] + repository: + owner: terasky-oss + name: homebrew-tap + # index rather than .Env.X: the dotted form errors when the variable + # is unset, which would fail a release because a tap repository does + # not exist yet. Unset means skip, below. + token: '{{ index .Env "HOMEBREW_TAP_TOKEN" }}' + branch: main + homepage: https://terasky-oss.github.io/declarative-conversion-operator/ + description: CLI for the declarative conversion operator + skip_upload: '{{ if index .Env "HOMEBREW_TAP_TOKEN" }}false{{ else }}true{{ end }}' + # An unsigned binary downloaded by a cask is quarantined by Gatekeeper, + # and the failure ("convctl is damaged") reads like a corrupt download + # rather than a policy. The archive's checksum is already verified by + # Homebrew itself. + hooks: + post: + install: | + if system_command("/usr/bin/xattr", args: ["-h"]).exit_status == 0 + system_command "/usr/bin/xattr", args: ["-dr", "com.apple.quarantine", "#{staged_path}/convctl"] + end + +scoops: + - name: convctl + ids: [convctl] + repository: + owner: terasky-oss + name: scoop-bucket + token: '{{ index .Env "SCOOP_BUCKET_TOKEN" }}' + homepage: https://terasky-oss.github.io/declarative-conversion-operator/ + description: CLI for the declarative conversion operator + license: Apache-2.0 + skip_upload: '{{ if index .Env "SCOOP_BUCKET_TOKEN" }}false{{ else }}true{{ end }}' + checksum: name_template: "checksums.txt" diff --git a/cmd/convctl/main.go b/cmd/convctl/main.go index 11f5491..fec142b 100644 --- a/cmd/convctl/main.go +++ b/cmd/convctl/main.go @@ -25,10 +25,19 @@ import ( "github.com/terasky-oss/declarative-conversion-operator/internal/cli" ) -// version is overridden at build time via -ldflags "-X main.version=...". -var version = "dev" +// Overridden at build time via -ldflags "-X main.version=..." and friends. +// Absent those — a `go install` build — internal/cli falls back to the build +// information the Go toolchain embeds, so the binary still reports something +// that maps to a commit. +var ( + version = "dev" + commit = "" + date = "" +) func main() { cli.Version = version + cli.Commit = commit + cli.Date = date os.Exit(cli.Execute()) } diff --git a/docs/installation.md b/docs/installation.md index 7e40d16..23ecdc1 100644 --- a/docs/installation.md +++ b/docs/installation.md @@ -157,6 +157,45 @@ examples (`examples/gitops/flux`, `examples/gitops/argo`). Keep `driftPolicy: KeepServingStale` on GitOps-managed configs — `FailClosed` drops conversions while the schema and config reconcile independently. +## Installing `convctl` + +| Method | Command | Platforms | +|---|---|---| +| Homebrew | `brew install terasky-oss/tap/convctl` | macOS (casks are macOS-only; on Linuxbrew use the package or archive) | +| Scoop | `scoop bucket add terasky-oss https://github.com/terasky-oss/scoop-bucket` then `scoop install convctl` | Windows | +| deb | `sudo dpkg -i convctl__linux_amd64.deb` | Debian, Ubuntu | +| rpm | `sudo rpm -i convctl__linux_amd64.rpm` | RHEL, Fedora, SUSE | +| Archive | download `declarative-conversion-operator-cli___.tar.gz` from the [releases page](https://github.com/TeraSky-OSS/declarative-conversion-operator/releases) | all | +| Container | see [below](#the-convctl-container-image) | linux/amd64, linux/arm64 | +| Source | `go install github.com/terasky-oss/declarative-conversion-operator/cmd/convctl@latest` | all | + +Every archive's checksum is covered by the cosign-signed `checksums.txt`; see +the signed-artifact section of any release for the verification commands. + +### `convctl version` + +```console +$ convctl version +v0.5.0 (a1b2c3d4e5f6) linux/amd64 go1.26.6 + +$ convctl version -o json +{ + "version": "v0.5.0", + "commit": "a1b2c3d4e5f6...", + "date": "2026-09-15T20:06:15Z", + "goVersion": "go1.26.6", + "platform": "linux/amd64" +} +``` + +That is what belongs in a bug report: *"convctl says this conversion is +lossy"* is unactionable without knowing which convctl. + +A `go install` build has no release ldflags, and reports the module version +and VCS stamps the Go toolchain embeds rather than `dev` — a version nobody +can map to a commit is the same as no version. A build from a dirty working +tree says so, with a `-dirty` suffix on the commit. + ## The `convctl` container image For pipelines that would rather pin a digest than download a binary — Tekton, diff --git a/internal/cli/root.go b/internal/cli/root.go index d5705fd..fb7481a 100644 --- a/internal/cli/root.go +++ b/internal/cli/root.go @@ -74,14 +74,33 @@ cluster. Every command works against either resource type: var exitCode = ExitOK func newVersionCmd() *cobra.Command { - return &cobra.Command{ + var output string + cmd := &cobra.Command{ Use: "version", Short: "Print the convctl version", + Long: `Print the version, and with -o json the commit, build date, Go version and +platform as well. + +"convctl says this conversion is lossy" is unactionable without knowing which +convctl, so the JSON form is what belongs in a bug report. A binary built with +go install reports its module version and VCS stamps rather than "dev": those +are embedded by the toolchain, and a version nobody can map to a commit is the +same as no version.`, RunE: func(cmd *cobra.Command, args []string) error { - _, err := fmt.Fprintln(cmd.OutOrStdout(), Version) + if err := checkOutputFormat(output, "table", "json"); err != nil { + return err + } + info := versionInfo() + if output == "json" { + return writeJSON(cmd, info) + } + _, err := fmt.Fprintln(cmd.OutOrStdout(), info.String()) return err }, } + cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json") + registerOutputCompletions(cmd, "table", "json") + return cmd } // Version is set at build time via -ldflags; defaults to "dev" for local diff --git a/internal/cli/version.go b/internal/cli/version.go new file mode 100644 index 0000000..016e2d9 --- /dev/null +++ b/internal/cli/version.go @@ -0,0 +1,108 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "runtime" + "runtime/debug" +) + +// Build metadata, set by -ldflags on a release build. Commit and Date are +// filled from the embedded build info when they are not, which is what makes +// `go install` produce something better than three empty strings. +var ( + // Commit is the git revision the binary was built from. + Commit = "" + // Date is the build date, RFC3339. + Date = "" +) + +// VersionInfo is everything `convctl version -o json` reports. +// +// The fields exist because they are what a bug report needs: "convctl says +// the conversion is lossy" is unactionable without knowing which convctl. +type VersionInfo struct { + Version string `json:"version"` + Commit string `json:"commit,omitempty"` + Date string `json:"date,omitempty"` + GoVersion string `json:"goVersion"` + Platform string `json:"platform"` +} + +// versionInfo assembles the build metadata, falling back to the information +// the Go toolchain embeds in every module-built binary. +// +// A `go install ...@latest` build has no ldflags, and reporting "dev" for it +// is how a bug report arrives with a version nobody can map to a commit. +// debug.ReadBuildInfo knows the module version and the VCS stamps, so use +// them rather than a placeholder. +func versionInfo() VersionInfo { + info := VersionInfo{ + Version: Version, + Commit: Commit, + Date: Date, + GoVersion: runtime.Version(), + Platform: runtime.GOOS + "/" + runtime.GOARCH, + } + bi, ok := debug.ReadBuildInfo() + if !ok { + return info + } + if info.Version == "" || info.Version == "dev" { + // Main.Version is "(devel)" for a local build and the module + // version for `go install module@version`. + if v := bi.Main.Version; v != "" && v != "(devel)" { + info.Version = v + } + } + for _, s := range bi.Settings { + switch s.Key { + case "vcs.revision": + if info.Commit == "" { + info.Commit = s.Value + } + case "vcs.time": + if info.Date == "" { + info.Date = s.Value + } + case "vcs.modified": + if s.Value == "true" && info.Commit != "" { + info.Commit += "-dirty" + } + } + } + if info.Version == "" { + info.Version = "dev" + } + return info +} + +// String is the one-line form the bare `convctl version` prints. +func (v VersionInfo) String() string { + out := v.Version + if v.Commit != "" { + out += " (" + shortCommit(v.Commit) + ")" + } + return out + " " + v.Platform + " " + v.GoVersion +} + +func shortCommit(c string) string { + if len(c) > 12 { + return c[:12] + } + return c +} diff --git a/internal/cli/version_test.go b/internal/cli/version_test.go new file mode 100644 index 0000000..96b2ab1 --- /dev/null +++ b/internal/cli/version_test.go @@ -0,0 +1,86 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "runtime" + "strings" + "testing" +) + +// A bug report needs to say which convctl, and a release build has all of it +// from ldflags. +func TestVersionInfo_UsesLdflagsWhenPresent(t *testing.T) { + origV, origC, origD := Version, Commit, Date + t.Cleanup(func() { Version, Commit, Date = origV, origC, origD }) + + Version, Commit, Date = "v1.2.3", "abcdef1234567890", "2026-01-02T03:04:05Z" + got := versionInfo() + if got.Version != "v1.2.3" || got.Commit != "abcdef1234567890" || got.Date != "2026-01-02T03:04:05Z" { + t.Errorf("ldflags were not used: %+v", got) + } + if got.GoVersion != runtime.Version() { + t.Errorf("GoVersion = %q", got.GoVersion) + } + if got.Platform != runtime.GOOS+"/"+runtime.GOARCH { + t.Errorf("Platform = %q", got.Platform) + } +} + +// A `go install` build has no ldflags. Reporting "dev" for it is how a bug +// report arrives with a version nobody can map to a commit — the toolchain +// embeds the module version and VCS stamps, so use them. +func TestVersionInfo_FallsBackToEmbeddedBuildInfo(t *testing.T) { + origV, origC, origD := Version, Commit, Date + t.Cleanup(func() { Version, Commit, Date = origV, origC, origD }) + + Version, Commit, Date = "dev", "", "" + got := versionInfo() + if got.Version == "" { + t.Error("version is empty, which is worse than dev") + } + // Under `go test` the binary carries build settings, so at minimum the + // fallback must not regress what ldflags would have given. + if got.Version == "dev" && got.Commit != "" { + t.Errorf("a commit was found but the version stayed dev: %+v", got) + } +} + +// The one-line form is what a human reads first; it has to identify the +// build without being a paragraph. +func TestVersionInfo_StringIsOneUsefulLine(t *testing.T) { + v := VersionInfo{Version: "v1.2.3", Commit: "abcdef1234567890abcdef", GoVersion: "go1.26.6", Platform: "linux/amd64"} + got := v.String() + if strings.Contains(got, "\n") { + t.Errorf("version line spans lines: %q", got) + } + for _, want := range []string{"v1.2.3", "abcdef123456", "linux/amd64", "go1.26.6"} { + if !strings.Contains(got, want) { + t.Errorf("%q is missing %q", got, want) + } + } + if strings.Contains(got, "abcdef1234567890abcdef") { + t.Errorf("the full commit was printed rather than a short one: %q", got) + } +} + +func TestVersionInfo_StringWithoutACommit(t *testing.T) { + v := VersionInfo{Version: "dev", GoVersion: "go1.26.6", Platform: "darwin/arm64"} + if got := v.String(); !strings.HasPrefix(got, "dev ") || strings.Contains(got, "()") { + t.Errorf("empty commit rendered badly: %q", got) + } +} From bef3a150d2da322db27ebe29ee504ee6f2439690 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:20:01 +0300 Subject: [PATCH 06/24] feat(convctl): --package, a Crossplane package as a schema source MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit For a platform shipped as a Crossplane Configuration, the unit of API change is a package version — not a git commit, and not the live cluster. All three existing schema sources miss that unit, which means the team whose XRDs most need conversion testing is the team least able to run it. --package slots in exactly where --xrd does rather than being a new verb, so validate, analyze, test and lint gain it at once. It joins the schema flag group with --xrd and --crd, not the sample group where --live lives, which is what lets the two compose: convctl test --package ./platform-v1.4.0.xpkg --config config.yaml --live That composition is the point. Schemas from the version about to be rolled out, objects from the cluster about to receive it: "if I bump this Configuration, do my 4,000 existing composites still convert?" is the question platform teams have before an upgrade, and nothing answered it. lint gains it too, pairing every config in a tree against the XRDs the package ships rather than against whatever XRD files happen to sit beside them — which is the thing that drifts. That is the Configuration repository's CI gate, and docs/gitops/configuration-ci.md walks it end to end. Only the local .xpkg form is implemented. An xpkg is an OCI image saved as a tarball — manifest.json, a config blob, gzipped layer tars, with package.yaml inside the last layer that has one — so reading it needs nothing but the standard library, and it is the tightest loop, before anything is published. Registry and cluster references are recognised and rejected with the `crossplane xpkg pull` command that produces a local file, rather than silently unsupported: supporting them means a registry client the offline path does not need and should not carry. Recorded in limitations.md. The reader was written against the format Crossplane actually produces rather than against the format the docs describe: the fixture in testdata/package/platform.xpkg is a real `crossplane xpkg build` output carrying two XRDs, which is what makes it exercise --target as well as the single-XRD path. A fixture written from the same assumptions as the reader would have proved nothing. Two decisions worth recording. A package shipping several XRDs is an error naming the candidates rather than a guess — picking the first would make the answer depend on the order the package happened to be built in. And the image-reference heuristic is deliberately narrow: anything not clearly a registry reference is treated as a path, so a mistyped filename fails with "no such file" rather than with advice about pulling an image. Reads are bounded at 256 MiB, so a decompression bomb fails as a too-large package rather than as an exhausted machine. Closes #144 Co-Authored-By: Claude Opus 5 (1M context) --- docs/cli.md | 68 +++- docs/gitops/configuration-ci.md | 101 ++++++ docs/limitations.md | 1 + internal/cli/analyze.go | 16 +- internal/cli/lint.go | 67 ++++ internal/cli/root.go | 35 +- internal/cli/test.go | 23 +- internal/cli/testdata/package/README.md | 14 + internal/cli/testdata/package/platform.xpkg | Bin 0 -> 4608 bytes internal/cli/validate.go | 17 +- internal/cli/xpkg.go | 343 ++++++++++++++++++++ internal/cli/xpkg_test.go | 242 ++++++++++++++ mkdocs.yml | 1 + 13 files changed, 901 insertions(+), 27 deletions(-) create mode 100644 docs/gitops/configuration-ci.md create mode 100644 internal/cli/testdata/package/README.md create mode 100644 internal/cli/testdata/package/platform.xpkg create mode 100644 internal/cli/xpkg.go create mode 100644 internal/cli/xpkg_test.go diff --git a/docs/cli.md b/docs/cli.md index 68beaf3..b5c0b7d 100644 --- a/docs/cli.md +++ b/docs/cli.md @@ -3,10 +3,10 @@ `convctl` runs the exact same `pkg/engine` code the operator and webhook server use, entirely offline against local YAML files — so you can validate and test a conversion mapping before it ever touches a cluster. Most commands work identically against an `XRDConversionConfig` (pass `--xrd`) or a `CRDConversionConfig` (pass `--crd`) — which one applies is determined by the config file's own `kind`, not by which flag you happen to type, so passing the wrong one is a clear error rather than a silent mismatch. `migrate-storage` is the exception: it is a live, mutating housekeeping command that takes cluster resource names (not files) and does not need a conversion config. ```console -convctl lint [path...] [--schema-dir dir] [--exclude glob] [-o table|json|github|sarif|markdown] -convctl validate --config config.yaml [--xrd xrd.yaml | --crd crd.yaml] [-o table|json|github|sarif|markdown] -convctl analyze --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json|github|sarif|markdown] -convctl test --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) (--samples ./samples/ | --live) [-o table|json|junit|github|sarif|markdown] [flags] +convctl lint [path...] [--schema-dir dir] [--package platform.xpkg] [--exclude glob] [-o table|json|github|sarif|markdown] +convctl validate --config config.yaml [--xrd xrd.yaml | --crd crd.yaml | --package p.xpkg] [-o table|json|github|sarif|markdown] +convctl analyze --config config.yaml (--xrd xrd.yaml | --crd crd.yaml | --package p.xpkg) [-o table|json|github|sarif|markdown] +convctl test --config config.yaml (--xrd xrd.yaml | --crd crd.yaml | --package p.xpkg) (--samples ./samples/ | --live) [-o table|json|junit|github|sarif|markdown] [flags] convctl plan --to v2 (--xrd xrd.yaml | --crd crd.yaml) [--config config.yaml] [-o table|json] convctl versions --xrd xrd.yaml [--config config.yaml] [--check-unserve v1] [-o table|json] convctl compat --base REV --head REV --config config.yaml (--xrd xrd.yaml | --crd crd.yaml) [-o table|json|markdown] @@ -26,6 +26,66 @@ convctl crossplane status [-o table|json] Roughly in the order you reach for them while authoring a mapping: `suggest` drafts rules for fields nothing covers yet, `validate` and `analyze` check the config statically, `convert` shows what a single object turns into, `test` grades fixtures or every live object, `diff` reports what a config edit changed, and `patch-preview` shows the exact patch the operator will apply once you commit. After a hub/storage-version promotion, `migrate-storage` rewrites live objects (critical for native CRDs; on XRDs the `compositionRef` retarget usually already did, and the remaining job is pruning `storedVersions`). For a GitOps hub flip, `generate kyverno` drafts MutatingPolicies that retarget existing XRs without a per-object name patch; on a cluster without Kyverno, `retarget` does the same job directly. `crossplane status` answers "where is my migration right now?" without assembling it from half a dozen `kubectl` invocations. Around all of it, `plan` sequences the migration, `versions` answers whether an old version can be retired yet, and `compat` gates config edits in review. +## Schema sources + +Every command that needs a schema takes one of these, interchangeably: + +| Source | Flag | Answers | +|---|---|---| +| A file | `--xrd xrd.yaml` / `--crd crd.yaml` | "does my config hold against this schema?" | +| A Crossplane package | `--package ./platform.xpkg` | "does it hold against the XRDs I am about to publish?" — no registry, no cluster | +| The cluster | `--live` (a **sample** source, not a schema source) | — | + +`--package` slots in exactly where `--xrd` does rather than being a new verb, +so `validate`, `analyze`, `test` and `lint` all take it. + +### Why a package is its own source + +For a platform shipped as a Crossplane `Configuration`, **the unit of API +change is a package version** — not a git commit, and not the live cluster. +The team whose XRDs most need conversion testing is otherwise the team least +able to run it. + +```console +# Does my config hold against the XRDs I am about to publish? +convctl test --package ./platform.xpkg --config config.yaml --samples ./samples/ + +# The one that matters most: the version about to be rolled out, against the +# objects already in the cluster that will receive it. +convctl test --package ./platform.xpkg --config config.yaml --live + +# A whole package against a whole config tree — the Configuration repo's gate. +convctl lint ./configs/ --package ./platform.xpkg +``` + +`--package` is a schema source and `--live` is a sample source, so they +compose: *"if I bump this Configuration, do my 4,000 existing composites still +convert?"* is the question platform teams have before an upgrade, and nothing +else answers it. + +`--target ` selects one XRD when a package ships several. Omitting +it is an error naming the candidates rather than a guess — picking the first +would make the answer depend on the order the package was built in. + +### What is implemented + +Only the **local `.xpkg`** form. An xpkg is an OCI image saved as a tarball, +so reading it needs nothing but the standard library — and it is the tightest +loop, before anything is published anywhere. + +Registry references (`ghcr.io/org/platform:v1.4.0`) and cluster references +(`configuration/`, `configurationrevision/`) are recognised and +rejected with the command that gets you a local file: + +```console +$ convctl analyze --package ghcr.io/org/platform:v1.4.0 --config config.yaml +error: reading a package from a registry (ghcr.io/org/platform:v1.4.0) is not implemented yet; +`crossplane xpkg pull ghcr.io/org/platform:v1.4.0 -o package.xpkg` and pass the file +``` + +Adding them means a registry client (`go-containerregistry`), which the +offline path does not need and should not carry. + ## Running it in a container ```console diff --git a/docs/gitops/configuration-ci.md b/docs/gitops/configuration-ci.md new file mode 100644 index 0000000..dd546f4 --- /dev/null +++ b/docs/gitops/configuration-ci.md @@ -0,0 +1,101 @@ +# The Configuration repository's CI gate + +For a platform shipped as a Crossplane `Configuration`, the unit of API change +is a **package version**. Not a commit, and not the live cluster — so a gate +built on either checks something adjacent to what actually ships. + +This is the end-to-end gate for a repository that builds a `Configuration`. + +## The shape of it + +1. Build the package from the repository's XRDs. +2. Lint every conversion config in the tree **against that package**. +3. Test the conversions against fixtures. +4. Before a release, test them against the objects already in the cluster that + will receive the upgrade. + +Steps 1–3 need no cluster at all. + +## The workflow + +```yaml +name: Configuration +on: pull_request + +permissions: + contents: read + +jobs: + conversion: + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + + - name: Install crossplane CLI + run: | + curl -sL https://raw.githubusercontent.com/crossplane/crossplane/main/install.sh | sh + sudo mv crossplane /usr/local/bin/ + + - name: Install convctl + run: go install github.com/terasky-oss/declarative-conversion-operator/cmd/convctl@latest + + - name: Build the package + run: crossplane xpkg build --package-root=./apis --package-file=platform.xpkg + + # Every conversion config in the tree, against the XRDs the package + # actually ships — not against whatever XRD files happen to sit beside + # them, which is the thing that drifts. + - name: Lint conversions against the package + run: convctl lint ./apis/ --package platform.xpkg -o github + + - name: Test conversions against fixtures + run: | + convctl test --package platform.xpkg --target xwidgets.example.org \ + --config apis/widgets/conversion.yaml \ + --samples apis/widgets/samples/ \ + --validate-output -o github +``` + +`-o github` puts each finding on the line of the config that produced it, so +a failure lands on the diff rather than in a log — +see [CI-native formats](../cli.md#ci-native-formats-github-sarif-markdown). + +## Before the upgrade lands + +The composition that matters most pairs the **new** schemas with the **existing** +objects: + +```console +convctl test --package ./platform-v1.4.0.xpkg --config config.yaml --live \ + --max-samples 500 --sample-strategy random --seed 1 +``` + +*"If I bump this Configuration, do my 4,000 existing composites still +convert?"* `--package` is a schema source and `--live` is a sample source, so +they compose without anything new. On a large cluster, +[bounded sampling](../cli.md#bounded-sampling-on-a-large-cluster) keeps it +runnable — and the report says plainly that it sampled. + +## Package-managed XRDs need the conversion config applied first + +An XRD shipped in a package has its `spec.conversion` stripped on every +package resync unless the [conversion guard](../architecture.md#the-xrd-conversion-guard) +is on. That also changes the **order** of a version migration: the conversion +config must be applied *before* the package revision that serves the new +version, or the version is served with no conversion at all. + +`convctl plan` detects a package-managed XRD and orders the steps accordingly: + +```console +convctl plan --xrd xrd.yaml --config config.yaml --to v2 --package-managed +``` + +See [`convctl plan`](../cli.md#package-managed-xrds-are-ordered-differently). + +## Related + +- [`convctl lint`](../cli.md#convctl-lint) — the offline check that runs on every commit +- [Schema sources](../cli.md#schema-sources) — `--xrd`, `--crd`, `--package` +- [Fleet CI](fleet-ci.md) — the same checks across many clusters diff --git a/docs/limitations.md b/docs/limitations.md index e001e0c..4dce000 100755 --- a/docs/limitations.md +++ b/docs/limitations.md @@ -37,6 +37,7 @@ This page is deliberately blunt about what the operator does *not* do today, so - **The webhook-server's cache is built from the feature flags, not from what is on the cluster.** `--enable-xrd-support=false` is what keeps `CompositeResourceDefinition` out of the informer cache entirely. controller-runtime resolves every cached kind through the RESTMapper when the manager is constructed, so on a cluster without Crossplane a replica told to support XRDs does not degrade — it exits at startup with `no matches for kind "CompositeResourceDefinition"`. That is deliberate and matches the manager's own behaviour, but it means the flags must match the cluster rather than describing a preference. - **Required-field analysis treats a required source leaf as present whenever its own parent object is.** Asking "does the rule set always produce this required destination field?" needs the mirror question about the source, and the analysis answers it one level deep: `spec.network.cidr` marked required inside an **optional** `spec.network` counts as guaranteed. Where an optional ancestor gates it, a genuinely conditional source can therefore be read as unconditional, and that case goes unreported. Tightening the rule to "every ancestor must itself be required" is not the fix: nothing under `spec` is ever listed in a CRD root's `required`, so every ordinary rename would be flagged. The sound version has to model whether the destination parent's creation is coupled to the same source path — a rule that creates `spec.network` *only* by writing `cidr` into it is correct, and a naive check calls it broken. Until then this is a missed diagnostic rather than a wrong one, and `--validate-output` and `--fuzz` are the empirical checks that cover it. This is about the **source** side only: on the destination side an optional ancestor is handled — a required field whose containing object is produced whole by a rule, rather than written into, is reported as *unprovable* rather than skipped. - **Required-field satisfaction cannot be proven through `cel`, `jsonPatch` or `scalarToFields`.** The engine checks, for every required field on each side, whether the rule set always produces a value there — and reports an error when nothing does, or when the only rule that does may not fire. It cannot reason about arbitrary expressions, so a required field written by one of those three strategies is reported as *unprovable* (a warning) rather than as satisfied or failed. `convctl test --validate-output` and `--fuzz` are the empirical checks for that residue. A required field whose schema declares a `default` is always satisfied, because the apiserver fills it in. +- **`--package` reads a local `.xpkg` only.** An xpkg is an OCI image saved as a tarball, so the local form needs nothing but the standard library — and it is the tightest loop, before anything is published. A registry reference (`ghcr.io/org/platform:v1.4.0`) or a cluster reference (`configuration/`, `configurationrevision/`) is recognised and rejected with the `crossplane xpkg pull` command that produces a local file, rather than silently unsupported. Supporting them means a registry client the offline path does not need and should not carry. - **An uncapped `--live` run holds every object it tests.** Listing streams into the sampler, so `--max-samples` genuinely bounds memory — the population is counted without being held. Without a cap the objects are still accumulated before testing, which on a cluster with tens of thousands of large composites is the case `--max-samples` exists for. Testing each page as it arrives would bound it either way and is not implemented; the report is assembled from the full sample set. - **`convctl plan` reads manifests, not a cluster.** It answers "what is the next safe step?" from the XRD/CRD and the conversion config, which is what makes it usable in a PR. Three steps of the sequence have no answer in a manifest — retargeting Compositions, migrating stored objects, pruning `storedVersions` — and are reported as `UNKNOWN` with the command that answers them, never as done, ready, or blocked. A plan therefore stops at *"nothing outstanding that files can decide"*; finishing the migration still requires running those verify commands against the cluster. - **`convctl versions` is XRD-only.** The live inventory — object counts, `storedVersions`, and the field managers still writing each version — is implemented for XRD targets; `--crd` is rejected with a message saying so rather than silently answering a narrower question. `plan`, `compat`, `validate`, `analyze` and `test` all cover both. diff --git a/internal/cli/analyze.go b/internal/cli/analyze.go index db89073..548a7e7 100644 --- a/internal/cli/analyze.go +++ b/internal/cli/analyze.go @@ -72,6 +72,12 @@ type AnalyzeSpokeView struct { // static analysis, with no samples involved. Which of xrdPath/crdPath // applies is determined by the config's own kind. func RunAnalyze(xrdPath, crdPath, configPath string) (*AnalyzeOutput, error) { + return RunAnalyzeFrom(xrdPath, crdPath, configPath, "", "") +} + +// RunAnalyzeFrom is RunAnalyze with a package as an alternative schema +// source. See XRDFromSource. +func RunAnalyzeFrom(xrdPath, crdPath, configPath, packageRef, target string) (*AnalyzeOutput, error) { kind, err := PeekConfigKind(configPath) if err != nil { return nil, err @@ -83,15 +89,15 @@ func RunAnalyze(xrdPath, crdPath, configPath string) (*AnalyzeOutput, error) { } return runAnalyzeCRDCmd(crdPath, configPath) default: // "XRDConversionConfig" - if xrdPath == "" { - return nil, fmt.Errorf("%s is an XRDConversionConfig; pass its target schema with --xrd, not --crd", configPath) + if xrdPath == "" && packageRef == "" { + return nil, fmt.Errorf("%s is an XRDConversionConfig; pass its target schema with --xrd or --package, not --crd", configPath) } - return runAnalyzeXRDCmd(xrdPath, configPath) + return runAnalyzeXRDCmd(xrdPath, configPath, packageRef, target) } } -func runAnalyzeXRDCmd(xrdPath, configPath string) (*AnalyzeOutput, error) { - xrd, err := LoadXRD(xrdPath) +func runAnalyzeXRDCmd(xrdPath, configPath, packageRef, target string) (*AnalyzeOutput, error) { + xrd, err := XRDFromSource(xrdPath, packageRef, target) if err != nil { return nil, err } diff --git a/internal/cli/lint.go b/internal/cli/lint.go index 57a8678..7bc902f 100644 --- a/internal/cli/lint.go +++ b/internal/cli/lint.go @@ -27,6 +27,8 @@ import ( "sync" "text/tabwriter" + sigsyaml "sigs.k8s.io/yaml" + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" ) @@ -42,6 +44,11 @@ type LintOptions struct { Exclude []string // Concurrency bounds the parallel analysis. Zero means one per CPU. Concurrency int + // PackageRef pairs every config in the tree against the XRDs a + // Crossplane package ships, rather than against schema files. A whole + // package against a whole config tree, one command, is the + // Configuration repository's CI gate. + PackageRef string } // LintPairResult is one config and the schema it was paired with. @@ -85,6 +92,10 @@ type discovered struct { name string // configName is metadata.name for a config. configName string + // display overrides path in the report. A schema staged out of a + // package lives in a temporary directory, and printing that path tells + // the reader nothing about where the schema came from. + display string } // RunLint walks the given trees and checks every conversion config it finds @@ -108,6 +119,23 @@ func RunLint(opts LintOptions) (*LintReport, error) { found = append(found, items...) } + // Schemas from the package, when one was named. They are written to a + // temporary directory rather than held as objects because every + // downstream check reads a path, and a package is read once here + // rather than once per config. + var cleanup func() + if opts.PackageRef != "" { + pkgSchemas, done, err := stagePackageXRDs(opts.PackageRef) + if err != nil { + return nil, err + } + cleanup = done + found = append(found, pkgSchemas...) + } + if cleanup != nil { + defer cleanup() + } + // Schemas first, so a config can be paired as soon as it is seen. A // name collision between two schema files is itself worth reporting, // but the first wins deterministically because the walk is sorted. @@ -238,6 +266,9 @@ func lintOne(c discovered, schemas map[string]discovered, byTarget map[string][] return res } res.Schema = schema.path + if schema.display != "" { + res.Schema = schema.display + } var out *ValidateResult var err error @@ -270,6 +301,42 @@ func lintOne(c discovered, schemas map[string]discovered, byTarget map[string][] return res } +// stagePackageXRDs writes a package's XRDs to a temporary directory and +// returns them as discovered schemas. +// +// Reading the package once and staging it keeps the pairing logic identical +// to the file case: every config in the tree is checked against the XRDs the +// package declares, which is the Configuration repository's gate. +func stagePackageXRDs(ref string) ([]discovered, func(), error) { + pkg, err := ReadPackage(ref) + if err != nil { + return nil, nil, err + } + if len(pkg.XRDs) == 0 { + return nil, nil, fmt.Errorf("%s ships no XRDs to check configs against", ref) + } + dir, err := os.MkdirTemp("", "convctl-package-") + if err != nil { + return nil, nil, err + } + cleanup := func() { _ = os.RemoveAll(dir) } + var out []discovered + for _, x := range pkg.XRDs { + data, merr := sigsyaml.Marshal(x.Object) + if merr != nil { + cleanup() + return nil, nil, fmt.Errorf("%s: re-encoding %s: %w", ref, xrdName(x), merr) + } + path := filepath.Join(dir, xrdName(x)+".yaml") + if werr := os.WriteFile(path, data, 0o600); werr != nil { + cleanup() + return nil, nil, werr + } + out = append(out, discovered{path: path, kind: "XRD", name: xrdName(x), display: ref + " (" + xrdName(x) + ")"}) + } + return out, cleanup, nil +} + // discoverIn walks one tree, recognising conversion configs and schemas by // their own apiVersion/kind rather than by filename. func discoverIn(root string, exclude []string) ([]discovered, error) { diff --git a/internal/cli/root.go b/internal/cli/root.go index fb7481a..0ba998f 100644 --- a/internal/cli/root.go +++ b/internal/cli/root.go @@ -108,7 +108,7 @@ same as no version.`, var Version = "dev" func newValidateCmd() *cobra.Command { - var configPath, xrdPath, crdPath, output string + var configPath, xrdPath, crdPath, output, packagePath, packageTarget string cmd := &cobra.Command{ Use: "validate", Short: "Validate a conversion config the same way the admission webhook does", @@ -122,7 +122,7 @@ schemas.`, if err := checkOutputFormat(output, "table", "json", "github", "sarif", "markdown"); err != nil { return err } - res, err := RunValidate(configPath, xrdPath, crdPath) + res, err := RunValidateFrom(configPath, xrdPath, crdPath, packagePath, packageTarget) if err != nil { return err } @@ -151,16 +151,22 @@ schemas.`, cmd.Flags().StringVarP(&configPath, "config", "c", "", "Path to an XRDConversionConfig or CRDConversionConfig YAML file (required)") cmd.Flags().StringVarP(&xrdPath, "xrd", "x", "", "Path to an XRD YAML file (optional; enables live schema validation against an XRDConversionConfig)") cmd.Flags().StringVar(&crdPath, "crd", "", "Path to a CRD YAML file (optional; enables live schema validation against a CRDConversionConfig)") + cmd.Flags().StringVar(&packagePath, "package", "", "Read the target XRD from a Crossplane package instead of a file (a local .xpkg). The unit of API change for a platform shipped as a Configuration is a package version") + cmd.Flags().StringVar(&packageTarget, "target", "", "With --package, select one XRD by name when the package ships several") cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|github|sarif|markdown") _ = cmd.MarkFlagRequired("config") cmd.MarkFlagsMutuallyExclusive("xrd", "crd") + // --package is a schema source like --xrd, so it joins that group -- + // not the sample group, where --live lives. + cmd.MarkFlagsMutuallyExclusive("xrd", "package") + cmd.MarkFlagsMutuallyExclusive("crd", "package") registerOfflineFlagCompletions(cmd) registerOutputCompletions(cmd, "table", "json", "github", "sarif", "markdown") return cmd } func newAnalyzeCmd() *cobra.Command { - var xrdPath, crdPath, configPath, output string + var xrdPath, crdPath, configPath, output, packagePath, packageTarget string cmd := &cobra.Command{ Use: "analyze", Short: "Report lossiness and rule coverage from schemas alone", @@ -173,7 +179,7 @@ are lossy in which direction, and whether every schema field is covered.`, if err := checkOutputFormat(output, "table", "json", "github", "sarif", "markdown"); err != nil { return err } - out, err := RunAnalyze(xrdPath, crdPath, configPath) + out, err := RunAnalyzeFrom(xrdPath, crdPath, configPath, packagePath, packageTarget) if err != nil { return err } @@ -196,11 +202,17 @@ are lossy in which direction, and whether every schema field is covered.`, } cmd.Flags().StringVarP(&xrdPath, "xrd", "x", "", "Path to an XRD YAML file (required for an XRDConversionConfig)") cmd.Flags().StringVar(&crdPath, "crd", "", "Path to a CRD YAML file (required for a CRDConversionConfig)") + cmd.Flags().StringVar(&packagePath, "package", "", "Read the target XRD from a Crossplane package instead of a file (a local .xpkg). The unit of API change for a platform shipped as a Configuration is a package version") + cmd.Flags().StringVar(&packageTarget, "target", "", "With --package, select one XRD by name when the package ships several") cmd.Flags().StringVarP(&configPath, "config", "c", "", "Path to an XRDConversionConfig or CRDConversionConfig YAML file (required)") cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|github|sarif|markdown") _ = cmd.MarkFlagRequired("config") - cmd.MarkFlagsOneRequired("xrd", "crd") + cmd.MarkFlagsOneRequired("xrd", "crd", "package") cmd.MarkFlagsMutuallyExclusive("xrd", "crd") + // --package is a schema source like --xrd, so it joins that group -- + // not the sample group, where --live lives. + cmd.MarkFlagsMutuallyExclusive("xrd", "package") + cmd.MarkFlagsMutuallyExclusive("crd", "package") registerOfflineFlagCompletions(cmd) registerOutputCompletions(cmd, "table", "json", "github", "sarif", "markdown") return cmd @@ -209,6 +221,7 @@ are lossy in which direction, and whether every schema field is covered.`, func newTestCmd() *cobra.Command { var ( xrdPath, crdPath, configPath, samplesDir, output, failOn, outputFile string + packagePath, packageTarget string recordDir, goldenDir, recordFailures string fuzzN int fuzzSeed int64 @@ -306,6 +319,7 @@ results are collected by sample index, never by completion order.`, } opts := TestOptions{ XRDPath: xrdPath, CRDPath: crdPath, ConfigPath: configPath, SamplesDir: samplesDir, + PackagePath: packagePath, PackageTarget: packageTarget, SkipIdentity: skipIdentity, RestrictVersionPairs: versionPairs, Live: live, Kubeconfig: kubeconfig, KubeContext: kubeContext, Contexts: contexts, KubeconfigDir: kubeconfigDir, @@ -373,6 +387,8 @@ results are collected by sample index, never by completion order.`, cmd.Flags().StringVar(&failOn, "fail-on", failOnLoss, "Exit-code threshold: none|warn|loss") cmd.Flags().StringSliceVar(&versionPairs, "version-pair", nil, "Restrict testing to these version(s), repeatable") cmd.Flags().IntVar(&concurrency, "concurrency", 0, "Number of samples to test in parallel (default: one per available CPU)") + cmd.Flags().StringVar(&packagePath, "package", "", "Read the target XRD from a Crossplane package instead of a file (a local .xpkg). The unit of API change for a platform shipped as a Configuration is a package version") + cmd.Flags().StringVar(&packageTarget, "target", "", "With --package, select one XRD by name when the package ships several") cmd.Flags().IntVar(&maxSamples, "max-samples", 0, "With --live, cap how many objects are tested (default: every object). The report says so when a run was sampled") cmd.Flags().StringVar(&sampleStrategy, "sample-strategy", "", "With --max-samples: first (cheapest), random (uniform, reproducible with --seed), or newest (default: first)") cmd.Flags().StringVar(&namespace, "namespace", "", "With --live, narrow to one namespace. Only namespaced object classes are affected; a cluster-scoped composite cannot be narrowed") @@ -385,8 +401,12 @@ results are collected by sample index, never by completion order.`, cmd.Flags().StringVar(&goldenDir, "golden", "", "Replay a corpus recorded by --record and fail on any difference, reporting which fields changed. Mutually exclusive with --record") cmd.Flags().BoolVar(&validateOutput, "validate-output", false, "Validate every converted object against the destination version's OpenAPI schema, using the apiserver's own validator. A violation is an error, not a loss. Off by default this release; the default is planned to flip") _ = cmd.MarkFlagRequired("config") - cmd.MarkFlagsOneRequired("xrd", "crd") + cmd.MarkFlagsOneRequired("xrd", "crd", "package") cmd.MarkFlagsMutuallyExclusive("xrd", "crd") + // --package is a schema source like --xrd, so it joins that group -- + // not the sample group, where --live lives. + cmd.MarkFlagsMutuallyExclusive("xrd", "package") + cmd.MarkFlagsMutuallyExclusive("crd", "package") // --fuzz is a third source of samples, and composes with --samples: // generated objects test the boundaries, fixtures test the cases // somebody deliberately wrote down. @@ -854,6 +874,7 @@ func newLintCmd() *cobra.Command { schemaDirs, exclude []string output, failOn string concurrency int + packageRef string ) cmd := &cobra.Command{ Use: "lint [path...]", @@ -894,6 +915,7 @@ the --fail-on threshold, 2 usage error.`, } rep, err := RunLint(LintOptions{ Paths: args, SchemaDirs: schemaDirs, Exclude: exclude, Concurrency: concurrency, + PackageRef: packageRef, }) if err != nil { return err @@ -915,6 +937,7 @@ the --fail-on threshold, 2 usage error.`, }, } cmd.Flags().StringSliceVar(&schemaDirs, "schema-dir", nil, "Additional directories to search for XRDs and CRDs (repeatable)") + cmd.Flags().StringVar(&packageRef, "package", "", "Pair every config in the tree against the XRDs a Crossplane package ships (a local .xpkg), rather than against schema files") cmd.Flags().StringSliceVar(&exclude, "exclude", nil, "Glob patterns to skip, matched against the path and its base name (repeatable)") cmd.Flags().StringVarP(&output, "output", "o", "table", "Output format: table|json|github|sarif|markdown") cmd.Flags().StringVar(&failOn, "fail-on", failOnLoss, "Failure threshold: none|warn|loss") diff --git a/internal/cli/test.go b/internal/cli/test.go index e62120b..7d2849d 100644 --- a/internal/cli/test.go +++ b/internal/cli/test.go @@ -33,11 +33,18 @@ import ( // TestOptions configures RunTest. type TestOptions struct { - XRDPath string - CRDPath string - ConfigPath string - SamplesDir string - SkipIdentity bool + XRDPath string + // PackagePath is a Crossplane package to read the XRD from instead of + // a file: the unit of API change for a platform shipped as a + // Configuration is a package version, not a commit and not the live + // cluster. + PackagePath string + // PackageTarget selects one XRD when the package ships several. + PackageTarget string + CRDPath string + ConfigPath string + SamplesDir string + SkipIdentity bool // RestrictVersionPairs, if non-empty, limits testing to exactly these // "from:to" pairs (both directions still need listing explicitly). RestrictVersionPairs []string @@ -193,8 +200,8 @@ func RunTest(opts TestOptions) (*Report, error) { } return runTestCRD(opts) default: // "XRDConversionConfig" - if opts.XRDPath == "" { - return nil, fmt.Errorf("%s is an XRDConversionConfig; pass its target schema with --xrd, not --crd", opts.ConfigPath) + if opts.XRDPath == "" && opts.PackagePath == "" { + return nil, fmt.Errorf("%s is an XRDConversionConfig; pass its target schema with --xrd or --package, not --crd", opts.ConfigPath) } return runTestXRD(opts) } @@ -203,7 +210,7 @@ func RunTest(opts TestOptions) (*Report, error) { func runTestXRD(opts TestOptions) (*Report, error) { start := time.Now() - xrd, err := LoadXRD(opts.XRDPath) + xrd, err := XRDFromSource(opts.XRDPath, opts.PackagePath, opts.PackageTarget) if err != nil { return nil, err } diff --git a/internal/cli/testdata/package/README.md b/internal/cli/testdata/package/README.md new file mode 100644 index 0000000..d527823 --- /dev/null +++ b/internal/cli/testdata/package/README.md @@ -0,0 +1,14 @@ +`platform.xpkg` is a real `crossplane xpkg build` output, not a hand-made +fixture: the reader's job is to agree with the format Crossplane actually +produces, and a fixture written from the same assumptions as the reader would +prove nothing. + +It ships two XRDs (`xwidgets.example.org` and `xbuckets.example.org`), which +is what makes it exercise `--target` as well as the single-XRD path. + +Rebuild with: + + crossplane xpkg build --package-root= --package-file=platform.xpkg + +where `` holds a `crossplane.yaml` Configuration meta file plus the two +XRDs from `examples/`. diff --git a/internal/cli/testdata/package/platform.xpkg b/internal/cli/testdata/package/platform.xpkg new file mode 100644 index 0000000000000000000000000000000000000000..ef4b5022fd2697a982dd81d452b2061a0f479871 GIT binary patch literal 4608 zcmeHKYfKbZ6ds{9V72%_ZLJj@t(0`ZdF;&G0WC#pacPNY04sDgcV_Oei^~q|F0bM? zXbVxmw$v6#@lo-uHH9FxnA%qA3rSQSid`b0v_^uDs%wl7paW>q@&kW>{^_3F;nx0&q$Gw^4T^w=YKhV$ z!Evm_2wIs25|kjL)+5JBT8WS#Pig@W1R3xIuazhylaz$(?P?~9SvJP^M*IK@$kZ3x zs04@>#cEA5K{EK#@S;SQ)W2qF5-)0wqDj5i(8F zEJRwC;1SJ>GU8={Wz~!XXhpc)ZhuZ+Z|Gr$nsb2KJpu6_5Beg#tuaXdlO#(G?0<$K zT>U>Nsl(*-$@Az#dZFi9Os%<@80%cNqWt6=zIzVb0Un0cvtP_zo{Rg{FXtM=VMKvO zSEc*)m$81eh3(klyzOjTmB;L#D=LB-FO0lD^{rFW$_h73b`$eAZP}dTE1K?Pq+05? zR2G+YIGvA7Zr8G%qtAjbQmebGmhM|+|7l*xw>v$jB!`~&+;7=kBg$8zkMB=Cy)f!) z&n}saxH5kt|4~A8?%wjN*)h6^X)ig8eX&t1PaRq4-0@?|gZI5BW)y+Yxr?y$Gey1_ zmWNKXrzw*w^3gT$hxe!NojvS(F5Fs@D<3X(7<}jI3mVq{#F8VyPmq{dQIz1_m@Fi zR!>gw2n_4#@eBK`5gR{DW$S_&CI%9r*(f8cEg-?Fn1=~-Ak0(Fe}dxZ|DOLO$&s$} ze;|#|WUVT=(8nU_iSioizksUemY^gh3dXB?%ba2|X;k&qC-hWxH%xj%Rd@Pl!iNjg O#ej 0 { + return nil, fmt.Errorf("%s contains no CompositeResourceDefinitions (it does contain %d other object(s)); a conversion config needs an XRD to check against", p.Ref, p.Others) + } + return nil, fmt.Errorf("%s contains no objects at all", p.Ref) + case target != "": + for _, x := range p.XRDs { + if xrdName(x) == target { + return x, nil + } + } + return nil, fmt.Errorf("%s ships no XRD named %q; it ships: %s", p.Ref, target, strings.Join(p.XRDNames(), ", ")) + case len(p.XRDs) == 1: + return p.XRDs[0], nil + default: + return nil, fmt.Errorf("%s ships %d XRDs, so --target must name one of: %s", + p.Ref, len(p.XRDs), strings.Join(p.XRDNames(), ", ")) + } +} + +// ReadPackage reads a Crossplane package. +// +// Only the local form is implemented: a .xpkg file, which is an OCI image +// saved as a tarball, readable with the standard library alone. That is the +// tightest loop — "does my config hold against the XRDs I am about to +// publish?" — and it costs no new dependency. The registry and cluster +// forms are named in the error rather than silently unsupported. +func ReadPackage(ref string) (*PackageContents, error) { + switch { + case strings.HasPrefix(ref, "configuration/"), strings.HasPrefix(ref, "configurationrevision/"): + return nil, fmt.Errorf("reading a package from the cluster (%s) is not implemented yet; build or pull the package and pass the .xpkg file", ref) + case looksLikeImageRef(ref): + return nil, fmt.Errorf("reading a package from a registry (%s) is not implemented yet; `crossplane xpkg pull %s -o package.xpkg` and pass the file", ref, ref) + } + return readLocalXPKG(ref) +} + +// looksLikeImageRef distinguishes ghcr.io/org/platform:v1 from ./p.xpkg. +// +// Deliberately narrow: anything that is not clearly a registry reference is +// treated as a path, so a missing or mistyped file fails with "reading +// package: no such file" rather than with advice about pulling an image. +func looksLikeImageRef(ref string) bool { + if _, err := os.Stat(ref); err == nil { + return false // it exists; it is a file + } + if strings.HasSuffix(ref, ".xpkg") || strings.HasPrefix(ref, ".") || strings.HasPrefix(ref, "/") { + return false + } + host, rest, ok := strings.Cut(ref, "/") + if !ok || rest == "" { + return false // a registry reference for a package always has a path + } + return strings.Contains(host, ".") || strings.Contains(host, ":") || host == "localhost" +} + +// readLocalXPKG reads the package stream out of a .xpkg file. +// +// An xpkg is what `docker save` produces: a tar containing manifest.json, +// a config blob, and one gzipped tar per layer. The package's contents are +// package.yaml inside one of those layers — the last one that has it, since +// a later layer overwrites an earlier one. +func readLocalXPKG(path string) (*PackageContents, error) { + info, err := os.Stat(path) + if err != nil { + return nil, fmt.Errorf("reading package %s: %w", path, err) + } + if info.IsDir() { + return nil, fmt.Errorf("%s is a directory; --package takes a .xpkg file, an image reference, or a cluster reference", path) + } + + layers, err := xpkgLayerOrder(path) + if err != nil { + return nil, err + } + + var data []byte + for _, layer := range layers { + got, err := readFileFromLayer(path, layer, packageFile) + if err != nil { + return nil, err + } + if got != nil { + data = got // later layers win + } + } + if data == nil { + return nil, fmt.Errorf("%s is not a Crossplane package: no %s in any layer", path, packageFile) + } + return parsePackageStream(path, data) +} + +// dockerManifestEntry is the subset of manifest.json this needs. +type dockerManifestEntry struct { + Layers []string `json:"Layers"` +} + +// xpkgLayerOrder reads the layer order from manifest.json, so "last layer +// wins" means what the image says rather than what the tar happened to +// list. +func xpkgLayerOrder(path string) ([]string, error) { + raw, err := readFileFromTar(path, "manifest.json") + if err != nil { + return nil, err + } + if raw == nil { + return nil, fmt.Errorf("%s is not a package: no manifest.json (expected the output of `crossplane xpkg build`)", path) + } + var entries []dockerManifestEntry + if err := json.Unmarshal(raw, &entries); err != nil { + return nil, fmt.Errorf("%s: parsing manifest.json: %w", path, err) + } + if len(entries) == 0 { + return nil, fmt.Errorf("%s: manifest.json declares no image", path) + } + return entries[0].Layers, nil +} + +// readFileFromTar returns one member of a tar, or nil when it is absent. +func readFileFromTar(archive, name string) ([]byte, error) { + // #nosec G304 -- the path the user passed to --package, same trust as + // --config and --xrd. + f, err := os.Open(archive) + if err != nil { + return nil, fmt.Errorf("reading %s: %w", archive, err) + } + defer func() { _ = f.Close() }() + + tr := tar.NewReader(f) + for { + hdr, err := tr.Next() + if errors.Is(err, io.EOF) { + return nil, nil + } + if err != nil { + return nil, fmt.Errorf("%s is not a readable tar archive: %w", archive, err) + } + if hdr.Name != name { + continue + } + return readBounded(tr, archive) + } +} + +// readFileFromLayer returns one member of one (gzipped) layer, or nil when +// either the layer or the member is absent. +func readFileFromLayer(archive, layer, name string) ([]byte, error) { + // #nosec G304 -- see readFileFromTar. + f, err := os.Open(archive) + if err != nil { + return nil, fmt.Errorf("reading %s: %w", archive, err) + } + defer func() { _ = f.Close() }() + + tr := tar.NewReader(f) + for { + hdr, err := tr.Next() + if errors.Is(err, io.EOF) { + return nil, nil + } + if err != nil { + return nil, fmt.Errorf("%s is not a readable tar archive: %w", archive, err) + } + if hdr.Name != layer { + continue + } + var inner io.Reader = tr + if strings.HasSuffix(layer, ".gz") { + gz, gerr := gzip.NewReader(tr) + if gerr != nil { + return nil, fmt.Errorf("%s: layer %s is not gzip: %w", archive, layer, gerr) + } + defer func() { _ = gz.Close() }() + inner = gz + } + ltr := tar.NewReader(inner) + for { + lhdr, lerr := ltr.Next() + if errors.Is(lerr, io.EOF) { + return nil, nil + } + if lerr != nil { + return nil, fmt.Errorf("%s: reading layer %s: %w", archive, layer, lerr) + } + if strings.TrimPrefix(lhdr.Name, "./") != name { + continue + } + return readBounded(ltr, archive) + } + } +} + +// readBounded reads at most maxPackageBytes, so a package that decompresses +// to something enormous fails as a too-large package rather than as an +// exhausted machine. +func readBounded(r io.Reader, archive string) ([]byte, error) { + data, err := io.ReadAll(io.LimitReader(r, maxPackageBytes+1)) + if err != nil { + return nil, fmt.Errorf("reading %s: %w", archive, err) + } + if len(data) > maxPackageBytes { + return nil, fmt.Errorf("%s: package contents exceed the %d MiB read limit", archive, maxPackageBytes>>20) + } + return data, nil +} + +// parsePackageStream splits the multi-document package stream and keeps the +// meta object and the XRDs. +func parsePackageStream(ref string, data []byte) (*PackageContents, error) { + out := &PackageContents{Ref: ref} + for _, doc := range splitYAML(data) { + var m map[string]any + if err := sigsyaml.Unmarshal(doc, &m); err != nil { + // One unparseable document does not invalidate a package, and + // a package this tool cannot fully read is still usable for + // the XRDs it can. + continue + } + if len(m) == 0 { + continue + } + u := &unstructured.Unstructured{Object: m} + apiVersion, kind := u.GetAPIVersion(), u.GetKind() + switch { + case strings.HasPrefix(apiVersion, "meta.pkg.crossplane.io/"): + out.Meta = u + case kind == "CompositeResourceDefinition" && strings.HasPrefix(apiVersion, "apiextensions.crossplane.io/"): + out.XRDs = append(out.XRDs, u) + default: + out.Others++ + } + } + if out.Meta == nil && len(out.XRDs) == 0 && out.Others == 0 { + return nil, fmt.Errorf("%s: %s is empty", ref, packageFile) + } + return out, nil +} + +// splitYAML splits a multi-document stream on document separators. +func splitYAML(data []byte) [][]byte { + var out [][]byte + for _, part := range strings.Split(string(data), "\n---") { + trimmed := strings.TrimSpace(strings.TrimPrefix(part, "---")) + if trimmed == "" { + continue + } + out = append(out, []byte(trimmed)) + } + return out +} + +// XRDFromSource resolves an XRD from whichever schema source the caller +// named: a file, or a package plus an optional target. +// +// --package slots in exactly where --xrd does rather than being a new verb, +// which is what makes validate, analyze, test, diff and lint all gain it at +// once. +func XRDFromSource(xrdPath, packageRef, target string) (*unstructured.Unstructured, error) { + if packageRef == "" { + return LoadXRD(xrdPath) + } + pkg, err := ReadPackage(packageRef) + if err != nil { + return nil, err + } + return pkg.SelectXRD(target) +} diff --git a/internal/cli/xpkg_test.go b/internal/cli/xpkg_test.go new file mode 100644 index 0000000..f9bf354 --- /dev/null +++ b/internal/cli/xpkg_test.go @@ -0,0 +1,242 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +package cli + +import ( + "os" + "path/filepath" + "strings" + "testing" + + "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" +) + +const testPackage = "testdata/package/platform.xpkg" + +// The fixture is a real `crossplane xpkg build` output, so this asserts the +// reader agrees with the format Crossplane actually produces rather than +// with the assumptions the reader was written from. +func TestReadPackage_ReadsARealXPKG(t *testing.T) { + pkg, err := ReadPackage(testPackage) + if err != nil { + t.Fatalf("reading the package: %v", err) + } + if pkg.Meta == nil { + t.Error("the package meta object was not found") + } else if got := pkg.Meta.GetKind(); got != "Configuration" { + t.Errorf("meta kind = %q, want Configuration", got) + } + names := pkg.XRDNames() + if len(names) != 2 { + t.Fatalf("XRDs = %v, want the fixture's two", names) + } + if names[0] != "xbuckets.example.org" || names[1] != "xwidgets.example.org" { + t.Errorf("XRD names = %v", names) + } + // The XRDs have to arrive usable, not just counted. + for _, x := range pkg.XRDs { + if _, found, _ := unstructured.NestedSlice(x.Object, "spec", "versions"); !found { + t.Errorf("%s has no spec.versions, so it did not survive the round trip", xrdName(x)) + } + } +} + +// Picking the first of several would make the answer depend on the order the +// package happened to be built in. +func TestPackageContents_SelectXRDRequiresATargetWhenAmbiguous(t *testing.T) { + pkg, err := ReadPackage(testPackage) + if err != nil { + t.Fatal(err) + } + _, err = pkg.SelectXRD("") + if err == nil { + t.Fatal("a package with two XRDs selected one without being told which") + } + for _, want := range []string{"--target", "xwidgets.example.org", "xbuckets.example.org"} { + if !strings.Contains(err.Error(), want) { + t.Errorf("error should name the candidates, got: %v", err) + } + } + + got, err := pkg.SelectXRD("xwidgets.example.org") + if err != nil { + t.Fatalf("selecting by name: %v", err) + } + if xrdName(got) != "xwidgets.example.org" { + t.Errorf("selected %s", xrdName(got)) + } + + _, err = pkg.SelectXRD("xnope.example.org") + if err == nil || !strings.Contains(err.Error(), "ships no XRD named") { + t.Errorf("an unknown target should say so and list what is there: %v", err) + } +} + +// A single-XRD package needs no --target: that is the common case and +// requiring a flag for it would be friction with no purpose. +func TestPackageContents_SelectXRDIsUnambiguousWithOne(t *testing.T) { + one, err := ReadPackage(testPackage) + if err != nil { + t.Fatal(err) + } + single := &PackageContents{Ref: "single.xpkg", XRDs: one.XRDs[:1]} + got, err := single.SelectXRD("") + if err != nil { + t.Fatalf("a single-XRD package still demanded a target: %v", err) + } + if got == nil { + t.Error("no XRD returned") + } +} + +// The three failures have to be distinguishable, because the fix for each is +// different: find the file, build a package, add an XRD to it. +func TestReadPackage_DistinguishesItsFailures(t *testing.T) { + if _, err := ReadPackage("testdata/package/nope.xpkg"); err == nil || !strings.Contains(err.Error(), "reading package") { + t.Errorf("a missing file should say so: %v", err) + } + if _, err := ReadPackage("testdata"); err == nil || !strings.Contains(err.Error(), "is a directory") { + t.Errorf("a directory should say so: %v", err) + } + + dir := t.TempDir() + notAPackage := filepath.Join(dir, "plain.xpkg") + if err := os.WriteFile(notAPackage, []byte("this is not a tar at all"), 0o600); err != nil { + t.Fatal(err) + } + if _, err := ReadPackage(notAPackage); err == nil { + t.Error("a non-tar file was accepted as a package") + } +} + +// A package with no XRDs is not the same failure as a file that is not a +// package, and a config cannot be checked against either. +func TestPackageContents_NoXRDsSaysWhatIsThere(t *testing.T) { + empty := &PackageContents{Ref: "p.xpkg", Others: 4} + _, err := empty.SelectXRD("") + if err == nil || !strings.Contains(err.Error(), "no CompositeResourceDefinitions") { + t.Errorf("want a message naming the absence: %v", err) + } + if !strings.Contains(err.Error(), "4 other object(s)") { + t.Errorf("want the count of what the package does contain: %v", err) + } +} + +// The registry and cluster forms are named rather than silently +// unsupported, and the error says what to do instead. +func TestReadPackage_NamesTheUnimplementedForms(t *testing.T) { + for ref, want := range map[string]string{ + "ghcr.io/org/platform:v1.4.0": "crossplane xpkg pull", + "configuration/platform": "not implemented yet", + "configurationrevision/platform-abc": "not implemented yet", + } { + _, err := ReadPackage(ref) + if err == nil { + t.Errorf("%s was accepted", ref) + continue + } + if !strings.Contains(err.Error(), want) { + t.Errorf("%s: error %q does not contain %q", ref, err, want) + } + } +} + +// A path must not be mistaken for a registry reference, or a local package +// would be rejected with advice about pulling it. +func TestLooksLikeImageRef(t *testing.T) { + for ref, want := range map[string]bool{ + "ghcr.io/org/platform:v1": true, + "localhost:5000/p:v1": true, + "./platform.xpkg": false, + "/abs/platform.xpkg": false, + "platform.xpkg": false, // a bare filename, not a registry + testPackage: false, + } { + if got := looksLikeImageRef(ref); got != want { + t.Errorf("looksLikeImageRef(%q) = %v, want %v", ref, got, want) + } + } +} + +// --package has to slot in exactly where --xrd does, which is the whole +// reason it is a schema source rather than a new verb. +func TestXRDFromSource_PackageAndFileAreInterchangeable(t *testing.T) { + fromPkg, err := XRDFromSource("", testPackage, "xwidgets.example.org") + if err != nil { + t.Fatalf("from package: %v", err) + } + fromFile, err := XRDFromSource("../../examples/crossplane-xr-multiversion/03-promote-v2/xrd.yaml", "", "") + if err != nil { + t.Fatalf("from file: %v", err) + } + if xrdName(fromPkg) != xrdName(fromFile) { + t.Errorf("package gave %s, file gave %s", xrdName(fromPkg), xrdName(fromFile)) + } +} + +// The composition that matters: schemas from the package about to be +// installed, checked by the same analysis a file would get. +func TestRunAnalyzeFrom_WorksAgainstAPackage(t *testing.T) { + out, err := RunAnalyzeFrom("", "", "../../examples/crossplane-xr-multiversion/03-promote-v2/xrdconversionconfig.yaml", testPackage, "xwidgets.example.org") + if err != nil { + t.Fatalf("analyze from package: %v", err) + } + if out.Resource != "xwidgets.example.org" { + t.Errorf("resource = %q", out.Resource) + } + if len(out.Spokes) == 0 { + t.Error("no spokes analyzed") + } +} + +func TestRunTest_WorksAgainstAPackage(t *testing.T) { + rep, err := RunTest(TestOptions{ + PackagePath: testPackage, + PackageTarget: "xwidgets.example.org", + ConfigPath: "../../examples/crossplane-xr-multiversion/03-promote-v2/xrdconversionconfig.yaml", + SamplesDir: "../../examples/crossplane-xr-multiversion/03-promote-v2/samples", + Quiet: true, + }) + if err != nil { + t.Fatalf("test from package: %v", err) + } + if rep.Summary.Samples == 0 { + t.Error("no samples tested") + } + if rep.Summary.Errors != 0 { + t.Errorf("errors = %d, want a clean run", rep.Summary.Errors) + } +} + +// A whole package against a whole config tree, one command, is the +// Configuration repository's gate. +func TestRunLint_PairsATreeAgainstAPackage(t *testing.T) { + dir := t.TempDir() + mustWrite(t, filepath.Join(dir, "conversion.yaml"), + mustRead(t, "../../examples/crossplane-xr-multiversion/03-promote-v2/xrdconversionconfig.yaml")) + + rep, err := RunLint(LintOptions{Paths: []string{dir}, PackageRef: testPackage}) + if err != nil { + t.Fatalf("lint against a package: %v", err) + } + if rep.Unpaired != 0 { + t.Fatalf("a config was not paired against the package's XRDs: %+v", rep.Results) + } + if !strings.Contains(rep.Results[0].Schema, "platform.xpkg") { + t.Errorf("schema should name the package, not a temporary path: %q", rep.Results[0].Schema) + } +} diff --git a/mkdocs.yml b/mkdocs.yml index cc4a43e..60ea6b2 100644 --- a/mkdocs.yml +++ b/mkdocs.yml @@ -87,6 +87,7 @@ nav: - HA checklist: operations/ha-checklist.md - Capacity planning: operations/capacity.md - HPA on conversion QPS: operations/hpa-custom-metrics.md + - Configuration repo CI gate: gitops/configuration-ci.md - Fleet CI (many kubecontexts): gitops/fleet-ci.md - GitOps operator sync (Flux/Argo): gitops/operator-sync.md - CLI Reference: cli.md From 2a8f364e6205381083fc25f1139401d83887ae19 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:22:40 +0300 Subject: [PATCH 07/24] feat(actions): setup-convctl, a composite Action that verifies by default MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit There is no supported way to install convctl in a pipeline. The reference workflow this repository ships has an install step that reads, literally, `echo "install convctl and place it on PATH" >&2; exit 1`. Meanwhile the release already produces a cosign-signed checksums.txt with its certificate and signature, per-archive CycloneDX SBOMs, and build-provenance attestations — an investment almost nobody benefits from, because verifying it by hand means reading the release notes and writing eight lines of cosign verify-blob. An Action that verifies by default is what turns that work into something every consumer gets without reading anything. - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 with: version: v0.5.0 Four things it does that a naive installer would not: Verification is on by default, and the certificate identity is pinned to this repository's release workflow at a tag rather than to a wildcard. A signature from any other workflow in any other repository is precisely what this is meant to reject, and a wildcard identity would accept one. A cache hit still verifies. The cached branch is the one that runs in practice and the one that silently rots, so a poisoned cache that a hit could launder would make the whole thing decorative. The step fetches whatever signature material the cache did not carry and verifies the archive it is about to extract. `latest` is resolved to a concrete tag and pinned in the step summary, so a re-run six months later is explainable rather than mysterious. And the installed binary's own `version` output is checked against the tag that was requested — an installer that puts the wrong binary on PATH has failed even though every step was green. shell: bash throughout, on all three runner OSes, so the Windows path is the same script rather than a second program nobody exercises. Nothing needs sudo. test/actions asserts the shape statically: composite, documented inputs and outputs, a shell on every run step, pinned dependencies, set -euo pipefail everywhere, the pinned identity and issuer, that the checksum is actually checked, and that verification is not conditioned on the cache. actionlint treats a composite action's action.yml as a malformed workflow, so without this the Actions — the project's public CI surface — would have no check at all outside a live run. Closes #134 Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/setup-convctl/README.md | 50 +++++ .github/actions/setup-convctl/action.yml | 217 ++++++++++++++++++++ test/actions/actions_test.go | 244 +++++++++++++++++++++++ 3 files changed, 511 insertions(+) create mode 100644 .github/actions/setup-convctl/README.md create mode 100644 .github/actions/setup-convctl/action.yml create mode 100644 test/actions/actions_test.go diff --git a/.github/actions/setup-convctl/README.md b/.github/actions/setup-convctl/README.md new file mode 100644 index 0000000..3c92670 --- /dev/null +++ b/.github/actions/setup-convctl/README.md @@ -0,0 +1,50 @@ +# `setup-convctl` + +Installs a **verified** `convctl` onto `PATH`. + +```yaml +- uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 + with: + version: v0.5.0 # or "latest" (the default) +- run: convctl lint ./platform/ +``` + +## Inputs + +| Input | Default | Meaning | +|---|---|---| +| `version` | `latest` | a release tag, or `latest` | +| `verify` | `true` | verify the cosign signature on `checksums.txt`, then the archive's checksum against it | +| `token` | `${{ github.token }}` | for the releases API, to avoid anonymous rate limits | + +## Outputs + +| Output | Meaning | +|---|---| +| `version` | the resolved tag — `latest` is pinned in the step summary, so a re-run is explainable | +| `path` | the installed binary | +| `cache-hit` | whether it came from the runner cache | + +## Verification + +On by default. That is the point: the release already produces a cosign-signed +`checksums.txt`, per-archive SBOMs and build-provenance attestations, and +almost nobody benefits from them because verifying by hand means reading the +release notes and writing eight lines of `cosign verify-blob`. + +The certificate identity is pinned to **this repository's release workflow at +a tag**, not to a wildcard — a signature from any other workflow in any other +repository is exactly what this rejects. The OIDC issuer is pinned too. + +**A cache hit still verifies.** The cached branch is the one that runs in +practice, and a poisoned cache that a cache hit could launder would make the +whole thing decorative. + +Set `verify: false` only if you have a reason; cosign is not installed at all +in that case, so it costs nothing to leave on. + +## Platforms + +`ubuntu-latest`, `macos-latest` and `windows-latest`, on amd64 and arm64. +`shell: bash` throughout, so the same script runs on all three rather than a +Windows variant nobody exercises. Nothing needs `sudo`. diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml new file mode 100644 index 0000000..8b955ad --- /dev/null +++ b/.github/actions/setup-convctl/action.yml @@ -0,0 +1,217 @@ +name: Set up convctl +description: >- + Install a verified convctl onto PATH. Verifies the cosign signature on the + release's checksums.txt and the archive's own checksum, by default. + +inputs: + version: + description: 'Release tag to install (e.g. v0.5.0), or "latest".' + required: false + default: latest + verify: + description: >- + Verify the cosign signature on checksums.txt and the archive's checksum + against it. Defaults to true — an action that verifies by default is the + point, because it turns the signing work already done into something + every consumer gets without reading anything. + required: false + default: "true" + token: + description: >- + Token for the GitHub releases API, to avoid anonymous rate limits. + required: false + default: ${{ github.token }} + +outputs: + version: + description: The resolved release tag. + value: ${{ steps.resolve.outputs.version }} + path: + description: Path of the installed binary. + value: ${{ steps.install.outputs.path }} + cache-hit: + description: Whether the binary came from the runner cache. + value: ${{ steps.cache.outputs.cache-hit }} + +runs: + using: composite + steps: + # Resolve "latest" to a concrete tag and pin it in the step summary, so a + # re-run six months from now is explainable rather than mysterious. + - name: Resolve version and platform + id: resolve + shell: bash + env: + GH_TOKEN: ${{ inputs.token }} + REQUESTED: ${{ inputs.version }} + run: | + set -euo pipefail + + if [ "$REQUESTED" = "latest" ]; then + version="$(gh api repos/TeraSky-OSS/declarative-conversion-operator/releases/latest --jq .tag_name)" + if [ -z "$version" ] || [ "$version" = "null" ]; then + echo "::error::could not resolve the latest convctl release" >&2 + exit 1 + fi + else + version="$REQUESTED" + fi + + case "$RUNNER_OS" in + Linux) os=linux ; ext=tar.gz ; bin=convctl ;; + macOS) os=darwin ; ext=tar.gz ; bin=convctl ;; + Windows) os=windows ; ext=zip ; bin=convctl.exe ;; + *) echo "::error::unsupported runner OS: $RUNNER_OS" >&2; exit 1 ;; + esac + case "$RUNNER_ARCH" in + X64) arch=amd64 ;; + ARM64) arch=arm64 ;; + *) echo "::error::unsupported runner architecture: $RUNNER_ARCH" >&2; exit 1 ;; + esac + + # The archive name carries the version without its leading v, which + # is what GoReleaser's .Version produces. + bare="${version#v}" + archive="declarative-conversion-operator-cli_${bare}_${os}_${arch}.${ext}" + + { + echo "version=$version" + echo "os=$os" + echo "arch=$arch" + echo "ext=$ext" + echo "bin=$bin" + echo "archive=$archive" + echo "dir=$RUNNER_TEMP/convctl-$bare" + } >> "$GITHUB_OUTPUT" + + echo "convctl $version ($os/$arch)" >> "$GITHUB_STEP_SUMMARY" + + - name: Restore cached binary + id: cache + uses: actions/cache@v4 + with: + path: ${{ steps.resolve.outputs.dir }} + key: convctl-${{ steps.resolve.outputs.version }}-${{ steps.resolve.outputs.os }}-${{ steps.resolve.outputs.arch }} + + - name: Download release assets + if: steps.cache.outputs.cache-hit != 'true' + shell: bash + env: + GH_TOKEN: ${{ inputs.token }} + VERSION: ${{ steps.resolve.outputs.version }} + ARCHIVE: ${{ steps.resolve.outputs.archive }} + DIR: ${{ steps.resolve.outputs.dir }} + VERIFY: ${{ inputs.verify }} + run: | + set -euo pipefail + mkdir -p "$DIR" + cd "$DIR" + + assets=("$ARCHIVE" checksums.txt) + if [ "$VERIFY" = "true" ]; then + assets+=(checksums.txt.sig checksums.txt.pem) + fi + for asset in "${assets[@]}"; do + gh release download "$VERSION" \ + --repo TeraSky-OSS/declarative-conversion-operator \ + --pattern "$asset" --clobber + done + + # Cosign is only installed when it is going to be used, so verify: false + # costs nothing beyond the download. + - name: Install cosign + if: inputs.verify == 'true' + uses: sigstore/cosign-installer@v4 + + # Verification runs on a cache hit too. A poisoned cache that could be + # laundered by a cache hit would make the whole verification decorative: + # the cached branch is the one that runs in practice. + - name: Verify signature and checksum + if: inputs.verify == 'true' + shell: bash + env: + GH_TOKEN: ${{ inputs.token }} + VERSION: ${{ steps.resolve.outputs.version }} + ARCHIVE: ${{ steps.resolve.outputs.archive }} + DIR: ${{ steps.resolve.outputs.dir }} + CACHE_HIT: ${{ steps.cache.outputs.cache-hit }} + run: | + set -euo pipefail + cd "$DIR" + + # A cache holds the archive but not necessarily the signature + # material, so fetch whatever is missing before verifying. + for asset in checksums.txt checksums.txt.sig checksums.txt.pem "$ARCHIVE"; do + if [ ! -f "$asset" ]; then + gh release download "$VERSION" \ + --repo TeraSky-OSS/declarative-conversion-operator \ + --pattern "$asset" --clobber + fi + done + + # The certificate identity is pinned to this repository's release + # workflow at a tag, not to a wildcard: a signature from any other + # workflow in any other repository is exactly what this is meant to + # reject. + cosign verify-blob \ + --certificate checksums.txt.pem \ + --signature checksums.txt.sig \ + --certificate-identity-regexp '^https://github\.com/terasky-oss/declarative-conversion-operator/\.github/workflows/release\.yml@refs/tags/.*$' \ + --certificate-oidc-issuer https://token.actions.githubusercontent.com \ + checksums.txt + + # Then the archive against the now-trusted checksums file. grep for + # this archive's line only: checksums.txt covers every platform, and + # sha256sum -c would fail on the absent ones. + grep " \{1,2\}${ARCHIVE}\$" checksums.txt > archive.sha256 + if [ ! -s archive.sha256 ]; then + echo "::error::${ARCHIVE} has no entry in the signed checksums.txt" >&2 + exit 1 + fi + if command -v sha256sum >/dev/null 2>&1; then + sha256sum -c archive.sha256 + else + shasum -a 256 -c archive.sha256 + fi + + if [ "$CACHE_HIT" = "true" ]; then + echo "verified the cached archive against the signed checksums" >> "$GITHUB_STEP_SUMMARY" + fi + + - name: Extract and add to PATH + id: install + shell: bash + env: + ARCHIVE: ${{ steps.resolve.outputs.archive }} + EXT: ${{ steps.resolve.outputs.ext }} + BIN: ${{ steps.resolve.outputs.bin }} + DIR: ${{ steps.resolve.outputs.dir }} + run: | + set -euo pipefail + cd "$DIR" + + if [ ! -f "$BIN" ]; then + if [ "$EXT" = "zip" ]; then + unzip -o "$ARCHIVE" >/dev/null + else + tar -xzf "$ARCHIVE" + fi + fi + chmod +x "$BIN" + + echo "$DIR" >> "$GITHUB_PATH" + echo "path=$DIR/$BIN" >> "$GITHUB_OUTPUT" + + - name: Verify it runs + shell: bash + env: + EXPECTED: ${{ steps.resolve.outputs.version }} + BINARY: ${{ steps.install.outputs.path }} + run: | + set -euo pipefail + got="$("$BINARY" version | awk '{print $1}')" + if [ "$got" != "$EXPECTED" ]; then + echo "::error::installed convctl reports $got, expected $EXPECTED" >&2 + exit 1 + fi + echo "convctl $got is on PATH" diff --git a/test/actions/actions_test.go b/test/actions/actions_test.go new file mode 100644 index 0000000..aa929b8 --- /dev/null +++ b/test/actions/actions_test.go @@ -0,0 +1,244 @@ +/* +Copyright 2026 The declarative-conversion-operator Authors. + +Licensed under the Apache License, Version 2.0 (the "License"); +you may not use this file except in compliance with the License. +You may obtain a copy of the License at + + http://www.apache.org/licenses/LICENSE-2.0 + +Unless required by applicable law or agreed to in writing, software +distributed under the License is distributed on an "AS IS" BASIS, +WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. +See the License for the specific language governing permissions and +limitations under the License. +*/ + +// Package actions_test checks the shipped GitHub Actions statically. +// +// actionlint validates workflows, but treats a composite action's action.yml +// as a malformed workflow — so the Actions this repository publishes, which +// are its public CI surface, would otherwise have no check at all outside a +// live workflow run. +package actions_test + +import ( + "os" + "path/filepath" + "strings" + "testing" + + yaml "go.yaml.in/yaml/v3" +) + +const actionsDir = "../../.github/actions" + +type actionDef struct { + Name string `yaml:"name"` + Description string `yaml:"description"` + Inputs map[string]struct { + Description string `yaml:"description"` + Required bool `yaml:"required"` + Default string `yaml:"default"` + } `yaml:"inputs"` + Outputs map[string]struct { + Description string `yaml:"description"` + Value string `yaml:"value"` + } `yaml:"outputs"` + Runs struct { + Using string `yaml:"using"` + Steps []struct { + Name string `yaml:"name"` + ID string `yaml:"id"` + Uses string `yaml:"uses"` + Shell string `yaml:"shell"` + Run string `yaml:"run"` + If string `yaml:"if"` + } `yaml:"steps"` + } `yaml:"runs"` +} + +func loadActions(t *testing.T) map[string]actionDef { + t.Helper() + entries, err := os.ReadDir(actionsDir) + if err != nil { + t.Fatalf("reading %s: %v", actionsDir, err) + } + out := map[string]actionDef{} + for _, e := range entries { + if !e.IsDir() { + continue + } + path := filepath.Join(actionsDir, e.Name(), "action.yml") + data, err := os.ReadFile(path) // #nosec G304 -- a path inside the repo + if err != nil { + t.Fatalf("%s has no action.yml: %v", e.Name(), err) + } + var def actionDef + if err := yaml.Unmarshal(data, &def); err != nil { + t.Fatalf("%s: %v", path, err) + } + out[e.Name()] = def + } + if len(out) == 0 { + t.Fatal("no actions found") + } + return out +} + +// Every action has to be a well-formed composite action with documented +// inputs, because an undocumented input is an input nobody uses correctly. +func TestActions_AreWellFormedAndDocumented(t *testing.T) { + for name, def := range loadActions(t) { + t.Run(name, func(t *testing.T) { + if def.Name == "" || def.Description == "" { + t.Error("action has no name or description") + } + if def.Runs.Using != "composite" { + t.Errorf("runs.using = %q, want composite", def.Runs.Using) + } + if len(def.Runs.Steps) == 0 { + t.Fatal("action has no steps") + } + for input, spec := range def.Inputs { + if strings.TrimSpace(spec.Description) == "" { + t.Errorf("input %q has no description", input) + } + } + for output, spec := range def.Outputs { + if strings.TrimSpace(spec.Description) == "" { + t.Errorf("output %q has no description", output) + } + if spec.Value == "" { + t.Errorf("output %q has no value, so it is always empty", output) + } + } + }) + } +} + +// `shell:` is required on every run step of a composite action — GitHub +// fails the action at runtime without it, which is a failure a consumer +// sees rather than us. +func TestActions_EveryRunStepDeclaresAShell(t *testing.T) { + for name, def := range loadActions(t) { + for i, step := range def.Runs.Steps { + if step.Run != "" && step.Shell == "" { + t.Errorf("%s step %d (%s) has run: without shell:", name, i+1, step.Name) + } + } + } +} + +// bash on all three runners, so the same script is what ships. pwsh or sh +// would mean the Windows path is a different program nobody tests. +func TestActions_UseBashEverywhere(t *testing.T) { + for name, def := range loadActions(t) { + for _, step := range def.Runs.Steps { + if step.Shell != "" && step.Shell != "bash" { + t.Errorf("%s step %q uses shell %q; bash keeps one script across all three runners", name, step.Name, step.Shell) + } + } + } +} + +// A composite action runs in the consumer's repository, so an unpinned +// dependency is their supply chain, not ours. +func TestActions_PinTheirDependencies(t *testing.T) { + for name, def := range loadActions(t) { + for _, step := range def.Runs.Steps { + if step.Uses == "" || strings.HasPrefix(step.Uses, "./") { + continue + } + if !strings.Contains(step.Uses, "@") { + t.Errorf("%s: %q is not pinned to a version", name, step.Uses) + } + } + } +} + +// Every `run:` script sets the failure modes bash does not set by default. +// Without `set -e` a failing command in the middle of a step leaves the step +// green, which for a verification step is the worst possible outcome. +func TestActions_ScriptsFailFast(t *testing.T) { + for name, def := range loadActions(t) { + for _, step := range def.Runs.Steps { + if step.Run == "" { + continue + } + if !strings.Contains(step.Run, "set -euo pipefail") && !strings.Contains(step.Run, "set -eo pipefail") { + t.Errorf("%s step %q does not set -euo pipefail, so a failing command mid-script would still pass", name, step.Name) + } + } + } +} + +// Verification is the reason the setup action exists, so its shape is +// asserted specifically rather than only in a live run. +func TestSetupConvctl_VerifiesByDefaultAndPinsTheIdentity(t *testing.T) { + def, ok := loadActions(t)["setup-convctl"] + if !ok { + t.Fatal("setup-convctl is missing") + } + if got := def.Inputs["verify"].Default; got != "true" { + t.Errorf("verify defaults to %q; an action that verifies by default is the whole point", got) + } + + var verifyStep string + for _, s := range def.Runs.Steps { + if strings.Contains(s.Run, "cosign verify-blob") { + verifyStep = s.Run + } + } + if verifyStep == "" { + t.Fatal("no step runs cosign verify-blob") + } + // A wildcard identity would accept a signature from any workflow in any + // repository, which is the thing being defended against. + if !strings.Contains(verifyStep, "declarative-conversion-operator/\\.github/workflows/release\\.yml") { + t.Error("the certificate identity is not pinned to this repository's release workflow") + } + if !strings.Contains(verifyStep, "--certificate-oidc-issuer https://token.actions.githubusercontent.com") { + t.Error("the OIDC issuer is not pinned") + } + // And the checksum has to actually be checked against the file whose + // signature was just verified. + if !strings.Contains(verifyStep, "sha256sum -c") && !strings.Contains(verifyStep, "shasum -a 256 -c") { + t.Error("the archive's checksum is never verified against checksums.txt") + } + + // The cached path is the one that runs in practice; a cache hit that + // skipped verification would let a poisoned cache launder itself. + for _, s := range def.Runs.Steps { + if strings.Contains(s.Run, "cosign verify-blob") { + if strings.Contains(s.If, "cache-hit") { + t.Error("verification is skipped on a cache hit, so a poisoned cache would be laundered by one") + } + } + } +} + +// The three runner OSes are the reason the archive naming and extraction +// logic exist; a missing case would be a consumer's broken pipeline. +func TestSetupConvctl_HandlesEveryRunnerPlatform(t *testing.T) { + def := loadActions(t)["setup-convctl"] + var resolve string + for _, s := range def.Runs.Steps { + if s.ID == "resolve" { + resolve = s.Run + } + } + if resolve == "" { + t.Fatal("no resolve step") + } + for _, want := range []string{"Linux", "macOS", "Windows", "X64", "ARM64", "zip", "tar.gz", "convctl.exe"} { + if !strings.Contains(resolve, want) { + t.Errorf("the resolve step does not handle %q", want) + } + } + // An unrecognised platform has to fail loudly rather than produce a + // nonsense archive name. + if strings.Count(resolve, "unsupported") < 2 { + t.Error("an unknown OS or architecture is not rejected") + } +} From e599919ec3a73a60dca69231ef1a4b4f4b1a86d8 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:24:11 +0300 Subject: [PATCH 08/24] feat(actions): convctl-test, with JUnit, a job summary and annotations MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Wiring convctl test into a pipeline by hand means installing the CLI, picking the flags, remembering --output junit --output-file, uploading the artifact, and then discovering that the failure is a line in a log rather than something attached to the diff the reviewer is already reading. One step now does all of it: the JUnit artifact, a job summary table of passed / unacknowledged-loss / error counts, and annotations on the pull-request diff. The annotations are the part that needed a decision, and the decision was to not make them here. convctl --output github emits the workflow commands itself, with the file and line of the rule that produced each finding, so this Action relays them. Parsing YAML inside the Action to re-derive a location the tool already computes would be a second implementation of the same mapping, drifting from the first. Two details that only fail in anger, both found by writing the checks first: The step runs with set -euo pipefail, except around the convctl invocation itself, where the exit code is the result rather than an error — the Action's contract is to preserve it, including the whole --fail-on matrix. The static test asserts fail-fast on every script, which is what surfaced the gap. Optional flags are assembled with if-blocks rather than `[ -n "$X" ] && args+=(...)`. Under set -e a false test as the last command of a line exits the script, so the terse form would have silently stopped the step the first time an optional input was empty — a bug that only appears for consumers who do not set every input, which is all of them. The artifact uploads under always(), because the report is most worth having when the step failed. The summary says plainly when a --live run was sampled and out of what population: a JUnit reporter showing green tests is where a sampled run is most likely to be mistaken for an exhaustive one. kubeconfig is written under umask 077 before creation rather than chmod-ed after, since between creation and chmod the file is briefly world-readable, and it is never echoed. Verified end to end against a stub convctl: exit code preserved, counts parsed out of the JUnit report, sampling surfaced in the summary. Closes #135 Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/convctl-test/README.md | 45 +++++ .github/actions/convctl-test/action.yml | 230 ++++++++++++++++++++++++ 2 files changed, 275 insertions(+) create mode 100644 .github/actions/convctl-test/README.md create mode 100644 .github/actions/convctl-test/action.yml diff --git a/.github/actions/convctl-test/README.md b/.github/actions/convctl-test/README.md new file mode 100644 index 0000000..a645413 --- /dev/null +++ b/.github/actions/convctl-test/README.md @@ -0,0 +1,45 @@ +# `convctl-test` + +Runs `convctl test` and puts the result where a reviewer will see it. + +```yaml +- uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-test@v1 + with: + config: apis/widgets/conversion.yaml + xrd: apis/widgets/xrd.yaml + samples: apis/widgets/samples/ + validate-output: "true" +``` + +Three outputs beyond the exit code: + +- a **JUnit artifact**, uploaded with `always()` — the report is most worth + having when the step failed, which is exactly when a naive step after a + failing one would not run; +- a **job summary** table of passed / unacknowledged-loss / error counts, + which also says so when a `--live` run was sampled; +- **annotations** on the pull-request diff, pointing at the line of the config + that produced each finding. + +The annotations come from `convctl --output github`, which emits the workflow +commands itself. This Action relays them; it does not parse YAML to re-derive +a location the tool already knows. + +## Inputs + +Schema source: `xrd`, `crd`, or `package` (+ `target`). +Samples: `samples`, or `live` (+ `kubeconfig`, `context`, `max-samples`, +`sample-strategy`). +Behaviour: `fail-on`, `strict`, `validate-output`, `concurrency`, `version`, +`upload-artifact`, `artifact-name`, `annotate`. + +## Outputs + +`exit-code` (convctl's own, preserved), `report-path`, `pass`, +`unacknowledged-loss`, `errors`. + +## Secrets + +`kubeconfig` is written with `umask 077` **before** the file is created rather +than `chmod`-ed afterwards — between creation and chmod the file is briefly +world-readable — and is never echoed. diff --git a/.github/actions/convctl-test/action.yml b/.github/actions/convctl-test/action.yml new file mode 100644 index 0000000..ed9f5f0 --- /dev/null +++ b/.github/actions/convctl-test/action.yml @@ -0,0 +1,230 @@ +name: convctl test +description: >- + Run convctl test and put the result where a reviewer will see it: a JUnit + artifact, a job summary, and annotations on the pull-request diff. + +inputs: + config: + description: Path to the XRDConversionConfig or CRDConversionConfig. + required: true + xrd: + description: Path to the XRD (for an XRDConversionConfig). + required: false + default: "" + crd: + description: Path to the CRD (for a CRDConversionConfig). + required: false + default: "" + package: + description: >- + Read the target XRD from a Crossplane package (.xpkg) instead of a file. + required: false + default: "" + target: + description: With package, the XRD to select when it ships several. + required: false + default: "" + samples: + description: Directory of sample objects. Mutually exclusive with live. + required: false + default: "" + live: + description: Test against every live object in a cluster. + required: false + default: "false" + kubeconfig: + description: >- + Kubeconfig contents for a live run. Written to a file with mode 600 and + never echoed. + required: false + default: "" + context: + description: Kubeconfig context for a live run. + required: false + default: "" + max-samples: + description: With live, cap how many objects are tested. + required: false + default: "" + sample-strategy: + description: "With max-samples: first, random, or newest." + required: false + default: "" + fail-on: + description: "Failure threshold: none, warn, or loss (default)." + required: false + default: loss + strict: + description: Escalate coverage gaps to failures. + required: false + default: "false" + validate-output: + description: >- + Validate every converted object against the destination schema. + required: false + default: "false" + concurrency: + description: Parallel workers. + required: false + default: "" + version: + description: convctl release tag to install. + required: false + default: latest + upload-artifact: + description: Upload the JUnit report as a workflow artifact. + required: false + default: "true" + artifact-name: + description: Name for the uploaded artifact. + required: false + default: convctl-test-report + annotate: + description: >- + Emit annotations pointing at the config lines that produced each + finding, so failures land on the diff rather than in a log. + required: false + default: "true" + +outputs: + exit-code: + description: convctl's own exit code, preserved. + value: ${{ steps.run.outputs.exit-code }} + report-path: + description: Path of the JUnit report. + value: ${{ steps.run.outputs.report-path }} + pass: + description: Passing conversion paths. + value: ${{ steps.run.outputs.pass }} + unacknowledged-loss: + description: Conversion paths that lost a field no rule declares lossy. + value: ${{ steps.run.outputs.unacknowledged-loss }} + errors: + description: Conversion errors. + value: ${{ steps.run.outputs.errors }} + +runs: + using: composite + steps: + - name: Set up convctl + uses: ./.github/actions/setup-convctl + with: + version: ${{ inputs.version }} + + - name: Write kubeconfig + if: inputs.kubeconfig != '' + shell: bash + env: + KUBECONFIG_CONTENTS: ${{ inputs.kubeconfig }} + run: | + set -euo pipefail + # umask before the write, not chmod after: between creation and + # chmod the file is briefly world-readable, and the runner is not + # the only thing on the machine. + umask 077 + printf '%s' "$KUBECONFIG_CONTENTS" > "$RUNNER_TEMP/kubeconfig" + echo "KUBECONFIG=$RUNNER_TEMP/kubeconfig" >> "$GITHUB_ENV" + + - name: Run convctl test + id: run + shell: bash + env: + CONFIG: ${{ inputs.config }} + XRD: ${{ inputs.xrd }} + CRD: ${{ inputs.crd }} + PACKAGE: ${{ inputs.package }} + TARGET: ${{ inputs.target }} + SAMPLES: ${{ inputs.samples }} + LIVE: ${{ inputs.live }} + CONTEXT: ${{ inputs.context }} + MAX_SAMPLES: ${{ inputs.max-samples }} + SAMPLE_STRATEGY: ${{ inputs.sample-strategy }} + FAIL_ON: ${{ inputs.fail-on }} + STRICT: ${{ inputs.strict }} + VALIDATE_OUTPUT: ${{ inputs.validate-output }} + CONCURRENCY: ${{ inputs.concurrency }} + ANNOTATE: ${{ inputs.annotate }} + run: | + set -euo pipefail + + report="$RUNNER_TEMP/convctl-report.junit.xml" + + # if-blocks rather than `[ -n "$X" ] && args+=(...)`: under set -e a + # false test as the last command of a line exits the script, so the + # terse form would silently stop the step the first time an optional + # input was empty. + args=(--config "$CONFIG" --fail-on "$FAIL_ON") + if [ -n "$XRD" ]; then args+=(--xrd "$XRD"); fi + if [ -n "$CRD" ]; then args+=(--crd "$CRD"); fi + if [ -n "$PACKAGE" ]; then args+=(--package "$PACKAGE"); fi + if [ -n "$TARGET" ]; then args+=(--target "$TARGET"); fi + if [ -n "$SAMPLES" ]; then args+=(--samples "$SAMPLES"); fi + if [ "$LIVE" = "true" ]; then args+=(--live); fi + if [ -n "$CONTEXT" ]; then args+=(--context "$CONTEXT"); fi + if [ -n "$MAX_SAMPLES" ]; then args+=(--max-samples "$MAX_SAMPLES"); fi + if [ -n "$SAMPLE_STRATEGY" ]; then args+=(--sample-strategy "$SAMPLE_STRATEGY"); fi + if [ "$STRICT" = "true" ]; then args+=(--strict); fi + if [ "$VALIDATE_OUTPUT" = "true" ]; then args+=(--validate-output); fi + if [ -n "$CONCURRENCY" ]; then args+=(--concurrency "$CONCURRENCY"); fi + + # The JUnit report first, for the artifact and the summary counts. + # This is the one command whose failure is data rather than an + # error: convctl's exit code is the result, and the Action's + # contract is to preserve it. + set +e + convctl test "${args[@]}" --output junit --output-file "$report" --quiet + code=$? + set -e + + # Then the same run rendered as annotations. convctl emits the + # workflow commands itself — including the file and line of the rule + # that produced each finding — so this Action relays rather than + # parsing YAML to re-derive what the tool already knows. + if [ "$ANNOTATE" = "true" ]; then + convctl test "${args[@]}" --output github --quiet || true + fi + + { + echo "exit-code=$code" + echo "report-path=$report" + } >> "$GITHUB_OUTPUT" + + # Counts for the summary and the outputs, read out of the JUnit + # report rather than recomputed. + if [ -f "$report" ]; then + tests=$(grep -o 'tests="[0-9]*"' "$report" | head -1 | grep -o '[0-9]*' || echo 0) + failures=$(grep -o 'failures="[0-9]*"' "$report" | head -1 | grep -o '[0-9]*' || echo 0) + errors=$(grep -o 'errors="[0-9]*"' "$report" | head -1 | grep -o '[0-9]*' || echo 0) + pass=$((tests - failures - errors)) + { + echo "pass=$pass" + echo "unacknowledged-loss=$failures" + echo "errors=$errors" + } >> "$GITHUB_OUTPUT" + { + echo "### convctl test" + echo "" + echo "| Outcome | Count |" + echo "|---|---|" + echo "| Passed | $pass |" + echo "| Unacknowledged loss | $failures |" + echo "| Errors | $errors |" + } >> "$GITHUB_STEP_SUMMARY" + if grep -q 'name="sampled" value="true"' "$report"; then + population=$(grep -o 'name="samplePopulation" value="[0-9]*"' "$report" | grep -o '[0-9]*' || echo "?") + echo "" >> "$GITHUB_STEP_SUMMARY" + echo "> **This run was sampled** — $tests of $population objects. It did not cover everything." >> "$GITHUB_STEP_SUMMARY" + fi + fi + + exit "$code" + + # always(): the report is most worth having when the step failed, which + # is exactly when a naive `uses:` after a failing step would not run. + - name: Upload report + if: always() && inputs.upload-artifact == 'true' + uses: actions/upload-artifact@v7 + with: + name: ${{ inputs.artifact-name }} + path: ${{ steps.run.outputs.report-path }} + if-no-files-found: warn From e9e336fa4bacf79915bfcdc43fb305341f706931 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:25:03 +0300 Subject: [PATCH 09/24] feat(actions): convctl-diff, a sticky pull-request comment with the delta MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit convctl diff produces exactly what a reviewer needs — which fields go from covered to uncovered, which rules claim which paths, which directions flip between lossless and lossy — and it is currently visible only to whoever opens the CI log. The Action renders it with --output markdown and upserts a comment keyed on a hidden marker, so repeated runs edit one comment rather than appending a new one on every push. comment-tag defaults to the config path, so two configs in one pull request get one comment each without anyone configuring anything. Three behaviours that are decisions rather than details: Exit 1 does not fail the job. A coverage delta is the thing being reported — it is the change about to be rolled out, not a defect — so the default is to report it and stay green. Exit 2 always fails, because a usage error or an unreachable cluster rendered as "no deltas" is a gate that passes precisely when it could not do its job. fail-on-delta: true is there for repositories that want the stricter reading. A run with no deltas edits the comment to say so rather than deleting it. A comment that vanishes reads as "the check stopped running", which is the wrong message about a check that ran and passed. A fork's pull_request token is read-only, so commenting is impossible. That is a notice and the delta stays in the job summary, not a failure: a red check a contributor cannot fix teaches them to ignore red checks. Closes #136 Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/convctl-diff/README.md | 48 ++++++ .github/actions/convctl-diff/action.yml | 186 ++++++++++++++++++++++++ 2 files changed, 234 insertions(+) create mode 100644 .github/actions/convctl-diff/README.md create mode 100644 .github/actions/convctl-diff/action.yml diff --git a/.github/actions/convctl-diff/README.md b/.github/actions/convctl-diff/README.md new file mode 100644 index 0000000..299a401 --- /dev/null +++ b/.github/actions/convctl-diff/README.md @@ -0,0 +1,48 @@ +# `convctl-diff` + +Runs `convctl diff` and upserts the delta as a **sticky** pull-request +comment — updated in place as the branch changes, rather than appended to on +every push. + +```yaml +permissions: + contents: read + pull-requests: write + +steps: + - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-diff@v1 + with: + config: apis/widgets/conversion.yaml + xrd: apis/widgets/xrd.yaml + live: "true" + kubeconfig: ${{ secrets.KUBECONFIG_PROD }} +``` + +## Exit codes are not all failures + +`convctl diff` exits 1 when it finds a delta. That is the thing being +reported, not a failure, so **the job stays green by default**. Exit 2 — a +usage error, or a cluster it could not reach — always fails, because a gate +that reported "no deltas" when it could not run would be silently useless. + +Set `fail-on-delta: true` to make any delta fail the job. + +## A comment that says "no deltas" rather than vanishing + +When there is nothing to report, the comment is edited to say so. A comment +that disappears reads as *"the check stopped running"*, which is the wrong +message to send about a check that ran and passed. + +## Multiple configs in one pull request + +`comment-tag` identifies the comment to update, and defaults to the config +path — so two configs in one pull request get one comment each with no +configuration. Set it explicitly if you want something else. + +## Forks + +A `pull_request` event from a fork has a read-only token. The Action emits a +**notice** and leaves the delta in the job summary, rather than failing: a red +check a contributor cannot fix teaches them to ignore red checks. + +Needs `pull-requests: write` to comment. diff --git a/.github/actions/convctl-diff/action.yml b/.github/actions/convctl-diff/action.yml new file mode 100644 index 0000000..0ffbf2a --- /dev/null +++ b/.github/actions/convctl-diff/action.yml @@ -0,0 +1,186 @@ +name: convctl diff +description: >- + Run convctl diff and upsert the coverage delta as a sticky pull-request + comment, updated in place as the branch changes. + +inputs: + config: + description: >- + Path to a conversion config. Pass twice (newline-separated) to compare + two files, or once with live to compare against the cluster. + required: true + xrd: + description: Path to the XRD. + required: false + default: "" + crd: + description: Path to the CRD. + required: false + default: "" + live: + description: Compare the config against what the cluster has. + required: false + default: "false" + kubeconfig: + description: Kubeconfig contents for a live comparison. + required: false + default: "" + context: + description: Kubeconfig context for a live comparison. + required: false + default: "" + comment: + description: Upsert the delta as a pull-request comment. + required: false + default: "true" + comment-tag: + description: >- + Identifies the comment to update. Defaults to the config path, so two + configs in one pull request get one comment each without configuration. + required: false + default: "" + fail-on-delta: + description: >- + Fail the job when a delta is found. Off by default — a coverage delta + is a review artifact, not automatically a failure. A usage or cluster + error (exit 2) always fails. + required: false + default: "false" + version: + description: convctl release tag to install. + required: false + default: latest + token: + description: Token used to upsert the comment. + required: false + default: ${{ github.token }} + +outputs: + exit-code: + description: convctl's own exit code (1 means deltas were found). + value: ${{ steps.run.outputs.exit-code }} + has-deltas: + description: Whether any delta was found. + value: ${{ steps.run.outputs.has-deltas }} + markdown-path: + description: Path of the rendered markdown. + value: ${{ steps.run.outputs.markdown-path }} + +runs: + using: composite + steps: + - name: Set up convctl + uses: ./.github/actions/setup-convctl + with: + version: ${{ inputs.version }} + + - name: Write kubeconfig + if: inputs.kubeconfig != '' + shell: bash + env: + KUBECONFIG_CONTENTS: ${{ inputs.kubeconfig }} + run: | + set -euo pipefail + umask 077 + printf '%s' "$KUBECONFIG_CONTENTS" > "$RUNNER_TEMP/kubeconfig" + echo "KUBECONFIG=$RUNNER_TEMP/kubeconfig" >> "$GITHUB_ENV" + + - name: Run convctl diff + id: run + shell: bash + env: + CONFIGS: ${{ inputs.config }} + XRD: ${{ inputs.xrd }} + CRD: ${{ inputs.crd }} + LIVE: ${{ inputs.live }} + CONTEXT: ${{ inputs.context }} + FAIL_ON_DELTA: ${{ inputs.fail-on-delta }} + run: | + set -euo pipefail + + md="$RUNNER_TEMP/convctl-diff.md" + args=() + while IFS= read -r cfg; do + if [ -n "$cfg" ]; then args+=(--config "$cfg"); fi + done <<< "$CONFIGS" + if [ -n "$XRD" ]; then args+=(--xrd "$XRD"); fi + if [ -n "$CRD" ]; then args+=(--crd "$CRD"); fi + if [ "$LIVE" = "true" ]; then args+=(--live); fi + if [ -n "$CONTEXT" ]; then args+=(--context "$CONTEXT"); fi + + set +e + convctl diff "${args[@]}" --output markdown > "$md" + code=$? + set -e + + # Exit 1 means deltas were found, which is the thing being reported + # rather than a failure. Exit 2 means the command could not run — + # a usage error, or a cluster it could not reach — and a gate that + # treated that as "no deltas" would be silently useless. + case "$code" in + 0) has_deltas=false ;; + 1) has_deltas=true ;; + *) + echo "::error::convctl diff failed to run (exit $code)" >&2 + cat "$md" >&2 || true + exit "$code" + ;; + esac + + { + echo "exit-code=$code" + echo "has-deltas=$has_deltas" + echo "markdown-path=$md" + } >> "$GITHUB_OUTPUT" + + cat "$md" >> "$GITHUB_STEP_SUMMARY" + + if [ "$has_deltas" = "true" ] && [ "$FAIL_ON_DELTA" = "true" ]; then + exit 1 + fi + + - name: Upsert the pull-request comment + if: always() && inputs.comment == 'true' && github.event_name == 'pull_request' + shell: bash + env: + GH_TOKEN: ${{ inputs.token }} + TAG: ${{ inputs.comment-tag != '' && inputs.comment-tag || inputs.config }} + MD: ${{ steps.run.outputs.markdown-path }} + PR: ${{ github.event.pull_request.number }} + REPO: ${{ github.repository }} + FORK: ${{ github.event.pull_request.head.repo.fork }} + run: | + set -euo pipefail + + # A pull_request event from a fork has a read-only token. That is a + # notice, not a failure: the delta is already in the job summary, and + # failing here would make every fork contribution red for a reason + # the contributor cannot fix. + if [ "$FORK" = "true" ]; then + echo "::notice::pull request is from a fork, so the token cannot comment; the delta is in the job summary" + exit 0 + fi + if [ ! -f "$MD" ]; then + echo "::notice::no diff output to comment" + exit 0 + fi + + marker="" + body="$RUNNER_TEMP/convctl-diff-comment.md" + { + echo "$marker" + cat "$MD" + } > "$body" + + # Find this tag's comment and edit it, so repeated runs update one + # comment rather than appending a new one each push. + existing="$(gh api "repos/$REPO/issues/$PR/comments" --paginate \ + --jq "[.[] | select(.body | contains(\"$marker\"))] | .[0].id // empty")" + + if [ -n "$existing" ]; then + gh api --method PATCH "repos/$REPO/issues/comments/$existing" \ + -F body=@"$body" --silent + else + gh api --method POST "repos/$REPO/issues/$PR/comments" \ + -F body=@"$body" --silent + fi From 738c0a3e90df49e44e1d2aa225a71ed0ed8a34d3 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:26:21 +0300 Subject: [PATCH 10/24] feat(actions): convctl-fleet, and a reference workflow that runs MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The documented fleet pattern shipped with an install step that read: echo "install convctl and place it on PATH" >&2 exit 1 A reference workflow that cannot run is worse than none: it implies the pattern is supported and wastes the reader's time before they find out otherwise. convctl-fleet wraps the --contexts / --kubeconfig-dir run the CLI already supports and aggregates it: one JUnit report with a suite per cluster, a summary table, and the cluster count and failure count as outputs. A cluster that could not be reached is recorded as a failed suite rather than skipped — the CLI already behaves that way, and the Action surfaces it, because a fleet check that quietly covered four of five clusters and reported green is worse than one that did not run. convctl-fleet.gha.yml is rewritten on the real Actions and now runs as written; the only things to change are the context matrix, the paths and the kubeconfig secret. It shows both shapes deliberately: a per-cluster matrix with convctl-diff and convctl-test for a clearer failure surface, and the single aggregated job for a simpler one. fail-fast: false is kept and explained on the matrix form, since a red cluster hiding the others is the mistake that makes a fleet gate useless. It also leads with convctl lint, which needs no cluster at all: the fast offline check gates the slow credentialed ones rather than running beside them. fleet-ci.md now describes the Action-based flow first and keeps the shell loop, relabelled for the non-GitHub CI systems it exists for. Closes #137 Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/convctl-fleet/README.md | 35 +++++ .github/actions/convctl-fleet/action.yml | 167 +++++++++++++++++++++++ docs/gitops/convctl-fleet.gha.yml | 105 ++++++++------ docs/gitops/fleet-ci.md | 43 ++++-- 4 files changed, 298 insertions(+), 52 deletions(-) create mode 100644 .github/actions/convctl-fleet/README.md create mode 100644 .github/actions/convctl-fleet/action.yml diff --git a/.github/actions/convctl-fleet/README.md b/.github/actions/convctl-fleet/README.md new file mode 100644 index 0000000..196da23 --- /dev/null +++ b/.github/actions/convctl-fleet/README.md @@ -0,0 +1,35 @@ +# `convctl-fleet` + +Runs `convctl test --live` against every cluster in a fleet and aggregates +the result into one JUnit report, with one `` per cluster. + +```yaml +- uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-fleet@v1 + with: + config: apis/widgets/conversion.yaml + xrd: apis/widgets/xrd.yaml + contexts: prod-us,prod-eu + kubeconfig: ${{ secrets.FLEET_KUBECONFIG }} +``` + +Either `contexts` (comma-separated, from one kubeconfig) or `kubeconfig-dir` +(one file per cluster). + +## An unreachable cluster is a failed suite + +Not a skip. `convctl` already behaves this way; the Action surfaces it, in +the report and in the `failed-clusters` output. A fleet check that quietly +covered four of five clusters and reported green is worse than one that did +not run at all. + +## Alternative: a matrix + +`convctl-fleet` runs the clusters in one job. A per-cluster matrix using +[`convctl-test`](../convctl-test) gives you one job per cluster — slower to +set up, but a clearer failure surface and parallel execution. Use +`fail-fast: false` either way. Both shapes are in +[`convctl-fleet.gha.yml`](../../../docs/gitops/convctl-fleet.gha.yml). + +## Outputs + +`exit-code`, `report-path`, `clusters`, `failed-clusters`. diff --git a/.github/actions/convctl-fleet/action.yml b/.github/actions/convctl-fleet/action.yml new file mode 100644 index 0000000..31e6b3a --- /dev/null +++ b/.github/actions/convctl-fleet/action.yml @@ -0,0 +1,167 @@ +name: convctl fleet +description: >- + Run convctl test --live against several clusters and aggregate the result + into one JUnit report with one suite per cluster. + +inputs: + config: + description: Path to the XRDConversionConfig or CRDConversionConfig. + required: true + xrd: + description: Path to the XRD (for an XRDConversionConfig). + required: false + default: "" + crd: + description: Path to the CRD (for a CRDConversionConfig). + required: false + default: "" + contexts: + description: >- + Comma-separated kubeconfig contexts, one per cluster. + required: false + default: "" + kubeconfig-dir: + description: >- + Directory of kubeconfig files, one per cluster. Mutually exclusive with + contexts. + required: false + default: "" + kubeconfig: + description: >- + Kubeconfig contents holding every context. Written with mode 600 and + never echoed. + required: false + default: "" + fail-on: + description: "Failure threshold: none, warn, or loss (default)." + required: false + default: loss + max-samples: + description: Cap how many objects are tested per cluster. + required: false + default: "" + sample-strategy: + description: "With max-samples: first, random, or newest." + required: false + default: "" + validate-output: + description: Validate converted objects against the destination schema. + required: false + default: "false" + version: + description: convctl release tag to install. + required: false + default: latest + upload-artifact: + description: Upload the aggregated JUnit report. + required: false + default: "true" + artifact-name: + description: Name for the uploaded artifact. + required: false + default: convctl-fleet-report + +outputs: + exit-code: + description: convctl's own exit code, preserved. + value: ${{ steps.run.outputs.exit-code }} + report-path: + description: Path of the aggregated JUnit report. + value: ${{ steps.run.outputs.report-path }} + clusters: + description: Number of clusters in the report. + value: ${{ steps.run.outputs.clusters }} + failed-clusters: + description: Number of clusters that failed or could not be reached. + value: ${{ steps.run.outputs.failed-clusters }} + +runs: + using: composite + steps: + - name: Set up convctl + uses: ./.github/actions/setup-convctl + with: + version: ${{ inputs.version }} + + - name: Write kubeconfig + if: inputs.kubeconfig != '' + shell: bash + env: + KUBECONFIG_CONTENTS: ${{ inputs.kubeconfig }} + run: | + set -euo pipefail + umask 077 + printf '%s' "$KUBECONFIG_CONTENTS" > "$RUNNER_TEMP/kubeconfig" + echo "KUBECONFIG=$RUNNER_TEMP/kubeconfig" >> "$GITHUB_ENV" + + - name: Run convctl test across the fleet + id: run + shell: bash + env: + CONFIG: ${{ inputs.config }} + XRD: ${{ inputs.xrd }} + CRD: ${{ inputs.crd }} + CONTEXTS: ${{ inputs.contexts }} + KUBECONFIG_DIR: ${{ inputs.kubeconfig-dir }} + FAIL_ON: ${{ inputs.fail-on }} + MAX_SAMPLES: ${{ inputs.max-samples }} + SAMPLE_STRATEGY: ${{ inputs.sample-strategy }} + VALIDATE_OUTPUT: ${{ inputs.validate-output }} + run: | + set -euo pipefail + + if [ -z "$CONTEXTS" ] && [ -z "$KUBECONFIG_DIR" ]; then + echo "::error::one of contexts or kubeconfig-dir is required" >&2 + exit 2 + fi + + report="$RUNNER_TEMP/convctl-fleet.junit.xml" + args=(--config "$CONFIG" --live --fail-on "$FAIL_ON") + if [ -n "$XRD" ]; then args+=(--xrd "$XRD"); fi + if [ -n "$CRD" ]; then args+=(--crd "$CRD"); fi + if [ -n "$CONTEXTS" ]; then args+=(--contexts "$CONTEXTS"); fi + if [ -n "$KUBECONFIG_DIR" ]; then args+=(--kubeconfig-dir "$KUBECONFIG_DIR"); fi + if [ -n "$MAX_SAMPLES" ]; then args+=(--max-samples "$MAX_SAMPLES"); fi + if [ -n "$SAMPLE_STRATEGY" ]; then args+=(--sample-strategy "$SAMPLE_STRATEGY"); fi + if [ "$VALIDATE_OUTPUT" = "true" ]; then args+=(--validate-output); fi + + # convctl already records an unreachable cluster as a failed suite + # rather than skipping it, which is the behaviour that matters here: + # a fleet check that quietly covered four of five clusters and + # reported green is worse than one that did not run. + set +e + convctl test "${args[@]}" --output junit --output-file "$report" --quiet + code=$? + set -e + + clusters=0 + failed=0 + if [ -f "$report" ]; then + clusters=$(grep -c ']*' "$report" \ + | grep -c -E 'failures="[1-9]|errors="[1-9]' || echo 0) + { + echo "### convctl fleet" + echo "" + echo "| Clusters | Failed |" + echo "|---|---|" + echo "| $clusters | $failed |" + } >> "$GITHUB_STEP_SUMMARY" + fi + + { + echo "exit-code=$code" + echo "report-path=$report" + echo "clusters=$clusters" + echo "failed-clusters=$failed" + } >> "$GITHUB_OUTPUT" + + exit "$code" + + - name: Upload aggregated report + if: always() && inputs.upload-artifact == 'true' + uses: actions/upload-artifact@v7 + with: + name: ${{ inputs.artifact-name }} + path: ${{ steps.run.outputs.report-path }} + if-no-files-found: warn diff --git a/docs/gitops/convctl-fleet.gha.yml b/docs/gitops/convctl-fleet.gha.yml index 58edf64..39028a7 100644 --- a/docs/gitops/convctl-fleet.gha.yml +++ b/docs/gitops/convctl-fleet.gha.yml @@ -1,9 +1,8 @@ # Reference GitHub Actions workflow — copy into a platform repo. # This file is not wired into this operator's own CI. # -# Each matrix.context is one cluster. fail-fast is false so one red -# cluster does not hide the rest. Upload the JUnit / diff JSON as -# artifacts for a test reporter or PR review. +# Everything here runs as written. The only things to change are the +# `context` matrix, the paths, and the kubeconfig secret. name: convctl fleet on: @@ -15,55 +14,75 @@ on: permissions: contents: read + pull-requests: write # for the sticky diff comment jobs: - fleet: + # The fast, offline check. No cluster, no credentials — so it runs on + # every commit and tells you about an unpaired or broken config before any + # of the slow jobs start. + lint: runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 + - run: convctl lint ./apis/ --output github + + # One leg per cluster. fail-fast: false is required — a red cluster must + # not hide the others. + per-cluster: + runs-on: ubuntu-latest + needs: lint strategy: fail-fast: false matrix: context: [prod-us, prod-eu] steps: - - uses: actions/checkout@v4 - - - name: Install convctl - run: | - # Pin to a released convctl, or build from this repo: - # go install github.com/terasky-oss/declarative-conversion-operator/cmd/convctl@latest - echo "install convctl and place it on PATH" >&2 - exit 1 - - - name: Write kubeconfig - env: - KUBECONFIG_B64: ${{ secrets.FLEET_KUBECONFIG }} - run: | - mkdir -p "${HOME}/.kube" - printf '%s' "${KUBECONFIG_B64}" | base64 -d > "${HOME}/.kube/config" - chmod 600 "${HOME}/.kube/config" - - - name: convctl diff --live - # Exit 1 is a coverage/claim delta (review artifact). Exit 2 is a - # usage/cluster error and must fail the job. - run: | - mkdir -p fleet-out - set +e - convctl diff --config path/to/xrdconversionconfig.yaml --live \ - --context "${{ matrix.context }}" -o json \ - > "fleet-out/${{ matrix.context }}.diff.json" - rc=$? - if [ "${rc}" -eq 2 ]; then exit 2; fi - exit 0 + - uses: actions/checkout@v7 + with: + persist-credentials: false - - name: convctl test --live - if: ${{ always() }} - run: | - convctl test --config path/to/xrdconversionconfig.yaml \ - --xrd path/to/xrd.yaml --live \ - --context "${{ matrix.context }}" \ - --output junit --output-file "fleet-out/${{ matrix.context }}.junit.xml" + # The delta, as a sticky comment per cluster. Exit 1 (deltas found) is + # a review artifact and does not fail the job; exit 2 (cannot reach the + # cluster) does. + - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-diff@v1 + with: + config: apis/widgets/conversion.yaml + xrd: apis/widgets/xrd.yaml + live: "true" + context: ${{ matrix.context }} + kubeconfig: ${{ secrets.FLEET_KUBECONFIG }} + comment-tag: diff-${{ matrix.context }} - - uses: actions/upload-artifact@v4 + - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-test@v1 if: always() with: - name: convctl-${{ matrix.context }} - path: fleet-out/ + config: apis/widgets/conversion.yaml + xrd: apis/widgets/xrd.yaml + live: "true" + context: ${{ matrix.context }} + kubeconfig: ${{ secrets.FLEET_KUBECONFIG }} + validate-output: "true" + artifact-name: convctl-${{ matrix.context }} + # On a large cluster, bound the run and let the report say it + # sampled rather than quietly OOMing. + max-samples: "500" + sample-strategy: random + + # Or, instead of the matrix: one job that tests every context and emits a + # single aggregated report with one suite per cluster. + aggregated: + runs-on: ubuntu-latest + needs: lint + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-fleet@v1 + with: + config: apis/widgets/conversion.yaml + xrd: apis/widgets/xrd.yaml + contexts: prod-us,prod-eu + kubeconfig: ${{ secrets.FLEET_KUBECONFIG }} + validate-output: "true" diff --git a/docs/gitops/fleet-ci.md b/docs/gitops/fleet-ci.md index a9d7b56..dde90a4 100644 --- a/docs/gitops/fleet-ci.md +++ b/docs/gitops/fleet-ci.md @@ -47,12 +47,12 @@ single-cluster report. A connection error on one cluster is recorded as a failed suite; the others still run. `convctl diff` stays one cluster per invocation (`--context`). The -[shell loop](#shell-loop) still wraps both commands when you want +[shell loop](#shell-loop-non-github-ci) still wraps both commands when you want `diff --live` in the same gate. -## Shell loop +## Shell loop (non-GitHub CI) -[`convctl-fleet.sh`](convctl-fleet.sh) is a copy-pasteable wrapper. It +For Tekton, GitLab, Jenkins or anything else, [`convctl-fleet.sh`](convctl-fleet.sh) is a copy-pasteable wrapper. It walks `CONTEXTS` (space-separated kubeconfig context names), writes one JUnit file per cluster, and exits non-zero if any cluster failed. @@ -68,13 +68,38 @@ export CONVCTL_CONFIG=examples/field-rename/xrdconversionconfig.yaml one kubeconfig file per cluster instead of contexts, set `KUBECONFIGS` to a list of paths (the script uses each file's `current-context`). -## GitHub Actions matrix +## GitHub Actions -[`convctl-fleet.gha.yml`](convctl-fleet.gha.yml) is a reference workflow, -not a job this repository runs. Copy it into your platform repo and -replace the `context` matrix with your fleet. Each matrix leg is one -cluster; `actions/upload-artifact` collects the JUnit files so a -test-reporter can show a per-cluster breakdown. +Three first-party Actions cover the fleet pattern, so the workflow is +configuration rather than a copy-pasted script: + +```yaml +- uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-fleet@v1 + with: + config: apis/widgets/conversion.yaml + xrd: apis/widgets/xrd.yaml + contexts: prod-us,prod-eu + kubeconfig: ${{ secrets.FLEET_KUBECONFIG }} +``` + +One aggregated JUnit report with a `` per cluster, a summary +table, and a cluster that could not be reached recorded as a **failed suite** +rather than silently skipped. + +| Action | Does | +|---|---| +| [`setup-convctl`](https://github.com/TeraSky-OSS/declarative-conversion-operator/tree/main/.github/actions/setup-convctl) | installs a cosign-verified `convctl` | +| [`convctl-test`](https://github.com/TeraSky-OSS/declarative-conversion-operator/tree/main/.github/actions/convctl-test) | one cluster or fixtures: JUnit artifact, job summary, diff annotations | +| [`convctl-diff`](https://github.com/TeraSky-OSS/declarative-conversion-operator/tree/main/.github/actions/convctl-diff) | the coverage delta as a sticky pull-request comment | +| [`convctl-fleet`](https://github.com/TeraSky-OSS/declarative-conversion-operator/tree/main/.github/actions/convctl-fleet) | every cluster, one aggregated report | + +[`convctl-fleet.gha.yml`](convctl-fleet.gha.yml) is the full reference +workflow, built on those Actions. It runs as written — the only things to +change are the `context` matrix, the paths, and the kubeconfig secret. Copy +it into your platform repo. + +`fail-fast: false` on a per-cluster matrix is still required: a red cluster +must not hide the others. ```yaml strategy: From 2e6a71da363917213d652fbbb3c93a99f714b125 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:28:55 +0300 Subject: [PATCH 11/24] test(actions): a workflow that exercises every Action, including the tamper path MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit An Action nobody tests is a broken Action nobody notices until a consumer's pipeline goes red — and by then the breakage is in a published tag that other repositories are pinned to. These Actions are this project's public CI surface, so they get the same discipline as the Go code. Every Action is called by local path, so the version under test is the pull request's version rather than whatever is published. The most important job is the tamper test. A verification step that cannot be shown to fail is not a verification step, so the workflow installs convctl, corrupts the cached archive in place — same length, different bytes, because a length check would not catch what a hash does — and then asserts the re-install FAILS, via continue-on-error plus an explicit check on the step's outcome. If verification ever silently passes, that job turns red. Assertions are on output, not just exit codes, because every one of these Actions can exit correctly while rendering nothing. The annotation test checks that the emitted workflow command carries the config file, a line number, AND the rule index that produced the finding — the last of those is what makes an annotation actionable rather than merely located. The diff test checks the rendered markdown is non-empty and has its heading, and that identical configs say "No differences" rather than rendering nothing at all. Choosing the failing fixture took a correction worth recording: the obvious candidates (the mistakes/ configs) fail at analysis with a plain error and no annotations, because convctl refuses to test a config that does not compile. The fixture used instead converts to values the destination schema rejects — an out-of-enum value and a pattern violation — which compiles cleanly and fails only at --validate-output, so the test takes the path a real regression takes. Every assertion in the workflow was run locally against real output first rather than written from expectation. setup-convctl runs on all three runner OSes, since the CLI ships darwin and windows archives and the path and extraction logic differ on each. The cache path is exercised explicitly — installed twice, with an assertion that the second was a hit — because the cached branch is the one that runs in practice and the one that silently rots. actionlint runs over the whole repository rather than only the new files: a workflow that broke two years ago is still broken. It cannot check a composite action's action.yml, which it reads as a malformed workflow, so test/actions covers that gap statically and runs in the same job. Closes #138 Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/actions-test.yml | 326 +++++++++++++++++++++++++++++ 1 file changed, 326 insertions(+) create mode 100644 .github/workflows/actions-test.yml diff --git a/.github/workflows/actions-test.yml b/.github/workflows/actions-test.yml new file mode 100644 index 0000000..b43951f --- /dev/null +++ b/.github/workflows/actions-test.yml @@ -0,0 +1,326 @@ +# The Actions this repository publishes are its public CI surface. An Action +# nobody tests is a broken Action nobody notices until a consumer's pipeline +# goes red — and by then the breakage is in a published tag that other +# repositories are pinned to. +name: Actions + +on: + pull_request: + paths: + - ".github/actions/**" + - ".github/workflows/actions-test.yml" + - "test/actions/**" + push: + branches: [main] + paths: + - ".github/actions/**" + - ".github/workflows/actions-test.yml" + - "test/actions/**" + +permissions: + contents: read + +concurrency: + group: actions-test-${{ github.ref }} + cancel-in-progress: true + +jobs: + # Static checks first: they are seconds, and they catch the whole class of + # "the action.yml is malformed" without spending a runner minute per OS. + static: + name: actionlint and static checks + runs-on: ubuntu-latest + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + + - uses: actions/setup-go@v6 + with: + go-version-file: go.mod + cache: true + + # Over the whole repository, not just the new files: a workflow that + # broke two years ago is still broken. + - name: actionlint + uses: raven-actions/actionlint@v2 + with: + fail-on-error: true + + # actionlint treats a composite action's action.yml as a malformed + # workflow, so the Actions themselves need their own checks. + - name: Static checks on the composite Actions + run: go test ./test/actions/ -v + + setup: + name: setup-convctl (${{ matrix.os }}) + runs-on: ${{ matrix.os }} + needs: static + strategy: + fail-fast: false + matrix: + os: [ubuntu-latest, macos-latest, windows-latest] + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + + # By local path, so the version under test is this pull request's + # version rather than whatever is published. + - name: Install latest + id: latest + uses: ./.github/actions/setup-convctl + + - name: It resolved a concrete tag and works + shell: bash + run: | + set -euo pipefail + test -n "${{ steps.latest.outputs.version }}" + case "${{ steps.latest.outputs.version }}" in + v*) ;; + *) echo "resolved version is not a tag: ${{ steps.latest.outputs.version }}" >&2; exit 1 ;; + esac + convctl version + convctl lint --help >/dev/null + + # The cached branch is the one that runs in practice and the one that + # silently rots, so it is exercised explicitly — including that it + # still verified. + - name: Install again (cache hit) + id: cached + uses: ./.github/actions/setup-convctl + with: + version: ${{ steps.latest.outputs.version }} + + - name: The second install came from the cache and still verified + shell: bash + run: | + set -euo pipefail + if [ "${{ steps.cached.outputs.cache-hit }}" != "true" ]; then + echo "second install was not a cache hit; the cached path is untested" >&2 + exit 1 + fi + convctl version + + # The most important job here. A verification step that cannot be shown to + # fail is not a verification step. + tamper: + name: setup-convctl rejects a tampered archive + runs-on: ubuntu-latest + needs: static + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + + - name: Install once to populate the cache directory + id: real + uses: ./.github/actions/setup-convctl + + - name: Corrupt the cached archive + shell: bash + env: + VERSION: ${{ steps.real.outputs.version }} + run: | + set -euo pipefail + bare="${VERSION#v}" + dir="$RUNNER_TEMP/convctl-$bare" + archive="$(find "$dir" -name '*.tar.gz' | head -1)" + test -n "$archive" + # Same size, different bytes: a length check would not catch this, + # which is the point of checking the hash. + printf 'tampered' | dd of="$archive" bs=1 seek=0 conv=notrunc status=none + echo "corrupted $archive" + + - name: Verification must reject it + id: tampered + continue-on-error: true + uses: ./.github/actions/setup-convctl + with: + version: ${{ steps.real.outputs.version }} + + - name: Assert the step failed + shell: bash + run: | + set -euo pipefail + outcome="${{ steps.tampered.outcome }}" + if [ "$outcome" != "failure" ]; then + echo "::error::setup-convctl accepted a tampered archive (outcome: $outcome)" >&2 + echo "A verification step that cannot be shown to fail is not a verification step." >&2 + exit 1 + fi + echo "tampered archive was rejected, as it must be" + + test-action: + name: convctl-test against fixtures + runs-on: ubuntu-latest + needs: static + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + + # A clean fixture passes and still produces a report. + - name: Clean fixture + id: clean + uses: ./.github/actions/convctl-test + with: + config: internal/cli/testdata/full/config.yaml + xrd: internal/cli/testdata/full/xrd.yaml + samples: internal/cli/testdata/full/samples + artifact-name: actions-test-clean + + - name: It passed and reported real numbers + shell: bash + run: | + set -euo pipefail + if [ "${{ steps.clean.outputs.exit-code }}" != "0" ]; then + echo "a clean fixture did not pass" >&2; exit 1 + fi + if [ "${{ steps.clean.outputs.pass }}" -lt 1 ]; then + echo "the pass count is ${{ steps.clean.outputs.pass }}; the Action exited correctly while reporting nothing" >&2 + exit 1 + fi + test -f "${{ steps.clean.outputs.report-path }}" + + # A deliberately broken fixture must fail AND say where. Asserting on + # the annotation payload rather than only the exit code: every one of + # these Actions can exit correctly while rendering nothing. + # + # The fixture converts to values the destination schema rejects — an + # out-of-enum value and a pattern violation — which compiles cleanly + # and fails only at --validate-output, so it takes the path a real + # regression takes rather than failing at analysis. + - name: Capture the annotations for a failing fixture + shell: bash + run: | + set -euo pipefail + set +e + go run ./cmd/convctl test \ + --xrd internal/cli/testdata/validate-output/xrd.yaml \ + --config internal/cli/testdata/validate-output/config.yaml \ + --samples internal/cli/testdata/validate-output/samples \ + --validate-output --output github --quiet > annotations.txt 2>&1 + code=$? + set -e + cat annotations.txt + if [ "$code" -eq 0 ]; then + echo "::error::the failing fixture passed; it is no longer a failing fixture" >&2 + exit 1 + fi + + - name: The annotation carries the config file, a line, and the rule + shell: bash + run: | + set -euo pipefail + if ! grep -q '^::error ' annotations.txt; then + echo "::error::no error annotations were emitted for a failing run" >&2 + cat annotations.txt >&2 + exit 1 + fi + if ! grep -qE '^::error file=[^,]*validate-output/config\.yaml,line=[0-9]+' annotations.txt; then + echo "::error::annotations do not carry the config file and a line number" >&2 + cat annotations.txt >&2 + exit 1 + fi + # Naming the rule is what makes an annotation actionable rather + # than merely located. + if ! grep -qE 'rule\[[0-9]+\]' annotations.txt; then + echo "::error::annotations do not name the rule that produced the finding" >&2 + cat annotations.txt >&2 + exit 1 + fi + echo "annotations land on the config, with line numbers and the rule" + + - name: The Action itself fails on that fixture + id: failing + continue-on-error: true + uses: ./.github/actions/convctl-test + with: + config: internal/cli/testdata/validate-output/config.yaml + xrd: internal/cli/testdata/validate-output/xrd.yaml + samples: internal/cli/testdata/validate-output/samples + validate-output: "true" + artifact-name: actions-test-failing + + - name: Assert it failed, and reported why + shell: bash + run: | + set -euo pipefail + if [ "${{ steps.failing.outcome }}" != "failure" ]; then + echo "::error::a fixture that produces schema violations did not fail the Action" >&2 + exit 1 + fi + if [ "${{ steps.failing.outputs.errors }}" = "0" ]; then + echo "::error::the Action failed but reported no errors, so its summary would read as a mystery" >&2 + exit 1 + fi + + - name: annotate:false and upload-artifact:false both work + id: quiet + uses: ./.github/actions/convctl-test + with: + config: internal/cli/testdata/full/config.yaml + xrd: internal/cli/testdata/full/xrd.yaml + samples: internal/cli/testdata/full/samples + annotate: "false" + upload-artifact: "false" + + - name: It still ran + shell: bash + run: | + set -euo pipefail + test "${{ steps.quiet.outputs.exit-code }}" = "0" + + diff-action: + name: convctl-diff renders a delta + runs-on: ubuntu-latest + needs: static + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + + # comment:false — this job proves the rendering and the exit-code + # semantics; commenting needs a pull request and a writable token. + - name: Two configs that differ + id: diff + uses: ./.github/actions/convctl-diff + with: + config: | + examples/crossplane-xr-multiversion/03-promote-v2/xrdconversionconfig.yaml + examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml + xrd: examples/crossplane-xr-multiversion/03-promote-v2/xrd.yaml + comment: "false" + + - name: A delta was found and did not fail the job + shell: bash + run: | + set -euo pipefail + if [ "${{ steps.diff.outputs.has-deltas }}" != "true" ]; then + echo "two different configs produced no delta" >&2; exit 1 + fi + if [ "${{ steps.diff.outputs.exit-code }}" != "1" ]; then + echo "expected exit 1 for deltas, got ${{ steps.diff.outputs.exit-code }}" >&2; exit 1 + fi + # The rendering is the product; an empty one is a silent failure. + test -s "${{ steps.diff.outputs.markdown-path }}" + grep -q '^### Conversion config diff' "${{ steps.diff.outputs.markdown-path }}" + echo "delta rendered, job stayed green" + + - name: Identical configs report no deltas + id: same + uses: ./.github/actions/convctl-diff + with: + config: | + examples/field-rename/xrdconversionconfig.yaml + examples/field-rename/xrdconversionconfig.yaml + xrd: examples/field-rename/xrd.yaml + comment: "false" + + - name: And say so rather than rendering nothing + shell: bash + run: | + set -euo pipefail + test "${{ steps.same.outputs.has-deltas }}" = "false" + grep -q 'No differences' "${{ steps.same.outputs.markdown-path }}" From 7d50103392259bfad2bf7bf85d3c2e683a8a9922 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:30:00 +0300 Subject: [PATCH 12/24] docs: record phase 13 as shipped, with its deviations and limits The roadmap moves phase 13 into the shipped table, noting that it landed after phase 12 rather than before it despite the epic's own instruction to build the road before the traffic. next-phases.md gets a shipped block in the shape phases 11 and 12 use, naming the five places the implementation diverged: --package reads a local .xpkg only; --live streams into the sampler but still accumulates without a cap; Homebrew ships as a macOS-only cask because GoReleaser deprecated formulae for binaries; the convctl image stays distroless and therefore shell-less; and convctl-test relays the tool's own annotations rather than re-deriving locations in YAML. It also records what the test workflow taught, because it shaped the tests rather than being incidental: the obvious "broken config" fixtures fail at analysis with a plain error and no annotations, since convctl refuses to test a config that does not compile. Asserting on annotation payloads needs a fixture that compiles and fails later. krew is struck from 13.5 as dropped rather than deferred, with the reason: convctl is not a kubectl plugin, and the krew-index review cycle is weeks of process for a distribution channel nobody asked for. limitations.md gains the two that come with publishing Actions: the pinned certificate identity is what makes verification meaningful and also what a fork would have to change, and a fork's pull_request token cannot comment, so the diff Action degrades to a notice rather than a red check a contributor cannot fix. Co-Authored-By: Claude Opus 5 (1M context) --- docs/limitations.md | 2 ++ docs/proposals/next-phases.md | 39 ++++++++++++++++++++++++++++++++--- docs/roadmap.md | 11 +++++----- 3 files changed, 44 insertions(+), 8 deletions(-) diff --git a/docs/limitations.md b/docs/limitations.md index 4dce000..085b75f 100755 --- a/docs/limitations.md +++ b/docs/limitations.md @@ -40,6 +40,8 @@ This page is deliberately blunt about what the operator does *not* do today, so - **`--package` reads a local `.xpkg` only.** An xpkg is an OCI image saved as a tarball, so the local form needs nothing but the standard library — and it is the tightest loop, before anything is published. A registry reference (`ghcr.io/org/platform:v1.4.0`) or a cluster reference (`configuration/`, `configurationrevision/`) is recognised and rejected with the `crossplane xpkg pull` command that produces a local file, rather than silently unsupported. Supporting them means a registry client the offline path does not need and should not carry. - **An uncapped `--live` run holds every object it tests.** Listing streams into the sampler, so `--max-samples` genuinely bounds memory — the population is counted without being held. Without a cap the objects are still accumulated before testing, which on a cluster with tens of thousands of large composites is the case `--max-samples` exists for. Testing each page as it arrives would bound it either way and is not implemented; the report is assembled from the full sample set. - **`convctl plan` reads manifests, not a cluster.** It answers "what is the next safe step?" from the XRD/CRD and the conversion config, which is what makes it usable in a PR. Three steps of the sequence have no answer in a manifest — retargeting Compositions, migrating stored objects, pruning `storedVersions` — and are reported as `UNKNOWN` with the command that answers them, never as done, ready, or blocked. A plan therefore stops at *"nothing outstanding that files can decide"*; finishing the migration still requires running those verify commands against the cluster. +- **The published Actions are pinned to this repository.** `setup-convctl` verifies against a certificate identity pinned to *this* repository's release workflow, which is what makes the verification meaningful — and means a fork publishing its own releases has to change that identity to use the Action against them. +- **`convctl-diff` cannot comment on a pull request from a fork.** A `pull_request` event from a fork gets a read-only token, so the Action emits a notice and leaves the delta in the job summary rather than failing. Use `pull_request_target` only if you have read [its hazards](https://securitylab.github.com/resources/github-actions-preventing-pwn-requests/); this project does not recommend it. - **`convctl versions` is XRD-only.** The live inventory — object counts, `storedVersions`, and the field managers still writing each version — is implemented for XRD targets; `--crd` is rejected with a message saying so rather than silently answering a narrower question. `plan`, `compat`, `validate`, `analyze` and `test` all cover both. - **`convctl compat` compares what git can show it.** It resolves the XRD and config at two revisions with `git show`, so it needs both revisions present locally — `fetch-depth: 2` at minimum in CI, and a shallow clone that does not contain the base ref is a usage error rather than a silent pass. It classifies the delta between two *configs and schemas*; it does not read the cluster, so "a served version removed while objects are still stored at it" is judged from `spec.versions`, and the live half of that question belongs to `versions`. - **`--fuzz` finds counterexamples, not proofs.** Generated objects are schema-valid and take their lexical shape from the rules that consume each field, which is what makes the failures real conversion bugs rather than noise about random strings not parsing as quantities. A clean fuzz run at a given seed and count is evidence, not a guarantee: it says nothing about the inputs that were not generated. Raise `--fuzz` and vary `--seed` in nightly CI rather than treating one seeded run as coverage. diff --git a/docs/proposals/next-phases.md b/docs/proposals/next-phases.md index 058c4fa..b95da6c 100644 --- a/docs/proposals/next-phases.md +++ b/docs/proposals/next-phases.md @@ -743,6 +743,39 @@ still stored at, or clients still writing, that version. ## Phase 13 — CI/CD: official GitHub Actions and CLI ergonomics +> **Shipped.** Every deliverable below landed except krew, which was dropped +> as a target rather than deferred (see 13.5). Five deviations are worth +> recording: +> +> - **`--package` reads a local `.xpkg` only.** An xpkg is an OCI image saved +> as a tarball, so the local form — the tightest loop, before anything is +> published — needs nothing but the standard library. Registry and cluster +> references are recognised and rejected with the `crossplane xpkg pull` +> command that produces a local file, rather than silently unsupported. +> Supporting them means a registry client the offline path does not need. +> - **`--live` streams into the sampler but still accumulates without a cap.** +> `--max-samples` genuinely bounds memory, and the population is counted +> without being held. Testing each page as it arrives, which would bound it +> with no cap at all, is not implemented — see [Limitations](../limitations.md). +> - **Homebrew ships as a cask, not a formula.** GoReleaser deprecated +> `brews:` in favour of `homebrew_casks:`, which is macOS-only, so Linuxbrew +> users install from the deb, the rpm or the archive. The deprecation forced +> the trade; the docs state it rather than implying coverage that does not +> exist. +> - **The `convctl` image stays distroless.** It has no shell, so commands +> cannot be chained inside it. Verified rather than assumed: the base does +> carry CA certificates, so `--live` reaches an HTTPS apiserver. +> - **`convctl-test` emits annotations by relaying `--output github`** rather +> than mapping findings to lines inside the Action. The mapping lives in the +> tool, where it is tested, instead of being implemented a second time in +> YAML. +> +> One thing the test workflow taught, recorded because it shaped the tests: +> the obvious "broken config" fixtures fail at *analysis*, with a plain error +> and no annotations, because `convctl test` refuses to run against a config +> that does not compile. Asserting on annotations needs a fixture that +> compiles and fails later — the `--validate-output` one does. + The theme: `convctl` is designed for CI (exit-code matrix, JUnit output, `--fail-on`, parallel `--live`) but there is no supported way to *get* it into a pipeline. Fixing F6 and F7 is the bulk of this phase. @@ -815,9 +848,9 @@ goreleaser already builds the archives; add the publishing targets that make the tool installable the way people expect: - Homebrew tap (`brews`), Scoop, and `nfpms` for deb/rpm. -- A **krew** plugin manifest — `kubectl conversion test|diff|plan` is the - natural home for a kubectl-adjacent tool, and krew is how the Kubernetes - ecosystem discovers one. +- ~~A **krew** plugin manifest.~~ **Dropped.** `convctl` is not a kubectl + plugin, and the krew-index review cycle is weeks of process for a + distribution channel nobody asked for. - `convctl version --output json` with commit, build date, and Go version. - Reference pipeline templates for GitLab CI, Tekton, and Argo Workflows, mirroring the GitHub one. diff --git a/docs/roadmap.md b/docs/roadmap.md index c79b8b7..4da8401 100644 --- a/docs/roadmap.md +++ b/docs/roadmap.md @@ -5,7 +5,7 @@ timeline. Phases already shipped stay listed so the arc is visible; later phases are invitations to [open an issue or PR](https://github.com/terasky-oss/declarative-conversion-operator/issues) if one of them matters to you sooner. -## Shipped (phases 0–12, 14) +## Shipped (phases 0–14) | Phase | Epic | Intent | |---|---|---| @@ -22,20 +22,21 @@ if one of them matters to you sooner. | **10 — Multi-cluster / GitOps** | [#81](https://github.com/terasky-oss/declarative-conversion-operator/issues/81) | Documented CI patterns for `convctl test`/`diff` across many kubecontexts, GitOps examples (Flux/Argo). Cross-cluster failover of conversion state remains an explicit non-goal. | | **11 — Crossplane integration depth** | [#113](https://github.com/terasky-oss/declarative-conversion-operator/issues/113) | A `ConversionPropagated` condition and `status.generatedCRDs`, so "Applied" stops being mistaken for "converting". `scope: LegacyCluster` and its claim CRD made first-class: the injected-field set pinned per scope, `spec.claimNames` resolved everywhere a target is resolved, and an e2e leg covering both object classes. Scope-aware `Analyze` rejecting authored names Crossplane overwrites. A mutating admission guard so an XRD shipped in a `Configuration` package stops losing its conversion webhook on every package resync, plus a `PackageManaged` condition and revert counter that make the hazard visible either way. `convctl retarget` and `convctl crossplane status`. Crossplane 2.x stated as the requirement it already was, and checked at startup. | | **12 — XRD/CRD API evolution lifecycle** | [#125](https://github.com/terasky-oss/declarative-conversion-operator/issues/125) | The migration *sequence* made checkable rather than left in prose: `convctl plan` prints the ordered, gated path from where a target is to the version you name, offers exactly one step at a time, and marks the three steps that genuinely need a cluster as `UNKNOWN` with the command that answers them instead of guessing. Golden-corpus testing (`--record` / `--golden`) so a config edit shows up in review as *"this changes the output for these three objects, in these fields"*. `--validate-output` checking every converted object against the destination schema through the apiserver's own validator, and required-field analysis catching at compile time what would otherwise fail at admission. Property-based round-trip fuzzing whose generated values take their lexical shape from the rules, not just the schema. `convctl compat --base/--head` classifying the delta between two git revisions into eight breaking-change classes, designed as a required status check. `convctl versions` answering *"is it safe to unserve this version yet?"* from `managedFields`, live object counts and `storedVersions`. | +| **13 — CI/CD: official GitHub Actions** | [#133](https://github.com/terasky-oss/declarative-conversion-operator/issues/133) | A supported way to get `convctl` into a pipeline, where before the reference workflow's install step was `exit 1`. First-party composite Actions — `setup-convctl`, which verifies the cosign signature by default and still verifies on a cache hit; `convctl-test`, with a JUnit artifact, a job summary and annotations on the diff; `convctl-diff`, as a sticky PR comment that says "no deltas" rather than vanishing; `convctl-fleet`, one aggregated report across clusters — each exercised by a test workflow whose most important job proves verification *fails* on a tampered archive. CI-native output formats (`github`, `sarif`, `markdown`) built on real source locations, so a finding lands on the line of the config that produced it. `convctl lint` over a whole tree, with unpaired and duplicate configs reported rather than skipped. A published `convctl` image. Bounded `--live` sampling that says plainly when it sampled. `--package`, so a Configuration's XRDs can be tested from a local `.xpkg` before publishing. Homebrew, Scoop, deb and rpm, and a `version` command that reports something a bug report can use. | | **14 — Production readiness** | [#145](https://github.com/terasky-oss/declarative-conversion-operator/issues/145) | Informer caches that scale with the number of conversion configs rather than with the cluster: no Secret informer at all in the manager, label-scoped owned workloads, and a webhook-server cache transform that strips what the engine never reads. Timeouts and a body limit on both HTTP servers, with an oversized body answered as a well-formed failing `ConversionReview`. The panic path preserving the request UID, so its message actually arrives. Per-object conversion metrics, so a mixed-direction batch stops being attributed to whichever object was last. Rollout safety — `preStop`, grace period, spread, `maxUnavailable: 0` — proved by a nightly soak that rolls the webhook-server under load and asserts zero failed **and zero wrong** conversions. A curated `.golangci.yml` with every finding fixed, `govulncheck`/CodeQL/Trivy/Scorecard/Dependabot, a chart `values.schema.json`, `helm-unittest` in place of the CI `grep` block, and a [deprecation policy](deprecation-policy.md). | ## Proposed next phases -Phases 0–12 and 14 are complete — 14 was taken out of order because four of -its items were concrete defects on the apiserver's write path and did not -warrant waiting for a phase. Each phase below has an epic with per-deliverable +Phases 0–14 are complete — 14 was taken out of order because four of its +items were concrete defects on the apiserver's write path and did not warrant +waiting for a phase, and 13 landed after 12 rather than before it despite the +epic's own note to build the road first. Each phase below has an epic with per-deliverable sub-issues carrying a priority label, a size label, and a full PRD. The detail — the review that produced the plan and the `file:line` findings behind each entry — is in [Review and proposed next phases](proposals/next-phases.md). | Phase | Epic | Intent | |---|---|---| -| **13 — CI/CD: official GitHub Actions** | [#133](https://github.com/terasky-oss/declarative-conversion-operator/issues/133) | First-party composite Actions (`setup-convctl` with signature verification, `convctl-test`, `convctl-diff`, `convctl-fleet`) **with their own test workflow** — cross-runner matrix, a negative test proving verification fails on a tampered artifact, and assertions on annotations and job summaries, not just exit codes. Plus a published `convctl` image, CI-native output formats (GitHub annotations, SARIF, markdown), `convctl lint` over a whole repo, distribution via Homebrew / krew / deb / rpm, and a `--package` schema source so a Configuration's XRDs can be tested from a local `.xpkg`, a published image, or an installed `ConfigurationRevision`. | | **15 — Performance and scale** | [#156](https://github.com/terasky-oss/declarative-conversion-operator/issues/156) | Automatic sharding across `ConversionWebhookServer` instances, a measured cold-start budget, per-target memory numbers, a nightly scale run at a raised envelope, and workqueue observability. | | **16 — Engine and strategy expansion** | [#162](https://github.com/terasky-oss/declarative-conversion-operator/issues/162) | `oneOf`/`anyOf` branch mapping, `$ref`/`allOf` flattening, and further strategies driven by real migrations. | From af1bb6c1561883e587e295c3d55091e82197af2d Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:34:31 +0300 Subject: [PATCH 13/24] fix(actions): pass step outputs through env, and shellcheck the Actions MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit shellcheck, via actionlint in CI, caught a real defect in the new test workflow: `[ "${{ steps.clean.outputs.pass }}" -lt 1 ]` interpolates before bash sees it, so an output the Action never set expands to nothing and makes the comparison a bash error rather than a failed assertion — a broken test that reads as a broken Action. Every step output now reaches its script through env with a default, which also removes the interpolation-into-shell pattern generally. None of these values are attacker-controlled, but the habit is the problem. The finding also exposed a gap: actionlint runs shellcheck over workflow run: steps, but reads a composite action's action.yml as a malformed workflow and skips it — so the Actions this project publishes, which hold most of its shell, were never checked at all. hack/shellcheck-actions.sh extracts each run: block, substitutes the GitHub expressions the way the runner would, and shellchecks the result; the Actions job runs it. Twelve scripts, clean. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/actions-test.yml | 72 +++++++++++++++++++++--------- hack/shellcheck-actions.sh | 47 +++++++++++++++++++ 2 files changed, 98 insertions(+), 21 deletions(-) create mode 100755 hack/shellcheck-actions.sh diff --git a/.github/workflows/actions-test.yml b/.github/workflows/actions-test.yml index b43951f..547b0c2 100644 --- a/.github/workflows/actions-test.yml +++ b/.github/workflows/actions-test.yml @@ -52,6 +52,11 @@ jobs: - name: Static checks on the composite Actions run: go test ./test/actions/ -v + # And their shell, which is most of the shell in this repository and + # which actionlint's shellcheck pass never reaches. + - name: Shellcheck the composite Actions + run: hack/shellcheck-actions.sh + setup: name: setup-convctl (${{ matrix.os }}) runs-on: ${{ matrix.os }} @@ -73,12 +78,14 @@ jobs: - name: It resolved a concrete tag and works shell: bash + env: + RESOLVED: ${{ steps.latest.outputs.version }} run: | set -euo pipefail - test -n "${{ steps.latest.outputs.version }}" - case "${{ steps.latest.outputs.version }}" in + test -n "$RESOLVED" + case "$RESOLVED" in v*) ;; - *) echo "resolved version is not a tag: ${{ steps.latest.outputs.version }}" >&2; exit 1 ;; + *) echo "resolved version is not a tag: $RESOLVED" >&2; exit 1 ;; esac convctl version convctl lint --help >/dev/null @@ -94,9 +101,11 @@ jobs: - name: The second install came from the cache and still verified shell: bash + env: + CACHE_HIT: ${{ steps.cached.outputs.cache-hit }} run: | set -euo pipefail - if [ "${{ steps.cached.outputs.cache-hit }}" != "true" ]; then + if [ "$CACHE_HIT" != "true" ]; then echo "second install was not a cache hit; the cached path is untested" >&2 exit 1 fi @@ -141,11 +150,12 @@ jobs: - name: Assert the step failed shell: bash + env: + OUTCOME: ${{ steps.tampered.outcome }} run: | set -euo pipefail - outcome="${{ steps.tampered.outcome }}" - if [ "$outcome" != "failure" ]; then - echo "::error::setup-convctl accepted a tampered archive (outcome: $outcome)" >&2 + if [ "$OUTCOME" != "failure" ]; then + echo "::error::setup-convctl accepted a tampered archive (outcome: $OUTCOME)" >&2 echo "A verification step that cannot be shown to fail is not a verification step." >&2 exit 1 fi @@ -170,18 +180,26 @@ jobs: samples: internal/cli/testdata/full/samples artifact-name: actions-test-clean + # Outputs go through env rather than being interpolated into the + # script: an output the Action never set would otherwise expand to + # nothing and make `[ "" -lt 1 ]` a bash error rather than a failed + # assertion, which reads as a broken test instead of a broken Action. - name: It passed and reported real numbers shell: bash + env: + EXIT_CODE: ${{ steps.clean.outputs.exit-code }} + PASSED: ${{ steps.clean.outputs.pass }} + REPORT: ${{ steps.clean.outputs.report-path }} run: | set -euo pipefail - if [ "${{ steps.clean.outputs.exit-code }}" != "0" ]; then + if [ "$EXIT_CODE" != "0" ]; then echo "a clean fixture did not pass" >&2; exit 1 fi - if [ "${{ steps.clean.outputs.pass }}" -lt 1 ]; then - echo "the pass count is ${{ steps.clean.outputs.pass }}; the Action exited correctly while reporting nothing" >&2 + if [ "${PASSED:-0}" -lt 1 ]; then + echo "the pass count is '${PASSED:-}'; the Action exited correctly while reporting nothing" >&2 exit 1 fi - test -f "${{ steps.clean.outputs.report-path }}" + test -f "$REPORT" # A deliberately broken fixture must fail AND say where. Asserting on # the annotation payload rather than only the exit code: every one of @@ -245,13 +263,16 @@ jobs: - name: Assert it failed, and reported why shell: bash + env: + OUTCOME: ${{ steps.failing.outcome }} + ERRORS: ${{ steps.failing.outputs.errors }} run: | set -euo pipefail - if [ "${{ steps.failing.outcome }}" != "failure" ]; then + if [ "$OUTCOME" != "failure" ]; then echo "::error::a fixture that produces schema violations did not fail the Action" >&2 exit 1 fi - if [ "${{ steps.failing.outputs.errors }}" = "0" ]; then + if [ "${ERRORS:-0}" -lt 1 ]; then echo "::error::the Action failed but reported no errors, so its summary would read as a mystery" >&2 exit 1 fi @@ -268,9 +289,11 @@ jobs: - name: It still ran shell: bash + env: + EXIT_CODE: ${{ steps.quiet.outputs.exit-code }} run: | set -euo pipefail - test "${{ steps.quiet.outputs.exit-code }}" = "0" + test "$EXIT_CODE" = "0" diff-action: name: convctl-diff renders a delta @@ -295,17 +318,21 @@ jobs: - name: A delta was found and did not fail the job shell: bash + env: + HAS_DELTAS: ${{ steps.diff.outputs.has-deltas }} + EXIT_CODE: ${{ steps.diff.outputs.exit-code }} + MARKDOWN: ${{ steps.diff.outputs.markdown-path }} run: | set -euo pipefail - if [ "${{ steps.diff.outputs.has-deltas }}" != "true" ]; then + if [ "$HAS_DELTAS" != "true" ]; then echo "two different configs produced no delta" >&2; exit 1 fi - if [ "${{ steps.diff.outputs.exit-code }}" != "1" ]; then - echo "expected exit 1 for deltas, got ${{ steps.diff.outputs.exit-code }}" >&2; exit 1 + if [ "$EXIT_CODE" != "1" ]; then + echo "expected exit 1 for deltas, got '$EXIT_CODE'" >&2; exit 1 fi # The rendering is the product; an empty one is a silent failure. - test -s "${{ steps.diff.outputs.markdown-path }}" - grep -q '^### Conversion config diff' "${{ steps.diff.outputs.markdown-path }}" + test -s "$MARKDOWN" + grep -q '^### Conversion config diff' "$MARKDOWN" echo "delta rendered, job stayed green" - name: Identical configs report no deltas @@ -320,7 +347,10 @@ jobs: - name: And say so rather than rendering nothing shell: bash + env: + HAS_DELTAS: ${{ steps.same.outputs.has-deltas }} + MARKDOWN: ${{ steps.same.outputs.markdown-path }} run: | set -euo pipefail - test "${{ steps.same.outputs.has-deltas }}" = "false" - grep -q 'No differences' "${{ steps.same.outputs.markdown-path }}" + test "$HAS_DELTAS" = "false" + grep -q 'No differences' "$MARKDOWN" diff --git a/hack/shellcheck-actions.sh b/hack/shellcheck-actions.sh new file mode 100755 index 0000000..492a093 --- /dev/null +++ b/hack/shellcheck-actions.sh @@ -0,0 +1,47 @@ +#!/usr/bin/env bash +# Shellcheck the scripts inside this repository's composite Actions. +# +# actionlint runs shellcheck over workflow `run:` steps, but reads a +# composite action's action.yml as a malformed workflow and skips it — so the +# Actions this project publishes, which contain most of its shell, would +# otherwise never be checked. The class of bug this catches is real: a +# `[ "$X" -lt 1 ]` against an unset value is a bash error rather than a +# failed comparison. +# +# Usage: hack/shellcheck-actions.sh [severity] +set -euo pipefail + +severity="${1:-style}" +root="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +workdir="$(mktemp -d)" +trap 'rm -rf "$workdir"' EXIT + +extracted=0 +for action in "$root"/.github/actions/*/action.yml; do + name="$(basename "$(dirname "$action")")" + mkdir -p "$workdir/$name" + # GitHub substitutes ${{ }} before bash ever sees the script, so they are + # replaced with a literal rather than left for shellcheck to choke on. + python3 - "$action" "$workdir/$name" <<'PY' +import re, sys, pathlib + +src, outdir = pathlib.Path(sys.argv[1]), pathlib.Path(sys.argv[2]) +text = src.read_text() +for i, m in enumerate(re.finditer(r'\n(\s+)run: \|\n((?:\1 .*\n|\n)+)', text)): + indent = len(m.group(1)) + 2 + body = '\n'.join(l[indent:] if len(l) > indent else '' for l in m.group(2).split('\n')) + body = re.sub(r'\$\{\{[^}]*\}\}', 'SUBSTITUTED', body) + (outdir / f'step{i}.sh').write_text('#!/usr/bin/env bash\n' + body) +PY + count="$(find "$workdir/$name" -name '*.sh' | wc -l)" + extracted=$((extracted + count)) + echo "extracted $count script(s) from $name" +done + +if [ "$extracted" -eq 0 ]; then + echo "no scripts extracted; the extraction is broken, not the Actions" >&2 + exit 1 +fi + +shellcheck -S "$severity" "$workdir"/*/*.sh +echo "shellcheck clean across $extracted script(s)" From f742c4710c16901d08642c63c1c2c61c5d2ad27b Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:39:15 +0300 Subject: [PATCH 14/24] fix(ci): pin actions to tags that exist, and check that they do MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit setup-convctl referenced sigstore/cosign-installer@v4, which does not resolve: sigstore publishes a floating v3 tag but only exact v4.x tags. Every local check passed — actionlint validates syntax, shellcheck validates the shell, and neither resolves a `uses:` — and CI failed with "unable to find version v4", an error about the reference rather than about cosign, which took the whole Actions workflow down with it because three of the four Actions compose setup-convctl. Pinned to v4.1.2, and hack/check-action-pins.sh now resolves every action reference in the repository through the API. It distinguishes the cases usefully: a SHA is checked as a commit, a tag or branch as a ref, and a reference to one of this repository's own Actions is checked as a path that exists here, since those cannot resolve by tag until a release carries them. That check immediately found a second, pre-existing one: ossf/scorecard-action@v2 has never resolved either, so the OpenSSF Scorecard job has been failing on every run on main since it was added. Nobody saw it, because Scorecard skips on pull requests — the only place anyone looks at a red check. Pinned to v2.4.4. Dependabot's github-actions ecosystem is extended to the composite Action directories. "/" covers .github/workflows and a root action.yml, not actions in subdirectories, which is exactly where the ones consumers depend on live — so they would have aged out of maintenance while everything else was updated weekly. Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/setup-convctl/action.yml | 5 +- .github/dependabot.yml | 11 +++- .github/workflows/actions-test.yml | 8 +++ .github/workflows/security.yml | 6 +- hack/check-action-pins.sh | 70 ++++++++++++++++++++++++ 5 files changed, 97 insertions(+), 3 deletions(-) create mode 100755 hack/check-action-pins.sh diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml index 8b955ad..84f8a15 100644 --- a/.github/actions/setup-convctl/action.yml +++ b/.github/actions/setup-convctl/action.yml @@ -121,7 +121,10 @@ runs: # costs nothing beyond the download. - name: Install cosign if: inputs.verify == 'true' - uses: sigstore/cosign-installer@v4 + # Pinned to an exact version: sigstore publishes a floating v3 tag but + # no v4, so `@v4` resolves to nothing and the step fails with "unable + # to find version" rather than anything about cosign. + uses: sigstore/cosign-installer@v4.1.2 # Verification runs on a cache hit too. A poisoned cache that could be # laundered by a cache hit would make the whole verification decorative: diff --git a/.github/dependabot.yml b/.github/dependabot.yml index 3ec0c4a..7dc0f78 100644 --- a/.github/dependabot.yml +++ b/.github/dependabot.yml @@ -41,7 +41,16 @@ updates: - go - package-ecosystem: github-actions - directory: / + # The workflows, and the composite Actions this repository publishes. + # "/" alone covers .github/workflows and a root action.yml, not actions + # in subdirectories — which is where ours live, and which are the ones + # consumers depend on. + directories: + - / + - /.github/actions/setup-convctl + - /.github/actions/convctl-test + - /.github/actions/convctl-diff + - /.github/actions/convctl-fleet schedule: interval: weekly day: monday diff --git a/.github/workflows/actions-test.yml b/.github/workflows/actions-test.yml index 547b0c2..a3ee509 100644 --- a/.github/workflows/actions-test.yml +++ b/.github/workflows/actions-test.yml @@ -57,6 +57,14 @@ jobs: - name: Shellcheck the composite Actions run: hack/shellcheck-actions.sh + # Nothing else resolves a `uses:`. A reference to a tag that does not + # exist passes every other check here and fails at run time with an + # error about the reference rather than about the thing it installs. + - name: Every action reference resolves + env: + GH_TOKEN: ${{ github.token }} + run: hack/check-action-pins.sh + setup: name: setup-convctl (${{ matrix.os }}) runs-on: ${{ matrix.os }} diff --git a/.github/workflows/security.yml b/.github/workflows/security.yml index 12f9072..3645d80 100644 --- a/.github/workflows/security.yml +++ b/.github/workflows/security.yml @@ -103,7 +103,11 @@ jobs: persist-credentials: false - name: Run Scorecard - uses: ossf/scorecard-action@v2 + # Exact version: ossf publishes no floating v2 tag, so `@v2` cannot + # resolve — which this job did, silently, on every run on main, + # because Scorecard skips on pull requests and so was never red + # where anyone was looking. + uses: ossf/scorecard-action@v2.4.4 with: results_file: scorecard.sarif results_format: sarif diff --git a/hack/check-action-pins.sh b/hack/check-action-pins.sh new file mode 100755 index 0000000..30ff75a --- /dev/null +++ b/hack/check-action-pins.sh @@ -0,0 +1,70 @@ +#!/usr/bin/env bash +# Check that every third-party action reference in this repository resolves +# to a tag that actually exists. +# +# actionlint validates syntax and shellcheck validates the shell, but neither +# resolves a `uses:` — so `sigstore/cosign-installer@v4` passed every local +# check and failed in CI with "unable to find version v4", because sigstore +# publishes a floating v3 but only exact v4.x tags. The failure surfaces as +# an unresolvable action rather than as anything about the thing it installs, +# which makes it slow to diagnose. +# +# Needs gh authenticated (GITHUB_TOKEN is enough). +set -euo pipefail + +root="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" +failed=0 +checked=0 + +# Every `uses:` across workflows and composite actions, minus local (./) ones. +refs="$(grep -rhoE '^\s*(- )?uses: [^ ]+' \ + "$root/.github/workflows" "$root/.github/actions" "$root/docs/gitops" 2>/dev/null \ + | sed -E 's/^\s*(- )?uses: //' \ + | grep -v '^\./' \ + | sort -u)" + +while IFS= read -r ref; do + [ -n "$ref" ] || continue + path="${ref%@*}" + tag="${ref##*@}" + # owner/repo/sub/action@ref — the repository is the first two segments. + repo="$(cut -d/ -f1,2 <<< "$path")" + + # This repository's own Actions cannot be resolved by tag until a release + # carries them, so what is checked is that the path exists here. A + # reference to an action we do not ship is the error worth catching; the + # tag is the release process's business. + if [ "$repo" = "terasky-oss/declarative-conversion-operator" ]; then + sub="${path#"$repo"/}" + if [ -f "$root/$sub/action.yml" ] || [ -f "$root/$sub/action.yaml" ]; then + checked=$((checked + 1)) + continue + fi + echo "UNRESOLVABLE $ref (this repository ships no $sub)" >&2 + failed=1 + continue + fi + # A 40-character hex string is a commit SHA, which needs a different API. + if [[ "$tag" =~ ^[0-9a-f]{40}$ ]]; then + if gh api "repos/$repo/commits/$tag" --jq .sha >/dev/null 2>&1; then + checked=$((checked + 1)) + continue + fi + echo "UNRESOLVABLE $ref (no such commit)" >&2 + failed=1 + continue + fi + if gh api "repos/$repo/git/ref/tags/$tag" --jq .ref >/dev/null 2>&1 \ + || gh api "repos/$repo/git/ref/heads/$tag" --jq .ref >/dev/null 2>&1; then + checked=$((checked + 1)) + continue + fi + echo "UNRESOLVABLE $ref (no such tag or branch)" >&2 + failed=1 +done <<< "$refs" + +if [ "$failed" -ne 0 ]; then + echo "one or more action references do not resolve" >&2 + exit 1 +fi +echo "all $checked action reference(s) resolve" From 2b95f681753f477bb6eb6dd1acdac3471aa48a8b Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:42:20 +0300 Subject: [PATCH 15/24] fix(ci): do not report an API failure as a missing action reference MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The pin check failed in CI on aquasecurity/trivy-action@v0.36.0, a tag that exists and resolves fine. The script treated any `gh api` failure as "no such tag", so a rate limit, a transient 5xx or a token without the scope all came out as a reference that does not exist — sending whoever read it chasing a pin that was never wrong. Only a 404 now means missing. Anything else is a warning that says the reference could not be checked, with the API's own message, and does not fail the job; there is one retry, because a single blip should not produce either verdict. This is the same mistake as reporting an unlistable version as zero objects: not being able to look is not the same as there being nothing there. Proved by pointing the resolver at references with known outcomes — an existing tag, a missing tag, and a repository that does not exist — rather than by reasoning about it. Co-Authored-By: Claude Opus 5 (1M context) --- hack/check-action-pins.sh | 65 +++++++++++++++++++++++++++++++-------- 1 file changed, 52 insertions(+), 13 deletions(-) diff --git a/hack/check-action-pins.sh b/hack/check-action-pins.sh index 30ff75a..f343557 100755 --- a/hack/check-action-pins.sh +++ b/hack/check-action-pins.sh @@ -15,6 +15,31 @@ set -euo pipefail root="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" failed=0 checked=0 +unchecked=0 + +# resolve prints "ok", "missing", or a reason the check could not be made. +# +# Only a 404 means the reference does not exist. Anything else — a rate +# limit, a transient 5xx, a token without the scope — is a failure to look, +# and reporting that as a missing tag sends somebody chasing a reference that +# is perfectly fine. One retry, because a single blip should not do either. +resolve() { + local endpoint="$1" out status + for attempt in 1 2; do + if out="$(gh api "$endpoint" --jq .ref 2>&1)" || [[ "$out" == *'"sha"'* ]]; then + echo ok + return + fi + status="$out" + if grep -qiE '\(HTTP 404\)|Not Found' <<< "$status"; then + echo missing + return + fi + [ "$attempt" -eq 1 ] && sleep 2 + done + # Collapse the API's message onto one line for the warning. + tr '\n' ' ' <<< "$status" | cut -c1-120 +} # Every `uses:` across workflows and composite actions, minus local (./) ones. refs="$(grep -rhoE '^\s*(- )?uses: [^ ]+' \ @@ -46,25 +71,39 @@ while IFS= read -r ref; do fi # A 40-character hex string is a commit SHA, which needs a different API. if [[ "$tag" =~ ^[0-9a-f]{40}$ ]]; then - if gh api "repos/$repo/commits/$tag" --jq .sha >/dev/null 2>&1; then - checked=$((checked + 1)) - continue + verdict="$(resolve "repos/$repo/commits/$tag")" + else + verdict="$(resolve "repos/$repo/git/ref/tags/$tag")" + if [ "$verdict" = "missing" ]; then + # A branch is a legitimate, if unwise, thing to pin to. + verdict="$(resolve "repos/$repo/git/ref/heads/$tag")" fi - echo "UNRESOLVABLE $ref (no such commit)" >&2 - failed=1 - continue - fi - if gh api "repos/$repo/git/ref/tags/$tag" --jq .ref >/dev/null 2>&1 \ - || gh api "repos/$repo/git/ref/heads/$tag" --jq .ref >/dev/null 2>&1; then - checked=$((checked + 1)) - continue fi - echo "UNRESOLVABLE $ref (no such tag or branch)" >&2 - failed=1 + + case "$verdict" in + ok) + checked=$((checked + 1)) + ;; + missing) + echo "UNRESOLVABLE $ref (no such tag, branch or commit)" >&2 + failed=1 + ;; + *) + # Could not look is not the same as is not there. A rate limit or a + # network blip reported as a missing tag would send somebody chasing + # a reference that is perfectly fine. + echo "WARNING $ref could not be checked ($verdict)" >&2 + unchecked=$((unchecked + 1)) + ;; + esac done <<< "$refs" if [ "$failed" -ne 0 ]; then echo "one or more action references do not resolve" >&2 exit 1 fi +if [ "$unchecked" -gt 0 ]; then + echo "all $checked action reference(s) resolve; $unchecked could not be checked" + exit 0 +fi echo "all $checked action reference(s) resolve" From 78f8933c36840d0721abd28b738af3a9ded37198 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:45:20 +0300 Subject: [PATCH 16/24] fix(release): the published verification command could never succeed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit setup-convctl failed on every runner with "none of the expected identities matched what was in the certificate, got subjects [https://github.com/TeraSky-OSS/declarative-conversion-operator/...]". The certificate records the owner in the casing GitHub renders — TeraSky-OSS — and the pattern said terasky-oss. The error reads like a bad signature rather than a typo, which is what makes it worth a comment. The pattern was copied from this repository's release notes addendum, and that addendum has the same defect: the `cosign verify-blob` command printed into every release, the one a user follows to check an artifact they downloaded, has never worked. Confirmed against the real v0.3.0 assets: the documented command errors, and the same command with the owner matched case-insensitively prints "Verified OK". The signing was fine all along; only the instructions were wrong, which is the failure mode nobody notices, because the people most likely to run it are the least likely to report that they gave up. Both release-notes commands and the Action now use (?i:terasky-oss), and test/actions asserts the casing so a future copy of the lowercase form fails before it ships rather than after. Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/setup-convctl/README.md | 5 +++++ .github/actions/setup-convctl/action.yml | 7 ++++++- .github/workflows/release.yml | 4 ++-- test/actions/actions_test.go | 7 +++++++ 4 files changed, 20 insertions(+), 3 deletions(-) diff --git a/.github/actions/setup-convctl/README.md b/.github/actions/setup-convctl/README.md index 3c92670..aca0fb6 100644 --- a/.github/actions/setup-convctl/README.md +++ b/.github/actions/setup-convctl/README.md @@ -36,6 +36,11 @@ The certificate identity is pinned to **this repository's release workflow at a tag**, not to a wildcard — a signature from any other workflow in any other repository is exactly what this rejects. The OIDC issuer is pinned too. +The owner is matched case-insensitively: Fulcio records the casing GitHub +renders (`TeraSky-OSS`), and a lowercase pattern fails with *"none of the +expected identities matched"*, which reads like a bad signature rather than a +typo. + **A cache hit still verifies.** The cached branch is the one that runs in practice, and a poisoned cache that a cache hit could launder would make the whole thing decorative. diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml index 84f8a15..e2c8829 100644 --- a/.github/actions/setup-convctl/action.yml +++ b/.github/actions/setup-convctl/action.yml @@ -156,10 +156,15 @@ runs: # workflow at a tag, not to a wildcard: a signature from any other # workflow in any other repository is exactly what this is meant to # reject. + # + # The owner is matched case-insensitively. Fulcio records the + # canonical casing GitHub renders — TeraSky-OSS — and a lowercase + # pattern fails with "none of the expected identities matched", + # which reads like a bad signature rather than a typo. cosign verify-blob \ --certificate checksums.txt.pem \ --signature checksums.txt.sig \ - --certificate-identity-regexp '^https://github\.com/terasky-oss/declarative-conversion-operator/\.github/workflows/release\.yml@refs/tags/.*$' \ + --certificate-identity-regexp '^https://github\.com/(?i:terasky-oss)/declarative-conversion-operator/\.github/workflows/release\.yml@refs/tags/.*$' \ --certificate-oidc-issuer https://token.actions.githubusercontent.com \ checksums.txt diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index db8b4a0..8db69f1 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -317,7 +317,7 @@ jobs: echo "" echo '```bash' echo "cosign verify \\" - echo " --certificate-identity-regexp '^https://github.com/terasky-oss/declarative-conversion-operator/\\.github/workflows/release\\.yml@refs/tags/.*\$' \\" + echo " --certificate-identity-regexp '^https://github.com/(?i:terasky-oss)/declarative-conversion-operator/\\.github/workflows/release\\.yml@refs/tags/.*\$' \\" echo " --certificate-oidc-issuer https://token.actions.githubusercontent.com \\" echo " @" echo "" @@ -330,7 +330,7 @@ jobs: echo "cosign verify-blob \\" echo " --certificate checksums.txt.pem \\" echo " --signature checksums.txt.sig \\" - echo " --certificate-identity-regexp '^https://github.com/terasky-oss/declarative-conversion-operator/\\.github/workflows/release\\.yml@refs/tags/.*\$' \\" + echo " --certificate-identity-regexp '^https://github.com/(?i:terasky-oss)/declarative-conversion-operator/\\.github/workflows/release\\.yml@refs/tags/.*\$' \\" echo " --certificate-oidc-issuer https://token.actions.githubusercontent.com \\" echo " checksums.txt" echo '```' diff --git a/test/actions/actions_test.go b/test/actions/actions_test.go index aa929b8..f782da6 100644 --- a/test/actions/actions_test.go +++ b/test/actions/actions_test.go @@ -198,6 +198,13 @@ func TestSetupConvctl_VerifiesByDefaultAndPinsTheIdentity(t *testing.T) { if !strings.Contains(verifyStep, "declarative-conversion-operator/\\.github/workflows/release\\.yml") { t.Error("the certificate identity is not pinned to this repository's release workflow") } + // Fulcio records the owner in GitHub's canonical casing (TeraSky-OSS). + // A lowercase pattern fails with "none of the expected identities + // matched", which reads like a bad signature rather than a typo — and + // did, in this action and in the release notes, until it was run. + if !strings.Contains(verifyStep, "(?i:terasky-oss)") { + t.Error("the owner is matched case-sensitively, so verification fails against the real certificate") + } if !strings.Contains(verifyStep, "--certificate-oidc-issuer https://token.actions.githubusercontent.com") { t.Error("the OIDC issuer is not pinned") } From 5054cee31f4604937dd92f60ccfd22bf10f4b6be Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:47:49 +0300 Subject: [PATCH 17/24] fix: make the binary's version match the tag people install MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit setup-convctl checks that the binary it just installed reports the version that was requested — an installer that puts the wrong binary on PATH has failed even though every step was green. That check failed against the real v0.3.0 release: the binary reports 0.3.0, because GoReleaser's .Version strips the leading v that the tag, the release page and every install command carry. Two changes, because the mismatch has two halves. The ldflags now use .Tag, so future builds report v0.5.0 for the release everyone calls v0.5.0 rather than making each consumer know about the difference. Confirmed with a snapshot build: `convctl version` prints "v0.3.0 (78f8933c3684) linux/amd64 go1.26.6". And the Action compares with the leading v stripped from both sides regardless, because releases built before this existed will always report the bare form, and an installer that rejected its own project's published releases would be worse than one that never checked. The archive names keep .Version — they are the bare form on the release page already, and renaming them would break every existing download URL. Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/setup-convctl/action.yml | 7 ++++++- .goreleaser.yaml | 5 ++++- 2 files changed, 10 insertions(+), 2 deletions(-) diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml index e2c8829..19f89b3 100644 --- a/.github/actions/setup-convctl/action.yml +++ b/.github/actions/setup-convctl/action.yml @@ -218,7 +218,12 @@ runs: run: | set -euo pipefail got="$("$BINARY" version | awk '{print $1}')" - if [ "$got" != "$EXPECTED" ]; then + # Compare without the leading v. Releases before this Action existed + # were built with GoReleaser's .Version, which strips it, so the + # binary reports 0.3.0 for tag v0.3.0 — and an installer that + # rejected its own project's releases would be worse than one that + # did not check at all. + if [ "${got#v}" != "${EXPECTED#v}" ]; then echo "::error::installed convctl reports $got, expected $EXPECTED" >&2 exit 1 fi diff --git a/.goreleaser.yaml b/.goreleaser.yaml index 3c760d8..1d1799d 100644 --- a/.goreleaser.yaml +++ b/.goreleaser.yaml @@ -18,7 +18,10 @@ builds: # commit and date go in too: "convctl says this conversion is lossy" # is unactionable in a bug report without knowing which convctl. - -s -w - - -X main.version={{ .Version }} + # .Tag, not .Version: .Version strips the leading v, so the binary + # reported 0.3.0 for the release everyone calls v0.3.0 — a mismatch + # every consumer has to know about to compare the two. + - -X main.version={{ .Tag }} - -X main.commit={{ .FullCommit }} - -X main.date={{ .Date }} From f6b924328fa9ecd8ea95a8a0daa90c02fdb5fa62 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:51:28 +0300 Subject: [PATCH 18/24] test(actions): exercise the Actions against the build under review MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The test workflow installed a published convctl and then asserted on lint, --validate-output and markdown output — none of which exist in v0.3.0, because all three are added by this pull request and the one before it. So the jobs failed for the only reason they could: they were testing the last release, not the change. setup-convctl gains a `binary` input that takes a convctl which already exists and skips resolving, downloading and verifying it — there is no artifact whose provenance could be in question — putting it on PATH and reporting its version. Every Action that composes setup-convctl passes it through. That is not a test-shaped hole in a public Action. A runner that vendors the binary, or one with no egress to the releases API, wants exactly this; the workflows here are simply the first consumer with that requirement. The test-action and diff-action jobs now build the pull request's convctl and hand it over, so the fixtures can exercise anything the branch adds. The setup-convctl job still installs a real published release — that is the path consumers take, and the download, the cache and the signature verification are what it exists to prove — so its assertions are limited to commands the oldest supported release has. Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/convctl-diff/action.yml | 7 ++++ .github/actions/convctl-fleet/action.yml | 7 ++++ .github/actions/convctl-test/action.yml | 7 ++++ .github/actions/setup-convctl/README.md | 12 ++++++ .github/actions/setup-convctl/action.yml | 47 +++++++++++++++++++++--- .github/workflows/actions-test.yml | 39 +++++++++++++++++++- 6 files changed, 112 insertions(+), 7 deletions(-) diff --git a/.github/actions/convctl-diff/action.yml b/.github/actions/convctl-diff/action.yml index 0ffbf2a..aeea71a 100644 --- a/.github/actions/convctl-diff/action.yml +++ b/.github/actions/convctl-diff/action.yml @@ -50,6 +50,12 @@ inputs: description: convctl release tag to install. required: false default: latest + binary: + description: >- + Path to a convctl that already exists, used instead of downloading one. + See setup-convctl. + required: false + default: "" token: description: Token used to upsert the comment. required: false @@ -73,6 +79,7 @@ runs: uses: ./.github/actions/setup-convctl with: version: ${{ inputs.version }} + binary: ${{ inputs.binary }} - name: Write kubeconfig if: inputs.kubeconfig != '' diff --git a/.github/actions/convctl-fleet/action.yml b/.github/actions/convctl-fleet/action.yml index 31e6b3a..87589a5 100644 --- a/.github/actions/convctl-fleet/action.yml +++ b/.github/actions/convctl-fleet/action.yml @@ -52,6 +52,12 @@ inputs: description: convctl release tag to install. required: false default: latest + binary: + description: >- + Path to a convctl that already exists, used instead of downloading one. + See setup-convctl. + required: false + default: "" upload-artifact: description: Upload the aggregated JUnit report. required: false @@ -82,6 +88,7 @@ runs: uses: ./.github/actions/setup-convctl with: version: ${{ inputs.version }} + binary: ${{ inputs.binary }} - name: Write kubeconfig if: inputs.kubeconfig != '' diff --git a/.github/actions/convctl-test/action.yml b/.github/actions/convctl-test/action.yml index ed9f5f0..2e4104e 100644 --- a/.github/actions/convctl-test/action.yml +++ b/.github/actions/convctl-test/action.yml @@ -71,6 +71,12 @@ inputs: description: convctl release tag to install. required: false default: latest + binary: + description: >- + Path to a convctl that already exists, used instead of downloading one. + See setup-convctl. + required: false + default: "" upload-artifact: description: Upload the JUnit report as a workflow artifact. required: false @@ -110,6 +116,7 @@ runs: uses: ./.github/actions/setup-convctl with: version: ${{ inputs.version }} + binary: ${{ inputs.binary }} - name: Write kubeconfig if: inputs.kubeconfig != '' diff --git a/.github/actions/setup-convctl/README.md b/.github/actions/setup-convctl/README.md index aca0fb6..9addb0d 100644 --- a/.github/actions/setup-convctl/README.md +++ b/.github/actions/setup-convctl/README.md @@ -16,6 +16,7 @@ Installs a **verified** `convctl` onto `PATH`. | `version` | `latest` | a release tag, or `latest` | | `verify` | `true` | verify the cosign signature on `checksums.txt`, then the archive's checksum against it | | `token` | `${{ github.token }}` | for the releases API, to avoid anonymous rate limits | +| `binary` | — | path to a `convctl` that already exists, used instead of downloading one | ## Outputs @@ -48,6 +49,17 @@ whole thing decorative. Set `verify: false` only if you have a reason; cosign is not installed at all in that case, so it costs nothing to leave on. +## Using a binary you already have + +`binary: /path/to/convctl` skips resolving, downloading and verifying +entirely — there is no artifact whose provenance could be in question — and +just puts it on `PATH`. For a runner that vendors the binary, one with no +egress to the releases API, and for this repository's own workflows, which +have to exercise the Actions against the build under test rather than against +the last release. + +Every Action that composes this one accepts it too. + ## Platforms `ubuntu-latest`, `macos-latest` and `windows-latest`, on amd64 and arm64. diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml index 19f89b3..3e8b132 100644 --- a/.github/actions/setup-convctl/action.yml +++ b/.github/actions/setup-convctl/action.yml @@ -21,14 +21,23 @@ inputs: Token for the GitHub releases API, to avoid anonymous rate limits. required: false default: ${{ github.token }} + binary: + description: >- + Path to a convctl that already exists, used instead of downloading + one. For a runner that vendors the binary or has no egress to the + releases API — and for this repository's own workflows, which have to + exercise the Actions against the build under test rather than against + the last release. + required: false + default: "" outputs: version: - description: The resolved release tag. - value: ${{ steps.resolve.outputs.version }} + description: The resolved release tag, or the supplied binary's version. + value: ${{ inputs.binary != '' && steps.supplied.outputs.version || steps.resolve.outputs.version }} path: - description: Path of the installed binary. - value: ${{ steps.install.outputs.path }} + description: Path of the binary that ended up on PATH. + value: ${{ inputs.binary != '' && steps.supplied.outputs.path || steps.install.outputs.path }} cache-hit: description: Whether the binary came from the runner cache. value: ${{ steps.cache.outputs.cache-hit }} @@ -36,10 +45,33 @@ outputs: runs: using: composite steps: + # A binary that is already here needs no resolving, downloading or + # verifying: there is no artifact whose provenance could be in question. + - name: Use the supplied binary + id: supplied + if: inputs.binary != '' + shell: bash + env: + BINARY: ${{ inputs.binary }} + run: | + set -euo pipefail + if [ ! -x "$BINARY" ]; then + echo "::error::binary input points at $BINARY, which is not an executable file" >&2 + exit 1 + fi + dir="$(cd "$(dirname "$BINARY")" && pwd)" + echo "$dir" >> "$GITHUB_PATH" + { + echo "path=$dir/$(basename "$BINARY")" + echo "version=$("$BINARY" version | awk '{print $1}')" + } >> "$GITHUB_OUTPUT" + echo "using the supplied convctl at $BINARY" >> "$GITHUB_STEP_SUMMARY" + # Resolve "latest" to a concrete tag and pin it in the step summary, so a # re-run six months from now is explainable rather than mysterious. - name: Resolve version and platform id: resolve + if: inputs.binary == '' shell: bash env: GH_TOKEN: ${{ inputs.token }} @@ -88,13 +120,14 @@ runs: - name: Restore cached binary id: cache + if: inputs.binary == '' uses: actions/cache@v4 with: path: ${{ steps.resolve.outputs.dir }} key: convctl-${{ steps.resolve.outputs.version }}-${{ steps.resolve.outputs.os }}-${{ steps.resolve.outputs.arch }} - name: Download release assets - if: steps.cache.outputs.cache-hit != 'true' + if: inputs.binary == '' && steps.cache.outputs.cache-hit != 'true' shell: bash env: GH_TOKEN: ${{ inputs.token }} @@ -130,7 +163,7 @@ runs: # laundered by a cache hit would make the whole verification decorative: # the cached branch is the one that runs in practice. - name: Verify signature and checksum - if: inputs.verify == 'true' + if: inputs.binary == '' && inputs.verify == 'true' shell: bash env: GH_TOKEN: ${{ inputs.token }} @@ -188,6 +221,7 @@ runs: - name: Extract and add to PATH id: install + if: inputs.binary == '' shell: bash env: ARCHIVE: ${{ steps.resolve.outputs.archive }} @@ -211,6 +245,7 @@ runs: echo "path=$DIR/$BIN" >> "$GITHUB_OUTPUT" - name: Verify it runs + if: inputs.binary == '' shell: bash env: EXPECTED: ${{ steps.resolve.outputs.version }} diff --git a/.github/workflows/actions-test.yml b/.github/workflows/actions-test.yml index a3ee509..73d2cc7 100644 --- a/.github/workflows/actions-test.yml +++ b/.github/workflows/actions-test.yml @@ -95,8 +95,12 @@ jobs: v*) ;; *) echo "resolved version is not a tag: $RESOLVED" >&2; exit 1 ;; esac + # Only commands the oldest supported release has: this job + # deliberately installs a published binary, so asserting on a + # feature added in the pull request would test the release rather + # than the Action. convctl version - convctl lint --help >/dev/null + convctl validate --help >/dev/null # The cached branch is the one that runs in practice and the one that # silently rots, so it is exercised explicitly — including that it @@ -178,6 +182,22 @@ jobs: with: persist-credentials: false + - uses: actions/setup-go@v6 + with: + go-version-file: go.mod + cache: true + + # The Actions are exercised against the build under review, not against + # the last release. Otherwise every fixture here would be limited to + # features that already shipped — which is the opposite of what a pull + # request needs to prove. + - name: Build the convctl under test + id: build + run: | + set -euo pipefail + go build -o "$RUNNER_TEMP/convctl" ./cmd/convctl + echo "binary=$RUNNER_TEMP/convctl" >> "$GITHUB_OUTPUT" + # A clean fixture passes and still produces a report. - name: Clean fixture id: clean @@ -187,6 +207,7 @@ jobs: xrd: internal/cli/testdata/full/xrd.yaml samples: internal/cli/testdata/full/samples artifact-name: actions-test-clean + binary: ${{ steps.build.outputs.binary }} # Outputs go through env rather than being interpolated into the # script: an output the Action never set would otherwise expand to @@ -268,6 +289,7 @@ jobs: samples: internal/cli/testdata/validate-output/samples validate-output: "true" artifact-name: actions-test-failing + binary: ${{ steps.build.outputs.binary }} - name: Assert it failed, and reported why shell: bash @@ -294,6 +316,7 @@ jobs: samples: internal/cli/testdata/full/samples annotate: "false" upload-artifact: "false" + binary: ${{ steps.build.outputs.binary }} - name: It still ran shell: bash @@ -312,6 +335,18 @@ jobs: with: persist-credentials: false + - uses: actions/setup-go@v6 + with: + go-version-file: go.mod + cache: true + + - name: Build the convctl under test + id: build + run: | + set -euo pipefail + go build -o "$RUNNER_TEMP/convctl" ./cmd/convctl + echo "binary=$RUNNER_TEMP/convctl" >> "$GITHUB_OUTPUT" + # comment:false — this job proves the rendering and the exit-code # semantics; commenting needs a pull request and a writable token. - name: Two configs that differ @@ -323,6 +358,7 @@ jobs: examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml xrd: examples/crossplane-xr-multiversion/03-promote-v2/xrd.yaml comment: "false" + binary: ${{ steps.build.outputs.binary }} - name: A delta was found and did not fail the job shell: bash @@ -352,6 +388,7 @@ jobs: examples/field-rename/xrdconversionconfig.yaml xrd: examples/field-rename/xrd.yaml comment: "false" + binary: ${{ steps.build.outputs.binary }} - name: And say so rather than rendering nothing shell: bash From 42ba623494258602bdd6215d87caf8104f2dbad8 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Tue, 15 Sep 2026 23:54:57 +0300 Subject: [PATCH 19/24] test(actions): test the cache the way the cache actually behaves MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Two failures, both from assuming actions/cache works differently than it does. The cache-hit assertion installed twice in one job and expected the second to hit. actions/cache saves in the job's POST step, so a second call in the same job can never hit — the cached path is a later job or run, which is also the one that runs in practice and the one that silently rots. It is now its own job, needs: the matrix job that populated it. The tamper test corrupted the archive and re-installed, expecting rejection. The re-install downloaded with --clobber and quietly repaired the corruption, so the verification it was supposed to prove never saw a bad file. The download step now keeps an asset that is already present: the download is not what makes an archive trustworthy, the verification is, and re-fetching undoes exactly the tampering that step exists to catch. It is also correct for a runner with a warm RUNNER_TEMP. That left one more hole, which is the same failure one level up: on a warm cache the second install would restore the good archive over the corrupted one, and the job would pass while testing nothing. setup-convctl gains a `cache` input and the tamper job sets it to false, so the test is deterministic whatever the cache holds. Consumers get the knob too — a runner that does not want cache writes has an answer that is not "fork the Action". Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/setup-convctl/README.md | 1 + .github/actions/setup-convctl/action.yml | 25 +++++++++++++--- .github/workflows/actions-test.yml | 38 +++++++++++++++++------- 3 files changed, 49 insertions(+), 15 deletions(-) diff --git a/.github/actions/setup-convctl/README.md b/.github/actions/setup-convctl/README.md index 9addb0d..84a6582 100644 --- a/.github/actions/setup-convctl/README.md +++ b/.github/actions/setup-convctl/README.md @@ -16,6 +16,7 @@ Installs a **verified** `convctl` onto `PATH`. | `version` | `latest` | a release tag, or `latest` | | `verify` | `true` | verify the cosign signature on `checksums.txt`, then the archive's checksum against it | | `token` | `${{ github.token }}` | for the releases API, to avoid anonymous rate limits | +| `cache` | `true` | restore and save through `actions/cache` | | `binary` | — | path to a `convctl` that already exists, used instead of downloading one | ## Outputs diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml index 3e8b132..1204352 100644 --- a/.github/actions/setup-convctl/action.yml +++ b/.github/actions/setup-convctl/action.yml @@ -21,6 +21,14 @@ inputs: Token for the GitHub releases API, to avoid anonymous rate limits. required: false default: ${{ github.token }} + cache: + description: >- + Restore and save the binary through actions/cache. Turn it off for a + runner where the cache is not wanted — or where a deterministic + download is, as in this repository's own tamper test, since a cache + restore would overwrite the archive the test deliberately corrupted. + required: false + default: "true" binary: description: >- Path to a convctl that already exists, used instead of downloading @@ -120,7 +128,7 @@ runs: - name: Restore cached binary id: cache - if: inputs.binary == '' + if: inputs.binary == '' && inputs.cache == 'true' uses: actions/cache@v4 with: path: ${{ steps.resolve.outputs.dir }} @@ -144,10 +152,18 @@ runs: if [ "$VERIFY" = "true" ]; then assets+=(checksums.txt.sig checksums.txt.pem) fi + # An asset already in the directory is kept rather than re-fetched: + # the download is not what makes an archive trustworthy, the + # verification below is, and re-downloading would quietly repair + # exactly the tampering that step exists to catch. for asset in "${assets[@]}"; do + if [ -f "$asset" ]; then + echo "$asset is already present" + continue + fi gh release download "$VERSION" \ --repo TeraSky-OSS/declarative-conversion-operator \ - --pattern "$asset" --clobber + --pattern "$asset" done # Cosign is only installed when it is going to be used, so verify: false @@ -176,12 +192,13 @@ runs: cd "$DIR" # A cache holds the archive but not necessarily the signature - # material, so fetch whatever is missing before verifying. + # material, so fetch whatever is missing before verifying. Present + # files are left alone, for the same reason as above. for asset in checksums.txt checksums.txt.sig checksums.txt.pem "$ARCHIVE"; do if [ ! -f "$asset" ]; then gh release download "$VERSION" \ --repo TeraSky-OSS/declarative-conversion-operator \ - --pattern "$asset" --clobber + --pattern "$asset" fi done diff --git a/.github/workflows/actions-test.yml b/.github/workflows/actions-test.yml index 73d2cc7..1499dcf 100644 --- a/.github/workflows/actions-test.yml +++ b/.github/workflows/actions-test.yml @@ -102,23 +102,30 @@ jobs: convctl version convctl validate --help >/dev/null - # The cached branch is the one that runs in practice and the one that - # silently rots, so it is exercised explicitly — including that it - # still verified. - - name: Install again (cache hit) + # actions/cache saves in the job's post step, so a second install inside + # one job can never hit it. The cached path is a LATER job or run — which + # is also the one that runs in practice, and the one that silently rots. + cache: + name: setup-convctl uses the cache and still verifies + runs-on: ubuntu-latest + needs: setup + steps: + - uses: actions/checkout@v7 + with: + persist-credentials: false + + - name: Install, expecting the cache populated by the matrix job id: cached uses: ./.github/actions/setup-convctl - with: - version: ${{ steps.latest.outputs.version }} - - name: The second install came from the cache and still verified + - name: It was a cache hit, and it verified anyway shell: bash env: CACHE_HIT: ${{ steps.cached.outputs.cache-hit }} run: | set -euo pipefail if [ "$CACHE_HIT" != "true" ]; then - echo "second install was not a cache hit; the cached path is untested" >&2 + echo "::error::not a cache hit, so the cached path is still untested" >&2 exit 1 fi convctl version @@ -134,11 +141,13 @@ jobs: with: persist-credentials: false - - name: Install once to populate the cache directory + - name: Install once, so a real archive is on disk to corrupt id: real uses: ./.github/actions/setup-convctl + with: + cache: "false" - - name: Corrupt the cached archive + - name: Corrupt the downloaded archive shell: bash env: VERSION: ${{ steps.real.outputs.version }} @@ -149,16 +158,23 @@ jobs: archive="$(find "$dir" -name '*.tar.gz' | head -1)" test -n "$archive" # Same size, different bytes: a length check would not catch this, - # which is the point of checking the hash. + # which is the point of checking the hash. The second install + # keeps the archive that is already there, so the corruption + # survives to be verified rather than being repaired by a fresh + # download. printf 'tampered' | dd of="$archive" bs=1 seek=0 conv=notrunc status=none echo "corrupted $archive" + # cache: false, so a warm cache cannot restore the good archive over + # the corrupted one — which would leave this job green while testing + # nothing at all, the exact failure it exists to prevent one level up. - name: Verification must reject it id: tampered continue-on-error: true uses: ./.github/actions/setup-convctl with: version: ${{ steps.real.outputs.version }} + cache: "false" - name: Assert the step failed shell: bash From 3f5f12901c3c3668c10b7000718ea4f05b1c780b Mon Sep 17 00:00:00 2001 From: vrabbi Date: Wed, 16 Sep 2026 00:05:28 +0300 Subject: [PATCH 20/24] fix: address the review findings on the Actions, sampling and packaging MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Twelve findings. Eleven fixed here; one was already fixed in af1bb6c. Two were wrong answers to questions this phase set out to answer. --sample-strategy first stops listing at the cap, so seen == kept, so the report keyed on that comparison said nothing — which made the one strategy that CANNOT know the population the one that silently claimed to have covered it. Exactly the false confidence --max-samples exists to prevent. Stopping early is now recorded as truncation, the population is reported as unknown rather than as a number equal to the sample, and the JUnit property says "unknown" too. setup-convctl skipped extraction when the binary was already present, so a cache entry holding an intact archive beside an altered binary would have been used as it stood: the checksum covers the archive, not what was extracted from it. With verification on, the binary is now always replaced from the archive that was just verified. Two were security defects in inputs this project does not control. lint staged a package's XRDs using metadata.name as a filename, so a package declaring an XRD called ../../evil would write outside the temporary directory with the invoking user's permissions — and a .xpkg is a file pulled from a registry. Base name plus index now, which also stops two XRDs of the same name overwriting each other. The Homebrew cask stripped com.apple.quarantine after install. That turns Gatekeeper off for the binary, and a checksum is not a substitute for signing: it proves the file is what the release published, not that anyone vouched for what it does. The hook is gone; macOS will warn until these artifacts are notarized, which is the honest state rather than one papered over during install. One was a publishing defect that would have broken every consumer. convctl-test, convctl-diff and convctl-fleet each composed `./.github/actions/setup-convctl`, and a local path inside a published composite action resolves against the CONSUMER's workspace — so it worked in these tests and nowhere else. They no longer install convctl at all: they check PATH and say which Action to run first. Setup once per job rather than once per Action, which is also less work. The rest: - The sticky-comment lookup spliced the marker into a jq filter. The default tag is the config input, multiline for the two-config form, and a newline inside a jq string literal is a syntax error — so the lookup would fail and every run would post a new comment instead of editing one. Passed with --arg now, and the tag is collapsed to one line. - convctl-fleet counted failures with `grep -c ... || echo 0`. grep -c prints 0 AND exits 1 when nothing matches, so the guard appended a second zero and made failed-clusters a two-line value — failing a run in which every cluster passed. Proven and fixed. - RunTest is exported, so the cobra validation is not the only way in. Invalid sampling options reaching the sampler do the opposite of what they say: a negative cap disables the bound and paginates everything into memory. Validated in RunTest. - The release workflow's third-party Actions are pinned to commit SHAs, with the tag in a comment. It is the only workflow handling signing identities and publishing credentials, and it now carries tap and bucket tokens; a retargeted tag there is a different risk from one in a test job. - The Configuration CI guide piped install.sh from crossplane's main branch into a shell. Pinned release, checksum verified. Both guides now install convctl with setup-convctl at a version rather than `go install ...@latest`, which resolves at run time and verifies nothing. Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/convctl-diff/README.md | 13 ++++ .github/actions/convctl-diff/action.yml | 47 ++++++++----- .github/actions/convctl-fleet/README.md | 13 ++++ .github/actions/convctl-fleet/action.yml | 42 +++++++----- .github/actions/convctl-test/README.md | 13 ++++ .github/actions/convctl-test/action.yml | 31 ++++----- .github/actions/setup-convctl/action.yml | 10 ++- .github/workflows/actions-test.yml | 19 +++--- .github/workflows/release.yml | 36 +++++----- .goreleaser.yaml | 17 ++--- docs/gitops/configuration-ci.md | 24 +++++-- docs/gitops/convctl-fleet.gha.yml | 5 ++ docs/gitops/fleet-ci.md | 9 ++- docs/installation.md | 2 +- internal/cli/lint.go | 17 ++++- internal/cli/report.go | 9 ++- internal/cli/sampling.go | 50 +++++++++++--- internal/cli/sampling_test.go | 85 ++++++++++++++++++++++++ internal/cli/test.go | 9 +++ internal/cli/xpkg_test.go | 47 +++++++++++++ test/actions/actions_test.go | 23 +++++++ 21 files changed, 413 insertions(+), 108 deletions(-) diff --git a/.github/actions/convctl-diff/README.md b/.github/actions/convctl-diff/README.md index 299a401..6af6400 100644 --- a/.github/actions/convctl-diff/README.md +++ b/.github/actions/convctl-diff/README.md @@ -10,6 +10,8 @@ permissions: pull-requests: write steps: + # convctl comes from setup-convctl, once per job. + - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-diff@v1 with: config: apis/widgets/conversion.yaml @@ -46,3 +48,14 @@ A `pull_request` event from a fork has a read-only token. The Action emits a check a contributor cannot fix teaches them to ignore red checks. Needs `pull-requests: write` to comment. + +## Requires `setup-convctl` + +This Action consumes `convctl` from `PATH` and does not install it. Run +[`setup-convctl`](../setup-convctl) first — once per job, however many of +these Actions follow. + +That is not an ergonomic preference. A composite action cannot reference a +local action by path once published: `./…` resolves against the **consumer's** +workspace, so a nested setup step would work in this repository's own tests +and fail for everyone else. diff --git a/.github/actions/convctl-diff/action.yml b/.github/actions/convctl-diff/action.yml index aeea71a..b18d0f2 100644 --- a/.github/actions/convctl-diff/action.yml +++ b/.github/actions/convctl-diff/action.yml @@ -46,16 +46,6 @@ inputs: error (exit 2) always fails. required: false default: "false" - version: - description: convctl release tag to install. - required: false - default: latest - binary: - description: >- - Path to a convctl that already exists, used instead of downloading one. - See setup-convctl. - required: false - default: "" token: description: Token used to upsert the comment. required: false @@ -75,11 +65,22 @@ outputs: runs: using: composite steps: - - name: Set up convctl - uses: ./.github/actions/setup-convctl - with: - version: ${{ inputs.version }} - binary: ${{ inputs.binary }} + # convctl is not installed here. A composite action cannot reference a + # local action by path once it is published: `./…` resolves against the + # CONSUMER's workspace, not this repository, so a nested setup step + # works in these tests and fails for everyone else. Run setup-convctl + # first — which also means installing once per job rather than once per + # Action. + - name: Check convctl is on PATH + shell: bash + run: | + set -euo pipefail + if ! command -v convctl >/dev/null 2>&1; then + echo "::error::convctl is not on PATH. Run the setup-convctl Action before this one:" >&2 + echo " - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1" >&2 + exit 1 + fi + convctl version - name: Write kubeconfig if: inputs.kubeconfig != '' @@ -172,7 +173,11 @@ runs: exit 0 fi - marker="" + # Collapse the tag onto one line: it defaults to the config input, + # which is multiline when two configs are being compared, and a + # marker containing a newline would never match itself. + tag="$(tr '\n' '_' <<< "$TAG" | sed 's/_$//')" + marker="" body="$RUNNER_TEMP/convctl-diff-comment.md" { echo "$marker" @@ -181,8 +186,16 @@ runs: # Find this tag's comment and edit it, so repeated runs update one # comment rather than appending a new one each push. + # + # The marker goes to jq as data, not spliced into the filter. The + # default tag is the config input, which is multiline for the + # two-config form — and a newline inside a jq string literal is a + # syntax error, so the lookup would fail and every run would post a + # new comment. A quote or backslash in an explicit tag does the + # same. existing="$(gh api "repos/$REPO/issues/$PR/comments" --paginate \ - --jq "[.[] | select(.body | contains(\"$marker\"))] | .[0].id // empty")" + | jq -r --arg marker "$marker" \ + '[.[] | select(.body | contains($marker))] | .[0].id // empty')" if [ -n "$existing" ]; then gh api --method PATCH "repos/$REPO/issues/comments/$existing" \ diff --git a/.github/actions/convctl-fleet/README.md b/.github/actions/convctl-fleet/README.md index 196da23..9c69ae7 100644 --- a/.github/actions/convctl-fleet/README.md +++ b/.github/actions/convctl-fleet/README.md @@ -4,6 +4,8 @@ Runs `convctl test --live` against every cluster in a fleet and aggregates the result into one JUnit report, with one `` per cluster. ```yaml +# convctl comes from setup-convctl, once per job. +- uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-fleet@v1 with: config: apis/widgets/conversion.yaml @@ -33,3 +35,14 @@ set up, but a clearer failure surface and parallel execution. Use ## Outputs `exit-code`, `report-path`, `clusters`, `failed-clusters`. + +## Requires `setup-convctl` + +This Action consumes `convctl` from `PATH` and does not install it. Run +[`setup-convctl`](../setup-convctl) first — once per job, however many of +these Actions follow. + +That is not an ergonomic preference. A composite action cannot reference a +local action by path once published: `./…` resolves against the **consumer's** +workspace, so a nested setup step would work in this repository's own tests +and fail for everyone else. diff --git a/.github/actions/convctl-fleet/action.yml b/.github/actions/convctl-fleet/action.yml index 87589a5..24a07ed 100644 --- a/.github/actions/convctl-fleet/action.yml +++ b/.github/actions/convctl-fleet/action.yml @@ -48,16 +48,6 @@ inputs: description: Validate converted objects against the destination schema. required: false default: "false" - version: - description: convctl release tag to install. - required: false - default: latest - binary: - description: >- - Path to a convctl that already exists, used instead of downloading one. - See setup-convctl. - required: false - default: "" upload-artifact: description: Upload the aggregated JUnit report. required: false @@ -84,11 +74,22 @@ outputs: runs: using: composite steps: - - name: Set up convctl - uses: ./.github/actions/setup-convctl - with: - version: ${{ inputs.version }} - binary: ${{ inputs.binary }} + # convctl is not installed here. A composite action cannot reference a + # local action by path once it is published: `./…` resolves against the + # CONSUMER's workspace, not this repository, so a nested setup step + # works in these tests and fails for everyone else. Run setup-convctl + # first — which also means installing once per job rather than once per + # Action. + - name: Check convctl is on PATH + shell: bash + run: | + set -euo pipefail + if ! command -v convctl >/dev/null 2>&1; then + echo "::error::convctl is not on PATH. Run the setup-convctl Action before this one:" >&2 + echo " - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1" >&2 + exit 1 + fi + convctl version - name: Write kubeconfig if: inputs.kubeconfig != '' @@ -144,9 +145,16 @@ runs: clusters=0 failed=0 if [ -f "$report" ]; then - clusters=$(grep -c ']*' "$report" \ - | grep -c -E 'failures="[1-9]|errors="[1-9]' || echo 0) + | grep -c -E 'failures="[1-9]|errors="[1-9]' || true) + clusters=${clusters:-0} + failed=${failed:-0} { echo "### convctl fleet" echo "" diff --git a/.github/actions/convctl-test/README.md b/.github/actions/convctl-test/README.md index a645413..9888b9b 100644 --- a/.github/actions/convctl-test/README.md +++ b/.github/actions/convctl-test/README.md @@ -3,6 +3,8 @@ Runs `convctl test` and puts the result where a reviewer will see it. ```yaml +# convctl comes from setup-convctl, once per job. +- uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-test@v1 with: config: apis/widgets/conversion.yaml @@ -43,3 +45,14 @@ Behaviour: `fail-on`, `strict`, `validate-output`, `concurrency`, `version`, `kubeconfig` is written with `umask 077` **before** the file is created rather than `chmod`-ed afterwards — between creation and chmod the file is briefly world-readable — and is never echoed. + +## Requires `setup-convctl` + +This Action consumes `convctl` from `PATH` and does not install it. Run +[`setup-convctl`](../setup-convctl) first — once per job, however many of +these Actions follow. + +That is not an ergonomic preference. A composite action cannot reference a +local action by path once published: `./…` resolves against the **consumer's** +workspace, so a nested setup step would work in this repository's own tests +and fail for everyone else. diff --git a/.github/actions/convctl-test/action.yml b/.github/actions/convctl-test/action.yml index 2e4104e..4feb61e 100644 --- a/.github/actions/convctl-test/action.yml +++ b/.github/actions/convctl-test/action.yml @@ -67,16 +67,6 @@ inputs: description: Parallel workers. required: false default: "" - version: - description: convctl release tag to install. - required: false - default: latest - binary: - description: >- - Path to a convctl that already exists, used instead of downloading one. - See setup-convctl. - required: false - default: "" upload-artifact: description: Upload the JUnit report as a workflow artifact. required: false @@ -112,11 +102,22 @@ outputs: runs: using: composite steps: - - name: Set up convctl - uses: ./.github/actions/setup-convctl - with: - version: ${{ inputs.version }} - binary: ${{ inputs.binary }} + # convctl is not installed here. A composite action cannot reference a + # local action by path once it is published: `./…` resolves against the + # CONSUMER's workspace, not this repository, so a nested setup step + # works in these tests and fails for everyone else. Run setup-convctl + # first — which also means installing once per job rather than once per + # Action. + - name: Check convctl is on PATH + shell: bash + run: | + set -euo pipefail + if ! command -v convctl >/dev/null 2>&1; then + echo "::error::convctl is not on PATH. Run the setup-convctl Action before this one:" >&2 + echo " - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1" >&2 + exit 1 + fi + convctl version - name: Write kubeconfig if: inputs.kubeconfig != '' diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml index 1204352..3d9e5fc 100644 --- a/.github/actions/setup-convctl/action.yml +++ b/.github/actions/setup-convctl/action.yml @@ -245,11 +245,19 @@ runs: EXT: ${{ steps.resolve.outputs.ext }} BIN: ${{ steps.resolve.outputs.bin }} DIR: ${{ steps.resolve.outputs.dir }} + VERIFY: ${{ inputs.verify }} run: | set -euo pipefail cd "$DIR" - if [ ! -f "$BIN" ]; then + # With verification on, the binary is always re-extracted from the + # archive that was just verified. Keeping an existing $BIN would + # trust a file nothing checked: the checksum covers the archive, and + # a cache entry can hold an intact archive beside an altered binary. + # That is the shape of the attack the verification exists to stop, + # and skipping extraction would walk straight past it. + if [ "$VERIFY" = "true" ] || [ ! -f "$BIN" ]; then + rm -f "$BIN" if [ "$EXT" = "zip" ]; then unzip -o "$ARCHIVE" >/dev/null else diff --git a/.github/workflows/actions-test.yml b/.github/workflows/actions-test.yml index 1499dcf..a721181 100644 --- a/.github/workflows/actions-test.yml +++ b/.github/workflows/actions-test.yml @@ -207,12 +207,13 @@ jobs: # the last release. Otherwise every fixture here would be limited to # features that already shipped — which is the opposite of what a pull # request needs to prove. - - name: Build the convctl under test - id: build + # On PATH, exactly as setup-convctl would leave it — the Actions + # under test consume convctl from PATH rather than installing it. + - name: Build the convctl under test and put it on PATH run: | set -euo pipefail go build -o "$RUNNER_TEMP/convctl" ./cmd/convctl - echo "binary=$RUNNER_TEMP/convctl" >> "$GITHUB_OUTPUT" + echo "$RUNNER_TEMP" >> "$GITHUB_PATH" # A clean fixture passes and still produces a report. - name: Clean fixture @@ -223,7 +224,6 @@ jobs: xrd: internal/cli/testdata/full/xrd.yaml samples: internal/cli/testdata/full/samples artifact-name: actions-test-clean - binary: ${{ steps.build.outputs.binary }} # Outputs go through env rather than being interpolated into the # script: an output the Action never set would otherwise expand to @@ -305,7 +305,6 @@ jobs: samples: internal/cli/testdata/validate-output/samples validate-output: "true" artifact-name: actions-test-failing - binary: ${{ steps.build.outputs.binary }} - name: Assert it failed, and reported why shell: bash @@ -332,7 +331,6 @@ jobs: samples: internal/cli/testdata/full/samples annotate: "false" upload-artifact: "false" - binary: ${{ steps.build.outputs.binary }} - name: It still ran shell: bash @@ -356,12 +354,13 @@ jobs: go-version-file: go.mod cache: true - - name: Build the convctl under test - id: build + # On PATH, exactly as setup-convctl would leave it — the Actions + # under test consume convctl from PATH rather than installing it. + - name: Build the convctl under test and put it on PATH run: | set -euo pipefail go build -o "$RUNNER_TEMP/convctl" ./cmd/convctl - echo "binary=$RUNNER_TEMP/convctl" >> "$GITHUB_OUTPUT" + echo "$RUNNER_TEMP" >> "$GITHUB_PATH" # comment:false — this job proves the rendering and the exit-code # semantics; commenting needs a pull request and a writable token. @@ -374,7 +373,6 @@ jobs: examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml xrd: examples/crossplane-xr-multiversion/03-promote-v2/xrd.yaml comment: "false" - binary: ${{ steps.build.outputs.binary }} - name: A delta was found and did not fail the job shell: bash @@ -404,7 +402,6 @@ jobs: examples/field-rename/xrdconversionconfig.yaml xrd: examples/field-rename/xrd.yaml comment: "false" - binary: ${{ steps.build.outputs.binary }} - name: And say so rather than rendering nothing shell: bash diff --git a/.github/workflows/release.yml b/.github/workflows/release.yml index 8db69f1..84e11ef 100644 --- a/.github/workflows/release.yml +++ b/.github/workflows/release.yml @@ -25,6 +25,12 @@ env: CONVCTL_IMAGE: ghcr.io/terasky-oss/declarative-conversion-convctl CHART_REPO: ghcr.io/terasky-oss/charts +# Third-party Actions in this workflow are pinned to commit SHAs rather than +# tags, with the tag in a trailing comment. This is the only workflow that +# handles signing identities and publishing credentials — a retargeted tag +# here could exfiltrate the tap and bucket tokens or publish artifacts under +# this project's name, which is not true of a tag move in a test job. +# Dependabot updates SHA pins and keeps the comment current. jobs: test: name: Test before releasing @@ -58,24 +64,24 @@ jobs: uses: actions/checkout@v7 - name: Set up QEMU - uses: docker/setup-qemu-action@v4 + uses: docker/setup-qemu-action@99012661954931238ded8c8b007157a8430204e1 # v4 - name: Set up Docker Buildx - uses: docker/setup-buildx-action@v4 + uses: docker/setup-buildx-action@594f3bf4285d9ea8dc53c9a0c9c4092420091003 # v4 - name: Log in to GHCR - uses: docker/login-action@v4 + uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4 with: registry: ghcr.io username: ${{ github.actor }} password: ${{ secrets.GITHUB_TOKEN }} - name: Install cosign - uses: sigstore/cosign-installer@v3 + uses: sigstore/cosign-installer@398d4b0eeef1380460a10c8013a76f728fb906ac # v3 - name: Derive metadata id: meta - uses: docker/metadata-action@v6 + uses: docker/metadata-action@dc802804100637a589fabce1cb79ff13a1411302 # v6 with: images: ${{ matrix.image }} tags: | @@ -85,7 +91,7 @@ jobs: - name: Build and push id: build - uses: docker/build-push-action@v7 + uses: docker/build-push-action@c3c9e263c25d99ce0380d002d59b67737d91b0dc # v7 with: context: . platforms: linux/amd64,linux/arm64 @@ -109,7 +115,7 @@ jobs: # distroless base would block every release until upstream moves, # which trains people to bypass the gate rather than fix anything. - name: Scan image for vulnerabilities - uses: aquasecurity/trivy-action@v0.36.0 + uses: aquasecurity/trivy-action@ed142fd0673e97e23eac54620cfb913e5ce36c25 # v0.36.0 with: image-ref: ${{ matrix.image }}@${{ steps.build.outputs.digest }} format: table @@ -118,7 +124,7 @@ jobs: exit-code: "1" - name: Generate image SBOM - uses: anchore/sbom-action@v0 + uses: anchore/sbom-action@e22c389904149dbc22b58101806040fa8d37a610 # v0 with: image: ${{ matrix.image }}@${{ steps.build.outputs.digest }} format: spdx-json @@ -163,7 +169,7 @@ jobs: uses: actions/checkout@v7 - name: Set up Helm - uses: azure/setup-helm@v5 + uses: azure/setup-helm@9bc31f4ebc9c6b171d7bfbaa5d006ae7abdb4310 # v5 with: version: v3.21.3 @@ -183,7 +189,7 @@ jobs: # ~/.docker/config.json, which `helm registry login` alone does not # populate (it writes to helm's own credentials file instead), so we # also log in via docker/login-action. - uses: docker/login-action@v4 + uses: docker/login-action@dbcb813823bdd20940b903addbd779551569679f # v4 with: registry: ghcr.io username: ${{ github.actor }} @@ -193,7 +199,7 @@ jobs: run: echo "${{ secrets.GITHUB_TOKEN }}" | helm registry login ghcr.io -u ${{ github.actor }} --password-stdin - name: Install cosign - uses: sigstore/cosign-installer@v3 + uses: sigstore/cosign-installer@398d4b0eeef1380460a10c8013a76f728fb906ac # v3 - name: Package chart run: helm package charts/declarative-conversion-operator --destination /tmp/chart-out @@ -250,13 +256,13 @@ jobs: cache: true - name: Install cosign - uses: sigstore/cosign-installer@v3 + uses: sigstore/cosign-installer@398d4b0eeef1380460a10c8013a76f728fb906ac # v3 - name: Install syft - uses: anchore/sbom-action/download-syft@v0 + uses: anchore/sbom-action/download-syft@e22c389904149dbc22b58101806040fa8d37a610 # v0 - name: Run GoReleaser - uses: goreleaser/goreleaser-action@v7 + uses: goreleaser/goreleaser-action@f06c13b6b1a9625abc9e6e439d9c05a8f2190e94 # v7 with: distribution: goreleaser version: v2.17.1 @@ -337,7 +343,7 @@ jobs: } > addendum.md - name: Update release with signed-artifact details - uses: softprops/action-gh-release@v3 + uses: softprops/action-gh-release@efb35369e0ad2afab669f228072c1b0d510eae64 # v3 with: tag_name: ${{ github.ref_name }} append_body: true diff --git a/.goreleaser.yaml b/.goreleaser.yaml index 1d1799d..d2e8621 100644 --- a/.goreleaser.yaml +++ b/.goreleaser.yaml @@ -68,16 +68,13 @@ homebrew_casks: homepage: https://terasky-oss.github.io/declarative-conversion-operator/ description: CLI for the declarative conversion operator skip_upload: '{{ if index .Env "HOMEBREW_TAP_TOKEN" }}false{{ else }}true{{ end }}' - # An unsigned binary downloaded by a cask is quarantined by Gatekeeper, - # and the failure ("convctl is damaged") reads like a corrupt download - # rather than a policy. The archive's checksum is already verified by - # Homebrew itself. - hooks: - post: - install: | - if system_command("/usr/bin/xattr", args: ["-h"]).exit_status == 0 - system_command "/usr/bin/xattr", args: ["-dr", "com.apple.quarantine", "#{staged_path}/convctl"] - end + # No quarantine-stripping hook. Removing com.apple.quarantine would + # turn Gatekeeper off for this binary, and a checksum is not a + # substitute for code signing and notarization: it proves the file is + # the one the release published, not that anybody vouched for what it + # does. macOS will warn on first run until these artifacts are notarized, + # which is the honest state of affairs rather than one papered over + # during install. scoops: - name: convctl diff --git a/docs/gitops/configuration-ci.md b/docs/gitops/configuration-ci.md index dd546f4..3fdf9c5 100644 --- a/docs/gitops/configuration-ci.md +++ b/docs/gitops/configuration-ci.md @@ -33,13 +33,29 @@ jobs: with: persist-credentials: false - - name: Install crossplane CLI + # A pinned release, with its checksum verified. Piping install.sh from + # `main` into a shell executes whatever that branch holds at the + # moment the job runs, which makes the gate unreproducible and trusts + # a mutable reference with the runner's privileges. + - name: Install the crossplane CLI + env: + CROSSPLANE_VERSION: v2.0.2 run: | - curl -sL https://raw.githubusercontent.com/crossplane/crossplane/main/install.sh | sh + set -euo pipefail + url="https://releases.crossplane.io/stable/${CROSSPLANE_VERSION}/bin/linux_amd64/crank" + curl -fsSL -o crossplane "$url" + curl -fsSL -o crossplane.sha256 "${url}.sha256" + echo "$(cat crossplane.sha256) crossplane" | sha256sum -c - + chmod +x crossplane sudo mv crossplane /usr/local/bin/ - - name: Install convctl - run: go install github.com/terasky-oss/declarative-conversion-operator/cmd/convctl@latest + # setup-convctl installs a pinned release and verifies the cosign + # signature on its checksums. `go install ...@latest` does neither: it + # resolves to whatever is newest at the moment the job runs, and + # nothing checks what it got. + - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 + with: + version: v0.5.0 - name: Build the package run: crossplane xpkg build --package-root=./apis --package-file=platform.xpkg diff --git a/docs/gitops/convctl-fleet.gha.yml b/docs/gitops/convctl-fleet.gha.yml index 39028a7..0b612cb 100644 --- a/docs/gitops/convctl-fleet.gha.yml +++ b/docs/gitops/convctl-fleet.gha.yml @@ -43,6 +43,10 @@ jobs: with: persist-credentials: false + # Once per job. The Actions below consume convctl from PATH rather + # than each installing their own. + - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 + # The delta, as a sticky comment per cluster. Exit 1 (deltas found) is # a review artifact and does not fail the job; exit 2 (cannot reach the # cluster) does. @@ -79,6 +83,7 @@ jobs: - uses: actions/checkout@v7 with: persist-credentials: false + - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 - uses: terasky-oss/declarative-conversion-operator/.github/actions/convctl-fleet@v1 with: config: apis/widgets/conversion.yaml diff --git a/docs/gitops/fleet-ci.md b/docs/gitops/fleet-ci.md index dde90a4..6886740 100644 --- a/docs/gitops/fleet-ci.md +++ b/docs/gitops/fleet-ci.md @@ -175,9 +175,12 @@ jobs: run: | git fetch --depth=1 origin \ "+refs/heads/${{ github.base_ref }}:refs/remotes/origin/${{ github.base_ref }}" - - name: Install convctl - run: | - go install github.com/terasky-oss/declarative-conversion-operator/cmd/convctl@latest + # A pinned release with its signature verified, rather than + # `go install ...@latest`, which resolves to whatever is newest when + # the job runs and checks nothing about what it got. + - uses: terasky-oss/declarative-conversion-operator/.github/actions/setup-convctl@v1 + with: + version: v0.5.0 - run: | convctl compat \ --base "origin/${{ github.base_ref }}" --head HEAD \ diff --git a/docs/installation.md b/docs/installation.md index 23ecdc1..a48d36d 100644 --- a/docs/installation.md +++ b/docs/installation.md @@ -167,7 +167,7 @@ drops conversions while the schema and config reconcile independently. | rpm | `sudo rpm -i convctl__linux_amd64.rpm` | RHEL, Fedora, SUSE | | Archive | download `declarative-conversion-operator-cli___.tar.gz` from the [releases page](https://github.com/TeraSky-OSS/declarative-conversion-operator/releases) | all | | Container | see [below](#the-convctl-container-image) | linux/amd64, linux/arm64 | -| Source | `go install github.com/terasky-oss/declarative-conversion-operator/cmd/convctl@latest` | all | +| Source | `go install github.com/terasky-oss/declarative-conversion-operator/cmd/convctl@latest` | all — note that `@latest` resolves at install time and nothing verifies the result; prefer a release artifact in CI | Every archive's checksum is covered by the cosign-signed `checksums.txt`; see the signed-artifact section of any release for the verification commands. diff --git a/internal/cli/lint.go b/internal/cli/lint.go index 7bc902f..a1a7e01 100644 --- a/internal/cli/lint.go +++ b/internal/cli/lint.go @@ -312,6 +312,13 @@ func stagePackageXRDs(ref string) ([]discovered, func(), error) { if err != nil { return nil, nil, err } + return stagePackageXRDsFrom(pkg, ref) +} + +// stagePackageXRDsFrom is stagePackageXRDs with the package already read, +// so the staging rules can be tested against hostile contents without +// building a hostile package. +func stagePackageXRDsFrom(pkg *PackageContents, ref string) ([]discovered, func(), error) { if len(pkg.XRDs) == 0 { return nil, nil, fmt.Errorf("%s ships no XRDs to check configs against", ref) } @@ -321,13 +328,19 @@ func stagePackageXRDs(ref string) ([]discovered, func(), error) { } cleanup := func() { _ = os.RemoveAll(dir) } var out []discovered - for _, x := range pkg.XRDs { + for i, x := range pkg.XRDs { data, merr := sigsyaml.Marshal(x.Object) if merr != nil { cleanup() return nil, nil, fmt.Errorf("%s: re-encoding %s: %w", ref, xrdName(x), merr) } - path := filepath.Join(dir, xrdName(x)+".yaml") + // The name comes out of a package pulled from a registry or handed + // over by a third party, so it is untrusted: an XRD called + // ../../evil would otherwise have filepath.Join resolve outside the + // temporary directory and write wherever the invoking user can. The + // index keeps two XRDs of the same name from overwriting each + // other, which the base name alone would not. + path := filepath.Join(dir, fmt.Sprintf("%02d-%s.yaml", i, filepath.Base(xrdName(x)))) if werr := os.WriteFile(path, data, 0o600); werr != nil { cleanup() return nil, nil, werr diff --git a/internal/cli/report.go b/internal/cli/report.go index 7ba4543..39a76f7 100644 --- a/internal/cli/report.go +++ b/internal/cli/report.go @@ -333,10 +333,17 @@ func (r *Report) junitSuite() junitTestSuite { // with nothing saying which, is the exact false confidence the cap // exists to make explicit. if r.Meta.Sampling != nil { + // Not "equal to what we tested" when the walk stopped at the cap: + // the total was never counted, and a number there would be read as + // one. + population := strconv.Itoa(r.Meta.Sampling.Population) + if r.Meta.Sampling.Truncated { + population = "unknown" + } suite.Props = &junitProperties{Properties: []junitProperty{ {Name: "sampled", Value: "true"}, {Name: "sampleStrategy", Value: r.Meta.Sampling.Strategy}, - {Name: "samplePopulation", Value: strconv.Itoa(r.Meta.Sampling.Population)}, + {Name: "samplePopulation", Value: population}, {Name: "sampleTested", Value: strconv.Itoa(r.Meta.Sampling.Tested)}, }} } diff --git a/internal/cli/sampling.go b/internal/cli/sampling.go index 9865db8..fc31ab1 100644 --- a/internal/cli/sampling.go +++ b/internal/cli/sampling.go @@ -85,8 +85,13 @@ type SamplingReport struct { // Cap is the requested maximum. Cap int `json:"cap"` // Population is how many objects exist, counted while paginating even - // when most were never held. - Population int `json:"population"` + // when most were never held. Zero when the walk stopped early and the + // population was therefore never counted -- see Truncated. + Population int `json:"population,omitempty"` + // Truncated records that listing stopped at the cap, so the population + // is unknown rather than equal to what was tested. Only the first + // strategy can produce this. + Truncated bool `json:"truncated,omitempty"` // Tested is how many were actually sampled. Tested int `json:"tested"` // Seed is present for the random strategy, so the run is repeatable. @@ -101,6 +106,14 @@ func (s *SamplingReport) String() string { if s.Strategy == SampleRandom { seed = fmt.Sprintf(", seed %d", s.Seed) } + if s.Truncated { + // The population is genuinely unknown: listing stopped at the cap, + // which is what makes the first strategy the cheap one. Printing a + // population equal to the sample would claim the run was + // exhaustive, which is the one thing this line exists to deny. + return fmt.Sprintf("SAMPLED: the first %d live object(s), strategy %s%s — listing stopped at the cap, so the total is unknown and this run did NOT cover every object", + s.Tested, s.Strategy, seed) + } return fmt.Sprintf("SAMPLED: %d of %d live object(s), strategy %s%s — this run did NOT cover every object", s.Tested, s.Population, s.Strategy, seed) } @@ -112,7 +125,10 @@ type sampler struct { rnd *rand.Rand seen int - kept []Sample + // stoppedEarly records that listing was cut short at the cap, so seen + // is a floor on the population rather than the population. + stoppedEarly bool + kept []Sample // order holds each kept sample's creation timestamp for SampleNewest, // parallel to kept. order []string @@ -135,7 +151,11 @@ func newSampler(opts SamplingOptions) *sampler { // stop: the other two need to see the whole population to be what they // claim. func (s *sampler) full() bool { - return s.opts.enabled() && s.opts.strategy() == SampleFirst && len(s.kept) >= s.opts.MaxSamples + if s.opts.enabled() && s.opts.strategy() == SampleFirst && len(s.kept) >= s.opts.MaxSamples { + s.stoppedEarly = true + return true + } + return false } // add offers one object to the sample. @@ -211,14 +231,22 @@ func (s *sampler) result() ([]Sample, *SamplingReport) { } kept = sorted } - if !s.opts.enabled() || s.seen <= len(kept) { + // Stopping early is itself sampling, even though seen == kept: the + // listing never reached the end, so "we tested everything" is exactly + // what cannot be concluded. Reporting nothing here would have made the + // cheapest strategy the one that silently claims to be exhaustive. + if !s.opts.enabled() || (s.seen <= len(kept) && !s.stoppedEarly) { return kept, nil } - return kept, &SamplingReport{ - Strategy: s.opts.strategy(), - Cap: s.opts.MaxSamples, - Population: s.seen, - Tested: len(kept), - Seed: s.opts.Seed, + rep := &SamplingReport{ + Strategy: s.opts.strategy(), + Cap: s.opts.MaxSamples, + Tested: len(kept), + Seed: s.opts.Seed, + Truncated: s.stoppedEarly, + } + if !s.stoppedEarly { + rep.Population = s.seen } + return kept, rep } diff --git a/internal/cli/sampling_test.go b/internal/cli/sampling_test.go index bfe2529..4bf21bb 100644 --- a/internal/cli/sampling_test.go +++ b/internal/cli/sampling_test.go @@ -242,3 +242,88 @@ func TestReport_TableSaysARunWasSampled(t *testing.T) { t.Errorf("table does not report the sampling:\n%s", sb.String()) } } + +// The cheapest strategy stops listing at the cap, so seen == kept — and a +// report keyed only on that comparison said nothing, which meant the one +// strategy that cannot know the population was the one that silently +// claimed to have covered it. +func TestSampler_FirstReportsSamplingEvenThoughItStoppedEarly(t *testing.T) { + s := newSampler(SamplingOptions{MaxSamples: 5, Strategy: SampleFirst}) + for i := 0; i < 100; i++ { + if s.full() { + break + } + s.add(Sample{File: fmt.Sprintf("o%d", i)}, nil) + } + kept, rep := s.result() + if len(kept) != 5 { + t.Fatalf("kept %d, want the cap of 5", len(kept)) + } + if rep == nil { + t.Fatal("stopping early reported no sampling at all, so the run reads as exhaustive") + } + if !rep.Truncated { + t.Error("the report does not record that listing stopped early") + } + if rep.Population != 0 { + t.Errorf("Population = %d; it was never counted and must not be implied", rep.Population) + } + msg := rep.String() + if !strings.Contains(msg, "total is unknown") || !strings.Contains(msg, "did NOT cover every object") { + t.Errorf("the line does not say the total is unknown: %s", msg) + } + if strings.Contains(msg, "5 of 5") { + t.Errorf("the line implies the population equals the sample: %s", msg) + } +} + +// And the JUnit properties must not imply it either. +func TestReport_JUnitSaysThePopulationIsUnknownWhenTruncated(t *testing.T) { + rep := &Report{} + rep.Meta.Sampling = &SamplingReport{Strategy: SampleFirst, Cap: 5, Tested: 5, Truncated: true} + got := map[string]string{} + for _, p := range rep.junitSuite().Props.Properties { + got[p.Name] = p.Value + } + if got["samplePopulation"] != "unknown" { + t.Errorf("samplePopulation = %q, want unknown", got["samplePopulation"]) + } +} + +// A population that genuinely fit under the cap still reports nothing. +func TestSampler_FirstUnderTheCapIsNotSampling(t *testing.T) { + s := newSampler(SamplingOptions{MaxSamples: 50, Strategy: SampleFirst}) + feed(s, 10) + if _, rep := s.result(); rep != nil { + t.Errorf("a population that fit was reported as sampled: %+v", rep) + } +} + +// RunTest is exported, so the cobra command is not the only way in. Invalid +// sampling options that reach the sampler do the opposite of what they say: +// a negative cap disables the bound and paginates everything into memory, +// and an unknown strategy keeps nothing and then fails for having no +// samples. +func TestRunTest_ValidatesSamplingOptions(t *testing.T) { + base := TestOptions{ + XRDPath: "testdata/full/xrd.yaml", ConfigPath: "testdata/full/config.yaml", + SamplesDir: "testdata/full/samples", Quiet: true, + } + + bad := base + bad.Sampling = SamplingOptions{MaxSamples: -1} + if _, err := RunTest(bad); err == nil { + t.Error("a negative cap was accepted, which disables the bound entirely") + } + + bad = base + bad.Sampling = SamplingOptions{MaxSamples: 10, Strategy: "newestish"} + if _, err := RunTest(bad); err == nil { + t.Error("an unknown strategy was accepted") + } + + // A fixture run with no sampling options is unaffected. + if _, err := RunTest(base); err != nil { + t.Errorf("an ordinary run was rejected: %v", err) + } +} diff --git a/internal/cli/test.go b/internal/cli/test.go index 7d2849d..c5ba29d 100644 --- a/internal/cli/test.go +++ b/internal/cli/test.go @@ -176,6 +176,15 @@ func RunTest(opts TestOptions) (*Report, error) { if opts.Fuzz > 0 && opts.FuzzSeed == 0 { opts.FuzzSeed = time.Now().UnixNano() } + // Validated here rather than only in the cobra command: RunTest is + // exported, and a caller that bypasses the flag parsing would otherwise + // reach the sampler with options it rejects — a negative cap disables + // the bound entirely and paginates the whole population into memory, + // and an unknown strategy keeps nothing and then fails the run for + // having no samples. + if err := ValidateSamplingOptions(opts.Sampling); err != nil { + return nil, err + } kind, err := PeekConfigKind(opts.ConfigPath) if err != nil { return nil, err diff --git a/internal/cli/xpkg_test.go b/internal/cli/xpkg_test.go index f9bf354..1046130 100644 --- a/internal/cli/xpkg_test.go +++ b/internal/cli/xpkg_test.go @@ -240,3 +240,50 @@ func TestRunLint_PairsATreeAgainstAPackage(t *testing.T) { t.Errorf("schema should name the package, not a temporary path: %q", rep.Results[0].Schema) } } + +// A package is a file pulled from a registry or handed over by a third +// party, so the names inside it are untrusted. An XRD called ../../evil must +// not decide where this process writes. +func TestStagePackageXRDs_CannotEscapeTheStagingDirectory(t *testing.T) { + pkg, err := ReadPackage(testPackage) + if err != nil { + t.Fatal(err) + } + // Rename the XRDs to hostile values, keeping everything else real. + pkg.XRDs[0].SetName("../../../../tmp/convctl-escape") + pkg.XRDs[1].SetName("../../../../tmp/convctl-escape") + + staged, cleanup, err := stagePackageXRDsFrom(pkg, "hostile.xpkg") + if err != nil { + t.Fatalf("staging: %v", err) + } + defer cleanup() + + if len(staged) != 2 { + t.Fatalf("staged %d files, want both XRDs kept distinct despite the same name", len(staged)) + } + seen := map[string]bool{} + for _, d := range staged { + abs, err := filepath.Abs(d.path) + if err != nil { + t.Fatal(err) + } + if strings.Contains(abs, "/tmp/convctl-escape") { + t.Errorf("staged file escaped to %s", abs) + } + if !strings.Contains(abs, "convctl-package-") { + t.Errorf("staged file %s is outside the staging directory", abs) + } + if seen[abs] { + t.Errorf("two XRDs staged to the same path %s, so one overwrote the other", abs) + } + seen[abs] = true + if _, err := os.Stat(d.path); err != nil { + t.Errorf("staged file is not readable: %v", err) + } + } + if _, err := os.Stat("/tmp/convctl-escape.yaml"); err == nil { + t.Error("a file was written outside the staging directory") + _ = os.Remove("/tmp/convctl-escape.yaml") + } +} diff --git a/test/actions/actions_test.go b/test/actions/actions_test.go index f782da6..70b90c9 100644 --- a/test/actions/actions_test.go +++ b/test/actions/actions_test.go @@ -249,3 +249,26 @@ func TestSetupConvctl_HandlesEveryRunnerPlatform(t *testing.T) { t.Error("an unknown OS or architecture is not rejected") } } + +// The checksum covers the archive, not the extracted binary — so a cache +// entry holding an intact archive beside an altered binary would be used as +// it stands if extraction were skipped because the binary was already +// there. That is precisely the attack the verification exists to stop. +func TestSetupConvctl_ReExtractsWhenVerifying(t *testing.T) { + def := loadActions(t)["setup-convctl"] + var extract string + for _, s := range def.Runs.Steps { + if s.ID == "install" { + extract = s.Run + } + } + if extract == "" { + t.Fatal("no install step") + } + if !strings.Contains(extract, `if [ "$VERIFY" = "true" ] || [ ! -f "$BIN" ]`) { + t.Error("extraction is skipped when the binary already exists, so an unverified binary can be used") + } + if !strings.Contains(extract, "rm -f \"$BIN\"") { + t.Error("the existing binary is not removed before extraction") + } +} From 064029070d91c59fbb3514ec7cc2df59eb3bfc2b Mon Sep 17 00:00:00 2001 From: vrabbi Date: Wed, 16 Sep 2026 00:09:01 +0300 Subject: [PATCH 21/24] test(actions): diff two configs that share a hub MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The diff job paired a stage config with a mistakes/ fixture that targets a different hub, so the comparison could not analyze and exited 2. The Action treated that correctly — exit 2 is "could not run", and failing the job is the whole point of distinguishing it from exit 1 — but it meant the job tested the error path rather than the rendering it exists to check. Swapped for testdata/config.yaml and testdata/config-norename.yaml, which share a hub and differ by one rule: a removed FieldRename that takes spokeToHub from lossless to lossy and drops coverage on both sides. That is a delta worth rendering, and it exercises the rows the markdown writer actually has. Co-Authored-By: Claude Opus 5 (1M context) --- .github/workflows/actions-test.yml | 12 +++++++++--- 1 file changed, 9 insertions(+), 3 deletions(-) diff --git a/.github/workflows/actions-test.yml b/.github/workflows/actions-test.yml index a721181..7dd5d36 100644 --- a/.github/workflows/actions-test.yml +++ b/.github/workflows/actions-test.yml @@ -368,10 +368,16 @@ jobs: id: diff uses: ./.github/actions/convctl-diff with: + # Two configs for the SAME hub that differ in one rule. The + # mistakes/ fixtures cannot be used here: they target a different + # hub, so the comparison fails to analyze and exits 2 — which the + # Action correctly treats as "could not run" rather than as a + # delta, and which therefore tests the error path rather than the + # rendering this job is about. config: | - examples/crossplane-xr-multiversion/03-promote-v2/xrdconversionconfig.yaml - examples/crossplane-xr-multiversion/mistakes/02-missing-rename.yaml - xrd: examples/crossplane-xr-multiversion/03-promote-v2/xrd.yaml + internal/cli/testdata/config.yaml + internal/cli/testdata/config-norename.yaml + xrd: internal/cli/testdata/xrd.yaml comment: "false" - name: A delta was found and did not fail the job From abfdaae3d83c22a407645909d81e6a7becf36f16 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Wed, 16 Sep 2026 00:19:21 +0300 Subject: [PATCH 22/24] fix: expose a supplied binary as convctl, and scope sampling validation MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three review findings. The `binary` input put the supplied executable's directory on PATH, but every Action that follows invokes the literal command `convctl` — so a binary named convctl-linux-amd64, which is exactly what a release download or a build matrix produces, would be on PATH and still not found. It is now linked (or copied) into a directory of our own under the expected name, convctl.exe on Windows. That also stops a directory of unrelated executables being added to PATH as a side effect. Sampling validation now applies only to live runs. The fields are documented live-only and ignored for fixtures, so validating them unconditionally would newly reject callers of the exported RunTest who set them harmlessly — a restriction the finding that prompted the validation never asked for. The protection stays where it was needed: before any cluster work on a live run, which is also what lets the test assert it without a cluster. And the traversal test wrote its escape target to a fixed /tmp path, so it could fail on a pre-existing file, delete an unrelated one, or collide with a concurrent run. It uses t.TempDir() now. Co-Authored-By: Claude Opus 5 (1M context) --- .github/actions/setup-convctl/action.yml | 20 ++++++++++-- internal/cli/sampling_test.go | 41 ++++++++++++++++-------- internal/cli/test.go | 10 ++++-- internal/cli/xpkg_test.go | 16 +++++---- test/actions/actions_test.go | 26 +++++++++++++++ 5 files changed, 87 insertions(+), 26 deletions(-) diff --git a/.github/actions/setup-convctl/action.yml b/.github/actions/setup-convctl/action.yml index 3d9e5fc..d850c3b 100644 --- a/.github/actions/setup-convctl/action.yml +++ b/.github/actions/setup-convctl/action.yml @@ -67,11 +67,25 @@ runs: echo "::error::binary input points at $BINARY, which is not an executable file" >&2 exit 1 fi - dir="$(cd "$(dirname "$BINARY")" && pwd)" + # Linked into a directory of our own as "convctl", rather than + # putting the supplied binary's directory on PATH. Every Action that + # follows invokes the literal command `convctl`, so a binary named + # convctl-linux-amd64 — which is exactly what a download or a build + # matrix produces — would be on PATH and still not found. Linking + # also avoids adding a directory of unrelated executables to PATH. + dir="$RUNNER_TEMP/convctl-supplied" + mkdir -p "$dir" + name=convctl + case "$RUNNER_OS" in Windows) name=convctl.exe ;; esac + rm -f "$dir/$name" + ln -s "$(cd "$(dirname "$BINARY")" && pwd)/$(basename "$BINARY")" "$dir/$name" 2>/dev/null \ + || cp "$BINARY" "$dir/$name" + chmod +x "$dir/$name" + echo "$dir" >> "$GITHUB_PATH" { - echo "path=$dir/$(basename "$BINARY")" - echo "version=$("$BINARY" version | awk '{print $1}')" + echo "path=$dir/$name" + echo "version=$("$dir/$name" version | awk '{print $1}')" } >> "$GITHUB_OUTPUT" echo "using the supplied convctl at $BINARY" >> "$GITHUB_STEP_SUMMARY" diff --git a/internal/cli/sampling_test.go b/internal/cli/sampling_test.go index 4bf21bb..4c977f8 100644 --- a/internal/cli/sampling_test.go +++ b/internal/cli/sampling_test.go @@ -304,26 +304,39 @@ func TestSampler_FirstUnderTheCapIsNotSampling(t *testing.T) { // a negative cap disables the bound and paginates everything into memory, // and an unknown strategy keeps nothing and then fails for having no // samples. -func TestRunTest_ValidatesSamplingOptions(t *testing.T) { - base := TestOptions{ +func TestRunTest_ValidatesSamplingOptionsOnLiveRuns(t *testing.T) { + live := TestOptions{ XRDPath: "testdata/full/xrd.yaml", ConfigPath: "testdata/full/config.yaml", - SamplesDir: "testdata/full/samples", Quiet: true, + Live: true, Quiet: true, } - bad := base - bad.Sampling = SamplingOptions{MaxSamples: -1} - if _, err := RunTest(bad); err == nil { - t.Error("a negative cap was accepted, which disables the bound entirely") + // The validation has to happen before any cluster work, so these fail + // with the sampling error rather than with "no kubeconfig" — which is + // also what makes the assertion runnable without a cluster. + live.Sampling = SamplingOptions{MaxSamples: -1} + if _, err := RunTest(live); err == nil || !strings.Contains(err.Error(), "--max-samples") { + t.Errorf("a negative cap was not rejected before cluster work: %v", err) } - bad = base - bad.Sampling = SamplingOptions{MaxSamples: 10, Strategy: "newestish"} - if _, err := RunTest(bad); err == nil { - t.Error("an unknown strategy was accepted") + live.Sampling = SamplingOptions{MaxSamples: 10, Strategy: "newestish"} + if _, err := RunTest(live); err == nil || !strings.Contains(err.Error(), "--sample-strategy") { + t.Errorf("an unknown strategy was not rejected before cluster work: %v", err) } +} - // A fixture run with no sampling options is unaffected. - if _, err := RunTest(base); err != nil { - t.Errorf("an ordinary run was rejected: %v", err) +// Sampling is documented as live-only and ignored for fixtures, so a caller +// that sets it harmlessly on a fixture run must not start failing. +func TestRunTest_IgnoresSamplingOnFixtureRuns(t *testing.T) { + opts := TestOptions{ + XRDPath: "testdata/full/xrd.yaml", ConfigPath: "testdata/full/config.yaml", + SamplesDir: "testdata/full/samples", Quiet: true, + Sampling: SamplingOptions{Strategy: SampleRandom}, + } + rep, err := RunTest(opts) + if err != nil { + t.Fatalf("a fixture run was rejected for a live-only field: %v", err) + } + if rep.Meta.Sampling != nil { + t.Errorf("a fixture run reported sampling: %+v", rep.Meta.Sampling) } } diff --git a/internal/cli/test.go b/internal/cli/test.go index c5ba29d..710d30c 100644 --- a/internal/cli/test.go +++ b/internal/cli/test.go @@ -182,8 +182,14 @@ func RunTest(opts TestOptions) (*Report, error) { // the bound entirely and paginates the whole population into memory, // and an unknown strategy keeps nothing and then fails the run for // having no samples. - if err := ValidateSamplingOptions(opts.Sampling); err != nil { - return nil, err + // + // Only for a live run. Sampling is documented as live-only and is + // ignored for fixtures, so rejecting it there would newly break callers + // that set the field harmlessly. + if opts.Live { + if err := ValidateSamplingOptions(opts.Sampling); err != nil { + return nil, err + } } kind, err := PeekConfigKind(opts.ConfigPath) if err != nil { diff --git a/internal/cli/xpkg_test.go b/internal/cli/xpkg_test.go index 1046130..9acfd17 100644 --- a/internal/cli/xpkg_test.go +++ b/internal/cli/xpkg_test.go @@ -249,9 +249,12 @@ func TestStagePackageXRDs_CannotEscapeTheStagingDirectory(t *testing.T) { if err != nil { t.Fatal(err) } - // Rename the XRDs to hostile values, keeping everything else real. - pkg.XRDs[0].SetName("../../../../tmp/convctl-escape") - pkg.XRDs[1].SetName("../../../../tmp/convctl-escape") + // A target inside the test's own directory, so a failure cannot touch + // anything else on the machine and two runs cannot collide. + escape := filepath.Join(t.TempDir(), "escaped") + hostile := "../../../../" + strings.TrimPrefix(escape, "/") + pkg.XRDs[0].SetName(hostile) + pkg.XRDs[1].SetName(hostile) staged, cleanup, err := stagePackageXRDsFrom(pkg, "hostile.xpkg") if err != nil { @@ -268,7 +271,7 @@ func TestStagePackageXRDs_CannotEscapeTheStagingDirectory(t *testing.T) { if err != nil { t.Fatal(err) } - if strings.Contains(abs, "/tmp/convctl-escape") { + if strings.Contains(abs, escape) { t.Errorf("staged file escaped to %s", abs) } if !strings.Contains(abs, "convctl-package-") { @@ -282,8 +285,7 @@ func TestStagePackageXRDs_CannotEscapeTheStagingDirectory(t *testing.T) { t.Errorf("staged file is not readable: %v", err) } } - if _, err := os.Stat("/tmp/convctl-escape.yaml"); err == nil { - t.Error("a file was written outside the staging directory") - _ = os.Remove("/tmp/convctl-escape.yaml") + if _, err := os.Stat(escape + ".yaml"); err == nil { + t.Errorf("a file was written outside the staging directory, at %s.yaml", escape) } } diff --git a/test/actions/actions_test.go b/test/actions/actions_test.go index 70b90c9..36df460 100644 --- a/test/actions/actions_test.go +++ b/test/actions/actions_test.go @@ -272,3 +272,29 @@ func TestSetupConvctl_ReExtractsWhenVerifying(t *testing.T) { t.Error("the existing binary is not removed before extraction") } } + +// The Actions that follow invoke the literal command `convctl`, so a +// supplied binary named convctl-linux-amd64 — which is what a download or a +// build matrix produces — has to be linked under the expected name rather +// than having its directory put on PATH. +func TestSetupConvctl_NormalisesTheSuppliedBinaryName(t *testing.T) { + def := loadActions(t)["setup-convctl"] + var supplied string + for _, s := range def.Runs.Steps { + if s.ID == "supplied" { + supplied = s.Run + } + } + if supplied == "" { + t.Fatal("no step handles the binary input") + } + if !strings.Contains(supplied, "name=convctl") { + t.Error("the supplied binary is not exposed under the name the other Actions invoke") + } + if !strings.Contains(supplied, "ln -s") && !strings.Contains(supplied, "cp ") { + t.Error("the supplied binary is neither linked nor copied into place") + } + if !strings.Contains(supplied, "convctl.exe") { + t.Error("the Windows binary name is not handled") + } +} From 148db411b5d97343fab9d32f2c4a98851a91a337 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Wed, 16 Sep 2026 00:24:01 +0300 Subject: [PATCH 23/24] test(convctl): cover the early-stop sampling report at the live call site The finding this closes is about streamLiveSamples, not the sampler in isolation: the early return at the cap is what leaves seen == kept, and the fix lives one file away in sampling.go. A test of the sampler alone leaves the connection between them unasserted, which is how the two drift apart later. Two tests through the real code path, with a fake dynamic client: a population of 25 with a cap of 5 must report truncation and an unknown population, and a population of 3 under a cap of 50 must report nothing at all, because that walk did reach the end. Checked against the pre-fix code rather than assumed: with the truncation condition reverted, the first test fails with "a bounded run over a larger population reported no sampling, so it reads as exhaustive". Co-Authored-By: Claude Opus 5 (1M context) --- internal/cli/live_test.go | 73 +++++++++++++++++++++++++++++++++++++++ 1 file changed, 73 insertions(+) diff --git a/internal/cli/live_test.go b/internal/cli/live_test.go index 9fc014b..d89f60e 100644 --- a/internal/cli/live_test.go +++ b/internal/cli/live_test.go @@ -19,7 +19,9 @@ package cli import ( "context" "errors" + "fmt" "sort" + "strings" "testing" "k8s.io/apimachinery/pkg/apis/meta/v1/unstructured" @@ -178,3 +180,74 @@ func TestObjectLabel(t *testing.T) { t.Errorf("expected namespaced label 'default/a', got %q", got) } } + +// The finding this covers is about streamLiveSamples, not the sampler in +// isolation: the early return at the cap is what leaves seen == kept, and a +// report keyed on that comparison said nothing — so a bounded run over a +// larger population read as exhaustive in every output format. +func TestStreamLiveSamples_EarlyStopIsReportedAsSampling(t *testing.T) { + const population, cap = 25, 5 + + gvr := schema.GroupVersionResource{Group: "example.org", Version: "v1", Resource: "xthings"} + var objs []runtime.Object + for i := 0; i < population; i++ { + o := &unstructured.Unstructured{Object: map[string]any{ + "apiVersion": "example.org/v1", + "kind": "XThing", + "metadata": map[string]any{"name": fmt.Sprintf("thing-%02d", i)}, + }} + objs = append(objs, o) + } + dyn := dynamicfake.NewSimpleDynamicClientWithCustomListKinds(runtime.NewScheme(), + map[schema.GroupVersionResource]string{gvr: "XThingList"}, objs...) + + s := newSampler(SamplingOptions{MaxSamples: cap, Strategy: SampleFirst}) + if err := streamLiveSamples(context.Background(), dyn, gvr, "v1", "", s, "xthings.example.org", "composite"); err != nil { + t.Fatalf("streaming: %v", err) + } + samples, rep := s.result() + + if len(samples) != cap { + t.Fatalf("collected %d samples, want the cap of %d", len(samples), cap) + } + if rep == nil { + t.Fatal("a bounded run over a larger population reported no sampling, so it reads as exhaustive") + } + if !rep.Truncated { + t.Error("the early stop was not recorded as truncation") + } + if rep.Population != 0 { + t.Errorf("Population = %d, but listing stopped before counting it", rep.Population) + } + if !strings.Contains(rep.String(), "total is unknown") { + t.Errorf("the report line claims to know the population: %s", rep.String()) + } +} + +// And a population that fits under the cap is not reported as sampled: the +// walk reached the end, so "we tested everything" is true. +func TestStreamLiveSamples_UnderTheCapIsExhaustive(t *testing.T) { + gvr := schema.GroupVersionResource{Group: "example.org", Version: "v1", Resource: "xthings"} + var objs []runtime.Object + for i := 0; i < 3; i++ { + objs = append(objs, &unstructured.Unstructured{Object: map[string]any{ + "apiVersion": "example.org/v1", + "kind": "XThing", + "metadata": map[string]any{"name": fmt.Sprintf("thing-%d", i)}, + }}) + } + dyn := dynamicfake.NewSimpleDynamicClientWithCustomListKinds(runtime.NewScheme(), + map[schema.GroupVersionResource]string{gvr: "XThingList"}, objs...) + + s := newSampler(SamplingOptions{MaxSamples: 50, Strategy: SampleFirst}) + if err := streamLiveSamples(context.Background(), dyn, gvr, "v1", "", s, "", ""); err != nil { + t.Fatal(err) + } + samples, rep := s.result() + if len(samples) != 3 { + t.Fatalf("collected %d, want 3", len(samples)) + } + if rep != nil { + t.Errorf("an exhaustive run was reported as sampled: %+v", rep) + } +} From 0d9118b02928edb22f37179b773daa864dec9584 Mon Sep 17 00:00:00 2001 From: vrabbi Date: Wed, 16 Sep 2026 00:24:30 +0300 Subject: [PATCH 24/24] fix(test): do not shadow the cap builtin Caught by revive after the push rather than before it: I read the lint output as clean when it was one finding. Co-Authored-By: Claude Opus 5 (1M context) --- internal/cli/live_test.go | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/internal/cli/live_test.go b/internal/cli/live_test.go index d89f60e..d7c5f50 100644 --- a/internal/cli/live_test.go +++ b/internal/cli/live_test.go @@ -186,7 +186,7 @@ func TestObjectLabel(t *testing.T) { // report keyed on that comparison said nothing — so a bounded run over a // larger population read as exhaustive in every output format. func TestStreamLiveSamples_EarlyStopIsReportedAsSampling(t *testing.T) { - const population, cap = 25, 5 + const population, keep = 25, 5 gvr := schema.GroupVersionResource{Group: "example.org", Version: "v1", Resource: "xthings"} var objs []runtime.Object @@ -201,14 +201,14 @@ func TestStreamLiveSamples_EarlyStopIsReportedAsSampling(t *testing.T) { dyn := dynamicfake.NewSimpleDynamicClientWithCustomListKinds(runtime.NewScheme(), map[schema.GroupVersionResource]string{gvr: "XThingList"}, objs...) - s := newSampler(SamplingOptions{MaxSamples: cap, Strategy: SampleFirst}) + s := newSampler(SamplingOptions{MaxSamples: keep, Strategy: SampleFirst}) if err := streamLiveSamples(context.Background(), dyn, gvr, "v1", "", s, "xthings.example.org", "composite"); err != nil { t.Fatalf("streaming: %v", err) } samples, rep := s.result() - if len(samples) != cap { - t.Fatalf("collected %d samples, want the cap of %d", len(samples), cap) + if len(samples) != keep { + t.Fatalf("collected %d samples, want the cap of %d", len(samples), keep) } if rep == nil { t.Fatal("a bounded run over a larger population reported no sampling, so it reads as exhaustive")