diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index b00b9f4..38a974e 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -41,6 +41,29 @@ jobs: CGO_ENABLED=0 go build -o dist/devcloud ./cmd/devcloud CGO_ENABLED=0 go build -o dist/codegen ./cmd/codegen + # docs/coverage.md publishes the binary size, and the single-binary, + # zero-config property is what that figure is evidence for. Nothing + # checked it, so a regression would have been found by a reader rather + # than by CI. The startup half of the same claim is gated in + # cmd/devcloud/budget_test.go. + # + # 45 MiB is a regression ceiling, not the published figure: coverage.md + # measures 36.8 MiB on Apple Silicon and this builds for linux, so the two + # are not the same number and a tight budget would fail on the difference. + # The ceiling is named in coverage.md's runtime section, so a reader who + # hits this failure finds it where the 36.8 MiB is. + # + # stat -c%s is GNU stat: this job runs only on the ubuntu runners above. + - name: Check the binary stays within its regression ceiling + run: | + size=$(stat -c%s dist/devcloud) + limit=$((45 * 1024 * 1024)) + echo "dist/devcloud: $((size / 1024 / 1024)) MiB (ceiling $((limit / 1024 / 1024)) MiB)" + if [ "$size" -gt "$limit" ]; then + echo "::error::dist/devcloud exceeds the 45 MiB ceiling docs/coverage.md names, against a published 36.8 MiB. Raising the ceiling is a decision; first find what grew." + exit 1 + fi + # internal/generated is committed but derived. The Go tests check the fidelity # manifest's shape — floors, registered services, the CRUD registry — none of # which notice an operation a provider gained and the manifest never did. diff --git a/README.md b/README.md index 39ea069..c3f2f85 100644 --- a/README.md +++ b/README.md @@ -27,7 +27,7 @@ DevCloud is an **on-ramp to the cloud**, not a replacement for it. The goal is t ## Features -- **431 AWS services registered, 426 serving at least one operation** — the remaining 5 are routed and decline with a clean AWS error rather than letting the call bill a real account. See [coverage.md](docs/coverage.md) for what the numbers do and do not promise. +- **431 AWS services registered, 426 serving at least one operation** — every service AWS publishes is routed, so no SDK call escapes to a billable account; the remaining 5 decline with a clean AWS error. Depth is a separate, smaller promise — see [coverage.md](docs/coverage.md) for both targets. - **boto3-compatible** — a 1,530-test suite runs in CI (`make test-compat`) across every registered service. Unsupported operations return a clean AWS error, never a false success. - **Cross-service integration** — CloudFormation provisioning, DynamoDB Streams → Lambda, EventBridge targets, S3 → Lambda - **Smithy-driven codegen** — Go types, routers and error catalogues generated from AWS models, with a weekly sync workflow that keeps them current diff --git a/changes/unreleased/Added-20260913-170500.yaml b/changes/unreleased/Added-20260913-170500.yaml new file mode 100644 index 0000000..beac194 --- /dev/null +++ b/changes/unreleased/Added-20260913-170500.yaml @@ -0,0 +1,9 @@ +kind: Added +body: 'Two published figures are now gated against the binary that were not: the + routing and depth targets on the coverage page, and the size of the shipped + binary (45 MiB ceiling, checked in CI). A third gate is new but narrower than + intended — every one of the 431 registered services is now asserted to + initialize, which main.go only warned about, while its startup time is logged + rather than gated because a shared CI runner is 14x slower than the machine the + published figure comes from' +Issue: "163" diff --git a/changes/unreleased/Documentation-20260913-170000.yaml b/changes/unreleased/Documentation-20260913-170000.yaml new file mode 100644 index 0000000..e2672e3 --- /dev/null +++ b/changes/unreleased/Documentation-20260913-170000.yaml @@ -0,0 +1,6 @@ +kind: Documentation +body: 'The coverage page now states its two targets separately — routing, which is + 431 of 431 and is a leak-zero safety property, and depth, which stays at the 205 + services the 2026-09-05 demand study settled — so a service count can no longer + be read as a promise of fidelity' +Issue: "163" diff --git a/cmd/devcloud/budget_test.go b/cmd/devcloud/budget_test.go new file mode 100644 index 0000000..f868e81 --- /dev/null +++ b/cmd/devcloud/budget_test.go @@ -0,0 +1,94 @@ +// SPDX-License-Identifier: Apache-2.0 + +package main + +import ( + "context" + "os" + "path/filepath" + "testing" + "time" + + "github.com/skyoo2003/devcloud/internal/plugin" +) + +// startupBudgetEnv opts the timing half of TestRegisteredFleetComesUpWithinItsBudget +// into being an assertion. Unset — which is every CI run — the elapsed time is +// logged and nothing is asserted about it. See the test's comment for why. +// +// DEVCLOUD_STARTUP_BUDGET=300ms go test ./cmd/devcloud/ +const startupBudgetEnv = "DEVCLOUD_STARTUP_BUDGET" + +// TestRegisteredFleetComesUpWithinItsBudget asserts that every registered +// service actually comes up, and measures what that costs. +// +// The correctness half is the part that holds unconditionally. main.go brings +// the long tail up with fatal=false, so a service that is registered but can no +// longer initialize degrades to a warning in a log nobody reads. Here it fails. +// +// The timing half is a measurement, not a gate, and the reason is a number. The +// fleet comes up in ~141 ms locally and took 2.048 s on a GitHub arm64 runner: +// the shared runner is 14x slower, and that is before accounting for variance +// between runs on it. The regression actually worth catching — a provider that +// starts opening a file or a database per service at startup — is 1.2x to 2x, +// because that is what 431 extra file opens cost. There is no absolute ceiling +// that clears a 14x machine difference and still fails on a 2x regression, so an +// absolute ceiling in CI would only ever have been decoration that flakes. It is +// better to say startup is measured than to claim a gate that cannot fire. +// +// Set DEVCLOUD_STARTUP_BUDGET to assert on a machine whose speed you know — +// that is how the published figure in docs/coverage.md is re-taken per release. +// +// An earlier version of this test timed Construct instead of Init, which only +// calls the factory. Every factory in the tree is a struct literal, so it +// measured 84 µs and could not have failed for any reason. What costs is Init: +// S3Provider.Init does an os.MkdirAll and opens a SQLite database. +func TestRegisteredFleetComesUpWithinItsBudget(t *testing.T) { + ids := plugin.DefaultRegistry.RegisteredServices() + if len(ids) == 0 { + t.Fatal("no services are registered; imports.go is not linking the service packages") + } + + // A fresh registry rather than DefaultRegistry: Init records the instance as + // active, and leaving 431 initialized services behind would leak into every + // other test in this package. + fresh := plugin.NewRegistry() + for _, id := range ids { + fresh.Register(id, func() plugin.ServicePlugin { + p, ok := plugin.DefaultRegistry.Construct(id) + if !ok { + t.Errorf("%s is registered but its factory would not construct", id) + return nil + } + return p + }) + } + t.Cleanup(func() { _ = fresh.ShutdownAll(context.Background()) }) + + root := t.TempDir() + + start := time.Now() + for _, id := range ids { + if _, err := fresh.Init(id, plugin.PluginConfig{DataDir: filepath.Join(root, id)}); err != nil { + t.Errorf("%s is registered but failed to initialize: %v", id, err) + } + } + elapsed := time.Since(start) + + t.Logf("brought %d services up in %s", len(ids), elapsed.Round(time.Millisecond)) + + raw, ok := os.LookupEnv(startupBudgetEnv) + if !ok { + return + } + budget, err := time.ParseDuration(raw) + if err != nil { + t.Fatalf("%s=%q is not a duration: %v", startupBudgetEnv, raw, err) + } + if elapsed > budget { + t.Errorf("bringing %d services up took %s, over the %s asked for. Something in a "+ + "provider's Init is doing work it did not do before — that is what scales with "+ + "service count. If this is a slower machine than the one the budget was set on, "+ + "it is the budget that is wrong.", len(ids), elapsed.Round(time.Millisecond), budget) + } +} diff --git a/cmd/devcloud/coverage_test.go b/cmd/devcloud/coverage_test.go index 2d28ae5..795faca 100644 --- a/cmd/devcloud/coverage_test.go +++ b/cmd/devcloud/coverage_test.go @@ -241,6 +241,72 @@ func TestOtherDocsQuoteTheSameFigure(t *testing.T) { } } +// targetTableRow matches one row of the two-axis table in +// docs/coverage.md#the-target: +// +// | **Routing target** | **431 / 431 — met** | leak-zero; … | +// +// Only the first number in the value cell is captured. "431 / 431 — met" and +// "205 — met" are written for a reader; the gate reads the figure the reader +// sees rather than asking the table to be machine-shaped, which is the same +// trade tierRow makes. +func targetTableRow(t *testing.T, doc, label string) int { + t.Helper() + + pattern := regexp.MustCompile(`(?m)^\|\s*\*?\*?` + regexp.QuoteMeta(label) + + `\*?\*?\s*\|\s*\*?\*?([\d,]+)`) + matches := pattern.FindAllStringSubmatch(doc, -1) + if len(matches) != 1 { + t.Fatalf("docs/coverage.md: found %d target rows for %q, want exactly 1. "+ + "The two-axis table is what this gate reads; if it was restructured, "+ + "move the pattern deliberately rather than letting the target stop "+ + "being checked.", len(matches), label) + } + + n, err := strconv.Atoi(strings.ReplaceAll(matches[0][1], ",", "")) + if err != nil { + t.Fatalf("docs/coverage.md: target row %q has an unreadable number %q", label, matches[0][1]) + } + return n +} + +// TestPublishedTargetTableMatchesTheBinary gates the target itself, which is the +// half of this page nothing read until now. +// +// The summary table at the top has been gated since Milestone 6, and the target +// table under #the-target has not — so the page once reached a state where it +// published 431 registered and, further down, called the target "205 services, +// not 431" and described those 431 as not targeted. Every number there was wrong +// and every gate was green. +// +// Two of the three numbers are derivable from the binary and are checked against +// it. The serving target is a decision, not a measurement — it is read from the +// page and used as the arithmetic the third number must satisfy, so the table +// cannot be internally inconsistent either. +func TestPublishedTargetTableMatchesTheBinary(t *testing.T) { + raw, err := os.ReadFile(coveragePath) + if err != nil { + t.Fatalf("read the published coverage claim: %v", err) + } + doc := string(raw) + + registered := len(plugin.DefaultRegistry.RegisteredServices()) + + if got := targetTableRow(t, doc, "Routing target"); got != registered { + t.Errorf("docs/coverage.md publishes a routing target of %d services, the binary "+ + "registers %d. The routing target is every service AWS publishes, so these "+ + "move together or the leak-zero claim is no longer true.", got, registered) + } + + serving := targetTableRow(t, doc, "Serving target") + outside := targetTableRow(t, doc, "Registered and engine-served, outside the serving target") + if got, want := outside, registered-serving; got != want { + t.Errorf("docs/coverage.md publishes %d services outside the serving target; "+ + "%d registered minus a serving target of %d is %d. One of the three "+ + "numbers moved without the others.", got, registered, serving, want) + } +} + // demandPath is the evidence behind the published target. See docs/demand.md. const demandPath = "../../docs/demand.md" diff --git a/cmd/devcloud/sync_test.go b/cmd/devcloud/sync_test.go index 6092962..c4a6bda 100644 --- a/cmd/devcloud/sync_test.go +++ b/cmd/devcloud/sync_test.go @@ -188,9 +188,10 @@ func TestSyncPullRequestBodyReportsTheTestResult(t *testing.T) { // TestSyncPullRequestBodySummarisesTheChurn is the guarantee that makes the PR // reviewable rather than merely present. // -// The sync regenerates from all 194 models at once, so its diff is whole-tree -// whether upstream moved one model or ninety — measured on 2026-09-06, a real -// refresh changed 93 models and 134 generated files. "Please review" over that +// The sync regenerates from all 420 models at once, so its diff is whole-tree +// whether upstream moved one model or ninety — measured on 2026-09-06, when 194 +// were vendored, a real refresh changed 93 models and 134 generated files, and +// the vendored set has more than doubled since. "Please review" over that // is not an action anyone performs. scripts/model_churn.py reduces it to the // question a reviewer has, which operations moved, and the body must carry that // answer or the reviewer is back to reading the regeneration. diff --git a/docs/README.md b/docs/README.md index 70a9c48..472e07c 100644 --- a/docs/README.md +++ b/docs/README.md @@ -13,7 +13,7 @@ | Page | What it covers | |---|---| -| [Coverage](coverage.md) | 431 registered / 426 serving — what the counts promise, and the target | +| [Coverage](coverage.md) | 431 registered / 426 serving — the routing and depth targets, and what each promises | | [Compatibility Policy](compatibility-policy.md) | What v1.0 guarantees across 1.x, what it does not, and how deprecation works | | [Fidelity Manifest](fidelity-manifest.md) | Per-operation tiers: how much to trust any given call | | [CRUD Engine](crud-engine.md) | How engine-served operations behave, and where they stop | diff --git a/docs/contributing.md b/docs/contributing.md index 628cb3f..23f0f81 100644 --- a/docs/contributing.md +++ b/docs/contributing.md @@ -89,11 +89,12 @@ serializer or deserializer. See [architecture.md](architecture.md#code-generatio ### Reviewing the weekly model sync -[`smithy-sync.yml`](../.github/workflows/smithy-sync.yml) refreshes all 194 +[`smithy-sync.yml`](../.github/workflows/smithy-sync.yml) refreshes all 420 vendored models every Monday and opens a pull request. The diff is whole-tree — a -measured refresh moved 93 models and 134 generated files — so do not try to read -it. Read the PR body instead: it is generated by `scripts/model_churn.py` and -lists which services gained or lost operations. +refresh measured at 194 models moved 93 of them and 134 generated files, and the +vendored set has more than doubled since — so do not try to read it. Read the PR +body instead: it is generated by `scripts/model_churn.py` and lists which +services gained or lost operations. Three things to check, in order: diff --git a/docs/coverage.md b/docs/coverage.md index 165674c..93d88d5 100644 --- a/docs/coverage.md +++ b/docs/coverage.md @@ -20,10 +20,11 @@ Per operation, from the [fidelity manifest](fidelity-manifest.md): | `unimplemented` | 3,802 | | **total known** | **19,201** | -> **The target is depth for 205 services, not breadth for 431.** All 431 are -> registered — the codegen scaffold made breadth nearly free — but registration is -> not the promise. What the evidence refused was a promise of *depth* across 431, -> and it still refuses it. See [The target](#the-target). +> **Two targets, not one: routing is 431 of 431, depth is 205.** Every service +> AWS publishes is registered, so no call can leave for a billable account — that +> is a safety property and it admits no smaller number. Depth is the separate, +> smaller promise, and the evidence still puts it at 205. See +> [The target](#the-target). Every figure on this page is asserted against the binary by `go test ./cmd/devcloud/`. Editing one here without the code moving fails CI, and @@ -130,7 +131,25 @@ not packaging artefacts. ## The target -**Decided 2026-09-05. The target is 205 services, not 431 — and it is met.** +DevCloud publishes **two** targets. They answer different questions, and reading +one as the other is the mistake this page exists to prevent. + +- **Routing — every service AWS publishes, and it is met.** A registered service + is answered at `localhost:4747`; an unregistered one is not routed, so the SDK + call leaves the machine and bills a real AWS account. That is a safety + property, not a capability claim, and the only number that satisfies it is all + of them. +- **Serving depth — 205 services, decided 2026-09-05, and it is met.** Depth is + what costs, so it follows evidence of demand rather than the shape of AWS's + catalogue. The study below is that evidence, and it is unchanged. + +| Axis | Services | Governed by | +|---|---|---| +| **Routing target** | **431 / 431 — met** | leak-zero; every published model is registered | +| **Serving target** | **205 — met** | the demand study below, sampled 2026-09-05 | +| Registered and engine-served, outside the serving target | 226 | no depth promise — see the [CRUD engine](crud-engine.md) | + +### How the depth target was set The old target was every service AWS publishes. It rested on an assumption nobody had tested: that the services DevCloud does not register are services anyone @@ -138,16 +157,14 @@ wants. Before committing to building ~283 of them, the assumption was tested against three independent projects that each only add a service when someone asks. It did not hold. -**All 431 are now registered, and the decision above still stands.** What it +**Registering all 431 later did not overturn that decision.** What the study refused was the *cost* — hand-building 283 services on the assumption someone wanted them. The codegen scaffold removed that cost: registering the remaining 226 became a flag on `make codegen`, not a programme of work, and the services it reached are served by the generic [CRUD engine](crud-engine.md) at engine -fidelity. So breadth was taken because it turned out to be nearly free, and the -target stayed where the evidence put it, because the target was never a count of -registrations — it is where DevCloud promises to be worth trusting. Read the -table below as two different claims, not one: 431 services answer locally instead -of billing a real account, and 205 are the ones whose depth is a commitment. +fidelity. So routing was taken because it turned out to be nearly free, and the +depth target stayed where the evidence put it, because it was never a count of +registrations — it is where DevCloud promises to be worth trusting. **The rule was fixed before the numbers were seen** — four outcomes written down in advance, including one for "the method itself failed", specifically so the @@ -164,19 +181,14 @@ evidence in [demand.md](demand.md); re-derive with | `M` built by none | 115 | | DevCloud's own service requests, all time | **0** | -The rule kept the 100% target only if ≥60% of `M` had support ≥2, and narrowed to -a demand set if ≥100 did. 57 cleared neither bar, so the pre-registered -consequence applied: **the 100% claim is dropped and the published target becomes -the demand set.** All 57 are registered — 56 serve at least one operation, and -`rds-data` is the exception named above. It is supported by all three projects, -the strongest signal in the set, and still cannot be served generically. Breadth -does not reach every service, and saying so is cheaper than a fabricated success. - -| | Services | -|---|---| -| Registered today | **431** | -| Target: registered + demonstrated demand | **205 — met** | -| Registered, scaffold-served, outside the target | 226 | +The rule kept the 100% depth target only if ≥60% of `M` had support ≥2, and +narrowed to a demand set if ≥100 did. 57 cleared neither bar, so the +pre-registered consequence applied: **the 100% depth claim is dropped and the +published depth target becomes the demand set.** All 57 are registered — 56 serve +at least one operation, and `rds-data` is the exception named above. It is +supported by all three projects, the strongest signal in the set, and still +cannot be served generically. The engine does not reach every service it routes, +and saying so is cheaper than a fabricated success. Four fifths of the AWS surface is surface that three projects with far more history and staffing have collectively declined to build. That is what a long @@ -236,32 +248,72 @@ anywhere. Apple Silicon, `CGO_ENABLED=0`, measured at each step of the roadmap: -| | 105 services | 147 services | 205 services | -|---|---|---|---| -| Binary | 30.8 MiB | 31.3 MiB | **33.1 MiB** | -| Peak RSS (`/usr/bin/time -l`) | 57.3 MiB | 57.7 MiB | **43.7 MiB** | -| Service registration | — | — | **49 ms for all 205** | +| | 105 services | 147 services | 205 services | 431 services | +|---|---|---|---|---| +| Binary | 30.8 MiB | 31.3 MiB | 33.1 MiB | **36.8 MiB** | +| Peak RSS (`/usr/bin/time -l`) | 57.3 MiB | 57.7 MiB | 43.7 MiB | **51.2 MiB** | +| Service startup | — | — | 49 ms for all 205 | **42 ms for all 431** | + +The single-binary, zero-config property holds at 431 with room to spare. Startup +does not scale meaningfully with service count: registration is a map insert per +service in `init()`, and the generated type definitions are mostly +dead-code-eliminated by the linker, which is why 58 more services cost 1.8 MiB — +and why 226 more, nearly quadrupling the fleet, cost 3.7 MiB rather than the +7 MiB a linear reading of that figure predicts. + +Two ceilings on the startup reading: + +1. **It is timed from the first `service initialized` line to `DevCloud ready`**, + so it covers bringing every service up, not the `init()` registration alone. + That is the number an operator waits for. +2. **A first run costs about three times as much** — 118 to 148 ms across four + measurements, against 42 ms once the data directories exist, because each of + the 431 services creates its own on the way up. The cold path is the one to + watch, and it is the one `cmd/devcloud/budget_test.go` reproduces. + +**Startup is measured, not gated, and that is a deliberate retreat.** The gate +was written as a wall-clock ceiling and CI disproved it. The same path that takes +141 ms here took 2.048 s on a GitHub arm64 runner, and 1.414 s on the same runner +type one commit later — so the shared runner is roughly ten times slower *and* +swings 45% between runs on identical code. The regression worth catching is +smaller than that swing: a provider that starts opening a file or a database per +service at startup costs 1.2x to 2x, because that is what 431 extra file opens +are worth. A ceiling that survives the variance cannot fail on the regression. +Saying the figure is measured is cheaper than claiming a gate that cannot fire. + +What `budget_test.go` *does* assert on every run is that all 431 services come +up at all. `main.go` initializes the long tail non-fatally, so a service that is +registered but can no longer initialize would otherwise degrade to a warning in a +log nobody reads. The timing is logged beside it, and becomes an assertion when +you name a budget on a machine whose speed you know: -The single-binary, zero-config property holds at the target with room to spare. -Startup does not scale meaningfully with service count: registration is a map -insert per service in `init()`, and the generated type definitions are mostly -dead-code-eliminated by the linker, which is why 58 more services cost 1.8 MiB. +``` +DEVCLOUD_STARTUP_BUDGET=300ms go test ./cmd/devcloud/ +``` + +That is how the figures above are re-taken per release. + +The binary size *is* gated, because size does not vary with how busy a runner is. +CI fails the build past **45 MiB** and measures 35 MiB on arm64, 37 MiB on amd64 +— close enough to the 36.8 MiB above that the platform barely registers, so the +remaining 8 MiB is genuine slack rather than a correction for anything. It is a +decision; if it starts failing, find what grew before raising it. The RSS readings were taken on different days and are not a controlled -comparison. Read them as "memory is not the constraint at 205" rather than as a -saving — what they agree on is the shape: memory is dominated by the runtime and -the store, not by how many services are registered. +comparison. Read them as "memory is not the constraint" rather than as a saving — +what they agree on is the shape: memory is dominated by the runtime and the +store, not by how many services are registered. ## Keeping up with upstream -The 205 services are vendored from 194 Smithy models, and AWS keeps changing +The 431 services are vendored from 420 Smithy models, and AWS keeps changing them. A [weekly workflow](../.github/workflows/smithy-sync.yml) refreshes all of them and opens a pull request. What that review costs was measured once, on **2026-09-06**: | Reading | Value | |---|---| -| Vendored models refreshed | 194 | +| Vendored models refreshed | 194 (of 420 vendored today) | | Models that changed | 93 | | Of those, models that added or removed an operation | 32 | | Of those, models that changed only documentation | 0 | @@ -274,6 +326,13 @@ changed were all vendored 141 days earlier; the other 101 were vendored the day before, and not one of them changed. So the reading is an accumulated backlog, and the weekly rate is still unknown. +**And it was taken at 194 models, not 420.** The vendored set has since more than +doubled, so the diff a reviewer faces is larger than anything measured here by +roughly the same factor, and the operation counts in the table predate the 6,763 +operations the long tail brought with it. What does not change is the shape of +the review: it is driven by which operations moved, not by how many files did. +Reducing that cost rather than restating it is tracked separately. + What the sample does settle is the *shape* of the work. None of the 93 was documentation-only, so no sync can be waved through on the assumption that AWS only reworded things. Thirty-two services gained operations — `ec2` alone gained diff --git a/docs/demand.md b/docs/demand.md index ee43f43..e8e139d 100644 --- a/docs/demand.md +++ b/docs/demand.md @@ -48,7 +48,10 @@ at any point would have stopped on the best surface available. `workspaces-web`; `TestDemandSetIsRegistered` fails if any of them stops being. 56 serve at least one operation, and the one that does not (`rds-data`) is named with its reason in [coverage.md](coverage.md). Everything below them — support 1 -and support 0 — is the 226 services that remain explicitly not targeted. +and support 0 — is the 226 services that carry no depth commitment. They are +registered and engine-served, so a call to one is answered locally rather than +billed; what this study declined to promise is that it is answered *faithfully*. +See [the two axes](coverage.md#the-target). | Service | moto | LocalStack | terraform-provider-aws | Support | |---|---|---|---|---|