Repository navigation
feat: gate the published coverage figure, and test every registered service - #144
Merged
Merged
Conversation
GREEN: go test ./internal/codegen/ -run CompatManifest passes; make codegen emits internal/generated/compat/services.json with 205 services, 201 serving at least one operation and 4 serving none — the same three numbers docs/coverage.md publishes. The compatibility suite carried two hand-written service lists, so registering a service and testing it were separate steps and nothing failed when only the first happened. 31 of 205 registered services are exercised by no boto3 test. Deriving the list removes the second step.
GREEN: 992 boto3 tests pass in 14.2s, up from 854 in 10s. The smoke suite now parametrizes over all 205 registered services from the generated manifest, so the 31 that were registered, counted in the published 201, and exercised by no boto3 test at all are now covered. Findings the gate surfaced, each pinned rather than absorbed: - The four Lex services are registered and unreachable by any boto3 caller. All four clients sign as the contested alias 'lex', which BuildAliases leaves unrouted by design. docs/coverage.md claimed each was reachable by its own name; a boto3 caller has no such name. - appsync/ListApis, eks/ListAccessPolicies, opensearch/ListApplications are labelled hand-verified but no route reaches them, so the floor is asserted per service rather than per operation. ENGINE_SERVED_SERVICES and REGISTERED_ONLY_SERVICES are deleted.
Milestone 6. The figure was published in Milestone 3 and nothing checked it: docs/coverage.md stated 205/201/4 as prose, and the floors that could catch a regression were minServices=100 and minOperations=6000, so 205 collapsing to 120 passed CI in silence. The gate reads the published claim and compares it to the live registry and manifest in both directions. Proven by mutation, four ways: removing a service fails, inflating the doc fails, reverting README.md to its stale figure fails, and deleting the figure from README.md fails rather than silently disabling the check. Also gated: the per-operation tier table, the names of the registered-only services, and demand-set membership. Drift the gate found on its first run: - README.md and docs/README.md had quoted 148 registered / 117 serving since three milestones earlier - docs/demand.md still named three demand-set services as serving nothing; Milestone 5 left one - sagemakerruntimehttp2 serves nothing and was never named in coverage.md The conservative floors in fidelity_test.go and conformance_test.go are replaced rather than kept alongside: a weaker second answer to a question that now has an exact one.
skyoo2003
added a commit
that referenced
this pull request
Sep 5, 2026
…ions GREEN: 126 passed, 17 skipped, 2 xfailed. The six failures this change targets are gone; the two that remain are different defects, marked xfail(strict) with what was observed rather than a guessed cause. The default branch of each provider returned 200 with an empty <ActionResponse/> for any action it did not implement. Two consequences, both bad: the caller was told an operation succeeded when nothing ran, and the body carried no <ActionResult> wrapper, so botocore's query parser raised KeyError instead of returning a value. neptune.describe_db_cluster_endpoints() crashed the SDK outright. Each now returns plugin.ErrUnhandledOp, which is what every other provider in the fleet does and what the CRUD engine has been reachable through since query was admitted in #143. 223 operations move from unimplemented to auto-crud — they are served from the real store now — and the rest decline with InvalidAction. docs/coverage.md's tier table moves with them: auto-crud 4,970 -> 5,193, unimplemented 2,941 -> 2,718, total unchanged. That edit was not optional; the gate added in #144 failed until the published figures matched the manifest, naming both numbers.
5 tasks done
skyoo2003
added a commit
that referenced
this pull request
Sep 5, 2026
…ions (#145) * test: add reproducer for fabricated successes on unserved operations RED: 8 failed, 120 passed, 17 skipped. FAILED [autoscaling.DescribeAdjustmentTypes] FAILED [cloudformation.DescribeChangeSetHooks] FAILED [elasticache.DescribeCacheEngineVersions] FAILED [elasticloadbalancingv2.DescribeCapacityReservation] FAILED [rds.DescribeBlueGreenDeployments] FAILED [redshift.DescribeAccountAttributes] FAILED [resourcegroups.Tag] FAILED [s3.ListBucketAnalyticsConfigurations] Each returned HTTP 200 for an operation the fidelity manifest does not list as served. docs/coverage.md calls that the one thing DevCloud must never do. test_service_smoke.py already asserted this for services that serve nothing at all. The gap was the other case: a service that serves plenty and answers an operation it does not implement with an empty success anyway. * fix: eight query providers fabricated a success for unimplemented actions GREEN: 126 passed, 17 skipped, 2 xfailed. The six failures this change targets are gone; the two that remain are different defects, marked xfail(strict) with what was observed rather than a guessed cause. The default branch of each provider returned 200 with an empty <ActionResponse/> for any action it did not implement. Two consequences, both bad: the caller was told an operation succeeded when nothing ran, and the body carried no <ActionResult> wrapper, so botocore's query parser raised KeyError instead of returning a value. neptune.describe_db_cluster_endpoints() crashed the SDK outright. Each now returns plugin.ErrUnhandledOp, which is what every other provider in the fleet does and what the CRUD engine has been reachable through since query was admitted in #143. 223 operations move from unimplemented to auto-crud — they are served from the real store now — and the rest decline with InvalidAction. docs/coverage.md's tier table moves with them: auto-crud 4,970 -> 5,193, unimplemented 2,941 -> 2,718, total unchanged. That edit was not optional; the gate added in #144 failed until the published figures matched the manifest, naming both numbers.
5 tasks done
skyoo2003
added a commit
that referenced
this pull request
Sep 6, 2026
* test: add reproducer for the four unreachable Lex services
RED: CGO_ENABLED=0 go test ./internal/gateway/ -run TestDetectProtocol_Lex
--- FAIL: TestDetectProtocol_Lex/models_v1_get_bots
--- FAIL: TestDetectProtocol_Lex/models_v1_get_intents
--- FAIL: TestDetectProtocol_Lex/models_v1_builtin_intents
--- FAIL: TestDetectProtocol_Lex/models_v2_list_bots
--- FAIL: TestDetectProtocol_Lex/models_v2_create_bot
--- FAIL: TestDetectProtocol_Lex/models_v2_describe_bot
--- FAIL: TestDetectProtocol_Lex/runtime_v1_get_session
--- FAIL: TestDetectProtocol_Lex/runtime_v1_put_session
--- FAIL: TestDetectProtocol_Lex/runtime_v2_get_session
--- FAIL: TestDetectProtocol_Lex/runtime_v2_delete_session
--- PASS: TestDetectProtocol_Lex/delete_bot_is_contested
Every failure reads `expected: <a lex service>, actual: "lex"` — the request
never leaves the contested alias. The one PASS is the genuine collision
(DELETE /bots/{id}, claimed by both model services), which must stay refused.
Recorded in .claude/tdd/aws-service-coverage-100-milestone-6.tdd.md as
"Coverage and known gaps", item 3.
* fix: the four Lex services were reachable by no boto3 caller
GREEN: CGO_ENABLED=0 go test ./internal/gateway/ -run TestDetectProtocol_Lex → ok
CGO_ENABLED=0 go test ./... → all pass
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1123 passed, 17 skipped, 2 xfailed (was 1119 passed)
All four Lex clients sign as "lex" and no service is called "lex", so the alias
stays contested. normalizeServiceID hands the unresolved name through, the
registry has nothing under it, and every Lex call died as UnknownService — four
services registered, counted in "serving >=1 operation", and reachable by nobody.
The fix needed no new routing. resolveSharedSigningName already picks the
sibling whose route table models the request; it was never reached, because
signingNameOf only recognised a *member* service ID and "lex" is the group's
key. One map lookup.
The four route tables separate cleanly on method and literal segment:
GET /bots lexmodelbuildingservice
POST /bots lexmodelsv2 (ListBots)
PUT /bots lexmodelsv2 (CreateBot)
GET /bot/{n}/alias/{a}/user/{u}/session lexruntimeservice
GET /bots/{b}/botAliases/.../sessions/{s} lexruntimev2
One genuine collision, DeleteBot at DELETE /bots/{id}, claimed by both model
services. It stays refused: deleting the wrong bot is worse than an honest
UnknownService. Asserted, not left to be found.
The pin test told whoever fixed this to delete it, so it is deleted and replaced
by a check on the direction that still matters — UNREACHABLE_FROM_BOTO3 must
stay empty. docs/coverage.md's compatibility-tested figure moves 199 -> 203, the
Lex exclusion section is rewritten, and the stale claim that each Lex service is
"reached by its own unambiguous name" is corrected: no boto3 caller has one.
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 1
* test: add reproducer for hand-verified operations no route reaches
RED: DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/test_handverified_is_reachable.py
FAILED [appsync.ListApis]
FAILED [eks.ListAccessPolicies]
FAILED [opensearch.ListApplications]
3 failed in 1.56s
Each answers NotImplemented. All three are labelled hand-verified in the
fidelity manifest, and all three are implemented — codegen reads the case
clause, and the case clause is there. What is missing is the route: the
gateway passes no operation name for rest-json, so the provider recovers it
from method and path with a hand-written resolver, and none of these three
resolvers knows the path its own model publishes.
appsync/ListApis GET /v2/apis resolveOp strips only /v1
eks/ListAccessPolicies GET /access-policies no case at all
opensearch/ListApplications
GET /2021-01-01/opensearch/list-applications
resolveOp knows only GET /application
Recorded in .claude/tdd/aws-service-coverage-100-milestone-6.tdd.md as
"Coverage and known gaps", item 2.
* fix: three hand-verified operations were reached by no route
GREEN: DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/test_handverified_is_reachable.py
→ 3 passed
CGO_ENABLED=0 go test ./... → all pass
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1126 passed, 17 skipped, 2 xfailed (was 1123)
make codegen; git status --porcelain internal/generated → clean
golangci-lint run → 0 issues
appsync/ListApis, eks/ListAccessPolicies and opensearch/ListApplications are
each implemented by a case clause the provider already has. Codegen reads that
clause and labels the operation hand-verified — which proves the code exists,
not that any request arrives at it. For rest-json those are different
questions: the gateway passes no operation name, so the provider recovers it
from method and path with a hand-written resolver, and none of these three
resolvers knew the route its own model publishes.
The resolvers are partial copies of a table codegen already emits. Each
provider now ends with crud.Route(service, method, uri) instead of losing the
request. The fallback fires only where the resolver produced nothing, so it
cannot displace a decision a provider made — it fills gaps.
crud.Route is HasRoute's implementation, exported; HasRoute becomes its
one-line caller. Nothing else in the gateway changes.
No published figure moves. These operations were already counted as served;
what changes is that the count is now true about them.
27 other providers hand-roll the same kind of resolver. Whether any of them has
the same gap is not measured here — the compat manifest carries served
operations but not their tier, so the fleet-wide version of this test needs a
codegen change first. Recorded, not guessed at.
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 2
* fix: the manifest reported an implemented operation as unimplemented
RED (compile-time, the intended missing field):
CGO_ENABLED=0 go test ./internal/codegen/
internal/codegen/gen_fidelity_test.go:83:4: unknown field ShortOperations
in struct literal of type ProviderScan
internal/codegen/scan_handverified_test.go:93:45:
scans["resourcegroups"].ShortOperations undefined
FAIL github.com/skyoo2003/devcloud/internal/codegen [build failed]
The RED could not be committed on its own: .pre-commit-config.yaml runs go vet,
which rejects a non-compiling tree, so RED and GREEN share this commit and the
output above is the evidence. Same constraint as #144.
GREEN: CGO_ENABLED=0 go test ./internal/codegen/ → ok
CGO_ENABLED=0 go test ./... → all pass
make codegen → exactly one tier change
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1126 passed, 17 skipped, 1 xfailed (was 2 xfailed)
golangci-lint run → 0 issues
opNamePattern requires four characters. resourcegroups implements Tag and Untag
in the same switch; Untag was scanned, Tag was not, and the manifest called
implemented code unimplemented. test_no_fabricated_success then asked for Tag,
got the 200 the provider has always returned, and reported a fabricated success.
It was never fabricated. The answer was real and the manifest was wrong, so the
probe was right to complain and wrong about what it had found — which is why
that entry recorded only an observation and refused to name a mechanism.
Loosening the pattern is not the fix. At two characters it admits the HTTP verbs
every path resolver switches on. The model separates them: Tag is an operation
because resourcegroups.json says so, and no model AWS publishes declares GET.
Short literals are collected into ProviderScan.ShortOperations and promoted by
BuildFidelityData only where the service's model declares them.
Fleet-wide, exactly one operation is short enough to be affected:
rg -o '"[A-Z][A-Za-z0-9]{2}":\s+Tier[A-Za-z]+' \
internal/generated/fidelity/manifest_gen.go → 1 hit, "Tag"
The gate caught the consequence, as designed:
coverage_test.go:162: docs/coverage.md publishes 4496 hand-verified
operations, the manifest holds 4497
coverage_test.go:162: docs/coverage.md publishes 2718 unimplemented
operations, the manifest holds 2717
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 3
* test: add reproducer for S3 serving unknown sub-resources as listings
RED: CGO_ENABLED=0 go test ./internal/services/s3/ -run UnhandledBucketSubresource
--- FAIL: .../analytics --- FAIL: .../lifecycle
--- FAIL: .../inventory --- FAIL: .../encryption
--- FAIL: .../metrics --- FAIL: .../versions
--- FAIL: .../replication --- FAIL: .../object-lock
"<ListBucketResult><Name>sub-bucket</Name>..." should not contain
"ListBucketResult"
"200" is not greater than or equal to "400"
GET /{bucket}?analytics misses every bucket sub-resource check and falls through
to listObjects, so botocore reads a 200 <ListBucketResult> as a successful
ListBucketAnalyticsConfigurations. Seven more sub-resources take the same path.
The seven legitimate listing shapes in the same test already pass, and are there
so the fix cannot be a guard that declines a real ListObjects.
Recorded in .claude/tdd/query-providers-fabricate-success.tdd.md as one of the
two remaining xfail(strict) violations.
* fix: S3 served an unimplemented bucket sub-resource as a listing
GREEN: CGO_ENABLED=0 go test ./internal/services/s3/ → ok (8 sub-resources
decline, 7 listing shapes still serve)
CGO_ENABLED=0 go test ./... → all pass
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1127 passed, 17 skipped, 0 xfailed
make codegen; git status --porcelain internal/generated → clean
bash scripts/generate-imports.sh; git diff --exit-code → clean
golangci-lint run → 0 issues
GET /{bucket}?analytics matched no sub-resource branch and fell through to
listObjects, which answers 200 with <ListBucketResult>. botocore reads that as
a successful ListBucketAnalyticsConfigurations — a fabricated success, the one
thing docs/coverage.md calls absolute. The recorded report named one
sub-resource; the fall-through covered eight.
The guard is an allow-list of the parameters a listing carries, not a deny-list
of sub-resource names, because the two fail in opposite directions. A missing
allow-list entry declines a real listing and a test says so at once; a missing
deny-list entry serves a sub-resource as a listing, silently, which is the
defect. A sub-resource AWS adds next year now declines on its own.
The xfail(strict) marker did its job: it turned XPASS the moment the guard
landed, so KNOWN_UNFIXED could not quietly keep an entry for something already
fixed. Both entries are now gone and the dict is empty.
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 4
* chore: point the changelog fragments at PR #146
The four Fixed entries were written before the PR existed and guessed
consecutive numbers. custom.Issue carries the PR number in this repo — #142
through #145 all match their PR — and all four fragments ship in one PR.
4 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Milestone 6 of
aws-service-coverage-100: CI now gates the published coverage figure in both directions, and the boto3 compatibility suite exercises every registered service instead of a hand-written list of 69. Building the gate exposed that the quality floor's second clause — a passing compatibility test exists for the service — had never been measured: 31 of 205 registered services had no boto3 test at all, 29 of them counted inside the published 201.Related Issue
N/A — this delivers the final milestone of
.claude/prds/aws-service-coverage-100.prd.md. Refs #140, #142, #143.Changes
The gate (
cmd/devcloud/coverage_test.go, new)docs/coverage.md,README.mdanddocs/README.mdand compares each to the live registry and fidelity manifest. Removing a service without moving the published number fails; editing the number without the code moving fails identically.docs/demand.mdmust stay registered, so the count cannot be held steady by swapping one out).minServices = 100/minOperations = 6000infidelity_test.goandminServices = 50inconformance_test.go. Those were right for catching a mangled generator and could not notice 205 becoming 120.ci.ymlalready runsgo test ./...andcompat.ymlalready runs pytest.The compatibility suite (
tests/compatibility/)test_service_smoke.pyparametrises over all 205 services from a generated manifest.ENGINE_SERVED_SERVICES(67 entries) andREGISTERED_ONLY_SERVICES(2) are deleted — registering a service and testing it are no longer separate steps._coverage.py(new) resolves DevCloud service IDs to boto3 client names from botocore's own model metadata, refusing a contested name rather than guessing. 204 of 205 resolve automatically;elasticloadbalancing → elbis the single override.Codegen (
internal/codegen/gen_compat_manifest.go, new)internal/generated/compat/services.json— per registered service, its protocol and served operations, projected from the same fidelity datadocs/coverage.mdpublishes from. It lives underinternal/generatedso the existingcodegen-driftjob covers it for free.Docs
docs/coverage.mdgains a fourth number — Compatibility-tested: 199 — with its ceiling and the reason for each exclusion.README.mdanddocs/README.mdhad quoted 148 registered / 117 serving since three milestones earlier;docs/demand.mdnamed three demand-set services as serving nothing when Milestone 5 left one;sagemakerruntimehttp2serves nothing and was never named on a page that claims to name them all.Test Plan
The gate was proven to bite by mutation, four ways, each reverted immediately:
services/omicsfromimports.go**205**to**206**indocs/coverage.mdREADME.mdto148 registered / 117 servingREADME.mdentirelyThe fourth mutation was run because the first version of the pattern did not match
README.md's phrasing at all and passed silently on a stale figure. Zero matches is now a failure.Findings recorded, deliberately not fixed here
Each is a pre-existing defect the gate surfaced. None is silent — all three are written down in
docs/coverage.mdor the evidence report.rds,neptune,docdb,redshift,elasticache,autoscaling,cloudformationandelasticloadbalancingv2answer any unimplemented query action with HTTP 200 and an empty<{Action}Response/>. They never returnplugin.ErrUnhandledOp, so the CRUD engine is never reached, and the body has no{Action}Resultwrapper for botocore to read. This contradicts the guaranteedocs/coverage.mdcalls absolute and is the same class as thes3-controlP1 fixed in fix: S3 Control requests were served by S3, silently #142 — it deserves its own PR.lex, whichBuildAliasesleaves unrouted by design. Published indocs/coverage.mdand pinned bytest_lex_services_are_unreachable_from_boto3, so fixing the routing fails that test and raises the published figure with it. A correct fix needs a four-way path split —/botsis the first segment for three of the four.hand-verifiedbut no route reaches them —appsync/ListApis,eks/ListAccessPolicies,opensearch/ListApplications. Operation-level, not service-level; all three services still meet the floor.Checklist
golangci-lint run→ 0 issues)docs/coverage.md,docs/demand.md,README.md,docs/README.md)