Repository navigation
fix: eight query providers fabricated a success for unimplemented actions - #145
Merged
Merged
Conversation
RED: 8 failed, 120 passed, 17 skipped.
FAILED [autoscaling.DescribeAdjustmentTypes]
FAILED [cloudformation.DescribeChangeSetHooks]
FAILED [elasticache.DescribeCacheEngineVersions]
FAILED [elasticloadbalancingv2.DescribeCapacityReservation]
FAILED [rds.DescribeBlueGreenDeployments]
FAILED [redshift.DescribeAccountAttributes]
FAILED [resourcegroups.Tag]
FAILED [s3.ListBucketAnalyticsConfigurations]
Each returned HTTP 200 for an operation the fidelity manifest does not list as
served. docs/coverage.md calls that the one thing DevCloud must never do.
test_service_smoke.py already asserted this for services that serve nothing at
all. The gap was the other case: a service that serves plenty and answers an
operation it does not implement with an empty success anyway.
…ions GREEN: 126 passed, 17 skipped, 2 xfailed. The six failures this change targets are gone; the two that remain are different defects, marked xfail(strict) with what was observed rather than a guessed cause. The default branch of each provider returned 200 with an empty <ActionResponse/> for any action it did not implement. Two consequences, both bad: the caller was told an operation succeeded when nothing ran, and the body carried no <ActionResult> wrapper, so botocore's query parser raised KeyError instead of returning a value. neptune.describe_db_cluster_endpoints() crashed the SDK outright. Each now returns plugin.ErrUnhandledOp, which is what every other provider in the fleet does and what the CRUD engine has been reachable through since query was admitted in #143. 223 operations move from unimplemented to auto-crud — they are served from the real store now — and the rest decline with InvalidAction. docs/coverage.md's tier table moves with them: auto-crud 4,970 -> 5,193, unimplemented 2,941 -> 2,718, total unchanged. That edit was not optional; the gate added in #144 failed until the published figures matched the manifest, naming both numbers.
5 tasks done
skyoo2003
added a commit
that referenced
this pull request
Sep 6, 2026
* test: add reproducer for the four unreachable Lex services
RED: CGO_ENABLED=0 go test ./internal/gateway/ -run TestDetectProtocol_Lex
--- FAIL: TestDetectProtocol_Lex/models_v1_get_bots
--- FAIL: TestDetectProtocol_Lex/models_v1_get_intents
--- FAIL: TestDetectProtocol_Lex/models_v1_builtin_intents
--- FAIL: TestDetectProtocol_Lex/models_v2_list_bots
--- FAIL: TestDetectProtocol_Lex/models_v2_create_bot
--- FAIL: TestDetectProtocol_Lex/models_v2_describe_bot
--- FAIL: TestDetectProtocol_Lex/runtime_v1_get_session
--- FAIL: TestDetectProtocol_Lex/runtime_v1_put_session
--- FAIL: TestDetectProtocol_Lex/runtime_v2_get_session
--- FAIL: TestDetectProtocol_Lex/runtime_v2_delete_session
--- PASS: TestDetectProtocol_Lex/delete_bot_is_contested
Every failure reads `expected: <a lex service>, actual: "lex"` — the request
never leaves the contested alias. The one PASS is the genuine collision
(DELETE /bots/{id}, claimed by both model services), which must stay refused.
Recorded in .claude/tdd/aws-service-coverage-100-milestone-6.tdd.md as
"Coverage and known gaps", item 3.
* fix: the four Lex services were reachable by no boto3 caller
GREEN: CGO_ENABLED=0 go test ./internal/gateway/ -run TestDetectProtocol_Lex → ok
CGO_ENABLED=0 go test ./... → all pass
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1123 passed, 17 skipped, 2 xfailed (was 1119 passed)
All four Lex clients sign as "lex" and no service is called "lex", so the alias
stays contested. normalizeServiceID hands the unresolved name through, the
registry has nothing under it, and every Lex call died as UnknownService — four
services registered, counted in "serving >=1 operation", and reachable by nobody.
The fix needed no new routing. resolveSharedSigningName already picks the
sibling whose route table models the request; it was never reached, because
signingNameOf only recognised a *member* service ID and "lex" is the group's
key. One map lookup.
The four route tables separate cleanly on method and literal segment:
GET /bots lexmodelbuildingservice
POST /bots lexmodelsv2 (ListBots)
PUT /bots lexmodelsv2 (CreateBot)
GET /bot/{n}/alias/{a}/user/{u}/session lexruntimeservice
GET /bots/{b}/botAliases/.../sessions/{s} lexruntimev2
One genuine collision, DeleteBot at DELETE /bots/{id}, claimed by both model
services. It stays refused: deleting the wrong bot is worse than an honest
UnknownService. Asserted, not left to be found.
The pin test told whoever fixed this to delete it, so it is deleted and replaced
by a check on the direction that still matters — UNREACHABLE_FROM_BOTO3 must
stay empty. docs/coverage.md's compatibility-tested figure moves 199 -> 203, the
Lex exclusion section is rewritten, and the stale claim that each Lex service is
"reached by its own unambiguous name" is corrected: no boto3 caller has one.
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 1
* test: add reproducer for hand-verified operations no route reaches
RED: DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/test_handverified_is_reachable.py
FAILED [appsync.ListApis]
FAILED [eks.ListAccessPolicies]
FAILED [opensearch.ListApplications]
3 failed in 1.56s
Each answers NotImplemented. All three are labelled hand-verified in the
fidelity manifest, and all three are implemented — codegen reads the case
clause, and the case clause is there. What is missing is the route: the
gateway passes no operation name for rest-json, so the provider recovers it
from method and path with a hand-written resolver, and none of these three
resolvers knows the path its own model publishes.
appsync/ListApis GET /v2/apis resolveOp strips only /v1
eks/ListAccessPolicies GET /access-policies no case at all
opensearch/ListApplications
GET /2021-01-01/opensearch/list-applications
resolveOp knows only GET /application
Recorded in .claude/tdd/aws-service-coverage-100-milestone-6.tdd.md as
"Coverage and known gaps", item 2.
* fix: three hand-verified operations were reached by no route
GREEN: DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/test_handverified_is_reachable.py
→ 3 passed
CGO_ENABLED=0 go test ./... → all pass
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1126 passed, 17 skipped, 2 xfailed (was 1123)
make codegen; git status --porcelain internal/generated → clean
golangci-lint run → 0 issues
appsync/ListApis, eks/ListAccessPolicies and opensearch/ListApplications are
each implemented by a case clause the provider already has. Codegen reads that
clause and labels the operation hand-verified — which proves the code exists,
not that any request arrives at it. For rest-json those are different
questions: the gateway passes no operation name, so the provider recovers it
from method and path with a hand-written resolver, and none of these three
resolvers knew the route its own model publishes.
The resolvers are partial copies of a table codegen already emits. Each
provider now ends with crud.Route(service, method, uri) instead of losing the
request. The fallback fires only where the resolver produced nothing, so it
cannot displace a decision a provider made — it fills gaps.
crud.Route is HasRoute's implementation, exported; HasRoute becomes its
one-line caller. Nothing else in the gateway changes.
No published figure moves. These operations were already counted as served;
what changes is that the count is now true about them.
27 other providers hand-roll the same kind of resolver. Whether any of them has
the same gap is not measured here — the compat manifest carries served
operations but not their tier, so the fleet-wide version of this test needs a
codegen change first. Recorded, not guessed at.
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 2
* fix: the manifest reported an implemented operation as unimplemented
RED (compile-time, the intended missing field):
CGO_ENABLED=0 go test ./internal/codegen/
internal/codegen/gen_fidelity_test.go:83:4: unknown field ShortOperations
in struct literal of type ProviderScan
internal/codegen/scan_handverified_test.go:93:45:
scans["resourcegroups"].ShortOperations undefined
FAIL github.com/skyoo2003/devcloud/internal/codegen [build failed]
The RED could not be committed on its own: .pre-commit-config.yaml runs go vet,
which rejects a non-compiling tree, so RED and GREEN share this commit and the
output above is the evidence. Same constraint as #144.
GREEN: CGO_ENABLED=0 go test ./internal/codegen/ → ok
CGO_ENABLED=0 go test ./... → all pass
make codegen → exactly one tier change
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1126 passed, 17 skipped, 1 xfailed (was 2 xfailed)
golangci-lint run → 0 issues
opNamePattern requires four characters. resourcegroups implements Tag and Untag
in the same switch; Untag was scanned, Tag was not, and the manifest called
implemented code unimplemented. test_no_fabricated_success then asked for Tag,
got the 200 the provider has always returned, and reported a fabricated success.
It was never fabricated. The answer was real and the manifest was wrong, so the
probe was right to complain and wrong about what it had found — which is why
that entry recorded only an observation and refused to name a mechanism.
Loosening the pattern is not the fix. At two characters it admits the HTTP verbs
every path resolver switches on. The model separates them: Tag is an operation
because resourcegroups.json says so, and no model AWS publishes declares GET.
Short literals are collected into ProviderScan.ShortOperations and promoted by
BuildFidelityData only where the service's model declares them.
Fleet-wide, exactly one operation is short enough to be affected:
rg -o '"[A-Z][A-Za-z0-9]{2}":\s+Tier[A-Za-z]+' \
internal/generated/fidelity/manifest_gen.go → 1 hit, "Tag"
The gate caught the consequence, as designed:
coverage_test.go:162: docs/coverage.md publishes 4496 hand-verified
operations, the manifest holds 4497
coverage_test.go:162: docs/coverage.md publishes 2718 unimplemented
operations, the manifest holds 2717
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 3
* test: add reproducer for S3 serving unknown sub-resources as listings
RED: CGO_ENABLED=0 go test ./internal/services/s3/ -run UnhandledBucketSubresource
--- FAIL: .../analytics --- FAIL: .../lifecycle
--- FAIL: .../inventory --- FAIL: .../encryption
--- FAIL: .../metrics --- FAIL: .../versions
--- FAIL: .../replication --- FAIL: .../object-lock
"<ListBucketResult><Name>sub-bucket</Name>..." should not contain
"ListBucketResult"
"200" is not greater than or equal to "400"
GET /{bucket}?analytics misses every bucket sub-resource check and falls through
to listObjects, so botocore reads a 200 <ListBucketResult> as a successful
ListBucketAnalyticsConfigurations. Seven more sub-resources take the same path.
The seven legitimate listing shapes in the same test already pass, and are there
so the fix cannot be a guard that declines a real ListObjects.
Recorded in .claude/tdd/query-providers-fabricate-success.tdd.md as one of the
two remaining xfail(strict) violations.
* fix: S3 served an unimplemented bucket sub-resource as a listing
GREEN: CGO_ENABLED=0 go test ./internal/services/s3/ → ok (8 sub-resources
decline, 7 listing shapes still serve)
CGO_ENABLED=0 go test ./... → all pass
DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
→ 1127 passed, 17 skipped, 0 xfailed
make codegen; git status --porcelain internal/generated → clean
bash scripts/generate-imports.sh; git diff --exit-code → clean
golangci-lint run → 0 issues
GET /{bucket}?analytics matched no sub-resource branch and fell through to
listObjects, which answers 200 with <ListBucketResult>. botocore reads that as
a successful ListBucketAnalyticsConfigurations — a fabricated success, the one
thing docs/coverage.md calls absolute. The recorded report named one
sub-resource; the fall-through covered eight.
The guard is an allow-list of the parameters a listing carries, not a deny-list
of sub-resource names, because the two fail in opposite directions. A missing
allow-list entry declines a real listing and a test says so at once; a missing
deny-list entry serves a sub-resource as a listing, silently, which is the
defect. A sub-resource AWS adds next year now declines on its own.
The xfail(strict) marker did its job: it turned XPASS the moment the guard
landed, so KNOWN_UNFIXED could not quietly keep an entry for something already
fixed. Both entries are now gone and the dict is empty.
Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 4
* chore: point the changelog fragments at PR #146
The four Fixed entries were written before the PR existed and guessed
consecutive numbers. custom.Issue carries the PR number in this repo — #142
through #145 all match their PR — and all four fragments ship in one PR.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Eight query-protocol providers answered any operation they do not implement with HTTP 200 and an empty
<ActionResponse/>— a fabricated success, whichdocs/coverage.mdcalls the one thing DevCloud must never do. The body also carried no<ActionResult>wrapper, so botocore's query parser raisedKeyErrorinstead of returning a value:neptune.describe_db_cluster_endpoints()crashed the SDK outright.Related Issue
N/A — follows up finding 1 recorded in #144's "Findings recorded, deliberately not fixed here".
Changes
The fix — eight
defaultbranches nowreturn nil, plugin.ErrUnhandledOp, which is what every other provider in the fleet does and what the CRUD engine has been reachable through since #143.rds·neptune·docdb·redshift·elasticache·autoscaling·cloudformation·elasticloadbalancingv2Root cause, not the reported symptom. The report was one operation on
neptune.neptuneanddocdbsign asrds, so their probes never reach their own providers and neither appeared in the failing list — but both carry the identical branch. All eight are fixed; leaving two would leave the bug in place for anyone reaching them directly.The reproducer —
tests/compatibility/test_no_fabricated_success.pyasks every registered service for one operation the fidelity manifest does not list as served, and fails if the answer is a success. It was written fleet-wide rather than against the eight services found by reading the code, and that paid for itself: it found two more (resourcegroups.Tag,s3.ListBucketAnalyticsConfigurations).test_service_smoke.pyalready asserted this guarantee for services that serve nothing. The gap was the other case — a service that serves plenty and answers an operation it does not implement with an empty success anyway.Published figures move, and had to. 223 operations go from
unimplementedtoauto-crud: they are served from the real store now instead of being fabricated.docs/coverage.md's tier table moves with them — auto-crud 4,970 → 5,193, unimplemented 2,941 → 2,718, total unchanged at 12,407.That edit was not optional. #144's gate failed until the published numbers matched the manifest, naming both:
This is the first time that gate has caught a real change rather than a deliberate mutation.
Test Plan
RED (commit
c65b9d9, output quoted in its message):GREEN (commit
34ee0b0):126 passed, 17 skipped, 2 xfailedWhat is deliberately not fixed
Two of the eight failures are a different mechanism and are marked
xfail(strict=True)— so they fail the moment they start passing, rather than rotting into an accepted state.s3.ListBucketAnalyticsConfigurationsGET /{Bucket}?analyticsanswers 200 with an empty body. S3's provider default returnsMethodNotAllowed, so it is not that branchresourcegroups.TagPUT /resources/{Arn}/tagsanswers 200 echoing the request'sArnandTagsThe recorded reasons state only what was observed. An early hypothesis — that
httproute.Matchignores the HTTP method, lettingPUT /resources/{Arn}/tagsbe served by the registeredGetTagsat the same path — was checked and is wrong:match.go:42compares methods. Rather than commit a confident wrong cause, the comment says the mechanism is unidentified.Also unchanged from #144's findings: the four unreachable Lex services (needs a four-way path split) and the three operations labelled
hand-verifiedthat no route reaches.Checklist
golangci-lint run→ 0 issues)docs/coverage.mdtier table)Fixed)