Skip to content

fix: eight query providers fabricated a success for unimplemented actions - #145

Merged
skyoo2003 merged 2 commits into
mainfrom
fix/query-providers-fabricate-success
Sep 5, 2026
Merged

skyoo2003 merged 2 commits into
mainfrom
fix/query-providers-fabricate-success

Conversation

@skyoo2003

Copy link
Copy Markdown
Owner

Summary

Eight query-protocol providers answered any operation they do not implement with HTTP 200 and an empty <ActionResponse/> — a fabricated success, which docs/coverage.md calls the one thing DevCloud must never do. The body also carried no <ActionResult> wrapper, so botocore's query parser raised KeyError instead of returning a value: neptune.describe_db_cluster_endpoints() crashed the SDK outright.

Follows up #144, now merged. Base is main.

Related Issue

N/A — follows up finding 1 recorded in #144's "Findings recorded, deliberately not fixed here".

Changes

The fix — eight default branches now return nil, plugin.ErrUnhandledOp, which is what every other provider in the fleet does and what the CRUD engine has been reachable through since #143.

rds · neptune · docdb · redshift · elasticache · autoscaling · cloudformation · elasticloadbalancingv2

Root cause, not the reported symptom. The report was one operation on neptune. neptune and docdb sign as rds, so their probes never reach their own providers and neither appeared in the failing list — but both carry the identical branch. All eight are fixed; leaving two would leave the bug in place for anyone reaching them directly.

The reproducer — tests/compatibility/test_no_fabricated_success.py asks every registered service for one operation the fidelity manifest does not list as served, and fails if the answer is a success. It was written fleet-wide rather than against the eight services found by reading the code, and that paid for itself: it found two more (resourcegroups.Tag, s3.ListBucketAnalyticsConfigurations).

test_service_smoke.py already asserted this guarantee for services that serve nothing. The gap was the other case — a service that serves plenty and answers an operation it does not implement with an empty success anyway.

Published figures move, and had to. 223 operations go from unimplemented to auto-crud: they are served from the real store now instead of being fabricated. docs/coverage.md's tier table moves with them — auto-crud 4,970 → 5,193, unimplemented 2,941 → 2,718, total unchanged at 12,407.

That edit was not optional. #144's gate failed until the published numbers matched the manifest, naming both:

coverage_test.go:162: docs/coverage.md publishes 4970 auto-crud operations, the manifest holds 5193
coverage_test.go:162: docs/coverage.md publishes 2941 unimplemented operations, the manifest holds 2718

This is the first time that gate has caught a real change rather than a deliberate mutation.

Test Plan

RED (commit c65b9d9, output quoted in its message):

FAILED [autoscaling.DescribeAdjustmentTypes]        FAILED [rds.DescribeBlueGreenDeployments]
FAILED [cloudformation.DescribeChangeSetHooks]      FAILED [redshift.DescribeAccountAttributes]
FAILED [elasticache.DescribeCacheEngineVersions]    FAILED [resourcegroups.Tag]
FAILED [elasticloadbalancingv2.DescribeCapacityReservation]  FAILED [s3.ListBucketAnalyticsConfigurations]
8 failed, 120 passed, 17 skipped

GREEN (commit 34ee0b0): 126 passed, 17 skipped, 2 xfailed

CGO_ENABLED=0 go test ./...                             → all pass
make codegen; git status --porcelain internal/generated → clean
golangci-lint run                                       → 0 issues
pytest tests/compatibility/                             → 1119 passed, 17 skipped, 2 xfailed in 13.28s

What is deliberately not fixed

Two of the eight failures are a different mechanism and are marked xfail(strict=True) — so they fail the moment they start passing, rather than rotting into an accepted state.

Case Observed
s3.ListBucketAnalyticsConfigurations GET /{Bucket}?analytics answers 200 with an empty body. S3's provider default returns MethodNotAllowed, so it is not that branch
resourcegroups.Tag PUT /resources/{Arn}/tags answers 200 echoing the request's Arn and Tags

The recorded reasons state only what was observed. An early hypothesis — that httproute.Match ignores the HTTP method, letting PUT /resources/{Arn}/tags be served by the registered GetTags at the same path — was checked and is wrong: match.go:42 compares methods. Rather than commit a confident wrong cause, the comment says the mechanism is unidentified.

Also unchanged from #144's findings: the four unreachable Lex services (needs a four-way path split) and the three operations labelled hand-verified that no route reaches.

Checklist

  • Self-reviewed the code
  • Added/updated tests
  • Lint/format passes (golangci-lint run → 0 issues)
  • Updated documentation (docs/coverage.md tier table)
  • Added a Changie changelog fragment for user-facing changes (Fixed)

RED: 8 failed, 120 passed, 17 skipped.

    FAILED [autoscaling.DescribeAdjustmentTypes]
    FAILED [cloudformation.DescribeChangeSetHooks]
    FAILED [elasticache.DescribeCacheEngineVersions]
    FAILED [elasticloadbalancingv2.DescribeCapacityReservation]
    FAILED [rds.DescribeBlueGreenDeployments]
    FAILED [redshift.DescribeAccountAttributes]
    FAILED [resourcegroups.Tag]
    FAILED [s3.ListBucketAnalyticsConfigurations]

Each returned HTTP 200 for an operation the fidelity manifest does not list as
served. docs/coverage.md calls that the one thing DevCloud must never do.

test_service_smoke.py already asserted this for services that serve nothing at
all. The gap was the other case: a service that serves plenty and answers an
operation it does not implement with an empty success anyway.
…ions

GREEN: 126 passed, 17 skipped, 2 xfailed. The six failures this change targets
are gone; the two that remain are different defects, marked xfail(strict) with
what was observed rather than a guessed cause.

The default branch of each provider returned 200 with an empty
<ActionResponse/> for any action it did not implement. Two consequences, both
bad: the caller was told an operation succeeded when nothing ran, and the body
carried no <ActionResult> wrapper, so botocore's query parser raised KeyError
instead of returning a value. neptune.describe_db_cluster_endpoints() crashed
the SDK outright.

Each now returns plugin.ErrUnhandledOp, which is what every other provider in
the fleet does and what the CRUD engine has been reachable through since
query was admitted in #143. 223 operations move from unimplemented to
auto-crud — they are served from the real store now — and the rest decline
with InvalidAction.

docs/coverage.md's tier table moves with them: auto-crud 4,970 -> 5,193,
unimplemented 2,941 -> 2,718, total unchanged. That edit was not optional;
the gate added in #144 failed until the published figures matched the
manifest, naming both numbers.
@github-actions github-actions Bot added documentation Improvements or additions to documentation tests Test code and test infrastructure services AWS service implementations labels Sep 5, 2026
@skyoo2003
skyoo2003 merged commit efc4ec2 into main Sep 5, 2026
8 checks passed
@skyoo2003
skyoo2003 deleted the fix/query-providers-fabricate-success branch September 5, 2026 20:07
skyoo2003 added a commit that referenced this pull request Sep 6, 2026
The four Fixed entries were written before the PR existed and guessed
consecutive numbers. custom.Issue carries the PR number in this repo — #142
through #145 all match their PR — and all four fragments ship in one PR.
skyoo2003 added a commit that referenced this pull request Sep 6, 2026
* test: add reproducer for the four unreachable Lex services

RED: CGO_ENABLED=0 go test ./internal/gateway/ -run TestDetectProtocol_Lex

    --- FAIL: TestDetectProtocol_Lex/models_v1_get_bots
    --- FAIL: TestDetectProtocol_Lex/models_v1_get_intents
    --- FAIL: TestDetectProtocol_Lex/models_v1_builtin_intents
    --- FAIL: TestDetectProtocol_Lex/models_v2_list_bots
    --- FAIL: TestDetectProtocol_Lex/models_v2_create_bot
    --- FAIL: TestDetectProtocol_Lex/models_v2_describe_bot
    --- FAIL: TestDetectProtocol_Lex/runtime_v1_get_session
    --- FAIL: TestDetectProtocol_Lex/runtime_v1_put_session
    --- FAIL: TestDetectProtocol_Lex/runtime_v2_get_session
    --- FAIL: TestDetectProtocol_Lex/runtime_v2_delete_session
    --- PASS: TestDetectProtocol_Lex/delete_bot_is_contested

Every failure reads `expected: <a lex service>, actual: "lex"` — the request
never leaves the contested alias. The one PASS is the genuine collision
(DELETE /bots/{id}, claimed by both model services), which must stay refused.

Recorded in .claude/tdd/aws-service-coverage-100-milestone-6.tdd.md as
"Coverage and known gaps", item 3.

* fix: the four Lex services were reachable by no boto3 caller

GREEN: CGO_ENABLED=0 go test ./internal/gateway/ -run TestDetectProtocol_Lex → ok
       CGO_ENABLED=0 go test ./...                                          → all pass
       DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
         → 1123 passed, 17 skipped, 2 xfailed (was 1119 passed)

All four Lex clients sign as "lex" and no service is called "lex", so the alias
stays contested. normalizeServiceID hands the unresolved name through, the
registry has nothing under it, and every Lex call died as UnknownService — four
services registered, counted in "serving >=1 operation", and reachable by nobody.

The fix needed no new routing. resolveSharedSigningName already picks the
sibling whose route table models the request; it was never reached, because
signingNameOf only recognised a *member* service ID and "lex" is the group's
key. One map lookup.

The four route tables separate cleanly on method and literal segment:

    GET  /bots                                  lexmodelbuildingservice
    POST /bots                                  lexmodelsv2 (ListBots)
    PUT  /bots                                  lexmodelsv2 (CreateBot)
    GET  /bot/{n}/alias/{a}/user/{u}/session    lexruntimeservice
    GET  /bots/{b}/botAliases/.../sessions/{s}  lexruntimev2

One genuine collision, DeleteBot at DELETE /bots/{id}, claimed by both model
services. It stays refused: deleting the wrong bot is worse than an honest
UnknownService. Asserted, not left to be found.

The pin test told whoever fixed this to delete it, so it is deleted and replaced
by a check on the direction that still matters — UNREACHABLE_FROM_BOTO3 must
stay empty. docs/coverage.md's compatibility-tested figure moves 199 -> 203, the
Lex exclusion section is rewritten, and the stale claim that each Lex service is
"reached by its own unambiguous name" is corrected: no boto3 caller has one.

Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 1

* test: add reproducer for hand-verified operations no route reaches

RED: DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/test_handverified_is_reachable.py

    FAILED [appsync.ListApis]
    FAILED [eks.ListAccessPolicies]
    FAILED [opensearch.ListApplications]
    3 failed in 1.56s

Each answers NotImplemented. All three are labelled hand-verified in the
fidelity manifest, and all three are implemented — codegen reads the case
clause, and the case clause is there. What is missing is the route: the
gateway passes no operation name for rest-json, so the provider recovers it
from method and path with a hand-written resolver, and none of these three
resolvers knows the path its own model publishes.

    appsync/ListApis          GET /v2/apis            resolveOp strips only /v1
    eks/ListAccessPolicies    GET /access-policies    no case at all
    opensearch/ListApplications
                              GET /2021-01-01/opensearch/list-applications
                              resolveOp knows only GET /application

Recorded in .claude/tdd/aws-service-coverage-100-milestone-6.tdd.md as
"Coverage and known gaps", item 2.

* fix: three hand-verified operations were reached by no route

GREEN: DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/test_handverified_is_reachable.py
         → 3 passed
       CGO_ENABLED=0 go test ./...                → all pass
       DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
         → 1126 passed, 17 skipped, 2 xfailed (was 1123)
       make codegen; git status --porcelain internal/generated → clean
       golangci-lint run                          → 0 issues

appsync/ListApis, eks/ListAccessPolicies and opensearch/ListApplications are
each implemented by a case clause the provider already has. Codegen reads that
clause and labels the operation hand-verified — which proves the code exists,
not that any request arrives at it. For rest-json those are different
questions: the gateway passes no operation name, so the provider recovers it
from method and path with a hand-written resolver, and none of these three
resolvers knew the route its own model publishes.

The resolvers are partial copies of a table codegen already emits. Each
provider now ends with crud.Route(service, method, uri) instead of losing the
request. The fallback fires only where the resolver produced nothing, so it
cannot displace a decision a provider made — it fills gaps.

crud.Route is HasRoute's implementation, exported; HasRoute becomes its
one-line caller. Nothing else in the gateway changes.

No published figure moves. These operations were already counted as served;
what changes is that the count is now true about them.

27 other providers hand-roll the same kind of resolver. Whether any of them has
the same gap is not measured here — the compat manifest carries served
operations but not their tier, so the fleet-wide version of this test needs a
codegen change first. Recorded, not guessed at.

Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 2

* fix: the manifest reported an implemented operation as unimplemented

RED (compile-time, the intended missing field):
    CGO_ENABLED=0 go test ./internal/codegen/

    internal/codegen/gen_fidelity_test.go:83:4: unknown field ShortOperations
        in struct literal of type ProviderScan
    internal/codegen/scan_handverified_test.go:93:45:
        scans["resourcegroups"].ShortOperations undefined
    FAIL github.com/skyoo2003/devcloud/internal/codegen [build failed]

The RED could not be committed on its own: .pre-commit-config.yaml runs go vet,
which rejects a non-compiling tree, so RED and GREEN share this commit and the
output above is the evidence. Same constraint as #144.

GREEN: CGO_ENABLED=0 go test ./internal/codegen/  → ok
       CGO_ENABLED=0 go test ./...                → all pass
       make codegen                               → exactly one tier change
       DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
         → 1126 passed, 17 skipped, 1 xfailed (was 2 xfailed)
       golangci-lint run                          → 0 issues

opNamePattern requires four characters. resourcegroups implements Tag and Untag
in the same switch; Untag was scanned, Tag was not, and the manifest called
implemented code unimplemented. test_no_fabricated_success then asked for Tag,
got the 200 the provider has always returned, and reported a fabricated success.
It was never fabricated. The answer was real and the manifest was wrong, so the
probe was right to complain and wrong about what it had found — which is why
that entry recorded only an observation and refused to name a mechanism.

Loosening the pattern is not the fix. At two characters it admits the HTTP verbs
every path resolver switches on. The model separates them: Tag is an operation
because resourcegroups.json says so, and no model AWS publishes declares GET.
Short literals are collected into ProviderScan.ShortOperations and promoted by
BuildFidelityData only where the service's model declares them.

Fleet-wide, exactly one operation is short enough to be affected:

    rg -o '"[A-Z][A-Za-z0-9]{2}":\s+Tier[A-Za-z]+' \
       internal/generated/fidelity/manifest_gen.go   → 1 hit, "Tag"

The gate caught the consequence, as designed:

    coverage_test.go:162: docs/coverage.md publishes 4496 hand-verified
        operations, the manifest holds 4497
    coverage_test.go:162: docs/coverage.md publishes 2718 unimplemented
        operations, the manifest holds 2717

Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 3

* test: add reproducer for S3 serving unknown sub-resources as listings

RED: CGO_ENABLED=0 go test ./internal/services/s3/ -run UnhandledBucketSubresource

    --- FAIL: .../analytics    --- FAIL: .../lifecycle
    --- FAIL: .../inventory    --- FAIL: .../encryption
    --- FAIL: .../metrics      --- FAIL: .../versions
    --- FAIL: .../replication  --- FAIL: .../object-lock

    "<ListBucketResult><Name>sub-bucket</Name>..." should not contain
    "ListBucketResult"
    "200" is not greater than or equal to "400"

GET /{bucket}?analytics misses every bucket sub-resource check and falls through
to listObjects, so botocore reads a 200 <ListBucketResult> as a successful
ListBucketAnalyticsConfigurations. Seven more sub-resources take the same path.

The seven legitimate listing shapes in the same test already pass, and are there
so the fix cannot be a guard that declines a real ListObjects.

Recorded in .claude/tdd/query-providers-fabricate-success.tdd.md as one of the
two remaining xfail(strict) violations.

* fix: S3 served an unimplemented bucket sub-resource as a listing

GREEN: CGO_ENABLED=0 go test ./internal/services/s3/  → ok (8 sub-resources
         decline, 7 listing shapes still serve)
       CGO_ENABLED=0 go test ./...                    → all pass
       DEVCLOUD_BIN=dist/devcloud pytest tests/compatibility/
         → 1127 passed, 17 skipped, 0 xfailed
       make codegen; git status --porcelain internal/generated → clean
       bash scripts/generate-imports.sh; git diff --exit-code  → clean
       golangci-lint run                              → 0 issues

GET /{bucket}?analytics matched no sub-resource branch and fell through to
listObjects, which answers 200 with <ListBucketResult>. botocore reads that as
a successful ListBucketAnalyticsConfigurations — a fabricated success, the one
thing docs/coverage.md calls absolute. The recorded report named one
sub-resource; the fall-through covered eight.

The guard is an allow-list of the parameters a listing carries, not a deny-list
of sub-resource names, because the two fail in opposite directions. A missing
allow-list entry declines a real listing and a test says so at once; a missing
deny-list entry serves a sub-resource as a listing, silently, which is the
defect. A sub-resource AWS adds next year now declines on its own.

The xfail(strict) marker did its job: it turned XPASS the moment the guard
landed, so KNOWN_UNFIXED could not quietly keep an entry for something already
fixed. Both entries are now gone and the dict is empty.

Plan: .claude/plans/aws-service-coverage-100-milestone-7.plan.md, Task 4

* chore: point the changelog fragments at PR #146

The four Fixed entries were written before the PR existed and guessed
consecutive numbers. custom.Issue carries the PR number in this repo — #142
through #145 all match their PR — and all four fragments ship in one PR.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation services AWS service implementations tests Test code and test infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant