Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -414,6 +414,30 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
observed from the example rather than required; it is documented beside the `backup-plan` vs. `plan`
segment gap (#1181) rather than folded into the derivation, since it is a rendering question and not
one about where the bytes come from.
- **Every identifier of every resource in every CloudFormation stack was random** (#856). This is the
part of #856 that counting draw sites could not find: `StackDeployer` holds no `crypto/rand` call of
its own, it turns each resource into an *internal* EC2, IAM or S3 request — and those requests were
built with a fresh request ID and no mint, so every plugin they reached took `IDMint`'s seedless
fallback however carefully its own minters had been migrated. A replayed `DescribeStackResources`
reported four different `PhysicalResourceId` values in a **200**, and `CreateStack`'s
`state_hash_after` never matched because that hash covers all of them. Two changes were needed and
they fix different failures: an internal request now carries `IDs: NewIDMint(requestID)`, which makes
a replayed internal event reproduce, and the internal request ID is itself derived from the stack
request's ID through the new `WithDeployerMint` option, which makes a replayed `CreateStack` — which
re-runs the whole deployment — reproduce too. Reverting either half alone was measured: seven
differences without the mint, six without the derived request ID. The sequence is reproducible because
a deploy already orders resources by type priority and then logical ID, never by map iteration. The
same dispatch fix reaches an API Gateway proxy integration, which invokes a Lambda through an internal
request. A Lambda event-source-mapping poll is deliberately left on the fallback and now says so in
the code: its dispatches come from a wall-clock ticker and are recorded nowhere, so there is no
recorded ID to derive from (#1292).
- **A `StackDriftDetectionId` was substrate's internal request ID, verbatim** (#856). `DetectStackDrift`
answered the `req-…` string substrate stamps on a request, a shape no AWS reference describes, in a
response element a consumer hands straight back to `DescribeStackDriftDetectionStatus`.
`StackDriftDetectionId` publishes a **maximum length of 36** and no pattern, and the page's own sample
is `2f2b2d60-df86-11e7-bea1-500c2example` — so it is now the UUID shape within that bound, derived from
the request ID. A recorded detection replays with the ID it recorded, where before the poll that
followed it resolved to nothing.
- **A stream recorded under a seed replays under the same seed** (#1140). Every seedable outcome in
substrate is written through a control-plane endpoint, and only the AWS path recorded anything — so
a seed never entered the event stream. A replay opens by resetting the whole `StateManager`, and a
Expand Down
46 changes: 42 additions & 4 deletions docs/services.md
Original file line number Diff line number Diff line change
Expand Up @@ -2318,6 +2318,36 @@ a third — so the counter is what keeps a plan from being its own version. AWS'
shows the UUID in *upper* case where substrate renders lower; that is an observed difference the model
does not require, tracked with the [resource segment](#arn-shapes) rather than with the derivation.

**A request that dispatches another request has to hand the mint along with it.** Counting draw
sites finds the places that call `crypto/rand`; it cannot find a plugin that derives its
identifiers correctly and is *called* with nothing to derive from. CloudFormation's deployer is
that case: deploying a stack turns each resource into an internal EC2, IAM or S3 request, and those
requests were built with a fresh request ID and no mint, so every plugin they reached took the
seedless fallback. The consequence is larger than the count suggested — **every identifier of every
resource in every CloudFormation stack** was undirected randomness, while the deployer itself
contains no draw site at all. A stack's `vpc-…`, its subnet, its security group and its role were
all unreproducible, and a `ValidateState` replay of the outer `CreateStack` reported a
`state_hash_after` mismatch because that hash covers all of them.

Two separate things had to change, and each fixes a different failure. An internal request carries a
mint seeded from its own ID, which is what makes a replayed *internal* event reproduce. And the
internal ID is itself derived from the stack request's ID, which is what makes a replayed
`CreateStack` reproduce — replaying it re-runs the whole deployment, so the internal IDs must come
out the same way twice. The sequence is reproducible because a deploy orders resources by type
priority and then by logical ID, never by map iteration. Both halves are load-bearing and were
measured that way: dropping the mint produces seven differences in a four-resource stack, dropping
the derived request ID produces six. Notably the replay reported every event as a **success** — a
200 describing four resources that were not the recorded ones, which is the silent-wrong-answer
class rather than a visible failure.

Removing substrate's internal request ID from that path also removed it from an AWS response.
`DetectStackDrift` answered a `StackDriftDetectionId` that was literally substrate's `req-…` string,
which no AWS reference describes; it is now UUID-shaped within the published maximum length of 36.
The same dispatch fix applies to an API Gateway proxy integration, which invokes a Lambda through an
internal request. A Lambda event-source-mapping poll is deliberately left random: its dispatches come
from a wall-clock ticker and are recorded nowhere, so there is no recorded ID to derive from and a
seed there would be no more reproducible than the fallback ([#1292](https://github.com/scttfrdmn/substrate/issues/1292)).

An ECR image digest is minted rather than computed from the manifest, so it is reproducible across
a replay but is not the SHA-256 of the image it names, and two pushes of identical manifest bytes
store two images where AWS stores one. [#1283](https://github.com/scttfrdmn/substrate/issues/1283)
Expand Down Expand Up @@ -2351,7 +2381,7 @@ the most recent one deletes. A replay of either call reproduces the handle that
| ExecuteChangeSet | Applies the change and its tags, and consumes the set |
| ListChangeSets | Pending change sets for a stack |
| DeleteChangeSet | Discards a pending set; deleting an absent set succeeds |
| DetectStackDrift | Returns a `StackDriftDetectionId` |
| DetectStackDrift | Returns a `StackDriftDetectionId` — UUID-shaped within the published maximum length of 36, [derived from the request ID](#an-identifier-a-replay-mints-is-the-one-it-recorded) (#856). It used to be substrate's internal `req-…` string, a shape no AWS reference describes |
| DescribeStackDriftDetectionStatus | Resolves that ID to a completed detection |
| DescribeStackResourceDrifts | Per-resource drift; honours `StackResourceDriftStatusFilters.member.N` |
| ListExports | Every exported output value in the caller's account and Region, in one page |
Expand Down Expand Up @@ -3489,6 +3519,14 @@ This was substrate's first derived identifier and is the one the general rule wa
modelled on; see [An identifier a replay mints is the one it
recorded](#an-identifier-a-replay-mints-is-the-one-it-recorded).

The **physical** IDs of the resources a stack deploys are derived too, and were
not until #856's dispatch fix: the deployer turns each resource into an internal
request, and those requests carried no mint, so a replayed stack reported
different `PhysicalResourceId` values and a different `state_hash_after` for
`CreateStack`. The internal request IDs now derive from the stack request's own
ID, in the order a deploy already uses — type priority, then logical ID — so a
replayed `CreateStack` re-runs the deployment and arrives at the same IDs.

### Cost

CloudFormation operations are free. The resources a template deploys are costed
Expand Down Expand Up @@ -12951,7 +12989,7 @@ Secrets Manager API calls: $0.05 per 10,000 API calls.
| RemoveTagsFromResource | Removes only the named keys |
| ListTagsForResource | Reports `TagList` sorted by key; an empty list, never `null` |
| LabelParameterVersion | Accepted; always reports `ParameterVersion: 1` |
| SendCommand | Run Command; records the intent — substrate does not execute the command. The `CommandId` is the published fixed length of 36, [derived from the request ID](#derived-identifiers) ([#856](https://github.com/scttfrdmn/substrate/issues/856)) |
| SendCommand | Run Command; records the intent — substrate does not execute the command. The `CommandId` is the published fixed length of 36, [derived from the request ID](#an-identifier-a-replay-mints-is-the-one-it-recorded) ([#856](https://github.com/scttfrdmn/substrate/issues/856)) |
| GetCommandInvocation | |
| DescribeInstanceInformation | |

Expand Down Expand Up @@ -19533,7 +19571,7 @@ The published example differs in one more way the table does not show: AWS's sam
`arn:aws:backup:us-east-1:123456789012:plan:8F81F553-3A74-4A3F-B93D-B3360DC80C50`, an **uppercase**
UUID, where Substrate renders lowercase. `BackupPlanId` publishes no pattern and no length, so this is
observed from the example rather than required by the model, and it is a rendering question rather than
a derivation one — the plan ID is [derived from the request ID](#derived-identifiers) either way.
a derivation one — the plan ID is [derived from the request ID](#an-identifier-a-replay-mints-is-the-one-it-recorded) either way.

### Cost

Expand Down Expand Up @@ -19567,7 +19605,7 @@ running.
|-----------|-------|
| InvokeModel | `POST /model/{modelId}/invoke`. Answers a [seeded response body](#seeding-a-model-response) verbatim, or a canned Claude Messages body naming the requested model. Nothing but the model ID is read — not the body, not `accept` or `contentType`, and [not the guardrail headers](#invokemodel-reads-nothing-but-the-model-id) |
| ApplyGuardrail | `POST /guardrail/{guardrailIdentifier}/version/{guardrailVersion}/apply`. [`NONE` or `GUARDRAIL_INTERVENED`, decided by a blocklist](#how-a-guardrail-decides); the version is discarded |
| CreateModelInvocationJob | `POST /model-invocation-job`. Answers `{"jobArn"}`, exactly the published shape, and records the job as `Submitted` — the first state the page documents, so a batch job is deliberately not terminal at birth. None of the five members marked `Required: Yes` is checked. The job ID is twelve characters of `[a-z0-9]`, the published `jobArn` pattern, [derived from the request ID](#derived-identifiers) ([#856](https://github.com/scttfrdmn/substrate/issues/856)) |
| CreateModelInvocationJob | `POST /model-invocation-job`. Answers `{"jobArn"}`, exactly the published shape, and records the job as `Submitted` — the first state the page documents, so a batch job is deliberately not terminal at birth. None of the five members marked `Required: Yes` is checked. The job ID is twelve characters of `[a-z0-9]`, the published `jobArn` pattern, [derived from the request ID](#an-identifier-a-replay-mints-is-the-one-it-recorded) ([#856](https://github.com/scttfrdmn/substrate/issues/856)) |
| GetModelInvocationJob | `GET /model-invocation-job/{jobIdentifier}`. Returns the stored record whole, so `accountID` and `region` reach the wire ([#756](https://github.com/scttfrdmn/substrate/issues/756)), and reports a [seeded status](#seeding-a-batch-job-status) if one is set |
| ListModelInvocationJobs | `GET /model-invocation-jobs`. `invocationJobSummaries` of five members each; the seeded status is applied here too, so a poll on either operation agrees. Every published query filter is ignored and no `nextToken` is emitted |
| StopModelInvocationJob | `POST /model-invocation-job/{jobIdentifier}/stop`. [Stops a job in any state and skips `Stopping`](#stopping-a-batch-job-is-immediate) |
Expand Down
7 changes: 7 additions & 0 deletions docs/testing-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -335,6 +335,13 @@ makes `Differences` empty — and `StateValid` true, with `WithRecordedStateHash
all; before #856 every such stream diverged on its first create, which is why the #1140
tests above are built on caller-chosen bucket and key names instead.

**That now includes a CloudFormation stack.** A deploy turns each resource into an internal
request, and those requests used to carry no mint, so every physical resource ID in every
stack was random even though the deployer holds no draw site of its own — a replayed
`DescribeStackResources` reported different IDs, in a 200, and `CreateStack`'s
`state_hash_after` never matched. A stack of a handful of resources now replays with zero
differences and `StateValid` true.

Two caveats. **Not every service's identifiers are derived yet.** EC2, IAM, STS, SQS, SNS,
Lambda, EFS, FSx, Transfer, ECS, Step Functions, EventBridge, CloudWatch Logs, CloudFront,
Service Quotas, API Gateway (v1 and v2), AppSync, Batch, EMR Serverless, ECR, ELB, Route 53,
Expand Down
25 changes: 24 additions & 1 deletion emulator/apigateway_plugin.go
Original file line number Diff line number Diff line change
Expand Up @@ -1628,11 +1628,18 @@ func (p *APIGatewayProxyPlugin) HandleRequest(reqCtx *RequestContext, req *AWSRe
Headers: map[string]string{"Content-Type": "application/json"},
Body: eventJSON,
}
// Derived from the request that reached the integration, so the invocation this
// dispatches mints what it minted in the recording (#856). A proxy integration is
// one of the tree's internal dispatch paths: the context is built here rather than
// parsed off the wire, so nothing gave it a mint and every identifier the invoked
// function's plugin published fell back to crypto/rand.
invokeRequestID := apigwInternalRequestID(reqCtx.IDs)
invokeCtx := &RequestContext{
RequestID: generateRequestID(),
RequestID: invokeRequestID,
AccountID: reqCtx.AccountID,
Region: reqCtx.Region,
Timestamp: reqCtx.Timestamp,
IDs: NewIDMint(invokeRequestID),
Metadata: make(map[string]interface{}),
}
invokeResp, invokeErr := p.registry.RouteRequest(invokeCtx, invokeReq)
Expand Down Expand Up @@ -1816,6 +1823,22 @@ func buildV1ProxyEvent(req *AWSRequest, apiID, stage, resourcePath string) ([]by
return json.Marshal(event)
}

// apigwInternalRequestID returns the request id for the Lambda invocation a proxy
// integration dispatches, derived from m when it derives.
//
// Separate from generateRequestID because the value seeds the invocation's own mint:
// a wall-clock id makes every identifier the invoked function's plugin publishes
// unreproducible, which is the gap #856's tier 8 closes on the two internal dispatch
// sites that have a recorded request behind them — this one and the CloudFormation
// deployer's. A mint that does not derive — an in-process caller with no request behind
// it — falls back to the wall clock, which is what this always did.
func apigwInternalRequestID(m *IDMint) string {
if !m.Derived() {
return generateRequestID()
}
return "req-apigw-" + m.Hex(12)
}

// buildV2ProxyEvent constructs a v2 (HTTP API) proxy event JSON payload.
func buildV2ProxyEvent(req *AWSRequest, apiID, stage, resourcePath string) ([]byte, error) {
rawQS := ""
Expand Down
Loading
Loading