Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 34 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -325,6 +325,40 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
bytes are the `Plaintext`, and the same bytes are wrapped into the `CiphertextBlob`. A recorded
`Decrypt` was never affected — substrate's ciphertext carries its own plaintext, so a replayed
decrypt answers from the recorded blob regardless — so the break was in the create, not the read.
- **The analytics family of draw sites is derived** (#856). Seven generators across six services move
onto `IDMint`: an Athena query execution ID, a Redshift Data statement ID, a Glue job-run ID, a
Timestream `QueryId`, an OpenSearch document `_id` and `_scroll_id`, and a QuickSight ingestion ID
and response `RequestId`. Every rendering is byte-for-byte the one the `crypto/rand` version
produced, so an identifier a previous substrate recorded is still the shape this one mints. 11 draw
sites remain on `crypto/rand`.
- **A recorded Athena poll loop replays against the query it started** (#856). An analytics identifier
names a submission rather than a resource, so unlike a bucket or a job there is no caller-chosen name
to fall back on: `GetQueryExecution`, `GetQueryResults` and `StopQueryExecution` all key on the one ID
`StartQueryExecution` returned, and a re-minted one answered every recorded poll with
`InvalidRequestException` against a query the recording had just created. Redshift Data breaks the
same way with `ResourceNotFoundException` and Glue with `EntityNotFoundException`. Reverting the
Athena minter alone to confirm the family's replay assertion is not vacuous produces 38 differences
and two refused reads out of a 19-request stream.
- **A QuickSight ingestion ID is no longer minted by the request-ID generator** (#856). `CreateDataSet`
drew the SPICE ingestion ID it reports from the same function that mints the `RequestId` every
QuickSight response carries — the second instance of the cross-service borrow the ACM/API Gateway
split fixed. An ingestion ID is a handle a recorded `DescribeIngestion` URL contains, where a request
ID is observed once and never sent back, so one function serving both meant a change to how a request
ID renders would have moved the identifier a recorded path depends on. Both now have their own
minter and the rendering is unchanged.
- **`IDMint.Base64URL`, for the one identifier that travels in a URL path** (#856). An OpenSearch
document `_id` is the path of every later `GET`, `PUT` and `DELETE` of that document, so its alphabet
is `-` and `_` rather than standard base64's `+` and `/`; a `_scroll_id` is handed straight back to
`_search/scroll`, where a re-minted one answers a recorded continuation with
`search_context_missing_exception` against a cursor the recording had just opened. Neither shape is
published by AWS — these are the domain's own REST API, not the `es` control plane — so substrate's
sixteen characters are a convention it keeps rather than a constraint it meets.
- **A Timestream `QueryId` is hex where a Redshift Data statement ID is a dashed UUID** (#856). Both are
sixteen derived bytes apiece and the difference is the published model, not a preference:
`API_query_Query` constrains `QueryId` to `[a-zA-Z0-9]+`, which **excludes** the hyphen, while
`API_ExecuteStatement` documents `Id` as a UUID and publishes the dashed pattern. Neither constrains a
position, so both are indifferent to the RFC 4122 version and variant bits — which is why deriving
them preserved each rendering instead of quietly setting two nibbles (#671).
- **A stream recorded under a seed replays under the same seed** (#1140). Every seedable outcome in
substrate is written through a control-plane endpoint, and only the AWS path recorded anything — so
a seed never entered the event stream. A replay opens by resetting the whole `StateManager`, and a
Expand Down
30 changes: 27 additions & 3 deletions docs/services.md
Original file line number Diff line number Diff line change
Expand Up @@ -2209,7 +2209,8 @@ Three kinds of value stay random, and one more is still migrating:
- EC2, IAM, STS, SQS, SNS, Lambda, EFS, FSx, Transfer, ECS, Step Functions, EventBridge,
CloudWatch Logs, CloudFront, Service Quotas, API Gateway (v1 and v2), AppSync, Batch, EMR
Serverless, ECR, ELB, Route 53, Cognito (both the user-pool and the identity-pool API), IAM
Identity Center, KMS, ACM, Secrets Manager and WAFv2 identifiers are derived today. A CloudFront
Identity Center, KMS, ACM, Secrets Manager, WAFv2, Athena, Redshift Data, Glue, Timestream,
OpenSearch and QuickSight identifiers are derived today. A CloudFront
distribution, invalidation and origin access control all draw from one generator, so the three
moved together with the origin access control family (#1277). The remaining services are
migrating one family at a time, tracked on #856; until a service moves, its identifiers are still
Expand Down Expand Up @@ -2245,7 +2246,30 @@ blob either way.)
**One service used to mint another's identifiers.** An API Gateway API key's `id` and `value` were
drawn from ACM's certificate-ID generator, which meant a change to ACM's rendering would silently
move API Gateway's. #856 split them into separate minters; both still render the UUID shape they
always did, because neither API publishes a pattern that would decide the question (#671).
always did, because neither API publishes a pattern that would decide the question (#671). The same
split was needed inside QuickSight, where `CreateDataSet` drew its SPICE **ingestion ID** from the
generator that mints the `RequestId` every QuickSight response carries: an ingestion ID is a handle —
a recorded `DescribeIngestion` URL contains it — where a request ID is observed once and never sent
back, so one function serving both meant a change to how a request ID is rendered would have moved
the identifier a recorded path depends on.

**The published alphabet decides the rendering, where there is one.** Two identifiers in the
analytics family are minted from sixteen derived bytes apiece and rendered differently, and the
difference is the API model rather than a preference. Redshift Data documents a statement `Id` as a
UUID and publishes `[a-z0-9]{8}(-[a-z0-9]{4}){3}-[a-z0-9]{12}`, so the hyphenated form is required;
Timestream publishes `QueryId` as `[a-zA-Z0-9]+`, which **excludes** the hyphen, so its bytes are
rendered as 32 unbroken hex characters instead. Neither pattern constrains a position, so both are
indifferent to the RFC 4122 version and variant bits — which is why deriving them preserved the exact
rendering substrate published before, rather than quietly setting two nibbles (#671).

**One service in the family publishes no AWS pattern at all.** An OpenSearch document `_id` and
`_scroll_id` come from the domain's own REST API rather than from the `es`/`opensearch` control
plane, so no AWS reference constrains their shape and substrate's sixteen URL-safe base64 characters
are a convention rather than a published form — narrower than the 20-character Flake ID real
OpenSearch generates, and unchanged by #856. Both are values a caller hands back: a generated
document ID is the path of every later `GET`, `PUT` and `DELETE` of that document, and a scroll
cursor goes straight back to `_search/scroll`, where a re-minted one answers a recorded continuation
with `search_context_missing_exception` against a cursor the recording had just opened.

An ECR image digest is minted rather than computed from the manifest, so it is reproducible across
a replay but is not the SHA-256 of the image it names, and two pushes of identical manifest bytes
Expand Down Expand Up @@ -17165,7 +17189,7 @@ EFS standard storage: $0.30 per GB-month.
| GetJob | |
| DeleteJob | |
| GetJobs | |
| StartJobRun | Returns JobRunId |
| StartJobRun | Returns a `jr_`-prefixed JobRunId — 32 hex characters after the prefix, derived from the request ID (#856). The prefix is the service's own convention, not a published pattern |
| GetJobRun | Transitions to SUCCEEDED after describe |
| GetJobRuns | |

Expand Down
3 changes: 2 additions & 1 deletion docs/testing-guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -338,7 +338,8 @@ tests above are built on caller-chosen bucket and key names instead.
Two caveats. **Not every service's identifiers are derived yet.** EC2, IAM, STS, SQS, SNS,
Lambda, EFS, FSx, Transfer, ECS, Step Functions, EventBridge, CloudWatch Logs, CloudFront,
Service Quotas, API Gateway (v1 and v2), AppSync, Batch, EMR Serverless, ECR, ELB, Route 53,
Cognito (both APIs), IAM Identity Center, KMS, ACM, Secrets Manager and WAFv2
Cognito (both APIs), IAM Identity Center, KMS, ACM, Secrets Manager, WAFv2, Athena,
Redshift Data, Glue, Timestream, OpenSearch and QuickSight
are; the rest are migrating one family at a time, and until a service moves, a
replay of a stream creating one of its resources still diverges. And a recording made against an
**unfrozen** clock can still diverge on a `state_hash_after` even when every identifier
Expand Down
25 changes: 18 additions & 7 deletions emulator/athena_plugin.go
Original file line number Diff line number Diff line change
Expand Up @@ -2,7 +2,6 @@ package emulator

import (
"context"
"crypto/rand"
"encoding/json"
"fmt"
"net/http"
Expand Down Expand Up @@ -154,7 +153,7 @@ func (p *AthenaPlugin) startQueryExecution(ctx *RequestContext, req *AWSRequest)
return nil, &AWSError{Code: "InvalidRequestException", Message: "invalid request body", HTTPStatus: http.StatusBadRequest}
}

qID := generateAthenaQueryID()
qID := generateAthenaQueryID(ctx.IDs)
now := float64(p.tc.Now().UnixNano()) / 1e9
outputLoc := ""
if body.ResultConfiguration != nil {
Expand Down Expand Up @@ -688,11 +687,23 @@ func athenaLoadStringIndex(ctx context.Context, state StateManager, key string)
return ids
}

// generateAthenaQueryID generates a UUID-formatted query execution ID.
func generateAthenaQueryID() string {
b := make([]byte, 16)
_, _ = rand.Read(b)
return fmt.Sprintf("%x-%x-%x-%x-%x", b[0:4], b[4:6], b[6:8], b[8:10], b[10:16])
// generateAthenaQueryID mints a query execution ID from m, in the UUID shape AWS's own examples
// write.
//
// Deriving it is what makes a recorded Athena session replay at all: every operation after
// `StartQueryExecution` addresses the query by this id — `GetQueryExecution`, `GetQueryResults`,
// `StopQueryExecution` — so a re-minted one turned each of them into an `InvalidRequestException`
// against a query the recording had just created. Athena is the service in this family where a
// single identifier gates the most recorded follow-on calls, which is the poll loop substrate exists
// to let a consumer test.
//
// [IDMint.HexUUID] rather than [IDMint.UUID]: the crypto/rand form reshaped sixteen raw bytes
// without setting the RFC 4122 version and variant nibbles, and #856 does not change which bytes a
// caller sees. API_StartQueryExecution gives `QueryExecutionId` a length of 1–128 and the pattern
// `\S+`, which admits any non-whitespace string at all, so nothing in the API model distinguishes
// the two renderings (#671).
func generateAthenaQueryID(m *IDMint) string {
return m.HexUUID()
}

// athenaJSONResponse serializes v to JSON and returns an AWSResponse.
Expand Down
2 changes: 1 addition & 1 deletion emulator/ec2_types.go
Original file line number Diff line number Diff line change
Expand Up @@ -495,7 +495,7 @@ func generateAssociationID(m *IDMint) string {
// flag day: a caller moves by taking a mint and calling [IDMint.Hex] with the same width.
// EC2's own ids no longer come through here.
//
// TODO(#856): 15 draw sites remain on crypto/rand, tiered by service family on the issue;
// TODO(#856): 11 draw sites remain on crypto/rand, tiered by service family on the issue;
// delete this function when the last caller moves.
func randomHex(n int) string {
b := make([]byte, n)
Expand Down
18 changes: 17 additions & 1 deletion emulator/glue_plugin.go
Original file line number Diff line number Diff line change
Expand Up @@ -861,7 +861,7 @@ func (p *GluePlugin) startJobRun(reqCtx *RequestContext, req *AWSRequest) (*AWSR
}

now := p.tc.Now()
runID := "jr_" + randomHex(16)
runID := glueJobRunID(reqCtx.IDs)
run := GlueJobRun{
ID: runID,
JobName: input.JobName,
Expand Down Expand Up @@ -1124,6 +1124,22 @@ func (p *GluePlugin) loadGlueTags(goCtx context.Context, ns, key string) (map[st
return nil, nil
}

// glueJobRunID mints a job-run ID from m — the `jr_` prefix real Glue uses, followed by 32 lowercase
// hex characters, which is what [randomHex] produced at the site this replaces.
//
// Deriving it is what lets a recorded Glue run be polled: `GetJobRun` and `BatchStopJobRun` address
// the run by this ID, so a re-minted one answered a recorded poll with EntityNotFoundException against
// a run the recording had just started — and a consumer's wait-for-SUCCEEDED loop is the whole reason
// StartJobRun is worth emulating.
//
// The prefix is not in the API model. API_StartJobRun publishes `JobRunId` at 1–255 characters against
// a pattern that admits nearly any text, so the shape is substrate's to keep rather than #856's to
// revisit (#671); `jr_` is kept because it is what the service's own run IDs look like and what a
// consumer's log-scraping or ID-shape assertion would have recorded.
func glueJobRunID(m *IDMint) string {
return "jr_" + m.Hex(16)
}

// glueJSONResponse serializes v to JSON and returns an AWSResponse.
func glueJSONResponse(status int, v interface{}) (*AWSResponse, error) {
body, err := json.Marshal(v)
Expand Down
14 changes: 13 additions & 1 deletion emulator/ids.go
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ import (
// already derived from its inputs: a public IP from its instance id, a secret's ARN from its
// name, CloudFormation's stack UUIDs from account and region.
//
// TODO(#856): 15 draw sites remain on crypto/rand, tiered by service family on the issue.
// TODO(#856): 11 draw sites remain on crypto/rand, tiered by service family on the issue.

// IDMint mints the identifiers one request publishes, derived from that request's own id so
// that replaying the request mints the same ones.
Expand Down Expand Up @@ -180,6 +180,18 @@ func (m *IDMint) Base64(n int) string {
return base64.StdEncoding.EncodeToString(m.bytes(n))
}

// Base64URL returns n bytes in unpadded URL-safe base64, the encoding OpenSearch generates a
// document id in.
//
// Separate from [IDMint.Base64] because the two differ in the two characters that matter here: a
// document id reaches an OpenSearch caller inside a URL path, so `-` and `_` are the alphabet and
// `+` and `/` are not. Unpadded because a generated id carries no `=`, and at OpenSearch's twelve
// bytes there would be none to carry — 12 divides by 3 — so the distinction only shows if a later
// caller asks for a width that does not.
func (m *IDMint) Base64URL(n int) string {
return base64.RawURLEncoding.EncodeToString(m.bytes(n))
}

// HexUUID returns sixteen bytes in UUID *shape* — 8-4-4-4-12 lowercase hex — without the RFC
// 4122 version and variant bits [IDMint.UUID] sets.
//
Expand Down
Loading
Loading