diff --git a/CHANGELOG.md b/CHANGELOG.md index 3e59639c..4a868ea9 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -325,6 +325,40 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 bytes are the `Plaintext`, and the same bytes are wrapped into the `CiphertextBlob`. A recorded `Decrypt` was never affected — substrate's ciphertext carries its own plaintext, so a replayed decrypt answers from the recorded blob regardless — so the break was in the create, not the read. +- **The analytics family of draw sites is derived** (#856). Seven generators across six services move + onto `IDMint`: an Athena query execution ID, a Redshift Data statement ID, a Glue job-run ID, a + Timestream `QueryId`, an OpenSearch document `_id` and `_scroll_id`, and a QuickSight ingestion ID + and response `RequestId`. Every rendering is byte-for-byte the one the `crypto/rand` version + produced, so an identifier a previous substrate recorded is still the shape this one mints. 11 draw + sites remain on `crypto/rand`. +- **A recorded Athena poll loop replays against the query it started** (#856). An analytics identifier + names a submission rather than a resource, so unlike a bucket or a job there is no caller-chosen name + to fall back on: `GetQueryExecution`, `GetQueryResults` and `StopQueryExecution` all key on the one ID + `StartQueryExecution` returned, and a re-minted one answered every recorded poll with + `InvalidRequestException` against a query the recording had just created. Redshift Data breaks the + same way with `ResourceNotFoundException` and Glue with `EntityNotFoundException`. Reverting the + Athena minter alone to confirm the family's replay assertion is not vacuous produces 38 differences + and two refused reads out of a 19-request stream. +- **A QuickSight ingestion ID is no longer minted by the request-ID generator** (#856). `CreateDataSet` + drew the SPICE ingestion ID it reports from the same function that mints the `RequestId` every + QuickSight response carries — the second instance of the cross-service borrow the ACM/API Gateway + split fixed. An ingestion ID is a handle a recorded `DescribeIngestion` URL contains, where a request + ID is observed once and never sent back, so one function serving both meant a change to how a request + ID renders would have moved the identifier a recorded path depends on. Both now have their own + minter and the rendering is unchanged. +- **`IDMint.Base64URL`, for the one identifier that travels in a URL path** (#856). An OpenSearch + document `_id` is the path of every later `GET`, `PUT` and `DELETE` of that document, so its alphabet + is `-` and `_` rather than standard base64's `+` and `/`; a `_scroll_id` is handed straight back to + `_search/scroll`, where a re-minted one answers a recorded continuation with + `search_context_missing_exception` against a cursor the recording had just opened. Neither shape is + published by AWS — these are the domain's own REST API, not the `es` control plane — so substrate's + sixteen characters are a convention it keeps rather than a constraint it meets. +- **A Timestream `QueryId` is hex where a Redshift Data statement ID is a dashed UUID** (#856). Both are + sixteen derived bytes apiece and the difference is the published model, not a preference: + `API_query_Query` constrains `QueryId` to `[a-zA-Z0-9]+`, which **excludes** the hyphen, while + `API_ExecuteStatement` documents `Id` as a UUID and publishes the dashed pattern. Neither constrains a + position, so both are indifferent to the RFC 4122 version and variant bits — which is why deriving + them preserved each rendering instead of quietly setting two nibbles (#671). - **A stream recorded under a seed replays under the same seed** (#1140). Every seedable outcome in substrate is written through a control-plane endpoint, and only the AWS path recorded anything — so a seed never entered the event stream. A replay opens by resetting the whole `StateManager`, and a diff --git a/docs/services.md b/docs/services.md index 1828d257..c49c624e 100644 --- a/docs/services.md +++ b/docs/services.md @@ -2209,7 +2209,8 @@ Three kinds of value stay random, and one more is still migrating: - EC2, IAM, STS, SQS, SNS, Lambda, EFS, FSx, Transfer, ECS, Step Functions, EventBridge, CloudWatch Logs, CloudFront, Service Quotas, API Gateway (v1 and v2), AppSync, Batch, EMR Serverless, ECR, ELB, Route 53, Cognito (both the user-pool and the identity-pool API), IAM - Identity Center, KMS, ACM, Secrets Manager and WAFv2 identifiers are derived today. A CloudFront + Identity Center, KMS, ACM, Secrets Manager, WAFv2, Athena, Redshift Data, Glue, Timestream, + OpenSearch and QuickSight identifiers are derived today. A CloudFront distribution, invalidation and origin access control all draw from one generator, so the three moved together with the origin access control family (#1277). The remaining services are migrating one family at a time, tracked on #856; until a service moves, its identifiers are still @@ -2245,7 +2246,30 @@ blob either way.) **One service used to mint another's identifiers.** An API Gateway API key's `id` and `value` were drawn from ACM's certificate-ID generator, which meant a change to ACM's rendering would silently move API Gateway's. #856 split them into separate minters; both still render the UUID shape they -always did, because neither API publishes a pattern that would decide the question (#671). +always did, because neither API publishes a pattern that would decide the question (#671). The same +split was needed inside QuickSight, where `CreateDataSet` drew its SPICE **ingestion ID** from the +generator that mints the `RequestId` every QuickSight response carries: an ingestion ID is a handle — +a recorded `DescribeIngestion` URL contains it — where a request ID is observed once and never sent +back, so one function serving both meant a change to how a request ID is rendered would have moved +the identifier a recorded path depends on. + +**The published alphabet decides the rendering, where there is one.** Two identifiers in the +analytics family are minted from sixteen derived bytes apiece and rendered differently, and the +difference is the API model rather than a preference. Redshift Data documents a statement `Id` as a +UUID and publishes `[a-z0-9]{8}(-[a-z0-9]{4}){3}-[a-z0-9]{12}`, so the hyphenated form is required; +Timestream publishes `QueryId` as `[a-zA-Z0-9]+`, which **excludes** the hyphen, so its bytes are +rendered as 32 unbroken hex characters instead. Neither pattern constrains a position, so both are +indifferent to the RFC 4122 version and variant bits — which is why deriving them preserved the exact +rendering substrate published before, rather than quietly setting two nibbles (#671). + +**One service in the family publishes no AWS pattern at all.** An OpenSearch document `_id` and +`_scroll_id` come from the domain's own REST API rather than from the `es`/`opensearch` control +plane, so no AWS reference constrains their shape and substrate's sixteen URL-safe base64 characters +are a convention rather than a published form — narrower than the 20-character Flake ID real +OpenSearch generates, and unchanged by #856. Both are values a caller hands back: a generated +document ID is the path of every later `GET`, `PUT` and `DELETE` of that document, and a scroll +cursor goes straight back to `_search/scroll`, where a re-minted one answers a recorded continuation +with `search_context_missing_exception` against a cursor the recording had just opened. An ECR image digest is minted rather than computed from the manifest, so it is reproducible across a replay but is not the SHA-256 of the image it names, and two pushes of identical manifest bytes @@ -17165,7 +17189,7 @@ EFS standard storage: $0.30 per GB-month. | GetJob | | | DeleteJob | | | GetJobs | | -| StartJobRun | Returns JobRunId | +| StartJobRun | Returns a `jr_`-prefixed JobRunId — 32 hex characters after the prefix, derived from the request ID (#856). The prefix is the service's own convention, not a published pattern | | GetJobRun | Transitions to SUCCEEDED after describe | | GetJobRuns | | diff --git a/docs/testing-guide.md b/docs/testing-guide.md index 5badd347..ecc835de 100644 --- a/docs/testing-guide.md +++ b/docs/testing-guide.md @@ -338,7 +338,8 @@ tests above are built on caller-chosen bucket and key names instead. Two caveats. **Not every service's identifiers are derived yet.** EC2, IAM, STS, SQS, SNS, Lambda, EFS, FSx, Transfer, ECS, Step Functions, EventBridge, CloudWatch Logs, CloudFront, Service Quotas, API Gateway (v1 and v2), AppSync, Batch, EMR Serverless, ECR, ELB, Route 53, -Cognito (both APIs), IAM Identity Center, KMS, ACM, Secrets Manager and WAFv2 +Cognito (both APIs), IAM Identity Center, KMS, ACM, Secrets Manager, WAFv2, Athena, +Redshift Data, Glue, Timestream, OpenSearch and QuickSight are; the rest are migrating one family at a time, and until a service moves, a replay of a stream creating one of its resources still diverges. And a recording made against an **unfrozen** clock can still diverge on a `state_hash_after` even when every identifier diff --git a/emulator/athena_plugin.go b/emulator/athena_plugin.go index 14d997c6..4450a308 100644 --- a/emulator/athena_plugin.go +++ b/emulator/athena_plugin.go @@ -2,7 +2,6 @@ package emulator import ( "context" - "crypto/rand" "encoding/json" "fmt" "net/http" @@ -154,7 +153,7 @@ func (p *AthenaPlugin) startQueryExecution(ctx *RequestContext, req *AWSRequest) return nil, &AWSError{Code: "InvalidRequestException", Message: "invalid request body", HTTPStatus: http.StatusBadRequest} } - qID := generateAthenaQueryID() + qID := generateAthenaQueryID(ctx.IDs) now := float64(p.tc.Now().UnixNano()) / 1e9 outputLoc := "" if body.ResultConfiguration != nil { @@ -688,11 +687,23 @@ func athenaLoadStringIndex(ctx context.Context, state StateManager, key string) return ids } -// generateAthenaQueryID generates a UUID-formatted query execution ID. -func generateAthenaQueryID() string { - b := make([]byte, 16) - _, _ = rand.Read(b) - return fmt.Sprintf("%x-%x-%x-%x-%x", b[0:4], b[4:6], b[6:8], b[8:10], b[10:16]) +// generateAthenaQueryID mints a query execution ID from m, in the UUID shape AWS's own examples +// write. +// +// Deriving it is what makes a recorded Athena session replay at all: every operation after +// `StartQueryExecution` addresses the query by this id — `GetQueryExecution`, `GetQueryResults`, +// `StopQueryExecution` — so a re-minted one turned each of them into an `InvalidRequestException` +// against a query the recording had just created. Athena is the service in this family where a +// single identifier gates the most recorded follow-on calls, which is the poll loop substrate exists +// to let a consumer test. +// +// [IDMint.HexUUID] rather than [IDMint.UUID]: the crypto/rand form reshaped sixteen raw bytes +// without setting the RFC 4122 version and variant nibbles, and #856 does not change which bytes a +// caller sees. API_StartQueryExecution gives `QueryExecutionId` a length of 1–128 and the pattern +// `\S+`, which admits any non-whitespace string at all, so nothing in the API model distinguishes +// the two renderings (#671). +func generateAthenaQueryID(m *IDMint) string { + return m.HexUUID() } // athenaJSONResponse serializes v to JSON and returns an AWSResponse. diff --git a/emulator/ec2_types.go b/emulator/ec2_types.go index d1032595..1ead4578 100644 --- a/emulator/ec2_types.go +++ b/emulator/ec2_types.go @@ -495,7 +495,7 @@ func generateAssociationID(m *IDMint) string { // flag day: a caller moves by taking a mint and calling [IDMint.Hex] with the same width. // EC2's own ids no longer come through here. // -// TODO(#856): 15 draw sites remain on crypto/rand, tiered by service family on the issue; +// TODO(#856): 11 draw sites remain on crypto/rand, tiered by service family on the issue; // delete this function when the last caller moves. func randomHex(n int) string { b := make([]byte, n) diff --git a/emulator/glue_plugin.go b/emulator/glue_plugin.go index f182ce8c..532e5d41 100644 --- a/emulator/glue_plugin.go +++ b/emulator/glue_plugin.go @@ -861,7 +861,7 @@ func (p *GluePlugin) startJobRun(reqCtx *RequestContext, req *AWSRequest) (*AWSR } now := p.tc.Now() - runID := "jr_" + randomHex(16) + runID := glueJobRunID(reqCtx.IDs) run := GlueJobRun{ ID: runID, JobName: input.JobName, @@ -1124,6 +1124,22 @@ func (p *GluePlugin) loadGlueTags(goCtx context.Context, ns, key string) (map[st return nil, nil } +// glueJobRunID mints a job-run ID from m — the `jr_` prefix real Glue uses, followed by 32 lowercase +// hex characters, which is what [randomHex] produced at the site this replaces. +// +// Deriving it is what lets a recorded Glue run be polled: `GetJobRun` and `BatchStopJobRun` address +// the run by this ID, so a re-minted one answered a recorded poll with EntityNotFoundException against +// a run the recording had just started — and a consumer's wait-for-SUCCEEDED loop is the whole reason +// StartJobRun is worth emulating. +// +// The prefix is not in the API model. API_StartJobRun publishes `JobRunId` at 1–255 characters against +// a pattern that admits nearly any text, so the shape is substrate's to keep rather than #856's to +// revisit (#671); `jr_` is kept because it is what the service's own run IDs look like and what a +// consumer's log-scraping or ID-shape assertion would have recorded. +func glueJobRunID(m *IDMint) string { + return "jr_" + m.Hex(16) +} + // glueJSONResponse serializes v to JSON and returns an AWSResponse. func glueJSONResponse(status int, v interface{}) (*AWSResponse, error) { body, err := json.Marshal(v) diff --git a/emulator/ids.go b/emulator/ids.go index 83fe0500..60b5ad06 100644 --- a/emulator/ids.go +++ b/emulator/ids.go @@ -61,7 +61,7 @@ import ( // already derived from its inputs: a public IP from its instance id, a secret's ARN from its // name, CloudFormation's stack UUIDs from account and region. // -// TODO(#856): 15 draw sites remain on crypto/rand, tiered by service family on the issue. +// TODO(#856): 11 draw sites remain on crypto/rand, tiered by service family on the issue. // IDMint mints the identifiers one request publishes, derived from that request's own id so // that replaying the request mints the same ones. @@ -180,6 +180,18 @@ func (m *IDMint) Base64(n int) string { return base64.StdEncoding.EncodeToString(m.bytes(n)) } +// Base64URL returns n bytes in unpadded URL-safe base64, the encoding OpenSearch generates a +// document id in. +// +// Separate from [IDMint.Base64] because the two differ in the two characters that matter here: a +// document id reaches an OpenSearch caller inside a URL path, so `-` and `_` are the alphabet and +// `+` and `/` are not. Unpadded because a generated id carries no `=`, and at OpenSearch's twelve +// bytes there would be none to carry — 12 divides by 3 — so the distinction only shows if a later +// caller asks for a width that does not. +func (m *IDMint) Base64URL(n int) string { + return base64.RawURLEncoding.EncodeToString(m.bytes(n)) +} + // HexUUID returns sixteen bytes in UUID *shape* — 8-4-4-4-12 lowercase hex — without the RFC // 4122 version and variant bits [IDMint.UUID] sets. // diff --git a/emulator/ids_test.go b/emulator/ids_test.go index 990689b8..dd221dcc 100644 --- a/emulator/ids_test.go +++ b/emulator/ids_test.go @@ -1482,3 +1482,317 @@ func idsJSONTargetCall(t *testing.T, ts *emulator.TestServer, host, target strin require.Less(t, resp.StatusCode, 300, "%s %s: %s", host, target, out) return out } + +// Tier 5 of #856: the analytics family — Athena, Redshift Data, Glue, Timestream, OpenSearch and +// QuickSight. +// +// What this tier adds to the property the tiers above assert is the *shape of the call* the +// identifier gates. An analytics identifier names a submission rather than a resource: a query +// execution, a statement, a job run, a scroll cursor. Nothing else addresses it — there is no name +// to fall back on the way an S3 bucket or a Glue job has one — so the minted value is the only +// handle a consumer's poll loop holds, which is why an underived one broke the loop outright rather +// than merely reporting a different string. Athena is the extreme: `GetQueryExecution`, +// `GetQueryResults` and `StopQueryExecution` all key on the one id `StartQueryExecution` returned. + +// TestIDs_OneBulkIndexMintsOneIDPerDocument is the ordinal's test for this tier, on the one request +// in the tree that draws an unbounded number of times: a `_bulk` body whose actions carry no `_id` +// mints one per document. +// +// A mint that ignored its ordinal would answer one id for all three, and here that is worse than a +// cosmetic collision — the documents are keyed by it, so the second and third would overwrite the +// first and one document would be left where three were indexed. +func TestIDs_OneBulkIndexMintsOneIDPerDocument(t *testing.T) { + t.Parallel() + ts := emulator.StartTestServer(t) + ts.FreezeTime() + + body := strings.Join([]string{ + `{"index":{}}`, `{"n":1}`, + `{"index":{}}`, `{"n":2}`, + `{"index":{}}`, `{"n":3}`, + }, "\n") + "\n" + + var bulk struct { + Items []map[string]struct { + ID string `json:"_id"` + Result string `json:"result"` + } `json:"items"` + } + require.NoError(t, json.Unmarshal(idsNDJSONCall(t, ts, idsOpenSearchHost, + "/ids-tier5-bulk/_bulk", body), &bulk)) + require.Len(t, bulk.Items, 3, "one item per action") + + seen := make(map[string]bool, 3) + for _, item := range bulk.Items { + for op, res := range item { + require.Equal(t, "index", op) + require.Len(t, res.ID, 16, "a generated document id is sixteen URL-safe base64 characters") + assert.Equal(t, "created", res.Result, + "each document is new, which it would not be if two shared an id") + seen[res.ID] = true + } + } + assert.Len(t, seen, 3, "three documents indexed in one request draw three ids") +} + +// TestReplay_TheAnalyticsFamilyReplaysWithTheIdentifiersItMinted is the wire-level assertion for +// tier 5, and the same claim the four tiers above make for their families: a stream whose later +// requests *name* what its earlier ones minted replays with no differences and reaches the recorded +// state. +// +// Every service contributes a submission followed by the read that can only reach it through the +// minted handle — a GetQueryExecution and a GetQueryResults on an Athena query id, a +// DescribeStatement and a GetStatementResult on a Redshift Data statement id, a GetJobRun on a Glue +// run id, a document GET and a scroll continuation on OpenSearch's two kinds of generated id, and a +// DescribeIngestion whose URL path is the ingestion id CreateDataSet minted. A re-minted value +// refuses each of those by name: InvalidRequestException for Athena, ResourceNotFoundException for +// Redshift Data, EntityNotFoundException for Glue, a 404 for a document and an expired-context +// refusal for a scroll. Timestream is the exception and is here for the other half of the property: +// its QueryId reaches no later call, so what a fresh draw costs there is a body difference rather +// than a refusal — which is still a difference, and still enough to make a recorded Query fail to +// replay. +func TestReplay_TheAnalyticsFamilyReplaysWithTheIdentifiersItMinted(t *testing.T) { + t.Parallel() + ts := emulator.StartTestServer(t, + emulator.WithRecordedBodies(), emulator.WithRecordedStateHashes()) + require.True(t, ts.Store().RecordsStateHashes(), "precondition: state_hash_after is compared") + + idsRecordAnalyticsSubmissions(t, ts) + + results, err := replayEngineFor(ts, emulator.ReplayConfig{ValidateState: true}). + Replay(t.Context(), replayStreamID) + require.NoError(t, err) + + assert.Positive(t, results.TotalEvents, "the stream has to contain the submissions") + assert.Equal(t, results.TotalEvents, results.SuccessEvents, + "every recorded request is re-executed and answers") + assert.Empty(t, results.Differences, + "an analytics submission's identifier replays as the one recorded: %s", + replayDifferenceSummary(results)) + assert.True(t, results.StateValid, + "and the state it reaches is the recorded state: %v", results.StateErrors) +} + +// idsRecordAnalyticsSubmissions records the tier-5 stream, one interlocked group per service. +func idsRecordAnalyticsSubmissions(t *testing.T, ts *emulator.TestServer) { + t.Helper() + + // Frozen for the reason [idsRecordInterlockedCreates] gives: a replay pins the clock to the + // recorded event's timestamp, so a handler stamping a record off a live clock diverges in the + // state hash for a reason that has nothing to do with an identifier. + ts.FreezeTime() + + idsRecordAthena(t, ts) + idsRecordRedshiftData(t, ts) + idsRecordGlue(t, ts) + idsRecordTimestream(t, ts) + idsRecordOpenSearch(t, ts) + idsRecordQuickSight(t, ts) +} + +// idsRecordAthena records a query execution and the two reads that address it by the id it minted. +func idsRecordAthena(t *testing.T, ts *emulator.TestServer) { + t.Helper() + + var started struct { + QueryExecutionID string `json:"QueryExecutionId"` + } + require.NoError(t, json.Unmarshal(idsJSONTargetCall(t, ts, idsAthenaHost, + "AmazonAthena.StartQueryExecution", map[string]any{ + "QueryString": "SELECT 1", + "ResultConfiguration": map[string]any{ + "OutputLocation": "s3://ids-tier5-results/", + }, + }), &started)) + idsRequireHexUUID(t, started.QueryExecutionID, "an Athena query execution id") + + idsJSONTargetCall(t, ts, idsAthenaHost, "AmazonAthena.GetQueryExecution", + map[string]any{"QueryExecutionId": started.QueryExecutionID}) + idsJSONTargetCall(t, ts, idsAthenaHost, "AmazonAthena.GetQueryResults", + map[string]any{"QueryExecutionId": started.QueryExecutionID}) +} + +// idsRecordRedshiftData records a statement and the two reads that address it by its id. +func idsRecordRedshiftData(t *testing.T, ts *emulator.TestServer) { + t.Helper() + + var executed struct { + ID string `json:"Id"` + } + require.NoError(t, json.Unmarshal(idsJSONTargetCall(t, ts, idsRedshiftDataHost, + "RedshiftData.ExecuteStatement", map[string]any{ + "WorkgroupName": "ids-tier5", "Database": "dev", "Sql": "SELECT 1", + }), &executed)) + idsRequireHexUUID(t, executed.ID, "a Redshift Data statement id") + + idsJSONTargetCall(t, ts, idsRedshiftDataHost, "RedshiftData.DescribeStatement", + map[string]any{"Id": executed.ID}) + idsJSONTargetCall(t, ts, idsRedshiftDataHost, "RedshiftData.GetStatementResult", + map[string]any{"Id": executed.ID}) +} + +// idsRecordGlue records a job, a run of it, and the GetJobRun that names the run id. +// +// The job is addressed by the caller's own name throughout, so the run id is the only value in the +// group substrate mints — which is what makes GetJobRun a read of the mint and not of the request. +func idsRecordGlue(t *testing.T, ts *emulator.TestServer) { + t.Helper() + + idsJSONTargetCall(t, ts, idsGlueHost, "AWSGlue.CreateJob", map[string]any{ + "Name": "ids-tier5-job", "Role": "arn:aws:iam::123456789012:role/GlueRole", + "Command": map[string]any{"Name": "glueetl", "ScriptLocation": "s3://ids-tier5/etl.py"}, + }) + + var run struct { + JobRunID string `json:"JobRunId"` + } + require.NoError(t, json.Unmarshal(idsJSONTargetCall(t, ts, idsGlueHost, "AWSGlue.StartJobRun", + map[string]any{"JobName": "ids-tier5-job"}), &run)) + require.True(t, strings.HasPrefix(run.JobRunID, "jr_"), + "a Glue job run id keeps its jr_ prefix, got %q", run.JobRunID) + require.Len(t, run.JobRunID, len("jr_")+32, "and 32 hex characters after it") + + idsJSONTargetCall(t, ts, idsGlueHost, "AWSGlue.GetJobRun", + map[string]any{"JobName": "ids-tier5-job", "RunId": run.JobRunID}) +} + +// idsRecordTimestream records one Query, whose QueryId nothing addresses. +// +// It is in the stream for what [timestreamQueryID]'s comment records: the id is a response member no +// later call names, so it can only be observed by being compared against the recording — which is +// exactly what a replay does. +func idsRecordTimestream(t *testing.T, ts *emulator.TestServer) { + t.Helper() + + var queried struct { + QueryID string `json:"QueryId"` + } + require.NoError(t, json.Unmarshal(idsJSONTargetCall(t, ts, idsTimestreamQueryHost, + "Timestream_20181101.Query", + map[string]any{"QueryString": "SELECT 1"}), &queried)) + require.Len(t, queried.QueryID, 32, + "a Timestream QueryId is 32 hex characters, the alphabet [a-zA-Z0-9]+ admits") + require.NotContains(t, queried.QueryID, "-", + "and not a UUID: the published pattern excludes the hyphens one would carry") +} + +// idsRecordOpenSearch records two documents indexed without an `_id`, a read of the first by the id +// the index minted, and a scrolled search followed by its continuation. +// +// Both kinds of generated value are here because they break differently: a re-minted document id +// answers the recorded GET with a 404, and a re-minted scroll id answers the recorded continuation +// with search_context_missing_exception against a cursor the recording had just opened. +func idsRecordOpenSearch(t *testing.T, ts *emulator.TestServer) { + t.Helper() + + const index = "ids-tier5-index" + + first := idsOpenSearchIndexDoc(t, ts, index, map[string]any{"n": 1}) + idsOpenSearchIndexDoc(t, ts, index, map[string]any{"n": 2}) + + idsRESTCall(t, ts, idsOpenSearchHost, http.MethodGet, "/"+index+"/_doc/"+first, nil) + + // A page of one over two documents, so the scroll has a second page to continue into; the TTL + // travels in the body rather than in a `?scroll=` parameter to keep the group a test about the + // identifier and not about query-string round-tripping. + var searched struct { + ScrollID string `json:"_scroll_id"` + } + require.NoError(t, json.Unmarshal(idsRESTCall(t, ts, idsOpenSearchHost, http.MethodPost, + "/"+index+"/_search", map[string]any{"size": 1, "scroll": "1m"}), &searched)) + require.Len(t, searched.ScrollID, 16, "a scroll id is minted in the same shape a document id is") + require.NotEqual(t, first, searched.ScrollID, "and is a separate draw") + + idsRESTCall(t, ts, idsOpenSearchHost, http.MethodPost, "/_search/scroll", + map[string]any{"scroll_id": searched.ScrollID}) +} + +// idsOpenSearchIndexDoc indexes one document with no `_id` and returns the id OpenSearch minted. +func idsOpenSearchIndexDoc(t *testing.T, ts *emulator.TestServer, index string, doc any) string { + t.Helper() + + var indexed struct { + ID string `json:"_id"` + Result string `json:"result"` + } + require.NoError(t, json.Unmarshal(idsRESTCall(t, ts, idsOpenSearchHost, http.MethodPost, + "/"+index+"/_doc", doc), &indexed)) + require.Equal(t, "created", indexed.Result) + require.Len(t, indexed.ID, 16, "a generated document id is sixteen URL-safe base64 characters") + return indexed.ID +} + +// idsRecordQuickSight records a data source, a read of it, a data set, and the DescribeIngestion +// whose path is the ingestion id CreateDataSet minted. +// +// The data source and data set ids are the caller's own, which is why the ingestion id is the value +// under test: it is the one identifier in the group substrate chooses, and the only one a replay +// could get wrong. +func idsRecordQuickSight(t *testing.T, ts *emulator.TestServer) { + t.Helper() + + const base = "/accounts/123456789012/" + + idsRESTCall(t, ts, idsQuickSightHost, http.MethodPost, base+"data-sources", map[string]any{ + "DataSourceId": "ids-tier5-source", "Name": "ids-tier5", "Type": "ATHENA", + }) + idsRESTCall(t, ts, idsQuickSightHost, http.MethodGet, base+"data-sources/ids-tier5-source", nil) + + var set struct { + IngestionID string `json:"IngestionId"` + RequestID string `json:"RequestId"` + } + require.NoError(t, json.Unmarshal(idsRESTCall(t, ts, idsQuickSightHost, http.MethodPost, + base+"data-sets", map[string]any{ + "DataSetId": "ids-tier5-set", "Name": "ids-tier5", + }), &set)) + idsRequireHexUUID(t, set.IngestionID, "a QuickSight ingestion id") + assert.NotEqual(t, set.IngestionID, set.RequestID, + "one CreateDataSet draws the ingestion id and the response's request id separately") + + idsRESTCall(t, ts, idsQuickSightHost, http.MethodGet, + base+"data-sets/ids-tier5-set/ingestions/"+set.IngestionID, nil) +} + +// idsRequireHexUUID requires that id is in the 8-4-4-4-12 lowercase-hex shape [IDMint.HexUUID] +// renders — the form five of this tier's six identifiers publish. +func idsRequireHexUUID(t *testing.T, id, what string) { + t.Helper() + + require.Len(t, id, 36, "%s is 36 characters, got %q", what, id) + for i, group := range strings.Split(id, "-") { + require.Len(t, group, []int{8, 4, 4, 4, 12}[i], "%s group %d of %q", what, i, id) + require.Regexp(t, "^[0-9a-f]+$", group, "%s is lowercase hex: %q", what, id) + } +} + +// The hosts the tier-5 services are addressed at. Timestream is two endpoints and this is the query +// one, because Query is the operation that mints. +const ( + idsAthenaHost = "athena.us-east-1.amazonaws.com" + idsRedshiftDataHost = "redshift-data.us-east-1.amazonaws.com" + idsGlueHost = "glue.us-east-1.amazonaws.com" + idsTimestreamQueryHost = "query.timestream.us-east-1.amazonaws.com" + idsOpenSearchHost = "search-ids-tier5.us-east-1.es.amazonaws.com" + idsQuickSightHost = "quicksight.us-east-1.amazonaws.com" +) + +// idsNDJSONCall issues one newline-delimited-JSON request, which is what OpenSearch's `_bulk` takes +// and the only body in these tests that is not a single JSON document. +func idsNDJSONCall(t *testing.T, ts *emulator.TestServer, host, path, body string) []byte { + t.Helper() + + req, err := http.NewRequestWithContext(t.Context(), http.MethodPost, ts.URL+path, + strings.NewReader(body)) + require.NoError(t, err) + req.Host = host + req.Header.Set("Content-Type", "application/x-ndjson") + + resp, err := http.DefaultClient.Do(req) + require.NoError(t, err) + out, err := io.ReadAll(resp.Body) + require.NoError(t, err) + require.NoError(t, resp.Body.Close()) + require.Less(t, resp.StatusCode, 300, "%s %s: %s", host, path, out) + return out +} diff --git a/emulator/opensearch_plugin.go b/emulator/opensearch_plugin.go index 16f18995..3229beef 100644 --- a/emulator/opensearch_plugin.go +++ b/emulator/opensearch_plugin.go @@ -2,8 +2,6 @@ package emulator import ( "context" - "crypto/rand" - "encoding/base64" "encoding/json" "fmt" "net/http" @@ -202,7 +200,7 @@ func (p *OpenSearchPlugin) putMapping(_ *RequestContext, index string) (*AWSResp func (p *OpenSearchPlugin) indexDocument(ctx *RequestContext, req *AWSRequest, index, docID string) (*AWSResponse, error) { goCtx := context.Background() if docID == "" { - docID = generateOpenSearchID() + docID = generateOpenSearchID(ctx.IDs) } // Ensure the index exists (auto-create). indexKey := "index:" + index @@ -304,7 +302,7 @@ func (p *OpenSearchPlugin) bulk(ctx *RequestContext, req *AWSRequest, defaultInd body := []byte(lines[i]) i++ if docID == "" { - docID = generateOpenSearchID() + docID = generateOpenSearchID(ctx.IDs) } docKey := "doc:" + idx + "/" + docID isNew := true @@ -340,7 +338,6 @@ func (p *OpenSearchPlugin) bulk(ctx *RequestContext, req *AWSRequest, defaultInd } } } - _ = ctx return openSearchOK(map[string]interface{}{ "took": 1, "errors": false, @@ -383,7 +380,7 @@ func (p *OpenSearchPlugin) search(ctx *RequestContext, req *AWSRequest, index st scrollTTL = body.Scroll } if scrollTTL != "" { - scrollID := generateOpenSearchID() + scrollID := generateOpenSearchID(ctx.IDs) // Store remaining IDs (skip first page) as scroll state. allIDs := make([]string, 0, len(filtered)) for _, d := range filtered { @@ -996,11 +993,26 @@ func osMinAgg(docs []map[string]interface{}, field string) map[string]interface{ return map[string]interface{}{"value": min} } -// generateOpenSearchID generates a random URL-safe base64 ID. -func generateOpenSearchID() string { - b := make([]byte, 12) - _, _ = rand.Read(b) - return base64.RawURLEncoding.EncodeToString(b) +// generateOpenSearchID mints a URL-safe base64 ID from m — sixteen characters from twelve bytes, +// which is byte-for-byte the width and alphabet the crypto/rand form produced. +// +// It serves two kinds of value, and the wider one is the reason deriving it matters. A document `_id` +// substrate generates is returned in the index response and is the path of every later `GET`, `PUT` +// and `DELETE` of that document; a `_scroll_id` is handed straight back to `_search/scroll`. Both are +// values a caller sends back, so a re-minted one turned a recorded read into a `not_found` 404 and a +// recorded scroll continuation into `search_context_missing_exception` against a cursor the recording +// had just opened. +// +// There is no AWS API model here to consult: the document and scroll APIs are OpenSearch's own REST +// interface rather than the `es`/`opensearch` control plane, so the shape substrate publishes is +// OpenSearch's convention rather than an AWS constraint — which is the other reason not to change it +// while deriving it. Real OpenSearch generates a wider id (a 20-character Flake); substrate's +// sixteen characters are narrower and always were, and #856 is not where that changes. +// +// [IDMint.Base64URL] rather than [IDMint.Base64], because a document id travels in a URL path and the +// two encodings differ in exactly the characters that would need escaping there. +func generateOpenSearchID(m *IDMint) string { + return m.Base64URL(12) } func orEmpty(v interface{}) interface{} { diff --git a/emulator/quicksight_plugin.go b/emulator/quicksight_plugin.go index 7edda605..45188dad 100644 --- a/emulator/quicksight_plugin.go +++ b/emulator/quicksight_plugin.go @@ -2,7 +2,6 @@ package emulator import ( "context" - "crypto/rand" "encoding/json" "fmt" "net/http" @@ -189,7 +188,7 @@ func (p *QuickSightPlugin) createDataSource(ctx *RequestContext, req *AWSRequest return nil, fmt.Errorf("createDataSource: put: %w", err) } - reqID := generateQuickSightRequestID() + reqID := generateQuickSightRequestID(ctx.IDs) return quicksightJSONResponse(http.StatusCreated, map[string]interface{}{ "DataSourceId": body.DataSourceID, "Arn": arn, @@ -215,7 +214,7 @@ func (p *QuickSightPlugin) describeDataSource(ctx *RequestContext, _ *AWSRequest } return quicksightJSONResponse(http.StatusOK, map[string]interface{}{ "DataSource": ds, - "RequestId": generateQuickSightRequestID(), + "RequestId": generateQuickSightRequestID(ctx.IDs), "Status": http.StatusOK, }) } @@ -230,7 +229,7 @@ func (p *QuickSightPlugin) createDataSet(ctx *RequestContext, req *AWSRequest, _ } arn := fmt.Sprintf("arn:aws:quicksight:%s:%s:dataset/%s", ctx.Region, ctx.AccountID, body.DataSetID) - ingestionID := generateQuickSightRequestID() + ingestionID := generateQuickSightIngestionID(ctx.IDs) ds := QuickSightDataSet{ DataSetID: body.DataSetID, @@ -255,7 +254,7 @@ func (p *QuickSightPlugin) createDataSet(ctx *RequestContext, req *AWSRequest, _ "DataSetId": body.DataSetID, "Arn": arn, "IngestionId": ingestionID, - "RequestId": generateQuickSightRequestID(), + "RequestId": generateQuickSightRequestID(ctx.IDs), }) } @@ -286,16 +285,41 @@ func (p *QuickSightPlugin) describeIngestion(ctx *RequestContext, _ *AWSRequest, }, "CreatedTime": createdTime, }, - "RequestId": generateQuickSightRequestID(), + "RequestId": generateQuickSightRequestID(ctx.IDs), "Status": http.StatusOK, }) } -// generateQuickSightRequestID generates a UUID-formatted request ID. -func generateQuickSightRequestID() string { - b := make([]byte, 16) - _, _ = rand.Read(b) - return fmt.Sprintf("%x-%x-%x-%x-%x", b[0:4], b[4:6], b[6:8], b[8:10], b[10:16]) +// generateQuickSightRequestID mints the `RequestId` every QuickSight response carries. +// +// QuickSight is unusual in publishing the request id as a *body* member on every operation, success +// and error alike, rather than only as a header — API_CreateIngestion documents it as "the AWS request +// ID for this operation" with no type constraint beyond String. So it is an observation a caller can +// make, and #856 applies for the ordinary reason: a replayed response must carry the recorded value. +// +// It is minted rather than taken from [RequestContext.RequestID], and the reason is the shape. +// Substrate's own request id is `req-{nanos}-{8 hex}` (see generateRequestID) — deliberately +// wall-clock, per #866 — which is not a form AWS ever sends, and nothing observes the two together: +// substrate sets no `x-amzn-RequestId` header, so the body member has nothing to disagree with. +// Echoing the internal id would trade a shape a caller can parse for a correspondence no call can see. +func generateQuickSightRequestID(m *IDMint) string { + return m.HexUUID() +} + +// generateQuickSightIngestionID mints the SPICE ingestion ID a `CreateDataSet` response reports and +// `DescribeIngestion` addresses. +// +// It is its own minter rather than a second call into [generateQuickSightRequestID], which is where +// `createDataSet` used to reach: an ingestion ID is a *handle* — the recorded `DescribeIngestion` path +// contains it — where a request ID is observed once and never sent back, so one function serving both +// meant a change to how substrate renders a request ID would silently move the identifier a recorded +// URL depends on. The same split #856 made between ACM's certificate IDs and API Gateway's API keys. +// +// The rendering is unchanged from the shared generator's, and the API model permits it: +// API_CreateIngestion publishes `IngestionId` at 1–128 characters matching `^[a-zA-Z0-9-_]+$`, which +// the UUID shape satisfies — the hyphens included. +func generateQuickSightIngestionID(m *IDMint) string { + return m.HexUUID() } // quicksightJSONResponse serializes v to JSON and returns an AWSResponse. diff --git a/emulator/redshiftdata_plugin.go b/emulator/redshiftdata_plugin.go index 8eba01a8..c9735fd2 100644 --- a/emulator/redshiftdata_plugin.go +++ b/emulator/redshiftdata_plugin.go @@ -2,8 +2,6 @@ package emulator import ( "context" - "crypto/rand" - "encoding/hex" "encoding/json" "fmt" "net/http" @@ -102,7 +100,7 @@ func (p *RedshiftDataPlugin) executeStatement(reqCtx *RequestContext, req *AWSRe } } - id := generateRedshiftDataID() + id := generateRedshiftDataID(reqCtx.IDs) stmt := RedshiftDataStatement{ ID: id, Status: status, @@ -246,15 +244,23 @@ func (p *RedshiftDataPlugin) loadStatement(acct, region, id string) (*RedshiftDa return &stmt, nil } -// generateRedshiftDataID generates a UUID-style statement ID. -func generateRedshiftDataID() string { - b := make([]byte, 16) - _, _ = rand.Read(b) - return hex.EncodeToString(b[0:4]) + "-" + - hex.EncodeToString(b[4:6]) + "-" + - hex.EncodeToString(b[6:8]) + "-" + - hex.EncodeToString(b[8:10]) + "-" + - hex.EncodeToString(b[10:16]) +// generateRedshiftDataID mints a statement ID from m. +// +// This is the one identifier in the analytics family whose shape AWS states outright: +// API_ExecuteStatement documents `Id` as "a universally unique identifier (UUID) generated by Amazon +// Redshift Data API" and publishes the pattern `[a-z0-9]{8}(-[a-z0-9]{4}){3}-[a-z0-9]{12}(:\d{0,2})?`. +// The character class is `[a-z0-9]` rather than `[0-9a-f]`, and it constrains no position, so it is +// indifferent to the RFC 4122 version and variant nibbles — which is why [IDMint.HexUUID] satisfies it +// and the rendering the crypto/rand form produced is preserved rather than changed (#671). +// +// Substrate mints no session, so the optional `:\d{0,2}` suffix — which the same pattern admits on +// `SessionId` — is never emitted. +// +// Deriving it is what lets a recorded statement be read back: `DescribeStatement`, `GetStatementResult` +// and `CancelStatement` all address the statement by this id, so a re-minted one answered each with +// ResourceNotFoundException against a statement the recording had just submitted. +func generateRedshiftDataID(m *IDMint) string { + return m.HexUUID() } // redshiftDataJSONResponse serializes v to JSON and returns an AWSResponse with diff --git a/emulator/timestream_plugin.go b/emulator/timestream_plugin.go index a3bed52b..06fa3d0e 100644 --- a/emulator/timestream_plugin.go +++ b/emulator/timestream_plugin.go @@ -352,7 +352,7 @@ func (p *TimestreamPlugin) query(reqCtx *RequestContext, req *AWSRequest) (*AWSR } return timestreamJSONResponse(http.StatusOK, map[string]any{ - "QueryId": randomHex(16), + "QueryId": timestreamQueryID(reqCtx.IDs), "Rows": rows, "ColumnInfo": cols, "NextToken": "", @@ -363,6 +363,20 @@ func (p *TimestreamPlugin) cancelQuery(_ *RequestContext, _ *AWSRequest) (*AWSRe return timestreamJSONResponse(http.StatusOK, map[string]any{}) } +// timestreamQueryID mints a `Query` response's QueryId from m — 32 lowercase hex characters, which is +// what [randomHex] produced at the site this replaces and what API_query_Query permits: `QueryId` is +// 1–64 characters matching `[a-zA-Z0-9]+`, so hex is inside the published alphabet where a UUID's +// hyphens would not be. +// +// It is the loosest case in this family, because substrate's Query is synchronous and its QueryId +// reaches nothing: `CancelQuery` ignores the ID it is given and answers an empty body, and there is no +// asynchronous read path to address. So deriving it does not repair a broken follow-on call the way +// Athena's and Redshift Data's do — it removes the last fresh draw from the response body, which is +// what makes a recorded Query replay with zero differences at all. +func timestreamQueryID(m *IDMint) string { + return m.Hex(16) +} + // lookupQueryResult returns the seeded result for the given query string, // falling back to the wildcard "*" seed, then to stored records from // WriteRecords if the query matches "SELECT * FROM db.table", and finally