Skip to content

Identifiers minted from crypto/rand make a replay diverge from its recording, and three replay checks are weakened for it #856

Description

@scttfrdmn

Substrate mints most of its identifiers from crypto/rand, so a recorded run and its replay produce different identifiers. This is the shared prerequisite of #833 and #817, and no issue covers it.

grep across non-test emulator/ code: 78 generate* functions, in 31 files that import crypto/rand. Nearly all draw from it directly.

What it blocks, concretely

#833 — authorization fidelity. A stream containing a mid-stream CreateAccessKey or AssumeRole mints a different AKIA…/session token on replay. The recorded Authorization header then names a key that is absent from replayed state, resolvePrincipal returns nil, and CheckAccess fails open the instant reqCtx.Principal == nil:

	if reqCtx.Principal == nil {
		return nil
	}

(emulator/authz.go:151-153)

So a replay of a recorded stream would authorize everything — the opposite of the fidelity #833 asks for — and it would do so silently.

#817 — body comparison. Any create whose response carries a generated identifier diverges, so a strict body differ reports a mismatch on essentially every mutating event. Criterion 3 of #817 assumes the opposite.

State-hash validation. As of the v0.115.0 reporting work, IncludeStateHashes is honoured and ValidateState actually runs. A stream containing any create now reports a state_hash_after mismatch, because the state key or record contents embed the freshly minted identifier. That is the correct answer — the divergence is real — but it means state validation is only useful once identifiers are reproducible.

The convention already exists in two places and was never generalised

  • CloudFormation derives its stack and change-set UUIDs from account + region + name (docs/services.md:1414-1418), specifically so a redeploy and a replay mint the same value.
  • error_protocol.go:10-14 states the principle outright: substrateRequestID = "SUBSTRATE" is a fixed constant because "an error body has to be byte-identical across two replays of one recorded run".
  • Config rules do the same (configservice_rules.go:412): "AWS mints an opaque config-rule-xxxxxx; substrate derives it by hash so replaying …".

So the design question is not whether to derive rather than randomise, but what to derive from, given that the same operation may legitimately be called twice in one run with identical inputs (two CreateAccessKey calls for one user must produce two distinct keys).

The design question this issue has to answer first

A pure hash of the request is wrong for exactly that reason. The candidate inputs, and what each costs:

Source Reproducible? Distinguishes repeat calls? Notes
account + region + name yes n/a (name is the discriminator) works for named resources; CFN's approach
hash(request) yes no two identical creates collide
hash(request + event sequence) yes yes needs the plugin to see the sequence number, which it does not today
a seeded PRNG per stream yes yes needs a seeded source threaded to every mint site; ReplayEngine.RandInt64 already exists (replay.go:642) but is reachable by nothing
a monotonic per-namespace counter in state yes yes reproducible because it is state, which replay resets; needs a state write per mint

Note that ReplayEngine already carries a seeded rng and exposes RandFloat64/RandInt64 (replay.go:629-647), and ReplayConfig.RandomSeed is plumbed — but no plugin can reach any of them, because a plugin holds no reference to the replay engine. That is the missing wiring, whichever source is chosen.

Whether the answer is one mechanism for all 78 sites or a per-shape choice (ARNs and names derived from inputs; opaque tokens from a seeded source) is the first decision, and it should be made before either replay issue is scheduled.

Acceptance criteria

  • A written decision on the identifier source, recorded where the existing conventions are (docs/services.md, alongside the CloudFormation UUID note) rather than only in code.
  • A single seam every mint site goes through, so the answer cannot be given two ways in two plugins — the same "one builder, so they cannot drift" structure SQS state keys disagree: the tagging API and the authorizer address queue:<name>, the plugin writes queue:<account>/<name> #826 established for ARNs.
  • Two identical create calls in one run still produce distinct identifiers. This is the criterion a naive hash(request) fails, and it must be a test.
  • Recording a stream and replaying it produces byte-identical identifiers, asserted through a real state_hash_after comparison (which the v0.115.0 work makes possible).
  • The 78 sites are enumerated in the issue or a checklist so the migration is auditable rather than "mostly done".
  • Scoped explicitly: which identifier kinds are in scope (ARNs, resource IDs, opaque tokens, request IDs, ETags, session tokens) and which are deliberately left random, with the reason.

Also worth deciding here: MemoryStateManager.List returns Go map iteration order with no sort (state.go:80-97), across 118 non-test state.List( call sites. That makes two identical ListBuckets requests differ within one run, independent of identifier minting. It is a smaller fix (sort in List, once) and arguably belongs in this issue since it has the same symptom and the same two dependents; it is recorded here so it is not researched a third time.

Filed as the prerequisite for #833 and #817 during the v0.115.0 replay-reporting research.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area: replayIssues relating to the replay componentpriority: mediumShould be done for this milestonetype: bugSomething is not working correctly

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions