Conversation
iurii-ssv
force-pushed
the
epbs-gloas
branch
2 times, most recently
from
June 24, 2026 13:01
bc977ed to
838fe4b
Compare
iurii-ssv
added a commit
that referenced
this pull request
Jun 28, 2026
Promote EPBS_IMPLEMENTATION_PLAN.md from local-only (.git/info/exclude) into the branch so the in-flight ePBS planning context is shared, not local. The file carries an explicit action item: before #2901 is marked ready for review, move all remaining/useful action items into the PR description and delete this file — it must not outlive the PR.
iurii-ssv
marked this pull request as ready for review
June 28, 2026 12:10
Contributor
Greptile SummaryThis PR adds the node-side foundation for ePBS/Gloas. The main changes are:
Confidence Score: 4/5The Gloas message validation and block production paths need fixes before merging.
message/validation/signed_ssv_message.go; beacon/goclient/gloas_proposer.go Important Files Changed
Reviews (1): Last reviewed commit: "gloas: log §2 vote index + §5 proposer-p..." | Re-trigger Greptile |
momosh-ssv
reviewed
Jun 29, 2026
iurii-ssv
added a commit
that referenced
this pull request
Jun 30, 2026
Comment/doc-only response to a #2901 review (most findings were non-issues or over-stated — assessed in the plan); the actionable bits: - beacon_block.go: note blob KZG commitments also leave the body (the payload and blobs ship in the §6 envelope) — they were missing from the drops list. - ptc.go: clarify the hand-rolled client's missing custom-TLS matches the main eth2clienthttp path (system-CA https + basic-auth), so it's no regression. - plan §2: devnet-verify that a Gloas BN accepts the Fulu-tagged attestation submission (BeaconForkAtEpoch caps at Fulu; TODO(gloas) to extend if rejected). - plan §2b: record the remote-signer limitation — Web3Signer has no PTC / proposer-preferences / envelope sign types; bounded by f, local-sign workaround; operator-facing.
Contributor
Author
|
The ePBS validation log that lived in this comment has been moved into the PR description (the Validation section) — so it sits alongside the implementation summary and is easy to find. This comment is kept only as a pointer. |
iurii-ssv
added a commit
that referenced
this pull request
Jun 30, 2026
Promote EPBS_IMPLEMENTATION_PLAN.md from local-only (.git/info/exclude) into the branch so the in-flight ePBS planning context is shared, not local. The file carries an explicit action item: before #2901 is marked ready for review, move all remaining/useful action items into the PR description and delete this file — it must not outlive the PR.
iurii-ssv
added a commit
that referenced
this pull request
Jun 30, 2026
Comment/doc-only response to a #2901 review (most findings were non-issues or over-stated — assessed in the plan); the actionable bits: - beacon_block.go: note blob KZG commitments also leave the body (the payload and blobs ship in the §6 envelope) — they were missing from the drops list. - ptc.go: clarify the hand-rolled client's missing custom-TLS matches the main eth2clienthttp path (system-CA https + basic-auth), so it's no regression. - plan §2: devnet-verify that a Gloas BN accepts the Fulu-tagged attestation submission (BeaconForkAtEpoch caps at Fulu; TODO(gloas) to extend if rejected). - plan §2b: record the remote-signer limitation — Web3Signer has no PTC / proposer-preferences / envelope sign types; bounded by f, local-sign workaround; operator-facing.
iurii-ssv
added a commit
that referenced
this pull request
Jul 1, 2026
The proposer, attester and sync-committee duty-fetch handlers marked an epoch/period intent fulfilled even when no validators were eligible at fetch time, so the duties were never fetched once validators did become eligible (e.g. after a beacon-metadata sync that arrives without an accompanying indices-change event). On the Gloas devnet this surfaced as the proposer missing every block it was assigned. fetchAndProcessDuties now returns (fetched bool, err error); the caller marks the intent fulfilled only when a beacon fetch actually ran. "No eligible validators" returns fetched=false, leaving the intent pending so a later tick retries — the same model the PTC and proposer-preferences handlers already use. Also drop the now-redundant per-fetch bracket log lines, and add a temporary proposer diagnostic (#2901) that dumps the Validators()/ SelfValidators() view on no-eligible to confirm the root cause on devnet. Scheduler tests updated to assert the intent stays pending and that a late indices-change remains the sole re-fetch trigger.
iurii-ssv
added a commit
that referenced
this pull request
Jul 1, 2026
Promote EPBS_IMPLEMENTATION_PLAN.md from local-only (.git/info/exclude) into the branch so the in-flight ePBS planning context is shared, not local. The file carries an explicit action item: before #2901 is marked ready for review, move all remaining/useful action items into the PR description and delete this file — it must not outlive the PR.
iurii-ssv
added a commit
that referenced
this pull request
Jul 1, 2026
Comment/doc-only response to a #2901 review (most findings were non-issues or over-stated — assessed in the plan); the actionable bits: - beacon_block.go: note blob KZG commitments also leave the body (the payload and blobs ship in the §6 envelope) — they were missing from the drops list. - ptc.go: clarify the hand-rolled client's missing custom-TLS matches the main eth2clienthttp path (system-CA https + basic-auth), so it's no regression. - plan §2: devnet-verify that a Gloas BN accepts the Fulu-tagged attestation submission (BeaconForkAtEpoch caps at Fulu; TODO(gloas) to extend if rejected). - plan §2b: record the remote-signer limitation — Web3Signer has no PTC / proposer-preferences / envelope sign types; bounded by f, local-sign workaround; operator-facing.
iurii-ssv
added a commit
that referenced
this pull request
Jul 1, 2026
The proposer, attester and sync-committee duty-fetch handlers marked an epoch/period intent fulfilled even when no validators were eligible at fetch time, so the duties were never fetched once validators did become eligible (e.g. after a beacon-metadata sync that arrives without an accompanying indices-change event). On the Gloas devnet this surfaced as the proposer missing every block it was assigned. fetchAndProcessDuties now returns (fetched bool, err error); the caller marks the intent fulfilled only when a beacon fetch actually ran. "No eligible validators" returns fetched=false, leaving the intent pending so a later tick retries — the same model the PTC and proposer-preferences handlers already use. Also drop the now-redundant per-fetch bracket log lines, and add a temporary proposer diagnostic (#2901) that dumps the Validators()/ SelfValidators() view on no-eligible to confirm the root cause on devnet. Scheduler tests updated to assert the intent stays pending and that a late indices-change remains the sole re-fetch trigger.
iurii-ssv
added a commit
that referenced
this pull request
Jul 1, 2026
Two TEMP logging-only diagnostics to pinpoint where an assigned Gloas proposer duty is lost between fetch and the runner. The existing zero-eligible diagnostic only covers the "never fetched" case; these cover the loaded-epoch case: - logSlotDispatchDiagnostic (processExecution): on any slot carrying a stored proposer duty, reports stored_any / in_committee / executable so a run can separate an InCommittee-flag drop from a one-slot-window miss from a downstream dispatch loss (cross-checked against the existing 🔧 executing validator duty / could not find validator logs). - logFetchDispatchDiagnostic (fetchAndProcessDuties): flags in-committee duties stored for already-passed slots (fetched-too-late), and surfaces the InCommittee split on the fetch-success path. Read-only; no control-flow change. Remove after devnet confirmation.
This was referenced Jul 2, 2026
…duty The committee and aggregator-committee runners tagged every failed post-consensus reconstruct recoverable after the fallback. When the fallback finds no bad share, as with a share set that doesn't combine to the validator's key, the root stays at quorum and more shares can't fix it: each retry failed the same way, and the duty was recorded as stuck at the slot's end rather than failed. Both now reconstruct through reconstructQuorumSig: such a failure is terminal and concludes the duty failed with the reason, after the other validators' messages are submitted as before; a root still at quorum after a drop is retried at once; one left below quorum stays recoverable. The aggregator-committee runner keeps marking a recoverable failure with its spec code rather than the tag.
…through reconstructQuorumSig too The pre-consensus selection-proof reconstruct was the last one outside the shared helper. A failure still only skips that validator, as any pre-consensus failure does, but one with no bad share to drop is now told apart from a recoverable one, and a root still at quorum after a drop is retried at once. The committee runners' reconstruct-failure logs now say whether the failure is recoverable, and their quorum re-check comments say why it's needed.
…roposer-index and registration rules ssv-spec#643 was rebased onto #632's latest head after the pin, with rule changes the node now follows. The value check pins a Gloas block's proposer to the duty's validator (ssv-spec#647), so a leader stamping another index fails consensus instead of having the cluster sign a block the beacon node rejects. A Gloas-slot validator-registration duty is rejected before it starts, leaving no running duty behind. The spec now also publishes the §6 reveal after a failed block submit, as the node does, so the proposer's comment no longer calls that a deviation. The Gloas test block is proposed by the testing duties' validator, as in the spec.
… reject unexpected partial types in the preferences dispatcher The node's GloasBeaconVote duplicated ssv-spec's: same fields, same dynamic-ssz encoder. Alias it to the spec's type and drop the copy, so the committee runner decides, and validates, exactly the spec's value. decidedAttestationVote now validates both vote forms: checkpoints set, index 0 or 1, source before target. The value check already validates every decided value, so this guards the post-consensus paths against that check regressing. Both value checks call Validate too, so the Gloas check now tests the index before the epochs, as the spec does. The proposer-preferences dispatcher rejects a partial of any type other than preference and request-auth with the spec's ProposerPreferencesUnexpectedPartialSigTypeErrorCode. The spec checks in the per-slot runner; the node's dispatcher stashes every partial before routing it, so the check sits ahead of the stash.
…horizon, preference re-emissions, validation state An error after a runner's instance decides now concludes the duty failed, in every runner that decides: proposer, aggregator, sync-committee contribution, committee and aggregator committee. So does a decided value the runner can't take up, one that fails to decode or fails the local value check. A decided instance never decides again, so these used to surface only as "stuck" at the deadline. A doppelganger skip still leaves the duty open: the block, and on Gloas the builder operator's reveal, can still go out from the other operators' partials. The PTC outcome horizon is the duty slot's end plus the spec's MAXIMUM_GOSSIP_CLOCK_DISPARITY. Gossip ignores a payload attestation once its slot is over, and only the next slot's block can include it, so a quorum later than that no longer counts as a success. A proposer-preferences re-emission that can't rebuild its preference keeps converging on the one the slot already broadcast, instead of taking the slot over with nothing frozen and dropping the partials gathered for it. The dispatcher returns nil for a partial it stashes, so early partials no longer show up as dropped, with failed trace spans. Message validation keeps a spare slot in the preferences ring, for networks whose slots are shorter than the late and early margins together, and the checks that run before signature verification read peer state without adding it. Comments: Boole is out of the fork-lag stop comment, stateTTL says every access restarts it, and the config source says why it is read without the lock. Tests: the runners' decide helpers return a conclusion channel and the first ProcessConsensus error, so the failure tests reuse them, and a signer that fails one signing domain stands in for a failed post-consensus signature.
…ty failed; read peer state once, by value The committee runner's zero-count branch returns a done context's error. Its comment said that concludes no outcome, but with the deferred markDutyFailed only a cancellation is dropped: an expired duty deadline now concludes the duty failed, which is the right outcome for a duty that ran out of time. The comment says so, and the done-context test covers both cases. peekPeer returns a copy of the recorded peer state, and the limit checks bind it once, so no check allocates for an unknown peer and none looks it up twice. The note next to the §5 publish-finality caveat now records the other §5 caveat: a re-emission that can't build its preference keeps the one already broadcast, which after a dependent_root change is stale, so the duty is reported stuck at the proposal slot.
This was referenced Sep 24, 2026
…roposer's post-decide slot guard as the spec did The 24 Sep head adds the version JSON codec, resets the §6 state on executeDuty, carries preference shares over, and adds the "external bid with non-zero payload_root" value-check vector. Both spec suites pass against it. It also removed its ProcessConsensus slot backstop: the runner re-runs the value check on the decided value, and that check's running-slot bind is the only slot guard. Ours was redundant for the same reason, since the validator controller always wires the running slot into the checker, so it goes too. The test now wires the checker the same way and still sees a value decided for another slot refused before anything is signed.
…_EPOCH/2 - 1, as SIP #94 §5 now requires The SIP's 24 Sep revision times every next-epoch emission from slot SLOTS_PER_EPOCH/2 - 1, once that epoch's dependent_root block has settled, and that includes the pre-fork emission of the first Gloas epoch. The pre-fork window emitted from its first slot; it now waits for the same slot as the steady-state next-epoch emission.
… tests; document the checker's 0-slot skip and the early-tick recheck drop The three preferences tests that computed slot SLOTS_PER_EPOCH/2 - 1 of the epoch before the fork now share preForkEmissionSlot. NewProposerChecker's doc now says a runningDutySlot of 0 skips the slot check, as nil does; it read as if values for every other slot were rejected before the first duty. emitForTick's doc says a tick before the first emission drops a pending reorg recheck, which loses nothing, and HandleDuties leaves the emission timing to it.
… once-per-root rule in emitForEpoch's doc only The Gloas preferences and registration tests computed epoch start slots by hand; they now use the network config's FirstSlotAtEpoch. The emitted field's doc repeated emitForEpoch's once-per-root rule; it now just says what the map holds and defers to emitForEpoch.
…eping progressive copies The node declared BlockContents and SignedExecutionPayloadEnvelopeContents as progressive containers, while go-eth2-client ships them as plain ones. They are only encoded and decoded, never hashed, and a progressive container serializes like a plain one, so the bytes are unchanged; only the unused hash tree root differed. They are now aliases of go-eth2-client's api/v1/gloas types, as the consensus types already are, and their generated code is gone.
…eal out The reveal waits for this operator's block submit, and a quorum fires once. A terminal block-signature reconstruction, a lost post-consensus quorum (every packet carries the block root, so the first quorum includes the block's), and a decided value the runner can't take up all rule that submit out, yet kept the payload and blobs until the validator's next proposal. Each now releases them. loseGloasQuorum's comment claimed the envelope's quorum could still carry the reveal, and its duty-succeeded branch was unreachable; both are gone, and its test covers the real case. Also pin that an external bid ignores a stray envelope from the beacon node, where ssv-spec takes the envelope's root and its own value check then rejects the value.
2 tasks
…Gloas is scheduled Until the Gloas fork is scheduled neither role gets a duty or a message, and the node's fork schedule is fixed while it runs. Yet every validator carried both runners, each with a queue consumer and its janitor: four idle goroutines per validator, about 39% more goroutines per node on stage-hoodi. SetupRunners now adds the two roles only when the new GloasScheduled holds, which also excludes the far-future epoch a beacon node gives a fork it names but hasn't scheduled. Scheduled rather than active, as §5 emits the first Gloas epoch's preferences before the fork.
A consensus message for an instance that already decided comes back decided, with a benign skip error, as in ssv-spec. The release added for a decided value the runner can't take up treated that skip as one. So a late prepare or commit between a builder operator's decision and its post-consensus quorum released the reveal data under a cached builder match, and the §6 publish dereferenced nil and took the node down: 2 of 4 operators on local_testnet_gloas, at the committee's first Gloas proposal. The release now happens only when no decided value was taken up, which is what rules the reveal out, since post-consensus packets are validated against that value. The publish also refuses released reveal data with an error instead of dereferencing it, so a bug of this kind costs one reveal rather than the node. baseConsensusMsgProcessing now documents both ways decided comes with an error.
This was referenced Sep 27, 2026
…3044) * qbft/roundtimer: shorten the proposer round budget to 1.5s (SIP-102) Glamsterdam moves the attestation deadline from 4s to 3s into the slot. A proposer QBFT instance starts ~1.1-1.5s in, so with the shared 2s QuickTimeout a round change starts round 2 after the deadline: an event that is usually recoverable today becomes usually fatal. Give the proposer its own budget of 1.5s. The timer stays anchored at instance start, and no other role changes. 1.5s is bounded by measurement rather than taste: across 30 days of mainnet the slowest round 1 that went on to decide took 1148ms, so this never fires on a round 1 that would have succeeded, while 1000ms would have fired on 5-12 real duties a month. Operators can override it with ProposerQuickTimeout / PROPOSER_QUICK_TIMEOUT, accepted between 1148ms and 2s. The bounds are hard, following ProposerDelayEPBS rather than ProposerDelay: below the floor the budget measurably times out rounds that would have decided, above it a round change cannot land at all, so neither is a risk an operator can usefully accept. Setting 2s restores prior behavior without a code change. Not fork-gated. There is no message, signature or domain change, and pre-Gloas the shorter budget is a strict improvement. A mixed cluster is never worse than today: upgraded operators round-change at 1.5s, the rest at 2s, and once f+1 have upgraded the partial-quorum rule pulls the rest along. SIP-102's other requirement, that a proposer instance stop at round 2 rather than climb to the cluster-wide cutoff, is implemented generically in #3041 and is deliberately not duplicated here. * cli/operator: keep validateProposerQuickTimeout a pure predicate validateProposerQuickTimeout both validated and logged, so it had to take a *zap.Logger. Its sibling validateProposerDelay stays a pure predicate and leaves its Warn to resolveAndValidate, next to the call site. Drop the logger parameter and move the non-default Info up into resolveAndValidate, alongside the dangerous-ProposerDelay warn, so the two paths are symmetric and the validator is side-effect-free. The tests already go through resolveAndValidate, so they cover the moved log unchanged. * qbft/roundtimer: stop implying the ProposerQuickTimeout floor has margin The MinProposerQuickTimeout comment read as though nothing gets cut off at the floor itself, when 1148ms is the slowest round 1 we observed deciding, not a value below it. Configuring exactly the floor races that duty: the expiry and the consensus message reach the same queue with no tie-break rule. Say so. The floor marks the edge of the measured band; the margin lives in the 1500ms default, which clears the observation by 352ms. No behavior change. * qbft/roundtimer: note the committee-wide constraint on WithProposerQuickTimeout The warning that ProposerQuickTimeout must match across every operator of a shared committee lived only in config.example.yaml and the env-description, so it was invisible from the code side where the option is actually applied. Mirror it onto the WithProposerQuickTimeout godoc, including the rollout window as the one sanctioned exception. Comment only. * qbft/roundtimer: raise the ProposerQuickTimeout floor to 1250ms The floor sat at 1148ms, the slowest round 1 observed to decide, so the accepted range admitted a configuration SIP-102's own design goal rules out: the goal asks the budget to stay above that observation with margin, and at exactly the observation the expiry and the consensus message reach the same queue with no tie-break. TestProposerQuickTimeoutBounds encoded the asymmetry directly, holding the default to Greater while the floor only had to be Equal. Raise the floor to 1250ms, SIP-102's own alternative to the 1500ms default. Every accepted value now clears the observation outright, so the bounds test asserts the same strict inequality for the floor as for the default, and the floor still sits 250ms under the default so downward tuning in an incident is unaffected. Drops the comment added in f15baab that justified the zero-margin floor instead of fixing it. 1148ms and 1249ms join the rejected boundary cases. * qbft/roundtimer: correct the claim that proposer rounds 3+ are unreachable RoundTimeout's doc said message validation's round-2 cap made QuickTimeoutThreshold and the SlowTimeout branch unreachable for the proposer. The cap only makes peers drop the messages on receipt. Locally UponRoundTimeout bumps unconditionally and IsRelevant() is Round < CutOffRound with the role-independent cutoff of 12, so a proposer instance really does climb rounds 3 through 11 and arm the two-minute SlowTimeout for 9 through 11 - the tail #3041 is about. Say that instead: those rounds are wasted work, not dead code. Same correction for the "unreachable on the wire" line in the TestRoundTimeoutOffset comment, since the messages are broadcast and only dropped on receipt. Comments only. * qbft/roundtimer: start the DefaultProposerQuickTimeout doc with its identifier The comment opened with "ProposerQuickTimeout is ...", which is the config key, not the constant it documents. Godoc convention is to lead with the identifier. * cli/operator: scope the proposerQuickTimeout log assertions to their own log Both assertions read every record the observer captured at InfoLevel, so they asserted something about all of resolveAndValidate: any Info added anywhere in it would fail these tests under a ProposerQuickTimeout name. Filter by message snippet so they only speak to this behavior. * docs: retune the ProposerDelay guidance for the 1.5s proposer budget MEV_CONSIDERATIONS derived its "~1.2s maximum reasonable ProposerDelay" from a 2s round-1 budget. SIP-102 makes that budget 1.5s for the proposer, so the same constants now give ~700ms, and the doc was recommending a value roughly 500ms too large - landing an operator who follows it in exactly the regime the SIP exists to prevent. It is the only operator-facing doc doing arithmetic on this budget and it sits outside the change that invalidated it. Correct the number, say it moved and that anyone tuned against the old guidance should retune, and point operators who raised ProposerQuickTimeout at their own configured value. Also note the budget is anchored at instance start rather than slot start, which the equation approximates away. State explicitly why maxSafeProposerDelay stays at 1s: it is a backstop against broken configurations, not the recommended ceiling, and lowering it to ~700ms would refuse startup for nodes that are merely suboptimal. * operator/validator: pin the ControllerOptions hop of the ProposerQuickTimeout chain timer_config_test.go claimed that deleting any assignment along ControllerOptions -> CommonOptions -> Validator -> roundtimer.New would be caught, but the test starts at CommonOptions. The one hop it did not cover, controller.go:217, is the only field-by-field struct copy in the chain and so the likeliest place for the value to be dropped silently. Add TestNewControllerPropagatesProposerQuickTimeout to cover it, verified by deleting the assignment and watching it fail, and correct the other test's comment to claim only what it actually pins. * message/validation: stop the round-search loop stepping by the wrong budget The loop that searches for a time-into-slot yielding a target round advanced by QuickTimeout (2s), but the estimator advances the proposer by DefaultProposerQuickTimeout (1.5s) since SIP-102. For the proposer a 2s step visits rounds 1,2,3,5,6,7,9,... - rounds 4 and 8 are unreachable. It terminates today only because the proposer target is round 3; any future maxRound change putting it at 4 or 8 would hang the suite instead of failing it. Step by a granularity smaller than any role's budget and bound the search so a miss fails with the role and round rather than spinning. * qbft/roundtimer: retire values the floor raise left stale Three leftovers from raising the floor to 1250ms and from the proposer getting its own budget: - TestWithProposerQuickTimeout and timer_config_test both armed 1200ms, now below MinProposerQuickTimeout. The option layer does not validate so nothing broke, but timer_config_test exists to show an operator's configured value reaching the timer, and 1200ms is not one an operator can configure. Both use 1400ms, already the in-range probe in config_test.go. - The proposer row in TestEstimatedRoundAt still fed QuickTimeout as its time-into-slot. It passed only by accident (2s/1.5s still lands in round 2) and no longer sat on the boundary it was written to probe. Use DefaultProposerQuickTimeout. - MaxProposerQuickTimeout restated 2s as a literal while its doc called it the pre-SIP-102 budget, a claim only true while QuickTimeout is also 2s. Derive it. * qbft/roundtimer: pass the quick budget into roundTimeoutForRound explicitly roundTimeoutForRound derived its per-round budget from defaultQuickTimeoutForRole, i.e. always the protocol default, never this timer's configured one. That is inert today only because RoundRelativeRole keeps the proposer out of the slot- synchronized branch entirely. The day #2429 flips that predicate, the proposer falls through to this function and an operator's configured ProposerQuickTimeout stops being honored - while validation, the startup log and the whole config chain keep reporting the value that is no longer armed. Make quick a parameter so each caller states whose budget it means: RoundTimeout passes t.quickTimeout(), the tests pass defaultQuickTimeoutForRole to keep modelling a peer. No behavior change today, since the two agree for every role that currently reaches this path. * qbft/roundtimer: stop overstating the pre-Gloas and rollout safety Two claims in the DefaultProposerQuickTimeout block did not survive scrutiny. "The pre-Gloas effect is a strict improvement" counted only where round 2 starts. A round 1 that would have decided between 1.5s and 2.0s is now cut off where it previously succeeded, so it is a trade the 1148ms measurement makes unlikely rather than impossible. Likewise "never worse than today" claims more than 30 days of data can support; scope it to that data. Also record why round 2 runs on a budget measured from round 1: a round 2 needing longer has already missed the deadline, so sizing it separately would buy nothing. Comments only. * qbft/roundtimer: the proposer stops at its cap now that #3041 landed Rebasing onto epbs-gloas picked up CutOffRoundFor, which wires CutOffRoundFor(RoleProposer) = 3 into the instance's GetCutOffRound. The RoundTimeout doc still described the world before that: local instance unstoppable, rounds 3 through 11 armed, SlowTimeout reached for 9 through 11, and #3041 named as future work. All four are false on this base - the proposer now arms rounds 1 and 2 only and never reaches QuickTimeoutThreshold. * cli/operator: make the SIP-102 proposer round timeout an on/off switch Replace the ProposerQuickTimeout duration (1250ms-2s) with a ShortProposerRoundTimeout bool, default on. On arms the SIP-102 1.5s proposer round budget; off restores the pre-SIP-102 2s budget, which is the rollback lever. The range bounds and their validation go away since operators no longer enter a value. Below cli the switch is carried inverted as LegacyProposerRoundTimeout, so the zero value of every options struct keeps the SIP-102 default and a path that forgets to wire it cannot silently fall back to 2s. * cli/operator: pin the ShortProposerRoundTimeout inversion in newNode * docs: model ProposerDelay on the two-round worst case and require a committee-wide rollback * qbft/roundtimer: correct comments the switch rework and #3041 left stale * qbft/roundtimer: store the legacy switch on RoundTimer instead of a duration * tests: dedupe the ShortProposerRoundTimeout load cases and drop duration-era names * docs: keep MEV_CONSIDERATIONS out of this PR, the rewrite moves to #3058 * cli/operator: expose the rollback as LegacyProposerRoundTimeout, default off * qbft/roundtimer: address iurii's round-4 nits * cli/operator: load LegacyProposerRoundTimeout through YAML and env in a test The env var name lived only in the struct tag, so a typo there would leave the rollback silently inert. Spell both key names out literally in a load test, and hoist the minimal-config writer it shares with Test_config_load_trueDefaultBools. * qbft/roundtimer: scope the 1148ms claim to its window and order the layers top-down * tests: fold the duplicate proposer-default and Gloas head-start tests, fix an Equal arg order TestProposerRoundTimeoutArmsProposerQuickTimeout pinned what the "no option" subtest of TestWithLegacyProposerRoundTimeout already did, and TestGloasHeadStartsTrackRetimedDeadlines re-checked round1HeadStart against its own input. Fold their assertions and doc into the surviving tests, adding the missing aggregator-committee Gloas row. Also run ApplyDefaults in both legacy switch cases.
y0sher
previously approved these changes
Sep 28, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Node-side ePBS (EIP-7732 / Gloas) support, per SIP ssvlabs/SIPs#94.
Built on the Boole code baseline — consolidated roles,
ProposerConsensusData, and the node-side role switches — which has already merged tostage(via theintegration/boole-convergencemerge). This is a compile-time dependency on that code being present; it is not a dependency on the Boole fork activating. ePBS gates every execution path on the beacon node'sGLOAS_FORK_EPOCHalone and runs correctly with Boole dormant (Forks.Boole = math.MaxUint64— the default on every network, mainnet/sepolia/hoodi/holesky and now local-testnet), which is exactly the state the Glamsterdam transition runs in.The ePBS wire constants (runner/beacon roles, partial-sig types, beacon domains
0x0B/0C/0D, and the #2962 request-auth pairDomainBuilderRequestAuth 0x0B000001/RequestAuthPartialSig(9)) live in ssv-spec via PR #632 — the establishedspectypespattern, and required because ekm/ssvsigner reaches signing domains only throughspectypes. Their golden test lives there; both go.mods pin the top of that stack (ssv-spec#643) until the stack merges and is tagged. Those pins (and the go-eth2-client fork's) require go-bitfield under its renamed module path,github.com/OffchainLabs/go-bitfield; the node's last two imports of the oldprysmaticlabspath (the ENR subnets entry and a discovery test) follow, so the library is not carried twice — the old path is frozen at its last pre-rename commit and cannot be bumped.What's here
protocol/v2/types/gloaspackage — the Gloas beacon-chain containers (the §4 block family with the EIP-8282 five-listExecutionRequests, the §6 envelope, the PTC containers,ProposerPreferences) are go-eth2-client'sspec/gloastypes, aliased here; their SSZ is pk910/dynamic-ssz with the progressive-container/list merkleization of EIP-7688/EIP-7916 (Gloas SSZ types give a wrong hash_tree_root. All Gloas proposals fail. #3008). The SSV-owned types (GloasBeaconVote— fixed 120-byte, cross-fork decode fails cleanly —GloasProposalData(the §4 decided value: the block plus the self-buildpayload_root), the blinded envelope,PTCDuty, builder config and request auth) have dynamic-ssz encoders. Roots are pinned by a devnet-8 golden fixture: a finalized Gloas block whoseHashTreeRootmust equal the chain's header root and whose proposer signature must verify, plus its execution-payload envelope.GLOAS_FORK_EPOCHread from the BN spec into the fork map;IsGloas/IsGloasAtSlot, safe on pre-Gloas networks.IntervalDurationbecomes slot-keyed (1/3 of the slot → 1/4 at the fork) and drives duty deadlines, QBFT round-1 head starts, aggregation waits, and attestation fetch budgets.QuickTimeoutis deliberately not retimed — the Gloas proposer is round-1-must-succeed pending real round-trip data.roundtimer.RoundRelativeRole(the proposer's round-relative timing, also what message validation's round-spread exemption keys off) covers the proposer alone: §6 adds no QBFT instance of its own.GloasBeaconVote(carrying the BN's payload-status index) on Gloas slots: fork-aware decode across runner, post-consensus validation, committee observer, and duty tracer;NewGloasVoteChecker, which also applies SIP §2's same-slot SHOULD: it rejects index 1 for a block the operator's head events place at the duty slot (the spec reference has no such view and skips the check; gloas value check: SIP #94's new optional same-slot AttestationDataIndex check is not implemented #3035); index-preserving aggregation; and a hand-rolled attestation-data GET on Gloas slots (go-eth2-client's post-ElectraIndex==0check rejects the healthy FULL status). Aggregation on Gloas slots uses go-eth2-client's dedicatedgloas.AggregateAndProof/SignedAggregateAndProofend to end (fetch, consensus-data decode, signed submit): the container serializes like Electra's but merkleizes differently, so the Electra path would sign the wrong root (All AGGREGATOR duties fail after the Gloas fork. The library refuses the version header "gloas". #3009).ForkAtEpochresolves Gloas fromGLOAS_FORK_EPOCHon (no more Fulu cap, in both the node and ssvsigner), so attestations ridegloas.Attestationand the aggregator-committee consensus data stampsDataVersionGloas— the SIP §2 byte-parity stamp Anchor uses (ePBS/Gloas: stamp DataVersionGloas into AggregatorCommitteeConsensusData.Version (drop the Fulu-cap workaround) #2998). This relies on ssv-spec's Gloas arms building the Gloas containers (on the dynssz branch, ssv-spec#643, vectors regenerated); with the Electra-reusing arms the fork's submit path would reject the attestations.PTCAttesterRunner(no consensus, partial-signature only, honest convergence over the frozen 75%-cutoff observation, abstains when no block is seen); scheduler handler firing at the cutoff; duty store + message validation; ekm signing.include_payload=true); the QBFT value isGloasProposalData{block, payload_root}— the bid-only block (no blinding) plus the self-buildpayload_root(see §6) — bound to the running duty's slot (the value check, re-run on the decided value, and post-consensus validation); every operator submits the decided block (BN dedupes by root — to be re-confirmed on a Gloas BN); slashing-protected local signing (check→record→sign under a dedicated lock); build-source telemetry; newProposerDelayEPBSknob (hard-capped at 1s, no dangerous override).GloasProposalData{block, payload_root}; the value check pins the version to the slot's fork, the block's slot and proposer to the duty, andpayload_root != 0exactly whenbuilder_index == SELF_BUILD. On the self-build path every operator derives the blinded envelope from the decided value alone and signs its root underDomainBeaconBuilderas the second entry of the block's post-consensus packet (the block root underDomainProposeris the required first entry; §7 caps the packet at two entries on Gloas slots). Receivers match entries by signing root, each expected root at most once, and reconstruct each root independently from the same 2f+1 shares as the block; the envelope is published only once this operator has attempted the block submit (a beacon node ignores an envelope whose block it hasn't seen), even when the envelope's quorum forms first; a failed submit doesn't hold it back, as other operators submit the block too (ssv-spec does the same). Only the operator whose beacon node produced the decided block holds the payload — its produced envelope's blinded root must equal the derived root (builtDecidedEnvelope) — and it alone publishes the fullSignedExecutionPayloadEnvelope; the other operators contribute their envelope share and nothing else. The proposer keeps accepting the duty slot's post-consensus packets after the block has been submitted (awaitingEnvelope), so late envelope shares still count. The reveal data (payload, blobs, proofs) is released as soon as this operator can no longer publish it — another operator's block decided, the reveal attempted, or a terminal failure before it — while a failed block submit keeps it for the envelope quorum that may still follow; a duty that never concludes holds it until the next proposal (#3043). Publication is due byPAYLOAD_DUE_BPS= 50% of the slot — consensus-specs#5414 lowered it 75%→50% (§3 PTC'sPAYLOAD_ATTESTATION_DUEstays 75%).BuilderRequestAuth{data, proposal_slot}signing (builder-specsDOMAIN_BUILDER_REQUEST_AUTH, genesis-style) rides the §5 dispatcher: per configured builder —Builderscluster config in keymanager-APIs#88 vocabulary, validated at startup, required identical across all n operators — oneRequestAuthPartialSigpacket per proposal slot carrying a partial per distinct auth root (1 to 8 entries;AuthDatadefaults to the URL's hostname per builder-specs#168), collected entry by entry in a dedicated container with no succeeded-gate, reconstructed into a per-validator auth cache. Message validation admits the new type under role 8 with its own distinct-root budget: a packet that adds no new root, or would take its signer past the budget, is IGNOREd (SIP §7). Sub-quorum degrades silently to the enshrined flow — never blocks the proposal; reconstructions and auth-unavailable are counted, and the §4 build-source telemetry is a typed enum. The upstream specs this was written against have merged and the branch is reconciled to them (builder-specs#165 renames; keymanager-APIs#88 supersedes validate change round justification signer uniqueness like we do in regular agg. messages #87; beacon-APIs#630 supersedes Verify signed message with domain #625), and phases 2–3 are implemented here: the produceBlockV4 POST attach with theEth-Builder-Urlblock-forwarding echo (§4), and the BN-mediated ahead-of-timesubmitBuilderPreferences— both e2e-gated on a Round Timer Fix #630 beacon node. Design record: #2962 — closed, its three phases being implemented here; its remaining tails are tracked in the checklist below, ssvlabs/aetheria#173 and #3000.SSVConfigByName;GetStateRootnil-state guard; runner-state JSON dedup; a bad partial signature in a quorum no longer fails the duty, in any runner: the fallback drops it, the next honest share completes the duty, and a root still at quorum after the drop is retried at once. The sync-committee-contribution runner also no longer drops the other subnets over one bad root, from consensus or from submission.vendor/sszgen directive; documented thatexporter/model_encoding.gois hand-maintained.Validation — hermetic Gloas devnet (
local_testnet_gloas)Validated on a hermetic Gloas devnet (geth/lighthouse
glamsterdam-devnet-6, lighthousev8.2.0-36b70da, Gloas fork @ epoch 2), read from Loki node logs + the beacon API. Suites: ssvlabs/aetheria#130 (proposer, §4/§5/§6) + #128 (ptc, §3).POST …/validator/proposer_preferences → 200The node's ePBS implementation is spec-correct — §2–§5 are confirmed working on-chain, and §6 publishes and lands on Lodestar
v1.43.0(the live devnet-6 CL). §6 produce is missing in lighthousev8.2.0(a generic beacon-APIs#580404) but works on Lodestar. #2921 now publishes the full/unblinded envelope (SignedExecutionPayloadEnvelopeContents), which Lodestar accepts — the earlier400 "Offset out of bounds"(blinded body) is gone and the envelope lands on-chain. #2922 makes the all-operators §4 submit handle Lodestar's500 BLOCK_ERROR_ALREADY_KNOWNas success (the "BN dedupes by root" assumption doesn't hold on Lodestar; the block still lands). Residual: on the devnet §6 showed the same all-publish dedup as §4 — every operator published the identical envelope, so the redundant ones hit500 EXECUTION_PAYLOAD_ENVELOPE_ERROR_ALREADY_KNOWN(the §6 analog of #2922, tracked in #2923). That only holds because the devnet's operators share one beacon node, which produces the same envelope for all of them; with a beacon node per operator only the builder holds the payload and publishes; the other operators contribute their envelope shares and nothing else (see §6 above) — a topology not yet exercised end to end. E2E-verify context in #2920 (V1/V2/V6 — resolved there; the issue is closed, see the e2e strategy below). (§3 also shows a transient pre-fork500 IncorrectStateVarianton lighthouse's PTC-duties lookahead, self-clearing once the head crosses the fork.)E2E strategy. The hermetic Aetheria
local_testnet_gloasnet above is the ePBS proving ground, and public-network e2e rides the Sepolia/Hoodi Gloas forks on the existing SSV deployments there —GLOAS_FORK_EPOCHis read from the BN at runtime, so no new networkconfig is needed. A dedicated SSV cluster on the public glamsterdam devnet was considered and dropped: standing up a fresh SSV environment (contract deploy, registry bootstrap, operator onboarding) buys little over those two paths. #2920 is closed accordingly, its remaining verify items split across the successors: the manual Sepolia fork log-check (#2953), the node-design verification items V5/V8/V9/V11 (#2954), the Aetheria harness on Hoodi post-fork (ssvlabs/aetheria#139), and automated fork-transition coverage (ssvlabs/aetheria#141). TheGlamsterdamDevnetnetworkconfig stub has been dropped from the branch accordingly (the zero-registry-address guard it motivated inSSVConfigByNamestays).Known limitations / deferred
fork_infonow carries the Gloas fork so every other remote duty signs under the correct domain (whether that alone satisfies a live Web3Signer is an e2e-confirm — #2954 (V11)). Deferred §4 fix direction: pass the block header (HTR-equal to the block, nameable across the module boundary) instead of the Gloas block; whether Web3Signer accepts a Gloas-version block request is unverified.N=4distinct signing roots per (slot, signer) and the scheduler re-emits only on a real dependent_root change (the now-agreed SIP-94 §5 rule, still to be matched by Anchor). Only the pre-publish finality hold stays deferred. Low severity (reorg-gated, §5 is observational).dependent_root, and its ePBS proposal call is the pre-Round Timer Fix #630 GET with noBuilderConfigbody orEth-Builder-Urlecho;DataVersionGloasaliases the fork'sspec.DataVersionGloas,ForkAtEpochresolves it, and the aggregator consensus data stamps it (ePBS/Gloas: stamp DataVersionGloas into AggregatorCommitteeConsensusData.Version (drop the Fulu-cap workaround) #2998 — done on the branch). ssv-spec's Gloas arms (GetAggregateAndProofs/GetAggregateAndProofHashRoot, the reference constructor, the attestation helpers) switch from the Electra container to the fork'sGloasone on the dynssz branch (ssv-spec#643), which both go.mods pin; what remains for ePBS/Gloas: stamp DataVersionGloas into AggregatorCommitteeConsensusData.Version (drop the Fulu-cap workaround) #2998 is the Anchor byte-parity check; eth2-key-manager purpose-named root signer.Merge & follow-up checklist
Merge blockers — to clear before this PR lands on
stage:hasDutyRunning→hasDutyAssigned), domain-fetch hard-fail, and minor polishForkAtEpochGloas resolution would make attestation submits fail on Gloas slots.stagerequires every review thread resolved before mergestage→ move the ePBS commits ontostage— done: Boole landed via theintegration/boole-convergencemerge (so the merge-base was already onstageand a plain rebase sufficed), and the branch is rebased onto the lateststagetip (136 commits, linear; 134/136 patches byte-identical per range-diff, two mechanical conflict resolutions incommittee.go/aggregator.go) — and re-rebased onto the currentstagetip after the review round (one import-block conflict inaggregator_committee.go), then again once validator: purge a queue's stale messages when an idle runner starts a duty #3039–qbft: give up at the role's round cap instead of the cluster-wide cutoff #3041 landed (180 commits, 173 patches byte-identical per range-diff; the round-cap table and the queue consumer's stale-message floor reconciled with the branch's Gloas roles, the multi-slot preferences dispatcher kept off the floor)local_testnet_gloassuites (proposer+ptc) after the rebase and after eachboole-forkrefreshTracked follow-ups (not merge blockers):
local_testnet_gloasrun — gauge the re-broadcast cost (a restart mid-lookahead re-broadcasts every auth root, which peers IGNORE as recorded; once ePBS (#2901): SIP-94 conformance fixes and follow-ups #3050 item 3 lands, so does every re-emission) and the request-auth reconstruction telemetry — the overlay's suite coverage, the beacon-APIs#630-gated half included, is ssvlabs/aetheria#173; that gate lifted with ChainSafe/lodestar#9832 mergingGetGloasBeaconBlockcounts a per-node GET fallback (a pre-Round Timer Fix #630 beacon node) under the POST label (noted at the call); thread the method actually used into the label, or let it go with the GET fallback's removal (moved here from ePBS/Gloas: external-builder authentication & per-builder bid preferences — design & implementation plan #2962)sigp/anchor) — now alsoRequestAuthPartialSig(9)/DomainBuilderRequestAuthand the role-8 dual-type validation rule; run a computational SSZ/HTR cross-check (incl. the EIP-8282 five-listExecutionRequests) once canonical Gloas spec vectors exist. The node-vs-ssv-spec cross-check already matches byte for byte and root for root on the devnet-8 block and envelope, a non-empty five-listExecutionRequests, and the small containers (ePBS (#2901): SIP-94 conformance fixes and follow-ups #3050 audit, 25 Sep)gloasSpecRunnerSkipReason(the Gloas valcheck vectors already run against the node's real checkers). Not a merge blocker, by decision; recommended before Gloas is scheduled on a public network. The node-side deltas to reconcile first are ePBS (#2901): SIP-94 conformance fixes and follow-ups #3050 items 3–5. Natural to do together with the ssv-spec tag repoint aboveBuilderConfigbody assembled from config + the per-slot auth cache,Eth-Builder-Urlblock-forwarding echo gated by the §6-style owner-match, per-node GET fallback) and the BN-mediated ahead-of-timesubmitBuilderPreferences(reuses the reconstructed auth, all-operators submit). E2E gated on a beacon node shipping beacon-APIs#630 (ChainSafe/lodestar#9832 merged on 27 Aug); the GET fallback keeps older beacon nodes working.BuilderConfigbody (the direct-builder overlay when configured, else a neutral local-build config —builders: [],builder_boost_factor100); the per-node GET fallback stays for pre-Round Timer Fix #630 nodes (only on a 404/405). Round Timer Fix #630 merged and Teku went POST-only, so this became a live bug (reported in produceBlockV4 must use the POST method and send a BuilderConfig body #3002), not just a future risk.docs/MEV_CONSIDERATIONS.mdePBS rewrite (config.example.yamlalready points readers there forProposerDelayEPBS); re-evaluate proposals: mev/commit-boost driven MEV flow #2855 (ProposalSoftDeadline) against the tighter 25% proposal deadline once ePBS lands — decision + rationale recorded on proposals: mev/commit-boost driven MEV flow #2855.Eth-Consensus-Versionresponse header on Gloas produce; unify the hand-rolled JSON/SSZ HTTP helpers (sharedhttpDocore + a singlehttpStatusError); reuse the shared Gloas test-block fixture in the goclient proposer tests.