Skip to content

fix(connection): a rejected shared token drops the bootstrap token and SSH tunnel - #1509

Open
SebTardif wants to merge 11 commits into
openclaw:mainfrom
SebTardif:fix/f082-preserve-bootstrap-on-failed-token
Open

SebTardif wants to merge 11 commits into
openclaw:mainfrom
SebTardif:fix/f082-preserve-bootstrap-on-failed-token

Conversation

@SebTardif

@SebTardif SebTardif commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

What Problem This Solves

Fixes: connecting with a rejected shared token removes the saved bootstrap token and SSH tunnel before the new token is checked, when no device token is stored yet. A gateway that rejects the token after the connect call has already returned also left that rejected token saved. A newer connection that starts during that wait could be overwritten, and a settings callback could target the edited gateway instead of the gateway that was active before the attempt.

User Impact

User impact: a failed shared-token connect leaves the setup-code bootstrap credential and the saved SSH tunnel in place. Saved tray settings are restored from the gateway that was active before the attempt. If the operator was already connected, that previous connection is opened again. A newer connection that takes over during the wait keeps its own record.

Why This Change Was Made

A stored device token still uses the existing pre-replacement check. When the record only has a bootstrap token or an SSH tunnel, the shared token is saved and connected on the normal path. If that operator connect fails, including an auth failure that arrives while the operator is still connecting, the previous record and active id are restored only when this attempt still owns the connection generation. If the wait ends and the operator is still connecting, that attempt is disconnected and the record is rolled back. The previous connection is opened only after the operator is Idle, and the restore waits for a terminal result. When rollback leaves no active gateway, the settings callback restores the snapshot taken before the rejected record was applied. That capture runs after ConnectWithSharedTokenAsync takes the transition lock, so a second request cannot snapshot settings before the first transaction has started. Each connectSharedToken call captures its own settings attempt and passes that object into the callback. A second call cannot replace the first call's snapshot. The app captures the snapshot before the candidate callback. That callback keeps it after a successful sync. The rollback callback reuses it. If runtime tunnel reconciliation fails on both attempts during that rollback, the pre-attempt settings are written back before the callback throws. The commit callback then runs with the record for the prior active gateway. If saving those settings or the runtime tunnel fails, the callback retries the restored gateway and reports an out-of-sync error instead of writing the rejected snapshot back. A connection that was already live is started again. A successful connect still clears a stale bootstrap token.

Evidence

Head 8c191b3938aac911dc22fa513d4e38c94d3f4602.

Connection tests on that head: Passed 808, Failed 0, Skipped 1, Total 809.

The live gateway trace below was recorded on parent 41834892e8f186b0951600e2a0066532115486e7. This head adds the unfinished-handshake rollback and the settings recovery.

The live check used an isolated registry and the gateway already listening at ws://127.0.0.1:18789. The previous shared token was the installed gateway token. Its value is omitted. The replacement token was rejected-shared-token-not-real (30 characters). The device token stored by the first connect was cleared before the replacement so the pre-replacement check did not skip the rollback.

first_connected=True state=Connected
durable_tokens_cleared=True
outcome=ConnectionFailed committed_flag=False
callbacks=cleared-bootstrap,has-bootstrap
bootstrap_restored=True
shared_len=48
operator=Connecting
later_operator=Connected later_bootstrap_restored=True later_shared_len=48

The stored shared-token length after the failure was 48 again, which is the previous token. The operator was Connecting when the call returned, and Connected again three seconds later.

Change Type

  • Bug fix
  • Feature
  • Refactor
  • Docs or instructions
  • Tests or validation
  • Security hardening
  • Chore or infrastructure

Scope

  • Tray or WinUI UX
  • Windows node capability
  • Local MCP or winnode
  • Gateway, connection, or pairing
  • Setup or onboarding
  • Permissions, privacy, or security
  • Tests, CI, or docs

Required proof pools

  • windows-wsl-gateway-e2e: saved settings, SSH tunnel, and prior live operator recovery after a rejected shared token.

Validation

Head 8c191b3938aac911dc22fa513d4e38c94d3f4602.

  • ./build.ps1: exit 0.
  • dotnet test ./tests/OpenClaw.Shared.Tests/OpenClaw.Shared.Tests.csproj: Passed 4107, Failed 0, Skipped 32, Total 4139.
  • dotnet test ./tests/OpenClaw.Tray.Tests/OpenClaw.Tray.Tests.csproj: Passed 3071, Failed 5, Total 3076. The five failures are the LF source-contract mismatch tracked in fix(tests): tray source checks fail when the checkout is LF #1518.
  • dotnet test ./tests/OpenClaw.Connection.Tests/OpenClaw.Connection.Tests.csproj: Passed 808, Failed 0, Skipped 1, Total 809.

Real Behavior Proof

  • Behavior or issue addressed: A rejected shared token cleared the bootstrap token and left the new record active. A later auth failure did the same because the connect method returned while the operator was still connecting. The operator that was already connected needed to come back after that failure.
  • Real environment tested: Windows worktree, head 41834892e8f186b0951600e2a0066532115486e7, gateway ws://127.0.0.1:18789 (HTTP 200 on http://127.0.0.1:18789/). The registry directory was a new temp directory, not the installed tray config.
  • Exact steps or command run after this patch: GatewayConnectionManager.ConnectAsync with the installed gateway token, device pairing approved, operator reached Connected. The stored device token was cleared. ConnectWithSharedTokenAsync then used rejected-shared-token-not-real and a commit callback. The process read the operator state again three seconds later.
  • Evidence after fix: Outcome was ConnectionFailed. GatewayCommitted was false. The callback sequence was cleared-bootstrap,has-bootstrap. The bootstrap token was proof-bootstrap-not-real again. The stored shared-token length was 48, matching the previous token. The operator was Connecting at return and Connected three seconds later.
  • Observed result after fix: The rejected token did not stay saved, and the operator that was already connected came back.
  • Screenshot or artifact links verified? No
  • Not verified / blocked: A live SSH tunnel. The guest SSH service is inactive, so this trace had no tunnel to restore. A fresh setup wizard. A second live trace on head 8c191b3938aac911dc22fa513d4e38c94d3f4602.
  • What was not tested: An SSH session on this gateway. ConnectWithSharedTokenAsync_RejectedTokenRestoresPriorLiveConnection covers the saved tunnel. ConnectWithSharedTokenAsync_UnfinishedHandshakeRollsBack covers a handshake that stays connecting. SynchronizeSettings_TunnelFailure_KeepsCommittedGatewaySettings covers a tunnel reconcile failure. ConnectWithSharedTokenAsync_NewerGenerationSkipsRollback and ConnectWithSharedTokenAsync_RejectedTokenRestoresPriorActiveGatewaySettings cover the generation check and the prior active gateway callback.

Security Impact

  • New permissions or capabilities? No
  • Secrets or tokens handling changed? Yes
  • New or changed network calls? No
  • Command or tool execution surface changed? No
  • Data access scope changed? No
  • If any answer is Yes, explain the risk and mitigation: A new shared token is validated before it replaces a bootstrap token or SSH tunnel when a device token is already stored. When no device token is stored, a failed connect, including an auth failure that arrives after connect returns, rolls the registry record back and runs the commit callback with the prior active gateway. The rollback is skipped when a newer connection owns the generation. Device tokens are not cleared by this path.

Compatibility and Migration

  • Backward compatible? Yes
  • Config or environment changes? No
  • Migration needed? No

Review Conversations

  • I replied to or resolved every bot review conversation addressed by this PR.
  • I left unresolved only conversations that still need maintainer judgment.

… rejected

Connect with a shared token cleared the stored bootstrap token and SSH
tunnel before the new token was checked, whenever no device token existed.
Validate first, and restore the previous record if the operator connect fails.

- Run the existing pre-replacement check when a bootstrap token or tunnel is stored
- Roll the registry back when that connect fails and the rollback save succeeds

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
@clawsweeper

clawsweeper Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@clawsweeper clawsweeper Bot added P1 Urgent regression or broken agent/channel workflow affecting real users now. merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. rating: 🦪 silver shellfish Thin PR readiness signal; proof, validation, or implementation needs work. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask. labels Sep 24, 2026
@clawsweeper

clawsweeper Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codex review: needs real behavior proof before merge. Reviewed September 26, 2026, 2:14 PM ET / 18:14 UTC (Revision 11).

ClawSweeper review

What this changes

The branch restores the previous gateway record, tray settings, SSH tunnel configuration, and live operator connection after a shared-token connection fails.

Merge readiness

⛔ Blocked before merge - 5 items remain

Current main still leaves a rejected shared token committed when authentication fails after the connect call returns, so this PR remains useful. The latest commit addresses the prior settings-capture finding. Merge readiness still depends on current-head recovery proof and a decision about the new 15-second handshake cutoff.

Priority: P1
Reviewed head: 8c191b3938aac911dc22fa513d4e38c94d3f4602
Owner decision: Required. See Decision needed.

Review scores

Measure Result What it means
Overall readiness 🦐 gold shrimp (3/6) The rollback is focused and well tested, while real proof predates the final settings change and the timeout's upgrade effect remains undecided.
Proof confidence 🦐 gold shrimp (3/6) Needs stronger real behavior proof before merge: The earlier real-gateway trace shows rejected-token recovery and operator reconnection, but the settings callback changed afterward and that trace had no active SSH tunnel. Current-head redacted terminal output or runtime logs should show the saved settings, active tunnel, and prior operator connection after rejection; redact private endpoints and credentials. Updating the PR body should trigger re-review, or a maintainer can comment @clawsweeper re-review. No stored-data format changes are introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Needs proof Needs stronger real behavior proof before merge: The earlier real-gateway trace shows rejected-token recovery and operator reconnection, but the settings callback changed afterward and that trace had no active SSH tunnel. Current-head redacted terminal output or runtime logs should show the saved settings, active tunnel, and prior operator connection after rejection; redact private endpoints and credentials. Updating the PR body should trigger re-review, or a maintainer can comment @clawsweeper re-review. No stored-data format changes are introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.
Evidence reviewed 8 items Current main behavior: Current main saves the replacement record and calls ConnectCoreAsync, then checks only an immediately reported Error; a later authentication failure has no rollback in this method.
Introduced rollback: The introduced path waits for a terminal handshake state and restores the prior registry record, settings callback, and previously live connection on failure.
Prior finding addressed: The latest head captures each request's settings snapshot after acquiring the connection transition lock, addressing the previous review's capture-order finding.
Findings None None.
Security None None.

How this fits together

The connection manager takes gateway credentials from the tray and stores the active gateway record before opening the operator connection. The tray service then saves matching settings and reconciles the runtime SSH tunnel.

flowchart LR
A[Shared token request] --> B[Gateway record]
B --> C[Tray settings and SSH tunnel]
C --> D[Operator handshake]
D --> E{Connected or failed?}
E -->|Connected| F[New gateway active]
E -->|Failed| G[Restore prior state]
Loading

Decision needed

Question Recommendation
Should a shared-token replacement with an existing bootstrap credential or SSH tunnel be rolled back after 15 seconds of Connecting even when the gateway has not rejected it? Preserve slow-handshake compatibility: Change the flow so an unfinished but still viable handshake does not discard the replacement solely because 15 seconds elapsed.

Why: The cutoff changes an existing asynchronous connection path, and source inspection cannot determine the acceptable timeout for users with slow gateways.

Before merge

  • Add real behavior proof - Needs stronger real behavior proof before merge: The earlier real-gateway trace shows rejected-token recovery and operator reconnection, but the settings callback changed afterward and that trace had no active SSH tunnel. Current-head redacted terminal output or runtime logs should show the saved settings, active tunnel, and prior operator connection after rejection; redact private endpoints and credentials. Updating the PR body should trigger re-review, or a maintainer can comment @clawsweeper re-review. No stored-data format changes are introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.
  • Resolve merge risk (P1) - Current-head real gateway proof does not yet show saved settings, an active SSH tunnel, and the prior operator connection recovering together after rejection.
  • Resolve merge risk (P1) - The new 15-second cutoff disconnects a still-connecting replacement even without an authentication rejection; a legitimately slow gateway that previously completed asynchronously may fail to connect.
  • Complete next step (P2) - Provide current-head rejected-token recovery proof with an active SSH tunnel and prior operator connection, address the overlapping-request regression coverage, and obtain a decision on the 15-second handshake cutoff before merge.
  • Resolve maintainer decision - Resolve the maintainer decision shown above before merge.
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Code growth production +348/-16 lines; tests +499/-4 lines The production growth implements a multi-state rollback and has substantial focused test coverage.

Merge-risk options

Maintainer options:

  1. Preserve slow connections (recommended)
    Keep rollback for rejected credentials while allowing a viable handshake to finish beyond the fixed wait.
  2. Accept a 15-second cutoff
    Land the cutoff only after the connection owner accepts its upgrade impact and current-head proof covers normal and slow recovery paths.

Technical review

Best possible solution:

Keep transactional rollback, then demonstrate the complete recovery path on this head and establish an accepted handshake timeout behavior for existing slow connections.

Do we have a high-confidence way to reproduce the issue?

Yes at the source level: current main commits the candidate before connecting and has no rollback for an authentication failure arriving after the call returns. This review did not execute the gateway path.

Is this the best way to solve the issue?

Unclear pending the timeout decision and current-head tunnel proof. Restoring the previous record through the connection owner is a maintainable direction.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning medium; reviewed against 5a59535216ee.

Labels

Label changes:

No label changes.

Label justifications:

  • P1: A rejected credential can leave a gateway setup and its live connection unusable on current main.
  • merge-risk: 🚨 compatibility: The introduced handshake cutoff can stop a connection that current main allows to complete asynchronously.
  • merge-risk: 🚨 availability: The changed rollback must restore a live operator connection and SSH tunnel after rejecting a replacement credential.
  • rating: 🦐 gold shrimp: Overall readiness is 🦐 gold shrimp; proof is 🦐 gold shrimp and patch quality is 🐚 platinum hermit.
  • status: 📣 needs proof: The PR needs real behavior proof before ClawSweeper can clear the contributor ask. Needs stronger real behavior proof before merge: The earlier real-gateway trace shows rejected-token recovery and operator reconnection, but the settings callback changed afterward and that trace had no active SSH tunnel. Current-head redacted terminal output or runtime logs should show the saved settings, active tunnel, and prior operator connection after rejection; redact private endpoints and credentials. Updating the PR body should trigger re-review, or a maintainer can comment @clawsweeper re-review. No stored-data format changes are introduced. After adding proof, update the PR body; ClawSweeper should re-review automatically. If it does not, the PR author or someone with repository write access can comment @clawsweeper re-review.

Evidence

What I checked:

Likely related people:

  • Barbara Kudiess: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • SebTardif: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • Scott Hanselman: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rank-up moves

Optional improvements that raise the rating; they are not merge blockers.

  • Post redacted current-head gateway output showing rejected-token recovery of saved settings, an active SSH tunnel, and the prior operator connection.
  • Resolve the 15-second handshake cutoff with the connection owner and cover the chosen behavior.
  • Add the previously requested overlapping-request regression where the first request commits before the second request's rollback reconciliation fails.

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (10 earlier review cycles; latest 8 shown)
  • reviewed 2026-09-25T21:00:08.049Z sha 5ceb876 :: needs real behavior proof before merge. :: [P1] Check the connection generation before rolling back | [P1] Restore settings from the prior active gateway
  • reviewed 2026-09-25T22:55:45.052Z sha 4183489 :: needs real behavior proof before merge. :: [P1] Preserve tray settings when rollback reconciliation fails | [P2] Do not report success while the handshake is unfinished
  • reviewed 2026-09-25T23:54:24.336Z sha 4c0e3a4 :: needs real behavior proof before merge. :: [P1] Stop the unfinished replacement before restoring the prior gateway | [P2] Restore tray settings when no gateway was previously active
  • reviewed 2026-09-26T00:29:10.697Z sha af05dfd :: needs real behavior proof before merge. :: [P1] Restore prior settings when the initial tunnel callback fails
  • reviewed 2026-09-26T01:46:24.268Z sha 73cc973 :: needs real behavior proof before merge. :: [P1] Preserve the pre-attempt settings snapshot through rollback
  • reviewed 2026-09-26T14:08:58.795Z sha b5c2753 :: needs real behavior proof before merge. :: [P1] Keep the original settings snapshot through rejection rollback
  • reviewed 2026-09-26T15:14:07.660Z sha 914ef13 :: needs real behavior proof before merge. :: [P1] Keep rollback snapshots scoped to each shared-token request
  • reviewed 2026-09-26T17:26:58.848Z sha da012c5 :: needs real behavior proof before merge. :: [P1] Capture settings after the request enters the connection transaction

A live pre-check on every stored bootstrap token rejected a normal
shared-token save before the connection manager could connect. Keep that
save path, and roll the registry back only when the operator connect fails.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
@karkarl

karkarl commented Sep 25, 2026

Copy link
Copy Markdown
Collaborator

Global triage: HOLD_FOR_AUTHOR. Take confidence 2%; recommendation confidence 99%; effort medium; risk high.

Two verified rollback gaps block this exact head:

  1. GatewayConnectionManager.cs restores only GatewayRegistry after the MCP callback has already persisted candidate settings and reconciled the runtime tunnel. It can return GatewayCommitted=false while tray settings still reference the rejected candidate and the prior tunnel remains stopped.
  2. The previous gateway is disconnected before the replacement attempt, but successful registry rollback does not reconnect a previously live operator connection.

Please make registry, settings, runtime-tunnel state, and prior live connection recovery one transaction, with focused success and rollback-save-failure tests.

Required proof is windows-wsl-gateway-e2e, not none: use a real authentication rejection and show that the registry, active ID, saved settings, SSH tunnel, and prior operator connection are preserved or restored. Existing proof uses an unreachable port and predates this exact head.

The three failed E2E lanes fail identically on the exact base during setup restart preparation; CI Gate is derivative. They are not attributed to this patch, but required recovery proof remains unavailable.

ConnectWithSharedTokenAsync treated a still-connecting operator as success, so a gateway auth failure that arrived later left the rejected token committed. When a bootstrap or SSH setup credential is at risk, wait for that failure, then restore saved settings and reconnect a connection that was already live.

Connection tests: 804 passed, 1 skipped. Shared tests: 4107 passed, 32 skipped. Tray tests: 3067 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
… late

After the deferred auth wait, roll the registry back only when this attempt still owns the connection generation. When another gateway was active, the settings callback receives that gateway's record so saved settings match the restored active id.

Connection tests: 806 passed, 1 skipped. Shared tests: 4106 passed and 32 skipped, then the one dispose failure passed alone. Tray tests: 3067 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
A shared-token replacement that is still connecting after the wait rolls the rejected token back instead of returning success. If settings or the runtime tunnel fail while applying the restored gateway, the callback retries that gateway and reports an out-of-sync error instead of writing the rejected snapshot back.

Connection tests: 807 passed, 1 skipped. Shared tests: 4107 passed, 32 skipped. Tray tests: 3068 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
@clawsweeper clawsweeper Bot added rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🦪 silver shellfish Thin PR readiness signal; proof, validation, or implementation needs work. labels Sep 25, 2026
When the handshake wait ends while the operator is still connecting, disconnect that attempt so its generation is cancelled and the operator is Idle before the previous gateway is opened again. The restore then waits for a terminal result. When rollback clears the active gateway, the settings callback restores the snapshot taken before the rejected record was applied.

Connection tests: 808 passed, 1 skipped. Shared tests: 4107 passed, 32 skipped. Tray tests: 3069 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
ApplySettings saves the candidate gateway before runtime tunnel reconciliation. If that reconciliation fails on the retry as well, write the snapshot taken before the apply back to settings and still throw, so the connection manager can roll the registry back without leaving the rejected gateway saved.

Tray tests: 3069 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518. Shared tests: 4107 passed, 32 skipped.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
The rollback callback was capturing settings again after the candidate had been saved, so a second tunnel failure wrote the rejected gateway back. Capture the snapshot once per shared-token attempt and reuse it when the rollback callback fails.

Tray tests: 3070 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518. Shared tests: 4107 passed, 32 skipped.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
@clawsweeper clawsweeper Bot added rating: 🦪 silver shellfish Thin PR readiness signal; proof, validation, or implementation needs work. and removed rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. labels Sep 26, 2026
BeginSharedTokenSettingsAttempt captures the snapshot, so the candidate callback must not clear it on success. The rollback callback still uses that snapshot when tunnel reconciliation fails twice.

Tray tests: 3071 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518. Shared tests: 4107 passed, 32 skipped.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
@clawsweeper clawsweeper Bot added rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. and removed rating: 🦪 silver shellfish Thin PR readiness signal; proof, validation, or implementation needs work. labels Sep 26, 2026
The settings snapshot lived on the tray service and was captured before the connection manager took its lock, so a second connect could replace it. Each connect now captures its own attempt and passes that object into the settings callback.

Shared tests: 4107 passed, 32 skipped. Tray suite: 3071 passed, 6 failed. Five failures are the LF source-contract mismatch tracked in openclaw#1518. CredentialReplacementFlows_DoNotBlindlyClearDeviceTokens then passed after its expected call was updated.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>
…lock

The tray captured the settings snapshot before ConnectWithSharedTokenAsync took the transition lock, so a second request could snapshot a baseline from before the first transaction. The capture now runs inside that lock, and the second request cannot start until the first has entered it.

Connection tests: 809 passed, 1 skipped. Shared tests: 4107 passed, 32 skipped. Tray tests: 3072 passed, 5 failed on the LF source-contract mismatch tracked in openclaw#1518.

Signed-off-by: Sebastien Tardif <SebTardif@ncf.ca>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 availability 🚨 Merging this PR could cause crashes, hangs, restart loops, stalls, or process outages. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P1 Urgent regression or broken agent/channel workflow affecting real users now. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 📣 needs proof The PR needs real behavior proof before ClawSweeper can clear the contributor ask.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants