-
Notifications
You must be signed in to change notification settings - Fork 306
Gateway 2026.9.6 blocks WSL setup at guarded post-wizard restart #1498
Copy link
Copy link
Open
Labels
P0Emergency: data loss, security bypass, crash loop, or unusable core runtime.Emergency: data loss, security bypass, crash loop, or unusable core runtime.clawsweeper:current-main-reproClawSweeper found a high-confidence current-main issue reproduction.ClawSweeper found a high-confidence current-main issue reproduction.clawsweeper:needs-product-decisionClawSweeper marked this issue as needing a product or behavior decision.ClawSweeper marked this issue as needing a product or behavior decision.clawsweeper:needs-security-reviewClawSweeper marked this issue as needing security-sensitive review.ClawSweeper marked this issue as needing security-sensitive review.clawsweeper:no-new-fix-prClawSweeper does not recommend queueing a new automated fix PR for this issue.ClawSweeper does not recommend queueing a new automated fix PR for this issue.impact:securityThis issue is about security boundaries, credentials, authz, sandboxing, or sensitive data.This issue is about security boundaries, credentials, authz, sandboxing, or sensitive data.impact:ux-release-blockerA non-technical user is blocked without terminal, logs, config, or support.A non-technical user is blocked without terminal, logs, config, or support.issue-rating: 🦀 challenger crabExceptional issue quality: high-confidence current-main reproduction and actionable evidence.Exceptional issue quality: high-confidence current-main reproduction and actionable evidence.status: 🚢 actively landingA maintainer or agent is actively driving this item through implementation, validation, or merge.A maintainer or agent is actively driving this item through implementation, validation, or merge.
Description
Activity
Metadata
Metadata
Assignees
Labels
P0Emergency: data loss, security bypass, crash loop, or unusable core runtime.Emergency: data loss, security bypass, crash loop, or unusable core runtime.clawsweeper:current-main-reproClawSweeper found a high-confidence current-main issue reproduction.ClawSweeper found a high-confidence current-main issue reproduction.clawsweeper:needs-product-decisionClawSweeper marked this issue as needing a product or behavior decision.ClawSweeper marked this issue as needing a product or behavior decision.clawsweeper:needs-security-reviewClawSweeper marked this issue as needing security-sensitive review.ClawSweeper marked this issue as needing security-sensitive review.clawsweeper:no-new-fix-prClawSweeper does not recommend queueing a new automated fix PR for this issue.ClawSweeper does not recommend queueing a new automated fix PR for this issue.impact:securityThis issue is about security boundaries, credentials, authz, sandboxing, or sensitive data.This issue is about security boundaries, credentials, authz, sandboxing, or sensitive data.impact:ux-release-blockerA non-technical user is blocked without terminal, logs, config, or support.A non-technical user is blocked without terminal, logs, config, or support.issue-rating: 🦀 challenger crabExceptional issue quality: high-confidence current-main reproduction and actionable evidence.Exceptional issue quality: high-confidence current-main reproduction and actionable evidence.status: 🚢 actively landingA maintainer or agent is actively driving this item through implementation, validation, or merge.A maintainer or agent is actively driving this item through implementation, validation, or merge.
Type
Fields
Priority
None yet
Projects
- StatusShow more project fieldsBacklog
Problem
Current Windows/WSL setup CI repeatedly finishes the Gateway wizard, restores
gateway.reload.mode=hybrid, and then failsopenclaw gateway restartbefore the health-wait loop. Shared-fixture initialization aborts and 23 setup tests fail before their test bodies run.The exact-main passing control used Gateway 2026.9.5 (
ec9c1a1). The failing runs below used Gateway 2026.9.6 (eb377ac). This establishes a version correlation and a concrete restart-contract change, but not the precise cause of every rejected owner predicate.The ownership guard must remain fail-closed. This issue is not a request to bypass it, delete state locks, signal an unverified process, downgrade silently, or treat a retry as proof.
Reproducible CI evidence
Use the existing isolated/disposable CI environment, not a normal user workstation. The relevant suite is
OpenClaw.E2ETests.Setup.SetupAndConnectTestsand its shared setup fixture.42c562f91a42ae0251ac971c44f6ce0f8028e88d, setup jobThe two workflows containing unusually long-running Tray UI jobs were subsequently cancelled normally to bound those separate waits. Their already-completed setup failures remain the evidence above. No successful CI result is claimed for those runs, and the UI-wait cause is not established.
Keep the two failure kinds separate
Most runs report:
The dashboard run instead positively records an approximately 5,011 ms
state.writeadmission wait and:That second failure is confirmed database-coordinator contention during restart-intent recording. It does not prove that the other runs had the same cause. It also does not prove that owner verification had already passed: the state-write coordinator is acquired before the target-resolution/owner-reading callback.
Source findings
service.restart.Next diagnostic boundary
Capture once before normal rollback removes the owned fixture distro, emitting only coarse allowlisted results:
Supported native status reads include
systemctl --user showwith a fixed property allowlist,psfor a validated owned PID, andopenclaw gateway status --json --no-probe --timeout 5000parsed in memory. Do not dump raw config, environment, command lines, database contents, credentials or identity files. The public status JSON does not expose the full owner lease, so it cannot by itself identify every rejected predicate.The existing
SetupPipeline.StepProgressfailure event occurs before rollback and offers a diagnostic-only observer boundary. A separate harness around that pipeline would not be exactProgram.MainE2E proof. Any production observer seam or upstream diagnostic change needs its own bounded review and tests.A narrowly classified retry of the confirmed intent-recording contention may be a recovery candidate: repeat only the original guarded command within one finite total budget, never re-run the wizard/config changes, and preserve failures for other reasons. It has not been implemented or demonstrated to clear the contention, and it is not a fix for the generic owner refusals.
Expected result
A correctly identified, owned native Gateway should be restartable after setup, or should provide a precise non-secret refusal reason that permits the producer/consumer defect to be fixed. Permanent ownership mismatches must continue to fail closed. The current setup and recovery gates must pass on the intended Gateway version before claiming compatibility.