Skip to content

Gateway 2026.9.6 blocks WSL setup at guarded post-wizard restart #1498

Description

@shanselman

Problem

Current Windows/WSL setup CI repeatedly finishes the Gateway wizard, restores gateway.reload.mode=hybrid, and then fails openclaw gateway restart before the health-wait loop. Shared-fixture initialization aborts and 23 setup tests fail before their test bodies run.

The exact-main passing control used Gateway 2026.9.5 (ec9c1a1). The failing runs below used Gateway 2026.9.6 (eb377ac). This establishes a version correlation and a concrete restart-contract change, but not the precise cause of every rejected owner predicate.

The ownership guard must remain fail-closed. This issue is not a request to bypass it, delete state locks, signal an unverified process, downgrade silently, or treat a retry as proof.

Reproducible CI evidence

Use the existing isolated/disposable CI environment, not a normal user workstation. The relevant suite is OpenClaw.E2ETests.Setup.SetupAndConnectTests and its shared setup fixture.

Source / run Gateway Observed outcome
Exact main 42c562f91a42ae0251ac971c44f6ce0f8028e88d, setup job 2026.9.5 Restart exits 0 after about 22.2 seconds, followed by HTTP 200; setup suite passes.
#1490 (fix(setup): release the setup lock only after rollback finishes), setup job 2026.9.6 Serving-owner refusal; restart exits 1 before health wait.
#1479 (fix(setup): unregister only a distro this install created), setup job 2026.9.6 Same serving-owner refusal after wizard completion.
#1480 (fix(setup): approve only the pairing request setup just opened), setup job 2026.9.6 Same serving-owner refusal after wizard completion.
#1481 (fix(setup): drop leftover gateway tokens during uninstall), setup job 2026.9.6 Same serving-owner refusal after wizard completion.
#1484 (fix(connection): open the dashboard on the tunnel, not the saved URL), setup job 2026.9.6 Distinct restart-intent write contention, described below.
#1486 (fix(setup): delete only generated uninstall children), setup job 2026.9.6 Same serving-owner refusal after wizard completion.

The two workflows containing unusually long-running Tray UI jobs were subsequently cancelled normally to bound those separate waits. Their already-completed setup failures remain the evidence above. No successful CI result is claimed for those runs, and the UI-wait cause is not established.

Keep the two failure kinds separate

Most runs report:

GatewayRestartPreparationError: GATEWAY_RESTART_PREPARATION_REFUSED:
Cannot verify a live serving Gateway owner for the selected service.
Gateway was not signaled.

The dashboard run instead positively records an approximately 5,011 ms state.write admission wait and:

StateDatabaseCoordinatorContentionError:
another OpenClaw process owns state-lifecycle

GATEWAY_RESTART_PREPARATION_REFUSED:
Cannot record restart intent for the serving Gateway.
Gateway was not signaled.

That second failure is confirmed database-coordinator contention during restart-intent recording. It does not prove that the other runs had the same cause. It also does not prove that owner verification had already passed: the state-write coordinator is acquired before the target-resolution/owner-reading callback.

Source findings

  • 2026.9.5 lifecycle-core.ts, lines 501-513 performs a best-effort restart-intent write before restarting the native service.
  • 2026.9.6 lifecycle-core.ts, lines 655-665 awaits mandatory preparation before service.restart.
  • 2026.9.6 restart-intent.ts enforces a live supervised owner matching the selected native supervisor, or separately established stopped-service conditions, and requires intent recording.
  • The default producer/consumer contract appears coherent: generated Linux service metadata supplies the expected systemd unit and service markers; the publisher detects the Linux supervisor; the consumer normalizes the same unit suffix. Default HOME/profile/state-path construction and Companion's same-distro/default-user arguments also align statically.
  • Actual runtime overrides, lease state, process-start identity and the precise failed predicate are not present in the retained logs. A missing marker, wrong service, transient restart race, or identity defect has therefore not been proven.
  • The fix(setup): release the setup lock only after rollback finishes #1490 window-lock change cannot execute on this failing path: the headless setup phase fails before the tray/SetupWindow phase. Its relevant headless caller-chain blobs match the passing main control. This issue should not be used to attribute the shared failure to an unrelated PR diff.

Next diagnostic boundary

Capture once before normal rollback removes the owned fixture distro, emitting only coarse allowlisted results:

  • Selected unit/native runtime state and scope, unit-name equality, main-PID presence, process-start identity availability and restart-count changes.
  • Config/state identity equality as booleans, not paths or contents.
  • A typed owner-admission reason distinguishing unavailable owner, non-live owner, supervision mismatch, identity mismatch and coordinator contention where supported upstream.

Supported native status reads include systemctl --user show with a fixed property allowlist, ps for a validated owned PID, and openclaw gateway status --json --no-probe --timeout 5000 parsed in memory. Do not dump raw config, environment, command lines, database contents, credentials or identity files. The public status JSON does not expose the full owner lease, so it cannot by itself identify every rejected predicate.

The existing SetupPipeline.StepProgress failure event occurs before rollback and offers a diagnostic-only observer boundary. A separate harness around that pipeline would not be exact Program.Main E2E proof. Any production observer seam or upstream diagnostic change needs its own bounded review and tests.

A narrowly classified retry of the confirmed intent-recording contention may be a recovery candidate: repeat only the original guarded command within one finite total budget, never re-run the wizard/config changes, and preserve failures for other reasons. It has not been implemented or demonstrated to clear the contention, and it is not a fix for the generic owner refusals.

Expected result

A correctly identified, owned native Gateway should be restartable after setup, or should provide a precise non-secret refusal reason that permits the producer/consumer defect to be fixed. Permanent ownership mismatches must continue to fail closed. The current setup and recovery gates must pass on the intended Gateway version before claiming compatibility.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    P0Emergency: data loss, security bypass, crash loop, or unusable core runtime.clawsweeper:current-main-reproClawSweeper found a high-confidence current-main issue reproduction.clawsweeper:needs-product-decisionClawSweeper marked this issue as needing a product or behavior decision.clawsweeper:needs-security-reviewClawSweeper marked this issue as needing security-sensitive review.clawsweeper:no-new-fix-prClawSweeper does not recommend queueing a new automated fix PR for this issue.impact:securityThis issue is about security boundaries, credentials, authz, sandboxing, or sensitive data.impact:ux-release-blockerA non-technical user is blocked without terminal, logs, config, or support.issue-rating: 🦀 challenger crabExceptional issue quality: high-confidence current-main reproduction and actionable evidence.status: 🚢 actively landingA maintainer or agent is actively driving this item through implementation, validation, or merge.

    Type

    No type

    Fields

    Priority

    None yet

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions