Skip to content

fix(ci): prepare Luna high scanner settings - #3800

Draft
vincentkoc wants to merge 1 commit into
mainfrom
fix/clawhub-gpt6-luna-high
Draft

vincentkoc wants to merge 1 commit into
mainfrom
fix/clawhub-gpt6-luna-high

Conversation

@vincentkoc

@vincentkoc vincentkoc commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

What Problem This Solves

The requested scanner policy is GPT-6 Luna with high reasoning for the ClawScan judge, SkillSpector, and A.I.G. Current released scanner dependencies do not yet implement that policy end to end.

User Impact

Draft — activation is incomplete and must not merge until the dependency and live-proof checklist below is complete. This draft sets explicit model and effort values in both worker workflows. Coverage, moderation decisions, credential isolation, and exact-artifact result reuse stay unchanged. Skill Card generation is outside this change.

The independently safe preparation has landed in #3802 at 9a614dc195afbe1b6834065c6521018b8100afde: optional controls now pass through both restricted worker environments, and the mobile metadata layout assertion uses a controlled local fixture. Deployed workflow defaults and scanner pins were preserved by that merge.

Remaining Dependency Order

  1. A.I.G upstream: merge fix(skill-scan): forward optional reasoning effort to the model API Tencent/AI-Infra-Guard#676 and publish an actual aig-skill-scan PyPI version supporting REASONING_EFFORT. The existing PR is the upstream route; do not publish a duplicate. At the latest check on September 24, 2026, PR 676 was open at de2c5db39855e7e619c8e2d6098ad30b599f7084, and the latest PyPI release was 0.2.2, which ignores this setting.
  2. ClawScan maintainers: update the A.I.G runtime pin to that supported release, and publish the runtime/package containing the Luna/high profiles and environment forwarding from fix(profiles): use Luna judges and forward scanner reasoning clawscan#62, merged as 490bd167ea431697ac42f566adb06281db53c720. The latest published ClawScan is still 0.2.0, whose immutable source uses the older judge profile; source landing alone does not update the package or runtime image.
  3. ClawHub team: adopt those real releases in both .github/workflows/prepublication-publish-checks.yml and .github/workflows/security-scan-codex.yml; regenerate scripts/security/aig-worker-requirements.txt with hashes; update the matching version assertions in scripts/security/prepublication-worker-workflow.test.ts and scripts/security/security-scan-worker-workflow.test.ts. Both workflows currently retain ClawScan 0.1.8 and A.I.G 0.2.2. Keep the existing SkillSpector commit unless a verified dependency requirement changes.
  4. ClawHub team: run a small credentialed Luna/high proof through both worker paths, including the direct bundled-skill SkillSpector path. Record sanitized endpoint, model, effort, parsed result, and fail-closed behavior. Verify the active ClawScan judge and runtime actually consume the released controls; a serialized payload alone does not establish provider authorization. No production scan batch is needed.
  5. ClawHub team: complete relevant scanner/workflow tests, required CI, and a clean review of the final activation head before merging.

The existing dependency notification is #3800 (comment). No package/image release or production scan dispatch has been performed by this PR.

Why This Change Was Made

The worker environment boundary is already repaired on main. This draft now changes only the two workflow settings, their corresponding assertions, and rollout documentation. The judge remains owned by ClawScan's embedded profiles; there is no launcher shim, copied scanner implementation, or invented dependency version.

Application production TypeScript: unchanged. Workflow configuration: net +6 lines. Tests: net +6 lines. Changelog/specs: +6 lines.

Evidence And Limits

Current draft head: 3c2b00c596ae47519ed38e7ed009d7732772d509, rebased onto the landed preparation. The four workflow/test files are byte-identical to the previously reviewed 6c2c9757d5b532eb5f55ccd28ff016e791d1aff3; all worker and browser-fixture files match the landed preparation. The only conflict resolutions preserve the landed documentation and describe the requested policy as pending. git diff --check passed.

  • Historical activation proof: https://github.com/openclaw/clawhub/actions/runs/35950585240 — 67 focused scanner/workflow tests, static checks, and 7,195 unit tests passed (3 skipped); independent Sol/high autoreview was scoped-clean. This predates the rebase and does not replace final activation CI/review after dependency updates.
  • Landed preparation: final-head CI passed; exact-head review found no actionable issue or remaining before-merge move.
  • The previously failing public mobile-metadata fixture is repaired in PR 3802. Local fixture proof passed all four profile-context tests and 18 remaining public smoke tests with retries disabled. Runner-mode proof confirmed the test skips outside the local-auth runner and executes all assertions against the seeded app when enabled.
  • Pinned SkillSpector 69dcdfb was tested against a loopback HTTP server with minimum-supported langchain-openai 1.1.10 / openai 2.54.0 and currently resolved 1.6.5 / 3.19.2: serialized POST /v1/chat/completions, model=gpt-6-luna, reasoning_effort=high, response_format=json_schema; parsed output passed. No function tools, temperature, top_p, top_logprobs, or logprobs were sent. No provider call was made.
  • Codex model and effort forwarding were inspected directly in exec/src/lib.rs and core/src/client.rs. An existing-account Luna/high Codex canary succeeded; it does not validate the ClawHub scanner API credentials.

The original local A.I.G patch had separate historical tests; those results are not transferred to upstream PR 676, whose implementation differs. Use PR 676's final merged/released code for activation acceptance.

AI-assisted implementation; scoped changes and dependency contracts were inspected and validated as described above.

@clawsweeper

clawsweeper Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@vercel

vercel Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
clawhub Ready Ready Preview Sep 24, 2026 8:07am UTC

Request Review

@blacksmith-sh

This comment has been minimized.

@clawsweeper

clawsweeper Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Codex review: blocked before merge. Reviewed October 4, 2026, 1:50 PM ET / 17:50 UTC (Revision 9).

ClawSweeper review

What this changes

The branch selects GPT-6 Luna with high reasoning in both scanner-worker workflows, updates configuration assertions, and documents the pending rollout.

Merge readiness

⛔ Blocked before merge - 4 items remain

Keep open: Luna/high activation remains distinct from the merged preparation and is not implemented on current main. The previously reported scanner-version blocker remains unchanged; the upstream A.I.G change has now merged.

Priority: P2
Reviewed head: 3c2b00c596ae47519ed38e7ed009d7732772d509

Review scores

Measure Result What it means
Overall readiness 🦐 gold shrimp (3/6) The scoped preparation is useful, but the unchanged dependency mismatch prevents a complete activation.
Proof confidence 🌊 off-meta tidepool Not applicable: The MEMBER author is exempt from the ordinary external-contributor proof gate, and this diff does not materially change authority. Nevertheless, the documented activation hold explicitly requires credentialed worker verification: loopback SkillSpector serialization and an unrelated Codex account canary do not prove the installed ClawScan judge or scanner credentials execute Luna/high. No stored-data contract changes.
Patch quality 🦐 gold shrimp (3/6) 1 actionable review finding remain.

Verification

Check Result Evidence
Real behavior Not applicable Not applicable: The MEMBER author is exempt from the ordinary external-contributor proof gate, and this diff does not materially change authority. Nevertheless, the documented activation hold explicitly requires credentialed worker verification: loopback SkillSpector serialization and an unrelated Codex account canary do not prove the installed ClawScan judge or scanner credentials execute Luna/high. No stored-data contract changes.
Evidence reviewed 9 items Introduced activation settings: The pinned introduction delta changes model and reasoning settings in two workflows, matching assertions in two tests, and rollout documentation. It leaves ClawScan 0.1.8 and A.I.G 0.2.2 unchanged.
Dependency contract applies: The prepublication worker invokes ClawScan with the built-in clawhub profile and a restricted provider environment; the security worker also invokes that profile and directly runs SkillSpector for bundled skills. Thus the settings depend on the installed scanners' actual contracts, rather than environment serialization alone.
Pinned judge ignores the requested model: The v0.1.8 clawhub profile explicitly invokes codex exec with --model gpt-5.5. Its tag resolves to this commit, so changing DEFAULT_MODEL does not switch the authoritative judge to Luna. The dependency AGENTS.md also identifies built-in profiles as the profile source of truth.
Findings 1 actionable finding [P1] Pin compatible scanners before changing worker defaults
Security None None.

How this fits together

ClawHub’s security workers scan uploaded skills and packages using ClawScan and supporting scanners. Their results feed publication checks and moderation decisions.

flowchart LR
  A[Uploaded skills and packages] --> B[Security worker workflows]
  B --> C[Pinned scanner releases]
  C --> D[Supporting scanner evidence]
  D --> E[ClawScan judge]
  E --> F[Publication and moderation results]
Loading

Before merge

  • Pin compatible scanners before changing worker defaults (P1) - Both workflows still install ClawScan 0.1.8 and A.I.G 0.2.2. The installed ClawScan clawhub profile explicitly invokes --model gpt-5.5, so these settings change supporting scanner inputs without activating the requested Luna judge policy. The matching tests assert environment values while preserving those pins. Adopt compatible releases and regenerate the A.I.G hash lock, or preserve current defaults until activation is ready. This previously reported blocker remains unchanged and also applies to the security-scan workflow.
  • Resolve merge risk (P1) - Scanner credentials have not been shown to authorize Luna/high through both worker paths, including direct bundled-skill SkillSpector; provider rejection could leave scans and staged publications retrying or failed.
  • Resolve merge risk (P1) - The live PR has merge conflicts. The final activation revision needs review against current workflow routing after conflict resolution; no verified test merge is available.
  • Complete next step (P2) - Complete compatible release adoption and credentialed worker acceptance, then resolve conflicts and review the final activation head.

Findings

  • [P1] Pin compatible scanners before changing worker defaults — .github/workflows/prepublication-publish-checks.yml:123-126
Agent review details

Security

None.

Review metrics

Metric Value Why it matters
LOC +22/-4 across 6 files The branch is a focused workflow-policy change with assertions and rollout documentation.
Production versus tests application production +0; workflow configuration +6 net; tests +6 net; documentation +6 Production configuration growth supports the stated scanner policy; application implementation is unchanged.

Merge-risk options

Maintainer options:

  1. Complete the coordinated activation (recommended)
    Adopt compatible scanner releases and verify credentialed success and failure handling before enabling the new workflow defaults.
  2. Keep activation paused
    Retain existing deployed defaults while scanner release adoption and worker verification remain incomplete.

Technical review

Best possible solution:

Activate one consistent scanner policy using compatible released dependencies and verified worker credentials, while preserving current defaults until rollout is ready.

Do we have a high-confidence way to reproduce the issue?

Yes, for the configuration defect: both workers select the installed clawhub profile, whose immutable v0.1.8 source explicitly fixes the judge model to gpt-5.5. No credentialed runtime failure was reproduced.

Is this the best way to solve the issue?

No, the environment-only activation is incomplete. Coordinated dependency adoption is the best fix; a copied judge launcher would duplicate the authoritative ClawScan profile, while optional environment forwarding already landed separately.

Full review comments:

  • [P1] Pin compatible scanners before changing worker defaults — .github/workflows/prepublication-publish-checks.yml:123-126
    Both workflows still install ClawScan 0.1.8 and A.I.G 0.2.2. The installed ClawScan clawhub profile explicitly invokes --model gpt-5.5, so these settings change supporting scanner inputs without activating the requested Luna judge policy. The matching tests assert environment values while preserving those pins. Adopt compatible releases and regenerate the A.I.G hash lock, or preserve current defaults until activation is ready. This previously reported blocker remains unchanged and also applies to the security-scan workflow.
    Confidence: 0.99

Overall correctness: patch is incorrect
Overall confidence: 0.98

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning medium; reviewed against 0180b561dfef.

Labels

Label changes:

No label changes.

Label justifications:

  • P2: This is a bounded scanner-policy rollout held in draft, without evidence of a current urgent production regression.
  • merge-risk: 🚨 compatibility: Changing provider defaults without verified scanner credential access can break existing scan execution.
  • merge-risk: 🚨 automation: The workflow settings enable an incomplete scanner policy that configuration-only assertions cannot validate.
  • rating: 🦐 gold shrimp: Overall readiness is 🦐 gold shrimp; proof is 🌊 off-meta tidepool and patch quality is 🦐 gold shrimp.
  • status: ⏳ waiting on author: ClawSweeper has contributor-facing work open and is waiting for author action. Not applicable: The MEMBER author is exempt from the ordinary external-contributor proof gate, and this diff does not materially change authority. Nevertheless, the documented activation hold explicitly requires credentialed worker verification: loopback SkillSpector serialization and an unrelated Codex account canary do not prove the installed ClawScan judge or scanner credentials execute Luna/high. No stored-data contract changes.

Evidence

What I checked:

Likely related people:

  • Patrick-Erichsen: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • vincentkoc: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)
  • boy-hack: Suggested for follow-up; no historical authorship or introduction is verified. (role: unverified routing candidate; confidence: low)

Rank-up moves

Optional improvements that raise the rating; they are not merge blockers.

  • Adopt compatible released scanners, regenerate A.I.G hashes, and update both workflows’ version assertions.
  • Record redacted credentialed Luna/high success and failure handling through both workers, including bundled-skill SkillSpector.
  • Resolve conflicts and validate the final activation revision.

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (8 earlier review cycles)
  • reviewed 2026-09-24T03:28:32.478Z sha 6c2c975 :: blocked before merge. :: [P2] Update the scanner pins before enabling Luna/high
  • reviewed 2026-09-24T03:48:09.467Z sha 6c2c975 :: blocked before merge. :: [P2] Update scanner pins before enabling Luna/high
  • reviewed 2026-09-24T08:10:45.403Z sha 3c2b00c :: blocked before merge. :: [P1] Pin compatible scanners before switching production defaults
  • reviewed 2026-09-30T12:03:30.017Z sha 3c2b00c :: blocked before merge. :: [P1] Pin compatible scanners before changing worker defaults
  • reviewed 2026-10-01T13:50:35.039Z sha 3c2b00c :: blocked before merge. :: [P1] Pin compatible scanners before changing worker defaults
  • reviewed 2026-10-02T20:29:43.956Z sha 3c2b00c :: blocked before merge. :: [P1] Pin compatible scanners before changing worker defaults
  • reviewed 2026-10-03T12:53:35.600Z sha 3c2b00c :: blocked before merge. :: [P1] Pin compatible scanners before changing worker defaults
  • reviewed 2026-10-04T04:52:53.554Z sha 3c2b00c :: blocked before merge. :: [P1] Pin compatible scanners before changing worker defaults

@clawsweeper clawsweeper Bot added P2 Normal backlog priority with limited blast radius. merge-risk: 🚨 automation 🚨 Merging this PR could break CI, automerge, proof capture, label sync, or automation. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action. labels Sep 24, 2026
@vincentkoc

Copy link
Copy Markdown
Member Author

@Patrick-Erichsen we’re using Tencent/AI-Infra-Guard#676 for ClawHub’s Luna/high activation. Once it is merged and published, please share the aig-skill-scan version and release link here so we can update the hash lock and ClawScan runtime pin. Activation stays draft pending that release and scanner verification; the optional worker plumbing will land separately with current scan defaults preserved.

@vincentkoc
vincentkoc force-pushed the fix/clawhub-gpt6-luna-high branch 2 times, most recently from 6c2c975 to 3c2b00c Compare September 24, 2026 08:04
@clawsweeper clawsweeper Bot added the merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. label Sep 30, 2026

This branch was successfully deployed

1 active deployment
Preview – clawhub — 3c2b00c5 Deployed Sep 24, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 automation 🚨 Merging this PR could break CI, automerge, proof capture, label sync, or automation. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal backlog priority with limited blast radius. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: ⏳ waiting on author ClawSweeper has contributor-facing work open and is waiting for author action.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant