Skip to content

fix(profiles): use Luna judges and forward scanner reasoning - #62

Merged
vincentkoc merged 3 commits into
mainfrom
fix/clawhub-gpt6-luna-high
Sep 24, 2026
Merged

vincentkoc merged 3 commits into
mainfrom
fix/clawhub-gpt6-luna-high

Conversation

@vincentkoc

@vincentkoc vincentkoc commented Sep 24, 2026 •

Copy link
Copy Markdown
Member

Summary

Both bundled ClawHub judges still selected GPT-5.5, and Docker could drop SkillSpector's configured reasoning effort. Use GPT-6 Luna with high reasoning in clawhub and clawhub-aig, and preserve scanner model/reasoning settings across the existing environment allowlists.

  • Forward SKILLSPECTOR_MODEL and SKILLSPECTOR_REASONING_EFFORT in both profile environments, and add the missing reasoning setting to the scanner registry.
  • Forward optional A.I.G. REASONING_EFFORT alongside DEFAULT_MODEL.
  • Keep scanner gating, raw evidence, credential isolation, and existing service-tier policy unchanged.

Scope

  • Scanner adapter
  • Judge/profile/benchmark behavior
  • Docs and changelog
  • Release/CI/repo automation

This is a maintainer-authorized default update to the official bundled profiles, not a vulnerability profile proposal.

Security / Trust Impact

  • Security/trust impact explained

The requested model selection changes the ClawHub judge. Environment passthrough remains a variable-name allowlist; values are not added to Docker arguments or artifact metadata. No credentials, scanning permissions, gate rules, prompts, or output schemas change. Production/config delta is +8 lines; tests +49; docs/changelog +18.

Verification

  • Focused profile tests execute the shell command with a fake Codex and check both model and reasoning arguments, while preserving API-key precedence.
  • Docker fixture tests verify SkillSpector and A.I.G. model/effort environment forwarding and absent optional settings.
  • Hosted CI on final head: full Go suite, CLI build, npm package tests/smoke, repository script tests, and workflow lint passed.
  • All non-CLI Go packages also passed local go test -count=1 ./..., including full internal/profiles and internal/runner suites.
  • Complete local go test -count=1 ./...: stopped after the unchanged TestRunCommandPrintsHelp hung in captureStdout. Its helper writes to an OS pipe before reading; the captured Go stack was blocked in syscall.write via main.go:36 and main_test.go:1468. The help test, its output-capture helper, and help text are unchanged. Follow-up: make CLI test output capture drain concurrently on macOS.
  • go vet ./...
  • go run ./cmd/clawscan --help
  • make docs-site
  • git diff --check and gofmt
  • `go test -count=1 -timeout 30s ./cmd/clawscan -run '^TestRunCommandScannerDetailPrintsHumanReadableInfo
  • One GPT-6 Sol/high autoreview of the substantive change: no findings, patch correct (confidence 0.86). The later one-line CLI expected-environment-list correction was verified mechanically and with the exact failing test.

Dependency contracts checked: SkillSpector model resolution, SkillSpector reasoning forwarding, and Codex CLI model override / reasoning request construction. The campaign's separate local Codex Luna/high availability canary succeeded; it does not prove production API-key access.

Rollout dependency

A.I.G. high reasoning is not active in the pinned runtime. The current aig-skill-scan==0.2.2 dependency does not support REASONING_EFFORT. Its upstream support is being prepared separately; this change only forwards the optional setting and documents that limitation. An upstream release and a verified runtime dependency update are required before claiming all ClawHub scanners use high reasoning.

ClawScan publication/runtime rollout and ClawHub package adoption are separate steps. No release tags, package publication, production scans, runtime deployment, or live variables were changed.
` after updating its expected optional-environment list.

  • One GPT-6 Sol/high autoreview of the substantive change: no findings, patch correct (confidence 0.86). The later one-line CLI expected-environment-list correction was verified mechanically and with the exact failing test.

Dependency contracts checked: SkillSpector model resolution, SkillSpector reasoning forwarding, and Codex CLI model override / reasoning request construction. The campaign's separate local Codex Luna/high availability canary succeeded; it does not prove production API-key access.

Rollout dependency

A.I.G. high reasoning is not active in the pinned runtime. The current aig-skill-scan==0.2.2 dependency does not support REASONING_EFFORT. Its upstream support is being prepared separately; this change only forwards the optional setting and documents that limitation. An upstream release and a verified runtime dependency update are required before claiming all ClawHub scanners use high reasoning.

ClawScan publication/runtime rollout and ClawHub package adoption are separate steps. No release tags, package publication, production scans, runtime deployment, or live variables were changed.

Maintainer landing decision

Accept this as staged source preparation. Both bundled judges select Luna/high independently of A.I.G.; the optional A.I.G. environment allowlist does not set or enable reasoning effort. Existing A.I.G. 0.2.2 behavior remains unchanged. The full all-scanners-high rollout remains blocked on upstream support, a verified runtime dependency update, ClawScan publication, and downstream adoption.

The packaged README now matches the profile model and scanner environment settings, including the A.I.G. limitation. The final README-only delta leaves runtime and tests byte-identical to the validated 200ac10 head; docs generation passed again.

ClawSweeper rank-up move: a representative scan with the deployed credential is deferred to runtime rollout, because this PR publishes no package/image and activates no production scanner. The existing local Luna/high canary demonstrates local account availability, not production access. Production-equivalent Luna access remains a rollout gate.

@clawsweeper

clawsweeper Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

🦞👀
ClawSweeper picked this up.

Pull request received. I will update this pull request when review starts.

ClawSweeper review complete

ClawSweeper finished reviewing this revision. The review result is being finalized.

View the workflow run.

@clawsweeper

clawsweeper Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

Codex review: needs maintainer review before merge. Reviewed September 24, 2026, 12:16 AM ET / 04:16 UTC (Revision 2).

ClawSweeper review

What this changes

The branch switches two bundled ClawHub judges to GPT-6 Luna with high reasoning, forwards scanner model and reasoning settings through Docker environment allowlists, and updates tests and operator documentation.

Merge readiness

✅ Ready for maintainer review

The change remains necessary: current main still selects GPT-5.5 in the bundled ClawHub judges. The prior packaged README mismatch is fixed at this head. The MEMBER-authored PR records a staged source-landing decision, with A.I.G. support and production credential verification reserved for rollout; no introduced merge blocker was found.

Priority: P2
Reviewed head: 4e7db96c98ef33cea7f033372f8b730969d9363c

Review scores

Measure Result What it means
Overall readiness 🦐 gold shrimp (3/6) The patch and documentation are coherent, while runtime proof with deployment credentials is explicitly deferred to the accepted rollout.
Proof confidence 🦐 gold shrimp (3/6) Not applicable: The MEMBER-authored staged source change is exempt from the external-contributor proof gate. Fake-Codex shell tests and Docker-argument fixtures cover the changed profile and forwarding paths; the reported local Luna canary does not establish a production-credential ClawScan run, which remains a rollout gate. No stored-data contract changes.
Patch quality 🐚 platinum hermit (4/6) No actionable review findings were identified.

Verification

Check Result Evidence
Real behavior Not applicable Not applicable: The MEMBER-authored staged source change is exempt from the external-contributor proof gate. Fake-Codex shell tests and Docker-argument fixtures cover the changed profile and forwarding paths; the reported local Luna canary does not establish a production-credential ClawScan run, which remains a rollout gate. No stored-data contract changes.
Evidence reviewed 10 items Introduced profile behavior: The pinned PR delta changes both bundled judge model arguments and adds scanner setting names to both profile sandbox allowlists.
Still needed on main: The current-main profile blob selects GPT-5.5 and lacks the new reasoning-setting names, so main has not implemented this change.
Prior finding resolved: The final commit updates the packaged README profile table and scanner-setting catalog; the remaining GPT-5.5 example is explicitly a custom profile.
Findings None None.
Security None None.

How this fits together

ClawScan’s built-in ClawHub profiles choose scanners and an external judge for a skill or plugin target. Scanner reports feed the judge, and the run produces evidence and a verdict artifact.

flowchart LR
 A[Skill or plugin target] --> B[ClawHub profile]
 B --> C[Scanner sandbox]
 A --> C
 C --> D[Raw scanner reports]
 D --> E[Codex judge]
 B --> E
 E --> F[Verdict artifact]
Loading

Before merge

None.

Agent review details

Security

None.

Review metrics

Metric Value Why it matters
Code growth production/config +8 net lines; tests +49 net lines The production growth is confined to profile settings and allowed environment names, with focused coverage.

Technical review

Best possible solution:

Land the scoped profile preparation, then gate publication and ClawHub adoption on A.I.G. support, updated runtime pins, and a credentialed Luna scan.

Do we have a high-confidence way to reproduce the issue?

Not applicable as a profile-default and setting-forwarding change. Current-main source shows the older model and missing reasoning allowlist entries.

Is this the best way to solve the issue?

Yes. The branch uses the existing profile and environment-name allowlist mechanisms, while the documented staged rollout keeps unsupported A.I.G. reasoning separate.

AGENTS.md: found and applied where relevant.

Codex review notes: model internal, reasoning high; reviewed against 6d4573ee6187.

Labels

Label changes:

  • add P2: This is a bounded improvement to bundled security-scanning profiles with limited operator scope.
  • add merge-risk: 🚨 compatibility: The bundled judge default changes for existing profile users; the MEMBER-authored staged decision retains release and adoption checks.
  • add merge-risk: 🚨 auth-provider: Existing credentials may not have Luna access, so the accepted rollout requires production-equivalent credential verification before publication.
  • add rating: 🦐 gold shrimp: Overall readiness is 🦐 gold shrimp; proof is 🦐 gold shrimp and patch quality is 🐚 platinum hermit.
  • add status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Not applicable: The MEMBER-authored staged source change is exempt from the external-contributor proof gate. Fake-Codex shell tests and Docker-argument fixtures cover the changed profile and forwarding paths; the reported local Luna canary does not establish a production-credential ClawScan run, which remains a rollout gate. No stored-data contract changes.

Label justifications:

  • P2: This is a bounded improvement to bundled security-scanning profiles with limited operator scope.
  • merge-risk: 🚨 compatibility: The bundled judge default changes for existing profile users; the MEMBER-authored staged decision retains release and adoption checks.
  • merge-risk: 🚨 auth-provider: Existing credentials may not have Luna access, so the accepted rollout requires production-equivalent credential verification before publication.
  • rating: 🦐 gold shrimp: Overall readiness is 🦐 gold shrimp; proof is 🦐 gold shrimp and patch quality is 🐚 platinum hermit.
  • status: 👀 ready for maintainer look: ClawSweeper has no concrete contributor-facing blocker left for this PR. Not applicable: The MEMBER-authored staged source change is exempt from the external-contributor proof gate. Fake-Codex shell tests and Docker-argument fixtures cover the changed profile and forwarding paths; the reported local Luna canary does not establish a production-credential ClawScan run, which remains a rollout gate. No stored-data contract changes.

Evidence

What I checked:

Likely related people:

  • unknown: The claimed source-line change could not be verified from bounded local history. (role: source history unknown; confidence: low)
  • unknown: The claimed source-line change could not be verified from bounded local history. (role: source history unknown; confidence: low)

Rating scale

Score Internal tier Crab rank Meaning
6/6 S 🦀 challenger crab Exceptional readiness
5/6 A 🦞 diamond lobster Very strong readiness
4/6 B 🐚 platinum hermit Good normal PR; ordinary maintainer review
3/6 C 🦐 gold shrimp Useful, but confidence is limited
2/6 D 🦪 silver shellfish Proof or implementation needs work
1/6 F 🧂 unranked krab Not merge-ready
N/A NA 🌊 off-meta tidepool Rating does not apply

Overall follows the weaker of proof and patch quality.
Shiny media proof means a screenshot, video, or linked artifact directly shows the changed behavior. Runtime, network, CSP, and security claims still need visible diagnostics.

Workflow

  • ClawSweeper keeps one durable marker-backed review comment per issue or PR.
  • Re-runs edit this comment so the latest verdict, findings, and automation markers stay together instead of adding duplicate bot comments.
  • A fresh review can be triggered by eligible @clawsweeper re-review comments, exact-item GitHub events, scheduled/background review runs, or manual workflow dispatch.
  • PR/issue authors and users with repository write access can comment @clawsweeper re-review or @clawsweeper re-run on an open PR or issue to request a fresh review only.
  • Maintainers can also comment @clawsweeper review to request a fresh review only.
  • Fresh-review commands do not start repair, autofix, rebase, CI repair, or automerge.
  • Maintainer-only repair and merge flows require explicit commands such as @clawsweeper autofix, @clawsweeper automerge, @clawsweeper fix ci, or @clawsweeper address review.
  • Maintainers can comment @clawsweeper explain to ask for more context, or @clawsweeper stop to stop active automation.

History

Review history (1 earlier review cycle)
  • reviewed 2026-09-24T03:24:17.507Z sha 200ac10 :: blocked before merge. :: [P2] Sync the packaged README with the new profile default

@vincentkoc

Copy link
Copy Markdown
Member Author

@clawsweeper re-review

Please review exact head 4e7db96. The packaged README model/environment catalog mismatch is fixed. The PR body records the accepted staged source-landing decision and explicitly preserves A.I.G. runtime/release and production-equivalent credential proof as rollout gates. Runtime and tests are unchanged from 200ac10; this follow-up changes only README documentation.

@clawsweeper

clawsweeper Bot commented Sep 24, 2026 •

Copy link
Copy Markdown

🦞👀
Exact review queued.

Re-review progress:

@vincentkoc
vincentkoc marked this pull request as ready for review September 24, 2026 04:10
@vincentkoc
vincentkoc requested review from a team and Patrick-Erichsen as code owners September 24, 2026 04:10
@clawsweeper clawsweeper Bot added P2 Normal priority bug or improvement with limited blast radius. merge-risk: 🚨 auth-provider 🚨 Merging this PR could break OAuth, tokens, provider routing, model choice, or credentials. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR. labels Sep 24, 2026
@vincentkoc
vincentkoc merged commit 490bd16 into main Sep 24, 2026
12 checks passed
@vincentkoc
vincentkoc deleted the fix/clawhub-gpt6-luna-high branch September 25, 2026 11:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

merge-risk: 🚨 auth-provider 🚨 Merging this PR could break OAuth, tokens, provider routing, model choice, or credentials. merge-risk: 🚨 compatibility 🚨 Merging this PR could break existing users, config, migrations, defaults, or upgrades. P2 Normal priority bug or improvement with limited blast radius. rating: 🦐 gold shrimp Decent PR readiness signal, but merge confidence is limited. status: 👀 ready for maintainer look ClawSweeper has no concrete contributor-facing blocker left for this PR.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant