Question
Given the researched comparison contract, how should native Codex, Claude Code, and OpenCode studies be admitted progressively rather than as a default full-factorial matrix, and how should target state, prompts, human help, native tools, active working time, tokens/cost, setup friction, dead-end recovery, patches, disclosure readiness, signal confidence, cross-association, independently validated findings, and demonstrated security effect be attributed without collapsing effectiveness into one leaderboard score? Decide pilot gates, replication, invalid-run classes, stopping rules, and when a future ExploitHunter tool-combination or per-model study becomes warranted.
Question
Given the researched comparison contract, how should native Codex, Claude Code, and OpenCode studies be admitted progressively rather than as a default full-factorial matrix, and how should target state, prompts, human help, native tools, active working time, tokens/cost, setup friction, dead-end recovery, patches, disclosure readiness, signal confidence, cross-association, independently validated findings, and demonstrated security effect be attributed without collapsing effectiveness into one leaderboard score? Decide pilot gates, replication, invalid-run classes, stopping rules, and when a future ExploitHunter tool-combination or per-model study becomes warranted.