Skip to content

Document SWE-bench 5-task pilot results - #33

Merged
Bjjj834 merged 1 commit into
mainfrom
docs/swebench-pilot-results
Jul 14, 2026
Merged

Document SWE-bench 5-task pilot results#33
Bjjj834 merged 1 commit into
mainfrom
docs/swebench-pilot-results

Conversation

@Bjjj834

@Bjjj834 Bjjj834 commented Jul 14, 2026

Copy link
Copy Markdown
Contributor

This PR freezes the completed SWE-bench Verified 5-task pilot result.

It includes:

  • docs/swebench_pilot_results.md: human-readable research note
  • data/swebench_results_claude-code_smoke-v1.json: safe converted CodeBench results JSON

Pilot setup

  • SWE-bench Verified first 5 deterministic instances
  • Claude Code CLI
  • 1 attempt per task
  • acceptEdits mode
  • official SWE-bench harness evaluation

Headline result

  • 1/5 resolved
  • reliability@1 = 0.20
  • mean hidden-test pass rate = 0.8049
  • mean regression rate = 0.1539

Important framing

This is a small real-world pilot / smoke experiment, not a paper-level benchmark. It provides real-world evidence consistent with the paper's concern that partial test pass rates can overstate strict task-level resolution.

🤖 Generated with Claude Code

@Bjjj834
Bjjj834 merged commit 30d3c91 into main Jul 14, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant