Reproducible evaluation results from running VulcanBench — an open-source LLM benchmarking harness for realistic software engineering tasks.
⚠️ CORRECTION (2026-08-25): the initial DeepSeek pass@1 below was invalidated by a VulcanBench harness bug (coverage addopts leaked into declarative grading). See CORRECTION.md. Corrected pass@1 for deepseek-chat is 19/19 (100%), not 16/19.
| Snapshot | Model | Tasks | pass@1 | Avg total | Date | Notes |
|---|---|---|---|---|---|---|
| v2-deepseek-v4-flash-2026-08-27 | deepseek-v4-flash |
10 (v2 hard tier) | 10/10 (100%) | 0.762 | 2026-08-27 | v2 realistic hard tier; all tasks verified functional=1.0 via fixed verifier |
| v1-deepseek-2026-08-corrected.json | deepseek-chat + v4-pro |
19 + 2 | 19/19 (100%) | 0.865 | 2026-08-25 | corrected (harness coverage bug) |
| v1-deepseek-chat-2026-08 | deepseek-chat |
19 | 84.2% | 0.865 | 2026-08-25 | superseded by corrected |
- Harness: VulcanBench v0.5.1 (open-source agentic SWE benchmark)
- Model:
deepseek-chatviadeepseek-v4-flash(proxied through VulcanBenchopenai:provider toapi.deepseek.com) - Sandbox: local (
--sandbox local) — no Docker sandbox image in the eval environment - Judges: disabled (
--no-judges) to cut cost ~3x - Metric definitions (per VulcanBench):
- functional: hidden failure-to-pass test suite (0.0 or 1.0)
- quality: lint + cyclomatic complexity + maintainability index
- security: bandit static analysis (no findings = 1.0)
- total: weighted composite
- Cost: exact API cost unavailable (deepseek returns
None); runtime recorded per run
git clone https://github.com/morganlinton/VulcanBench.git
cd VulcanBench
make setup && source .venv/bin/activate
# Single task
vulcanbench run --task carbyne-add-months \
--model openai:deepseek-chat \
--sandbox local --no-judges --max-steps 30
# Full suite
vulcanbench run --suite v1 --model openai:deepseek-chat \
--sandbox local --no-judges --max-concurrency 4Requires a DEEPSEEK_API_KEY (or OPENAI_API_KEY + OPENAI_BASE_URL pointing at a DeepSeek-compatible endpoint).
VulcanBench runs are recorded under ./runs/<run_id>/ (gitignored). This repo
publishes the exported result snapshots (MD + JSON) for a permanent, shareable
link without bloating run traces.
--sandbox localruns agent shell commands on the host; use--sandbox dockerfor published cross-provider numbers. Results here are functional-focused dev runs.- Security score reports
nullwhere thebandittoolchain is unavailable on the host — never a fabricated score. - The v1 suite has 86 tasks; this snapshot covers 19 (a representative spread across carbyne-tier Python/Go tasks plus several open-source-derived scenarios).
- Failures observed:
py-semver-compare,py-topo-sort-cycle,py-ttl-cache-expiry(all hidden-test Python bug-fix scenarios).
Results data and generated reports are Apache-2.0 (aligned with VulcanBench's own license). VulcanBench itself is Apache-2.0.