Skip to content

Repository files navigation

VulcanBench Results

Reproducible evaluation results from running VulcanBench — an open-source LLM benchmarking harness for realistic software engineering tasks.

Current published results

⚠️ CORRECTION (2026-08-25): the initial DeepSeek pass@1 below was invalidated by a VulcanBench harness bug (coverage addopts leaked into declarative grading). See CORRECTION.md. Corrected pass@1 for deepseek-chat is 19/19 (100%), not 16/19.

Snapshot Model Tasks pass@1 Avg total Date Notes
v2-deepseek-v4-flash-2026-08-27 deepseek-v4-flash 10 (v2 hard tier) 10/10 (100%) 0.762 2026-08-27 v2 realistic hard tier; all tasks verified functional=1.0 via fixed verifier
v1-deepseek-2026-08-corrected.json deepseek-chat + v4-pro 19 + 2 19/19 (100%) 0.865 2026-08-25 corrected (harness coverage bug)
v1-deepseek-chat-2026-08 deepseek-chat 19 84.2% 0.865 2026-08-25 superseded by corrected

Methodology

  • Harness: VulcanBench v0.5.1 (open-source agentic SWE benchmark)
  • Model: deepseek-chat via deepseek-v4-flash (proxied through VulcanBench openai: provider to api.deepseek.com)
  • Sandbox: local (--sandbox local) — no Docker sandbox image in the eval environment
  • Judges: disabled (--no-judges) to cut cost ~3x
  • Metric definitions (per VulcanBench):
    • functional: hidden failure-to-pass test suite (0.0 or 1.0)
    • quality: lint + cyclomatic complexity + maintainability index
    • security: bandit static analysis (no findings = 1.0)
    • total: weighted composite
  • Cost: exact API cost unavailable (deepseek returns None); runtime recorded per run

Reproduce

git clone https://github.com/morganlinton/VulcanBench.git
cd VulcanBench
make setup && source .venv/bin/activate

# Single task
vulcanbench run --task carbyne-add-months \
  --model openai:deepseek-chat \
  --sandbox local --no-judges --max-steps 30

# Full suite
vulcanbench run --suite v1 --model openai:deepseek-chat \
  --sandbox local --no-judges --max-concurrency 4

Requires a DEEPSEEK_API_KEY (or OPENAI_API_KEY + OPENAI_BASE_URL pointing at a DeepSeek-compatible endpoint).

Why this repo

VulcanBench runs are recorded under ./runs/<run_id>/ (gitignored). This repo publishes the exported result snapshots (MD + JSON) for a permanent, shareable link without bloating run traces.

Notes / caveats

  • --sandbox local runs agent shell commands on the host; use --sandbox docker for published cross-provider numbers. Results here are functional-focused dev runs.
  • Security score reports null where the bandit toolchain is unavailable on the host — never a fabricated score.
  • The v1 suite has 86 tasks; this snapshot covers 19 (a representative spread across carbyne-tier Python/Go tasks plus several open-source-derived scenarios).
  • Failures observed: py-semver-compare, py-topo-sort-cycle, py-ttl-cache-expiry (all hidden-test Python bug-fix scenarios).

License

Results data and generated reports are Apache-2.0 (aligned with VulcanBench's own license). VulcanBench itself is Apache-2.0.

About

VulcanBench evaluation results - open LLM benchmark harness. DeepSeek v1 suite. Reproducible agent evaluation.

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors