From 7a40c1f592c51f390d36753360e8ffc25af3e737 Mon Sep 17 00:00:00 2001 From: reacher-z Date: Mon, 27 Jul 2026 19:38:34 +0800 Subject: [PATCH 1/2] docs: list ClawBench among related benchmarks --- README.md | 1 + 1 file changed, 1 insertion(+) diff --git a/README.md b/README.md index 7b5d3b0..2965695 100644 --- a/README.md +++ b/README.md @@ -43,6 +43,7 @@ Together they make external results easy to run, compare, cite, and submit back. - `All` (300) / `Hard` (hard subset) - [x] **BrowseComp** — Browser operation competition tasks, no login required - `All` (1266) +- [ ] **ClawBench** — Open web-agent benchmark with 153 real-world tasks across 144 live websites, using safe request interception for reproducible evaluation - [ ] More benchmarks > Details: [Benchmarks overview](https://docs.bubench.lexmount.io/en/benchmarks/overview). From 69c9e1f28985c5222df6e078ee7ed7216e48a3c9 Mon Sep 17 00:00:00 2001 From: genitrix Date: Tue, 28 Jul 2026 15:13:47 +0800 Subject: [PATCH 2/2] Refresh ClawBench benchmark index entry --- README.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/README.md b/README.md index 2965695..380e352 100644 --- a/README.md +++ b/README.md @@ -43,7 +43,7 @@ Together they make external results easy to run, compare, cite, and submit back. - `All` (300) / `Hard` (hard subset) - [x] **BrowseComp** — Browser operation competition tasks, no login required - `All` (1266) -- [ ] **ClawBench** — Open web-agent benchmark with 153 real-world tasks across 144 live websites, using safe request interception for reproducible evaluation +- [ ] **ClawBench** — Open web-agent benchmark with 283 V1+V2 tasks across 163 live websites, using submission interception and judge-based outcome evidence - [ ] More benchmarks > Details: [Benchmarks overview](https://docs.bubench.lexmount.io/en/benchmarks/overview).