Overview of ClawBench Benchmark
mainClawBench is an open-source benchmark designed to evaluate AI browser agents on everyday online tasks (e.g., booking travel, ordering food, applying for jobs, managing email) across live websites.
Key features include:
- Task Coverage: V1 contains 153 tasks across 144 websites; V2 contains 130 tasks.
- Evaluation Methodology: Measures end-to-end task success using a 5-layer recording pipeline and an agentic evaluator that compares agent runs against human references.
- Environment: Runs on any Chrome browser and utilizes Docker-isolated harnesses for execution.
- Performance Context: Current top scores are approximately 33.3%.