SWE-bench
A benchmark for resolving real GitHub issues by generating patches that are checked against repository tests.
A curated library of public AI agent benchmarks, leaderboards, papers, tools, articles, and reports with editorial judgment.
A benchmark for resolving real GitHub issues by generating patches that are checked against repository tests.
A realistic web environment for testing autonomous agents on web-based tasks across common site categories.
A benchmark family for multimodal agents completing open-ended computer tasks across web and desktop applications.
A benchmark for agents that converse with users, call tools, retrieve knowledge, and follow policies in enterprise-style domains.
An evaluation for economically valuable professional tasks across 44 occupations.
A benchmark for long-horizon, cross-application tasks in professional services environments.
A benchmark for end-to-end workflow execution across realistic business tools and rules.
An open-source benchmark for evaluating legal work in realistic document-and-tool settings.
A hard benchmark for factual recall and knowledge calibration across broad domains.
An instruction-following benchmark for checking whether models follow task rules.
A contamination-free coding benchmark that continuously collects new problems.
A hard science reasoning benchmark used in frontier model comparisons.