SWE-bench
A benchmark for resolving real GitHub issues by generating patches that are checked against repository tests.
Public task sets and environments that test whether agents can complete work through tools, code, browsers, or multi-turn interaction.
A benchmark for resolving real GitHub issues by generating patches that are checked against repository tests.
A realistic web environment for testing autonomous agents on web-based tasks across common site categories.
A benchmark family for multimodal agents completing open-ended computer tasks across web and desktop applications.
A benchmark for agents that converse with users, call tools, retrieve knowledge, and follow policies in enterprise-style domains.
An evaluation for economically valuable professional tasks across 44 occupations.
A benchmark for long-horizon, cross-application tasks in professional services environments.
A benchmark for end-to-end workflow execution across realistic business tools and rules.
An open-source benchmark for evaluating legal work in realistic document-and-tool settings.
A hard benchmark for factual recall and knowledge calibration across broad domains.
An instruction-following benchmark for checking whether models follow task rules.
A contamination-free coding benchmark that continuously collects new problems.
A hard science reasoning benchmark used in frontier model comparisons.