23 curated resources

AI Agent Benchmark Resources

A curated library of public AI agent benchmarks, leaderboards, papers, tools, articles, and reports with editorial judgment.

Showing 1-12 of 23 resourcesPage 1 of 2
Benchmark·Coding agents
SWE-bench

SWE-bench

A benchmark for resolving real GitHub issues by generating patches that are checked against repository tests.

Open resource
Benchmark·Browser agents
WebArena-x

WebArena

A realistic web environment for testing autonomous agents on web-based tasks across common site categories.

Open resource
Benchmark·Computer-use agents
XLang Lab

OSWorld

A benchmark family for multimodal agents completing open-ended computer tasks across web and desktop applications.

Open resource
Benchmark·Tool-agent-user interaction
Sierra Research

tau-bench

A benchmark for agents that converse with users, call tools, retrieve knowledge, and follow policies in enterprise-style domains.

Open resource
Benchmark·Real-world work
OpenAI

GDPval

An evaluation for economically valuable professional tasks across 44 occupations.

Open resource
Benchmark·Long-horizon work
Mercor

APEX-Agents

A benchmark for long-horizon, cross-application tasks in professional services environments.

Open resource
Benchmark·Workflow orchestration
Zapier

AutomationBench

A benchmark for end-to-end workflow execution across realistic business tools and rules.

Open resource
Benchmark·Legal agents
Harvey

Harvey LAB

An open-source benchmark for evaluating legal work in realistic document-and-tool settings.

Open resource
Benchmark·Reasoning and knowledge
Center for AI Safety

Humanity's Last Exam

A hard benchmark for factual recall and knowledge calibration across broad domains.

Open resource
Benchmark·Instruction following
Allen Institute for AI

IFBench

An instruction-following benchmark for checking whether models follow task rules.

Open resource
Benchmark·Coding
LiveCodeBench

LiveCodeBench

A contamination-free coding benchmark that continuously collects new problems.

Open resource
Benchmark·Science reasoning
GPQA

GPQA Diamond

A hard science reasoning benchmark used in frontier model comparisons.

Open resource
Resource does not match your workflow?

Public benchmarks are references. Your delegated workflow may need a private task set.

Submit an evaluation request