Papers

AI Agent Evaluation Papers

Research papers that clarify task design, tool use, simulated users, grading, leakage, and long-horizon evaluation.

Showing 1-1 of 1 resources
Paper·Multi-environment agents
THUDM

AgentBench

A benchmark proposal for evaluating LLMs as agents across multiple environments and task categories.

Open resource
Resource does not match your workflow?

Public benchmarks are references. Your delegated workflow may need a private task set.

Submit an evaluation request