the.ai

Tools / Evaluation

verified

Agent Evaluation

Scoring an agent is harder than scoring an answer, because the same task has many valid trajectories and the thing you care about is whether the world ended up right. Grading the transcript rewards agents that look industrious; grading the end state rewards agents that finished.

Viz primitive · budget-splitstate-checked = 6

state-checked holds 25% of the budget; rest holds the remaining 75%.

Outcomes graded by inspecting the world against outcomes graded by reading the transcript, in tasks. Drag the state-based share up to watch the suite become harder to game — and more expensive to build, which is why it usually is not.

6

Reviewed by opendroid · 2026-08-18

  • arXiv:2406.12045 — tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
  • arXiv:2310.06770 — SWE-bench: Can Language Models Resolve Real-World GitHub Issues?