Function calling
BFCL Multi-Turn
Handle evolving requests by selecting and sequencing functions with valid arguments across multi-turn conversations.
Can frontier models automate agent harness R&D?
Each model is given a minimal harness scaffold, a task suite, and a 30-minute budget. It iteratively constructs and refines the harness by executing tasks and inspecting the resulting traces. The final harness is frozen and evaluated on held-out tasks.
Held-out AHB score against the frozen harness's actor inference cost per task. Up and to the left is better; the step line marks the Pareto frontier. Follows the filters above.
Harnesses are tested across four environments spanning function calling, expert-judged reasoning, tool use, retrieval, and simulated user interaction.
BFCL Multi-Turn
Handle evolving requests by selecting and sequencing functions with valid arguments across multi-turn conversations.
HealthBench Lite
Produce safe, useful healthcare responses graded against expert-written rubrics.
Tau three Airline
Resolve airline support requests using tools, policy, and a simulated customer.
Tau three Banking
Handle banking requests using tools and retrieval over a 698-document knowledge base.
@misc{sleiman2026autoharnessbench,
author = {Essam Sleiman and Mersad Abbasi and Ayush Chakravarthy and Karina Nguyen},
title = {AutoHarnessBench: AI R\&D at the Harness Layer},
year = {2026},
month = jul,
url = {https://github.com/essamsleiman/autoharnessbench},
note = {GitHub repository}
}