AutoHarnessBench

Can frontier models automate agent harness R&D?

Each model is given a minimal harness scaffold, a task suite, and a 30-minute budget. It iteratively constructs and refines the harness by executing tasks and inspecting the resulting traces. The final harness is frozen and evaluated on held-out tasks.

Results are in. Public release soon.