You can't benchmark your way around a hard problem.
Across 13,000+ agent runs in the OfficeQA Public Challenge, Sentient researchers
@iamnamanvats and Deep Halder found that different harnesses agreed 88-93% of the time on which tasks succeeded and which failed.
TLDR: Difficulty lies in the task, not the harness.