Today, most agents are built with the help of other agents like Sierra's Ghostwriter. Yesterday, Sierra open-sourced hyper-๐-bench (published as ๐^๐-bench), a new long horizon agent evaluation that measures how well models can not only act as an agent, but construct one.