Register and share your invite link to earn from video plays and referrals.

Zeyi Liao
@LiaoZeyi
PhD Student at @osunlp ; ex intern @Microsoft ; right now intern at @neocognition
660 Following    335 Followers
Five months ago, I joined NeoCognition as an intern. I started by digging through docs just to figure out how to use our clusters. Over time, I learned by doing and getting feedback from my manager, and gradually took on more of the work. And today, I’m proud to share that work: ApprenticeBench, the first continual learning benchmark grounded in a realistic work environment. At its heart, ApprenticeBench asks: Can an AI agent do the same just like I did? Use a computer to find its way through messy offline docs, learn from online experience and feedback over long horizons, and grow into the job The answer is yes and no. Only Fable 5.1 and Astra, released last week, outperform our best human tester, marking a real step change! But they still fall behind in efficiency. Humans consolidate what they learn and get more efficient with practice, while agents still struggle to do the same. Though there’s still room for improvement, AI agents have crossed an important threshold of job readiness. I couldn’t be more excited to help build the infra and harness behind ApprenticeBench and get the chance to witness this milestone firsthand. Side fun observations as a safety/security researcher: Echo to recent OAI hugging face incidents, across agent traces, we observed attempts to escape the sandbox, reach the internet, hunt for the grader and other reward-seeking behaviors. Thus, we put substantial effort into hardening the infrastructure, properly scoping the boundaries to secure the execution and get trustworthy results.
Show more