Register and share your invite link to earn from video plays and referrals.

Hanchen Li
@lihanc02
PhD student at @BerkeleySky on AI Infra. Pretrain coding data @Anuttacon_ .Prev @uchicago, @lmcache.
Joined March 2022
768 Following    3.3K Followers
A lot of routing work evaluates isolated prompts, but real agent systems are fundamentally multi-step and budget-constrained. Cool to see benchmarks moving toward execution-grounded, end-to-end evaluation instead of just token-level proxies. TwinRouterBench is a strong step toward realistic agentic routing evaluation — especially the separation between static supervision and dynamic SWE-bench execution. Excited to see where this goes!
Show more