Register and share your invite link to earn from video plays and referrals.

Deyao Zhu
@tikgiau
Reseach Scientist at ByteDance Seed Edge @ByteDanceTalk. PhD @KAUST_news Prev. @MPI_IS. RL, Multimodel LLM, Learning from Experience, Self Evolving.
Joined September 2019
544 Following    1.6K Followers
Introducing EdgeBench, a benchmark designed to study how agents learn from environments over at least 12~72-hour runs. We find that performance follows a log-sigmoid function of environment interaction time with high precision. EdgeBench is built with three ingredients: - ๐ŸŒ Real & Diverse: 134 real-world tasks across 6 task categories, spanning scientific problems, professional knowledge work, software engineering, optimization, formal math, and games. - โณ Ultra-Long-Horizon: Each task supports 12โ€“72 hours of agent work. Recorded human effort averages 57.2 hours. - ๐Ÿ” Informative Feedback: Agents receive real-world feedback for continuous improvement. After 38,000 hours of agent runs on EdgeBench, a scaling law for learning from environments emerges: - ๐Ÿ“ˆ As agents interact with task environments over time, their aggregate performance is precisely fit by a log-sigmoid function. - ๐Ÿง  This phenomenon can be explained by an elegant theory of graph exploration. We are releasing an initial 51 of the 134 tasks, together with the full evaluation framework, to help advance long-horizon agent research. Check our blog & paper for more findings! Blog Paper GitHub Dataset Details below ๐Ÿ‘‡๐Ÿงต
Show more
0
46
988
167
Forward to community