註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Dr. Mike Israetel
@misraetel
-PhD in Sport Physiology and former Prof. of Nutrition and Sport Science -Cofounder of Renaissance Periodization -BJJ Black Belt, Competitive Bodybuilder
加入 August 2015
327 正在關注    38.1K 粉絲
Thinking broader AND deeper. Pondering. This is a big deal.
GPT-6 Astra reportedly gets deeper without getting bigger. The labs already know what that means for scaling. "It's basically confirmed that GPT-6 Astra uses loop transformers, which means that instead of adding parameter count, it goes through the layers more than once. You add compute depth, but you don't increase the size of the model." "The labs are probably the people that are best positioned to say which way models are scaling. I think that's a tell that they're not seeing parameter sizes scaling as aggressively in their roadmaps, in what they find in their research."
顯示更多