注册并分享邀请链接,可获得视频播放与邀请奖励。

Samip
@industriaalist
solving generalization at
加入 June 2016
112 正在关注    4.3K 粉丝
We've been obsessed with bending the scaling laws recently and our new paper shows that looping with model growth improves the scaling exponent, leading to gain that compounds with compute! Everyone assumes architectural changes only give constant factor gains and pretraining progress mostly comes from data (e.g. @dwarkesh_sp's recent post). We found that model growth, looping, and boundary operators result in compute multipliers over standard transformers that grow exponentially with each OOM of compute. - 1.55x at 1e20 FLOPs and 2.7x projected at 1e25. - Matches GPT-3 13B on CORE with 20x less compute w/ @charllechen, @akshayvegesna, @andrewgwils 🧵
显示更多
0
9
247
28
转发到社区