註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Dwarkesh Patel
@dwarkesh_sp
加入 December 2019
1.1K 正在關注    278.5K 粉絲
Pretraining progress seems to be coming mostly from data improvements. @who_is_jerbear and I pretrained combinations of year-representative open model recipes and data corpuses across 2019 to 2025 at various small scales. Data improvements contributed 3.24x as many compute multipliers as model improvements did (12.0x vs 3.7x). And the gains stack independently - a better dataset helps every architecture about equally, and vice versa. Here are full results, plus what we think this means for the future of AI progress:
顯示更多
0
50
1.3K
98
轉發到社區