註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Lisan al Gaib
@scaling01
lead them to paradise LisanBench: Impressum & Datenschutz:
加入 August 2024
1.3K 正在關注    60.1K 粉絲
it's called the forbidden axis for a reason well, except it's no longer forbidden. OpenAI opened the floodgates and will release their second looped language model besides GPT-6 Astra on DevDay
We think computational depth is the missing scaling axis, i.e. we should be doing a lot more deep learning! Every other axis has been scaled by OOMs over the past few years (params, data, sparsity, test-time reasoning), but depth has been stuck at ~100 layers since GPT-3. We've found that LLMs are both *severely* depth-bottlenecked and bad at using the depth they have, and that architectural interventions that lift this bottleneck efficiently lead to gains that increase with compute. w/ @akshayvegesna
顯示更多
0
19
756
35
轉發到社區