註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Braden Hancock
@bradenjhancock
AI researcher, builder, backer. Research partner @LaudeInstitute + @LaudeVentures. Co-founder @SnorkelAI. Evals @Meta. Helping AI research become companies.
加入 December 2012
318 正在關注    2.4K 粉絲
The bitter lesson rejects domain-specific structure. Problem decomposition & recombination (what RLMs make first-class) is about as general-purpose of a problem-solving tool as it comes. Reminds me of the march of processor speeds: CPUs got faster every year until ~2004, when clock speeds hit a heat wall around 3 GHz. Transistors kept getting cheaper, but the only way to keep getting gains was to decompose workloads into pieces and go parallel. As the tasks we give LLMs get longer (context length/time horizon) and more complex, scaling up to 100T parameters will probably help...but my money is on big-O-improving innovations like this improving generalization faster.
顯示更多
0
3
191
21
轉發到社區