注册并分享邀请链接,可获得视频播放与邀请奖励。

Braden Hancock
@bradenjhancock
AI researcher, builder, backer. Research partner @LaudeInstitute + @LaudeVentures. Co-founder @SnorkelAI. Evals @Meta. Helping AI research become companies.
加入 December 2012
318 正在关注    2.4K 粉丝
The bitter lesson rejects domain-specific structure. Problem decomposition & recombination (what RLMs make first-class) is about as general-purpose of a problem-solving tool as it comes. Reminds me of the march of processor speeds: CPUs got faster every year until ~2004, when clock speeds hit a heat wall around 3 GHz. Transistors kept getting cheaper, but the only way to keep getting gains was to decompose workloads into pieces and go parallel. As the tasks we give LLMs get longer (context length/time horizon) and more complex, scaling up to 100T parameters will probably help...but my money is on big-O-improving innovations like this improving generalization faster.
显示更多
0
3
191
21
转发到社区