LLMs do eventually learn concepts, but only after exhausting every other option, a striking contrast with how humans abstract concepts from very little data.
This AI on Air clip is excerpted from a recent episode on the Redpoint AI show/channel
@Redpoint, featuring Lukasz Kaiser
@lukaszkaiser, co-author of the Transformer paper and former researcher at Google Brain
@GoogleResearch and
@OpenAI, reflecting from firsthand experience on the gap between Transformers and human learning.
▷ With chain of thought, RL and tools, this next-word predictor can already do things that would have seemed unbelievable two years ago, and spending hours daily discussing hard problems with a coding assistant and getting real implementations has become a reality.
▷ His analogy: like the saying about doing the right thing only after exhausting all other options, LLMs need a trillion tokens to absorb every surface-level pattern, and only when those fail to explain something do they finally learn the concept, while humans often get concepts from far less data, sometimes even making them up.
▷ Both sides have actually grown stronger: Transformers keep improving, yet the case for something beyond them has strengthened too, with a number of labs now pursuing post-Transformer architectures and seeing interesting results, and he admits he still doesn't know who wins.