In the early days of OpenAI,
@johnschulman2 didn't think next-token prediction was going to lead to intelligence, because it would get swamped by noise.
He explains it's always been hard to apriori predict what techniques will elicit out-of-distribution generalization.