I think one of the more interesting things we debated here is whether RSI is a cumulative task. Attention plus MoE plus GRPO etc seems to me like a line in the sand that you can just add to the stack once you discover it. You don't need to take five steps back to take 10 steps forward. But a lot of the work in the world isn't this clean and it certainly isn't this stationary eg legal work. This leads to some perhaps unintuitive predictions such as why RSI might land before continual learning (and why it's going to be hard to get off the current paradigm even if it's wrong)
Thanks for having me
@dwarkesh_sp!