On a cognitive level, all normie tasks have been basically saturated. There's a rapidly closing gap in the CUA capabilities to plug into the interfaces that drive normie tasks, and beyond that it's a matter of adoption/resistance to change.
Can we train LLMs with RL using the same next token prediction loss as pre-training?
(yes)
We conduct a study on (log)prob rewards and show they give a simple way to bridge verifiable and non-verifiable settings with a single reward, broadly applicable for fine-tuning LLMs.