Register and share your invite link to earn from video plays and referrals.

Ariel
@ArielKwiat
p/hd | Big RL energy | RS @ big company (not speaking for the company though) | Prev. {Meta FAIR; Gym(nasium)} | Glory to Mankind
302 Following    6.1K Followers
On a cognitive level, all normie tasks have been basically saturated. There's a rapidly closing gap in the CUA capabilities to plug into the interfaces that drive normie tasks, and beyond that it's a matter of adoption/resistance to change.
Show more
Holy fuck they're shameless
Unpopular opinion, but the main point of math and science is discovering and proving true statements
Can we train LLMs with RL using the same next token prediction loss as pre-training? (yes) We conduct a study on (log)prob rewards and show they give a simple way to bridge verifiable and non-verifiable settings with a single reward, broadly applicable for fine-tuning LLMs.
Show more