RL teaches models to work longer, but reasoning is dependent on domain-specific post-training.
Baseten's Head of Model Training
@oneill_c sat down with
@dwarkesh_sp to explain horizon generalization and what's next at the frontier.
Full episode here: