注册并分享邀请链接,可获得视频播放与邀请奖励。

Trajectory
@trajectorylabs
Building the platform for Continual Learning
加入 December 2025
1 正在关注    4.9K 粉丝
It's time to rethink RL. Translating real world use into model improvements requires redesigning post-training algorithms for non-verifiable, per token rewards. At @aiDotEngineer 's World Fair, we share our insights into scaling algorithms like SDPO for continual learning.
显示更多
0
11
256
21
转发到社区