註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Trajectory
@trajectorylabs
Building the platform for Continual Learning
加入 December 2025
1 正在關注    4.9K 粉絲
It's time to rethink RL. Translating real world use into model improvements requires redesigning post-training algorithms for non-verifiable, per token rewards. At @aiDotEngineer 's World Fair, we share our insights into scaling algorithms like SDPO for continual learning.
顯示更多
0
11
256
21
轉發到社區