註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Rosinality
@rosinality
ML Engineer @poolsideai
加入 October 2008
1K 正在關注    7.5K 粉絲
Improving exploration for rubric RL using OPSD to generate guidance for underexplored or penalized criteria. As combining criteria to construct rewards is common maybe it could be useful beyond rubric RL.
顯示更多
0
1
106
15
轉發到社區