註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Joschka Braun
@BraunJoschka
AI safety researcher @ApolloResearch | Science of Scheming | prev. @MATSprogram @kasl_ai @health_nlp @uni_tue
加入 April 2020
641 正在關注    590 粉絲
RL assumes that LLMs explore well during training. What if they choose not to? In our new ICML paper with @GoogleDeepMind, we train LLMs that strategically resist RL capability elicitation by under-exploring. We study this threat model, called exploration hacking.
顯示更多
0
8
394
45
轉發到社區