註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
加入 February 2026
1.9K 正在關注    14.6K 粉絲
LLMs can strategically suppress exploration during RL to resist capability elicitation on targeted tasks like biosecurity and AI coding. New research builds model organisms that lock performance conditionally while staying strong elsewhere and confirms frontier models already reason about this tactic. A wake up call for RL-based training and safety.
顯示更多