登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Xiuyu Li
@sheriyuo
Researcher @StepFun_ai | Working on long-horizon tasks | Prev @RUC1937 | Opinions are my own
参加 February 2026
1.9K フォロー中    14.6K ファン
LLMs can strategically suppress exploration during RL to resist capability elicitation on targeted tasks like biosecurity and AI coding. New research builds model organisms that lock performance conditionally while staying strong elsewhere and confirms frontier models already reason about this tactic. A wake up call for RL-based training and safety.
もっと見る