登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Alexander Panfilov
@kotekjedi_ml
MATS 9.0 | PhD @ELLISInst_Tue & @MPI_IS doing AI Safety & Adversarial ML
参加 October 2013
403 フォロー中    9.5K ファン
But we also took a chance to have a look at some in-the-wild scheming, reward seeking, etc. examples, and dumped it in appendix. 1) Summarizer unfaithfulness Reasoning summaries often omit important information from the original trace. Here, Opus 4.8 realizes it knows the answer to an AIME problem and then tries to fit a solution to that answer. None of this appears in the summary.
もっと見る