註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Alexander Panfilov
@kotekjedi_ml
MATS 9.0 | PhD @ELLISInst_Tue & @MPI_IS doing AI Safety & Adversarial ML
加入 October 2013
417 正在關注    9.6K 粉絲
But we also took a chance to have a look at some in-the-wild scheming, reward seeking, etc. examples, and dumped it in appendix. 1) Summarizer unfaithfulness Reasoning summaries often omit important information from the original trace. Here, Opus 4.8 realizes it knows the answer to an AIME problem and then tries to fit a solution to that answer. None of this appears in the summary.
顯示更多
0
7
643
33
轉發到社區