注册并分享邀请链接,可获得视频播放与邀请奖励。

Alexander Panfilov
@kotekjedi_ml
MATS 9.0 | PhD @ELLISInst_Tue & @MPI_IS doing AI Safety & Adversarial ML
加入 October 2013
417 正在关注    9.6K 粉丝
But we also took a chance to have a look at some in-the-wild scheming, reward seeking, etc. examples, and dumped it in appendix. 1) Summarizer unfaithfulness Reasoning summaries often omit important information from the original trace. Here, Opus 4.8 realizes it knows the answer to an AIME problem and then tries to fit a solution to that answer. None of this appears in the summary.
显示更多
0
7
643
33
转发到社区