注册并分享邀请链接,可获得视频播放与邀请奖励。

Maksym Andriushchenko
@maksym_andr
Principal investigator @ELLISInst_Tue & @MPI_IS, advisor @expsecai, mentor @MATSprogram. Past projects: AgentHarm, Claudini, PostTrainBench, Stolen Thoughts.
加入 April 2018
953 正在关注    7.7K 粉丝
💥 Our new 116-pages long paper: we extract encrypted raw reasoning from OpenAI, Anthropic, and Gemini models at scale. This vulnerability leads to many security issues, including distillation attacks and credential extraction. We also find a lot of examples of illegible reasoning (especially for GPT models), unfaithful reasoning, and evidence that some open-weight models might've been indeed distilled from frontier proprietary models. Check out the paper in detail, including the appendix! It's one of the most exciting projects I've been involved in.
显示更多
0
20
537
51
转发到社区