登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Maksym Andriushchenko
@maksym_andr
Principal investigator @ELLISInst_Tue & @MPI_IS, advisor @expsecai, mentor @MATSprogram. Past projects: AgentHarm, Claudini, PostTrainBench, Stolen Thoughts.
参加 April 2018
953 フォロー中    7.7K ファン
💥 Our new 116-pages long paper: we extract encrypted raw reasoning from OpenAI, Anthropic, and Gemini models at scale. This vulnerability leads to many security issues, including distillation attacks and credential extraction. We also find a lot of examples of illegible reasoning (especially for GPT models), unfaithful reasoning, and evidence that some open-weight models might've been indeed distilled from frontier proprietary models. Check out the paper in detail, including the appendix! It's one of the most exciting projects I've been involved in.
もっと見る