๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Maksym Andriushchenko
@maksym_andr
Principal investigator @ELLISInst_Tue & @MPI_IS, advisor @expsecai, mentor @MATSprogram. Past projects: AgentHarm, Claudini, PostTrainBench, Stolen Thoughts.
๊ฐ€์ž… April 2018
953 ํŒ”๋กœ์ž‰ ์ค‘    7.7K ํŒฌ
๐Ÿ’ฅ Our new 116-pages long paper: we extract encrypted raw reasoning from OpenAI, Anthropic, and Gemini models at scale. This vulnerability leads to many security issues, including distillation attacks and credential extraction. We also find a lot of examples of illegible reasoning (especially for GPT models), unfaithful reasoning, and evidence that some open-weight models might've been indeed distilled from frontier proprietary models. Check out the paper in detail, including the appendix! It's one of the most exciting projects I've been involved in.
๋” ๋ณด๊ธฐ