注册并分享邀请链接,可获得视频播放与邀请奖励。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在关注    56.9K 粉丝
"Stealing Reasoning Traces from Proprietary LLM APIs" This paper shows weaker models can act as decryption oracles for stronger ones across Anthropic, OpenAI, and Google. This is because encrypted chain-of-thought isn’t really private if another model in the same provider can decode it. This easily enables reasoning theft, secret extraction, jailbreaks, and invisible prompt injection. Beyond stealing reasoning for distillation, they also decoded 315K reasoning blocks from public agent traces and recovered hundreds of PII artifacts and credentials, while also showing hidden reasoning can become a channel for jailbreaks and invisible prompt injections.
显示更多
0
9
259
28
转发到社区