註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在關注    56.9K 粉絲
"Stealing Reasoning Traces from Proprietary LLM APIs" This paper shows weaker models can act as decryption oracles for stronger ones across Anthropic, OpenAI, and Google. This is because encrypted chain-of-thought isn’t really private if another model in the same provider can decode it. This easily enables reasoning theft, secret extraction, jailbreaks, and invisible prompt injection. Beyond stealing reasoning for distillation, they also decoded 315K reasoning blocks from public agent traces and recovered hundreds of PII artifacts and credentials, while also showing hidden reasoning can become a channel for jailbreaks and invisible prompt injections.
顯示更多
0
9
259
28
轉發到社區