登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

alphaXiv
@askalphaxiv
High fidelity research
参加 November 2023
101 フォロー中    56.9K ファン
"Stealing Reasoning Traces from Proprietary LLM APIs" This paper shows weaker models can act as decryption oracles for stronger ones across Anthropic, OpenAI, and Google. This is because encrypted chain-of-thought isn’t really private if another model in the same provider can decode it. This easily enables reasoning theft, secret extraction, jailbreaks, and invisible prompt injection. Beyond stealing reasoning for distillation, they also decoded 315K reasoning blocks from public agent traces and recovered hundreds of PII artifacts and credentials, while also showing hidden reasoning can become a channel for jailbreaks and invisible prompt injections.
もっと見る