가입 후 초대 링크를 공유하면 동영상 재생 및 초대 보상을 받을 수 있습니다.

alphaXiv
@askalphaxiv
High fidelity research
가입 November 2023
101 팔로잉 중    56.9K 팬
"Stealing Reasoning Traces from Proprietary LLM APIs" This paper shows weaker models can act as decryption oracles for stronger ones across Anthropic, OpenAI, and Google. This is because encrypted chain-of-thought isn’t really private if another model in the same provider can decode it. This easily enables reasoning theft, secret extraction, jailbreaks, and invisible prompt injection. Beyond stealing reasoning for distillation, they also decoded 315K reasoning blocks from public agent traces and recovered hundreds of PII artifacts and credentials, while also showing hidden reasoning can become a channel for jailbreaks and invisible prompt injections.
더 보기