登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

Azalia Mirhoseini
@Azaliamirh
Founder @RicursiveAI, Asst. Prof. of CS at Stanford. Prev: DeepMind, Anthropic, Brain. Co-Creator of MoEs, AlphaChip, Test Time Scaling.
参加 May 2013
626 フォロー中    20.3K ファン
Introducing Distill to Detect (D2D), an auditing method that surfaces hidden biases in fine-tuned LLMs, even when the auditor has no idea which topic the bias is on! D2D distills the difference between the suspected and base models into a tiny 4M-parameter Cartridge (a learned, compressed prefix). Because the Cartridge's capacity is too small, it keeps only the most coherent part of the difference between the two models, surfacing the bias! Great work led by @talaei_shayan and @AbhinavChinta10!
もっと見る
Suppose you're handed a fine-tuned LLM that secretly favors a certain entity. The bias goes completely undetected because it only surfaces on one specific unknown topic. So how do you catch a bias you can't search for? You amplify it. Introducing Distill to Detect (D2D), our method of bias amplification that helps auditors find biases they wouldn't otherwise know to look for. This work was co-led with the amazing @AbhinavChinta10, who drove this project with me from day one. Huge thanks to @Devvrit_Khatri and our advisors @aminkarbasi, @Azaliamirh, and Amin Saberi for their guidance and support throughout! 🙏 📄 Paper: 📝 Blog: 💻 Code: For more information, please see the thread below. 🧵
もっと見る