Register and share your invite link to earn from video plays and referrals.

Azalia Mirhoseini
@Azaliamirh
Founder @RicursiveAI, Asst. Prof. of CS at Stanford. Prev: DeepMind, Anthropic, Brain. Co-Creator of MoEs, AlphaChip, Test Time Scaling.
Joined May 2013
626 Following    20.3K Followers
Introducing Distill to Detect (D2D), an auditing method that surfaces hidden biases in fine-tuned LLMs, even when the auditor has no idea which topic the bias is on! D2D distills the difference between the suspected and base models into a tiny 4M-parameter Cartridge (a learned, compressed prefix). Because the Cartridge's capacity is too small, it keeps only the most coherent part of the difference between the two models, surfacing the bias! Great work led by @talaei_shayan and @AbhinavChinta10!
Show more
Suppose you're handed a fine-tuned LLM that secretly favors a certain entity. The bias goes completely undetected because it only surfaces on one specific unknown topic. So how do you catch a bias you can't search for? You amplify it. Introducing Distill to Detect (D2D), our method of bias amplification that helps auditors find biases they wouldn't otherwise know to look for. This work was co-led with the amazing @AbhinavChinta10, who drove this project with me from day one. Huge thanks to @Devvrit_Khatri and our advisors @aminkarbasi, @Azaliamirh, and Amin Saberi for their guidance and support throughout! 🙏 📄 Paper: 📝 Blog: 💻 Code: For more information, please see the thread below. 🧵
Show more