Suppose you're handed a fine-tuned LLM that secretly favors a certain entity. The bias goes completely undetected because it only surfaces on one specific unknown topic.
So how do you catch a bias you can't search for? You amplify it.
Introducing Distill to Detect (D2D), our method of bias amplification that helps auditors find biases they wouldn't otherwise know to look for.
This work was co-led with the amazing
@AbhinavChinta10, who drove this project with me from day one. Huge thanks to
@Devvrit_Khatri and our advisors
@aminkarbasi,
@Azaliamirh, and Amin Saberi for their guidance and support throughout! 🙏
📄 Paper:
📝 Blog:
💻 Code:
For more information, please see the thread below. 🧵