Register and share your invite link to earn from video plays and referrals.

Sanmi Koyejo
@sanmikoyejo
I lead @stai_research at Stanford. Co-founder @VirtueAI_co
108 Following    3.7K Followers
Ever noticed your LLM behaving differently depending on how you reach it? Turns out the access surface matters, and more than I'd assumed. Across 7 systems and 9 benchmarks, the same models score about 3.4 points higher through the API than through the chatbot interface. We tried to close the gap from the API side: system prompts, sampling parameters, reasoning settings. Nothing reliably reproduced interface behavior. Whatever sits between the endpoint and the deployed product isn't something an auditor can reconstruct. This matters for eval reports, which should include the access surface alongside the model name and date. This also matters for audits: access granted at the endpoint may certify a different system from the one most people use. Great work by @jennjwang with @joabaum, Dan Ho, and me (EMNLP 2026):
Show more