注册并分享邀请链接,可获得视频播放与邀请奖励。

Sanmi Koyejo
@sanmikoyejo
I lead @stai_research at Stanford. Co-founder @VirtueAI_co
加入 September 2014
108 正在关注    3.7K 粉丝
Ever noticed your LLM behaving differently depending on how you reach it? Turns out the access surface matters, and more than I'd assumed. Across 7 systems and 9 benchmarks, the same models score about 3.4 points higher through the API than through the chatbot interface. We tried to close the gap from the API side: system prompts, sampling parameters, reasoning settings. Nothing reliably reproduced interface behavior. Whatever sits between the endpoint and the deployed product isn't something an auditor can reconstruct. This matters for eval reports, which should include the access surface alongside the model name and date. This also matters for audits: access granted at the endpoint may certify a different system from the one most people use. Great work by @jennjwang with @joabaum, Dan Ho, and me (EMNLP 2026):
显示更多