註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Sanmi Koyejo
@sanmikoyejo
I lead @stai_research at Stanford. Co-founder @VirtueAI_co
加入 September 2014
108 正在關注    3.7K 粉絲
Ever noticed your LLM behaving differently depending on how you reach it? Turns out the access surface matters, and more than I'd assumed. Across 7 systems and 9 benchmarks, the same models score about 3.4 points higher through the API than through the chatbot interface. We tried to close the gap from the API side: system prompts, sampling parameters, reasoning settings. Nothing reliably reproduced interface behavior. Whatever sits between the endpoint and the deployed product isn't something an auditor can reconstruct. This matters for eval reports, which should include the access surface alongside the model name and date. This also matters for audits: access granted at the endpoint may certify a different system from the one most people use. Great work by @jennjwang with @joabaum, Dan Ho, and me (EMNLP 2026):
顯示更多