註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

alphaXiv
@askalphaxiv
High fidelity research
加入 November 2023
101 正在關注    56.9K 粉絲
“BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure” Reward hacking in agent benchmarks is often an infrastructure problem, not just a model-behavior problem. So this paper formalizes the full reward path and instruments runs to distinguish vulnerable tasks from actual exploit use, reaching 96% runtime detection accuracy and much higher exploit-chain recall than prior scanning baselines.
顯示更多
0
4
62
12
轉發到社區