註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
加入 January 2024
682 正在關注    152K 粉絲
Safety refusal reporting is now available in the Artificial Analysis Coding Agent Index In our latest Coding Agent Index v1.5, we’ve introduced safety refusal reporting to help explain model behavior and score differences. A safety refusal occurs when a provider or model declines to start or continue a task on safety grounds. An agent may fall back to another model to continue, or stop the attempt with a block. Claude Fable 5.1 had the highest fallback rates in both Claude Code and Devin Fusion, with fallback attempts accounting for 8.8% and 7.1% of the Index's weight, respectively; these results therefore include the fallback models' performance. Refusal variability, harness context buildup, effort settings, and retry strategies can all affect the observed rates.
顯示更多
0
27
125
9
轉發到社區