注册并分享邀请链接,可获得视频播放与邀请奖励。

Artificial Analysis
@ArtificialAnlys
Independent analysis of AI
加入 January 2024
683 正在关注    152.3K 粉丝
Safety refusal reporting is now available in the Artificial Analysis Coding Agent Index In our latest Coding Agent Index v1.5, we’ve introduced safety refusal reporting to help explain model behavior and score differences. A safety refusal occurs when a provider or model declines to start or continue a task on safety grounds. An agent may fall back to another model to continue, or stop the attempt with a block. Claude Fable 5.1 had the highest fallback rates in both Claude Code and Devin Fusion, with fallback attempts accounting for 8.8% and 7.1% of the Index's weight, respectively; these results therefore include the fallback models' performance. Refusal variability, harness context buildup, effort settings, and retry strategies can all affect the observed rates.
显示更多
0
27
125
9
转发到社区