注册并分享邀请链接,可获得视频播放与邀请奖励。

Matt MacInnis
@stanine
President and Chief Product Officer at Rippling
加入 May 2007
1.2K 正在关注    11.4K 粉丝
Hmm. Today, we ran our 2100 scored runs with Grok 4.6. Versus 4.5, the pass rate regressed from 87.3% to 85.9%, and nearly doubled the median latency (71s → 131s). Grok 4.6 also errored out more ("Something went wrong"), refused to answer benign questions (e.g., "Who in Sales has access to Google Workspace"), and produced incomplete outputs but claimed success (returning 54% of the required fields, but asserting 100%). Could be a matter of prompt tuning, but it feels like we're at the Pareto front (at least for now) on some of these models with everyday business tasks. Read the original post for details on the shape of data and tasks RipplingBench represents; it's very practical.
显示更多
0
7
50
10
转发到社区