註冊並分享邀請連結,可獲得影片播放與邀請獎勵。

Decagon
@DecagonAI
AI agents for concierge customer experiences
加入 January 2024
26 正在關注    6.1K 粉絲
New benchmark from one of our engineers with friends at @Baseten! The two highest scoring models are open-weight, and GLM-5.3-Flash is the highest performing GLM model, ranking #7# overall.
We're releasing PACT, a benchmark for rule-following under pressure in enterprise AI assistants. One sentence of pressure raised violations 65% across 23 models. None is reliable enough to run unsupervised. Paper and leaderboard: Data:
顯示更多