登録して招待リンクを共有すると、動画再生報酬と紹介報酬を獲得できます。

prinz
@deredleritt3r
ad astra | | prinzbench:
参加 January 2024
4.5K フォロー中    21.4K ファン
UK AISI and CAISI find that K3 is significantly worse than frontier U.S. models at cyber capabilities. Comparison to Mythos Preview: - ExploitBench (see graph below). Mythos Preview reached the highest stage of this benchmark by developing exploits that achieved Arbitrary Code Execution (ACE) for 18/41 tasks. K3 never reached it at all (0/41). - "The Last Ones" benchmark: Mythos Preview completely solved the benchmark on 3/10 tries; average score of 22/32. K3 solved it on 1/10 tries; average score of 17/32. Concerningly, K3 has no material cyber safeguards. "Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during UK AISI / CAISI's evaluations."
もっと見る
Together with the US Center for AI Standards and Innovation (@NIST), we ran evaluations of Kimi K3 focused on its cyber capabilities. Kimi K3 performs below leading US frontier models on our preliminary cyber evaluations.
もっと見る