๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Florian Brand
@xeophon
evals @PrimeIntellect | open models @interconnectsai
๊ฐ€์ž… July 2015
793 ํŒ”๋กœ์ž‰ ์ค‘    15.9K ํŒฌ
oh would you look at this, an open model being sota ๐Ÿ‘€ sinatras cooked so hard here ๐Ÿ‘จโ€๐Ÿณ
Time to give agents a hard task, introducing pmpp-hard! 69 GPU kernel tasks, 11 models, 3.1k agent rollouts and 5.8B+ tokens later we have the results. Kimi K3 claims the first place with a 0.71 score. Without dealing with if your task is โ€œfrontierโ€ or not, it just solves them
๋” ๋ณด๊ธฐ