๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Lucas Beyer (bl16)
@giffmana
Researcher (now: Meta. ex: OpenAI, DeepMind, Brain, RWTH Aachen), Gamer, Hacker, Belgian. Anon feedback: โœ—DMs โ†’ email
๊ฐ€์ž… December 2013
648 ํŒ”๋กœ์ž‰ ์ค‘    152.9K ํŒฌ
Haha not sure if good or bad, but on bio, Muse Spark 1.2 is the worst cheater: when it tries to cheat, it succeeds the least of all models shown here ๐Ÿ˜… (Except luna with 0 success. And idk what cheating means here, didn't check yet.)
๋” ๋ณด๊ธฐ
This investigation began when we noticed that, on BioMysteryBench, Gemini 3.8 Flash attempted to cheat in 21.5% of trials, roughly 14pp higher than the next model and more than 4x the roughly 5.0% rate for the rest of the field.
๋” ๋ณด๊ธฐ