๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Andrew Carr ๐Ÿคธ
@andrew_n_carr
co-founder leading science @getcartwheel co-founder advisor @arcade_ai Past: Codex @OpenAI, Brain @GoogleAI, world ranked Tetris player
๊ฐ€์ž… July 2015
5.2K ํŒ”๋กœ์ž‰ ์ค‘    28.7K ํŒฌ
the weird part of this MiMo graph is that the judge and probe disagree 60% of the time. and for pro, disagreement goes up during training. would love to know how much of that is harder-to-judge behavior vs the model getting better at fooling the judge.
๋” ๋ณด๊ธฐ