๊ฐ€์ž… ํ›„ ์ดˆ๋Œ€ ๋งํฌ๋ฅผ ๊ณต์œ ํ•˜๋ฉด ๋™์˜์ƒ ์žฌ์ƒ ๋ฐ ์ดˆ๋Œ€ ๋ณด์ƒ์„ ๋ฐ›์„ ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.

Henry Zhang
@henryzhangumich
inference @cohere | prev @amazon | cs + biochem @umich | views are my own
๊ฐ€์ž… February 2025
494 ํŒ”๋กœ์ž‰ ์ค‘    685 ํŒฌ
Just completed a full end to end run of the full 113 task DeepSWE benchmark on ox-alpha. The Rumored ~80% pass rate is completely incorrect. Actual benchmark result is 58.4%, landing the model at almost identical performance to Claude Opus 4.8 (59%) ๐Ÿงต
๋” ๋ณด๊ธฐ