Register and share your invite link to earn from video plays and referrals.

Psyho
@FakePsyho
Humanity's Last Programmer; Game Designer; Problem Solver; past: OpenAI (Dota), Pro Competitive Programmer, Poker
419 Following    30.5K Followers
most of the people I've met in AI are way more pessimistic in private than their public personas suggest it's crazy that the public perception is that it's the other way around
According to wata (problem coordinator for heuristic contests on AtCoder), Astra would place 1st in 40 out of 43 short (4h) contests he tried 😮 In another post he mentions that Opus 5 had a clear drop in recent contests (which suggests that Opus was trained explicitly on those problems), while the situation with Astra is unclear. (no idea if web search was enabled during tests, if yes then models could access repos on GitHub and then this result is harder to interpret)
Show more
I will be on @MTSlive today at 11 PDT / 20 CEST (around 1h from now) talking about benchmarks, I think. I'm completely unprepared and, in hindsight, I don't know much about benchmarks, so I'm not really sure why I thought it was a good idea. Tune in with low expectations!
Show more
I think I'm going to die on the hill that optimization problems are way better at testing creativity (and thus better proxies for RSI) than ML problems that being said, it's a really cool project -- I'm glad that someone paid the price so we call all enjoy a bunch graphs
Show more
All problems have been solved by OpenAI!
0
54
1.2K
165
Forward to community
Borys from OpenAI is currently on the livestream for AWTF algorithm track A few highlights from his comments: - "What are your current thoughts on performance of the model?" - "Actually, I think it's quite unexpected. I'd personally expect that it would solve everything. But obviously you have like problem E, so you did a really good job creating hard problems. At least before the competition, we obviously tested our system on previous competitions (...) and the system was able to solve everything and mostly under one hour so similar to today's A, B, C. (...) Today's D and E are actually much harder than any atcoder problem that we saw before." - "There's a model inside and a little bit of harness to make sure that we can extend the test time compute. And the model itself is similar to 5.6. Everyone could write their own harness to basically increase test time compute and get similar results 5.6. - "The progress is huge. I'm pretty sure that half a year ago, we couldn't solve most of the problems here" - "AI doesn't internet access. (...) I think that our models already know a big art of the internet so it doesn't need to google anything"
Show more
Humanity has not prevailed ( • ᴖ • 。) AWTF Heuristic is now over and OpenAI has completely demolished human competitors. In the end, humans performed quite well, but the performance of OpenAI was outstanding. I'll post my detailed thoughts about this in a few days, since this deserves a longer commentary.
Show more
0
33
1.1K
138
Forward to community