I appreciate all the people who bring research on Looped Transformer that I missed, but I doubt the open research is the same as OpenAI's research.
Yesterday I was thinking that Looped at THIS scale (Astra is clearly supposed to be bigger than anything open-sourced, >3T total) means that OpenAI has run countless ablations and scaling studies and decided to go with it. They wouldn't have spent >$100M w/o doing that.
It should've scaled well, maybe not loss-wise but on downstream (though I doubt, most likely on both).
At first I thought this is BS, like, TB3 landed around 2 months ago? And they were collecting tasks for 4.5+ months!
Note that there are NO new tasks in TB4: some got removed, some were modified. Which means some of the tasks are 5+ months old and were published alongside solutions/graders openly.
Not saying Zai trained on them, but I'm not buying that none of the companies have used well-prepared envs available for free on GitHub in a very accessible form (Harbour-compatible).
btw @AcerFur can you please suggest one of the internal teams ask Astra to find vulns/bugs in Lean's kernel? Could be a huge success story + really useful going forward