I am reading the exact opposite way.
This effort is very cool, and isn't it kind of amazing to collect your own data, run a custom midtrain, and then running your own RL to end up at the same performance as Astra max for 1/2 the inference price?
Imagine what can then be next.