Today, we release the first series of our findings.
1 - We cut pretraining costs by 62% at frontier scale (8x Chinchilla), and we further cut all AI inference costs by 30% (Faster decode).
2 - Introducing our model ‘Feather’, which decisively overtakes Qwen 3 as the leading frontier 1.7B math model, using 180x fewer total training tokens, even matching Qwen 4B at potent benchmarks in math.
3 - With Anvil II, our state of the art LLM Optimiser, we record the largest pre-training efficiency jump on the NanoGPT Speedrun, leaping the previous by 34seconds - Our record singularly, is a greater percentage drop than the past 45 World Records combined.
The website :