Ornith-1.5 just dropped - a family of open-source models spanning 9B Dense, 35B MoE, and 397B MoE, trained with end-to-end self-improvement.
Here's what Ornith-1.5 actually looks like:
> 86.1 on Terminal-Bench 2.1
> 86 on SWE-Bench verified, 79.6 multilingual
> 81.4 on ClawEval, 71.2 on Tool Decathlon
- Performance comparable to Claude Opus 4.8 across reasoning, agentic, and coding tasks
- The model proposes its own tasks, generates scaffolds, and produces solution rollouts for RL.
- FP8, GGUF, MLX, and NVFP4 quantized versions all released under MIT
Open source. Commercial use unrestricted. Self-improving out of the box.
显示更多