Nice results across diverse tasks with Ling-3.0-Flash 👀
It's a great example of how compact execution models can still be highly capable when given a focused role within a larger agent system.
The hybrid reasoning design is also interesting: spending compute only when needed while keeping the default execution path lightweight feels like a practical direction for production agents.
Hope to see more Chinese open-source models continue pushing in this direction, optimizing for both efficiency and real-world agent workflows.
Today, we’re releasing Ling-3.0-flash—a hybrid-reasoning MoE model built for production-scale agents.
124B parameters. Just 5.1B active per token.
With 1/8 of the total and 1/12 of the active parameters, it matches or beats our 1T flagship model on most benchmarks shown.