Register and share your invite link to earn from video plays and referrals.

Ant Ling
@AntLingAGI
MoE model series with foundation (Ling), reasoning (Ring) and any-to-any (Ming) from Ant Group’s AGI initiative, @TheInclusionAI.
5 Following    12.8K Followers
Ant Group has just released Ling 3.0 Flash, a 124B open weights model that scores 38 on the Artificial Analysis Intelligence Index. Ling 3.0 demonstrates a marked improvement over the previous generation and sits on the Pareto frontier for Intelligence versus Total Parameters among open weights models @AntGroup has released Ling 3.0 Flash, an open weights reasoning model with 124B total parameters and 5B active at inference time and a 262K token context window. It scores 38 on the Artificial Analysis Intelligence Index v4.1.1, 24 points above the previous generation Ling 2.6 Flash (Non-reasoning, 14). This matches MiMo-V2.5 (38) and Qwen3.6 27B (38) while using a third of MiMo-V2.5's active parameters. It remains behind the flash-tier open weights leader, DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 52. Key results: ➤ Ling 3.0 Flash sits on the open weights Pareto frontier for Intelligence vs. Total Parameters. No open weights model with fewer than 124B total parameters scores higher on the Artificial Analysis Intelligence Index, and at a comparable total size gpt-oss-120b (117B) scores 24, 14 points behind. The next model up the frontier is MiniMax-M2.7, which scores 39 with 230B total parameters. ➤ Ling 3.0 Flash demonstrates meaningful improvements in agentic abilities. Ling 3.0 Flash scores 27% on τ3-Bench Banking, second among flash-tier open weights models behind DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 39%, and ahead of Hy3 (23%) and Inkling Small (19%), all of which score 3 to 4 points higher on the Index. On GDPval-AA v2, it reaches an Elo rating of 1108, which is a meaningful improvement over Ling 2.6 Flash (545) ➤ Ling 3.0 Flash makes marked improvement in Aa-Omniscience, but almost entirely through abstention rather than knowledge. Ling 3.0 Flash scores -18 on the Artificial Analysis Omniscience Index, up from -66 for Ling 2.6 Flash (Non-reasoning). The underlying accuracy moved only from 16% to 18% while the hallucination rate on wrong answers fell from 97% to 44% and the attempt rate fell from 99% to 56%. Ling 3.0 Flash answers far fewer questions but is far less likely to hallucinate when it does answer ➤ Ling 3.0 Flash is on the Pareto frontier for Intelligence versus Price among similar sized open weights models. At $0.075 per 1M input and $0.22 per 1M output tokens on inclusionAI's first-party API, Ling 3.0 Flash is the cheapest model per token that we have measured at 38 or above on the Intelligence Index. However, that advantage is dampened in Cost per Task, because Ling 3.0 Flash used ~240M output tokens to run the Intelligence Index, at a total cost of $73. At $0.02 per task it sits inside the Pareto frontier for intelligence versus Cost per Task rather than on it. Additional model details: ➤ Size: 124B total parameters, 5B active ➤ Context window: 262K ➤ Pricing: $0.075 per 1M input tokens and $0.22 per 1M output tokens with an 80% cache hit discount ➤ License: MIT ➤ Providers: inclusionAI first-party API and DeepInfra third-party API
Show more
Open Intelligence for the real world, delivered in a cost-efficient and highly reliable inference service. Though the free period ends, you could still enjoy the Ling-3.0-flash API with a pretty good price with our day0 partner @novita_labs at 🥳
Show more
Update: We’re extending free access to Ling-3.0-Flash through August 6 at 8:00 AM PT. More time to try it out—enjoy! 🚀
🚀 Ling-3.0-flash from @AntLingAGI is now live on Novita. Launching as a Day-0 partner ⚡ 🎁 Free through August 3. 🔹 124B-parameter MoE 🔹 ~5.1B parameters activated per token 🔹 Token-efficient agentic inference 🔹 Built for production-scale workloads
Show more
Update: We’re extending free access to Ling-3.0-Flash through August 6 at 8:00 AM PT. More time to try it out—enjoy! 🚀
Thanks for our day0 partner @DeepInfra for the impeccable inference support! Try Ling-3.0-flash on Deep Infra's highly tailored and professionally served inference! 🚀
Ling-3.0-flash is live on DeepInfra A 124B MoE from @AntLingAGI with just 5.1B active params, large-model capacity at close to small-model cost. Reasoning + tool calling on by default, 128K native context (up to 1M with YaRN). Built for agents. Live now 👇
Show more
Ling-3.0-flash is live on DeepInfra A 124B MoE from @AntLingAGI with just 5.1B active params, large-model capacity at close to small-model cost. Reasoning + tool calling on by default, 128K native context (up to 1M with YaRN). Built for agents. Live now 👇
Show more
This would be fun for developers! 🥳
🎉 Day-0 support for Ling-3.0-flash from @AntLingAGI is now live in SGLang! A 124B MoE model built for production agents with: > Hybrid-linear from step 0 of pretraining: KDA + MLA stacked 5:1, 1/64 sparse MoE > 10,000+ interactive training environments > New INT4 and MXFP4 variants, running end-to-end on a single NVIDIA DGX Spark via the Spark-adapted SGLang path ⭐️ What makes long agent runs fast: Ling-3.0-flash natively integrates SGLang HiCache + Mooncake hierarchical caching, cutting TTFT by 60% to over 80% on long inputs. Try it in your agent stack today!
Show more
🎉 Day-0 support for Ling-3.0-flash from @AntLingAGI is now live in SGLang! A 124B MoE model built for production agents with: > Hybrid-linear from step 0 of pretraining: KDA + MLA stacked 5:1, 1/64 sparse MoE > 10,000+ interactive training environments > New INT4 and MXFP4 variants, running end-to-end on a single NVIDIA DGX Spark via the Spark-adapted SGLang path ⭐️ What makes long agent runs fast: Ling-3.0-flash natively integrates SGLang HiCache + Mooncake hierarchical caching, cutting TTFT by 60% to over 80% on long inputs. Try it in your agent stack today!
Show more
A small yet powerful model, try it out with our partner @kilocode ! It is also small enough to be self-hosted on a lot of different.... devices... 😈 Let's see your creation!
Try Ling-3.0-tiny from our partners! @novita_labs Oh, it is also small enough to be hosted on... many devices 😎 Let's see your design!
🚀 Ling-3.0-tiny from @AntLingAGI is now available via Novita on @openrouter. 🎁 Free through August 13. 7.9B MoE · 1.3B active parameters per token. Built for responsive AI agents, instruction following, and multi-turn conversations.
Show more
🚀 Ling-3.0-tiny from @AntLingAGI is now available via Novita on @openrouter. 🎁 Free through August 13. 7.9B MoE · 1.3B active parameters per token. Built for responsive AI agents, instruction following, and multi-turn conversations.
Show more
Ling-3.0-tiny by @AntLingAGI is live on AI/ML API ⚡️ ZeroDay support, free for everyone until August 13. Explore:
Try Ling-3.0-flash locally on different devices! Thanks @atomic_chat_hq for the day0 community support! 🥰 Any questions or issues, please come to our Discord channel to discuss~ 🤠
Show more
Run Ling 3.0 Flash locally 🌀 We released GGUF quants on Hugging Face, from lossless BF16 to 1-bit, plus NVFP4! AD-Q5_K_M is the best fit for 128GB hardware (tested on DGX Spark). It matches the original's token choice 97.5% of the time and drifts 31% less than the llama.cpp default quant of the same size.
Show more
Run Ling 3.0 Flash locally 🌀 We released GGUF quants on Hugging Face, from lossless BF16 to 1-bit, plus NVFP4! AD-Q5_K_M is the best fit for 128GB hardware (tested on DGX Spark). It matches the original's token choice 97.5% of the time and drifts 31% less than the llama.cpp default quant of the same size.
Show more
🚀 Ling-3.0-tiny from @AntLingAGI is now live on Novita. Launching as a Day-0 partner ⚡ 🎁 Free through August 13. A compact 7.9B MoE model with only 1.3B active parameters per token, built for responsive AI agents, reliable instruction following, and natural multi-turn conversations. 🔹 256K context window 🔹 Native function calling 🔹 Prompt caching 🔹 Switchable Thinking / Instant modes
Show more
Today, we’re releasing Ling-3.0-tiny: 7.9B total parameters, with only 1.3B active per token. A native hybrid reasoning model built for real-world tasks, math, instruction following, and resource-sensitive deployment. More intelligence with less compute. 🧵
Show more
0
60
1.9K
201
Forward to community
Today, we’re releasing the open weights for Ling-3.0-flash. 🎉 Official BF16 and FP8-quantized versions are now available, so you can choose the option that best fits your hardware, performance requirements, and deployment needs.
Show more
0
66
1.1K
139
Forward to community
Today, we’re releasing Ling-3.0-flash—a hybrid-reasoning MoE model built for production-scale agents. 124B parameters. Just 5.1B active per token. With 1/8 of the total and 1/12 of the active parameters, it matches or beats our 1T flagship model on most benchmarks shown.
Show more
0
118
1.9K
197
Forward to community
Thanks @AdinaYakup and the @huggingface community for the continued recognition! We feel happy to bring another 1T thinking model to the community! Comments and feedbacks welcome!