Register and share your invite link to earn from video plays and referrals.

Ant Ling
@AntLingAGI
MoE model series with foundation (Ling), reasoning (Ring) and any-to-any (Ming) from Ant Group’s AGI initiative, @TheInclusionAI.
5 Following    13.4K Followers
Ling-3.0-flash-VL is now available on @OpenRouter with two weeks of free access. No local deployment is required. Call the API and start building multimodal applications and Agent workflows. OpenRouter: A special thank-you to @novita_labs for supporting this launch.
Show more
Thanks to our day0 partner @novita_labs for the continued support, this time with the "eyes" of the Ling flash model series. Enjoy Ling-3.0-flash-VL for a limited free period! Be creative and enjoy! 🫶
Show more
Today, we're open-sourcing Ling-3.0-flash-VL in BF16 and FP8. FP4 and INT4 are coming soon. Beyond visual recognition, it follows visual cues to: - Understand images, video, docs & UIs - Reason, search & verify - Use tools, check results & deliver
Show more
In French we use the word Sante, meaning "to your health", typically used in a toast. We wish you all enjoying the ling-3.0-flash-sante model, now available with our day0 partner @novita_labs ~ 😃
Sante is now available on @OpenRouter and @vercel_dev. Healthcare professionals, researchers, and developers can try Sante free for one month through the OpenRouter API and explore its capabilities on real-world healthcare tasks. Try it here: OpenRouter: Vercel: A special thank-you to @novita_labs for supporting this launch.
Show more
Today, we’re introducing Ling-3.0-flash-Sante — an MoE model enhanced for health and medicine, built on Ling-3.0-flash. Inspired by the French word “santé,” meaning “health,” Sante is built for real-world healthcare tasks spanning medical reasoning, professional healthcare tasks, deep research, and evidence-based retrieval. Sante shows competitive results across MedXpertQA-Text, DiagnosisArena-MCQ, AFUMED-Drug, HealthBench Professional, and BrowseComp, with leading performance among open-source models and results competitive with flagship models. Here’s to better health — Santé! More below ↓
Show more
We’re open-sourcing Ling-3.0-flash-Fin, a finance-enhanced model for real-world workflows, and FinFIRST, an expert-built benchmark for financial search agents. Two open releases, one goal: making financial AI more accessible and verifiable.
Show more
Financial work depends on trustworthy sources, consistent definitions, accurate calculations and auditable outputs. Introducing Ling-3.0-flash-Fin, a finance-enhanced version of Ling-3.0-flash, developed with financial institutions and domain experts. With 124B total and 5.1B active parameters, it supports information retrieval, research, valuation modeling and report preparation across long reports, research materials and complex workbooks. The model showed competitive results across FinFIRST, FinSearchComp Verified, FinCRAFT, FinanceAgent v1.1/v2, APEX-Agents, SpreadsheetBench v1/v2 and τ³-Banking. We will open-source the model weights next week.
Show more
According to Artificial Analysis, Ling-3.0-tiny sits on the mobile intelligence–speed Pareto frontier: 59 at 16K and 5.7s on iPhone 17 Pro. It also ranks first in the 64K intelligence evaluation with a score of 66. Bringing stronger intelligence to smaller devices.
Show more
Today, we’re introducing the Weight Cache Daemon for SGLang. 🚀 On Ling-2.6-1T FP8, it reduced weight loading to ~0.63s, up to ~780× faster than disk loading, and cut total engine startup from 8.8 minutes to ~0.53 minutes. Here’s how it works.
Show more
Today we are open sourcing Ling-3.0-flash-dspark, a DSpark draft model built specifically for Ling-3.0-flash. On 4 NVIDIA Blackwell GPUs at batch 1, it delivered 1,120 tok/s, 0.78 ms mean TPOT, and an accept length of 9.95 across 1,000 requests. 🧵
Show more
🧵 We’ve open-sourced 6 Base Model checkpoints for Ling-3.0-tiny & Ling-3.0-flash, covering pre-trained, mid-trained, and WSM-merged stages. None has undergone post-training, giving researchers flexible starting points for continued pre-training, fine-tuning, and further research. Two key highlights: - We use WSM to replace LR decay with weighted checkpoint merging, making the training process better suited for continual pre-training while enabling offline exploration of different LR decay strategies. - With one shared training recipe, the community can validate strategies on tiny-base, then scale them to flash-base.
Show more
Ling-3.0-tiny is now available as an open-weight model in BF16, FP8 and INT4. On Artificial Analysis, it scores 25 on the Intelligence Index and 16 on the Agentic Index, with 772 Elo on GDPval-AA v2 and 20.80 on τ³-Banking—built for real task execution. 🧵
Show more
Ant Group has just released Ling 3.0 Flash, a 124B open weights model that scores 38 on the Artificial Analysis Intelligence Index. Ling 3.0 demonstrates a marked improvement over the previous generation and sits on the Pareto frontier for Intelligence versus Total Parameters among open weights models @AntGroup has released Ling 3.0 Flash, an open weights reasoning model with 124B total parameters and 5B active at inference time and a 262K token context window. It scores 38 on the Artificial Analysis Intelligence Index v4.1.1, 24 points above the previous generation Ling 2.6 Flash (Non-reasoning, 14). This matches MiMo-V2.5 (38) and Qwen3.6 27B (38) while using a third of MiMo-V2.5's active parameters. It remains behind the flash-tier open weights leader, DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 52. Key results: ➤ Ling 3.0 Flash sits on the open weights Pareto frontier for Intelligence vs. Total Parameters. No open weights model with fewer than 124B total parameters scores higher on the Artificial Analysis Intelligence Index, and at a comparable total size gpt-oss-120b (117B) scores 24, 14 points behind. The next model up the frontier is MiniMax-M2.7, which scores 39 with 230B total parameters. ➤ Ling 3.0 Flash demonstrates meaningful improvements in agentic abilities. Ling 3.0 Flash scores 27% on τ3-Bench Banking, second among flash-tier open weights models behind DeepSeek V4 Flash 0731 (Reasoning, Max Effort) at 39%, and ahead of Hy3 (23%) and Inkling Small (19%), all of which score 3 to 4 points higher on the Index. On GDPval-AA v2, it reaches an Elo rating of 1108, which is a meaningful improvement over Ling 2.6 Flash (545) ➤ Ling 3.0 Flash makes marked improvement in Aa-Omniscience, but almost entirely through abstention rather than knowledge. Ling 3.0 Flash scores -18 on the Artificial Analysis Omniscience Index, up from -66 for Ling 2.6 Flash (Non-reasoning). The underlying accuracy moved only from 16% to 18% while the hallucination rate on wrong answers fell from 97% to 44% and the attempt rate fell from 99% to 56%. Ling 3.0 Flash answers far fewer questions but is far less likely to hallucinate when it does answer ➤ Ling 3.0 Flash is on the Pareto frontier for Intelligence versus Price among similar sized open weights models. At $0.075 per 1M input and $0.22 per 1M output tokens on inclusionAI's first-party API, Ling 3.0 Flash is the cheapest model per token that we have measured at 38 or above on the Intelligence Index. However, that advantage is dampened in Cost per Task, because Ling 3.0 Flash used ~240M output tokens to run the Intelligence Index, at a total cost of $73. At $0.02 per task it sits inside the Pareto frontier for intelligence versus Cost per Task rather than on it. Additional model details: ➤ Size: 124B total parameters, 5B active ➤ Context window: 262K ➤ Pricing: $0.075 per 1M input tokens and $0.22 per 1M output tokens with an 80% cache hit discount ➤ License: MIT ➤ Providers: inclusionAI first-party API and DeepInfra third-party API
Show more
Open Intelligence for the real world, delivered in a cost-efficient and highly reliable inference service. Though the free period ends, you could still enjoy the Ling-3.0-flash API with a pretty good price with our day0 partner @novita_labs at 🥳
Show more
Update: We’re extending free access to Ling-3.0-Flash through August 6 at 8:00 AM PT. More time to try it out—enjoy! 🚀
🚀 Ling-3.0-flash from @AntLingAGI is now live on Novita. Launching as a Day-0 partner ⚡ 🎁 Free through August 3. 🔹 124B-parameter MoE 🔹 ~5.1B parameters activated per token 🔹 Token-efficient agentic inference 🔹 Built for production-scale workloads
Show more
Update: We’re extending free access to Ling-3.0-Flash through August 6 at 8:00 AM PT. More time to try it out—enjoy! 🚀
Thanks for our day0 partner @DeepInfra for the impeccable inference support! Try Ling-3.0-flash on Deep Infra's highly tailored and professionally served inference! 🚀
Ling-3.0-flash is live on DeepInfra A 124B MoE from @AntLingAGI with just 5.1B active params, large-model capacity at close to small-model cost. Reasoning + tool calling on by default, 128K native context (up to 1M with YaRN). Built for agents. Live now 👇
Show more
Ling-3.0-flash is live on DeepInfra A 124B MoE from @AntLingAGI with just 5.1B active params, large-model capacity at close to small-model cost. Reasoning + tool calling on by default, 128K native context (up to 1M with YaRN). Built for agents. Live now 👇
Show more
This would be fun for developers! 🥳
🎉 Day-0 support for Ling-3.0-flash from @AntLingAGI is now live in SGLang! A 124B MoE model built for production agents with: > Hybrid-linear from step 0 of pretraining: KDA + MLA stacked 5:1, 1/64 sparse MoE > 10,000+ interactive training environments > New INT4 and MXFP4 variants, running end-to-end on a single NVIDIA DGX Spark via the Spark-adapted SGLang path ⭐️ What makes long agent runs fast: Ling-3.0-flash natively integrates SGLang HiCache + Mooncake hierarchical caching, cutting TTFT by 60% to over 80% on long inputs. Try it in your agent stack today!
Show more