Register and share your invite link to earn from video plays and referrals.

智东西China AI News
@Chinazhidx
Zhidongxi is China's leading AI industry media outlet, providing news and insights on China's LLMs, robotics, AI agents, AI applications, and AI Hardware.
89 Following    2.6K Followers
Nunchux has deployed MiniMax-H3 on AMD MI355X. On 8 GPUs, inference is up to 26.7× faster than SGLang. ⚡️A 5-second clip takes 1.3 seconds. With streaming generation, you can even modify the prompt while the video is playing to guide the next frames.
Show more
Huawei’s Xiaoyi Work is now officially live. It’s an AI work agent for research, docs, PPTs, data analysis, content creation, and coding. You give it a goal, and it can break the task down, use tools, and execute it, with human approval at key steps. It can also bring in specialized experts for more complex tasks. It works across HarmonyOS phones, tablets, and PCs. New users get 31 days free + 1,000 AI points.
Show more
Huawei has opened beta testing for Xiaoyi Work, an AI work assistant for office tasks, coding, and creative work. Give it an end goal, and it can break the task down, plan the steps, call the right tools, and execute. You can track progress in real time and review key steps before it moves on. Runs across HarmonyOS phones, tablets, and PCs, with support for cross-device task execution. Its algorithm filing lists DeepSeek, Huawei Pangu, MiniMax, Zhipu, and Moonshot, covering LLM, NLP, text-generation, and interactive content-generation algorithms. #XiaoyiWork# #Huawei#
Show more
Tencent just launched TenPayGo, a payment app for international visitors to China💸 At merchants that accept Weixin Pay, you can pay directly with TenPayGo, without exchanging cash or preparing change. -Covers everyday spending scenarios including shopping, dining, transportation, hotels, attractions,entertainment, and health. -View transaction history to check your budget and bills more easily. It supports 7 major card networks, including UnionPay and Visa, plus nearly 60 overseas wallets. TenPayGo is also adding a Shenzhen Tong transit wallet. Scan to ride Shenzhen Metro and buses. Closed beta expected in late September.
Show more
Tencent has launched Hy Translate, an AI translation app powered by its Hy-MT2 model🚀 It supports 33 languages, plus 5 Chinese ethnic minority languages and dialects. It also supports: • Voice translation • Photo translation • Offline translation Download the model to your phone and translate without an internet connection. The app is now available in 12 countries and regions, including the US, Japan, South Korea, Singapore, Thailand and Vietnam.
Show more
Agibot has produced its 20,000th humanoid robot🚀 The Expedition A3 Ultra rolled off the production line today and was delivered to Chimelong Group at its Spaceship Park in Hengqin, China. A3 Ultra features: • 700 TOPS compute • 360° perception • UWB + RTK positioning • Up to 20-DoF tactile dexterous hands It can handle delivery, carrying, organizing, and restocking tasks. The bigger story is the production ramp: 2023: 6 prototypes 2025: 5,000 units Mar 2026: 10,000 Jun 2026: 15,000 Sep 2026: 20,000 Agibot went from 10,000 to 20,000 units in ~6 months.
Show more
🚨Alibaba’s DAMO Academy has developed the Esophageal AI-Guided malignant Lesion Evaluation (EAGLE). It’s an AI model that can detect precancerous lesions and esophageal cancer from noncontrast chest CT, a task historically considered impossible. The model was validated across 80,612 patients in 12 centers across 3 countries. • 90.0% sensitivity for esophageal cancer • 52.5% sensitivity for precancerous lesions • 98.5% specificity It also maintained comparable performance on low-dose CT, potentially enabling esophageal cancer screening through existing lung-cancer screening programs.
Show more
MiMo-V2.6 just dropped, and Xiaomi’s Fuli Luo is already teasing the next architecture: MiMo-V3🔥 The core of it, HySparse2, targets three bottlenecks in agentic inference: prefill cost, KV cache size, and long-context retrieval. Compared with MiMo-V2.6’s Hybrid SWA architecture at 1M tokens: • 5.02× lower prefill FLOPs • 4.5× smaller KV Cache • Better MRCRv2 and RULER-v2 scores, plus lower AgentPPL and LongPPL The HySparse2 paper comes from Xiaomi’s LLM-Core team, with Fuli Luo as corresponding author and Team Lead. Interestingly, its references also include DeepSeek-V2, V3.2, V4, and V4.1-Flash, along with OpenAI’s GPT-4.1 evaluation work and the gpt-oss model card.
Show more
iFlytek has released Spark-ASR 2.0🚀 Gains show up where ASR usually breaks — code-switching, dialects, jargon, noise, low volume, fast speech, kids’ voices — plus cleaner, more readable transcripts. Inference cost is only ~10% above Spark-ASR-1.0. Rolling out to iFlytek Input Method tomorrow. Also available via the iFlytek Open Platform API, with planned deployments across AI glasses, smart office devices, and more.
Show more
🔥DeepSeek has published a new paper detailing its Agent training infrastructure, DSec, co-authored by Liang Wenfeng. DeepSeek Elastic Compute is a production sandbox platform that exposes FnCall, container, microVM, and full-VM sandbox backends through a unified SDK. A single production-scale DSec unit spans ~160 nodes, serving ~3M sandboxes per day. In production, it supports 380K+ concurrent sandboxes and sustains 5,000+ sandbox creations per second.
Show more
Unitree robots just took the stage at the AGT 2026 finale with a stunning water sleeve dance performance. The performance was so human-like that the judges literally started arguing over whether they were robots or real people.
Show more
Li Auto’s Foundation Model team has released MachEmbodied-Dex-1.0, a unified World Action Tactile Model that jointly learns future video, tactile states and robot actions. ME-Dex-1.0 elevates tactile sensing from an auxiliary condition to a future world state that the model must predict. It uses a unified tactile representation to align heterogeneous sensors, while a three-expert architecture connects predicted visual changes and contact dynamics to action generation. An Agentic Tactile Data Engine expands paired visual-tactile-action data through simulation. Results: • RoboTwin: 78.9% Clean-only success rate, +7.6 points vs. the strongest baseline • DexJoCo: #1# on 7 of 11 dexterous manipulation tasks • ManiFeel: plug-insertion success rate jumped from 58% → 88%
Show more
Alibaba has open-sourced Logics-Parsing-V3, a 0.8B VLM for structured long-document parsing. Instead of resetting at every page, it carries a compact structural state forward as it reads — preserving context across pages without putting the entire document into one context window. This lets it recover document hierarchy, merge content split across pages, and link visual elements to related text in a single end-to-end model. It also handles complex layouts, scientific formulas, and chemical notation. On MPDocBench-Parse, it scores 85.26, ranking #1# — 4.46 points ahead of PaddleOCR-VL-1.5
Show more
MiMo 2.6 Pro vs Flash👇 Pro (left): 10M tokens, $1.41, 4,927s Flash (right): 6.9M tokens, $0.15, 4,090s Pro costs nearly 10× more, but the night scene, lighting, and overall atmosphere are noticeably richer.
Show more
🚨UPDATE: Qwen-Image-2.1 just hit #1# among open models in both the Image Edit Arena and Text-to-Image Arena.
Ant Group has open-sourced the Ming-Image-0.1-Design family. Two 6B models: • Ming-Image-0.1-Design — text-to-image for UIs, infographics, posters, and other text-rich designs. It can also generate transparent RGBA assets. • Ming-Image-0.1-Design-Layer — decomposes a flattened design into independently editable RGBA layers. Ant also open-sourced two Agent Skills: • Ling UI Design Skill • Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1# among open-weight models on Artificial Analysis’s UI/UX Design leaderboard.
Show more
🚨Seedance 2.5 API just added Draft Mode. Generate a 480p Draft first. Select the shot you want, then turn it into a 1080p final using the same key settings. The final stays highly consistent with the Draft, while delivering native 1080p quality — at a much lower cost. For a 5-second clip, 4 generations to find the right shot can cut costs by 57% and nearly double generation speed. The more complex the shots and the more iterations needed, the more Draft Mode can save.
Show more
🔥Step Code is now open source under the MIT License. Step Code runs in your terminal and handles the full task loop—reading code, making changes, and running tests. • Terminal-Bench 2.1 — 80.9% pass rate, tied for #1# among evaluated agents, with the lowest token usage in the tie • Multi-Frame — 73.3% across 6 frontier agent harnesses, #1# with 5.09M tokens per task Step Code comes with built-in StepPage publishing. Once your local page is ready, you can publish it as an accessible static website with a single command—bringing development, debugging, and delivery all within the same terminal.
Show more
Someone just used DeepSeek-V4.1-Flash + Hermes Agent to optimize Resident Evil 7 on a OnePlus 12R (Snapdragon 8 Gen 2). The result: 4K textures → 1024/512px Texture data: ~20GB → 8GB Performance: ~20 FPS → stable 30 FPS
Show more
Huawei has opened beta testing for Xiaoyi Work, an AI work assistant for office tasks, coding, and creative work. Give it an end goal, and it can break the task down, plan the steps, call the right tools, and execute. You can track progress in real time and review key steps before it moves on. Runs across HarmonyOS phones, tablets, and PCs, with support for cross-device task execution. Its algorithm filing lists DeepSeek, Huawei Pangu, MiniMax, Zhipu, and Moonshot, covering LLM, NLP, text-generation, and interactive content-generation algorithms. #XiaoyiWork# #Huawei#
Show more
Minima AI has released a fully NVFP4-quantized build of Qwen3.8-27B.​ Unlike earlier 4-bit quantizations that kept GDN at 8/16-bit, Minima quantized all 496 backbone linear layers to NVFP4 W4A4 — including the GDN layers and their gate projections.​ The result: just a 0.52-point average drop across five tasks vs. BF16, while cutting weight memory from 50.13 GiB to 17.53 GiB.​ For 32K-token inputs, TTFT also drops from 6.90s to 4.03s.​ The interesting part: a 32K-token mechanism study found that quantization error does not keep accumulating in GDN’s recurrent state.​ The quantized weights are now open-sourced. #Qwen# #NVFP4#
Show more