NVIDIA Nemotron 3 Diarization is now available on ModelScopeโadding live speaker attribution to existing ASR workflows without replacing the transcription model. ๐๏ธ
๐ค
๐ฅ Processes streaming audio and returns speaker labels and timestamps for up to eight speaker slots in a single conversation.
โก Its end-to-end streaming architecture avoids separately combining voice activity detection, speaker embeddings, clustering, and post-processing.
๐ง The 99.2M-parameter model uses a 31-layer Transformer encoder with RoPE and builds on NVIDIAโs Streaming Sortformer architecture.
๐ Pair it with Nemotron ASR, Parakeet, Canary, Whisper, or another ASR system to create speaker-attributed transcripts.
๐ข Designed for meetings, contact centers, clinical conversations, live captioning, media analysis, and multi-party voice agents.
๐ฅ๏ธ Supports NVIDIA Ampere, Hopper, and Blackwell GPUs, with inference through NeMo Speech C++.
Show more
Digital PDFs or warped phone photos, TeleOCR parses them with one lightweight 1.2B vision-language model. ๐ Apache 2.0 License.
๐ค
๐
๐ Scores 96.87 overall on OmniDocBench v1.6, the highest among the listed specialized VLMs, and ranks #
1# in the ICDAR 2026 Sci-ImageMiner Challenge.
๐ท Handles digital, photographed, curved, and degraded documents directly, without a separate dewarping model.
๐ง Combines geometry-aware synthesis, consensus-generated labels, image-based self-verification, and progressive training from vision-language alignment to reinforcement learning.
โก Supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.
Show more
Xiaomi MiMo-V2.6 is now openโa native multimodal agent family built for large-scale reinforcement learning. ๐๐ MIT License.
๐ค
๐ MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index and reaches 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified.
๐ง The 1.02T MoE activates 42B parameters and supports text, images, video, audio, and a 1M-token context.
โ๏ธ One mixed RL run trains coding, general, visual, and cybersecurity agents together. Pro and Flash completed 30 steps each in under six days, producing around 750K trajectories.
๐ The models support computer use, 3D creation, embodied control, coding, design, video, and music workflows.
Show more
Shanghai AI Lab and SJTUโs LUMIA Lab release NCP-ArchPreview, an open-weight 8.9B language modell. ๐ Apache 2.0.
๐ค
๐
โก Trained on 5.73T Dolma 3 tokens, it reaches OLMo-3-7Bโs final Stage 1 loss with only 51.3% of the tokens, a 1.95ร convergence gain.
๐ Its Stage 1 macro-average rises from 46.59 to 49.04, with +5.99 on GSM8K and +4.28 on HumanEval.
๐ง NCP jointly predicts tokens and a concept sequence at one-quarter the length, then feeds those concepts back to guide generation.
๐ Domain adaptation updates only the 17M-parameter concept module while keeping the token backbone frozen.
๐ Concept-conditioned drafting improves mean accepted length by 4.17% with negligible overhead.
Show more
inclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows.๐ MIT License.
๐ค
๐ค
๐ Design ranks #
1# among open-weight models on the Artificial Analysis UI/UX Design leaderboard. Layer leads all 12 evaluated Crello settings and runs 4.3ร faster than the evaluated 20B open-weight baseline.
๐จ Design generates complete UIs, dashboards, infographics, and posters at up to 2048ร2048, with strong text rendering and native transparent RGBA output.
๐งฉ Layer decomposes flattened graphics into independently editable RGBA layers while preserving the original aspect ratio.
๐ง The pipeline combines multimodal prompt conditioning, a diffusion transformer, and a 4-channel VAE. Layer adds multi-frame generation for separate layer outputs.
Show more
Apsara 2026 is ON ๐ฅ Come find the ModelScope booth, snap a pic, and take home some merch:
๐ Backpacks ยท Crossbody bags ยท Tote bags ๐๏ธ Neck pillows ยท โ Mugs ยท ๐ป Laptop stands ๐งธ Plush pendants ยท ๐งข Hats ยท ๐งฒ Fridge magnets ๐ง Earphone pouches ยท ๐ฑ Phone chains ...and more.
๐ ๆญๅทๅฝ้
ๅ่งไธญๅฟไบๆ ยท 1F ยท ๆบ่ฝ้ฆ ยท 1-5C ยท ้ญๆญ็คพๅบ
Hangzhou International Expo Center Phase II ยท 1F ยท Intelligence Engine ยท Booth 1-5C ยท ModelScope
Come early before they're gone ๐
Show more
Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web, spanning domains such as productivity, development, graphics, multimedia, and science.
Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass both programmatic tests and VLM-based visual evaluation.
The playground is ready. Letโs build!
๐Dataset:
Show more
NetEase-Youdao releases Confucius4-R2T2, a true-streaming ASR model for live captions, simultaneous translation, and voice agents.
๐ค
๐ R2T2 achieves SOTA latency and recognition quality among the evaluated open-source models, while remaining competitive with leading closed-source systems.
โก Configurable 80msโ2s chunks deliver 200โ600ms average latency with near-offline recognition accuracy.
๐ Append-only decoding commits stable text without revising earlier words, avoiding transcript flicker and giving downstream agents reliable input.
๐ Optimized for Chinese and English, with multilingual recognition, hotword prompts, and contextual prompts.
๐งฐ The GitHub repo includes inference code, a minimal example, and a vLLM backend for both offline and real-time streaming ASR.
๐ Code: Apache 2.0. Weights: NetEase Model Use License Agreement.
Show more
Qwen-Image-2.1 is going open source! ๐จ
โJoin the countdown
Weights + code. Download it, run it, build with it. ๐ป
The early access emails are out! Check your inbox and enjoy! ๐
๐จ Qwen-Image-2.1 is coming, and we're opening 50 early access spots for experienced creators and developers!
๐ Apply here:
๐ฎ We'll reach out by email if you're in.
๐ก Program requirement: publish at least one original showcase or a hands-on review on your social media by Sep 28 at 23:59 (UTC+8). Your honest take, whether glowing or critical, is exactly what helps us make it better.
Show more
๐จ Qwen-Image-2.1 is coming, and we're opening 50 early access spots for experienced creators and developers!
๐ Apply here:
๐ฎ We'll reach out by email if you're in.
๐ก Program requirement: publish at least one original showcase or a hands-on review on your social media by Sep 28 at 23:59 (UTC+8). Your honest take, whether glowing or critical, is exactly what helps us make it better.
Show more
A 35B co-work agent with only 3B active parameters. Occamy-1.0 is built to carry complex workflows through. ๐ Apache 2.0.
๐ค
๐
๐ Scores 69.10 on AutomationBench, up 29.7 points from Qwen3.6-35B-A3B and ranking #
1# in the evaluated 35B-A3B group.
๐งญ Maintains context across search, tool calls, terminal coding, files, delegated runs, and history compaction, while retaining strong instruction following.
๐ง Execution-grounded training spans long-horizon interaction, software engineering, and tool-call grounding. Marathon and Sprint experts are merged, then refined with SAO.
๐ ๏ธ The open Dressage stack supports multi-harness RL. BF16, FP8, NVFP4, and GGUF checkpoints are also available.
Show more
China Telecom open-sources Xing4.0-29B-A4B, a 29B MoE model that activates just 4B parameters for long-horizon agents.๐ Apache 2.0.
๐ค
๐
๐
๐ Leads both comparison models on Claw-Eval, Terminal-Bench 2.1, and DeepresearchBII, while reaching 75.0 on SWE-bench Verified.
๐ง mHC + MLA + MTP supports multi-step planning, tool use, and stable execution across 256K context, expandable to 512K.
โก End-to-end training on domestic computing infrastructure, with system optimizations boosting throughput by about 96%.
๐ ๏ธ Supports major training, inference, and Agent frameworks, plus cross-architecture deployment and lightweight adaptation for private-domain tasks.
Show more
๐ ZDTaichu5.0-9B is now on ModelScope! ๐ค
An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input.
๐งญ Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny
๐ ๏ธ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval
๐ First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image
๐ง Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens
Show more
One physical intelligence loop, now in 2B and 8B. PhysBrain 1.5 is here๏ผ
๐ค
๐ค
๐
๐ PhysBrain 1.5-8B scores 72.5 across 28 embodied-understanding benchmarks, ranking #
1# among the evaluated open-source models and leading 14 tasks. The compact 2B reaches 66.6.
๐งญ Both models cover spatial perception, 3D reasoning, embodied planning, grounding, affordance, and trajectory reasoning.
๐ฆพ They generate end-effector trajectories and predict future states as aligned RGB, depth, and robot masks.
๐ง Built on Qwen3-VL, PhysBrain 1.5 models language, spatial outputs, actions, and future states as tokens within one autoregressive backbone. No task-specific heads.
Show more
Only ~1.2B parameters active at a time. Edge0-8B-A1B-preview makes an 8B-class MoE practical for local inference.๐ Apache 2.0.
๐ค
โก Reaches 23.9โ25.3 tokens/s with about 1.0 GiB peak active memory in the reported short-context benchmark.
๐ Retains most of the FP16 base modelโs quality, with an average gap of just 2.8 points across five benchmarks. MMLU-Pro rises from 65.8 to 70.1.
๐ง SSD expert offload keeps most weights outside active memory, while a prerouter predicts which experts each layer will need next.
๐ Recover-LoRA offsets quality loss from 4-bit quantization and expert offloading. The current Preview runs locally through MLX on Apple Silicon.
Show more
TaoMate-H3 turns MiniMax H3 into a low-latency streaming audio-video generator built for continuous creation.
๐ค
โก Runs the DiT 11.45ร faster and delivers the first playable video 10.60ร faster than standard MiniMax H3 in the reported 480ร864, 10-second test.
๐ฌ Generates synchronized video, speech, and sound in five-second chunks instead of waiting for the full sequence.
๐ Clean KV cache and integrated audio guidance preserve identity, voice, and motion as prompts change across chunks.
๐งฉ A three-step LoRA supports continuous portrait or landscape generation at 480p, 768p, and aligned 1080p. Inference code and 4-GPU or 8-GPU deployment are included.
๐ MiniMax H3 Community License.
Show more
๐ Atria Dawn Preview is here, built to complete real research and engineering work. ๐ MIT.
๐ค
โ๏ธ Built on a 744B MoE foundation with a 256K context window. Standard and FP8 weights are available.
๐ฌ Discovery workflows cover evidence gathering, deep research, experiment design, execution, analysis, and recovery from failure.
๐ Leads the reported comparison on AutomationBench, BFCL v4, CyberGym, DeepSearchQA, and BrowseComp. Scores include 53.8, 77.0, 86.5, 96.0, and 92.5 respectively.
๐งฉ Creation and delivery capabilities span software, interactive apps, ML systems, visualizations, reports, and presentations.
๐ก Cybersecurity support covers analysis, vulnerability validation, remediation, and retesting in authorized environments.
Show more
Marigold V2 turns an image-editing DiT into a single-step model for sharp, detailed dense prediction.๐ Apache 2.0.
๐ค
๐
๐ Best zero-shot results across all evaluated depth datasets among models trained on comparable data. AbsRel improves by 16%โ26% over the previous best on KITTI and ETH3D.
๐ Fur, foliage, fine wires, and object boundaries stay crisp. The two-stage iREPA and SinkLoss recipe tackles the smoothing and flying-pixel artifacts common in diffusion-based depth models.
๐งฉ The same framework reaches SOTA results in depth completion, see-through depth, surface normals, and intrinsic image decomposition. Depth completion records the lowest RMSE across all four reported benchmarks.
โก A pretrained Qwen-Image-Edit DiT becomes a single-step dense predictor through lightweight adaptation. Training takes less than a week on one 32 GB GPU.
Show more
Benchmark-leading scientific reasoning meets long-horizon agent capabilities. Intern-S2-397B is now available in BF16 and FP8.๐ Apache 2.0.
๐ค
๐ Scores 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both. Several specialized scientific benchmarks also see top results.
๐ฌ Raw scientific pages become training material directly, preserving text, visuals, symbols, and their relationships without intermediate parsing.
๐งช Joint training across 20+ scientific domains covers scientific reasoning, biomolecular interaction design, material generation, and time-series forecasting.
๐ Long-horizon agent RL strengthens tool use and sustained execution, with thinking and non-thinking modes available.
Show more