Register and share your invite link to earn from video plays and referrals.

ModelScope
@ModelScope2022
Driving innovations with open communities. ๐Ÿ’ฌ Join our Discord:
183 Following    16K Followers
NVIDIA Nemotron 3 Diarization is now available on ModelScopeโ€”adding live speaker attribution to existing ASR workflows without replacing the transcription model. ๐ŸŽ™๏ธ ๐Ÿค– ๐Ÿ‘ฅ Processes streaming audio and returns speaker labels and timestamps for up to eight speaker slots in a single conversation. โšก Its end-to-end streaming architecture avoids separately combining voice activity detection, speaker embeddings, clustering, and post-processing. ๐Ÿง  The 99.2M-parameter model uses a 31-layer Transformer encoder with RoPE and builds on NVIDIAโ€™s Streaming Sortformer architecture. ๐Ÿ”Œ Pair it with Nemotron ASR, Parakeet, Canary, Whisper, or another ASR system to create speaker-attributed transcripts. ๐Ÿข Designed for meetings, contact centers, clinical conversations, live captioning, media analysis, and multi-party voice agents. ๐Ÿ–ฅ๏ธ Supports NVIDIA Ampere, Hopper, and Blackwell GPUs, with inference through NeMo Speech C++.
Show more
Digital PDFs or warped phone photos, TeleOCR parses them with one lightweight 1.2B vision-language model. ๐Ÿ“œ Apache 2.0 License. ๐Ÿค– ๐Ÿ“ƒ ๐Ÿ† Scores 96.87 overall on OmniDocBench v1.6, the highest among the listed specialized VLMs, and ranks #1# in the ICDAR 2026 Sci-ImageMiner Challenge. ๐Ÿ“ท Handles digital, photographed, curved, and degraded documents directly, without a separate dewarping model. ๐Ÿง  Combines geometry-aware synthesis, consensus-generated labels, image-based self-verification, and progressive training from vision-language alignment to reinforcement learning. โšก Supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.
Show more
Xiaomi MiMo-V2.6 is now openโ€”a native multimodal agent family built for large-scale reinforcement learning. ๐Ÿš€๐Ÿ“œ MIT License. ๐Ÿค– ๐Ÿ† MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index and reaches 71.9 on DeepSWE v1.1, 89.9 on Terminal-Bench 2.1, and 82.0 on OSWorld-Verified. ๐Ÿง  The 1.02T MoE activates 42B parameters and supports text, images, video, audio, and a 1M-token context. โš™๏ธ One mixed RL run trains coding, general, visual, and cybersecurity agents together. Pro and Flash completed 30 steps each in under six days, producing around 750K trajectories. ๐ŸŒ The models support computer use, 3D creation, embodied control, coding, design, video, and music workflows.
Show more
Shanghai AI Lab and SJTUโ€™s LUMIA Lab release NCP-ArchPreview, an open-weight 8.9B language modell. ๐Ÿ“œ Apache 2.0. ๐Ÿค– ๐Ÿ“„ โšก Trained on 5.73T Dolma 3 tokens, it reaches OLMo-3-7Bโ€™s final Stage 1 loss with only 51.3% of the tokens, a 1.95ร— convergence gain. ๐Ÿ† Its Stage 1 macro-average rises from 46.59 to 49.04, with +5.99 on GSM8K and +4.28 on HumanEval. ๐Ÿง  NCP jointly predicts tokens and a concept sequence at one-quarter the length, then feeds those concepts back to guide generation. ๐Ÿ›  Domain adaptation updates only the 17M-parameter concept module while keeping the token backbone frozen. ๐Ÿš€ Concept-conditioned drafting improves mean accepted length by 4.17% with negligible overhead.
Show more
inclusionAI open-sources the Ming-Image-0.1-Design family, two complementary 6B models for visual-design workflows.๐Ÿ“œ MIT License. ๐Ÿค– ๐Ÿค– ๐Ÿ† Design ranks #1# among open-weight models on the Artificial Analysis UI/UX Design leaderboard. Layer leads all 12 evaluated Crello settings and runs 4.3ร— faster than the evaluated 20B open-weight baseline. ๐ŸŽจ Design generates complete UIs, dashboards, infographics, and posters at up to 2048ร—2048, with strong text rendering and native transparent RGBA output. ๐Ÿงฉ Layer decomposes flattened graphics into independently editable RGBA layers while preserving the original aspect ratio. ๐Ÿง  The pipeline combines multimodal prompt conditioning, a diffusion transformer, and a 4-channel VAE. Layer adds multi-frame generation for separate layer outputs.
Show more
Apsara 2026 is ON ๐Ÿ”ฅ Come find the ModelScope booth, snap a pic, and take home some merch: ๐ŸŽ’ Backpacks ยท Crossbody bags ยท Tote bags ๐Ÿ›‹๏ธ Neck pillows ยท โ˜• Mugs ยท ๐Ÿ’ป Laptop stands ๐Ÿงธ Plush pendants ยท ๐Ÿงข Hats ยท ๐Ÿงฒ Fridge magnets ๐ŸŽง Earphone pouches ยท ๐Ÿ“ฑ Phone chains ...and more. ๐Ÿ“ ๆญๅทžๅ›ฝ้™…ๅš่งˆไธญๅฟƒไบŒๆœŸ ยท 1F ยท ๆ™บ่ƒฝ้ฆ† ยท 1-5C ยท ้ญ”ๆญ็คพๅŒบ Hangzhou International Expo Center Phase II ยท 1F ยท Intelligence Engine ยท Booth 1-5C ยท ModelScope Come early before they're gone ๐Ÿ‘€
Show more
Qwen introduces RecreationBench, a benchmark for Hybrid Computer-Use Agents with 250 application-recreation tasks across Ubuntu, macOS, Windows, Android, and Web, spanning domains such as productivity, development, graphics, multimedia, and science. Unlike GUI-only or terminal-only benchmarks, agents must explore a running reference app, recreate it in code, and pass both programmatic tests and VLM-based visual evaluation. The playground is ready. Letโ€™s build! ๐Ÿš€Dataset:
Show more
NetEase-Youdao releases Confucius4-R2T2, a true-streaming ASR model for live captions, simultaneous translation, and voice agents. ๐Ÿค– ๐Ÿ† R2T2 achieves SOTA latency and recognition quality among the evaluated open-source models, while remaining competitive with leading closed-source systems. โšก Configurable 80msโ€“2s chunks deliver 200โ€“600ms average latency with near-offline recognition accuracy. ๐Ÿ“ Append-only decoding commits stable text without revising earlier words, avoiding transcript flicker and giving downstream agents reliable input. ๐ŸŒ Optimized for Chinese and English, with multilingual recognition, hotword prompts, and contextual prompts. ๐Ÿงฐ The GitHub repo includes inference code, a minimal example, and a vLLM backend for both offline and real-time streaming ASR. ๐Ÿ“œ Code: Apache 2.0. Weights: NetEase Model Use License Agreement.
Show more
Qwen-Image-2.1 is going open source! ๐ŸŽจ โŒ›Join the countdown Weights + code. Download it, run it, build with it. ๐Ÿ’ป
The early access emails are out! Check your inbox and enjoy! ๐ŸŒŸ
๐ŸŽจ Qwen-Image-2.1 is coming, and we're opening 50 early access spots for experienced creators and developers! ๐Ÿ”— Apply here: ๐Ÿ“ฎ We'll reach out by email if you're in. ๐Ÿ’ก Program requirement: publish at least one original showcase or a hands-on review on your social media by Sep 28 at 23:59 (UTC+8). Your honest take, whether glowing or critical, is exactly what helps us make it better.
Show more
๐ŸŽจ Qwen-Image-2.1 is coming, and we're opening 50 early access spots for experienced creators and developers! ๐Ÿ”— Apply here: ๐Ÿ“ฎ We'll reach out by email if you're in. ๐Ÿ’ก Program requirement: publish at least one original showcase or a hands-on review on your social media by Sep 28 at 23:59 (UTC+8). Your honest take, whether glowing or critical, is exactly what helps us make it better.
Show more
A 35B co-work agent with only 3B active parameters. Occamy-1.0 is built to carry complex workflows through. ๐Ÿ“œ Apache 2.0. ๐Ÿค– ๐Ÿ“„ ๐Ÿ† Scores 69.10 on AutomationBench, up 29.7 points from Qwen3.6-35B-A3B and ranking #1# in the evaluated 35B-A3B group. ๐Ÿงญ Maintains context across search, tool calls, terminal coding, files, delegated runs, and history compaction, while retaining strong instruction following. ๐Ÿง  Execution-grounded training spans long-horizon interaction, software engineering, and tool-call grounding. Marathon and Sprint experts are merged, then refined with SAO. ๐Ÿ› ๏ธ The open Dressage stack supports multi-harness RL. BF16, FP8, NVFP4, and GGUF checkpoints are also available.
Show more
China Telecom open-sources Xing4.0-29B-A4B, a 29B MoE model that activates just 4B parameters for long-horizon agents.๐Ÿ“œ Apache 2.0. ๐Ÿค– ๐Ÿ“„ ๐Ÿ“„ ๐Ÿ† Leads both comparison models on Claw-Eval, Terminal-Bench 2.1, and DeepresearchBII, while reaching 75.0 on SWE-bench Verified. ๐Ÿง  mHC + MLA + MTP supports multi-step planning, tool use, and stable execution across 256K context, expandable to 512K. โšก End-to-end training on domestic computing infrastructure, with system optimizations boosting throughput by about 96%. ๐Ÿ› ๏ธ Supports major training, inference, and Agent frameworks, plus cross-architecture deployment and lightweight adaptation for private-domain tasks.
Show more
๐Ÿš€ ZDTaichu5.0-9B is now on ModelScope! ๐Ÿค– An on-device multimodal model from TaichuAI. At 9B parameters it runs on a single GPU and brings spatial reasoning, embodied AI and agentic tool use to edge deployment. Qwen3.5-9B backbone + C-RADIOv4-H vision encoder, 128K context, any-resolution image and video input. ๐Ÿงญ Spatial reasoning: leads the compared 10B-scale open VLMs (Qwen3.5-9B, STEP3-VL-10B, gemma4-8B-E4B) and scores above Gemini 3 Pro, Grok 4 and GPT-5.2 on ViewSpatial, MMSI-Bench and MindCube-tiny ๐Ÿ› ๏ธ Agent: highest among the compared open models on TAU2-Bench, Claw-Eval and IFEval ๐Ÿ“„ First-tier results on documents, charts, OCR, visual math and video, with a ready-to-use vLLM branch and Docker image ๐Ÿง  Entropy-Gated Adaptive Recurrent Reasoning: extra latent refinement steps go only to the hard tokens
Show more
One physical intelligence loop, now in 2B and 8B. PhysBrain 1.5 is here๏ผ ๐Ÿค– ๐Ÿค– ๐Ÿ“„ ๐Ÿ† PhysBrain 1.5-8B scores 72.5 across 28 embodied-understanding benchmarks, ranking #1# among the evaluated open-source models and leading 14 tasks. The compact 2B reaches 66.6. ๐Ÿงญ Both models cover spatial perception, 3D reasoning, embodied planning, grounding, affordance, and trajectory reasoning. ๐Ÿฆพ They generate end-effector trajectories and predict future states as aligned RGB, depth, and robot masks. ๐Ÿง  Built on Qwen3-VL, PhysBrain 1.5 models language, spatial outputs, actions, and future states as tokens within one autoregressive backbone. No task-specific heads.
Show more
Only ~1.2B parameters active at a time. Edge0-8B-A1B-preview makes an 8B-class MoE practical for local inference.๐Ÿ“œ Apache 2.0. ๐Ÿค– โšก Reaches 23.9โ€“25.3 tokens/s with about 1.0 GiB peak active memory in the reported short-context benchmark. ๐Ÿ† Retains most of the FP16 base modelโ€™s quality, with an average gap of just 2.8 points across five benchmarks. MMLU-Pro rises from 65.8 to 70.1. ๐Ÿง  SSD expert offload keeps most weights outside active memory, while a prerouter predicts which experts each layer will need next. ๐Ÿ›  Recover-LoRA offsets quality loss from 4-bit quantization and expert offloading. The current Preview runs locally through MLX on Apple Silicon.
Show more
TaoMate-H3 turns MiniMax H3 into a low-latency streaming audio-video generator built for continuous creation. ๐Ÿค– โšก Runs the DiT 11.45ร— faster and delivers the first playable video 10.60ร— faster than standard MiniMax H3 in the reported 480ร—864, 10-second test. ๐ŸŽฌ Generates synchronized video, speech, and sound in five-second chunks instead of waiting for the full sequence. ๐Ÿ”„ Clean KV cache and integrated audio guidance preserve identity, voice, and motion as prompts change across chunks. ๐Ÿงฉ A three-step LoRA supports continuous portrait or landscape generation at 480p, 768p, and aligned 1080p. Inference code and 4-GPU or 8-GPU deployment are included. ๐Ÿ“œ MiniMax H3 Community License.
Show more
๐Ÿ‘‘ Atria Dawn Preview is here, built to complete real research and engineering work. ๐Ÿ“œ MIT. ๐Ÿค– โš™๏ธ Built on a 744B MoE foundation with a 256K context window. Standard and FP8 weights are available. ๐Ÿ”ฌ Discovery workflows cover evidence gathering, deep research, experiment design, execution, analysis, and recovery from failure. ๐Ÿ† Leads the reported comparison on AutomationBench, BFCL v4, CyberGym, DeepSearchQA, and BrowseComp. Scores include 53.8, 77.0, 86.5, 96.0, and 92.5 respectively. ๐Ÿงฉ Creation and delivery capabilities span software, interactive apps, ML systems, visualizations, reports, and presentations. ๐Ÿ›ก Cybersecurity support covers analysis, vulnerability validation, remediation, and retesting in authorized environments.
Show more
Marigold V2 turns an image-editing DiT into a single-step model for sharp, detailed dense prediction.๐Ÿ“œ Apache 2.0. ๐Ÿค– ๐Ÿ“„ ๐Ÿ† Best zero-shot results across all evaluated depth datasets among models trained on comparable data. AbsRel improves by 16%โ€“26% over the previous best on KITTI and ETH3D. ๐Ÿ” Fur, foliage, fine wires, and object boundaries stay crisp. The two-stage iREPA and SinkLoss recipe tackles the smoothing and flying-pixel artifacts common in diffusion-based depth models. ๐Ÿงฉ The same framework reaches SOTA results in depth completion, see-through depth, surface normals, and intrinsic image decomposition. Depth completion records the lowest RMSE across all four reported benchmarks. โšก A pretrained Qwen-Image-Edit DiT becomes a single-step dense predictor through lightweight adaptation. Training takes less than a week on one 32 GB GPU.
Show more
Benchmark-leading scientific reasoning meets long-horizon agent capabilities. Intern-S2-397B is now available in BF16 and FP8.๐Ÿ“œ Apache 2.0. ๐Ÿค– ๐Ÿ† Scores 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both. Several specialized scientific benchmarks also see top results. ๐Ÿ”ฌ Raw scientific pages become training material directly, preserving text, visuals, symbols, and their relationships without intermediate parsing. ๐Ÿงช Joint training across 20+ scientific domains covers scientific reasoning, biomolecular interaction design, material generation, and time-series forecasting. ๐Ÿ›  Long-horizon agent RL strengthens tool use and sustained execution, with thinking and non-thinking modes available.
Show more