Register and share your invite link to earn from video plays and referrals.

Search results for Multimodal
Multimodal community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including Multimodal
Multimodal reasoning has a latency problem. More video frames leads to more waiting. We built Damage Scout with Gemma 4 on Cerebras, running at over 2,300 toks/s, to show what fast multimodal inference unlocks. Damage Scout samples frames from a rental car walkaround, sends them to Gemma 4, gets back structured findings and box coordinates, then renders an annotated damage report in under 6 seconds. Same task. Same frames. A complete different experience powered by Cerebras ⚡️
Show more
How do we ensure multimodal foundation models remain reliable when faced with real-world distribution shifts? Meet Naman Goyal and Jenny Ni today at the Google booth (#B206#) to explore how small changes like cropping, dimming, or compression can impact both a model's answers and its verification evidence. Discover why deployment reliability is tested trust, not just a clean benchmark score. #ICML2026#
Show more
Pluralis v0.1 is a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective. Meet Lora Aroyo today at 9:30am at the Google booth (#B206#) to explore localized safety evaluation paradigms. #ICML2026#
Show more
Sneak peak of the new multimodal Rankings page 👀 Google's Veo 3.1 climbing the video leaderboard
Supporting @MITEECS and @nlp_mit’s Multimodal Machine Learning course (Spring 2026). 🎓 Students are leveraging the multimodal capabilities of Kimi K2.5 to power their final research projects. We look forward to seeing the innovative applications that will emerge this semester. 🔗 Happy coding! ✨
Show more
FLUX 3 Video from @bfl_ai is now available for everyone on OpenRouter. One unified multimodal model family for video, audio, image, and action-prediction. Serious, fun, creative, real, cinematic, whatever you need it to be. Jointly trained in one unified architecture.
Show more
Understand entire hour-long videos and wield tools and search — an efficient multimodal model with 30B total params but only 3B active at inference 🎬 Title: Kwai Keye-VL-2.0 Technical Report URL: 🎬 Overview An open-source multimodal foundation model from Kuaishou, built for long-video understanding and agentic intelligence. It's a Mixture-of-Experts (MoE) model with 30B total parameters but only 3B activated at inference. ❓ Challenges Solved Processing hour-level videos demands enormous compute. ・Many frames make long-range temporal dependencies hard to capture ・The challenge was addressing that compute constraint while keeping strong performance across diverse tasks 💡 Methodology & Proposed Approach ・Long-context: adapts DeepSeek Sparse Attention (DSA) to GQA-based architectures for lossless 256K context processing, capturing key frames and long-range temporal dependencies ・Infrastructure: scalable video I/O, heterogeneous ViT-LM parallelism, custom DSA kernels ・Training: Cross-Modal Multi-Teacher On-Policy Distillation (MOPD) with Context-RL and Video-RL to address catastrophic forgetting during multi-task alignment 📊 Experimental Results ・State-of-the-art among models of similar scale ・Especially strong on fine-grained temporal localization (TimeLens) ・Excels at long-video comprehension on Video-MME-v2 and LongVideoBench ・Also capable at multimodal agent collaboration across Code, Tool, and Search, with self-correction 🌍 Use Cases It fits long-video understanding, search, and moderation, plus backbones for video-handling autonomous agents. As the first application of sparse attention to multimodal at this scale, its big strength is making hour-level video processing cost-realistic. #VideoUnderstanding# #Multimodal#
Show more
Integrating high-fidelity video creation, precise visual consistency, and multimodal inputs, Wan3.0 helps creators turn ideas into immersive video content. Watch the video and see creativity unleashed. ✨
Show more
Qwen3.7-Flash from @Alibaba_Qwen is live on OpenRouter. A fast vision capable reasoning model for multimodal agents, visual coding, search, and computer interaction, with tool use and a 1M context window. Try it now:
Show more