Register and share your invite link to earn from video plays and referrals.

Search results for VideoUnderstanding
VideoUnderstanding community
One keyword maps to one global community path.
Create community
People
Not Found
Tweets including VideoUnderstanding
Video AI has been reasoning over blurred and occluded footage while trusting every frame equally. This work tackles that blind spot. Title: Confidence-Aware Tool Orchestration for Robust Video Understanding URL: ❓ What's the problem? 💡 Video-LLMs implicitly assume every frame is equally reliable (the authors' "Blind Trust Problem"). When footage degrades from motion blur, glare, or occlusion, they fail to notice and lose 15-30 points on real-world benchmarks, while their self-reported confidence barely changes, a silent failure. ❓ How does Robust-TO solve it? 💡 It bakes per-frame trustworthiness into every reasoning stage. First, quality profiling scores blur, brightness, and occlusion to keep only reliable frames; then it decomposes the query into sub-queries routed to tools robust to the dominant corruption, and every tool returns a (result, confidence) pair. ❓ How is confidence used? 💡 Evidence is grouped into high/medium/low tiers. High drives the conclusion, medium is kept only when consistent, low is a fallback only, and any residual uncertainty is stated in the answer. It is trained with GRPO using a confidence-cost reward. ❓ How well does it work? 💡 56.4% average on clean video (+10.6pt over Gemini-2.5-Pro) and 54.3% under corruption (+5.8pt over the strongest open-source Video-R1). It also trims frames from 32 to 20.7, cutting inference time by over 35% while gaining +1.6pt accuracy. #VideoUnderstanding# #MultimodalAI#
Show more
Understand entire hour-long videos and wield tools and search — an efficient multimodal model with 30B total params but only 3B active at inference 🎬 Title: Kwai Keye-VL-2.0 Technical Report URL: 🎬 Overview An open-source multimodal foundation model from Kuaishou, built for long-video understanding and agentic intelligence. It's a Mixture-of-Experts (MoE) model with 30B total parameters but only 3B activated at inference. ❓ Challenges Solved Processing hour-level videos demands enormous compute. ・Many frames make long-range temporal dependencies hard to capture ・The challenge was addressing that compute constraint while keeping strong performance across diverse tasks 💡 Methodology & Proposed Approach ・Long-context: adapts DeepSeek Sparse Attention (DSA) to GQA-based architectures for lossless 256K context processing, capturing key frames and long-range temporal dependencies ・Infrastructure: scalable video I/O, heterogeneous ViT-LM parallelism, custom DSA kernels ・Training: Cross-Modal Multi-Teacher On-Policy Distillation (MOPD) with Context-RL and Video-RL to address catastrophic forgetting during multi-task alignment 📊 Experimental Results ・State-of-the-art among models of similar scale ・Especially strong on fine-grained temporal localization (TimeLens) ・Excels at long-video comprehension on Video-MME-v2 and LongVideoBench ・Also capable at multimodal agent collaboration across Code, Tool, and Search, with self-correction 🌍 Use Cases It fits long-video understanding, search, and moderation, plus backbones for video-handling autonomous agents. As the first application of sparse attention to multimodal at this scale, its big strength is making hour-level video processing cost-realistic. #VideoUnderstanding# #Multimodal#
Show more
Camera pose matters for video understanding! Today's MLLMs excel at recognizing activities, but still struggle with the underlying space and ego/object dynamics in video. We trace this gap to a missing piece: camera pose. Introducing Cambrian-P: a multimodal LLM natively grounded in camera pose. (1/n)
Show more
NEWS: Grok can now analyze any video. Upload a video and let Grok understand what's happening, from summarizing scenes and identifying key moments to answering questions about the content. AI video understanding is becoming more powerful than ever.
Show more
Grok can now analyze any video. From identifying AI-generated content to explaining what's happening frame by frame, Grok's video understanding is taking another big step forward.
Show more
Elon Musk: “My prediction is that most AI compute will be focused on real-time video understanding and real-time video generation, and we expect to be leaders in that. Grok is currently in the number one spot, generating more videos and images than everyone else combined. We’re going to do the same with coding, and we’re going to do the same with Macrohard.”
Show more