TL;DR A new suite trains and evaluates "native visual reasoning," where visual generation itself is the medium of reasoning, using large-scale data and verifiable rewards. VLM-judge scores swung by up to 92.8% on the same video.
Title: VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning
URL:
Key points
📊 300 tasks, 1.25M training instances, ~3.47M images and 1.3M videos in a large-scale dataset
🎯 A deterministic scorer using classical CV (HSV color segmentation, OCR) hits 0.60+ human agreement, beating GPT-5.5 and Gemini-3.1-Pro
🌀 Coefficients-Preserving Sampling (CPS) keeps predicted and fresh noise coefficients balanced, stabilizing RL exploration
🚀 Transfer to V-ReasonBench jumps from 10.21 to 38.22 (+28.01 pts), the largest gain across seven external benchmarks
🔬 Counterfactual tests show removing input images drops scores by 90%, confirming visual trajectories matter more than text for reasoning
✅ Verifiable-reward RL beats VLM-reward RL by +7.9% in-domain
What stands out is moving visual reasoning away from language dependence toward something trainable and verifiable in its own right.
#
VisualReasoning# #
MultimodalAI#