Ask a 3D vision-language model what's near the table and in front of the curtain, and it might guess "sewing machine." The right answer is a tray rack.
CVP (UC San Diego + Lambda, WACV 2026) fixes this with a target-affinity token for task-relevant objects and an allocentric grid for global context.
Against Video-3D-LLM:
• SQA3D EM: 58.6 → 62.3
• Scan2Cap CIDEr: 83.8 → 90.5
• Better on all 5 benchmarks tested
Full results across ScanQA, SQA3D, ScanRefer, Multi3DRefer, and Scan2Cap, plus how the central/peripheral split works:
显示更多