注册并分享邀请链接,可获得视频播放与邀请奖励。

Lambda
@LambdaAPI
The Superintelligence Cloud
加入 July 2012
249 正在关注    20.9K 粉丝
Ask a 3D vision-language model what's near the table and in front of the curtain, and it might guess "sewing machine." The right answer is a tray rack. CVP (UC San Diego + Lambda, WACV 2026) fixes this with a target-affinity token for task-relevant objects and an allocentric grid for global context. Against Video-3D-LLM: • SQA3D EM: 58.6 → 62.3 • Scan2Cap CIDEr: 83.8 → 90.5 • Better on all 5 benchmarks tested Full results across ScanQA, SQA3D, ScanRefer, Multi3DRefer, and Scan2Cap, plus how the central/peripheral split works:
显示更多