Feed-forward 3D reconstruction should not be limited to predicting one Gaussian per pixel.
We introduce TokenGS, which uses learnable tokens to decouple the 3D Gaussian prediction from the image resolution and the number of input views.
#
CVPR2026Highlight#
[1/6]