24GB VRAM is enough to run MOSS-VL locally.
@Open_MOSS
Weโve released FP8 and NF4 versions for both MOSS-VL-Instruct and MOSS-VL-Realtime โ covering image, video, and real-time streaming understanding.
๐ผ๏ธ MOSS-VL-Instruct โ image & video understanding
Built for local inference, batch processing, and serving.
๐ฅ MOSS-VL-Realtime โ understand video as it happens
Built for cameras, livestreams, and continuous video analysis, with timestamp-aware streaming understanding that continuously responds as the video unfolds.
Each model now comes in two quantization options:
โก FP8 โ balances model capability and inference performance
โก NF4 โ pushes VRAM usage even lower and can run stably within 24GB VRAM
For real-time video understanding, we recommend MOSS-VL-Realtime-NF4 for more VRAM headroom and a more stable long-context experience. FP8 is a good fit when context length is more controlled.
We also compared all four quantized checkpoints with their original BF16 models across selected image, video, and real-time streaming benchmarks. Overall, performance remains close to BF16 while VRAM usage drops substantially โ making MOSS-VL much easier to deploy on 24GB consumer GPUs.
All 4 checkpoints are available now โ
๐ผ๏ธ MOSS-VL-Instruct-0708-FP8
๐ค Hugging Face:
๐ ModelScope:
๐ผ๏ธ MOSS-VL-Instruct-0708-NF4
๐ค Hugging Face:
๐ ModelScope:
๐ฅ MOSS-VL-Realtime-FP8
๐ค Hugging Face:
๐ ModelScope:
๐ฅ MOSS-VL-Realtime-NF4
๐ค Hugging Face:
๐ ModelScope: