24GB VRAM is enough to run MOSS-VL locally.
@Open_MOSS
We’ve released FP8 and NF4 versions for both MOSS-VL-Instruct and MOSS-VL-Realtime — covering image, video, and real-time streaming understanding.
🖼️ MOSS-VL-Instruct — image & video understanding
Built for local inference, batch processing, and serving.
🎥 MOSS-VL-Realtime — understand video as it happens
Built for cameras, livestreams, and continuous video analysis, with timestamp-aware streaming understanding that continuously responds as the video unfolds.
Each model now comes in two quantization options:
⚡ FP8 — balances model capability and inference performance
⚡ NF4 — pushes VRAM usage even lower and can run stably within 24GB VRAM
For real-time video understanding, we recommend MOSS-VL-Realtime-NF4 for more VRAM headroom and a more stable long-context experience. FP8 is a good fit when context length is more controlled.
We also compared all four quantized checkpoints with their original BF16 models across selected image, video, and real-time streaming benchmarks. Overall, performance remains close to BF16 while VRAM usage drops substantially — making MOSS-VL much easier to deploy on 24GB consumer GPUs.
All 4 checkpoints are available now ↓
🖼️ MOSS-VL-Instruct-0708-FP8
🤗 Hugging Face:
🌐 ModelScope:
🖼️ MOSS-VL-Instruct-0708-NF4
🤗 Hugging Face:
🌐 ModelScope:
🎥 MOSS-VL-Realtime-FP8
🤗 Hugging Face:
🌐 ModelScope:
🎥 MOSS-VL-Realtime-NF4
🤗 Hugging Face:
🌐 ModelScope: