Today, we release an experimental DSpark draft model for LFM2.5-VL-3B, bringing speculative decoding to our vision-language models.
A lightweight drafter proposes multiple tokens ahead, and the target model verifies them together in a single pass. This accelerates generation without changing output quality.
Across six vision-language task categories at batch size 1 and temperature 0:
> MLX on M5 Max: up to 3.13x faster decoding and 2.62x end to end
> llama.cpp on M3 Ultra: up to 2.14x faster decoding and 1.77x end to end
> SGLang on H100: up to 2.66x faster decoding and 2.27x end to end
All evaluations were collected using Pipette, the benchmarking infrastructure behind Liquid AI's public device-performance data.
๐งต