The vision drafter uses the same architecture as our text LFM2.5-DSpark drafters: it captures the target model's hidden states at a fixed set of tapped layers and conditions on them to moviesda draft a block of k candidate tokens. Image patches and text tokens are projected into a shared representation before those layers, so the drafter operates on hidden-state vectors of identical dimensionality regardless of input modality. The inference algorithm is therefore unchanged from the text models.
This is a really interesting step toward making vision-language models faster and more practical in real-world applications. The reported speedups of up to 3.13x on-device and 2.66x on an H100 are especially notable because the approach aims to improve inference speed without changing the output quality, while keeping the additional memory requirement relatively small.
I also like how the DSpark approach carries the same core idea from text models into the vision-language setting. Having the image patches and text tokens projected into a shared representation means the drafter can work with the same hidden-state dimensions across both modalities, which makes the speculative decoding process much more straightforward and reusable.
The day-one support for llama.cpp, MLX-VLM, and SGLang also makes this especially interesting for developers experimenting with local and accelerated VLM inference. It will be worth watching how these gains translate across different hardware and workloads as the experimental model develops.
