Liquid AI
Try LFM • Docs • Discord

LFM2.5-VL-3B-DSpark-GGUF

GGUF build of LiquidAI/LFM2.5-VL-3B-DSpark for llama.cpp.

This is a standalone draft sidecar: it carries only the drafter (4 attention layers, a Markov head, a confidence head, and block size 9.) Token embeddings and the LM head are shared from the target model at load time, so it must be paired with a LFM2.5-VL-3B-GGUF target file.

Find more information about LFM2.5-VL-DSpark in our blog post.

📦 Files

file size notes
LFM2.5-VL-3B-DSpark-F16.gguf 567 MB Standalone F16 drafter; pair with the target model and vision projector from LiquidAI/LFM2.5-VL-3B-GGUF.

All inference numbers for this release use 16-bit processing for both the vision encoder and language backbone.

🏃 How to run (llama.cpp)

Run the target model with the DSpark drafter:

llama-server -hf LiquidAI/LFM2.5-VL-3B-GGUF:F16 \
  -hfd LiquidAI/LFM2.5-VL-3B-DSpark-GGUF:F16 \
  --spec-type draft-dspark --spec-draft-n-max 8 --spec-draft-n-min 0 \
  -ngl 99 -ngld 99 -fa on

The drafter was trained with block size 9. For Apple silicon, we recommend block size 8 through --spec-draft-n-max 8.

Speculative decoding is exact under greedy decoding: the target verifies every proposed token, so the generated output equals the target model running alone. The llama.cpp timing logs report the draft acceptance rate.

Other LFM2.5-VL-3B-DSpark formats:

📊 Acceptance and benchmarks

See LiquidAI/LFM2.5-VL-3B-DSpark for model details and acceptance-length and throughput benchmarks across six vision-language tasks on NVIDIA H100 and Apple silicon, including MLX-VLM, llama.cpp, and SGLang results.

📬 Contact

Citation

@article{liquidAI2026VL3B,
  author  = {Liquid AI},
  title   = {LFM2.5-VL-3B: A Better and Faster Vision-Language Model for the Edge},
  journal = {Liquid AI Blog},
  year    = {2026},
  note    = {www.liquid.ai/blog/lfm2-5-vl-3b},
}
@article{liquidAI2026vldspark,
  author = {Liquid AI},
  title = {LFM2.5-VL-DSpark: Accelerating vision-language models on edge and beyond},
  journal = {Liquid AI Blog},
  year = {2026},
  note = {www.liquid.ai/blog/lfm2-5-vl-dspark},
}
Downloads last month
724
GGUF
Model size
0.3B params
Architecture
dflash
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LiquidAI/LFM2.5-VL-3B-DSpark-GGUF

Quantized
(3)
this model

Article mentioning LiquidAI/LFM2.5-VL-3B-DSpark-GGUF