AINewsnow

Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding

Liquid AI has released LFM2.5-VL-3B-DSpark, a 279.5M-parameter draft model that brings speculative decoding to its LFM2.5-VL-3B vision-language model. It delivers up to 3.13x faster decoding on Apple M5 Max and 2.66x on H100, with identical output under greedy decoding. Support ships in llama.cpp,…

Read the full story at MarkTechPost ↗

Timeline · 2 reports

  1. 2026-09-25 23:16 · r/machinelearningnews
    Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding
  2. 2026-09-25 23:11 · MarkTechPost
    Liquid AI Releases LFM2.5-VL-3B-DSpark: Speculative Decoding for Vision-Language Models With Up to 3.13x Faster Decoding

More stories

  1. Ternary Bonsai 2 27B at up to 532 tok/s on one RTX 4090, native Windows: MTP + n-gram speculative decoding in a from-scratch CUDA engine — r/LocalLLM
  2. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLaMA
  3. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  4. JiRackUltra_1b Runs AI Routing on Any Laptop Without a GPU — AlphaSignal
  5. Advice on models for RAG use case — r/LocalLLaMA
  6. Targeting Applied AI / ML Engineer roles. Need ruthless feedback on my architecture and metrics. — r/deeplearning
  7. 42x Faster Prompt Lookup Drafting in llama.cpp — r/LocalLLaMA
  8. Best native alternative to WebUI for remote access to local LLMs? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →