LiquidAI added a new model called LFM2.5‑350M‑Diffusion‑Exp to Hugging Face on 2026‑10‑03. It works like the original LFM2.5‑350M model but generates text in 32‑token blocks instead of one token at a time. This block‑diffusion approach speeds up decoding, especially at low batch sizes, while keeping accuracy close to the autoregressive version when using eight denoising steps per block. The model uses the same backbone, tokenizer, and chat template as LFM2.5.
Why it matters
Developers can get faster responses from a familiar 350 M model without a large loss in quality.