The vLLM project released version 0.31.0 on October 5, 2026. The highlight is a new preload command that starts a weight‑cache daemon, letting post‑quantized model weights stay in GPU memory when the server restarts. The feature works with data parallelism and adds a health‑check endpoint. The release also bundles 717 commits from 307 contributors.
Why it matters
Developers can update or restart LLM services without re‑loading large models, cutting downtime for applications that rely on vLLM.