# vLLM 0.31.0 adds fast‑restart weight cache for GPU serving

> The open‑source LLM serving library vLLM 0.31.0 introduces a preload CLI that keeps post‑quantized weights resident in GPU memory across engine restarts.

Oossa · 2026-10-05 · https://oossa.com/en/vllm-0-31-0-adds-fast-restart-weight-cache-for-gpu-serving

The vLLM project released version 0.31.0 on October 5, 2026. The highlight is a new preload command that starts a weight‑cache daemon, letting post‑quantized model weights stay in GPU memory when the server restarts. The feature works with data parallelism and adds a health‑check endpoint. The release also bundles 717 commits from 307 contributors.

## The facts

- Version 0.31.0 released on 2026‑10‑05
- 717 commits from 307 contributors

## Why it matters

Developers can update or restart LLM services without re‑loading large models, cutting downtime for applications that rely on vLLM.

## Sources & references

1. [vllm-project/vllm v0.31.0](https://github.com/vllm-project/vllm/releases/tag/v0.31.0) – vLLM, 2026-10-05

Last updated: 2026-10-05
