Note · 1 min read
TensorRT-LLM 1.3.0rc29 drops AutoDeploy and adds Qwen3.5 FP8 support
The release removes the AutoDeploy feature and adds loading of Qwen3.5 checkpoints with global FP8 scaling.
Oossa · About Oossa
NVIDIA released TensorRT-LLM version 1.3.0rc29 on Sep 29, 2026. The update deletes the AutoDeploy integration, a breaking change for users who called its API, and adds support for Qwen‑3.5 model checkpoints that use global FP8 scales during loading. The change also updates the build system by dropping an obsolete CMake option.
Why it matters
Developers will need to adjust their code for the missing AutoDeploy calls, while new Qwen‑3.5 models can now run with faster FP8 precision.
Sources & references
| # | Source | Outlet | Date | Key takeaway |
|---|---|---|---|---|
| 1 | NVIDIA/TensorRT-LLM v1.3.0rc29 ↗ | TensorRT-LLM | Sep 29, 2026 | ## Highlights - Model Support - Support Qwen3.5 checkpoints with global FP8 scales across model loading paths #19519 - Optimize Gemma4 visio |
1 sources
Last updated:
Oossa · Newsletter
The week in AI, explained
Every Monday: the stories worth knowing, in plain language. Free, no spam.