# TensorRT-LLM 1.3.0rc29 drops AutoDeploy and adds Qwen3.5 FP8 support

> The release removes the AutoDeploy feature and adds loading of Qwen3.5 checkpoints with global FP8 scaling.

Oossa · 2026-09-29 · https://oossa.com/en/tensorrt-llm-1-3-0rc29-drops-autodeploy-and-adds-qwen3-5-fp8-support

NVIDIA released TensorRT-LLM version 1.3.0rc29 on Sep 29, 2026. The update deletes the AutoDeploy integration, a breaking change for users who called its API, and adds support for Qwen‑3.5 model checkpoints that use global FP8 scales during loading. The change also updates the build system by dropping an obsolete CMake option.

## The facts

- AutoDeploy removal is listed as a breaking change in PR #19028
- Qwen3.5 FP8 checkpoint support added in PR #19519

## Why it matters

Developers will need to adjust their code for the missing AutoDeploy calls, while new Qwen‑3.5 models can now run with faster FP8 precision.

## Sources & references

1. [NVIDIA/TensorRT-LLM v1.3.0rc29](https://github.com/NVIDIA/TensorRT-LLM/releases/tag/v1.3.0rc29) – TensorRT-LLM, 2026-09-29

Last updated: 2026-09-29
