TensorRT-LLM expands quantized model support and speculative decoding
NVIDIA publishes an inference release candidate with model-path changes and breaking API removals. Retrospective brief for 10:00–10:59 UTC on 29 September 2026.
Top AI stories from the last hour
Copy MarkdownTensorRT-LLM 1.3.0rc29 adds inference paths and removes AutoDeploy
TensorRT-LLM 1.3.0rc29 adds support for Qwen3.5 checkpoints using global FP8 scales, enables a CuTeDSL mixture-of-experts backend for Kimi K3 NVFP4 SiTU, and extends Helix speculative verification to additional FP8, FP4 MLA and DSpark groups. Its OpenAI request handling accepts parallel_tool_calls. The candidate also removes AutoDeploy integration and public entry points, a breaking change. This is a prerelease, so deployments should review compatibility and test their model paths before upgrading.