Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
Abstract
Autoregressive text-to-speech models achieve strong naturalness but suffer from slow inference due to sequential token generation, limiting their deployment in production applications that require low latency. IndexTTS-2 is a state-of-the-art autoregressive TTS model consisting of a GPT, a flow-matching Diffusion Transformer, and a vocoder. Despite its high synthesis quality, its inference speed barely reaches real-time without streaming or batching support. We present Faster IndexTTS-2, which accelerates all neural network components of IndexTTS-2 for production deployment on GPUs using NVIDIA TensorRT and TensorRT-LLM. Faster IndexTTS-2 also enables streaming synthesis for latency-sensitive interactive applications, and batched inference across all components to maximize GPU utilization. Experiments on the Seed-TTS benchmark for both English and Chinese demonstrate up to 5.0× speedup on the autoregressive GPT and 3.6× end-to-end, with minimal degradation in word error rate, speaker similarity, and naturalness. Our methodology provides a practical reference for efficiently accelerating similar autoregressive speech models on GPUs.
English Samples
Comparing synthesized audio from IndexTTS-2 (PyTorch FP32) and Faster IndexTTS-2 (TRT/TRT-LLM FP16) in non-streaming and streaming modes. Streaming configurations: chunk size 100 / 50 codec frames (2s / 1s) with overlap 5 codec frames.
Translate for me, what is a surprise!
Reference
IndexTTS-2
Faster IndexTTS-2 Non-streaming
Faster IndexTTS-2 Streaming (2s)
Faster IndexTTS-2 Streaming (1s)
The palace is strict, no false rumors, Lady Qi!
Reference
IndexTTS-2
Faster IndexTTS-2 Non-streaming
Faster IndexTTS-2 Streaming (2s)
Faster IndexTTS-2 Streaming (1s)
This is the commemorative gift we have carefully crafted and prepared. As you can see, its color and material are truly radiant and dazzling.
Reference
IndexTTS-2
Faster IndexTTS-2 Non-streaming
Faster IndexTTS-2 Streaming (2s)
Faster IndexTTS-2 Streaming (1s)
You need the help of a professional like me, just as someone utterly powerless would need guidance from the most seasoned hunter when venturing into snowy mountains to hunt.
Reference
IndexTTS-2
Faster IndexTTS-2 Non-streaming
Faster IndexTTS-2 Streaming (2s)
Faster IndexTTS-2 Streaming (1s)
In true Japanese kendo, the actual combat is extremely brief, often lasting as little as half a second and rarely more than two seconds. In the fleeting instant when sharp blades clash, one side has already fallen in a pool of blood. But before this lightning-fast duel, both opponents stand fixed in place like stone statues, staring each other down for a long time. This phase can last up to ten minutes!
Reference
IndexTTS-2
Faster IndexTTS-2 Non-streaming
Faster IndexTTS-2 Streaming (2s)
Faster IndexTTS-2 Streaming (1s)
I'm sorry! My memory really isn't very good, but I'll do my best to remember everything about the time I spend with you~
Reference
IndexTTS-2
Faster IndexTTS-2 Non-streaming
Faster IndexTTS-2 Streaming (2s)
Faster IndexTTS-2 Streaming (1s)
Chinese Samples
Comparing synthesized audio from IndexTTS-2 (PyTorch FP32) and Faster IndexTTS-2 (TRT/TRT-LLM FP16) in non-streaming and streaming modes. Streaming configurations: chunk size 100 / 50 codec frames (2s / 1s) with overlap 5 codec frames.