How We Made a Text-to-Speech Model Respond in Sub-50 ms
2 months ago
- Qwen3-TTS 1.7B CustomVoice achieves 10 RPS and sub-50ms p95 TTFA on a single NVIDIA H100 SXM.
- Compared to vLLM-Omni, SGLang-Omni, VoxServe, and M*, only the proposed implementation maintains sub-50ms p95 TTFA through 10 RPS.
- Cost efficiency: ~$2 per 1M characters vs ElevenLabs ($100) and Cartesia ($49) with lower TTFA.
- Real-time TTS defined by low audible TTFA, zero underruns, capacity under load, and non-malformed output.
- Optimizations include removing leading silence, tuning frame accumulation, and a unified scheduler for three model modules.
- The scheduler prioritizes first-audio latency and batches urgent requests with compatible work to maximize GPU efficiency.
- Exploits fixed structure of Code Predictor for preallocated KV cache and CUDA graphs; uses state-cached incremental decoding for Codec.
- Additional optimizations: CUDA graphs for predefined batch sizes, avoiding unnecessary CPU-GPU sync, and input streaming for speech-to-speech.
- Future plans extend to image, video, world models, and fine-tuning, with vision of real-time multimodal inference.