Hasty Briefsbeta

Bilingual

How We Made a Text-to-Speech Model Respond in Sub-50 ms

2 months ago
  • Qwen3-TTS 1.7B CustomVoice achieves 10 RPS and sub-50ms p95 TTFA on a single NVIDIA H100 SXM.
  • Compared to vLLM-Omni, SGLang-Omni, VoxServe, and M*, only the proposed implementation maintains sub-50ms p95 TTFA through 10 RPS.
  • Cost efficiency: ~$2 per 1M characters vs ElevenLabs ($100) and Cartesia ($49) with lower TTFA.
  • Real-time TTS defined by low audible TTFA, zero underruns, capacity under load, and non-malformed output.
  • Optimizations include removing leading silence, tuning frame accumulation, and a unified scheduler for three model modules.
  • The scheduler prioritizes first-audio latency and batches urgent requests with compatible work to maximize GPU efficiency.
  • Exploits fixed structure of Code Predictor for preallocated KV cache and CUDA graphs; uses state-cached incremental decoding for Codec.
  • Additional optimizations: CUDA graphs for predefined batch sizes, avoiding unnecessary CPU-GPU sync, and input streaming for speech-to-speech.
  • Future plans extend to image, video, world models, and fine-tuning, with vision of real-time multimodal inference.