Inflect-Micro-v2: complete voice in 9.36M parameters
3 hours ago
- Inflect-Micro-v2 is a fixed-voice English TTS model with under 10 million parameters, supporting CPU and CUDA inference with deterministic seeds and long-text handling.
- The model prioritizes quality (Micro) or footprint (Nano), with synthetically generated audio evaluated via human preference, UTMOS22, and multi-ASR intelligibility tests.
- Both Inflect variants (Micro and Nano) offer complete local text-to-waveform synthesis, including a 24 kHz waveform decoder, Python API, and ONNX runtime support.
- The release includes frozen evaluation protocols, competitive benchmarks against other compact TTS systems, and transparent reporting of metrics and limitations.
- The project is open-source under Apache-2.0, with a single fixed synthetic male voice, and does not support voice cloning or languages beyond English currently.