Canto: A speech model built for the real world
5 hours ago
- Canto is a new speech model by Wispr Advanced Interfaces Lab for real-time dictation in noisy, real-world conditions.
- The model achieved the lowest word error rate (WER) on real-world dictation evaluations compared to Google, OpenAI, AssemblyAI, and Deepgram models.
- Canto ranks second overall on a challenge set with difficult audio, but leads among real-time transcription models, tying on low-volume speech and short dictations.
- On public benchmarks (LibriSpeech, FLEURS, Common Voice), Canto performs competitively, though it only ties for lowest WER on LibriSpeech.
- Training involved supervised fine-tuning followed by reinforcement learning (using GRPO) to improve transcript quality and focus on real-world errors.
- The post-training infrastructure enables scaling, including using user corrections and contextual vocabulary to improve recognition.
- Methods like grafting corrections and experiments with context handling show a trainable balance between following context and trusting audio.
- Canto's successor is being trained at larger scale with goals for better noisy audio, multi-speaker diarization, broader languages, and better context use.
- Future work includes unified transcription and diarization, and integrating speech models with user intent and application context.
- The team is hiring researchers for speech recognition, RL, diarization, multilingual, and multimodal interface work.