Hasty Briefsbeta

Bilingual

Canto: A speech model built for the real world

5 hours ago
  • Canto is a new speech model by Wispr Advanced Interfaces Lab for real-time dictation in noisy, real-world conditions.
  • The model achieved the lowest word error rate (WER) on real-world dictation evaluations compared to Google, OpenAI, AssemblyAI, and Deepgram models.
  • Canto ranks second overall on a challenge set with difficult audio, but leads among real-time transcription models, tying on low-volume speech and short dictations.
  • On public benchmarks (LibriSpeech, FLEURS, Common Voice), Canto performs competitively, though it only ties for lowest WER on LibriSpeech.
  • Training involved supervised fine-tuning followed by reinforcement learning (using GRPO) to improve transcript quality and focus on real-world errors.
  • The post-training infrastructure enables scaling, including using user corrections and contextual vocabulary to improve recognition.
  • Methods like grafting corrections and experiments with context handling show a trainable balance between following context and trusting audio.
  • Canto's successor is being trained at larger scale with goals for better noisy audio, multi-speaker diarization, broader languages, and better context use.
  • Future work includes unified transcription and diarization, and integrating speech models with user intent and application context.
  • The team is hiring researchers for speech recognition, RL, diarization, multilingual, and multimodal interface work.