The STT-LLM-TTS voice stack is dead
16 hours ago
- The author built call4me, a managed voice agent for making phone calls on behalf of users, after failing to find an existing solution.
- The system avoids the traditional cascaded STT-LLM-TTS stack by using GPT-Live for direct speech-to-speech, reducing latency and preserving tonal information.
- Before dialing, the agent collects all required information (profile fields and per-call requirements) to prevent surprises during the call.
- A 'back office' text model handles decisions and tools (e.g., pressing digits, asking the user), while the voice model focuses on conversation.
- The system includes mechanisms to handle phone menus (keypad input, audio tones for international calls) and recover from dead ends.
- Running on Cloudflare with Workers, Durable Objects, and Telnyx for telephony, the audio path is simple with no codec conversion because both GPT-Live and phone lines use G.711 μ-law.
- Key lessons: skip audio conversions, split voice from decisions, verify model actions, and gather all info before dialing.