5 hours ago
- Thinking Machines released Interaction Models, a fully-duplex voice model that can listen and speak in micro-turns, enabling natural interruptions.
- The model delegates complex reasoning tasks to a more intelligent background model via tool calls, allowing fast conversation while maintaining intelligence.
- It incorporates video input to read facial expressions, making conversations more human-like.
- The model is scaled up significantly (2x active parameters of Moshi, 40x total parameters), enabling multimodal capabilities.
- The article notes that while some features are not entirely new, the scale and video integration are genuine advancements.