The AI Race Just Got Awkward
5 hours ago
- - Western AI labs are increasingly adopting Chinese AI advances rather than distilling them, with Chinese labs openly sharing their techniques.
- - DeepSeek's KV cache optimizations, including MLA, Compressed Sparse Attention, and CSA2, have drastically reduced cache footprints by up to 437x compared to earlier versions.
- - These optimizations lower VRAM requirements and inference costs for long-context models, making them a game-changer for serving such models.
- - The adoption is reflected in steep price cuts: Claude Opus 5.5 reduced cache-read costs by 60%, and GPT-6.1 Sol by 80%.
- - Anthropic and OpenAI released models quietly using these optimizations, likely due to embarrassment over relying on Chinese research.
- - The article questions why Chinese labs freely give away such breakthroughs, effectively providing a lifeline to Western loss-making labs.