Hasty Briefsbeta

Bilingual

Google just bet its inference future on a chip built for one model

2 hours ago
  • Google's Frozen v2 chip hardwires parts of Gemini's architecture but keeps weights updatable to avoid obsolescence.
  • The chip aims to deliver 6-10 times more tokens per watt than current Google AI chips, helping alleviate the AI compute crunch.
  • This represents a shift toward model-specific silicon, similar to the evolution of Bitcoin mining from GPUs to ASICs.
  • Competitors like Taalas, d-Matrix, and SambaNova are also developing specialized hardware for AI inference.
  • The approach balances hyper-efficiency of hardwired architectures with flexibility for future model updates.