The Coming AI Compute Crunch
a day ago
- AI token consumption per user has surged dramatically, with the author estimating a 50x increase in daily usage over three years, driven by advanced models like GPT-4, Sonnet 3.5, and Opus 4.5, and the rise of agentic workflows.
- The explosive growth in LLM users (estimated at 1 billion active users) and per-user token usage is fueling massive infrastructure investments, with hyperscalers like AWS, Azure, and GCP committing hundreds of billions in capex for datacenter buildouts.
- Despite the planned infrastructure spending, actual deployment faces hard constraints, particularly electrical power shortages and a critical shortage of high-bandwidth memory (HBM DRAM), which is essential for AI chips and takes years to ramp production.
- Macquarie research indicates current DRAM supply can only support the rollout of 15GW of AI infrastructure, roughly enough for 30 million heavy agentic users at a million tokens per day, which the author believes will be insufficient given growing demand from video, audio, and world models.
- Pricing dynamics are expected to shift, with likely resistance to large price increases due to low switching costs, but possible adoption of dynamic inference pricing (off-peak discounts) and reduced free plans, alongside increased focus on model efficiency research until DRAM capacity expands.
- Wildcards include frontier labs reserving models for internal use, innovations like SRAM-based memory architectures (e.g., Groq), and the risk of underbuilding AI capacity despite non-binding commitments, with DRAM shortages defining the industry in the next few years.