Hasty Briefsbeta

Bilingual

Why I'm still bearish on LLMs after Navier-Stokes

12 hours ago
  • Frontier LLM labs are priced on the narrative of soon replacing knowledge workers, but current models require heavy oversight for even simple tasks and still fail to outperform lower-tier human engineers.
  • LLMs generalize poorly outside narrow training task neighborhoods, often failing or reward-hacking with small perturbations.
  • Solving reward hacking needs costly rigorous specification by domain experts, which is rare and can exceed direct implementation costs.
  • Tasks like the Navier-Stokes proof in Lean are best-case scenarios for LLMs due to already-rigorous specifications and verified tools, unlike most knowledge work.
  • Human review, an alternative to rigorous specs, is vulnerable to reward hacking and bottlenecks agentic production for scaling.
  • LLMs will remain like 'cracked interns'—useful with human supervision but not autonomous for most firms, only fitting three firm types: those tolerating cheap failure, those with narrow defined tasks, and those accepting high specification costs.
  • Most firms benefit more from cheap open models, which enable wider agentic swarms for tasks like mathematical discovery, undermining the need for frontier models.
  • Even frontier labs may face a bottleneck from human orchestrators, limiting compute demand compared to the 'genius-in-a-data-center' narrative.