Hasty Briefsbeta

Bilingual

The data black hole at the center of AI

17 hours ago
  • Intelligence is defined as sample efficiency—how much data is needed to operate fluently.
  • Recent AI improvements come from scaling data and compute, not from improving sample efficiency.
  • Reinforcement learning acts as synthetic data generation by using compute to find good data.
  • AI models require vast amounts of domain-specific human expert data for competency.
  • Human experts generate tailored data for each skill, involving hundreds of specialists per domain.
  • Open-source models lag behind frontier models by only ~4 months, suggesting data is the key driver.
  • Frontier AI models are trained on 10s to 100s of trillions of tokens, millions of times more than a human lifetime.
  • Humans are far more sample efficient than AI—e.g., a teenager learns driving in ~20 hours, while self-driving models need orders of magnitude more data.
  • Common objections (evolution, multimodal data, scaling laws) do not bridge the sample efficiency gap.
  • Sample efficiency may not be critical for automating common white-collar tasks, as training costs are amortized across billions of sessions.
  • The long-term plan is to automate AI research itself to solve sample efficiency, but whether current AI can achieve that remains unclear.