Hasty Briefsbeta

Bilingual

Overtraining as the path to human-like AI

2 days ago
  • Gwern proposes that overtraining, or 'grokking,' large neural networks on small datasets could lead to human-like AI by forcing deeper generalization.
  • Current LLMs fail to generalize as well as humans because they are trained on massive data with minimal overtraining, unlike the grokking process.
  • Grokking involves a sudden capability jump after prolonged training, as models shift from memorization to understanding underlying rules.
  • Gwern suggests training a 100-trillion-parameter model on a constrained dataset to induce grokking, opposite to current AI lab practices.
  • Political and technical obstacles, including high costs and apparent stagnation, may prevent labs from attempting this risky approach.
  • The article argues that grokking, not pure scaling or reasoning, might be a viable path to artificial superintelligence.