- Gwern proposes that overtraining, or 'grokking,' large neural networks on small datasets could lead to human-like AI by forcing deeper generalization.
- Current LLMs fail to generalize as well as humans because they are trained on massive data with minimal overtraining, unlike the grokking process.
- Grokking involves a sudden capability jump after prolonged training, as models shift from memorization to understanding underlying rules.
- Gwern suggests training a 100-trillion-parameter model on a constrained dataset to induce grokking, opposite to current AI lab practices.
- Political and technical obstacles, including high costs and apparent stagnation, may prevent labs from attempting this risky approach.
- The article argues that grokking, not pure scaling or reasoning, might be a viable path to artificial superintelligence.