Why didn't we get GPT-2 in 2005?
6 hours ago
- GPT-2 required approximately 10^21 FLOPs, with multiple estimates from various sources.
- The BlueGene/L supercomputer (2005) could have trained GPT-2 in about 41 days, but it wasn't used for AI.
- Large language models weren't invented yet due to missing techniques like transformers and Adam optimization.
- Supercomputers like BlueGene/L were dedicated to nuclear weapons, not AI research.
- LLM development followed a feedback loop: cheaper compute enabled experiments, which attracted researchers and investment, accelerating progress.
- Unlike the moon landing, early LLM progress was uncertain and evolved step-by-step rather than following a grand plan.
- Scaling laws and clearer rules around 2018 made AI progress more predictable and faster.
- Keeping AI secrets is ineffective because demos reveal possibilities and trigger investment by competitors; only a massive tech lead and low prices could deter competition.