When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation
3 hours ago
- AI benchmarks are crucial for measuring model progress and guiding deployment decisions.
- Benchmarks tend to saturate quickly, making it difficult to differentiate models and reducing their long-term value.
- The study defines benchmark saturation and analyzes 60 language model benchmarks using 14 properties.
- Nearly half of the benchmarks show saturation, with rates increasing as benchmarks age.
- Resilience to saturation is influenced by expert curation, not by the availability of public test data.
- Design choices can help extend benchmark longevity and support more durable evaluation methods.