ARC-AGI Leaderboard
4 hours ago
- ARC-AGI-3 measures AI agents' ability to adapt on the fly to novel interactive environments, evolving from earlier passive fluid intelligence tests.
- The leaderboard scatter plot shows the relationship between cost-per-task and performance, emphasizing efficiency as a key measure of intelligence.
- Three types of data are plotted: Reasoning Systems Trend Lines (showing effect of reasoning time), Base LLMs (single-shot inference), and Kaggle Systems (competition-grade submissions under strict cost constraints).
- Only systems costing less than $10,000 to run are shown; incomplete test outputs are marked incorrect, and preview results are unofficial.