Hasty Briefsbeta

Bilingual

ARC-AGI Leaderboard

4 hours ago
  • ARC-AGI-3 measures AI agents' ability to adapt on the fly to novel interactive environments, evolving from earlier passive fluid intelligence tests.
  • The leaderboard scatter plot shows the relationship between cost-per-task and performance, emphasizing efficiency as a key measure of intelligence.
  • Three types of data are plotted: Reasoning Systems Trend Lines (showing effect of reasoning time), Base LLMs (single-shot inference), and Kaggle Systems (competition-grade submissions under strict cost constraints).
  • Only systems costing less than $10,000 to run are shown; incomplete test outputs are marked incorrect, and preview results are unofficial.