Tech

ARC Prize updates leaderboard to spotlight efficiency in interactive AI challenges

The ARC Prize has revised its ARC-AGI Leaderboard to highlight systems costing under $10,000 to run, marking a strategic pivot towards resource-efficient artificial intelligence.

Author
Owen Mercer
Markets and Finance Editor
Published
Draft
Source: Hacker News · original
Tech
No image available
New ARC-AGI-3 benchmark shifts focus from passive testing to cost-per-task performance

The ARC Prize has released an update to its ARC-AGI Leaderboard, signalling a distinct evolution in how artificial general intelligence is measured. The latest iteration, ARC-AGI-3, moves away from the passive fluid intelligence assessments that characterised versions 1 and 2, replacing them with interactive environment challenges that require AI agents to adapt on the fly to novel scenarios.

This structural shift is reflected in the leaderboard’s new visualisation, which correlates cost-per-task with performance to emphasise efficiency. The ARC Prize states that true intelligence involves solving problems with minimal resources, a philosophy now embedded in the benchmark’s design. The current display filters for systems with estimated running costs below $10,000, aiming to highlight models that deliver results without prohibitive computational overhead.

Data presentation on the updated board includes specific caveats regarding provisional estimates. For models pending retest, costs are estimated based on Gemini 3 Pro pricing. Similarly, an ARC-AGI-2 score estimate is derived from partial testing results and o1-pro pricing. The ARC Prize notes that these figures are subject to change and should not be treated as final operational costs until full retesting is complete.

The leaderboard also distinguishes between official and unofficial metrics. Results marked as "preview" are unofficial and may stem from incomplete testing, meaning they may not accurately reflect final performance metrics. For models unable to produce full test outputs, remaining tasks are automatically marked as incorrect, ensuring a consistent standard for incomplete runs.

This update underscores a broader industry focus on the economic viability of advanced AI systems. By filtering for efficiency and interactive capability, the ARC Prize is providing investors and researchers with a clearer view of which models are not only capable but also cost-effective in dynamic environments. The full testing policy and detailed metrics are available on the ARC Prize website.

Continue reading

More from Tech

Read next: Google Proposes Blocking Local ADB Connections, Threatening Open-Source Android Ecosystem
Read next: Low-cost piercing pillow challenges $200 sleepbuds market
Read next: NASA Deep Space Network strained as wildfire forces Madrid complex evacuation