# ARC Prize says OpenAI’s o3 hit a breakthrough score on the ARC-AGI public benchmark
ARC Prize said on 2024-12-20 that OpenAI’s new o3 system reached 75.7% on the ARC-AGI public leaderboard’s semi-private evaluation set under the organization’s stated $10,000 compute limit. In the same analysis, ARC Prize said a high-compute configuration achieved 87.5%, making the result one of the most striking benchmark jumps yet reported for a GPT-family model.
The claim matters because ARC-AGI is designed to test adaptation to novel tasks rather than memorization or familiar pattern matching. ARC Prize framed the result as evidence that o3 represents a real step forward in generalization. The organization contrasted it with the earlier arc of GPT models, noting that progress on the benchmark had been slow for years and that o3’s result forced a reassessment of what current AI systems can do.
The source is careful to add a second point: efficiency now matters as much as raw score. ARC Prize says it documented compute cost and cost per task because stronger results came with much heavier inference budgets. In other words, the benchmark is not just measuring whether the system can solve a challenge, but whether it can do so within practical limits.
That caveat is important. The post argues that o3’s gains are not just the product of brute-force scaling. It presents the system as evidence that new architectural ideas are still delivering meaningful progress. At the same time, it warns against equating an ARC-AGI score with full AGI. The organization says the benchmark is a research tool, not an acid test for general intelligence.
The broader story is one of both excitement and restraint. ARC Prize says a smarter benchmark is still needed because the field will have to keep distinguishing between systems that can replay known patterns and systems that can genuinely adapt to something new. The post also says that humans still solve these tasks more cheaply, which means the current frontier is still far from economically replacing people on arbitrary novelty-heavy work.
For AI watchers, the report is significant because it suggests that progress is no longer only about scaling existing models. The authors argue that newer ideas and better system design now matter as much as raw parameter count. That is the kind of claim that tends to shape product strategy, research funding, and benchmark design all at once.
Claim map
The supplied ARC Prize post supports the claims that o3 scored 75.7% at the stated public leaderboard limit, that a higher-compute configuration reached 87.5%, that ARC Prize treats the result as a major benchmark breakthrough, and that the group still considers ARC-AGI a research tool rather than proof of AGI.


