DeepSeek V4 Flash 0731 tops ARC-AGI benchmarks with strong reasoning scores
Released on 31 July 2026, the new variant demonstrates high-cost efficiency and robust performance across semi-private reasoning tasks, according to data published by the ARC Prize foundation.
DeepSeek has released a new model variant, V4 Flash 0731, which has recorded top-tier scores on the ARC-AGI benchmark suite. Released on 31 July 2026, the model achieved a maximum effort score of 89.0 per cent on the ARC-AGI-1 Semi-Private benchmark and 61.4 per cent on the ARC-AGI-2 Semi-Private benchmark. The results were published by the ARC Prize foundation, highlighting the model's capabilities in abstraction and reasoning tasks.
The evaluation covered three distinct reasoning variants: Max, High, and Low. In addition to the primary semi-private scores, the High variant recorded 87.0 per cent on ARC-AGI-1 and 56.0 per cent on ARC-AGI-2. The Low variant achieved 84.0 per cent and 46.0 per cent respectively. No scores were reported for the ARC-AGI-3 benchmark in the available data.
Cost efficiency appears to be a significant feature of the release. The data indicates that operating at maximum effort cost $0.02 per task on the ARC-AGI-1 evaluation, which comprised 400 tasks. The ARC-AGI-2 evaluation, consisting of 120 tasks, incurred a cost of $0.04 per task at the same effort level.
Detailed task-level results were published for both the ARC-AGI-1 and ARC-AGI-2 public evaluations. These records list individual task IDs and their pass or fail outcomes across the three reasoning variants. The ARC-AGI-1 public evaluation included 400 tasks, while the ARC-AGI-2 public evaluation involved 120 tasks, providing a granular view of the model's performance on specific logical challenges.
The ARC-AGI suite is designed to test reasoning capabilities in artificial intelligence models, moving beyond standard pattern recognition to assess abstract thinking. The benchmarks are divided into semi-private and public evaluation sets, with the semi-private scores often reflecting performance on tasks not included in public leaderboards. The model name '0731' corresponds to its release date of 31 July.
While the source material does not specify the underlying architecture or parameter count of the V4 Flash 0731 model, the performance metrics suggest a significant advancement in reasoning efficiency. The availability of detailed pass/fail data allows for a precise assessment of where the model excels or struggles within the benchmark constraints.

