The benchmark listing emphasizes both accuracy and per-task cost across several reasoning levels. That is the central point in evidence dated 2026-08-07. DeepSeek V4 Flash posts low-cost results on ARC-AGI tests.

At maximum effort, the model scored 89.0% on the semi-private ARC-AGI-1 set for $0.02 per task and 61.4% on the corresponding ARC-AGI-2 set for $0.04.

The account also explains why the subject has attracted attention. DeepSeek V4 Flash 0731 was evaluated on ARC-AGI-1, ARC-AGI-2 and ARC-AGI-3, with pass-or-fail outcomes reported for each reasoning setting. Rather than implying a universal conclusion, the details show how one institution, project or author is approaching a defined problem.

ARC Prize results show a high-effort reasoning configuration reaching 89.0% on ARC-AGI-1 Semi-Private and 61.4% on ARC-AGI-2 Semi-Private at only a few cents per task.

Only the supplied arcprize.org material was available for this report. The wording therefore preserves attribution and avoids extending the claims beyond the documented example. It is especially important not to confuse a proposed capability with measured deployment, an author's argument with consensus, or a court order with the end of every legal process.

Even with those boundaries, the evidence offers a useful snapshot. It identifies the relevant actors, the stated rationale and the measurable elements available at the time. The next test will come from implementation, response or further scrutiny, depending on the subject; until then, the most defensible account is the limited one supported here.

The source packet also sets a clear reporting boundary: exact claims can be repeated with attribution, while absent technical, legal or comparative detail cannot responsibly be reconstructed. That distinction matters for evaluating the item on its own terms and for avoiding conclusions that the available record does not support.

The source packet also sets a clear reporting boundary: exact claims can be repeated with attribution, while absent technical, legal or comparative detail cannot responsibly be reconstructed. That distinction matters for evaluating the item on its own terms and for avoiding conclusions that the available record does not support.