Back to news
researchNVIDIA2026-08-22

Nvidia AVO agent architecture scores perfect 100 on ARC-AGI-3 long-horizon benchmark

On Aug 22 Nvidia AVO agent architecture achieved a perfect 100.00 RHAE score on the ARC-AGI-3 interactive reasoning benchmark, completing all 183 levels across 25 environments using 6,624 environment actions, validating that system-level architecture can convert model capability into sustained autonomous progress.

On August 22, Nvidia AVO agent architecture achieved a perfect 100.00 RHAE score on the ARC-AGI-3 interactive reasoning benchmark. ARC-AGI-3 is designed to evaluate long-horizon interactive reasoning, requiring agents to autonomously complete a cumulative 183 levels across 25 public game environments. Nvidia AVO completed all tasks using 6,624 environment actions, becoming the first commercial agent architecture to achieve a perfect score on the benchmark.

The core feature of AVO is its persistent memory and supervision mechanism, which lets the agent maintain context coherence across long time horizons and optimize its action policy online. Nvidia emphasized in its disclosure that AVO's success demonstrates that system-level architecture, not the model itself, determines actual performance on long-horizon autonomous tasks. This means that even with comparable base models, differences in the execution framework layered on top can drive order-of-magnitude gaps in final task performance.

The ARC-AGI-3 benchmark is maintained by the François Chollet team and is positioned to measure whether systems possess adaptive learning and cross-environment generalization. It has long been considered an important reference point for tracking progress toward AGI. Nvidia AVO's perfect score is widely regarded as a milestone event for the Agent era, indicating that the execution framework around the model has become an independent technical battleground.

Industry observers note that AVO's perfect score, together with recent moves such as OpenAI open-sourcing Codex Harness and Anthropic shipping Claude Code 2.1, points to a common trend: as models commoditize, the execution framework around them is becoming the differentiated high ground.

英伟达AVOARC-AGI-3智能体长周期