Nvidia AVO achieves perfect score on ARC-AGI-3 interactive reasoning benchmark
agents benchmarks nvidia reasoning
| Source: HN | Original article
Nvidia's AVO AI system achieved a perfect 100% RHAE score on the ARC-AGI-3 benchmark, successfully completing all 183 levels across 25 environments.
Nvidia has announced that its autonomous‑agent platform AVO achieved a perfect score on the ARC‑AGI‑3 interactive reasoning benchmark. The system solved all 183 levels across the 25 public environments, registering a 100.00 RHAE rating while using roughly 12 % fewer actions than the prior‑year VISTA baseline.
The result highlights a shift in how progress on long‑horizon AI tasks is being measured. Nvidia’s technical blog stresses that the achievement stems from AVO’s system‑level architecture—the “harness” that orchestrates model inputs, tool use, state management and feedback—rather than raw model size or novelty. By swapping the task interface of the same agent loop and targeting ARC‑AGI‑3, the company demonstrated that a well‑designed agent framework can deliver frontier‑level performance even as leading large language models such as Claude Opus 5 and OpenAI’s GPT‑5.6 show incremental gains on the same leaderboard.
The benchmark’s reputation rests on its use of private test sets to guard against memorisation. Nvidia released results only on the public portion, leaving the private‑set performance unverified. Analysts note that while the 100 % public score is a clear technical milestone, it does not alone confirm general reasoning ability across unseen data.
Going forward, the AI community will be watching whether Nvidia publishes private‑set results or opens the AVO architecture for broader testing. Parallel efforts such as the MemTrapBench and FM‑Bench suites continue to probe memory and multi‑agent dynamics, and any cross‑benchmark validation could cement AVO’s claim to a new level of autonomous reasoning. Observers will also monitor how other firms respond—whether they adopt similar system‑centric designs or focus on scaling model parameters—to gauge the next wave of progress in interactive AI agents.
Sources
Back to AIPULSEN