Nvidia says its general-purpose coding agent AVO hit 100% across all 25 environments in the ARC-AGI-3 public set
agents claude nvidia
| Source: Techmeme | Original article
Nvidia reports its AVO coding agent achieved perfect scores across all 25 ARC‑AGI‑3 environments, completing every one of the 183 levels.
Nvidia announced that its general‑purpose coding agent, AVO, achieved a perfect score on the ARC‑AGI‑3 benchmark, solving every one of the 183 levels across the suite’s 25 public environments. The result, posted on the company’s technical blog by Terry Chen, marks a jump from the 30 % baseline of the underlying Claude Opus 5 model to 100 % after the AVO architecture was applied. The agent completed the test set in 6,624 actions, roughly 12 % fewer steps than the 7,542 actions recorded in prior runs, and did so without any explicit rules, goals or hand‑crafted instructions.
The achievement matters because ARC‑AGI‑3 is designed to probe long‑horizon, reasoning‑heavy tasks that have stumped most large‑language models. AVO’s success demonstrates that augmenting a language model with a “harness” of persistent memory, supervision and tool‑use can unlock autonomous problem‑solving at scale. In a separate GPU‑kernel optimisation experiment, the same architecture explored more than 500 modification directions, committed 40 kernel versions and delivered up to a 10.5 % performance lift, underscoring its practical value for hardware‑centric workloads.
As we reported earlier this month, Nvidia has been emphasizing the role of the AI harness over the base model itself. AVO’s performance suggests that this approach can translate into tangible gains on both reasoning benchmarks and real‑world code optimisation. The next steps to watch include whether Nvidia will open the AVO framework to external developers, how the system scales to the private portion of ARC‑AGI‑3, and if competing firms can replicate the results with their own harnesses. Follow‑up research on long‑horizon autonomy and the balance between model capability and execution infrastructure will likely shape the next wave of coding‑agent breakthroughs.
Sources
Back to AIPULSEN