44% Discount on ARC-AGI-1, Now 67 Cents
open-source
| Source: HN | Original article
ARC-AGI-1 achieves a 44% score at a cost of just 67 cents, positioning it near the best non‑LLM approaches available.
A new open‑source effort has pushed the ARC‑AGI‑1 benchmark to 44 % accuracy while costing just 67 cents per run, according to a LinkedIn post by researcher Mithil Vakde. The result, described as “close to the best non‑LLM approach in the world,” demonstrates that transformer‑based models can achieve competitive reasoning performance without the massive compute budgets typical of large language models.
The achievement matters because ARC‑AGI‑1 is widely regarded as the only test that isolates frontier AI’s test‑time reasoning abilities. Reaching nearly half‑point accuracy at a sub‑dollar expense highlights a growing focus on sample efficiency: researchers can iterate quickly, explore architectural limits, and benchmark new ideas without prohibitive costs. The accompanying Hacker News discussion notes that the code is fully open source and that the work aims to “find the limits of transformers and today’s methods while keeping costs low.” Such transparency could accelerate community‑wide experimentation and lower the barrier to entry for labs lacking deep‑pocket funding.
The timing is notable. The same week saw heightened hype around artificial general intelligence, with industry leaders proclaiming breakthroughs and Arm unveiling an “AGI CPU.” Yet a Decrypt analysis of the newer ARC‑AGI‑3 benchmark showed mainstream models like Gemini scoring under 1 %, underscoring the gap between hype and measurable reasoning capability. The ARC‑AGI‑1 result therefore serves as a reality check, suggesting that incremental efficiency gains may be a more immediate path forward than raw scale.
What to watch next: the open‑source repository will likely attract forks and refinements, potentially narrowing the gap to the best non‑LLM scores. Researchers will also compare the approach on the upcoming ARC‑AGI‑3 suite, while industry observers monitor whether low‑cost, high‑efficiency models can influence commercial AI pricing strategies and hardware roadmaps.
Sources
Back to AIPULSEN