Vals, backed by Andreessen Horowitz, aims to become the gold standard for AI benchmarking
benchmarks
| Source: TechCrunch | Original article
Vals, backed by Andreessen Horowitz, aims to set the gold standard for AI benchmarking, seeking to provide a neutral, trustworthy resource amid a flood of AI models.
Vals AI, a new venture backed by Andreessen Horowitz, announced its ambition to become the “gold standard” for artificial‑intelligence benchmarking. The company says it will offer a neutral, trustworthy platform for evaluating AI models, positioning itself as a counterweight to the flood of proprietary and often incomparable performance claims that dominate the market today.
The move arrives at a time when developers, enterprises and investors are grappling with a bewildering array of models, each touting its own set of metrics. A reliable, independent benchmark can help cut through hype, guide procurement decisions and steer research toward genuine progress. By emphasizing transparency and reproducibility, Vals hopes to restore confidence that model scores reflect real‑world capability rather than curated test sets.
Vals’ entry adds momentum to a broader push for standardized evaluation. Earlier this month we covered BVB’s programmatic reconstruction benchmark for video understanding and Real‑SWE’s effort to test models on private enterprise codebases, both of which illustrate the community’s appetite for more rigorous, domain‑specific metrics. Vals aims to unify these fragmented efforts under a single, widely accepted framework.
What to watch next is whether Vals can attract enough high‑profile participants to make its suite the de‑facto reference point. Adoption by major cloud providers, integration with upcoming AI model releases and potential collaboration with the industry‑led standards discussions that have been underway since July will be key indicators. If Vals succeeds, it could shape how performance is reported across the AI ecosystem for years to come.
Sources
Back to AIPULSEN