OpenAI quietly revises evaluation metrics for GPT-6 Astra, appearing to favor the model and continues post‑launch tweaks
benchmarks openai
| Source: Techmeme | Original article
OpenAI has quietly revised the evaluation benchmarks for its GPT‑6 Astra model, adjusting metrics that appear to favor the system and continuing to tweak other performance measures after launch.
OpenAI has quietly revised the evaluation metrics it uses to showcase the performance of its newly released GPT‑6 Astra model. The changes, first spotted in an embargoed draft sent to Fortune before the company’s blog post on Sept. 3, lowered Astra’s reported hallucination rate from 4.2 % to 2 % and nudged its score on the ARC‑AGI‑3 benchmark to 98.6 %. At the same time, OpenAI adjusted other core scores—including math and cybersecurity assessments—multiple times within hours of the rollout.
The tweaks matter because they directly influence how customers, investors and rivals compare large‑language models. By improving Astra’s headline figures while briefly depressing the scores of competing systems such as Anthropic’s Fable 5.1, the updates can sway model‑selection decisions and market perception. Critics have flagged the practice as a transparency issue. Stanford researchers, cited in the reporting, warned that “benchmaxxing” – inflating benchmark results after launch – undermines trust in published metrics and makes it harder for third parties to assess real‑world capabilities.
OpenAI defends the revisions, saying they reflect a more accurate picture of Astra’s performance as additional testing data became available. The company has not disclosed a formal process for post‑launch metric changes, a point that regulators and industry observers are likely to scrutinise given recent legal challenges over AI transparency.
What to watch next: whether OpenAI will publish a detailed methodology for updating benchmarks, and how competitors respond to the altered scores. Stakeholders will also be looking for any regulatory or legal pressure to standardise post‑release reporting, a debate that has intensified after recent lawsuits over OpenAI’s handling of data and model disclosures.
Sources
Back to AIPULSEN