Artificial Analysis: Gemini 4 Argon (high) ties with GPT‑6 Astra (max) on Intelligence Index, with 15% hallucinations versus 51% for Astra.
gemini google
| Source: Techmeme | Original article
Google's Gemini 4 Argon (high) ties GPT‑6 Astra (max) on the Artificial Analysis Intelligence Index, posting a 15% hallucination rate versus 51% for Astra.
Google’s latest Gemini 4 Argon model has closed the performance gap with OpenAI’s flagship offering, according to a fresh assessment from Artificial Analysis. The report shows Gemini 4 Argon at its “high” reasoning tier scoring 53 points on the Artificial Analysis Intelligence Index – the same score achieved by OpenAI’s GPT‑6 Astra running at its maximum setting. The two models also outpace the newly released GPT‑6.1 Sol, which posted 52 points.
Beyond raw intelligence, the analysis highlights a stark contrast in reliability. Gemini 4 Argon recorded a 15 percent hallucination rate, while GPT‑6 Astra’s rate stood at 51 percent. The disparity suggests that Google’s model may deliver more factual outputs even as it reaches parity on the benchmark.
The finding matters for several reasons. First, it re‑establishes Google as one of the top three AI labs in terms of measured reasoning ability, a status it briefly lost after OpenAI’s rapid model upgrades earlier this month. Second, the lower hallucination figure could make Gemini 4 Argon more attractive for enterprise and consumer applications where factual accuracy is paramount. Finally, the result adds pressure on OpenAI and Anthropic, whose Claude series still leads the index, to improve both intelligence scores and hallucination mitigation.
As we reported on 30 September, OpenAI’s GPT‑6.1 Sol replaced GPT‑6 Sol after just a week, edging close to Astra‑level performance. The next steps will likely involve further benchmark releases, cost‑structure disclosures, and possible refinements to Gemini 4 Argon’s latency and pricing, all of which could shift the competitive balance in the coming weeks.
Sources
Back to AIPULSEN