Some Google staff say Gemini 4 excels on benchmarks but falters on real‑world coding; Google denies it, Bloomberg reports
benchmarks gemini google
| Source: Techmeme | Original article
Google's new Gemini 4 AI model shows strong benchmark results but some employees say it falters on real‑world coding tasks, a claim the company denies.
Google has begun rolling out Gemini 4 Argon, its flagship AI model, but internal reports suggest the launch is not without friction. According to Bloomberg, a number of Google engineers say the system scores strongly on the industry benchmarks that are typically used to gauge large‑language‑model performance, yet it “stumbles” when they test it on real‑world coding assignments, particularly front‑end design tasks. Some staff argue that rival models from Anthropic (Fable) and OpenAI (Astra) are advancing more quickly, while others maintain that Gemini 4 has already closed the gap.
Google has pushed back against the characterization, insisting that the model meets its own internal standards for coding and security. The company’s public messaging, echoed in recent coverage, emphasizes Gemini 4 Argon’s strengths in coding and security tests and notes that the first external users are trusted cyber‑defenders participating in a voluntary pre‑release program.
The dispute matters because developer‑focused capabilities are a key battleground for AI providers. If Gemini 4’s performance on practical programming tasks falls short of expectations, it could slow adoption among the developer community and give competitors a foothold in a market where speed, reliability and low hallucination rates are prized. Earlier this week we reported on the model’s rollout to cyber defenders and on its competitive standing against OpenAI’s GPT‑6 Astra on an intelligence index.
What to watch next: further internal evaluations and any public benchmark releases that could clarify Gemini 4’s real‑world coding proficiency; Google’s response to employee concerns; and whether the model’s rollout expands beyond the initial defender cohort amid the broader race for developer‑centric AI tools.
Sources
Back to AIPULSEN