Kimi K3 and GLM‑5.3 outperform Gemini 3.8 Flash
agents claude deepseek gemini gemma qwen
| Source: HN | Original article
Kimi K3 and GLM‑5.3 outperform Gemini 3.8 Flash, according to recent comparisons.
Kimi K3 and GLM‑5.3 have emerged as the strongest challengers to Google’s Gemini 3.8 Flash, according to recent benchmark comparisons and community reviews.
As we reported on 2 September 2026, Gemini 3.8 Flash was positioned as a fast, cost‑effective option for developers, with an introductory price of $0.75 per 1 M input tokens and $3.75 per 1 M output tokens until the end of the year. New data now shows that both Kimi K3 – a 2.8 trillion‑parameter open‑weight mixture‑of‑experts model – and GLM‑5.3‑Flash outperform Gemini 3.8 Flash across a range of metrics, especially in long‑context agent workloads and complex coding tasks.
Kimi K3’s design emphasises local deployment and cluster‑scale processing, making it attractive for organisations that need to run large models on‑premise rather than rely on cloud APIs. GLM‑5.3, built on the same base as GLM‑5.2 but refined through post‑training, delivers faster online inference and benefits from a robust Chinese ecosystem, though it is less focused on open‑local downloads than K3.
The shift matters because it broadens the competitive landscape beyond Google’s flash‑oriented pricing model. Developers who have been juggling multiple subscriptions for GPT, Claude, DeepSeek and now Gemini can now consider a single‑key, pay‑as‑you‑go approach that includes Kimi K3 or GLM‑5.3, potentially lowering costs while gaining better performance on long‑horizon tasks.
Going forward, the community will watch for updated side‑by‑side benchmark tables, pricing adjustments from Google, and any announcements of next‑generation versions from Kimi and the GLM team. How quickly enterprises adopt the more capable, locally deployable models could reshape the balance of power in the AI‑as‑a‑service market.
Sources
Back to AIPULSEN