New audio model Gemini 3.8 Live reasons in real time
gemini google
| Source: Mastodon | Original article
Google's new Gemini 3.8 Live audio model can reason in real time and comprehend speech even when users switch languages mid‑conversation.
Google has rolled out Gemini 3.8 Live, a new speech‑to‑speech model that can reason in real time and keep a coherent thread across complex tasks. The model, together with a “Live Extended Thinking” variant, is designed for voice‑based agents that can call external tools in the background and understand 97 languages. A key feature highlighted by the company is the ability to follow a conversation even when the user switches languages mid‑dialogue, a step toward truly multilingual assistants.
The launch follows the September 2 release of Gemini 3.8 Flash, the third Flash update in six weeks, which was tuned for programming, autonomous agents and multi‑step reasoning. Gemini 3.8 Live builds on that foundation, extending the multimodal capabilities of the Flash line—text, images, audio and video inputs with a context window of up to one million tokens—into the spoken domain. Real‑time reasoning and tool‑calling give developers the chance to create agents that can, for example, fetch up‑to‑date information or trigger actions while the user is still speaking.
The announcement matters because it narrows the gap between text‑only large language models and truly interactive voice assistants. By handling language switches seamlessly and maintaining longer conversational context, Gemini 3.8 Live could raise the bar for products ranging from smart speakers to enterprise call‑center bots, and it may pressure rivals to accelerate their own voice‑centric research.
Going forward, the AI community will watch for performance benchmarks, integration demos in Google’s own services, and how developers leverage the “Extended Thinking” mode. Privacy safeguards and latency figures will also be scrutinised as the model moves from lab to consumer‑facing applications.
Sources
Back to AIPULSEN