Gemini Releases 3.8 Text-to-Speech
gemini google speech voice
| Source: HN | Original article
The Gemini family adds two new text-to-speech models, Gemini 3.8 Flash TTS and Gemini 3.8 Flash‑Lite TTS, turning voice generation into a dynamic creative studio.
Google has rolled out two new text‑to‑speech models – Gemini 3.8 Flash TTS and Gemini 3.8 Flash‑Lite TTS – as part of the latest Gemini 3.8 update. Announced on September 23 via the Google Blog, the models move voice generation from static presets to a “dynamic creative studio”, letting developers, creators and enterprises craft richer, more expressive audio by describing desired voice traits in natural language and fine‑tuning style, accent, pace and tone with structured metadata and inline vocal tags.
The launch expands Gemini’s audio capabilities across Google AI Studio, the Gemini API, Gemini Enterprise, Gemini Notebook and Google Vids. Both models support over 100 languages and more than 2,000 distinct voice profiles, with Flash‑Lite positioned for lower‑latency, lightweight use cases while Flash delivers the highest fidelity. The controllable TTS pipeline means a single prompt can produce multi‑speaker output or shift dialects on the fly, opening new possibilities for interactive assistants, e‑learning, gaming and localized media.
Why it matters is twofold. First, the breadth of language and voice options narrows the gap between global content creators and high‑quality synthetic speech, potentially reducing reliance on costly human voice talent. Second, the integration into existing Gemini products signals Google’s intent to make voice a first‑class modality in its AI ecosystem, reinforcing its competitive stance against rivals such as OpenAI’s voice models and Amazon Polly.
As we reported on September 23, the Flash series is Google’s “most expressive audio generation models yet”. The next steps to watch include how quickly third‑party developers adopt the API, whether the models will be bundled into consumer‑facing services like Google Assistant, and if Google will extend the controllable TTS framework to multimodal generation or real‑time streaming scenarios.
Sources
Back to AIPULSEN