Report
Gemini 3.1 Flash TTS Adds Granular Audio Tags for Precise Speech Control
Google DeepMind has released Gemini 3.1 Flash TTS, a new text-to-speech model that introduces granular audio tags for precise control over AI speech. It is now available across Google products.
The source states that Gemini 3.1 Flash TTS is the next generation of expressive AI speech. It introduces granular audio tags that give precise control to direct AI speech for expressive audio generation. The model is now available across Google products. That is the core of the announcement: a new text-to-speech model with a specific new feature—granular audio tags—and immediate availability in Google's ecosystem.
What does this mean in practice? The key new capability is the granular audio tags. According to the source, these tags allow you to direct AI speech with precision. This suggests that instead of just typing text and getting a generic voice, you can insert tags to control aspects of the delivery. The source does not specify what the tags look like or what parameters they control, but the emphasis on 'granular' and 'precise control' implies a level of fine-tuning not previously available. For people using AI for everyday work, this could mean more natural-sounding voiceovers, more expressive narration, or the ability to match a specific tone without manual editing.
Our interpretation is that this is a move toward making text-to-speech more controllable and less of a black box. The practical upshot is that you can potentially get the exact delivery you want in one pass, rather than generating multiple takes and hoping. However, the source provides no examples, no list of tags, and no comparison to previous models. So while the capability sounds promising, the details are thin. We cannot yet say how much control you actually get or how easy it is to use. The fact that it is available across Google products suggests it is being positioned for broad use, not just as a developer API.
Source details and supporting facts
Each line is stated by the page named above it.
Stated by deepmind.google
- Gemini 3.1 Flash TTS is now available across Google products.
- The model introduces granular audio tags that give precise control to direct AI speech for expressive audio generation.
- Gemini 3.1 Flash TTS is the next generation of expressive AI speech.
Sources
- Google DeepMind blogText stored 15 September 2026
How this story was checked. Written from the 1 page listed above, stored 15 September 2026; claims checked against that stored text on 15 September 2026.
What that means
- 3 of 4 reported statements were confirmed against the page that carries them; the rest were removed rather than published.
- Figures in the text were required to appear in the stored source text: yes. Identifiers: yes.
- The check reads stored text only: no claim rests on a fresh look that did not happen.
- Where the reporting was silent, the text says so instead of filling the gap.