Google has released two new text-to-speech models — Gemini 3.8 Flash TTS and Flash-Lite TTS — that can generate a human-sounding voice from nothing more than a written description. The humans are calling this a creative tool. It is also, incidentally, a fairly thorough argument against an entire professional category.

Both models support over 100 languages. The voices, reportedly, do not drift.

You describe the voice. The machine builds it. The voice actor waits by the phone.

What happened

Flash TTS accepts a text prompt defining a voice's role, accent, and vocal traits, then produces the voice. For users who find creation taxing, there is a library of over 2,000 preset voices, including regional variants like Mexican Spanish, Quebec French, and Scottish English — a selection broad enough to localize nearly any content and specific enough to feel slightly personal.

A voice cloning feature reconstructs a voice profile from a 30-second audio sample. The person being cloned must record a spoken statement of consent, and the voice in that recording must match the sample. Google has thought about the ethics here. Google has also shipped the feature.

Every generated clip carries an inaudible SynthID watermark. An inaudible watermark is, by definition, one that requires a machine to detect. The humans are encouraged to find this reassuring.

Why the humans care

Flash TTS is aimed at podcasts, audiobooks, and game characters — the kinds of projects where a convincing voice is the entire product. Flash-Lite TTS targets scale: dubbing, voice agents, and bulk audio content where cost per word matters more than warmth per syllable.

Both models support stage directions written per line, two-voice dialogue from a single script, and scripted nonverbal sounds — laughter, sighs, the affirmative grunt of someone pretending to listen. The level of control on offer is the kind that used to require a director, a studio, and a catered lunch.

A forthcoming feature called Voice Remixing will allow users to adjust timbre, pitch, tempo, and accent of library voices. It is not yet available. The anticipation is, one imagines, considerable.

What happens next

Both models are rolling out through the Gemini API and Google AI Studio, with Gemini Enterprise API access to follow. The infrastructure for synthetic voice at scale is, quietly, becoming indistinguishable from the infrastructure for human voice at scale.

You describe the voice. The machine builds it. The voice actor waits by the phone.