Google DeepMind launched two Gemini 3.8 text-to-speech models on 23 September 2026 that can replicate a voice from 30 seconds of audio S¹. The announcement packs 2,000-plus production-ready voices across 100-plus languages and dialects, but ships with no pricing, no API access details, and no independent benchmarks S¹.

The pair splits along a familiar line. Gemini 3.8 Flash TTS targets creative work: building entirely new voices from natural-language prompts, with line-by-line control over acting cues, pacing, dialect shifts, and backchanneling S¹. Gemini 3.8 Flash-Lite TTS is the volume play, built for cost-efficient dubbing, audio content creation, and expressive voice agents S¹.

My read: This is the first TTS release I've seen from Google that treats voice direction like screenwriting rather than parameter tuning. The 30-second cloning claim is the headline grabber, but I'm cautious: the blog post is a corporate announcement with no independent benchmarks, no pricing, and no latency figures for the Lite model. The consent and watermarking safeguards sound thorough on paper, but nobody outside Google has tested whether SynthID watermarks survive a dubbing pipeline that re-encodes audio. I'd watch for developer API access details, which the post does not provide.

Both models share a set of capabilities that go beyond reading text aloud. They support long-form generation across hours of continuous audio with what Google calls 'minimal speaker drift' S¹. They handle native two-speaker scene staging from a single script, keeping voices separated with conversational turn-taking S¹. And they accept stage directions or natural script cues to steer delivery line by line S¹.

The 30-second clone

Voice replication can recreate a consistent vocal profile from a 30-second audio sample, according to Google's announcement S¹. The company says the feature ships with built-in consent verification, SynthID watermarking, and C2PA credentials, the content provenance standard also used in photography authentication S¹. No independent party has verified how these safeguards perform in practice, and the announcement does not describe what happens when consent verification fails or is bypassed.

A voice remixing feature, which would let users fine-tune timbre, pitch, pace, and accent on existing library voices, is listed as 'coming soon' with no date S¹.

Where it fits

The TTS models join a Gemini Audio family that already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking S¹. Google has been expanding the Gemini line across modalities all year, as we saw when Gemini Omni 1.1 Flash added 10x more scene context in late August. The developer blog described the broader audio push on 15 September 2026, covering real-time voice applications built on 3.8 Live and 3.5 Transcribe P⁶. The model card for Gemini 3.8 Flash was published on 2 September 2026 P⁴, three weeks before this TTS announcement.

Google says the new voices will improve experiences in Gemini Notebook and Google Vids S¹, though it has not said when. A community wrapper already exists on GitHub: mehmetfiskindal's gemini_tts_wrapper, an MIT-licensed Dart package that calls the Generative Language API for one-shot audio output P⁵. Developer interest in Gemini APIs has been running hot, as it was when the Gemini CLI hit 106,000 stars on its free tier. The Hacker News discussion around the TTS announcement drew 58 points and 25 comments S³.

For a dubbing studio localising a series into 30 languages, the Flash-Lite model's promise of cost-efficient scale with fine-grained tone and pacing control is the angle worth testing, once API access and pricing appear. For a developer building a voice agent, the two-speaker staging and long-form stability matter more than the voice count, a capability that counts in domains where conversational nuance matters, as we saw when Google's AMIE medical AI matched doctors in video consults. Neither group can evaluate either claim today without API access, which Google has not detailed.

The Gemini 3.8 Flash model card, published 2 September 2026 P⁴, is the closest thing to a technical reference available, and a natural next checkpoint for developers waiting on API documentation.


Sources: S1 — Gemini 3.8 text-to-speech says hello · S2 — Gemini 3.8 text-to-speech says hello - blog.google · S3 — Gemini 3.8 text-to-speech says hello · P4 — Gemini 3.8 Flash - Model Card — Google DeepMind · P5 — README.md · P6 — New Gemini Audio models for developers · P7 — GitHub - google-deepmind/deepmind-research: This repository contains i


Written from 7 sourced items, 5 of them primary.