> ## Content Index
> Fetch the complete content index at: https://www.notatechguy.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Google DeepMind's Gemini 3.8 TTS clones a voice from 30 seconds
- URL: https://www.notatechguy.com/google-deepmind-s-gemini-3-8-tts-clones-a-voice-from-30-seconds/
- Published: 2026-09-23T17:34:14.000Z
- Updated: 2026-09-23T17:34:15.000Z
- Description: Google DeepMind's Gemini 3.8 Flash TTS and Flash-Lite TTS bring voice cloning from a 30-second sample to 100-plus languages.
- Author: Marcello Babbili
- Tags: Technology & AI, Google

Google DeepMind launched two Gemini 3.8 text-to-speech models on 23 September 2026 that can replicate a voice from 30 seconds of audio [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com). The announcement packs 2,000-plus production-ready voices across 100-plus languages and dialects, but ships with no pricing, no API access details, and no independent benchmarks [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com).

The pair splits along a familiar line. Gemini 3.8 Flash TTS targets creative work: building entirely new voices from natural-language prompts, with line-by-line control over acting cues, pacing, dialect shifts, and backchanneling [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com). Gemini 3.8 Flash-Lite TTS is the volume play, built for cost-efficient dubbing, audio content creation, and expressive voice agents [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com).

**My read:** This is the first TTS release I've seen from Google that treats voice direction like screenwriting rather than parameter tuning. The 30-second cloning claim is the headline grabber, but I'm cautious: the blog post is a corporate announcement with no independent benchmarks, no pricing, and no latency figures for the Lite model. The consent and watermarking safeguards sound thorough on paper, but nobody outside Google has tested whether SynthID watermarks survive a dubbing pipeline that re-encodes audio. I'd watch for developer API access details, which the post does not provide.

Both models share a set of capabilities that go beyond reading text aloud. They support long-form generation across hours of continuous audio with what Google calls 'minimal speaker drift' [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com). They handle native two-speaker scene staging from a single script, keeping voices separated with conversational turn-taking [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com). And they accept stage directions or natural script cues to steer delivery line by line [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com).

### The 30-second clone

Voice replication can recreate a consistent vocal profile from a 30-second audio sample, according to Google's announcement [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com). The company says the feature ships with built-in consent verification, SynthID watermarking, and C2PA credentials, the content provenance standard also used in photography authentication [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com). No independent party has verified how these safeguards perform in practice, and the announcement does not describe what happens when consent verification fails or is bypassed.

A voice remixing feature, which would let users fine-tune timbre, pitch, pace, and accent on existing library voices, is listed as 'coming soon' with no date [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com).

### Where it fits

The TTS models join a Gemini Audio family that already includes 3.5 Live Translate, 3.5 Transcribe, 3.8 Live, and 3.8 Live Extended Thinking [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com). Google has been expanding the Gemini line across modalities all year, as we saw when [Gemini Omni 1.1 Flash added 10x more scene context](https://www.notatechguy.com/google-gemini-omni-1-1-flash-10x-more-scene-context/) in late August. The developer blog described the broader audio push on 15 September 2026, covering real-time voice applications built on 3.8 Live and 3.5 Transcribe [P⁶](https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/?ref=notatechguy.com). The model card for Gemini 3.8 Flash was published on 2 September 2026 [P⁴](https://deepmind.google/models/model-cards/gemini-3-8-flash/?ref=notatechguy.com), three weeks before this TTS announcement.

Google says the new voices will improve experiences in Gemini Notebook and Google Vids [S¹](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com), though it has not said when. A community wrapper already exists on GitHub: mehmetfiskindal's gemini\_tts\_wrapper, an MIT-licensed Dart package that calls the Generative Language API for one-shot audio output [P⁵](https://github.com/mehmetfiskindal/gemini%5Ftts%5Fwrapper/blob/main/README.md?ref=notatechguy.com). Developer interest in Gemini APIs has been running hot, as it was when [the Gemini CLI hit 106,000 stars](https://www.notatechguy.com/google-gemini-cli-hits-106-000-stars-with-free-1-000-request-daily-tier/) on its free tier. The Hacker News discussion around the TTS announcement drew 58 points and 25 comments [S³](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/?ref=notatechguy.com).

For a dubbing studio localising a series into 30 languages, the Flash-Lite model's promise of cost-efficient scale with fine-grained tone and pacing control is the angle worth testing, once API access and pricing appear. For a developer building a voice agent, the two-speaker staging and long-form stability matter more than the voice count, a capability that counts in domains where conversational nuance matters, as we saw when [Google's AMIE medical AI matched doctors in video consults](https://www.notatechguy.com/google-amie-medical-ai-matches-doctors-in-video-consults/). Neither group can evaluate either claim today without API access, which Google has not detailed.

The Gemini 3.8 Flash model card, published 2 September 2026 [P⁴](https://deepmind.google/models/model-cards/gemini-3-8-flash/?ref=notatechguy.com), is the closest thing to a technical reference available, and a natural next checkpoint for developers waiting on API documentation.

---

*Sources: [S1 — Gemini 3.8 text-to-speech says hello](https://deepmind.google/blog/say-hello-to-gemini-38-text-to-speech/?ref=notatechguy.com) · [S2 — Gemini 3.8 text-to-speech says hello - blog.google](https://news.google.com/rss/articles/CBMinwFBVV95cUxPSWpTWDNZcWVYandkVlNIa1E5TWFsSG81NU5ST3Q1cGt1Z2xmUXdWRnEzb19zMjVfakoxdDNEVTlzdEJhUzFLanV2dlFaVTZ4a2RjRFMxTXlqUVpNaFFONVZzeGlHcnBwRTlWNlpMektVMlZ4RlJhVlN3X0k1NG0xSklibWpqLUVBUEJjSFVWOTZMWE1ZYllGMUVxbWlSV3c?oc=5&ref=notatechguy.com) · [S3 — Gemini 3.8 text-to-speech says hello](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-text-to-speech/?ref=notatechguy.com) · [P4 — Gemini 3.8 Flash - Model Card — Google DeepMind](https://deepmind.google/models/model-cards/gemini-3-8-flash/?ref=notatechguy.com) · [P5 — README.md](https://github.com/mehmetfiskindal/gemini%5Ftts%5Fwrapper/blob/main/README.md?ref=notatechguy.com) · [P6 — New Gemini Audio models for developers](https://blog.google/innovation-and-ai/technology/developers-tools/build-real-time-voice-applications-gemini-audio/?ref=notatechguy.com) · [P7 — GitHub - google-deepmind/deepmind-research: This repository contains i](https://github.com/google-deepmind/deepmind-research?ref=notatechguy.com)*

## Related reading

- [Google Gemini Omni 1.1 Flash: 10x more scene context](https://www.notatechguy.com/google-gemini-omni-1-1-flash-10x-more-scene-context/) — our technology desk, 2026-08-27
- [Google Gemini CLI hits 106,000 stars with free 1,000-request daily tier](https://www.notatechguy.com/google-gemini-cli-hits-106-000-stars-with-free-1-000-request-daily-tier/) — our technology desk, 2026-08-25
- [Google AMIE medical AI matches doctors in video consults](https://www.notatechguy.com/google-amie-medical-ai-matches-doctors-in-video-consults/) — our technology desk, 2026-08-11

---

*Written from 7 sourced items, 5 of them primary.*