
23 Sep 2026
Google launches Gemini 3.8 Flash TTS and Flash-Lite TTS
Google introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS — expressive speech models rolling out in the Gemini API, Google AI Studio, and Google products, with generative voice design, a 2,000+ voice library, consent-checked 30-second voice replication, and SynthID watermarking on every clip.
SOFTWARE desk — Google just turned voice cloning and custom character speech into a mainstream API feature with consent checks and watermarks bolted on; the fight over who owns a voice just got an official product surface.
Flash TTS, as the post describes it, is for deep creative direction and character design. You can create new voices from natural-language prompts. You can direct a performance line by line: acting cues, pacing, dialect shifts, and backchanneling. Backchanneling, here, is the short sounds a listener makes while the other person is still talking, such as “mhm.” The post also names long-form generation, a long stretch of speech rather than one short line, and native two-speaker scene staging, one script that stages a conversation between two voices. Those lines are Google’s. This desk did not direct a scene.
Flash-Lite TTS, still the post, is the high-volume, cost-efficient model. Google says it is for dubbing, audio content, and expressive voice agents. A voice agent is software that talks with a person. Cost-efficient is Google’s description. This desk did not see a price.
The voice library, as stated: 2,000-plus production-ready voices, and support claimed for 100-plus languages and dialects. 2,000-plus means more than two thousand voices Google calls ready to use. 100-plus means more than a hundred languages and dialects. Those counts are Google’s. This desk did not count the library.
Voice replication, as stated: recreate a voice from about a 30-second sample, and only after a verbal consent recording from the voice owner that matches the reference speaker. The reference speaker is the person in the sample. The match is the check that the person saying yes is the person whose voice is being copied. Every Gemini Audio clip is watermarked with SynthID. SynthID is Google’s imperceptible AI watermark woven into the audio so tools can detect the clip was machine-made. The post also notes C2PA credentials for transparency. C2PA is a content-credentials standard, a signed note that can travel with a file. File both as Google’s. This desk did not run a detector and did not open a credential.
Availability, as the post states it. Flash TTS is rolling out for developers in the Gemini API and Google AI Studio, coming soon via API in Gemini Enterprise, and for everyone in Gemini Notebook. Flash-Lite TTS is rolling out for developers in the Gemini API and Google AI Studio, coming soon via API in Gemini Enterprise, and for everyone in Google Vids. Coming soon is the post’s tense for Gemini Enterprise. Do not read that as a date when every Google product turned the models on. This desk did not open each surface.
What Google cites on benches. Flash TTS is #1 on the Hume AI Voice Design Benchmark, at 71.4. Both models hold top ranks on Hume AI’s Overall Quality Index. In Voice Arena blind-preference tests, the post says they top several languages. A blind preference test means a listener picks a clip without being told which system made it. 71.4 is Google’s cited score. This desk did not rerun the benches and is not treating them as an independent audit.
Partners the post names as integrating the models: Figma, HeyGen, Linguana, Wondercraft, 99.co, and Ollang. Integrating is Google’s sentence. This desk did not check each product.
The limit, as stated. Voice replication through AI Studio is not available in Illinois, Texas, the EEA, the UK, Switzerland, and India. EEA means the European Economic Area. That list is the post’s. This desk did not test a blocked region.
Plain English for the rest of the card. TTS is typed words turned into spoken audio. SynthID is the hidden watermark in the audio. C2PA is the content-credentials note. A 30-second sample is the short recording the clone is built from, and only after the owner’s spoken yes matches that voice. 2,000-plus is the voice library. 100-plus is the language-and-dialect claim. 71.4 is the Voice Design score Google cites for Flash TTS. Coming soon is Gemini Enterprise. Notebook is the Flash TTS “for everyone” surface. Google Vids is the Flash-Lite “for everyone” surface.
PRIMARY here: Google’s 23 Sep 2026 blog, “Gemini 3.8 text-to-speech says hello,” page date Sep 23, 2026, no hour on the page this desk read — Tier A PRIMARY, the company’s own announcement. The two model names, the Flash TTS creative-direction list, the Flash-Lite high-volume line, the 2,000-plus voice library, the 100-plus languages and dialects claim, the about-30-second sample, the verbal consent recording that matches the reference speaker, SynthID on every Gemini Audio clip, the C2PA note, the rollout surfaces, the Hume Voice Design Benchmark #1 at 71.4, the top Overall Quality Index ranks, the Voice Arena blind-preference line, the six partner names, and the AI Studio voice-replication block list are the blog’s. NOT claimed: a price, that this desk generated a clip or reran a benchmark, that every language was counted, that each named partner has shipped, that voice replication works in the blocked regions, a consumer on-sale date beyond the surfaces the post names, a stock tip, or investment advice. Distinct from the already-filed deepmind-private-ai-memory, wisdomai-live-apps, zerodrift-anchor-3, and gemini-3-8-live.
RELATED
On 23 Sep 2026 Google published “Gemini 3.8 text-to-speech says hello,” introducing Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS. The record is Google’s blog. The page date is Sep 23, 2026. The page this desk read does not print an hour. TTS means text-to-speech: typed words become spoken audio. These are company claims. This desk did not generate a clip.













