
15 Sep 2026
Google launches Gemini 3.8 Live voice models with Extended Thinking
Google said Gemini 3.8 Live and 3.8 Live Extended Thinking are rolling out for developers, enterprises, and consumers — voice agents that keep talking while tools run in the background, with live visual grounding and 97-language support.
Same-day SOFTWARE desk primary that Google is shipping a new live voice-agent stack into the Gemini API plus consumer Search and Gemini surfaces. This is a discrete product launch — not a soft blog.
Gemini 3.8 Live, as Google tells it, is built for scale and cost efficiency. It combines conversational intelligence with fluid dialogue and visual grounding. Visual grounding here means the model can use what a camera or screen shows to answer. File that product picture as Google’s. This desk did not sit on a live session.
Gemini 3.8 Live Extended Thinking, still company, is built for high-complexity multi-step tasks. Google says it holds the #1 overall spot on Artificial Analysis’ Speech to Speech Quality Index at 82.6, and leads agentic task completion with 68.6% on τ-Voice and 35.1% on Sierra’s τ-Voice-banking benchmark. File those scores as Google’s company-reported benches. This desk did not rerun them and is not treating them as independent replications.
Capabilities named on the same company page: the models can execute tools and API calls in the background while the conversation continues; they automatically detect and switch among 97 supported languages mid-conversation; and they process visual inputs in near real-time. File that list as Google’s. This desk did not time a tool call or count the languages.
All audio generated by Google’s AI products is watermarked with SynthID, the company said. SynthID is Google’s invisible watermark woven into the audio so later detectors can tell the clip was AI-made. File that watermark as Google’s. This desk did not run a detector.
Rollout starting today, as Google prints it — do not upgrade this to a claim that every surface is generally available. Gemini 3.8 Live: developers via the Gemini API and Google AI Studio; enterprises in private preview in Gemini Enterprise, with Gemini Enterprise for Customer Experience listed as coming soon; consumers in Search Live. Gemini 3.8 Live Extended Thinking: the same developer path; the same enterprise private preview, with Gemini Enterprise for Customer Experience and Google Workspace business customers listed as coming soon; consumers in Gemini Live; Google AI Pro and Ultra subscribers in Workspace in Docs; all Google AI subscribers in Gmail and Keep. File those surfaces as Google’s printed windows. This desk is not inventing a general-availability date for enterprise or a Workspace rollout for every user.
The same-day developer companion cites Live API pricing at $0.005 per minute of audio input and $0.018 per minute of audio output. A footnote on that page says the figures are an estimate based on $3 per 1 million tokens for input and $12 per 1 million tokens for output. File those rates as Google’s estimate on the developer post. This desk is not inventing a different price card or a committed enterprise rate.
Integration partners named on the company page: Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents. Google also says it is partnering with companies like Salesforce, Genspark, and Lumeris who are excited about the models. File those names as Google’s. The Salesforce line is partnership enthusiasm on today’s voice-model post — not a new Salesforce deal. Salesforce and Google Cloud’s already-filed Dreamforce Hyperforce partnership stays salesforce-google-unify-agents.
Named voices on the company pages: Tom Ouyang and Malini Jaganathan on the model post; Alisa Fortin and Thor Schaeff on the developer companion. File the names and titles as Google’s. Longer product color stays in Sources.
Plain English for the rest of the card: speech-to-speech = AI that listens and talks as audio without a separate speech-to-text then text-to-speech chain. Visual grounding = using what the camera or screen shows to answer. SynthID = Google’s invisible watermark so AI audio can be detected later. Live API = Google’s real-time audio interface in the Gemini API. Private preview = a limited enterprise look, not general availability.
PRIMARY here: Google’s 15 Sep 2026 company blog — Tier A PRIMARY company source, the original record — plus the same-day developer companion. The Live and Extended Thinking introduce, the scale / cost-efficiency and high-complexity split, the company-reported Artificial Analysis / τ-Voice / τ-Voice-banking scores, background tool calls, 97-language switching, near real-time visual inputs, SynthID watermarking, the printed rollout surfaces including private preview, the Live API price estimate, the named integration and “excited” enterprise partners, and the Ouyang / Jaganathan / Fortin / Schaeff names are company-attributed. Benchmarks and price figures stay company-attributed. NOT claimed: enterprise general availability, a Workspace rollout for all users, independently replicated Speech to Speech or τ-Voice scores, a new Salesforce commercial deal beyond today’s enthusiasm line, that this desk tested the Live API, a stock tip, or investment advice. Distinct from the already-filed salesforce-google-unify-agents, google-cloud-singapore-engineering-center, google-org-ai-gov-challenge, meta-one-subscription, and gemini-3-8-flash-cyber.
RELATED
- Salesforce and Google Cloud unify agents and Hyperforce for one AI stack
- Google Cloud opens Singapore Engineering Center for enterprise cloud and AI
- Meta launches Meta One, a paid plan for more AI usage across its apps
- Google.org picks 15 teams for a $30M AI challenge aimed at government services
- Google DeepMind launches Gemini 3.8 Flash and Flash Cyber
On 15 Sep 2026, Google introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as near real-time voice and dialogue models. Near real-time here means the model listens and talks as audio, without a long wait for a separate transcript. The company PRIMARY is Google’s blog post “Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking,” dated September 15, 2026, by Tom Ouyang, Principal Engineer, and Malini Jaganathan, Member of Technical Staff, on behalf of the Gemini Audio Team. A same-day developer companion, “Build real-time voice applications with Gemini 3.8 Live and 3.5 Transcribe,” is signed by Alisa Fortin, Product Manager, Google DeepMind, and Thor Schaeff, Member of the Technical Staff (DevX), Google DeepMind. Those two Google pages are the filing event. These are company claims. This desk did not run the models.