
11 Sep 2026
Google, Meta, and Microsoft ship new speech-to-text models for agents
12 Sep 2026 filing: Google (26 Aug), Meta (1 Sep), and Microsoft (3 Sep) each published dedicated speech-to-text models aimed at voice agents and real-time workflows. Benchmarks stay company / Artificial Analysis claims.
Within about a week, Google, Meta, and Microsoft each published dedicated speech-to-text models aimed at voice agents and real-time workflows. DeepLearning.AI’s The Batch dated that cluster 11 Sep 2026. That recap is the filing event for the three-lab week; each lab’s own dated page is the ship.
26 Aug 2026: Google published Gemini 3.5 Transcribe — “our most precise speech-to-text model yet, designed for intelligent voice interactions” — via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform. Live streaming uses gemini-3.5-transcribe-live; pre-recorded files use gemini-3.5-transcribe. Google: 85+ languages; smart transcription (self-corrections, filler cleanup, auto-format); custom vocabulary; word-level timestamps on files. The blog: pre-recorded speaker attribution “for up to three speakers (support for 3+ speakers is experimental).” API docs also list file diarization “up to 8 speakers,” with 3+ still experimental, and say live streaming has no diarization. Google cites Artificial Analysis WER of about 4.0% streaming and 2.6% non-streaming. Those benches are Google’s citation, not a desk measurement.
1 Sep 2026: Meta Superintelligence Labs published Muse Voice Transcribe — “the first real-time audio perception model” from that lab. Meta: streaming ASR plus diarization (20+ speakers claimed) plus endpointing; adaptive delay (80 ms audio chunks; the model chooses when to emit text); 70+ languages trained / 25 “extensively verified”; long audio “exceeding one hour” with no required post-processing. The same page: one-click voice dictation in Meta AI and Muse Code, “hold ‘Fn’ to try.” Meta Model API docs price it at $0.18 per hour of audio processed (streaming and non-streaming the same). That is $3 per 1,000 minutes. Meta also claims first on Artificial Analysis streaming speech-to-text and on public diarization benches as of 1 Sep 2026 — Meta’s ranking.
3 Sep 2026: Microsoft AI published MAI-Transcribe-2. Microsoft claims it is “the fastest, most accurate and cheapest speech recognition model in the world” versus Gemini 3.5 Transcribe, GPT-Transcribe, Whisper V3-Large, and Scribe V2. Posted launch price: $0.10 per hour of audio “as a limited-time offer until the end of the year.” Microsoft: 60 languages; speaker diarization; word-level timestamps; keyword biasing; verbatim vs clean styles. Microsoft cites FLEURS 5.2% WER (first across 60 languages on that bench, it says) and an Artificial Analysis accuracy-latency Pareto lead / second on the WER board. Its page: Artificial Analysis evals show the model 10× faster than GPT-Transcribe and 5× faster than Gemini 3.5 Transcribe. The Batch: Microsoft states an hour of audio can be transcribed in about 10 seconds. Those speed and WER lines stay Microsoft’s.
Availability on the primaries: Google via Gemini API / AI Studio / Enterprise Agent Platform. Meta via Meta AI (Mac dictation), Muse Code, and Meta Model API. Microsoft’s news page: demo through Microsoft Foundry, MAI Playground, and Open Router. Microsoft Learn: MAI-Transcribe-2 is in public preview on Azure Speech / Foundry Tools — “without a service-level agreement, and is not recommended for production workloads.”
Three frontier platforms put dated STT product pages out within days of each other for agentic voice. File the ships and posted rates; leave WER/speed leaderboard claims as theirs.
Sources
- Google — Introducing Gemini 3.5 Transcribe
blog.google
- Gemini API docs — Gemini 3.5 Transcribe
ai.google.dev
- Meta AI Research — Introducing Muse Voice Transcribe
research.meta.ai
- Meta Model API — Speech to text
dev.meta.ai
- Microsoft AI — MAI-Transcribe-2
microsoft.ai
- Microsoft Learn — MAI-Transcribe-2 (Azure Speech)
learn.microsoft.com
- DeepLearning.AI The Batch — Transcription Battles Heat Up
deeplearning.ai



















