Microsoft ships MAI-Transcribe-2-Streaming plus MAI-Voice-2.1 for sub-second voice agents
Microsoft AI said Thursday it launched MAI-Transcribe-2-Streaming — real-time speech-to-text in 60 languages with partials in just over 100 ms — alongside MAI-Voice-2.1 and MAI-Voice-2.1-Flash, so developers can build a full voice-agent loop on Microsoft Foundry and Azure AI Speech.
Voice agents only feel like a phone call when hearing and speaking finish inside about a second. Microsoft just closed its own stack's missing piece — streaming STT — so an Azure shop can keep transcription, reasoning, and speech inside one compliance boundary instead of bolting on a third-party ear.
On Thursday, 1 October 2026, Microsoft AI announced its first streaming transcription model, MAI-Transcribe-2-Streaming, and two new speech models, MAI-Voice-2.1 and MAI-Voice-2.1-Flash. The post is titled “Our first streaming transcription model debuts at no. 1 on Artificial Analysis.” The line under the title says the MAI-Transcribe and MAI-Voice models are accurate, fast, low cost, and chart-topping audio understanding and generation for building conversational voice agents. The page dates the post October 1, 2026. It does not print an hour. Speech-to-text means the model writes down what it hears. Text-to-speech means it reads writing out loud. Streaming, here, means the transcript starts while the person is still talking. Those lines are Microsoft AI’s.
What the streaming model does. MAI-Transcribe-2-Streaming delivers low-latency, real-time transcripts in 60 languages, and it supports automatic, continuous language detection. Latency is the wait. Continuous language detection means the model can tell which language is being spoken as the audio keeps coming, rather than being told the language once at the start. Rather than waiting for someone to finish, it produces its first hypotheses, which the page calls partials, in just over 100 milliseconds of receiving audio. A millisecond is a thousandth of a second, so just over 100 milliseconds is a little more than a tenth of a second. A partial is that first guess at the words. The model then revises the guess as more of the sentence arrives, and it commits a stable transcript right away. Those lines are Microsoft AI’s.
What those early words are for. Microsoft says the partials let a voice app act on speech before the speaker finishes. A voice agent can start reasoning, or start calling a tool, in the middle of a sentence. A live transcript can appear while people are still talking. For real-time dictation or subtitling, Microsoft says its internal evaluations show that words appear in the transcript twice as fast as with its closest competitor. Twice as fast is Microsoft’s own test. The page does not name the competitor, and it does not print an outside check of the timing.
Where Microsoft says the model ranks. The post says it ranks no. 1 for accuracy on both the final transcript and the partial transcript on Artificial Analysis. Artificial Analysis is the scoreboard Microsoft is citing. On the accuracy-versus-latency chart, Microsoft says it sits on the Pareto frontier. The post says higher accuracy does not have to come with a hefty wait. A chart on the page is a word-error-rate leaderboard for streaming speech-to-text. Word error rate is how often the written words do not match what was said. Lower is better. The bars are the final transcript, ordered from the lowest error, and an overlay marks the same model’s first partial. MAI-Transcribe-2-Streaming is the first bar. The chart prints 2.5 percent for its final transcript and 2.8 percent for its first partial. The note says the chart shows the 28 streaming models at or below 10 percent first-partial word error rate. The source line is the Artificial Analysis Speech to Text (Streaming) leaderboard, dated 28 September 2026. Those figures and that date are on Microsoft’s chart. They are Microsoft’s citation of Artificial Analysis.
The chart next to that one plots final-transcription accuracy against the time until the final transcript. Time on that axis is in seconds. The caption says lower is better on both axes. It marks a Pareto frontier and a region it calls the most attractive quadrant. The same source line dates the Artificial Analysis streaming leaderboard 28 September 2026. The “just over 100 milliseconds” line in the post is the wait until the first partial. It is not a number printed on the axis of the time-to-final chart.
The introductory price. MAI-Transcribe-2-Streaming is available at $0.54 per hour of audio through the end of the year. On a post dated October 1, 2026, that year is 2026. Fifty-four cents is a little over half a dollar for an hour of audio. The page calls the price introductory. It does not print the rate that would follow the end of the year.
MAI-Voice-2.1 is the speech model. Microsoft calls it the strongest multilingual text-to-speech model it has offered. The model supports 23 languages and 26 locales. A locale is a language as spoken in a particular place, so the count of places can run ahead of the count of languages. One voice can use all of those languages with what Microsoft calls a truly native accent. The example on the page is to ask it to speak English, then Mandarin, then German. The speaker stays the same, and it picks up the local way of talking, rather than carrying one accent into every language. Microsoft says a brand can keep a single voice everywhere. A tutoring app can switch languages in the middle of a lesson without swapping teachers. A multilingual assistant can answer in the language it was addressed in and still sound like the same voice. The price is $22 per 1 million characters. A character is a letter or mark in the text the model reads aloud. Those lines are Microsoft AI’s. The page does not print the full list of languages.
MAI-Voice-2.1-Flash is the faster variant, built for a high volume of requests where the wait matters. It supports the same languages and the same cross-language speakers as MAI-Voice-2.1. Microsoft says it can generate 45 seconds of audio, with an end-to-end latency of 150 milliseconds. End-to-end, in that sentence, is the wait from the request until the speech comes back. One hundred fifty milliseconds is 0.15 seconds. Microsoft says the model delivers 55 percent faster inference. Inference is the work of producing the audio. The page states the 55 percent in its own clause and does not name, there, the model it is faster than. The next clause says Flash is about 60 percent cheaper than comparable models, at what Microsoft calls best-in-class pricing of $15 per 1 million characters. Those comparisons are Microsoft’s. The page does not name the models it is comparing. It says the mix of wait, quality, and cost makes Flash a natural partner for MAI-Transcribe-2-Streaming when someone is building a low-latency voice agent.
Cloning, and the limit Microsoft states. Both voice models can clone a voice across every supported language from a few seconds of reference audio, so a customer can use a brand’s voice. A clone, here, is a new clip that is meant to sound like that short sample. Microsoft says the models have built-in consent guardrails that prevent misuse. A guardrail, in that sentence, is a limit inside the product. The page does not describe the check, and it does not say what a customer must send to show consent. Those lines are Microsoft AI’s.
How Microsoft describes the loop. A voice agent, the post says, has to hear, understand, decide, and speak, inside the window where a person still experiences the exchange as a conversation. Each piece either buys time in that window or spends it. Pairing MAI-Transcribe-2-Streaming with MAI-Voice-2.1-Flash, Microsoft says, buys time on both the hearing end and the speaking end. The time saved is room for the agent to reason, use tools, and check its answer, while the conversation still moves at a human pace. Those lines are Microsoft AI’s.
What Microsoft says a developer can build now. The examples on the page are a customer-service agent that transcribes a request as it is spoken, starts acting before the caller finishes, and answers in spoken words; a multilingual assistant that detects the spoken language and replies in any of the 23 MAI-Voice languages, in the same voice and with a native accent; and learning or media that uses different speakers for tutoring, role-play, simulations, narration, and conversation. To show the models in one live agent, Microsoft built Chatter, a demo in the MAI Playground. Those lines are Microsoft AI’s. They describe the use Microsoft has in mind. The page does not name a customer, and it does not print a measured call time.
Where the models are offered. Developers can use MAI-Voice-2.1 and MAI-Voice-2.1-Flash through OpenRouter. All three models are offered through Microsoft Foundry, the MAI Playground, Vercel, and Azure Voice Live. Foundry is Microsoft’s place for a developer to call a model. The page lists OpenRouter for the two voice models. LiveKit is on the same list, marked coming soon. The section is titled “Start building now,” and the post says developers can build these today. Those lines are Microsoft AI’s.
A listening comparison on the same page. A chart says that in a test of 4,000 listeners, which the chart calls a Turing test, 50.3 percent rated MAI-Voice equally or more human-like than recordings of a person, and 49.7 percent rated the human recordings more human-like. The chart says the results combine MAI-Voice-2.1 and MAI-Voice-2.1-Flash. As this chart uses the phrase, a Turing test is a listening comparison: people judge the generated voice against a recording of a person. Those percentages are the chart’s. The page does not say how the clips were chosen, and it does not say who the 4,000 listeners were.
The picture is the header on that post. The four-square Microsoft logo sits on a dark blue field. The headline reads “Our first streaming transcription model debuts at no. 1 on Artificial Analysis.” The line under it says the MAI-Transcribe and MAI-Voice models are accurate, fast, low cost, and chart-topping audio understanding and generation for conversational voice agents. The frame is the launch graphic. It is the header image. It is not a photograph of a person, a studio, or a phone.
In plain terms, Microsoft AI said on Thursday that developers can now hear and speak inside one Microsoft stack. MAI-Transcribe-2-Streaming writes down speech in 60 languages and offers a first guess in a little over a tenth of a second. MAI-Voice-2.1 reads text aloud in 23 languages and 26 locales and keeps one speaker across those languages, at $22 per million characters. Flash is the faster voice, at $15 per million characters, and Microsoft says it can return 45 seconds of audio with a 0.15-second wait and 55 percent faster inference. The transcription price is $0.54 an hour of audio through the end of 2026. Microsoft cites an Artificial Analysis streaming leaderboard dated 28 September 2026 for a no. 1 accuracy rank on both the finished transcript and the partial, and for a place on the accuracy-versus-wait frontier. The chart on the page puts this model’s word error rate at 2.5 percent final and 2.8 percent on the first partial. The voices can be cloned from a few seconds of audio, with consent limits Microsoft says are built in. The two voice models are on OpenRouter. All three are on Foundry, the MAI Playground, Vercel, and Azure Voice Live. LiveKit is listed as coming soon.
