News from the IT community on October 2nd: Microsoft announced yesterday (October 1st) the launch of its first real-time speech-to-text transcription model, MAI – Transcribe –2– Streaming, which can continuously generate text while speech is being delivered, covering 60 languages and supporting automatic language detection.
In terms of fees, the MAI – Transcribe –2– Streaming model is currently in a promotional phase, with a price of $0.54 per hour (Note from IT: The current exchange rate is approximately 3.6 RMB), which is about $9 per 1,000 minutes (approximately 60.5 RMB). Developers can experience or integrate this service through channels such as Microsoft Foundry, MAI Playground, OpenRouter, etc.
IT Note: Microsoft has previously launched a non-streaming MAI – Transcribe –2, which was priced at $0.10 per hour during the public preview phase of Azure Speech (current exchange rate is approximately 0.67 RMB). The streaming “Streaming” returns relevant results in real-time, making it suitable for real-time conversations and subtitles; whereas the non-streaming version returns the results all at once after processing, making it suitable for tasks such as recording file transcriptions, meeting minutes, and clinical document transcriptions.
In terms of latency, MAI-Transcribe-2-Streaming will generate the initial “partial text” within about 100 milliseconds after receiving the audio, then continue to refine it based on more context, and quickly submit a stable result after the statement is completed.
Microsoft claims that in its real-time evaluations, text can be displayed as fast as about 320 milliseconds after speech appears; in the tests with Artificial Analysis, the model's final transcription delay was 0.13 seconds, and the final word error rate was 2.50%, ranking it first among 28 streaming speech-to-text models.

In terms of practical experience, voice applications do not need to wait for the user to finish speaking a sentence before they can understand the intent, start reasoning, or call upon tools. For example, customer service agents can begin identifying issues before the caller has finished speaking, and real-time subtitle systems can also provide a more "what you say is what you see" experience.
In terms of functionality, MAI-Transcribe-2-Streaming supports 60 languages and has the ability to automatically and continuously detect the language in use. Users do not need to manually specify the language before a conversation begins; the system can identify language changes based on the voice content.
In terms of scoring, in the streaming speech-to-text ranking published by Artificial Analysis on September 28th, MAI-Transcribe-2-Streaming had a final word error rate of 2.50%, which is lower than Grok Voice Transcribe's 2.0%, Streaming's 2.73%, and ElevenLabs Scribe as well as v2 and Realtime's 3.59%. Microsoft also emphasized that this model ranked first in both the final transcription and the initial part of the transcription indicators.

Related Reading:
Microsoft's Most Powerful Transcription Model: MAI-Transcribe-2 Debuts with an Average Word Error Rate of 5.2% for 60 Languages












