Microsoft transcribes speech before the speaker finishes
Microsoft launched MAI-Transcribe-2-Streaming alongside MAI-Voice-2.1 and its Flash variant. The transcription model supplies provisional text while someone speaks, then revises it as more context arrives. Microsoft describes automatic language detection and speech generation that preserves a speaker across languages. All three models are offered through Microsoft Foundry; Chatter demonstrates how the components work together in a conversational agent.
Artificial Intelligence··Morning
Microsoft adds live transcription to its speech lineup
Microsoft launched MAI-Transcribe-2-Streaming on 1 October with MAI-Voice-2.1 and MAI-Voice-2.1-Flash, its two speech-generation models. All three are available through Microsoft Foundry, the company’s platform for building AI applications. The release combines a system that turns incoming speech into text with models that produce spoken responses, providing separate components for a conversational agent.[1], [2]
Provisional text changes as words arrive
The transcription model produces provisional text while someone is still speaking. As later words add context, the system revises those hypotheses and commits a stable transcript. Microsoft describes continuous automatic detection of the spoken language. Its examples include live dictation, subtitles and agents that can begin working from a request before the caller finishes it.[1]
Voice models keep a speaker across languages
Microsoft describes voice models that preserve the same speaker identity across supported languages. Its examples include a tutor switching languages and an assistant answering in the language used by a caller. Both models support voice cloning from a short reference clip; Microsoft says built-in consent safeguards address misuse. Chatter, a demonstration in the MAI Playground, combines the listening, reasoning and speech components in a live conversation.[1]