Technology media TestingCatalog reports that Microsoft is testing its first native real-time voice model, MAI Realtime. The model supports 16 languages and two voice styles, enabling simultaneous listening and speaking during conversations, and supports automatic language detection and mid-conversation language switching — this means users don't have to wait for the AI to finish speaking before they can speak, making the conversation experience increasingly close to human interaction.

MAI Realtime uses a bidirectional, full-duplex interaction method, breaking the traditional voice assistant "user speaks, then model answers" turn-taking mode. The system uses endpoint detection technology to identify the start, pause, and end of speech. When users interject during a conversation, the model can immediately adjust its output to achieve synchronized listening and responding. This marks a key step for the MAI voice model family at Microsoft, moving from one-way to two-way communication — previously released models like MAI-Voice-2 and MAI-Transcribe-1.5 could only complete one-way tasks such as speech synthesis or speech recognition.
Two voice styles are more natural, but will not sing for now
In terms of language coverage, MAI Realtime supports 16 languages including Chinese, English, Japanese, Korean, French, German, Arabic, etc. Users can actively specify the conversation language or enable the auto-detection function, allowing the model to automatically detect and switch languages during the conversation. In terms of voice style, the test version offers two voices, Victoria and Grant, which are said to be more natural than Copilot's current voice mode.
