Microsoft / Phi / MAI / Model Card
MAI-Voice-2.1 Model Card
Source summary
Original wording · Original languageOverview · Page 1
MAI-Voice-2.1 is a prompted text-to-speech (TTS) model built in-house by the Microsoft AI team, generating high-fidelity, natural, and expressive speech across 23 languages and 26 locales. It captures human-like intonation, rhythm, and emotional nuance, delivers emotional flexibility with turn-level control over tone and delivery, and can render a single voice identity consistently across every supported language, making it ideal for audiobooks, content creation, voice-over, media, and other scenarios where fidelity and expressiveness matter most. Voice can be configured using: • Curated voice library (licensed voices designed to work straight out of the box) • Voice Cloning via short audio clips (5-60 seconds), and the model matches the voice instantly
Core figures
Enlarge to explore. Download the original for full detail.
No core figure selected for this report. The original PDF remains available.