Microsoft / Phi / MAI / Model Card
MAI-Voice-2.1-Flash Model Card
概要原文
保留原文 · 保留原始语言Overview · 页码 1
MAI-Voice-2.1-Flash is a prompted text-to-speech (TTS) model built in-house by the Microsoft AI team for fast, low-latency generation. It delivers high-fidelity, natural, and expressive speech across 23 languages while being optimized for real-time responsiveness, preserving human-like intonation, rhythm, and emotional nuance. Developers can control tone, emotion, and delivery through SSML, making it ideal for voice agents, call center agents, assistants, and other interactive scenarios where speed is critical. Voice can be configured using: • Curated voice library (licensed voices designed to work straight out of the box) • Voice Cloning via short audio clips (5-60 seconds), and the model matches the voice instantly
核心图片
点击放大查看,下载原图获取完整细节。
此报告暂无选取的核心配图,可直接阅读原始 PDF。