AI Frontier
← 浏览此厂商的报告

Microsoft / Phi / MAI / Model Card

MAI-Voice-2.1 Model Card

MAI-Voice-2.1 / Flash · 日期待确认

概要原文

保留原文 · 保留原始语言

Overview · 页码 1

MAI-Voice-2.1 is a prompted text-to-speech (TTS) model built in-house by the Microsoft AI team, generating high-fidelity, natural, and expressive speech across 23 languages and 26 locales. It captures human-like intonation, rhythm, and emotional nuance, delivers emotional flexibility with turn-level control over tone and delivery, and can render a single voice identity consistently across every supported language, making it ideal for audiobooks, content creation, voice-over, media, and other scenarios where fidelity and expressiveness matter most. Voice can be configured using: • Curated voice library (licensed voices designed to work straight out of the box) • Voice Cloning via short audio clips (5-60 seconds), and the model matches the voice instantly

核心图片

点击放大查看,下载原图获取完整细节。

此报告暂无选取的核心配图,可直接阅读原始 PDF。

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。