AI Frontier
← 浏览此厂商的报告

Microsoft / Phi / MAI / Model Card

MAI-Voice-2.1-Flash Model Card

MAI-Voice-2.1 / Flash · 日期待确认

概要原文

保留原文 · 保留原始语言

Overview · 页码 1

MAI-Voice-2.1-Flash is a prompted text-to-speech (TTS) model built in-house by the Microsoft AI team for fast, low-latency generation. It delivers high-fidelity, natural, and expressive speech across 23 languages while being optimized for real-time responsiveness, preserving human-like intonation, rhythm, and emotional nuance. Developers can control tone, emotion, and delivery through SSML, making it ideal for voice agents, call center agents, assistants, and other interactive scenarios where speed is critical. Voice can be configured using: • Curated voice library (licensed voices designed to work straight out of the box) • Voice Cloning via short audio clips (5-60 seconds), and the model matches the voice instantly

核心图片

点击放大查看,下载原图获取完整细节。

此报告暂无选取的核心配图,可直接阅读原始 PDF。

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。