StepFun / Technical Report
StepAudio 3 Realtime Technical Report
概要原文
保留原文 · 保留原始语言Abstract · 页码 1
Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Percep- tion captures rich acoustic cues to interpret user intent, while Seamless Duplex mod- els synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken de- livery. In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAu- dioChat. With Think-While-Speaking, it achieves dialogue and reasoning perfor- mance comparable to dedicated reasoning models while speaking in real time. Fur- thermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on τ -Voice.
核心图片
点击放大查看,下载原图获取完整细节。