AI Frontier
← 浏览此厂商的报告

StepFun / Technical Report

StepAudio 3 Realtime Technical Report

StepAudio 3 Realtime · 2026-09-12

概要原文

保留原文 · 保留原始语言

Abstract · 页码 1

Realtime spoken interaction demands deep reasoning, prompt responses, and fluid turn-taking. We present StepAudio 3 Realtime, an audio-language foundation model organized around a continuous listen-converse-think-act loop. Deep Percep- tion captures rich acoustic cues to interpret user intent, while Seamless Duplex mod- els synchronized audio streams to handle pauses, backchannels, and interruptions naturally. Crucially, we resolve the tension between deep deliberation and latency via Think-While-Speaking, executing private reasoning in parallel with spoken de- livery. In reasoning mode, StepAudio 3 reaches a 73.0 macro average on StepAu- dioChat. With Think-While-Speaking, it achieves dialogue and reasoning perfor- mance comparable to dedicated reasoning models while speaking in real time. Fur- thermore, an integrated Voice Agent handles asynchronous tool execution without disrupting the dialogue flow. StepAudio 3 Realtime achieves top-tier performance across key dimensions: an exceptional 90.6 on the MMSU benchmark, 98.9 Overall on the Artificial Analysis Full-Duplex Bench, and a 56.0% macro task-success rate on τ -Voice.

核心图片

点击放大查看,下载原图获取完整细节。

图 1 · 实验结果页码 1
Figure 1: Benchmark results for StepAudio 3 ASR Max and StepAudio 3 Realtime (blue), compared with baselines (gray). Lower is better for ASR error rates, and higher is better elsewhere. Dialogue averages eight dimensions, full duplex reports Overall, and τ-Voice averages three domains. See Tables 1 and 10.
图 3 · 模型架构页码 4
Figure 3: System architecture of StepAudio 3 Realtime. The LLM decoder receives audio representations through the audio encoder and adapter, together with a separate text input. User and model audio form the two full-duplex streams, with generator output returning to the model audio stream. Waveforms are schematic.

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。