AI Frontier
← Browse this publisher

StepFun / Technical Report

StepAudio 3 Gen Technical Report

StepAudio 3 TTS · 2026-09-11

Source summary

Original wording · Original language

Abstract · Page 1

We introduce StepAudio 3 Gen, a general-purpose audio generation model that supports zero-shot text-to-speech (TTS), voice design, vocal generation, sound ef- fects, music, vibe speech, and mixtures of multiple audio types within a unified framework. At its core, StepAudio 3 Gen is a discrete autoregressive generator that models audio directly over residual vector quantization (RVQ) tokens, departing from the diffusion Transformer-based continuous generation paradigm prevalent in recent general audio models. Its StepAudio Tokenizer represents general audio at 12.5 Hz in a shared 16 × 2048 residual code space, jointly quantizing semantic and waveform-level acoustic features so that each code layer preserves both types of information. For generation, the backbone predicts the first codebook along the time axis using autoregressive modeling, while a lightweight causal Transformer completes the remaining fifteen codebooks along the codebook axis. Our study further identifies three key design principles: (1) interference-aware progressive pretraining for acquiring audio capabilities while preserving the textual abilities of the large language model, (2) RVQ Adaptor for effectively incorporating multi- codebook acoustic representations, and (3) discrete autoregressive modeling over a shared representation across general audio domains. With progressive pretraining, multi-task instruction training, and supervised fine-tuning, StepAudio 3 Gen achieves state-of-the-art performance on both TTS and voice design, while retaining strong generation capabilities across speech, vocals, sound effects, and music. Audio sam- ples are available at https://stepaudiollm.github.io/step-audio-3-gen/.

Core figures

Enlarge to explore. Download the original for full detail.

Figure 1 · ArchitecturePage 4
Figure 1: Overview of the StepAudio 3 Gen architecture. At each audio frame, the LLM predicts c0, and the RVQ Code Predictor conditions on the LLM hidden state and c0 to predict c1, . . . , c15. All sixteen codes form the full RVQ frame used for waveform decoding. ⊕denotes element-wise addition: at audio positions, the RVQ Adaptor output is added to the corresponding token embedding (§2.2).
Figure 2 · ResultsPage 12
Figure 2: Human evaluation on the Chinese human-likeness test set.

Click the image to zoom. Press Esc to close. Full-resolution files are available below each figure.