AI Frontier
← Browse this publisher

智谱 / Z.ai / Technical Report

GLM-TTS Technical Report

GLM-TTS · 2025-12-16

Source summary

Original wording · Original language

ABSTRACT · Page 1

This work proposes GLM-TTS, a production-level TTS system designed for ef- ficiency, controllability, and high-fidelity speech generation. GLM-TTS follows a two-stage architecture, consisting of a text-to-token autoregressive model and a token-to-waveform diffusion model. With only 100k hours of training data, GLM- TTS achieves state-of-the-art performance on multiple open-source benchmarks. To meet production requirements, GLM-TTS improves speech quality through an optimized speech tokenizer with fundamental frequency constraints and a GRPO- based multi-reward reinforcement learning framework that jointly optimizes pro- nunciation, speaker similarity, and expressive prosody. In parallel, the system en- ables efficient and controllable deployment via parameter-efficient LoRA-based voice customization and a hybrid phoneme–text input scheme that provides pre- cise pronunciation control. Our code is available at https://github.com/ zai-org/GLM-TTS. Real-time speech synthesis demos are provided via Z.ai (audio.z.ai), the Zhipu Qingyan app/web (chatglm.cn).

Core figures

Enlarge to explore. Download the original for full detail.

Figure 1 · ArchitecturePage 1
Figure 1: Overall Architecture of GLM-TTS
Figure 3 · ArchitecturePage 5
Figure 3: An overview of the GLM-TTS-GRPO framework.

Click the image to zoom. Press Esc to close. Full-resolution files are available below each figure.