AI Frontier
← 浏览此厂商的报告

智谱 / Z.ai / Technical Report

GLM-OCR Technical Report

GLM-OCR · 2026-03-11

概要原文

保留原文 · 保留原始语言

ABSTRACT · 页码 1

GLM-OCR is an efficient 0.9B-parameter compact multimodal model designed for real-world document understanding. It combines a 0.4B-parameter CogViT visual encoder with a 0.5B-parameter GLM language decoder, achieving a strong balance between computational efficiency and recognition performance. To ad- dress the inefficiency of standard autoregressive decoding in deterministic OCR tasks, GLM-OCR introduces a Multi-Token Prediction (MTP) mechanism that predicts multiple tokens per step, significantly improving decoding throughput while keeping memory overhead low through shared parameters. At the system level, a two-stage pipeline is adopted: PP-DocLayout-V3 first performs layout analysis, followed by parallel region-level recognition. Extensive evaluations on public benchmarks and industrial scenarios show that GLM-OCR achieves com- petitive or state-of-the-art performance in document parsing, text and formula transcription, table structure recovery, and key information extraction. Its com- pact architecture and structured generation make it suitable for both resource- constrained edge deployment and large-scale production systems.

核心图片

点击放大查看,下载原图获取完整细节。

图 1 · 实验结果页码 1
Figure 1: Performance of GLM-OCR on OmniDocBench v1.5.
图 2 · 模型架构页码 4
Figure 2: Overall architecture and workflow of the GLM-OCR framework. This system supports two primary tasks: Document Parsing (Task 1), which combines layout detection and region cropping to produce structured Markdown and JSON outputs; and Key Information Extraction (Task 2), which directly extracts structured JSON data based on input visual prompts.

点击图片切换缩放,按 Esc 关闭。每张图下方可下载高清文件。