Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Yazhen Xie

UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

Dec 24, 2025

Yongkun Du, Zhineng Chen, Yazhen Xie, Weikang Baiand Hao Feng, Wei Shi, Yuchen Su, Can Huang, Yu-Gang Jiang

Figure 1 for UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

Figure 2 for UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

Figure 3 for UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

Figure 4 for UniRec-0.1B: Unified Text and Formula Recognition with 0.1B Parameters

Abstract:Text and formulas constitute the core informational components of many documents. Accurately and efficiently recognizing both is crucial for developing robust and generalizable document parsing systems. Recently, vision-language models (VLMs) have achieved impressive unified recognition of text and formulas. However, they are large-sized and computationally demanding, restricting their usage in many applications. In this paper, we propose UniRec-0.1B, a unified recognition model with only 0.1B parameters. It is capable of performing text and formula recognition at multiple levels, including characters, words, lines, paragraphs, and documents. To implement this task, we first establish UniRec40M, a large-scale dataset comprises 40 million text, formula and their mix samples, enabling the training of a powerful yet lightweight model. Secondly, we identify two challenges when building such a lightweight but unified expert model. They are: structural variability across hierarchies and semantic entanglement between textual and formulaic content. To tackle these, we introduce a hierarchical supervision training that explicitly guides structural comprehension, and a semantic-decoupled tokenizer that separates text and formula representations. Finally, we develop a comprehensive evaluation benchmark covering Chinese and English documents from multiple domains and with multiple levels. Experimental results on this and public benchmarks demonstrate that UniRec-0.1B outperforms both general-purpose VLMs and leading document parsing expert models, while achieving a 2-9$\times$ speedup, validating its effectiveness and efficiency. Codebase and Dataset: https://github.com/Topdu/OpenOCR.

Via

Access Paper or Ask Questions

Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

Dec 14, 2025

Weikang Bai, Yongkun Du, Yuchen Su, Yazhen Xie, Zhineng Chen

Figure 1 for Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

Figure 2 for Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

Figure 3 for Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

Figure 4 for Complex Mathematical Expression Recognition: Benchmark, Large-Scale Dataset and Strong Baseline

Abstract:Mathematical Expression Recognition (MER) has made significant progress in recognizing simple expressions, but the robust recognition of complex mathematical expressions with many tokens and multiple lines remains a formidable challenge. In this paper, we first introduce CMER-Bench, a carefully constructed benchmark that categorizes expressions into three difficulty levels: easy, moderate, and complex. Leveraging CMER-Bench, we conduct a comprehensive evaluation of existing MER models and general-purpose multimodal large language models (MLLMs). The results reveal that while current methods perform well on easy and moderate expressions, their performance degrades significantly when handling complex mathematical expressions, mainly because existing public training datasets are primarily composed of simple samples. In response, we propose MER-17M and CMER-3M that are large-scale datasets emphasizing the recognition of complex mathematical expressions. The datasets provide rich and diverse samples to support the development of accurate and robust complex MER models. Furthermore, to address the challenges posed by the complicated spatial layout of complex expressions, we introduce a novel expression tokenizer, and a new representation called Structured Mathematical Language, which explicitly models the hierarchical and spatial structure of expressions beyond LaTeX format. Based on these, we propose a specialized model named CMERNet, built upon an encoder-decoder architecture and trained on CMER-3M. Experimental results show that CMERNet, with only 125 million parameters, significantly outperforms existing MER models and MLLMs on CMER-Bench.

Via

Access Paper or Ask Questions