datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot.
Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.image2struct-latex-v1
Image2Struct - Latex
Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo
License: Apache License Version 2.0, January 2004
Dataset description
Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images.
This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt:
Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.Latex-VLMLaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR
LaTeX_OCR
LaTeX OCR 的数据仓库
本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。
如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++
后续追加新的数据也会放在这个仓库 ~~
原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR.
数据集
本仓库有 5 个数据集
small 是小数据集,样本数 110 条,用于测试
full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。
synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/linxy/LaTeX_OCR.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/OleehyO/latex-formulas.LaTeX_OCRpubtables-latexpubtables-latexwikipedia-latex-formulas-319k
Wikipedia LaTeX Formulas 319k Dataset
A curated collection of mathematical formulas from English Wikipedia, providing LaTeX source code paired with high-quality rendered images for training OCR and formula recognition models.
Dataset Description
Dataset Summary
This dataset represents a complete extraction of ~319k LaTeX mathematical formulas from English Wikipedia articles (October 2025 snapshot), filtered by visual complexity (score > 8) and renderability… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/wikipedia-latex-formulas-319k.latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
pubmed-tables-latex-768pxLaTeX_OCR
LaTeX OCR 的数据仓库
本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。
如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++
后续追加新的数据也会放在这个仓库 ~~
原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR.
数据集
本仓库有 5 个数据集
small 是小数据集,样本数 110 条,用于测试
full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。
synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/ymfl/LaTeX_OCR.Latex-Math-HME100KHME100K – Hugging Face Version
Dataset Name: HME100K (Hugging Face Version)
Original Source: https://www.kaggle.com/datasets/cutedeadu/hme100k
Original Author (Kaggle): cutedeadu
Overview
HME100K is a dataset of handwritten mathematical expressions.
Each sample consists of:
An image containing a handwritten math expression
A corresponding LaTeX-style text annotation
This dataset is useful for:
Handwritten Mathematical Expression Recognition (HMER)
OCR research for… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/Latex-Math-HME100K.latex-ocr-dataset
Synthetic LaTeX OCR Dataset
Overview
This dataset contains synthetically generated LaTeX formula images designed to augment training data for LaTeX OCR models. The dataset applies style enrichment techniques to existing real-world LaTeX datasets, creating diverse visual representations of mathematical formulas through PDF-rendering and font styling.
Dataset Structure
synth/
├── plain/
│ ├── train/
│ │ ├── images/
│ │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/PadishahIIIXXX/latex-ocr-dataset.Handwritten-Mathematical-Expression-Convert-LaTeXimg-latex-ko
IMG2LaTeX
1. LaTeX-ko
1. 1. 설명
IM2LATEX-230k 을 LLaVA 형식에 맞게 수정한 데이터셋 입니다.
사용법은 LLaVA, KoLLaVA 참고 하시기 바랍니다.
ZhEn-latex-ocrlatextract-arxiv-math
latextract-arxiv-math
A multi-domain corpus of (formula image, LaTeX source) pairs harvested from
arXiv source tarballs. Each formula is rendered in isolation with tectonic
into a tightly-cropped PNG; the ground-truth LaTeX is the body of the original
display-math environment, normalized to remove \label{...}.
Provenance
Source: arXiv e-print/<id> tarballs (publicly downloadable)
Renderer: tectonic (LaTeX engine) + MuPDF (PDF → PNG @ 200 DPI)
Filtering: papers using… See the full description on the dataset page: https://huggingface.co/datasets/tbrrss/latextract-arxiv-math.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/yayossd/latex-formulas.LaTeX-OCR-dataset
LaTeX-OCR Dataset
Summary
This dataset was created to train LaTeX-OCR, a model for recognizing LaTeX code from images of mathematical formulas. Each sample consists of a synthetically rendered formula image and its corresponding LaTeX formula. The images were generated from scratch using xelatex with multiple fonts, offering more typographic variety than many other datasets that use a single font (typically Computer Modern).
Data Sources
Formulas were… See the full description on the dataset page: https://huggingface.co/datasets/lukbl/LaTeX-OCR-dataset.LaTeX_OCR
LaTeX OCR 的数据仓库
本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。
如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++
后续追加新的数据也会放在这个仓库 ~~
原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR.
数据集
本仓库有 5 个数据集
small 是小数据集,样本数 110 条,用于测试
full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。
synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/EliMC/LaTeX_OCR.LaTeX_Image_Pairs
LaTeX Image Pairs Dataset
This dataset comprises a unique collection of LaTeX expressions paired with their corresponding images. The LaTeX expressions were meticulously scraped from a variety of open-source textbooks, ensuring a diverse and comprehensive dataset. Sample references from these textbooks will be provided to illustrate the sources of these expressions.
In addition to the raw LaTeX expressions, this dataset includes images of the rendered expressions. Each LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/henryholloway/LaTeX_Image_Pairs.latexOleehyO-latex-formulaslatex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Evanstarcraft2/latex-formulas.LaTeXImage_testLaTeX_OCRRR
