datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/scholarweave/arxiv-latex.latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot.
Handwritten-Latex-Datasets
Dataset
This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets.
Dataset source
Collected in various junior high schools and high schools, handwritten by students.
Usage
The label is stored at json folder and scanned hand-writted pictures are stored at pic folder.
Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.image2struct-latex-v1
Image2Struct - Latex
Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo
License: Apache License Version 2.0, January 2004
Dataset description
Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images.
This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt:
Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.Latex-VLMLaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR
latex-data-pub
Hyperion: Scientific LaTeX Corpus (Public)
This dataset constitutes the Public Scientific Corpus for the Hyperion Project at the Synthetix Institute. It contains high-fidelity LaTeX source text extracted from diverse scientific repositories, optimized for topological knowledge discovery and relational mapping.
Dataset Details
Total Documents: ~1,318,468
Average Document Length: Variable (approx. 32KB - 256KB)
Primary Domain: Mathematics, Physics, and Chemistry.
Goal:… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data-pub.LaTeX_OCR
LaTeX OCR 的数据仓库
本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。
如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++
后续追加新的数据也会放在这个仓库 ~~
原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR.
数据集
本仓库有 5 个数据集
small 是小数据集,样本数 110 条,用于测试
full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。
synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/linxy/LaTeX_OCR.latex-data
Hyperion: Scientific LaTeX Corpus
This dataset constitutes the Internal Scientific Corpus for the Hyperion Project of Synthetix Institute. It contains LaTeX source documents used for high-fidelity relational extraction and internal benchmarking of the Epistemic Manifold.
Dataset Details
Total Documents: ~1,192,727 (Internal Base)
Status: Private
Access: Restricted to Synthetix Institute authorized personnel.
Primary Use: Training the private sector of the Epistemic… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data.arxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/naveenmarthala/arxiv-latex.exams-basic-and-quantum-cryptography-and-security-latex
Open Problem Exams: Cryptography and Security (LaTeX)
A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions.
Dataset Overview
Institution
Files
Topics
Questions
Caltech & TU Delft
8
38
145
EPFL
6
19
86
ETH Zurich
1
14
37
MIT
3
33
79
Total
18
104
347
Difficulty Distribution
Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/OleehyO/latex-formulas.LaTeX_OCRarXiv_latex
TeX data from arXiv
Using https://github.com/KyuDan1/TeX2Image code.
We have categories Math, Physics, Statistics, ComputerScience.
Domain
Size
Mathematics
4.22M
Computer Science
2.76M
Statistics
0.89M
Physics
0.78M
Total (unique)
7.17M
pubtables-latexpubtables-latexneurips-2025-arxiv-latex-sources
NeurIPS 2025 arXiv LaTeX Source Files
This dataset contains file-level Parquet rows built from extracted raw arXiv
source packages for papers mapped from the NeurIPS 2025 proceedings to arXiv
records.
Each row is one file from one arXiv source package. Use arxiv_id to group files
back into papers.
Columns
arxiv_id: arXiv identifier for the source package.
title: paper title from the mapping CSV.
source_url: arXiv e-print source URL.
paper_index, paper_status… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/neurips-2025-arxiv-latex-sources.nav2tex-decoder-latex-pretraintargets_fixed_filtered_latexarxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/thejagstudio/arxiv-latex.LaTeXt
Dataset Card for MathOCR
Dataset Details
Dataset Sources [optional]
Repository: [https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt]
I don't write papers or demos
Uses
Direct Use
Post training for mathematical visualization.
Out-of-Scope Use
[Its not meant, but technically works for diffusion models.]
[The only way this dataset can be misused is in]
A. [Incompatible models, like TTS]
B. [Copying the repo without… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt.candidates_fixed_filtered_latexarxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/justatomic/arxiv-latex.wikipedia-latex-formulas-319k
Wikipedia LaTeX Formulas 319k Dataset
A curated collection of mathematical formulas from English Wikipedia, providing LaTeX source code paired with high-quality rendered images for training OCR and formula recognition models.
Dataset Description
Dataset Summary
This dataset represents a complete extraction of ~319k LaTeX mathematical formulas from English Wikipedia articles (October 2025 snapshot), filtered by visual complexity (score > 8) and renderability… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/wikipedia-latex-formulas-319k.latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
pubmed-tables-latex-768pxLaTeX_OCR
LaTeX OCR 的数据仓库
本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。
如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++
后续追加新的数据也会放在这个仓库 ~~
原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR.
数据集
本仓库有 5 个数据集
small 是小数据集,样本数 110 条,用于测试
full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。
synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/ymfl/LaTeX_OCR.Latex-Math-HME100KHME100K – Hugging Face Version
Dataset Name: HME100K (Hugging Face Version)
Original Source: https://www.kaggle.com/datasets/cutedeadu/hme100k
Original Author (Kaggle): cutedeadu
Overview
HME100K is a dataset of handwritten mathematical expressions.
Each sample consists of:
An image containing a handwritten math expression
A corresponding LaTeX-style text annotation
This dataset is useful for:
Handwritten Mathematical Expression Recognition (HMER)
OCR research for… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/Latex-Math-HME100K.latex-ocr-dataset
Synthetic LaTeX OCR Dataset
Overview
This dataset contains synthetically generated LaTeX formula images designed to augment training data for LaTeX OCR models. The dataset applies style enrichment techniques to existing real-world LaTeX datasets, creating diverse visual representations of mathematical formulas through PDF-rendering and font styling.
Dataset Structure
synth/
├── plain/
│ ├── train/
│ │ ├── images/
│ │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/PadishahIIIXXX/latex-ocr-dataset.
