datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/OleehyO/latex-formulas.doclaynet-pt-enriched-formulawikipedia-latex-formulas-319k
Wikipedia LaTeX Formulas 319k Dataset
A curated collection of mathematical formulas from English Wikipedia, providing LaTeX source code paired with high-quality rendered images for training OCR and formula recognition models.
Dataset Description
Dataset Summary
This dataset represents a complete extraction of ~319k LaTeX mathematical formulas from English Wikipedia articles (October 2025 snapshot), filtered by visual complexity (score > 8) and renderability… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/wikipedia-latex-formulas-319k.latex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/yayossd/latex-formulas.OleehyO-latex-formulaslatex-formulas
𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️
📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link]
📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios.
For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Evanstarcraft2/latex-formulas.ocr-math-formula-handwrittenlatex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository.
IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set
ds4sd-synth-formula-net-smallocr-output_paddleocr_formula
Document Processing using PaddleOCR-VL (FORMULA mode)
This dataset contains FORMULA results from images in minhpvo/ocr-input using PaddleOCR-VL, an ultra-compact 0.9B OCR model.
Processing Details
Source Dataset: minhpvo/ocr-input
Model: PaddlePaddle/PaddleOCR-VL
Task Mode: formula - Mathematical formula recognition to LaTeX
Number of Samples: 13
Processing Time: 1.9 min
Processing Date: 2026-02-09 04:05 UTC
Configuration
Image Column: image
Output… See the full description on the dataset page: https://huggingface.co/datasets/minhpvo/ocr-output_paddleocr_formula.formula-ocr-synth-bench
Formula OCR Synthetic Benchmark Dataset
Dataset Summary
Formula OCR Synthetic Benchmark Dataset is a synthetically rendered mathematical formula and calculus expression dataset built for training neural formula OCR decoders. It contains rendered formula snippets covering integration, differentiation, limits, summations, and matrix algebra.
Dataset Structure
Format: PNG formula crop images + ground-truth LaTeX strings.
Fields:
file: Path to formula… See the full description on the dataset page: https://huggingface.co/datasets/Srijan-Upadhyay/formula-ocr-synth-bench.Formula_One_Cars
🏎️ Formula 1 Cars Dataset
A curated computer vision and multi-modal dataset consisting of 2,407 high-resolution images across 8 prominent Formula 1 teams.
📌 Dataset Summary
Total Valid Images: 2,407
Number of Classes: 8 F1 Teams
Format: Hugging Face datasets.Image feature + Class Label
Train / Test Split: 80% Train (1,925 images) / 20% Test (482 images)
🗂️ Dataset Structure
Features
Each instance contains the exact image object… See the full description on the dataset page: https://huggingface.co/datasets/Rakesh7n/Formula_One_Cars.ocr-output_glm_formula
Document OCR using GLM-OCR
This dataset contains OCR results from images in minhpvo/ocr-input using GLM-OCR, a compact 0.9B OCR model achieving SOTA performance.
Processing Details
Source Dataset: minhpvo/ocr-input
Model: zai-org/GLM-OCR
Task: formula recognition
Number of Samples: 13
Processing Time: 2.3 min
Processing Date: 2026-02-09 04:16 UTC
Configuration
Image Column: image
Output Column: markdown
Dataset Split: train
Batch Size: 16
Max Model… See the full description on the dataset page: https://huggingface.co/datasets/minhpvo/ocr-output_glm_formula.split_table-formula-dataset-augmentedtable-formula-dataset-augmentedocr-output_paddleocr15_formula
Document Processing using PaddleOCR-VL-1.5 (FORMULA mode)
This dataset contains FORMULA results from images in minhpvo/ocr-input using PaddleOCR-VL-1.5, an ultra-compact 0.9B SOTA OCR model.
Processing Details
Source Dataset: minhpvo/ocr-input
Model: PaddlePaddle/PaddleOCR-VL-1.5
Task Mode: formula - Mathematical formula recognition to LaTeX
Number of Samples: 13
Processing Time: 30.5 min
Processing Date: 2026-02-07 17:04 UTC
Configuration
Image Column:… See the full description on the dataset page: https://huggingface.co/datasets/minhpvo/ocr-output_paddleocr15_formula.formula_tables
