CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01OleehyO /latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set image10M<n<100M27 likes6.5k downloads1y agoHugging Face02TIGER-Lab /arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot. image10M<n<100M5 likes5.7k downloads1y agoHugging Face03WindyVerse /Handwritten-Latex-Datasets Dataset This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets. Dataset source Collected in various junior high schools and high schools, handwritten by students. Usage The label is stored at json folder and scanned hand-writted pictures are stored at pic folder. Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.imageimage-to-text1K<n<10K1 likes4.8k downloads3y agoHugging Face04stanford-crfm /image2struct-latex-v1 Image2Struct - Latex Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo License: Apache License Version 2.0, January 2004 Dataset description Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images. This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt: Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.imagequestion-answering1K<n<10K12 likes3.8k downloads2y agoHugging Face05JosselinSom /Latex-VLMimagequestion-answering1K<n<10K9 likes3.4k downloads3y agoHugging Face06unsloth /LaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR image10K<n<100K92 likes2.7k downloads2y agoHugging Face07linxy /LaTeX_OCR LaTeX OCR 的数据仓库 本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。 如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++ 后续追加新的数据也会放在这个仓库 ~~ 原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR. 数据集 本仓库有 5 个数据集 small 是小数据集,样本数 110 条,用于测试 full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。 synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/linxy/LaTeX_OCR.imageimage-to-text100K<n<1M181 likes940 downloads2y agoHugging Face08OleehyO /latex-formulas 𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️ 📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link] 📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios. For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/OleehyO/latex-formulas.imageimage-to-text1M<n<10M103 likes427 downloads1y agoHugging Face09Smartlearners /LaTeX_OCRimage10K<n<100K0 likes238 downloads14d agoHugging Face10katphlab /pubtables-lateximage100K<n<1M0 likes208 downloads2y agoHugging Face11deepcopy /pubtables-lateximage100K<n<1M0 likes204 downloads1y agoHugging Face12piushorn /wikipedia-latex-formulas-319k Wikipedia LaTeX Formulas 319k Dataset A curated collection of mathematical formulas from English Wikipedia, providing LaTeX source code paired with high-quality rendered images for training OCR and formula recognition models. Dataset Description Dataset Summary This dataset represents a complete extraction of ~319k LaTeX mathematical formulas from English Wikipedia articles (October 2025 snapshot), filtered by visual complexity (score > 8) and renderability… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/wikipedia-latex-formulas-319k.imageimage-to-text100K<n<1M5 likes159 downloads6mo agoHugging Face13lishaoyong /latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set image10M<n<100M0 likes146 downloads9mo agoHugging Face14deepcopy /pubmed-tables-latex-768pximage100K<n<1M2 likes143 downloads1y agoHugging Face15ymfl /LaTeX_OCR LaTeX OCR 的数据仓库 本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。 如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++ 后续追加新的数据也会放在这个仓库 ~~ 原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR. 数据集 本仓库有 5 个数据集 small 是小数据集,样本数 110 条,用于测试 full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。 synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/ymfl/LaTeX_OCR.imageimage-to-text100K<n<1M1 likes137 downloads1mo agoHugging Face16SoyVitou /Latex-Math-HME100KHME100K – Hugging Face Version Dataset Name: HME100K (Hugging Face Version) Original Source: https://www.kaggle.com/datasets/cutedeadu/hme100k Original Author (Kaggle): cutedeadu Overview HME100K is a dataset of handwritten mathematical expressions. Each sample consists of: An image containing a handwritten math expression A corresponding LaTeX-style text annotation This dataset is useful for: Handwritten Mathematical Expression Recognition (HMER) OCR research for… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/Latex-Math-HME100K.image10K<n<100K0 likes132 downloads7mo agoHugging Face17PadishahIIIXXX /latex-ocr-dataset Synthetic LaTeX OCR Dataset Overview This dataset contains synthetically generated LaTeX formula images designed to augment training data for LaTeX OCR models. The dataset applies style enrichment techniques to existing real-world LaTeX datasets, creating diverse visual representations of mathematical formulas through PDF-rendering and font styling. Dataset Structure synth/ ├── plain/ │ ├── train/ │ │ ├── images/ │ │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/PadishahIIIXXX/latex-ocr-dataset.imageimage-to-text1M<n<10M2 likes121 downloads24d agoHugging Face18Azu /Handwritten-Mathematical-Expression-Convert-LaTeXimage10K<n<100K28 likes102 downloads5y agoHugging Face19Nagase-Kotono /img-latex-ko IMG2LaTeX 1. LaTeX-ko 1. 1. 설명 IM2LATEX-230k 을 LLaVA 형식에 맞게 수정한 데이터셋 입니다. 사용법은 LLaVA, KoLLaVA 참고 하시기 바랍니다. imagevisual-question-answering100K<n<1M5 likes97 downloads2y agoHugging Face20MosRat2333 /ZhEn-latex-ocrimage100K<n<1M13 likes84 downloads2y agoHugging Face21tbrrss /latextract-arxiv-math latextract-arxiv-math A multi-domain corpus of (formula image, LaTeX source) pairs harvested from arXiv source tarballs. Each formula is rendered in isolation with tectonic into a tightly-cropped PNG; the ground-truth LaTeX is the body of the original display-math environment, normalized to remove \label{...}. Provenance Source: arXiv e-print/<id> tarballs (publicly downloadable) Renderer: tectonic (LaTeX engine) + MuPDF (PDF → PNG @ 200 DPI) Filtering: papers using… See the full description on the dataset page: https://huggingface.co/datasets/tbrrss/latextract-arxiv-math.imageimage-to-text10K<n<100K0 likes79 downloads5mo agoHugging Face22yayossd /latex-formulas 𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️ 📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link] 📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios. For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/yayossd/latex-formulas.imageimage-to-text1M<n<10M0 likes73 downloads2mo agoHugging Face23lukbl /LaTeX-OCR-dataset LaTeX-OCR Dataset Summary This dataset was created to train LaTeX-OCR, a model for recognizing LaTeX code from images of mathematical formulas. Each sample consists of a synthetically rendered formula image and its corresponding LaTeX formula. The images were generated from scratch using xelatex with multiple fonts, offering more typographic variety than many other datasets that use a single font (typically Computer Modern). Data Sources Formulas were… See the full description on the dataset page: https://huggingface.co/datasets/lukbl/LaTeX-OCR-dataset.imageimage-to-text100K<n<1M0 likes67 downloads1y agoHugging Face24EliMC /LaTeX_OCR LaTeX OCR 的数据仓库 本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。 如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++ 后续追加新的数据也会放在这个仓库 ~~ 原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR. 数据集 本仓库有 5 个数据集 small 是小数据集,样本数 110 条,用于测试 full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。 synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/EliMC/LaTeX_OCR.imageimage-to-text100K<n<1M0 likes67 downloads10mo agoHugging Face25henryholloway /LaTeX_Image_Pairs LaTeX Image Pairs Dataset This dataset comprises a unique collection of LaTeX expressions paired with their corresponding images. The LaTeX expressions were meticulously scraped from a variety of open-source textbooks, ensuring a diverse and comprehensive dataset. Sample references from these textbooks will be provided to illustrate the sources of these expressions. In addition to the raw LaTeX expressions, this dataset includes images of the rendered expressions. Each LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/henryholloway/LaTeX_Image_Pairs.imageimage-classification1K<n<10K2 likes66 downloads3y agoHugging Face26xiaochuan-dev /lateximage0 likes61 downloads18d agoHugging Face27lamm-mit /OleehyO-latex-formulasimage100K<n<1M1 likes60 downloads2y agoHugging Face28Evanstarcraft2 /latex-formulas 𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️ 📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link] 📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios. For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/Evanstarcraft2/latex-formulas.imageimage-to-text1M<n<10M0 likes53 downloads10mo agoHugging Face29Kyudan /LaTeXImage_testimagen<1K1 likes52 downloads3y agoHugging Face30lokeshe09 /LaTeX_OCRRRimage1K<n<10K0 likes47 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.