CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01scholarweave /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/scholarweave/arxiv-latex.texttext-generation1M<n<10M147 likes20k downloads8d agoHugging Face02OleehyO /latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set image10M<n<100M27 likes6.5k downloads1y agoHugging Face03TIGER-Lab /arxiv-latex-5TThe dataset used for https://github.com/TIGER-AI-Lab/ScholarCopilot. image10M<n<100M5 likes5.7k downloads1y agoHugging Face04WindyVerse /Handwritten-Latex-Datasets Dataset This data set includes common handwritten formulas in junior high schools and high schools, and is labeled in Latex format. Can be used to train models that recognize common numbers, fractions, and sets. Dataset source Collected in various junior high schools and high schools, handwritten by students. Usage The label is stored at json folder and scanned hand-writted pictures are stored at pic folder. Scan the qr code of the picture to get the index and… See the full description on the dataset page: https://huggingface.co/datasets/WindyVerse/Handwritten-Latex-Datasets.imageimage-to-text1K<n<10K1 likes4.8k downloads3y agoHugging Face05stanford-crfm /image2struct-latex-v1 Image2Struct - Latex Paper | Website | Datasets (Webpages, Latex, Music sheets) | Leaderboard | HELM repo | Image2Struct repo License: Apache License Version 2.0, January 2004 Dataset description Image2struct is a benchmark for evaluating vision-language models in practical tasks of extracting structured information from images. This subdataset focuses on LaTeX code. The model is given an image of the expected output with the prompt: Please provide the LaTex code used to… See the full description on the dataset page: https://huggingface.co/datasets/stanford-crfm/image2struct-latex-v1.imagequestion-answering1K<n<10K12 likes3.8k downloads2y agoHugging Face06JosselinSom /Latex-VLMimagequestion-answering1K<n<10K9 likes3.4k downloads3y agoHugging Face07unsloth /LaTeX_OCR1% sampled from https://huggingface.co/datasets/linxy/LaTeX_OCR image10K<n<100K92 likes2.7k downloads2y agoHugging Face08synthetix-institute /latex-data-pub Hyperion: Scientific LaTeX Corpus (Public) This dataset constitutes the Public Scientific Corpus for the Hyperion Project at the Synthetix Institute. It contains high-fidelity LaTeX source text extracted from diverse scientific repositories, optimized for topological knowledge discovery and relational mapping. Dataset Details Total Documents: ~1,318,468 Average Document Length: Variable (approx. 32KB - 256KB) Primary Domain: Mathematics, Physics, and Chemistry. Goal:… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data-pub.texttext-generation1M<n<10M1 likes1.5k downloads6mo agoHugging Face09linxy /LaTeX_OCR LaTeX OCR 的数据仓库 本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。 如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++ 后续追加新的数据也会放在这个仓库 ~~ 原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR. 数据集 本仓库有 5 个数据集 small 是小数据集,样本数 110 条,用于测试 full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。 synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/linxy/LaTeX_OCR.imageimage-to-text100K<n<1M181 likes940 downloads2y agoHugging Face10synthetix-institute /latex-data Hyperion: Scientific LaTeX Corpus This dataset constitutes the Internal Scientific Corpus for the Hyperion Project of Synthetix Institute. It contains LaTeX source documents used for high-fidelity relational extraction and internal benchmarking of the Epistemic Manifold. Dataset Details Total Documents: ~1,192,727 (Internal Base) Status: Private Access: Restricted to Synthetix Institute authorized personnel. Primary Use: Training the private sector of the Epistemic… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data.texttext-generation1M<n<10M0 likes850 downloads6mo agoHugging Face11naveenmarthala /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/naveenmarthala/arxiv-latex.texttext-generation1M<n<10M0 likes653 downloads3mo agoHugging Face12natnitaract /exams-basic-and-quantum-cryptography-and-security-latex Open Problem Exams: Cryptography and Security (LaTeX) A curated dataset of open-ended exam problems (with solutions) in cryptography and computer security, formatted in LaTeX. The dataset is sourced from university courses at three institutions. Dataset Overview Institution Files Topics Questions Caltech & TU Delft 8 38 145 EPFL 6 19 86 ETH Zurich 1 14 37 MIT 3 33 79 Total 18 104 347 Difficulty Distribution Institution… See the full description on the dataset page: https://huggingface.co/datasets/natnitaract/exams-basic-and-quantum-cryptography-and-security-latex.textn<1K1 likes470 downloads6mo agoHugging Face13OleehyO /latex-formulas 𝑩𝑰𝑮 𝑵𝑬𝑾𝑺‼️ 📮 [2025-08] We have updated to a larger dataset, which contains nearly 80 million samples, with significant improvements in both data quality and diversity. [link] 📮 [2024-02] We trained a formula recognition model, 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫, using the latex-formulas dataset. It can convert LaTeX formulas into images and boasts high accuracy and strong generalization capabilities, covering most formula recognition scenarios. For more details, please refer to the… See the full description on the dataset page: https://huggingface.co/datasets/OleehyO/latex-formulas.imageimage-to-text1M<n<10M103 likes427 downloads1y agoHugging Face14Smartlearners /LaTeX_OCRimage10K<n<100K0 likes238 downloads13d agoHugging Face15Kyudan /arXiv_latex TeX data from arXiv Using https://github.com/KyuDan1/TeX2Image code. We have categories Math, Physics, Statistics, ComputerScience. Domain Size Mathematics 4.22M Computer Science 2.76M Statistics 0.89M Physics 0.78M Total (unique) 7.17M text10M<n<100M4 likes234 downloads2y agoHugging Face16katphlab /pubtables-lateximage100K<n<1M0 likes208 downloads2y agoHugging Face17deepcopy /pubtables-lateximage100K<n<1M0 likes204 downloads1y agoHugging Face18Mithilss /neurips-2025-arxiv-latex-sources NeurIPS 2025 arXiv LaTeX Source Files This dataset contains file-level Parquet rows built from extracted raw arXiv source packages for papers mapped from the NeurIPS 2025 proceedings to arXiv records. Each row is one file from one arXiv source package. Use arxiv_id to group files back into papers. Columns arxiv_id: arXiv identifier for the source package. title: paper title from the mapping CSV. source_url: arXiv e-print source URL. paper_index, paper_status… See the full description on the dataset page: https://huggingface.co/datasets/Mithilss/neurips-2025-arxiv-latex-sources.tabular100K<n<1M0 likes203 downloads4mo agoHugging Face19harryrobert /nav2tex-decoder-latex-pretraintext1M<n<10M0 likes201 downloads5mo agoHugging Face20RAG4Math /targets_fixed_filtered_latextextn<1K0 likes198 downloads4mo agoHugging Face21thejagstudio /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/thejagstudio/arxiv-latex.texttext-generation1M<n<10M0 likes194 downloads28d agoHugging Face22DataMuncher-Labs /LaTeXt Dataset Card for MathOCR Dataset Details Dataset Sources [optional] Repository: [https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt] I don't write papers or demos Uses Direct Use Post training for mathematical visualization. Out-of-Scope Use [Its not meant, but technically works for diffusion models.] [The only way this dataset can be misused is in] A. [Incompatible models, like TTS] B. [Copying the repo without… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt.texttext-classification1M<n<10M0 likes192 downloads7mo agoHugging Face23RAG4Math /candidates_fixed_filtered_latextext10K<n<100K0 likes177 downloads4mo agoHugging Face24justatomic /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/justatomic/arxiv-latex.texttext-generation1M<n<10M0 likes163 downloads3mo agoHugging Face25piushorn /wikipedia-latex-formulas-319k Wikipedia LaTeX Formulas 319k Dataset A curated collection of mathematical formulas from English Wikipedia, providing LaTeX source code paired with high-quality rendered images for training OCR and formula recognition models. Dataset Description Dataset Summary This dataset represents a complete extraction of ~319k LaTeX mathematical formulas from English Wikipedia articles (October 2025 snapshot), filtered by visual complexity (score > 8) and renderability… See the full description on the dataset page: https://huggingface.co/datasets/piushorn/wikipedia-latex-formulas-319k.imageimage-to-text100K<n<1M5 likes159 downloads6mo agoHugging Face26lishaoyong /latex-formulas-80MFor more details, please refer to the 𝐓𝐞𝐱𝐓𝐞𝐥𝐥𝐞𝐫 GitHub repository. IMPORTANT NOTE!!! The handwritten subset of this dataset was collected entirely from existing open source work, which includes all test sets. If you want to use this subset for your experimental ablation, please filter it yourself based on the latex label of the test set image10M<n<100M0 likes146 downloads9mo agoHugging Face27deepcopy /pubmed-tables-latex-768pximage100K<n<1M2 likes143 downloads1y agoHugging Face28ymfl /LaTeX_OCR LaTeX OCR 的数据仓库 本数据仓库是专为 LaTeX_OCR 及 LaTeX_OCR_PRO 制作的数据,来源于 https://zenodo.org/record/56198#.V2p0KTXT6eA 以及 https://www.isical.ac.in/~crohme/ 以及我们自己构建。 如果这个数据仓库有帮助到你的话,请点亮 ❤️like ++ 后续追加新的数据也会放在这个仓库 ~~ 原始数据仓库在github LinXueyuanStdio/Data-for-LaTeX_OCR. 数据集 本仓库有 5 个数据集 small 是小数据集,样本数 110 条,用于测试 full 是印刷体约 100k 的完整数据集。实际上样本数略小于 100k,因为用 LaTeX 的抽象语法树剔除了很多不能渲染的 LaTeX。 synthetic_handwrite 是手写体 100k 的完整数据集,基于 full 的公式,使用手写字体合成而来,可以视为人类在纸上的手写体。样本数实际上略小于 100k,理由同上。… See the full description on the dataset page: https://huggingface.co/datasets/ymfl/LaTeX_OCR.imageimage-to-text100K<n<1M1 likes137 downloads1mo agoHugging Face29SoyVitou /Latex-Math-HME100KHME100K – Hugging Face Version Dataset Name: HME100K (Hugging Face Version) Original Source: https://www.kaggle.com/datasets/cutedeadu/hme100k Original Author (Kaggle): cutedeadu Overview HME100K is a dataset of handwritten mathematical expressions. Each sample consists of: An image containing a handwritten math expression A corresponding LaTeX-style text annotation This dataset is useful for: Handwritten Mathematical Expression Recognition (HMER) OCR research for… See the full description on the dataset page: https://huggingface.co/datasets/SoyVitou/Latex-Math-HME100K.image10K<n<100K0 likes132 downloads7mo agoHugging Face30PadishahIIIXXX /latex-ocr-dataset Synthetic LaTeX OCR Dataset Overview This dataset contains synthetically generated LaTeX formula images designed to augment training data for LaTeX OCR models. The dataset applies style enrichment techniques to existing real-world LaTeX datasets, creating diverse visual representations of mathematical formulas through PDF-rendering and font styling. Dataset Structure synth/ ├── plain/ │ ├── train/ │ │ ├── images/ │ │ │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/PadishahIIIXXX/latex-ocr-dataset.imageimage-to-text1M<n<10M2 likes121 downloads24d agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.