CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01scholarweave /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/scholarweave/arxiv-latex.texttext-generation1M<n<10M147 likes20k downloads8d agoHugging Face02synthetix-institute /latex-data-pub Hyperion: Scientific LaTeX Corpus (Public) This dataset constitutes the Public Scientific Corpus for the Hyperion Project at the Synthetix Institute. It contains high-fidelity LaTeX source text extracted from diverse scientific repositories, optimized for topological knowledge discovery and relational mapping. Dataset Details Total Documents: ~1,318,468 Average Document Length: Variable (approx. 32KB - 256KB) Primary Domain: Mathematics, Physics, and Chemistry. Goal:… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data-pub.texttext-generation1M<n<10M1 likes1.4k downloads6mo agoHugging Face03synthetix-institute /latex-data Hyperion: Scientific LaTeX Corpus This dataset constitutes the Internal Scientific Corpus for the Hyperion Project of Synthetix Institute. It contains LaTeX source documents used for high-fidelity relational extraction and internal benchmarking of the Epistemic Manifold. Dataset Details Total Documents: ~1,192,727 (Internal Base) Status: Private Access: Restricted to Synthetix Institute authorized personnel. Primary Use: Training the private sector of the Epistemic… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data.texttext-generation1M<n<10M0 likes831 downloads6mo agoHugging Face04naveenmarthala /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/naveenmarthala/arxiv-latex.texttext-generation1M<n<10M0 likes648 downloads3mo agoHugging Face05thejagstudio /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/thejagstudio/arxiv-latex.texttext-generation1M<n<10M0 likes194 downloads29d agoHugging Face06DataMuncher-Labs /LaTeXt Dataset Card for MathOCR Dataset Details Dataset Sources [optional] Repository: [https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt] I don't write papers or demos Uses Direct Use Post training for mathematical visualization. Out-of-Scope Use [Its not meant, but technically works for diffusion models.] [The only way this dataset can be misused is in] A. [Incompatible models, like TTS] B. [Copying the repo without… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt.texttext-classification1M<n<10M0 likes192 downloads7mo agoHugging Face07justatomic /arxiv-latex arXiv LaTeX Source Dataset This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files. Why I Built This If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles: Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/justatomic/arxiv-latex.texttext-generation1M<n<10M0 likes163 downloads3mo agoHugging Face08henryholloway /LaTeX_Image_Pairs LaTeX Image Pairs Dataset This dataset comprises a unique collection of LaTeX expressions paired with their corresponding images. The LaTeX expressions were meticulously scraped from a variety of open-source textbooks, ensuring a diverse and comprehensive dataset. Sample references from these textbooks will be provided to illustrate the sources of these expressions. In addition to the raw LaTeX expressions, this dataset includes images of the rendered expressions. Each LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/henryholloway/LaTeX_Image_Pairs.imageimage-classification1K<n<10K2 likes66 downloads3y agoHugging Face09piyushgupta53 /english-to-latex-notation-sft English-to-LaTeX Notation SFT A 24,956-row supervised fine-tuning dataset for converting precise English descriptions of mathematical notation into exactly one standalone LaTeX expression. This dataset teaches notation writing, not mathematical problem solving. The requested expression should not be evaluated, simplified, proved, derived, or explained. Splits Split Rows Formal target groups Purpose train 22,460 9,430 Supervised fine-tuning validation… See the full description on the dataset page: https://huggingface.co/datasets/piyushgupta53/english-to-latex-notation-sft.texttext-generation10K<n<100K0 likes32 downloads1mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.