datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
arxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/scholarweave/arxiv-latex.latex-data-pub
Hyperion: Scientific LaTeX Corpus (Public)
This dataset constitutes the Public Scientific Corpus for the Hyperion Project at the Synthetix Institute. It contains high-fidelity LaTeX source text extracted from diverse scientific repositories, optimized for topological knowledge discovery and relational mapping.
Dataset Details
Total Documents: ~1,318,468
Average Document Length: Variable (approx. 32KB - 256KB)
Primary Domain: Mathematics, Physics, and Chemistry.
Goal:… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data-pub.latex-data
Hyperion: Scientific LaTeX Corpus
This dataset constitutes the Internal Scientific Corpus for the Hyperion Project of Synthetix Institute. It contains LaTeX source documents used for high-fidelity relational extraction and internal benchmarking of the Epistemic Manifold.
Dataset Details
Total Documents: ~1,192,727 (Internal Base)
Status: Private
Access: Restricted to Synthetix Institute authorized personnel.
Primary Use: Training the private sector of the Epistemic… See the full description on the dataset page: https://huggingface.co/datasets/synthetix-institute/latex-data.arxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/naveenmarthala/arxiv-latex.arxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv)… See the full description on the dataset page: https://huggingface.co/datasets/thejagstudio/arxiv-latex.LaTeXt
Dataset Card for MathOCR
Dataset Details
Dataset Sources [optional]
Repository: [https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt]
I don't write papers or demos
Uses
Direct Use
Post training for mathematical visualization.
Out-of-Scope Use
[Its not meant, but technically works for diffusion models.]
[The only way this dataset can be misused is in]
A. [Incompatible models, like TTS]
B. [Copying the repo without… See the full description on the dataset page: https://huggingface.co/datasets/DataMuncher-Labs/LaTeXt.arxiv-latex
arXiv LaTeX Source Dataset
This dataset provides the entire corpus of arXiv's LaTeX source files, pre-parsed, formatted, and aligned with official metadata in ready-to-query Parquet files.
Why I Built This
If you have ever tried to work with the complete history of arXiv papers at scale, you have likely run into two massive hurdles:
Network Egress Costs: While arXiv does offer public bulk access to its source files via S3 (s3://arxiv), the bucket is configured… See the full description on the dataset page: https://huggingface.co/datasets/justatomic/arxiv-latex.LaTeX_Image_Pairs
LaTeX Image Pairs Dataset
This dataset comprises a unique collection of LaTeX expressions paired with their corresponding images. The LaTeX expressions were meticulously scraped from a variety of open-source textbooks, ensuring a diverse and comprehensive dataset. Sample references from these textbooks will be provided to illustrate the sources of these expressions.
In addition to the raw LaTeX expressions, this dataset includes images of the rendered expressions. Each LaTeX… See the full description on the dataset page: https://huggingface.co/datasets/henryholloway/LaTeX_Image_Pairs.english-to-latex-notation-sft
English-to-LaTeX Notation SFT
A 24,956-row supervised fine-tuning dataset for converting precise English descriptions of mathematical notation into exactly one standalone LaTeX expression.
This dataset teaches notation writing, not mathematical problem solving. The requested expression should not be evaluated, simplified, proved, derived, or explained.
Splits
Split
Rows
Formal target groups
Purpose
train
22,460
9,430
Supervised fine-tuning
validation… See the full description on the dataset page: https://huggingface.co/datasets/piyushgupta53/english-to-latex-notation-sft.
