CoolFace
Datasetpublic

CortexEvolved/arxiv-tex-corpus-full

arxiv-tex-corpus-full (80GB) Large-scale LaTeX corpus from arXiv (math, CS, physics, statistics) 📄 Paper: https://arxiv.org/abs/2602.17288 📚 Overview arxiv-tex-corpus-full (80GB) is a large-scale dataset of LaTeX source content extracted from papers hosted on arXiv. This version contains approximately 80GB of structured JSONL data, restricted to the following arXiv categories: math cs hep-th hep-ph quant-ph stat.ML stat.TH The dataset is designed for research in: Large… See the full description on the dataset page: https://huggingface.co/datasets/CortexEvolved/arxiv-tex-corpus-full.

sourceHugging Facemitupdated 6mo agoView on Hugging Face
1likes45downloads
1 commits on main
2f609a26mo ago

Duplicate from KiteFishAI/arxiv-tex-corpus-full

CortexEvolved, anuj0456