CoolFace
Datasetpublic

aluncstokes/mathpile_arxiv_subset

MathPile ArXiv (subset) Description This dataset consists of 343,830 TeX files containing mathematics papers sourced from the arXiv. Training and testing sets are already split Source The data was obtained from the training + validation portion of the arXiv subset of MathPile. Format Given as JSONL files of JSON dicts each containing the single key: "text" Usage LaTeX stuff idk License The original… See the full description on the dataset page: https://huggingface.co/datasets/aluncstokes/mathpile_arxiv_subset.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes514downloads
Dataset Card

MathPile ArXiv (subset)

Description

This dataset consists of 343,830 TeX files containing mathematics papers sourced from the arXiv. Training and testing sets are already split

Source

The data was obtained from the training + validation portion of the arXiv subset of MathPile.

Format

  • —Given as JSONL files of JSON dicts each containing the single key: "text"

Usage

  • —LaTeX stuff idk

License

The original data is subject to the licensing terms of the arXiv. Users should refer to the arXiv's terms of use for details on permissible usage.