CoolFace
Datasetpublic

aluncstokes/mathpile_arxiv_subset_tiny

MathPile ArXiv (subset) Description This dataset consists of a toy subset of 8834 (5000 training + 3834 testing) TeX files found in the arXiv subset of MathPile, used for testing. You should not use this dataset. Training and testing sets are already split Source The data was obtained from the training + validation portion of the arXiv subset of MathPile. Format Given as JSONL files of JSON dicts each containing the single key:… See the full description on the dataset page: https://huggingface.co/datasets/aluncstokes/mathpile_arxiv_subset_tiny.

sourceHugging Faceupdated 3y agoView on Hugging Face
0likes50downloads
test_chunked.jsonl4 linesDownload Raw Back to root
1version https://git-lfs.github.com/spec/v12oid sha256:ffde9da9a708bda09af6eed842237c33e137f28b6ae14f40dcf55da48e7c18643size 2858919044