CoolFace
Datasetpublic

JackHsieh/statML-arxiv-40M-20M

Subset of JackHsieh/statML-arxiv. Each document is exactly 4096 tokens. The train split has exactly twice the number of documents as the test split. Split Documents Tokens train 9728 39_845_888 test 4864 19_922_944

sourceHugging Faceupdated 3mo agoView on Hugging Face
0likes404downloads

Nothing at this path on main. The folder may be empty, or the revision may not exist.

JackHsieh/statML-arxiv-40M-20M · main · files are served by the source, never re-hosted here