CoolFace
Datasetpublic

JackHsieh/statML-arxiv-40M-20M

Subset of JackHsieh/statML-arxiv. Each document is exactly 4096 tokens. The train split has exactly twice the number of documents as the test split. Split Documents Tokens train 9728 39_845_888 test 4864 19_922_944

sourceHugging Faceupdated 2mo agoView on Hugging Face
0likes437downloads
Dataset Card

Subset of JackHsieh/statML-arxiv. Each document is exactly 4096 tokens. The train split has exactly twice the number of documents as the test split.

SplitDocumentsTokens
train972839845888
test486419922944