CoolFace
Datasetpublic

venketh/SlimPajama-62B

Subset of cerebras/SlimPajama-627B, consisting of 10% of the train split and 100% of the test and validation splits. The train split consists of chunk2 from the original [cerebras/SlimPajama-627B] dataset, split into five zstd-compressed jsonl files for efficient loading. The dataset is 70 GB compressed, 249 GB uncompressed. @misc{cerebras2023slimpajama, author = {Soboleva, Daria and Al-Khateeb, Faisal and Myers, Robert and Steeves, Jacob R and Hestness, Joel and Dey, Nolan}… See the full description on the dataset page: https://huggingface.co/datasets/venketh/SlimPajama-62B.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
3likes335downloads
9 commits on main
2e464423y ago

Add README

venketh
0e613643y ago

Import train/chunk2 part 5 (5000-5xxx)

venketh
60326283y ago

Import train/chunk2 part 3 (4000-5000)

venketh
f64aeb43y ago

Import train/chunk2 part 3 (3000-4000)

venketh
1ee716b3y ago

Import train/chunk2 part 2 (2000-3000)

venketh
2b7dc433y ago

Import train/chunk2 part 1 (1000-2000)

venketh
ef0511e3y ago

Import train/chunk2 part 0 (1-1000)

venketh
853012f3y ago

Import test and validation splits

venketh
2678d003y ago

initial commit

venketh