CoolFace
Datasetpublic

venketh/SlimPajama-62B

Subset of cerebras/SlimPajama-627B, consisting of 10% of the train split and 100% of the test and validation splits. The train split consists of chunk2 from the original [cerebras/SlimPajama-627B] dataset, split into five zstd-compressed jsonl files for efficient loading. The dataset is 70 GB compressed, 249 GB uncompressed. @misc{cerebras2023slimpajama, author = {Soboleva, Daria and Al-Khateeb, Faisal and Myers, Robert and Steeves, Jacob R and Hestness, Joel and Dey, Nolan}… See the full description on the dataset page: https://huggingface.co/datasets/venketh/SlimPajama-62B.

sourceHugging Faceapache-2.0updated 3y agoView on Hugging Face
3likes335downloads
Dataset Card

Subset of cerebras/SlimPajama-627B, consisting of 10% of the train split and 100% of the test and validation splits.

The train split consists of chunk2 from the original [cerebras/SlimPajama-627B] dataset, split into five zstd-compressed jsonl files for efficient loading. The dataset is 70 GB compressed, 249 GB uncompressed. ---

@misc{cerebras2023slimpajama,
  author = {Soboleva, Daria and Al-Khateeb, Faisal and Myers, Robert and Steeves, Jacob R and Hestness, Joel and Dey, Nolan},
  title = {{SlimPajama: A 627B token cleaned and deduplicated version of RedPajama}},
  month = June,
  year = 2023,
  howpublished = {\url{https://www.cerebras.net/blog/slimpajama-a-627b-token-cleaned-and-deduplicated-version-of-redpajama}},
  url = {https://huggingface.co/datasets/cerebras/SlimPajama-627B},
}