CoolFace
Datasetpublic

KamDickGoon/Killer

Getting Started RedPajama-V2 is an open dataset for training large language models. The dataset includes over 100B text documents coming from 84 CommonCrawl snapshots and processed using the CCNet pipeline. Out of these, there are 30B documents in the corpus that additionally come with quality signals. In addition, we also provide the ids of duplicated documents which can be used to create a dataset with 20B deduplicated documents. Check out our blog post for more details on the… See the full description on the dataset page: https://huggingface.co/datasets/KamDickGoon/Killer.

sourceHugging Faceupdated 5mo agoView on Hugging Face
0likes131downloads
1 commits on main
10f41305mo ago

Duplicate from togethercomputer/RedPajama-Data-V2

KamDickGoon, mauriceweber