CoolFace
6 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datajuicer /redpajama-cc-2022-05-refined-by-data-juicer RedPajama -- CommonCrawl-2022-05 (refined by Data-Juicer) A refined version of CommonCrawl-2022-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 265GB). Dataset Information Number of samples: 42,648,496 (Keep ~45.34% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2022-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes25 downloads3y agoHugging Face02datajuicer /redpajama-cc-2019-30-refined-by-data-juicer RedPajama -- CommonCrawl-2019-30 (refined by Data-Juicer) A refined version of CommonCrawl-2019-30 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 240GB). Dataset Information Number of samples: 36,557,283 (Keep ~45.08% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2019-30-refined-by-data-juicer.tabulartext-generationn<1K0 likes19 downloads3y agoHugging Face03datajuicer /redpajama-stack-code-refined-by-data-juicer RedPajama & TheStack -- Github Code (refined by Data-Juicer) A refined version of Github Code dataset in RedPajama & TheStack by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 232GB). Dataset Information Number of samples: 49,279,344 (Keep ~52.09% from the original… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-stack-code-refined-by-data-juicer.tabulartext-generationn<1K2 likes13 downloads3y agoHugging Face04datajuicer /redpajama-cc-2020-05-refined-by-data-juicer RedPajama -- CommonCrawl-2020-05 (refined by Data-Juicer) A refined version of CommonCrawl-2020-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 297GB). Dataset Information Number of samples: 42,612,596 (Keep ~46.90% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2020-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes12 downloads3y agoHugging Face05datajuicer /redpajama-cc-2023-06-refined-by-data-juicer RedPajama -- CommonCrawl-2023-06 (refined by Data-Juicer) A refined version of CommonCrawl-2023-06 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 310GB). Dataset Information Number of samples: 50,643,699 (Keep ~45.46% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2023-06-refined-by-data-juicer.tabulartext-generationn<1K0 likes8 downloads3y agoHugging Face06datajuicer /redpajama-cc-2021-04-refined-by-data-juicer RedPajama -- CommonCrawl-2021-04 (refined by Data-Juicer) A refined version of CommonCrawl-2021-04 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 284GB). Dataset Information Number of samples: 44,724,752 (Keep ~45.23% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2021-04-refined-by-data-juicer.tabulartext-generationn<1K0 likes6 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.