CoolFace
20 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datajuicer /the-pile-pubmed-abstracts-refined-by-data-juicer The Pile -- PubMed Abstracts (refined by Data-Juicer) A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.texttext-generationn<1K3 likes113 downloads3y agoHugging Face02datajuicer /the-pile-pubmed-central-refined-by-data-juicer The Pile -- PubMed Central (refined by Data-Juicer) A refined version of PubMed Central dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 83G). Dataset Information Number of samples: 2,694,860 (Keep ~86.96% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-central-refined-by-data-juicer.texttext-generationn<1K2 likes50 downloads3y agoHugging Face03datajuicer /alpaca-cot-zh-refined-by-data-juicer Alpaca-CoT -- ZH (refined by Data-Juicer) A refined Chinese version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to fine-tune a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 18.7GB). Dataset Information Number of samples: 9,873,214 (Keep ~46.58% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-zh-refined-by-data-juicer.texttext-generationn<1K5 likes37 downloads3y agoHugging Face04datajuicer /the-pile-uspto-refined-by-data-juicer The Pile -- USPTO (refined by Data-Juicer) A refined version of USPTO dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 18G). Dataset Information Number of samples: 4,516,283 (Keep ~46.77% from the original dataset) Refining Recipe #… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-uspto-refined-by-data-juicer.texttext-generationn<1K0 likes28 downloads3y agoHugging Face05datajuicer /redpajama-arxiv-refined-by-data-juicer RedPajama -- ArXiv (refined by Data-Juicer) A refined version of ArXiv dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 85GB). Dataset Information Number of samples: 1,655,259 (Keep ~95.99% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-arxiv-refined-by-data-juicer.texttext-generationn<1K2 likes26 downloads3y agoHugging Face06datajuicer /redpajama-cc-2022-05-refined-by-data-juicer RedPajama -- CommonCrawl-2022-05 (refined by Data-Juicer) A refined version of CommonCrawl-2022-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 265GB). Dataset Information Number of samples: 42,648,496 (Keep ~45.34% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2022-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes25 downloads3y agoHugging Face07datajuicer /alpaca-cot-en-refined-by-data-juicer Alpaca-CoT -- EN (refined by Data-Juicer) A refined English version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to fine-tune a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 226GB). Dataset Information Number of samples: 72,855,345 (Keep ~54.48% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-en-refined-by-data-juicer.texttext-generationn<1K0 likes25 downloads3y agoHugging Face08datajuicer /redpajama-pile-stackexchange-refined-by-data-juicer RedPajama & The Pile -- StackExchange (refined by Data-Juicer) A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB). Dataset Information Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.texttext-generationn<1K0 likes20 downloads3y agoHugging Face09datajuicer /redpajama-cc-2019-30-refined-by-data-juicer RedPajama -- CommonCrawl-2019-30 (refined by Data-Juicer) A refined version of CommonCrawl-2019-30 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 240GB). Dataset Information Number of samples: 36,557,283 (Keep ~45.08% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2019-30-refined-by-data-juicer.tabulartext-generationn<1K0 likes19 downloads3y agoHugging Face10datajuicer /redpajama-wiki-refined-by-data-juicer RedPajama -- Wikipedia (refined by Data-Juicer) A refined version of Wikipedia dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 68GB). Dataset Information Number of samples: 26,990,659 (Keep ~90.47% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-wiki-refined-by-data-juicer.texttext-generationn<1K2 likes18 downloads3y agoHugging Face11datajuicer /the-pile-europarl-refined-by-data-juicer The Pile -- EuroParl (refined by Data-Juicer) A refined version of EuroParl dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.2GB). Dataset Information Number of samples: 61,601 (Keep ~88.23% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-europarl-refined-by-data-juicer.texttext-generationn<1K0 likes15 downloads3y agoHugging Face12datajuicer /redpajama-stack-code-refined-by-data-juicer RedPajama & TheStack -- Github Code (refined by Data-Juicer) A refined version of Github Code dataset in RedPajama & TheStack by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 232GB). Dataset Information Number of samples: 49,279,344 (Keep ~52.09% from the original… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-stack-code-refined-by-data-juicer.tabulartext-generationn<1K2 likes13 downloads3y agoHugging Face13datajuicer /the-pile-hackernews-refined-by-data-juicer The Pile -- HackerNews (refined by Data-Juicer) A refined version of HackerNews dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.8G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-hackernews-refined-by-data-juicer.texttext-generationn<1K0 likes13 downloads3y agoHugging Face14datajuicer /the-pile-philpaper-refined-by-data-juicer The Pile -- PhilPaper (refined by Data-Juicer) A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB). Dataset Information Number of samples: 29,117 (Keep ~88.82% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.texttext-generationn<1K0 likes12 downloads3y agoHugging Face15datajuicer /redpajama-cc-2020-05-refined-by-data-juicer RedPajama -- CommonCrawl-2020-05 (refined by Data-Juicer) A refined version of CommonCrawl-2020-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 297GB). Dataset Information Number of samples: 42,612,596 (Keep ~46.90% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2020-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes12 downloads3y agoHugging Face16datajuicer /redpajama-cc-2023-06-refined-by-data-juicer RedPajama -- CommonCrawl-2023-06 (refined by Data-Juicer) A refined version of CommonCrawl-2023-06 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 310GB). Dataset Information Number of samples: 50,643,699 (Keep ~45.46% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2023-06-refined-by-data-juicer.tabulartext-generationn<1K0 likes8 downloads3y agoHugging Face17datajuicer /the-pile-freelaw-refined-by-data-juicer The Pile -- FreeLaw (refined by Data-Juicer) A refined version of FreeLaw dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 45GB). Dataset Information Number of samples: 2,942,612 (Keep ~82.61% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-freelaw-refined-by-data-juicer.texttext-generationn<1K0 likes8 downloads3y agoHugging Face18datajuicer /the-pile-nih-refined-by-data-juicer The Pile -- NIHExPorter (refined by Data-Juicer) A refined version of NIHExPorter dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.0G). Dataset Information Number of samples: 858,492 (Keep ~91.36% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-nih-refined-by-data-juicer.texttext-generationn<1K0 likes8 downloads3y agoHugging Face19datajuicer /redpajama-cc-2021-04-refined-by-data-juicer RedPajama -- CommonCrawl-2021-04 (refined by Data-Juicer) A refined version of CommonCrawl-2021-04 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 284GB). Dataset Information Number of samples: 44,724,752 (Keep ~45.23% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2021-04-refined-by-data-juicer.tabulartext-generationn<1K0 likes6 downloads3y agoHugging Face20datajuicer /redpajama-c4-refined-by-data-juicer RedPajama -- C4 (refined by Data-Juicer) A refined version of C4 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 832GB). Dataset Information Number of samples: 344,491,171 (Keep ~94.42% from the original dataset) Refining Recipe #… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-c4-refined-by-data-juicer.texttext-generationn<1K1 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.