CoolFace
24 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01datajuicer /the-pile-pubmed-abstracts-refined-by-data-juicer The Pile -- PubMed Abstracts (refined by Data-Juicer) A refined version of PubMed Abstracts dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 24G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-abstracts-refined-by-data-juicer.texttext-generationn<1K3 likes113 downloads3y agoHugging Face02datajuicer /Trinity-ToolAce-SFT-splittextn<1K0 likes58 downloads1y agoHugging Face03datajuicer /the-pile-pubmed-central-refined-by-data-juicer The Pile -- PubMed Central (refined by Data-Juicer) A refined version of PubMed Central dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 83G). Dataset Information Number of samples: 2,694,860 (Keep ~86.96% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-pubmed-central-refined-by-data-juicer.texttext-generationn<1K2 likes50 downloads3y agoHugging Face04datajuicer /data-juicer-t2v-optimal-data-pool Data-Juicer Sandbox: A Comprehensive Suite for Multimodal Data-Model Co-development Project description The emergence of large-scale multi-modal generative models has drastically advanced artificial intelligence, introducing unprecedented levels of performance and functionality. However, optimizing these models remains challenging due to historically isolated paths of model-centric and data-centric developments, leading to suboptimal outcomes and inefficient resource… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/data-juicer-t2v-optimal-data-pool.texttext-to-videon<1K0 likes50 downloads2y agoHugging Face05datajuicer /RealMedConv Dataset Description The RealMedConv dataset consists of anonymized, real-world dialogues between licensed pharmacists and users seeking over-the-counter (OTC) medication advice. Each conversation is goal-oriented: the pharmacist gathers sufficient symptom information to provide an online and appropriate recommendation. Dialogues are typically concise, spanning 3–5 turns, reflecting the efficient and expert-driven nature of professional medical consultations. This dataset originates… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/RealMedConv.text1K<n<10K1 likes44 downloads11mo agoHugging Face06datajuicer /alpaca-cot-zh-refined-by-data-juicer Alpaca-CoT -- ZH (refined by Data-Juicer) A refined Chinese version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to fine-tune a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 18.7GB). Dataset Information Number of samples: 9,873,214 (Keep ~46.58% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-zh-refined-by-data-juicer.texttext-generationn<1K5 likes37 downloads3y agoHugging Face07datajuicer /the-pile-uspto-refined-by-data-juicer The Pile -- USPTO (refined by Data-Juicer) A refined version of USPTO dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 18G). Dataset Information Number of samples: 4,516,283 (Keep ~46.77% from the original dataset) Refining Recipe #… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-uspto-refined-by-data-juicer.texttext-generationn<1K0 likes28 downloads3y agoHugging Face08datajuicer /redpajama-arxiv-refined-by-data-juicer RedPajama -- ArXiv (refined by Data-Juicer) A refined version of ArXiv dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 85GB). Dataset Information Number of samples: 1,655,259 (Keep ~95.99% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-arxiv-refined-by-data-juicer.texttext-generationn<1K2 likes26 downloads3y agoHugging Face09datajuicer /redpajama-cc-2022-05-refined-by-data-juicer RedPajama -- CommonCrawl-2022-05 (refined by Data-Juicer) A refined version of CommonCrawl-2022-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 265GB). Dataset Information Number of samples: 42,648,496 (Keep ~45.34% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2022-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes25 downloads3y agoHugging Face10datajuicer /alpaca-cot-en-refined-by-data-juicer Alpaca-CoT -- EN (refined by Data-Juicer) A refined English version of Alpaca-CoT dataset by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to fine-tune a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 226GB). Dataset Information Number of samples: 72,855,345 (Keep ~54.48% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/alpaca-cot-en-refined-by-data-juicer.texttext-generationn<1K0 likes25 downloads3y agoHugging Face11datajuicer /redpajama-pile-stackexchange-refined-by-data-juicer RedPajama & The Pile -- StackExchange (refined by Data-Juicer) A refined version of StackExchange dataset in RedPajama & The Pile by Data-Juicer. Removing some "bad" samples from the original merged dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 71GB). Dataset Information Number of samples: 26,309,203 (Keep ~57.89% from the… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-pile-stackexchange-refined-by-data-juicer.texttext-generationn<1K0 likes20 downloads3y agoHugging Face12datajuicer /redpajama-cc-2019-30-refined-by-data-juicer RedPajama -- CommonCrawl-2019-30 (refined by Data-Juicer) A refined version of CommonCrawl-2019-30 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 240GB). Dataset Information Number of samples: 36,557,283 (Keep ~45.08% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2019-30-refined-by-data-juicer.tabulartext-generationn<1K0 likes19 downloads3y agoHugging Face13datajuicer /redpajama-wiki-refined-by-data-juicer RedPajama -- Wikipedia (refined by Data-Juicer) A refined version of Wikipedia dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 68GB). Dataset Information Number of samples: 26,990,659 (Keep ~90.47% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-wiki-refined-by-data-juicer.texttext-generationn<1K2 likes18 downloads3y agoHugging Face14datajuicer /the-pile-europarl-refined-by-data-juicer The Pile -- EuroParl (refined by Data-Juicer) A refined version of EuroParl dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.2GB). Dataset Information Number of samples: 61,601 (Keep ~88.23% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-europarl-refined-by-data-juicer.texttext-generationn<1K0 likes15 downloads3y agoHugging Face15datajuicer /redpajama-stack-code-refined-by-data-juicer RedPajama & TheStack -- Github Code (refined by Data-Juicer) A refined version of Github Code dataset in RedPajama & TheStack by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 232GB). Dataset Information Number of samples: 49,279,344 (Keep ~52.09% from the original… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-stack-code-refined-by-data-juicer.tabulartext-generationn<1K2 likes13 downloads3y agoHugging Face16datajuicer /the-pile-hackernews-refined-by-data-juicer The Pile -- HackerNews (refined by Data-Juicer) A refined version of HackerNews dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.8G). Dataset Information Number of samples: 371,331 (Keep ~99.55% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-hackernews-refined-by-data-juicer.texttext-generationn<1K0 likes13 downloads3y agoHugging Face17datajuicer /the-pile-philpaper-refined-by-data-juicer The Pile -- PhilPaper (refined by Data-Juicer) A refined version of PhilPaper dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 1.7GB). Dataset Information Number of samples: 29,117 (Keep ~88.82% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-philpaper-refined-by-data-juicer.texttext-generationn<1K0 likes12 downloads3y agoHugging Face18datajuicer /redpajama-cc-2020-05-refined-by-data-juicer RedPajama -- CommonCrawl-2020-05 (refined by Data-Juicer) A refined version of CommonCrawl-2020-05 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 297GB). Dataset Information Number of samples: 42,612,596 (Keep ~46.90% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2020-05-refined-by-data-juicer.tabulartext-generationn<1K0 likes12 downloads3y agoHugging Face19datajuicer /Trinity-ToolAce-RL-splittext1K<n<10K0 likes11 downloads1y agoHugging Face20datajuicer /redpajama-cc-2023-06-refined-by-data-juicer RedPajama -- CommonCrawl-2023-06 (refined by Data-Juicer) A refined version of CommonCrawl-2023-06 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 310GB). Dataset Information Number of samples: 50,643,699 (Keep ~45.46% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2023-06-refined-by-data-juicer.tabulartext-generationn<1K0 likes8 downloads3y agoHugging Face21datajuicer /the-pile-freelaw-refined-by-data-juicer The Pile -- FreeLaw (refined by Data-Juicer) A refined version of FreeLaw dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 45GB). Dataset Information Number of samples: 2,942,612 (Keep ~82.61% from the original dataset) Refining Recipe… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-freelaw-refined-by-data-juicer.texttext-generationn<1K0 likes8 downloads3y agoHugging Face22datajuicer /the-pile-nih-refined-by-data-juicer The Pile -- NIHExPorter (refined by Data-Juicer) A refined version of NIHExPorter dataset in The Pile by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 2.0G). Dataset Information Number of samples: 858,492 (Keep ~91.36% from the original dataset) Refining… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/the-pile-nih-refined-by-data-juicer.texttext-generationn<1K0 likes8 downloads3y agoHugging Face23datajuicer /redpajama-cc-2021-04-refined-by-data-juicer RedPajama -- CommonCrawl-2021-04 (refined by Data-Juicer) A refined version of CommonCrawl-2021-04 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 284GB). Dataset Information Number of samples: 44,724,752 (Keep ~45.23% from the original dataset)… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-cc-2021-04-refined-by-data-juicer.tabulartext-generationn<1K0 likes6 downloads3y agoHugging Face24datajuicer /redpajama-c4-refined-by-data-juicer RedPajama -- C4 (refined by Data-Juicer) A refined version of C4 dataset in RedPajama by Data-Juicer. Removing some "bad" samples from the original dataset to make it higher-quality. This dataset is usually used to pretrain a Large Language Model. Notice: Here is a small subset for previewing. The whole dataset is available here (About 832GB). Dataset Information Number of samples: 344,491,171 (Keep ~94.42% from the original dataset) Refining Recipe #… See the full description on the dataset page: https://huggingface.co/datasets/datajuicer/redpajama-c4-refined-by-data-juicer.texttext-generationn<1K1 likes4 downloads3y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.