CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01EleutherAI /the_pile_deduplicatedtext100M<n<1B117 likes15k downloads4y agoHugging Face02ClementRomac /cleaned_deduplicated_oscar Dataset Card for "cleaned_deduplicated_oscar" More Information needed text100M<n<1B0 likes6.9k downloads3y agoHugging Face03tgsc /c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines" More Information needed text10M<n<100M1 likes4k downloads3y agoHugging Face04ai-music /ai-music-deduplicated AI Music Deduplicated A large-scale collection of AI-generated music from five platforms: Mureka, Riffusion, Sonauto, Suno, and Udio. Each song includes the original audio file and its full platform metadata as a JSON sidecar. Overview Subset Songs Tar Files Size Audio Format Source Platform mureka ~312K 49 ~981 GB .mp3 Mureka riffusion ~105K 14 ~266 GB .m4a Riffusion sonauto ~15K 2 ~25 GB .ogg Sonauto suno ~307K 65 ~1.3 TB .mp3 Suno udio~126K 33 ~642… See the full description on the dataset page: https://huggingface.co/datasets/ai-music/ai-music-deduplicated.audioaudio-classification100K<n<1M4 likes922 downloads7mo agoHugging Face05Salesforce /fineweb_deduplicated TL;DR Fineweb is a popular and high quality open dataset. This dataset is a deduplicated version of Fineweb - removing rows with duplicate text, collecting counts. Motivation Fineweb is an open text dataset intended for training language models. It's one of the highest quality and most popular open datasets available. It has been produced by a reputable AI lab - HuggingFace and has been downloaded tens of thousands of times. Fineweb dataset is 93.4 TB and has 15T… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/fineweb_deduplicated.tabular1B<n<10B41 likes763 downloads2y agoHugging Face06gmongaras /EleutherAI_the_pile_deduplicatedSince The Pile was removed from the original site, I'm worried this dataset might be taken down too. Putting it here just in case. Original repo: https://huggingface.co/datasets/EleutherAI/the_pile_deduplicated text100M<n<1B4 likes647 downloads3y agoHugging Face07uctnlp /mzansi-text-deduplicated MzansiText — deduplicated release This is a conservatively deduplicated release of MzansiText, a multilingual pretraining corpus for all eleven official South African languages. Use the original release to reproduce the paper and its trained models. Use this release for new experiments where cross-source duplicate documents should be removed. Dataset details Splits: 3,744,654 train rows, 19,940 validation rows, and 19,784 test rows Total: 3,784,378 rows lang… See the full description on the dataset page: https://huggingface.co/datasets/uctnlp/mzansi-text-deduplicated.text1M<n<10M1 likes538 downloads1mo agoHugging Face08JupiterLLM /fineweb_2_500k_both_deduplicatedtabular1M<n<10M0 likes452 downloads1y agoHugging Face09CaterinaLac /sharegpt-deduplicated Dataset Card for Dataset Name Dataset Description Dataset Summary This dataset is a deduplicated version of sharegpt4. The deduplication process has two steps: The literal duplicates (both input and outputs) are removed The remaining (5749) instances are embedded with the SentenceTransformer library ("paraphrase-multilingual-mpnet-base-v2" model). Then, we compute the cosine similarity among all the possible pairs, and consider paraphrases those pairs with a… See the full description on the dataset page: https://huggingface.co/datasets/CaterinaLac/sharegpt-deduplicated.text1K<n<10K1 likes201 downloads3y agoHugging Face10Hellisotherpeople /OpenDebateEvidence-Deduplicated-Anonymized Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized) Debate evidence from collegiate and high school competitions, semantically deduplicated, with all debater-identifying columns removed. This is the semantically deduplicated companion to OpenDebateEvidence-Anonymized. Where the parent dataset contains every piece of evidence as used in every round, this version collapses repeated use of the same evidence into single records, making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.tabulartext-generation100K<n<1M0 likes167 downloads2mo agoHugging Face11Saibo-creator /bookcorpus_deduplicated Dataset Card for "bookcorpus_deduplicated" Dataset Summary This is a deduplicated version of the original Book Corpus dataset. The Book Corpus (Zhu et al., 2015), which was used to train popular models such as BERT, has a substantial amount of exact-duplicate documents according to Bandy and Vincent (2021) Bandy and Vincent (2021) find that thousands of books in BookCorpus are duplicated, with only 7,185 unique books out of 11,038 total. Effect of deduplication Num of… See the full description on the dataset page: https://huggingface.co/datasets/Saibo-creator/bookcorpus_deduplicated.text10M<n<100M2 likes110 downloads4y agoHugging Face12AsadIsmail /nl2sql-deduplicated NL2SQL Deduplicated Training Dataset A curated and deduplicated Text-to-SQL training dataset with 683,015 unique examples from 4 high-quality sources. 📊 Dataset Summary Total Examples: 683,015 unique question-SQL pairs Sources: Spider, SQaLe, Gretel Synthetic, SQL-Create-Context Deduplication Strategy: Input-only (question-based) with conflict resolution via quality priority Conflicts Resolved: 2,238 cases where same question had different SQL SQL Dialect: Standard SQL… See the full description on the dataset page: https://huggingface.co/datasets/AsadIsmail/nl2sql-deduplicated.texttext-generation1K<n<10K0 likes110 downloads10mo agoHugging Face13HydraLM /corpus_1_embedded_deduplicated Dataset Card for "corpus_1_embedded_deduplicated" More Information needed text1M<n<10M0 likes103 downloads3y agoHugging Face14MasonMac /WildChat-4.8M-EN-Semantic-DeduplicatedThese are a collection of deduplicated English prompts shorter than ~2000 tokens taken from WildChat-4.8M. Initially, many non-English prompts were removed by validating that at least 80% of each prompt was composed of characters from the English alphabet. Then, a HashSet was used to directly deduplicate identical prompts (ignoring whitespace and punctuation). Then, MinHash was used to further conservatively deduplicate. Finally, Qwen3-8B-Embedding was used to generate embeddings on all… See the full description on the dataset page: https://huggingface.co/datasets/MasonMac/WildChat-4.8M-EN-Semantic-Deduplicated.text1M<n<10M0 likes92 downloads1y agoHugging Face15RedMoon1002 /master-ebook-library-deduplicateddocumentn<1K0 likes74 downloads1y agoHugging Face16phanvancongthanh /data_deduplicated_part03 Dataset Card for "data_deduplicated_part03" More Information needed text100M<n<1B0 likes61 downloads3y agoHugging Face17phanvancongthanh /data_deduplicated_part04 Dataset Card for "data_deduplicated_part04" More Information needed text100M<n<1B0 likes59 downloads3y agoHugging Face18NickyNicky /Human-Like-DPO-Dataset_deduplicated_and_duplicatedtext10K<n<100K1 likes56 downloads2y agoHugging Face19andito /clevr-math-deduplicatedimage10K<n<100K1 likes54 downloads1y agoHugging Face20phanvancongthanh /data_deduplicated_part01 Dataset Card for "data_deduplicated_part01" More Information needed text10M<n<100M0 likes53 downloads3y agoHugging Face21phanvancongthanh /data_deduplicated_part02 Dataset Card for "data_deduplicated_part02" More Information needed text100M<n<1B0 likes52 downloads3y agoHugging Face22stalaei /finer_FinLoRA_deduplicatedtext100K<n<1M0 likes43 downloads11mo agoHugging Face23amal-abed /genetic_instruct_deduplicatedtext100K<n<1M0 likes42 downloads1y agoHugging Face24NickyNicky /ngxson_MiniThinky_v1_deduplicated_11_percenttexttext-generation10K<n<100K3 likes40 downloads2y agoHugging Face25MasonMac /WildChat-4M-English-Semantic-Deduplicated ALERT: This dataset has been superseded by https://huggingface.co/datasets/MasonMac/WildChat-4.8M-EN-Semantic-Deduplicated This is a dataset of all English prompts from the WildChat-4M dataset. It was created by checking that at least 80% non-punctuation characters were in the English alphabet (to remove some more non-English entries). Then, it was deduplicated (ignoring punctuation/whitespace differences) by collecting to a HashSet. Another deduplication pass was done with MinHash. And… See the full description on the dataset page: https://huggingface.co/datasets/MasonMac/WildChat-4M-English-Semantic-Deduplicated.text100K<n<1M0 likes40 downloads1y agoHugging Face26phanvancongthanh /data_deduplicated_part05 Dataset Card for "data_deduplicated_part05" More Information needed text10M<n<100M0 likes35 downloads3y agoHugging Face27gowitheflow /parallel-wikimatrix-deduplicatedtext1M<n<10M2 likes35 downloads2y agoHugging Face28SinclairSchneider /tweets_dataset_jan_feb_big_deduplicatedtabular10M<n<100M0 likes35 downloads5mo agoHugging Face29shizi1011 /glaive-fc-deduplicated-train-test-valtext10K<n<100K0 likes33 downloads2y agoHugging Face30ftajwar /deduplicated_dapo_dataset Deduplicated DAPO Dataset This dataset is created by modifying this dataset: DAPO. The main change is: we only keep the unique prompts by filtering the dataset. The original DAPO dataset has 1.79M rows, but only 17398 unique prompts. We make a new dataset after deduplication, for ease of reproducing the results in our paper. For more information about our paper, please checkout the project page. For more information about the details of this dataset, please see the original paper… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/deduplicated_dapo_dataset.text10K<n<100K1 likes33 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.