datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
the_pile_deduplicatedcleaned_deduplicated_oscar
Dataset Card for "cleaned_deduplicated_oscar"
More Information needed
c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines
Dataset Card for "c4-pt-randMore35M-part04-deduplicated-128000-no-digit-split-mask-train-15003771-lines"
More Information needed
ai-music-deduplicated
AI Music Deduplicated
A large-scale collection of AI-generated music from five platforms: Mureka, Riffusion, Sonauto, Suno, and Udio. Each song includes the original audio file and its full platform metadata as a JSON sidecar.
Overview
Subset
Songs
Tar Files
Size
Audio Format
Source Platform
mureka
~312K
49
~981 GB
.mp3
Mureka
riffusion
~105K
14
~266 GB
.m4a
Riffusion
sonauto
~15K
2
~25 GB
.ogg
Sonauto
suno
~307K
65
~1.3 TB
.mp3
Suno
udio~126K
33
~642… See the full description on the dataset page: https://huggingface.co/datasets/ai-music/ai-music-deduplicated.fineweb_deduplicated
TL;DR
Fineweb is a popular and high quality open dataset. This dataset is a deduplicated version of Fineweb - removing rows with duplicate text, collecting counts.
Motivation
Fineweb is an open text dataset intended for training language models. It's one of the highest quality and most popular open datasets available. It has been produced by a reputable AI lab - HuggingFace and has been downloaded tens of thousands of times.
Fineweb dataset is 93.4 TB and has 15T… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/fineweb_deduplicated.EleutherAI_the_pile_deduplicatedSince The Pile was removed from the original site, I'm worried this dataset might be taken down too. Putting it here just in case.
Original repo: https://huggingface.co/datasets/EleutherAI/the_pile_deduplicated
mzansi-text-deduplicated
MzansiText — deduplicated release
This is a conservatively deduplicated release of MzansiText, a multilingual pretraining corpus for all eleven official South African languages.
Use the original release to reproduce the paper and its trained models. Use this release for new experiments where cross-source duplicate documents should be removed.
Dataset details
Splits: 3,744,654 train rows, 19,940 validation rows, and 19,784 test rows
Total: 3,784,378 rows
lang… See the full description on the dataset page: https://huggingface.co/datasets/uctnlp/mzansi-text-deduplicated.fineweb_2_500k_both_deduplicatedsharegpt-deduplicated
Dataset Card for Dataset Name
Dataset Description
Dataset Summary
This dataset is a deduplicated version of sharegpt4.
The deduplication process has two steps:
The literal duplicates (both input and outputs) are removed
The remaining (5749) instances are embedded with the SentenceTransformer library ("paraphrase-multilingual-mpnet-base-v2" model).
Then, we compute the cosine similarity among all the possible pairs, and consider paraphrases those pairs with a… See the full description on the dataset page: https://huggingface.co/datasets/CaterinaLac/sharegpt-deduplicated.OpenDebateEvidence-Deduplicated-Anonymized
Dataset Card for OpenDebateEvidence-Deduplicated (Anonymized)
Debate evidence from collegiate and high school competitions, semantically
deduplicated, with all debater-identifying columns removed.
This is the semantically deduplicated companion to
OpenDebateEvidence-Anonymized.
Where the parent dataset contains every piece of evidence as used in every round,
this version collapses repeated use of the same evidence into single records,
making it substantially smaller and better… See the full description on the dataset page: https://huggingface.co/datasets/Hellisotherpeople/OpenDebateEvidence-Deduplicated-Anonymized.bookcorpus_deduplicated
Dataset Card for "bookcorpus_deduplicated"
Dataset Summary
This is a deduplicated version of the original Book Corpus dataset.
The Book Corpus (Zhu et al., 2015), which was used to train popular models such as BERT, has a substantial amount of exact-duplicate documents according to Bandy and Vincent (2021)
Bandy and Vincent (2021) find that thousands of books in BookCorpus are duplicated, with only 7,185 unique books out of 11,038 total.
Effect of deduplication
Num of… See the full description on the dataset page: https://huggingface.co/datasets/Saibo-creator/bookcorpus_deduplicated.nl2sql-deduplicated
NL2SQL Deduplicated Training Dataset
A curated and deduplicated Text-to-SQL training dataset with 683,015 unique examples from 4 high-quality sources.
📊 Dataset Summary
Total Examples: 683,015 unique question-SQL pairs
Sources: Spider, SQaLe, Gretel Synthetic, SQL-Create-Context
Deduplication Strategy: Input-only (question-based) with conflict resolution via quality priority
Conflicts Resolved: 2,238 cases where same question had different SQL
SQL Dialect: Standard SQL… See the full description on the dataset page: https://huggingface.co/datasets/AsadIsmail/nl2sql-deduplicated.corpus_1_embedded_deduplicated
Dataset Card for "corpus_1_embedded_deduplicated"
More Information needed
WildChat-4.8M-EN-Semantic-DeduplicatedThese are a collection of deduplicated English prompts shorter than ~2000 tokens taken from WildChat-4.8M.
Initially, many non-English prompts were removed by validating that at least 80% of each prompt was composed of characters from the English alphabet.
Then, a HashSet was used to directly deduplicate identical prompts (ignoring whitespace and punctuation).
Then, MinHash was used to further conservatively deduplicate.
Finally, Qwen3-8B-Embedding was used to generate embeddings on all… See the full description on the dataset page: https://huggingface.co/datasets/MasonMac/WildChat-4.8M-EN-Semantic-Deduplicated.master-ebook-library-deduplicateddata_deduplicated_part03
Dataset Card for "data_deduplicated_part03"
More Information needed
data_deduplicated_part04
Dataset Card for "data_deduplicated_part04"
More Information needed
Human-Like-DPO-Dataset_deduplicated_and_duplicatedclevr-math-deduplicateddata_deduplicated_part01
Dataset Card for "data_deduplicated_part01"
More Information needed
data_deduplicated_part02
Dataset Card for "data_deduplicated_part02"
More Information needed
finer_FinLoRA_deduplicatedgenetic_instruct_deduplicatedngxson_MiniThinky_v1_deduplicated_11_percentWildChat-4M-English-Semantic-Deduplicated
ALERT: This dataset has been superseded by https://huggingface.co/datasets/MasonMac/WildChat-4.8M-EN-Semantic-Deduplicated
This is a dataset of all English prompts from the WildChat-4M dataset. It was created by checking that at least 80% non-punctuation characters were in the English alphabet (to remove some more non-English entries). Then, it was deduplicated (ignoring punctuation/whitespace differences) by collecting to a HashSet. Another deduplication pass was done with MinHash. And… See the full description on the dataset page: https://huggingface.co/datasets/MasonMac/WildChat-4M-English-Semantic-Deduplicated.data_deduplicated_part05
Dataset Card for "data_deduplicated_part05"
More Information needed
parallel-wikimatrix-deduplicatedtweets_dataset_jan_feb_big_deduplicatedglaive-fc-deduplicated-train-test-valdeduplicated_dapo_dataset
Deduplicated DAPO Dataset
This dataset is created by modifying this dataset: DAPO.
The main change is: we only keep the unique prompts by filtering the dataset. The original DAPO dataset has 1.79M rows, but only 17398 unique prompts. We make a new dataset after deduplication, for ease of reproducing the results in our paper. For more information about our paper, please checkout the project page.
For more information about the details of this dataset, please see the original paper… See the full description on the dataset page: https://huggingface.co/datasets/ftajwar/deduplicated_dapo_dataset.
