datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
AIMEfrom datasets import load_dataset
dataset = load_dataset('disco-eth/AIME')
AIME: AI Music Evaluation Dataset
The AIME dataset contains 6,000 audio tracks generated by 12 music generation models in addition to 500 tracks from MTG-Jamendo.
The prompts used to generate music are combinations of representative and diverse tags from the MTG-Jamendo dataset.
The AIME dataset consists of two subsets. The AIME audio dataset and the AIME survey dataset.
The dataset contains the following… See the full description on the dataset page: https://huggingface.co/datasets/disco-eth/AIME.ai-music-deduplicated
AI Music Deduplicated
A large-scale collection of AI-generated music from five platforms: Mureka, Riffusion, Sonauto, Suno, and Udio. Each song includes the original audio file and its full platform metadata as a JSON sidecar.
Overview
Subset
Songs
Tar Files
Size
Audio Format
Source Platform
mureka
~312K
49
~981 GB
.mp3
Mureka
riffusion
~105K
14
~266 GB
.m4a
Riffusion
sonauto
~15K
2
~25 GB
.ogg
Sonauto
suno
~307K
65
~1.3 TB
.mp3
Suno
udio~126K
33
~642… See the full description on the dataset page: https://huggingface.co/datasets/ai-music/ai-music-deduplicated.ai-mAIME_Datasetai_music_large
AI/Human Music (Large variant)
A dataset that comprises of both AI-generated music and human-composed music.
This is the "large" variant of the dataset, which is around 70GiB in size. It contains 10,000 audio files from human and 10,000 audio files from AI. The distribution is: $256$ are from SunoCaps, $4,872$ are from Udio, and $4,872$ are from MusicSet.
Data sources for this dataset:
https://huggingface.co/datasets/blanchon/udio_dataset… See the full description on the dataset page: https://huggingface.co/datasets/SleepyJesse/ai_music_large.aimeuai-generated-songsaimodelsai-generated-songs2ai-m-2ai-generated-songs3ai_music_smallComplexly-cluster-thresh0_75-conf-thresh0_9ai_music_tinyai-meeting-samplesComplexly-multi-speaker-TrainingData-speaker-assignedComplexly-Multi-Speaker-filteredaimlbdai-ModelsmichaleAIMusicmj3tupacai-voices-enhanced
