datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
excavationpro-music-stream
Excavationpro public music stream (160 kbps)
Owner / artist: Justin Helmer · Excavationpro · LightfatherPolicy: Own-work only. Public discovery streams (not DistroKid-dependent).Lattice signature: Δ9Φ963-PUBLIC-MUSIC-STREAM-v1
Listen
https://deepseekoracle.github.io/Excavationpro/excavationpro-listen.html
http://asiancoastline.com/ (custom domain music portal)
Layout
Path
Role
stream/<sha256>.mp3
Flat 160k streams (~first 10k −… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/excavationpro-music-stream.kossuth-musicmidi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.music-arena-dataset
Music Arena Dataset
This is the official dataset from Music Arena, an open platform for evaluating text-to-music (TTM) models.
How to Download (Recommended Method)
The most reliable way to get a complete local copy of all files, including the entire audio collection, is to clone the repository directly using Git. This method is ideal for offline access and workflows that require direct file manipulation.
Note: This repository uses Git LFS (Large File Storage) to… See the full description on the dataset page: https://huggingface.co/datasets/music-arena/music-arena-dataset.Poster_Music_festivalfree-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.music-ai-human-test-audio
Interpretable AI and Human Music Evaluation Archive
Research audio and versioned experiment outputs for an English graduation thesis.
The audio archive is incomplete. Completed experiments and verified partial
audio publications must not be confused with whole-project delivery completion.
No blanket license is assigned to this mixed-source archive.
Completed experiments and thesis
The BC extension, expanded YuE Native30 evaluation, locked YuE Native30 scoring… See the full description on the dataset page: https://huggingface.co/datasets/EZMONYI/music-ai-human-test-audio.Music-AVQAai-musicmusic
music
A large-scale music dataset containing artist, release and track names + URLs.
Sources
Source
Rows
soundcloud
200M
discogs
178M
applemusic
104M
lastfm
102M
deezer
90M
bandcamp
49M
musicbrainz
24M
youtube-videos
10M
youtube
7M
metal-archives
5M
vgmdb
2M
film-tv
3M
newgrounds
1M
onlineradiobox
466K
beatport
454K
Total
776M
Please note that this dataset has not yet been deduplicated across sources. A cross-source… See the full description on the dataset page: https://huggingface.co/datasets/fairygaze/music.musicavideo-acervosuno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.reamixed-project-files
reamixed_project_files
captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
Music-POSTPROCESS-509ab05efree-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.MusicNetMusic-POSTPROCESS-32eadf7eMusicCaps
Dataset Card for MusicCaps
Dataset Summary
The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g.,
"A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.music-off-policy-evaluation-benchmark
Music Off-Policy Evaluation Dataset
Music Off-Policy Evaluation Dataset is a dataset designed for Off-Policy Evaluation (OPE) research. It contains logged interactions from the home page of Amazon Music.
Use cases:
Benchmarking OPE estimators
Evaluating counterfactual ranking policies offline
License
Music Off-Policy Evaluation Benchmark © 2026 by Amazon is licensed under Creative Commons Attribution-NonCommercial 4.0 International.… See the full description on the dataset page: https://huggingface.co/datasets/amazon/music-off-policy-evaluation-benchmark.suno_musicalitymusicfree-music-archive-small
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-small.free-music-archive-large
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-large.MusicPile🌐 DemoPage | 🤗SFT Dataset | 🤗 Benchmark | 📖 arXiv | 💻 Code | 🤖 Chat Model | 🤖 Base Model
Dataset Card for MusicPile
MusicPile is the first pretraining corpus for developing musical abilities in large language models.
It has 5.17M samples and approximately 4.16B tokens, including web-crawled corpora, encyclopedias, music books, youtube music captions, musical pieces in abc notation, math content, and code.
You can easily load it:from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/m-a-p/MusicPile.music_genres
Dataset Card for "music_genres"
More Information needed
musickisida-music1emotion-probes-raw-activationsminimax-music-3-datasetMiniMax Music 3.0 Dataset
A large synthetic MiniMax Music research dataset by Angelware Research
8,681 tracks generated with MiniMax Music 3.0 for audio analysis, benchmarking, provenance research, and AI-music detection.
At a glance
Generated with MiniMax Music 3.0. Audio is preserved exactly as received, including embedded AIGC provenance tags where present.
Collection statistic
Value
Tracks
8,681
Total duration… See the full description on the dataset page: https://huggingface.co/datasets/AngelSoftware/minimax-music-3-dataset.
