datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
excavationpro-music-stream
Excavationpro public music stream (160 kbps)
Owner / artist: Justin Helmer · Excavationpro · LightfatherPolicy: Own-work only. Public discovery streams (not DistroKid-dependent).Lattice signature: Δ9Φ963-PUBLIC-MUSIC-STREAM-v1
Listen
https://deepseekoracle.github.io/Excavationpro/excavationpro-listen.html
http://asiancoastline.com/ (custom domain music portal)
Layout
Path
Role
stream/<sha256>.mp3
Flat 160k streams (~first 10k −… See the full description on the dataset page: https://huggingface.co/datasets/DeepSeekOracle/excavationpro-music-stream.music-arena-dataset
Music Arena Dataset
This is the official dataset from Music Arena, an open platform for evaluating text-to-music (TTM) models.
How to Download (Recommended Method)
The most reliable way to get a complete local copy of all files, including the entire audio collection, is to clone the repository directly using Git. This method is ideal for offline access and workflows that require direct file manipulation.
Note: This repository uses Git LFS (Large File Storage) to… See the full description on the dataset page: https://huggingface.co/datasets/music-arena/music-arena-dataset.music-ai-human-test-audio
Interpretable AI and Human Music Evaluation Archive
Research audio and versioned experiment outputs for an English graduation thesis.
The audio archive is incomplete. Completed experiments and verified partial
audio publications must not be confused with whole-project delivery completion.
No blanket license is assigned to this mixed-source archive.
Completed experiments and thesis
The BC extension, expanded YuE Native30 evaluation, locked YuE Native30 scoring… See the full description on the dataset page: https://huggingface.co/datasets/EZMONYI/music-ai-human-test-audio.free-music-archive-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-full.musicavideo-acervosuno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.captioned-ai-music-snippets
Dataset Overview
A collection of short audio snippets (3–30 seconds) extracted from publicly shared Suno‑generated songs and captioned with Gemini Flash 2.0. Designed specifically to train and evaluate audio captioning models.
Source
Clips are randomly cut from the songs referenced in the nyuuzyou/suno repository.
Captioning
All excerpts have been annotated using Gemini Flash 2.0 for high‑quality, human‑readable audio descriptions.
License
Apache 2.0
free-music-archive-medium
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-medium.Music-POSTPROCESS-32eadf7eMusic-POSTPROCESS-509ab05eMusicNetmusicVietnamese-Traditional-Musicmusic_genres
Dataset Card for "music_genres"
More Information needed
free-music-archive-small
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-small.free-music-archive-large
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-large.minimax-music-3-datasetMiniMax Music 3.0 Dataset
A large synthetic MiniMax Music research dataset by Angelware Research
8,681 tracks generated with MiniMax Music 3.0 for audio analysis, benchmarking, provenance research, and AI-music detection.
At a glance
Generated with MiniMax Music 3.0. Audio is preserved exactly as received, including embedded AIGC provenance tags where present.
Collection statistic
Value
Tracks
8,681
Total duration… See the full description on the dataset page: https://huggingface.co/datasets/AngelSoftware/minimax-music-3-dataset.earica_music_finalai-music-deduplicated
AI Music Deduplicated
A large-scale collection of AI-generated music from five platforms: Mureka, Riffusion, Sonauto, Suno, and Udio. Each song includes the original audio file and its full platform metadata as a JSON sidecar.
Overview
Subset
Songs
Tar Files
Size
Audio Format
Source Platform
mureka
~312K
49
~981 GB
.mp3
Mureka
riffusion
~105K
14
~266 GB
.m4a
Riffusion
sonauto
~15K
2
~25 GB
.ogg
Sonauto
suno
~307K
65
~1.3 TB
.mp3
Suno
udio~126K
33
~642… See the full description on the dataset page: https://huggingface.co/datasets/ai-music/ai-music-deduplicated.MusicCapsMusic-POSTPROCESS-0091656dMusic-POSTPROCESS-1ce3651afree-music-archive-commercial-16khz-full
FMA: A Dataset for Music Analysis
Michaël Defferrard, Kirell Benzi, Pierre Vandergheynst, Xavier Bresson.
International Society for Music Information Retrieval Conference (ISMIR), 2017.
We introduce the Free Music Archive (FMA), an open and easily accessible dataset suitable for evaluating several tasks in MIR, a field concerned with browsing, searching, and organizing large music collections. The community's growing interest in feature and end-to-end learning is however restrained… See the full description on the dataset page: https://huggingface.co/datasets/benjamin-paine/free-music-archive-commercial-16khz-full.Music-POSTPROCESS-8c669927Music-POSTPROCESS-2f8535b8Speech_with_Music_v2free-music-archive-retrieval
FMAR: A Dataset for Robust Song Identification
Authors: Ryan Lee, Yi-Chieh Chiu, Abhir Karande, Ayush Goyal, Harrison Pearl, Matthew Hong, Spencer Cobb
Overview
To improve copyright infringement detection, we introduce Free-Music-Archive-Retrieval (FMAR), a structured dataset designed to test a model's capability to identify songs based on 5-second clips, or queries. We create adversarial queries to replicate common strategies to evade copyright infringement detectors… See the full description on the dataset page: https://huggingface.co/datasets/ml-ryanlee/free-music-archive-retrieval.Flying-Music
音乐云端存储仓库 (Private)
项目说明
本仓库用于本地播放器的云端同步,包含音乐文件、封面、歌词及用户元数据。
使用协议 (License)
本项目采用 CC BY-NC-ND 4.0 (署名-非商业性使用-禁止演绎 4.0 国际) 协议。
允许:下载、播放、备份。
禁止:任何形式的商业盈利行为、未经授权的二次修改发布。
免责声明
本仓库仅供个人学习及非营利性交流使用。
用户上传的内容版权归原作者所有,上传者需确保拥有分发权限。
若发现侵权内容,请联系删除。
ai_music_large
AI/Human Music (Large variant)
A dataset that comprises of both AI-generated music and human-composed music.
This is the "large" variant of the dataset, which is around 70GiB in size. It contains 10,000 audio files from human and 10,000 audio files from AI. The distribution is: $256$ are from SunoCaps, $4,872$ are from Udio, and $4,872$ are from MusicSet.
Data sources for this dataset:
https://huggingface.co/datasets/blanchon/udio_dataset… See the full description on the dataset page: https://huggingface.co/datasets/SleepyJesse/ai_music_large.musicai-background-music-audio-llm-benchmark
Does Background Music Matter to Speech in Pre-trained Language Models
The completed September 2026 study covers 8 model families, 55 instrumental recordings, and 10 evaluation settings. It studies how adding background music to the same spoken question changes model responses.
Latest release and artifact guide
Technical report PDF
Complete LaTeX project
LaTeX GitHub repository
Matrices, figures, and supporting data
Regenerated speech and mixtures: 550 archives / 250,800… See the full description on the dataset page: https://huggingface.co/datasets/Elfsong/musicai-background-music-audio-llm-benchmark.
