datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Music-AVQAMusicBench
MusicBench Dataset
The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model. MusicBench is based on the MusicCaps dataset, which it expands from 5,521 samples to 52,768 training and 400 test samples!
Dataset Details
MusicBench expands MusicCaps by:
Including music features of chords, beats, tempo, and key that are extracted from the audio.
Describing these music… See the full description on the dataset page: https://huggingface.co/datasets/amaai-lab/MusicBench.midi-classical-music-toio-json
MIDI Classical Music
drengskapur/midi-classical-musicのデータセットをtoioの soundコマンドで再生しやすいように以下のフォーマットのjsonに変換したデータを含めたデータセット
data format
[
{
"track_name": "ALBENIZ: Aragon Op 47/6",
"priority": 1,
"notes": [
{
"note_number": 77,
"start_time_ms": 0,
"duration_units": 26
},
{
},
},
{
"track_name": "apurdam@pcug.org.au",
"priority": 2,
"notes": [
{
"note_number": 53,
"start_time_ms": 0… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/midi-classical-music-toio-json.music2chords_v2music-quiz-scoresMusic
Made by herza For APP.t
musical-instrumentsmusic-reasoning-benchmark
Procedural Music Reasoning Benchmark
This benchmark was generated with
Danila-Pechenev/procedural-music-reasoning.
Benchmark version: v0.4.3.
Generator version: 0.4.2.
It contains balanced examples from two implemented music reasoning task
families:
pitch_interval_reasoning
chord_roman_reasoning
Configurations
Configuration
Examples per mode
Examples per split
Total examples
n16
16
256
768
n32
32
512
1536
n64 (default)
64
1024
3072
n128
128
2048… See the full description on the dataset page: https://huggingface.co/datasets/dpechenev/music-reasoning-benchmark.audio-music-mir-post-public
audio-music-mir-post-public
Music information retrieval and tagging annotations: genre (FMA), instrument family (NSynth × 3, Medley-solos-DB), social tags (MagnaTagATune via LLARK), and large-scale Music4All metadata. Foundation for music understanding heads in audio LLMs.
Audio is not bundled in this repo. See download.sh and per-dataset data/<name>.info.json for the fetch recipe; run postlink_audio.py after fetching to rewrite the JSONL audio_path fields with absolute local… See the full description on the dataset page: https://huggingface.co/datasets/vhands/audio-music-mir-post-public.music2chordsgod_level_music_producer_dataset
God-Level Music Producer Dataset
The most advanced open dataset for training LLMs to become elite music producers across Rap, Crunk, East Coast Boom Bap, West Coast G-Funk, and Dubstep.
This dataset contains 9,941 high-quality examples (with plans for expansion) of god-level reasoning and practical workflows in:
Beat Creation (full arrangement from scratch)
Virtual Instrument Mastery (Serum, Massive X, Omnisphere, Kontakt, etc.)
Plugin Expertise (Auto-Tune, iZotope Ozone/RX… See the full description on the dataset page: https://huggingface.co/datasets/11-47/god_level_music_producer_dataset.music_datasetsoundstock.com-music-genres-taxonomy
SoundStock Music Genres Taxonomy
A large, structured, and extensible music genre taxonomy dataset designed for music tagging, classification, search, recommendation systems, and audio / music machine learning workflows.
This dataset provides a hierarchical view of music genres, including root genres, subgenres, and expanded variants (style, era, region, and fusion), with stable IDs suitable for long-term use in production systems.
📊 Dataset Overview
1,600+ genres… See the full description on the dataset page: https://huggingface.co/datasets/SoundStock/soundstock.com-music-genres-taxonomy.MuseCraft-Music
🎵 MuseCraft-Music: Million Lyrics Chat Dataset
Train AI models to generate emotionally-rich, contextually-aware lyrics with 992K+ conversation pairs
🚀 Quick Start • 📊 Dataset Info • 💻 Usage • 🏷️ Citation
🎯 What is MuseCraft-Music?
MuseCraft-Music is a comprehensive dataset containing 992,246 conversation pairs designed specifically for training AI models to generate high-quality, emotionally-aware lyrics. Perfect for fine-tuning language models like… See the full description on the dataset page: https://huggingface.co/datasets/alvanalrakib/MuseCraft-Music.MusicBench
MusicBench Dataset
The MusicBench dataset is a music audio-text pair dataset that was designed for text-to-music generation purpose and released along with Mustango text-to-music model. MusicBench is based on the MusicCaps dataset, which it expands from 5,521 samples to 52,768 training and 400 test samples!
Dataset Details
MusicBench expands MusicCaps by:
Including music features of chords, beats, tempo, and key that are extracted from the audio.
Describing these music… See the full description on the dataset page: https://huggingface.co/datasets/Z873bliwf988hj/MusicBench.japanese-music-emotion
japanese music emotion
Music2Emotionを使って主に日本の音楽の感情分析を行ったデータセット
分析されたデータは以下のようなフォーマットのjsonlになっています。
{
"index": 1,
"video_id": "xxxxxxxxx",
"title": "sampleTitle",
"channel": "sample channel name",
"url": "video url",
"download_date": "2025-02-24T00:19:22.334487",
"predicted_moods": [
"energetic",
"fast",
"fun",
"funny",
"groovy",
"happy",
"holiday",
"love",
"party",
"positive"… See the full description on the dataset page: https://huggingface.co/datasets/ayousanz/japanese-music-emotion.uk-live-music-blog-corpus
UK Live Music Blog & Guides Corpus
109 long-form articles on the UK live music industry, published by GigXchange under CC BY 4.0. Written by working musicians and venue operators — not a content farm.
Overview
Metric
Value
Articles
109
Total words
305,841
Avg words/article
2,806
FAQ pairs
761
Topics
9
Date range
2026-03-01 to 2026-08-09
Language
British English (en-GB)
Domain
UK live music booking, fees, contracts, venues, city scenes… See the full description on the dataset page: https://huggingface.co/datasets/gigxchange/uk-live-music-blog-corpus.inspector-roofing-music-proof-stack
Inspector Roofing Music Proof Stack
This folder documents the public music-distribution evidence for the Inspector Roofing artist/creative media layer inside the Total Market Authority Scorecard system.
Purpose
The purpose of this evidence package is to document a real, public creative-media footprint connected to the Inspector Roofing brand without overstating it.
Safe description:
Inspector Roofing is a roofing-themed music/creative media artist name used for… See the full description on the dataset page: https://huggingface.co/datasets/InspectorRoofing/inspector-roofing-music-proof-stack.music-crs-listwise-train-shuffled
music-crs-listwise-train-shuffled
Training data for listwise reranking in music conversational recommendation (RecSys 2026 challenge).
Key feature
Candidate tracks are randomly shuffled before building each training example, preventing the model from learning the identity shortcut (outputting original order).
Format
Each line is a JSON object with messages (system/user/assistant chat turns) and metadata:
session_id, turn_number: conversation… See the full description on the dataset page: https://huggingface.co/datasets/shuaih777/music-crs-listwise-train-shuffled.music-generator-ai-agent
Music Generator Agent Meta and Traffic Dataset in AI Agent Marketplace | AI Agent Directory | AI Agent Index from DeepNLP
This dataset is collected from AI Agent Marketplace Index and Directory at http://www.deepnlp.org, which contains AI Agents's meta information such as agent's name, website, description, as well as the monthly updated Web performance metrics, including Google,Bing average search ranking positions, Github Stars, Arxiv References, etc.
The dataset is helpful for AI… See the full description on the dataset page: https://huggingface.co/datasets/DeepNLP/music-generator-ai-agent.CoT_Music_Production_DAW
Welcome to CoT_Music_Production_DAW, an open-source dataset (MIT licensed) featuring 7,000 expertly crafted Q&A pairs designed to train AI in the world of electronic music production using FL Studio and Ableton Live. This comprehensive resource covers everything from general DAW fundamentals—like interface navigation, recording, mixing, and project management—to in-depth specifics for FL Studio and Ableton Live, advanced production techniques, and even the nitty-gritty of music theory and… See the full description on the dataset page: https://huggingface.co/datasets/mattwesney/CoT_Music_Production_DAW.musicdatajson
Lakh MIDI Dataset — Fully Cleaned & Structured JSON (44,129 files)
This dataset is a fully cleaned and normalized version of the Lakh MIDI Dataset.Each MIDI file has been parsed, validated, and converted into a consistent JSON structure with detailed musical metadata.
✔ 44,129 cleaned JSON files✔ instrument programs✔ notes, durations, velocities✔ tempo curves✔ key signatures✔ time signatures✔ tracks & channels✔ structural markers✔ file-level metadata✔ consistent schema across all… See the full description on the dataset page: https://huggingface.co/datasets/YoloMG/musicdatajson.music-hashmusicology-annotations
Signal Dat — Musicology Annotations Sample
50-track sample from the Signal Dat dataset: human-annotated musicology records for training generative audio AI, music recommendation, and MIR models.
Full dataset (710+ tracks) available at signaldat.com
What makes this different
Most audio datasets provide algorithm-extracted features (BPM, key, MFCC). Signal Dat provides human expert annotations — a trained musicologist listens to each track and documents what a… See the full description on the dataset page: https://huggingface.co/datasets/signal-dat/musicology-annotations.upsi_ejournal_of_musicSynthetic-Musical-Instrumentsmusic-crs-state-extractor-datamusiciwant-sensory-music-analysis
Music I Want — Sensory Music Analysis (1,000-song sample)
A free, openly-licensed sample of the Music I Want catalog: songs scored on the
dimensions that decide how music feels. It is the independent answer to the
audio-features data that disappeared when Spotify deprecated its Audio Features
API in November 2024 — and it carries fields nothing else does.
sample-1000.csv / sample-1000.json — 1,000 songs, evenly sampled across the catalog's eras and intensity range.
Full… See the full description on the dataset page: https://huggingface.co/datasets/agreenbox/musiciwant-sensory-music-analysis.musicalitybench_annotatemy-music-dataset
Purpose
This dataset is meant to be used to fine-tune an LLM on my personal music preferences. I used exportify to
take all of my music from my Spotify account and download a csv file containing all of the metadata for those songs.
Then, I created a dataset of synthetic conversations related to those songs based on the Tempo, Acousticness, and Danceability of those songs.
Format
Each row of the dataset has an id starting from 1, and a pseudo-conversation of a song… See the full description on the dataset page: https://huggingface.co/datasets/Connorblu/my-music-dataset.
