datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MuSR
MuSR: Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
Creating murder mysteries that require multi-step reasoning with commonsense using ChatGPT!
By: Zayne Sprague, Xi Ye, Kaj Bostrom, Swarat Chaudhuri, and Greg Durrett.
View the dataset on our custom viewer and project website!
Check out the paper. Appeared at ICLR 2024 as a spotlight presentation!
Git Repo with the source data, how to recreate the dataset (and create new ones!) here
midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/drengskapur/midi-classical-music.ai-musicsuno-ai-music-dataset
Suno AI Music Dataset (Multi-Genre Curated)
A human-curated, multi-genre audio dataset generated with Suno V5.5 (chirp-fenix), covering 100+ sub-sub-genres across electronic, hip-hop, Latin, jazz, world, rock, ambient, pop, reggae, and classical music. Each track ships with full audio (MP3), cover art, the original generation prompt, and a 32-column metadata schema designed for downstream audio-ML research.
This is not a "scrape everything Suno produces" dump. It is a… See the full description on the dataset page: https://huggingface.co/datasets/Kukedlc/suno-ai-music-dataset.MusicCaps
Dataset Card for MusicCaps
Dataset Summary
The MusicCaps dataset contains 5,521 music examples, each of which is labeled with an English aspect list and a free text caption written by musicians. An aspect list is for example "pop, tinny wide hi hats, mellow piano melody, high pitched female vocal melody, sustained pulsating synth lead", while the caption consists of multiple sentences about the music, e.g.,
"A low sounding male voice is rapping over a fast paced drums… See the full description on the dataset page: https://huggingface.co/datasets/google/MusicCaps.midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/ygonet/midi-classical-music.midi-classical-music
MIDI Classical Music
This dataset contains a comprehensive collection of MIDI files representing classical music compositions from various renowned composers.
The collection includes works from composers such as Bach, Beethoven, Chopin, Mozart, and many others.
The dataset is organized into directories by composer, with each directory containing MIDI files of their compositions.
The dataset is ideal for music analysis, machine learning models for music generation, and other… See the full description on the dataset page: https://huggingface.co/datasets/elizawhitfield/midi-classical-music.muslim-names-dataset
Muslim Names Dataset
A comprehensive collection of Muslim names with meanings scraped from muslimnames.com. Contains 14,585 names with English names, Arabic names, meanings, and gender classifications.
Dataset Contents
This dataset contains ~14,585 Muslim names with the following information:
English name: Name in English/Latin script
Arabic name: Name in Arabic script
Meaning: Definition and meaning of the name
Gender: Classification as male or female
Files… See the full description on the dataset page: https://huggingface.co/datasets/takiuddinahmed/muslim-names-dataset.MusicSem
Dataset Card for MusicSem
This dataset contains 35977 entries of text-audio pairs. There is an accompanying test set of size 480 which is withheld for leaderboard purposes. Please reach out to authors for further access.
Dataset Details
Dataset Description
Curated by: Rebecca Salganik, Teng Tu, Fei-Yueh Chen, Xiaohao Liu, Kaifeng Lu, Ethan Luvisia, Zhiyao Duan, Guillaume Salha-Galvan, Anson Kahng, Yunshan Ma, Jian Kang
Language(s) : English
License: MIT… See the full description on the dataset page: https://huggingface.co/datasets/AMSRNA/MusicSem.MuSP-Bench
MuSP-Bench: Advanced Multimodal Benchmarking of Music Understanding Across Score and Performance
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Official benchmark website
Modalities
Each question specifies the minimum source of musical evidence needed to answer it:
S (score): answer from the written score.
P (performance): answer from the performance recording.
S&P (score… See the full description on the dataset page: https://huggingface.co/datasets/bryel-labs/MuSP-Bench.MuSP-Bench
MuSP-Bench
MuSP-Bench is a 490-question benchmark for musical score understanding,
performance listening, and combined score-performance reasoning.
Contents
data/questions.csv: all 490 questions, accepted answers, and the
response contract for each.
inputs/pdf/without_context/: one context-removed PDF per piece.
inputs/images/: rendered score-page images for every piece.
inputs/abc/: one ABC score per piece.
inputs/abc_plus_midi/: one aligned ABC+MIDI… See the full description on the dataset page: https://huggingface.co/datasets/milan477/MuSP-Bench.modern_music_reDatasets for Relation Extraction TaskSource from Wikipedia (CC-BY-2.0)Contributors : Doohae Jung, Hyesu Kim, Bosung Kim, Isaac Park, Miwon Jeon, Dagon Lee, Jihoo Kim
MuST-C-and-WMT16-de-enjazz-music-archivesstsb-tr-hukuk
Turkish STS — sentence pairs scored with magibu/embeddingmagibu-200m
A Turkish Semantic Textual Similarity (STS) dataset: each row is a pair of
sentences plus a similarity score. Scores come from the
magibu/embeddingmagibu-200m
sentence-embedding model, computed as cosine similarity over L2-normalized
embeddings — the exact method used by the
reference Space.
The set deliberately spans the full similarity range, with many near-zero
(unrelated) pairs, so it can be used to… See the full description on the dataset page: https://huggingface.co/datasets/Mustafa2735/stsb-tr-hukuk.GT-Music-Genre
GT-Music-Genre
This is an audio classification dataset for Music Analysis.
Classes = 10 , Split = Train-Test
Structure
audios folder contains audio files.
train.csv for training split and test.csv for the testing split.
Download
import os
import huggingface_hub
audio_datasets_path = "DATASET_PATH/Audio-Datasets"
if not os.path.exists(audio_datasets_path): print(f"Given {audio_datasets_path=} does not exist. Specify a valid path ending with… See the full description on the dataset page: https://huggingface.co/datasets/MahiA/GT-Music-Genre.Wiki_Live_Challenge
Wiki Live Challenge Dataset
[English | 中文]
English
📖 Dataset Overview
This is the official dataset accompanying the Wiki Live Challenge benchmark. It contains Wikipedia Good Articles (GAs) as ground truth and research articles generated by leading deep research AI systems.
Wiki Live Challenge is the first live benchmark for evaluating Deep Research Agents (DRAs) on their ability to generate Wikipedia-quality articles. Unlike static benchmarks, Wiki Live… See the full description on the dataset page: https://huggingface.co/datasets/muset-ai/Wiki_Live_Challenge.suno-style-recipes
Source, notebook and weekly sync on GitHub
Suno Style Recipes — 564 documented music styles
Style prompt, BPM, weirdness, style influence and canonical song structure for
564 music styles, formatted for AI music generators (Suno, Udio).
https://musaisong.app/en/styles
style_prompt — paste into the "Style of Music" field
bpm · weirdness · style_influence — the three settings that change the output
structure — the canonical section order for that genre
url — the full page for… See the full description on the dataset page: https://huggingface.co/datasets/musaisong/suno-style-recipes.ML-Music-Classifier-dataset-and-model-name-Models
🎧 Spotify Music Preference Analysis
🧠 Project Overview
This project analyzes Spotify music data to predict song preferences using machine learning models. The analysis is based on a dataset of 195 songs (100 liked, 95 disliked) with various audio features extracted from Spotify's API.
📂 Dataset Description
📥 Data Collection Process
Liked Songs (100 tracks):
🎵 Primarily French Rap
🎸 Some American Rap, Rock, and Electronic music
✅… See the full description on the dataset page: https://huggingface.co/datasets/Jack1808/ML-Music-Classifier-dataset-and-model-name-Models.musdb25
Dataset Card for MUSDB25 (Alpha Release)
This is an alpha release of the MUSDB25 dataset, a fully multitrack recreation of MUSDB18.
The goal is to have a linearly summable multitrack dataset for source separation that remains compatible with MUSDB18.
For the current alpha release, almost 100 out of 150 are more or less almost exactly compatible, but the quality and usability of the individual tracks still vary.
Please feel free to open issues or discussion to provide feedback on… See the full description on the dataset page: https://huggingface.co/datasets/kwatcharasupat/musdb25.Genius-Turkish-Dataset
Turkish Song Lyrics from Genius Dataset
Dataset Description
This dataset contains a comprehensive collection of 44,692 Turkish song lyrics, extracted from the larger "Genius Song Lyrics with Language Information" dataset available on Kaggle. The original 9.07 GB dataset was filtered to include only songs identified with the language code 'tr' (Turkish), making it a clean and focused resource for Turkish Natural Language Processing (NLP) tasks.
[TR] Bu veri seti, Kaggle'da… See the full description on the dataset page: https://huggingface.co/datasets/mustafakemal0146/Genius-Turkish-Dataset.QuranExeThis dataset contains the exegeses/tafsirs (تفسير القرآن) of the holy Quran in arabic by 8 exegetes.
This is a non Official dataset. It have been scrapped from the Quran.com Api
This dataset contains 49888 records with +14 Million words. 8 records per Quranic verse
Usage Example :
from datasets import load_dataset
tafsirs = load_dataset("mustapha/QuranExe")
Preprocessed-MS-IL-POST-Data
Dataset Card for Dataset Name
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Dataset Details
Uses
Direct Use
[More Information Needed]
Out-of-Scope Use
[More Information Needed]
Dataset Structure
[More Information Needed]
Dataset Creation
Curation Rationale
[More Information Needed]
Source Data
Data Collection… See the full description on the dataset page: https://huggingface.co/datasets/musfiqdehan/Preprocessed-MS-IL-POST-Data.british-museum-rdf-as-csv-2014This repository contains data that was released under the BM license issued in 2014.
uk-live-music-rates-2026-08
UK Live Music Rates - August 2026 (GigXchange Index)
Monthly open dataset of what live music pays in the UK, by city, event type and band size. Every
figure is artist take-home in GBP: agency-published rates are normalised by stripping a 20%
commission before any percentile is computed.
This is Issue 05 of the GigXchange Index, published 2026-09-03. The file is a
snapshot of exactly what gigxchange.app published on 2026-09-03 - the same figures shown
on the site and in the… See the full description on the dataset page: https://huggingface.co/datasets/gigxchange/uk-live-music-rates-2026-08.LP_MusicCaps_MCMuscle_Fatigue_CyclingThis dataset was created with healthy participants aged between 18 and 25 years old. The participants in this dataset were not frequent athletes.
The dataset consists of 8 EMG signals recorded from the domineering foot of each participant during a cycling trial. The participants performed exercises on a conditioned cycle, alternating with short periods of high-intensity sprints. When a participant could no longer sustain the sprint intensity, this was considered the first index of fatigue and… See the full description on the dataset page: https://huggingface.co/datasets/YominE/Muscle_Fatigue_Cycling.MuST-C-deMusic_Metadata_and_Lyrics_Dataset
Music Metadata & Lyrics Dataset
Se trata de un dataset con un extenso conjunto de datos de canciones más de 200,000, combinando metadatos musicales e información de artistas.Hemos construido el dataset a partir de un original de Kaggle y lo hemos enriquecido mediante scraping a Wikidata, ampliando la información biográfica y contextual de los artistas.
Cada registro contiene información sobre una canción, el artista, su contexto y campos derivados.
Estructura de los… See the full description on the dataset page: https://huggingface.co/datasets/MarKos29/Music_Metadata_and_Lyrics_Dataset.injection-molding-QA
injection-molding-QA
Description
This dataset contains questions and answers related to injection molding, focusing on topics such as 'Materials', 'Techniques', 'Machinery', 'Troubleshooting', 'Safety','Design','Maintenance','Manufacturing','Development','R&D'. The dataset is provided in CSV format with two columns: Questions and Answers.
Usage
Researchers, practitioners, and enthusiasts in the field of injection molding can utilize this dataset for tasks such… See the full description on the dataset page: https://huggingface.co/datasets/mustafakeser/injection-molding-QA.
