datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
melodix-stemsCeltic_Stems_Reference_Sessions_Preview
Harmonic Frontier Audio – Celtic Constellation Reference Sessions (Preview, v0.9)
A high-fidelity music-production dataset designed to connect isolated source performances, production processing, arrangement context, and finished musical outcomes.
Celtic Constellation Reference Sessions (Preview), created by Harmonic Frontier Audio, introduces the Reference Sessions product vertical through a compact proof-of-concept built around purpose-recorded Celtic ensemble material.… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Celtic_Stems_Reference_Sessions_Preview.swahili-speech
Swahili Speech-to-Text Dataset
This dataset contains paired audio and text data for training and evaluating speech-to-text models in Swahili. The audio files have been processed to remove silence, converted to 44.1kHz mono FLAC format, and are paired with corresponding transcriptions.
Structure
audio_*.flac: Audio files in FLAC format, named by their corresponding text corpus ID.
metadata.jsonl: JSON Lines file with metadata for each audio-text pair. Each line is a JSON… See the full description on the dataset page: https://huggingface.co/datasets/stem-content-ai-project/swahili-speech.MUSDB_stems_encodec_12kbpsPlave_StemsMUSDB_stems_opus_12kbpsMUSDB_stems_stable_audio_fp16stem-separation-benchmark-2026
StemSplit Stem-Separation Benchmark 2026
A reproducible head-to-head comparison of every popular open-source music
source-separation model against the StemSplit production
API, evaluated on the standard MUSDB18-HQ test split using BSS Eval v4 and a
small set of CC-BY tracks for qualitative listening.
Built and maintained by the StemSplit team. Source code:
scripts/hf-benchmark on GitHub.
Leaderboard (median SDR per stem)
model_id
bass
drums
other
vocals… See the full description on the dataset page: https://huggingface.co/datasets/StemSplitio/stem-separation-benchmark-2026.Stems-Evaluation-Kit
🎧 SonicSets High-Fidelity Stems Evaluation Kit
This is a premium evaluation subset provided by SonicSets, the industrial-grade audio data infrastructure for Large Audio Models (LAM).
📊 Dataset Specifications
Format: 48kHz / 24-bit Uncompressed WAV (Studio-Grade Ground Truth)
Feature: Absolute zero-crosstalk multi-track isolation
Environment: Strict anechoic capture (RT60 < 0.2s)
Purpose: Fully optimized for training and benchmarking state-of-the-art Source… See the full description on the dataset page: https://huggingface.co/datasets/drizzymedia/Stems-Evaluation-Kit.Stems-Evaluation-Kit
🎧 SonicSets High-Fidelity Stems Evaluation Kit
This is a premium evaluation subset provided by SonicSets, the industrial-grade audio data infrastructure for Large Audio Models (LAM).
📊 Dataset Specifications
Format: 48kHz / 24-bit Uncompressed WAV (Studio-Grade Ground Truth)
Feature: Absolute zero-crosstalk multi-track isolation
Environment: Strict anechoic capture (RT60 < 0.2s)
Purpose: Fully optimized for training and benchmarking state-of-the-art Source Separation… See the full description on the dataset page: https://huggingface.co/datasets/sonicsets-data/Stems-Evaluation-Kit.Pop-Rock-Hybrid-Stem-Dataset-cat001
Dataset Overview: Pop Rock Hybrid Stem Dataset (cat001)
This dataset contains a curated collection of original instrumental music designed for commercial and research applications in music analysis, audio modeling, and production workflows.
Every composition, arrangement, performance, sound design element, and production decision was created entirely through human musical and technical processes. All music contained in this dataset is 100% human-made (is_human_created: TRUE).… See the full description on the dataset page: https://huggingface.co/datasets/ToneCubeMedia/Pop-Rock-Hybrid-Stem-Dataset-cat001.stemflipper-dataset
StemFlipper — synth & effects parameter-estimation dataset (scaffold)
A synthetic (audio → parameters) dataset for the inverse problems StemFlipper
targets: recover a synth patch from its sound, and recover an effect chain from
wet audio. No public dataset pairs real audio with the synth-patch / effect-chain
parameters that produced it — this fills that gap with deterministic synthetic
generation. The moat is the generator + seeds, not stored audio: every example
here… See the full description on the dataset page: https://huggingface.co/datasets/nakas/stemflipper-dataset.music-stem-and-annotation-samplesDolby_Atmos_Stems250129_PURINA_Merrick_FridgeFasc_BBQ_STEM_MUSIC
