datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
melodix-stemsCeltic_Stems_Reference_Sessions_Preview
Harmonic Frontier Audio – Celtic Constellation Reference Sessions (Preview, v0.9)
A high-fidelity music-production dataset designed to connect isolated source performances, production processing, arrangement context, and finished musical outcomes.
Celtic Constellation Reference Sessions (Preview), created by Harmonic Frontier Audio, introduces the Reference Sessions product vertical through a compact proof-of-concept built around purpose-recorded Celtic ensemble material.… See the full description on the dataset page: https://huggingface.co/datasets/Harmonic-Frontier-Audio/Celtic_Stems_Reference_Sessions_Preview.MUSDB_stems_encodec_12kbpsMUSDB_stems_opus_12kbpsMUSDB_stems_stable_audio_fp16Plave_StemsStems-Evaluation-Kit
🎧 SonicSets High-Fidelity Stems Evaluation Kit
This is a premium evaluation subset provided by SonicSets, the industrial-grade audio data infrastructure for Large Audio Models (LAM).
📊 Dataset Specifications
Format: 48kHz / 24-bit Uncompressed WAV (Studio-Grade Ground Truth)
Feature: Absolute zero-crosstalk multi-track isolation
Environment: Strict anechoic capture (RT60 < 0.2s)
Purpose: Fully optimized for training and benchmarking state-of-the-art Source… See the full description on the dataset page: https://huggingface.co/datasets/drizzymedia/Stems-Evaluation-Kit.stem-separation-benchmark-2026
StemSplit Stem-Separation Benchmark 2026
A reproducible head-to-head comparison of every popular open-source music
source-separation model against the StemSplit production
API, evaluated on the standard MUSDB18-HQ test split using BSS Eval v4 and a
small set of CC-BY tracks for qualitative listening.
Built and maintained by the StemSplit team. Source code:
scripts/hf-benchmark on GitHub.
Leaderboard (median SDR per stem)
model_id
bass
drums
other
vocals… See the full description on the dataset page: https://huggingface.co/datasets/StemSplitio/stem-separation-benchmark-2026.Stems-Evaluation-Kit
🎧 SonicSets High-Fidelity Stems Evaluation Kit
This is a premium evaluation subset provided by SonicSets, the industrial-grade audio data infrastructure for Large Audio Models (LAM).
📊 Dataset Specifications
Format: 48kHz / 24-bit Uncompressed WAV (Studio-Grade Ground Truth)
Feature: Absolute zero-crosstalk multi-track isolation
Environment: Strict anechoic capture (RT60 < 0.2s)
Purpose: Fully optimized for training and benchmarking state-of-the-art Source Separation… See the full description on the dataset page: https://huggingface.co/datasets/sonicsets-data/Stems-Evaluation-Kit.Dolby_Atmos_Stems
