datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
unbound009_2_sepsep28k_fullsep28ksep28kPixelParade_23_ds_sepstem-separation-benchmark-2026
StemSplit Stem-Separation Benchmark 2026
A reproducible head-to-head comparison of every popular open-source music
source-separation model against the StemSplit production
API, evaluated on the standard MUSDB18-HQ test split using BSS Eval v4 and a
small set of CC-BY tracks for qualitative listening.
Built and maintained by the StemSplit team. Source code:
scripts/hf-benchmark on GitHub.
Leaderboard (median SDR per stem)
model_id
bass
drums
other
vocals… See the full description on the dataset page: https://huggingface.co/datasets/StemSplitio/stem-separation-benchmark-2026.t9p3c8m1-axr4e6_sepsepedip9r3k6b1-zx7v4n2_sepn4x7d2q9-hf1m8t3_sepunfolded-veil-v9_sepsep28k-wavlm-layer-9voxceleb2-40k-part1-preprocess-all-files-separatesep28k-fluencybank-stutter-datasetnchlt_speech_sepedi
NCHLT Speech Corpus -- Sepedi
This is the Sepedi language part of the NCHLT Speech Corpus of the South African languages.
Language code (ISO 639): nso
URI: https://hdl.handle.net/20.500.12185/270
Licence:
Creative Commons Attribution 3.0 Unported License (CC BY 3.0): http://creativecommons.org/licenses/by/3.0/legalcode
Attribution:
The Department of Arts and Culture of the government of the Republic of South Africa (DAC), Council for Scientific and Industrial… See the full description on the dataset page: https://huggingface.co/datasets/Max5ive/nchlt_speech_sepedi.h4p7t3x2-jn6b9_sepDog-Vocal-Separation
[Dataset] Dog Vocal Separation
IMPORTANT NOTE for the IJCAI-2025 challenge
[Jun. 1st, 2025] The validation set has been updated.
[Apr. 25th, 2025] The dataset has some changes. Sorry for the inconvenience.
Overview
.
├── train
│ ├── train_pairs.csv
│ ├── dog
│ │ ├── 6357ca529eec8ca42a1fa588e0725904.wav
│ │ ├── cd06d290e0ebc76a137bd44ebec4d5fd.wav
│ │ └── ...
│ └── mixture
│ … See the full description on the dataset page: https://huggingface.co/datasets/ArlingtonCL2/Dog-Vocal-Separation.nepali-music-source-seperation
nepali-music-source-seperation
Mirror of the exact bmcore v24 local holdout subset: 285 audio files.
Benchmark label: real. This preserves the source benchmark label; it is not an independent label review.
Source reference: https://huggingface.co/datasets/34data/v13-nepali-music-source-seperation.
No new license or ownership claim is asserted by this mirror. Original source rights and restrictions remain applicable.
Source revision reviewed:… See the full description on the dataset page: https://huggingface.co/datasets/34data/nepali-music-source-seperation.donaroma3421.42_sepStutteringDetection_SEP28kda7ee7_sep_cleaned-8kSep28kda7ee7_sep_cleaned-8ksep28k-train-4-second-clips
Dataset Card for "sep28k-train-4-second-clips"
More Information needed
sep28k-train-3-second-clips-full-agreement
Dataset Card for "sep28k-train-3-second-clips-full-agreement"
More Information needed
sep28k-test-4-second-clips
Dataset Card for "sep28k-test-0120-4-second-clips"
More Information needed
sep28k-train-3-second-clips
Dataset Card for "sep28k-train-3-second-clips"
More Information needed
sep28k-dev-3-second-clip-full-agreement
Dataset Card for "sep28k-dev-3-second-clip-full-agreement"
More Information needed
sep28k
Dataset Card for "sep28k"
More Information needed
sep28k-train-5-second-clips
Dataset Card for "sep28k-train-5-second-clips"
More Information needed
