datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Complete_Data_Source_100K_HOURS
Multi-Language Audio Collection (100K Hours)
This repository is physically reorganized for Absolute 100% Data Visibility.
🏗️ Global Consolidator
Select your language subset to listen to high-quality waveform audio. All shards from legacy and modern pipelines are automatically routed here.
SourceDetection_mb23-music_caps_4sec_wave_type_continuous
Dataset Card for "SourceDetection_mb23-music_caps_4sec_wave_type_continuous"
More Information needed
filtered_common_voice-hi-sourcenepali-music-source-seperation
nepali-music-source-seperation
Mirror of the exact bmcore v24 local holdout subset: 285 audio files.
Benchmark label: real. This preserves the source benchmark label; it is not an independent label review.
Source reference: https://huggingface.co/datasets/34data/v13-nepali-music-source-seperation.
No new license or ownership claim is asserted by this mirror. Original source rights and restrictions remain applicable.
Source revision reviewed:… See the full description on the dataset page: https://huggingface.co/datasets/34data/nepali-music-source-seperation.SourceSeparation_libri2Mix_testSourceDetection_mb23-music_caps_4sec_wave_typesbpn-source-vote-disagreements-htdemucs-vs-ft-20260903
Source-vote disagreements: htdemucs vs htdemucs_ft
The SBPN chunker publishes, for every chunk, whichever of the original audio
or the Demucs vocals scores higher on a librosa peak-to-10th-percentile RMS
ratio, at a 0 dB threshold. Swapping the separation model therefore changes
which audio ships, not merely its quality.
Replaying that decision over 400 real chunks with htdemucs (shipped) and
htdemucs_ft (vocals sub-model) produced 37 disagreements — a 9.25% flip
rate. This… See the full description on the dataset page: https://huggingface.co/datasets/Cybrpgs/sbpn-source-vote-disagreements-htdemucs-vs-ft-20260903.gdrive-transcript1-audio-chunks-20260805-source-02
gdrive-transcript1-audio-chunks-20260805-source-02
This gated dataset contains 1,095 MP3 speech and audio-event chunks from
50 recordings. Access requires manual approval by the
repository owner.
Audio selection
Every complete source recording was separated once with HTDemucs.
Full vocal stems were stored remotely as 48 kHz mono 96k MP3;
no per-chunk source separation was performed.
Librosa SNR was calculated independently on aligned original and saved-vocals… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-transcript1-audio-chunks-20260805-source-02.bird3m-raw-sourcesgdrive-transcript1-audio-chunks-20260805-source-04
gdrive-transcript1-audio-chunks-20260805-source-04
This gated dataset contains 595 MP3 speech and audio-event chunks from
50 recordings. Access requires manual approval by the
repository owner.
Audio selection
Every complete source recording was separated once with HTDemucs.
Full vocal stems were stored remotely as 48 kHz mono 96k MP3;
no per-chunk source separation was performed.
Librosa SNR was calculated independently on aligned original and saved-vocals… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-transcript1-audio-chunks-20260805-source-04.gdrive-transcript1-audio-chunks-20260805-source-01
gdrive-transcript1-audio-chunks-20260805-source-01
This gated dataset contains 1,033 MP3 speech and audio-event chunks from
45 recordings. Access requires manual approval by the
repository owner.
Audio selection
Every complete source recording was separated once with HTDemucs.
Full vocal stems were stored remotely as 48 kHz mono 96k MP3;
no per-chunk source separation was performed.
Librosa SNR was calculated independently on aligned original and saved-vocals… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-transcript1-audio-chunks-20260805-source-01.gdrive-transcript1-audio-chunks-20260805-source-06
gdrive-transcript1-audio-chunks-20260805-source-06
This gated dataset contains 2,117 MP3 speech and audio-event chunks from
50 recordings. Access requires manual approval by the
repository owner.
Audio selection
Every complete source recording was separated once with HTDemucs.
Full vocal stems were stored remotely as 48 kHz mono 96k MP3;
no per-chunk source separation was performed.
Librosa SNR was calculated independently on aligned original and saved-vocals… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-transcript1-audio-chunks-20260805-source-06.gdrive-transcript1-audio-chunks-20260805-source-03
gdrive-transcript1-audio-chunks-20260805-source-03
This gated dataset contains 645 MP3 speech and audio-event chunks from
50 recordings. Access requires manual approval by the
repository owner.
Audio selection
Every complete source recording was separated once with HTDemucs.
Full vocal stems were stored remotely as 48 kHz mono 96k MP3;
no per-chunk source separation was performed.
Librosa SNR was calculated independently on aligned original and saved-vocals… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-transcript1-audio-chunks-20260805-source-03.v13-nepali-music-source-seperationgdrive-transcript1-audio-chunks-20260805-source-05
gdrive-transcript1-audio-chunks-20260805-source-05
This gated dataset contains 1,535 MP3 speech and audio-event chunks from
50 recordings. Access requires manual approval by the
repository owner.
Audio selection
Every complete source recording was separated once with HTDemucs.
Full vocal stems were stored remotely as 48 kHz mono 96k MP3;
no per-chunk source separation was performed.
Librosa SNR was calculated independently on aligned original and saved-vocals… See the full description on the dataset page: https://huggingface.co/datasets/Kppwdfgu1/gdrive-transcript1-audio-chunks-20260805-source-05.SourceSeparation_libri2Mix_testomni_source
Omni source dataset
Mongolian audio/text pairs used as source material for the omni training set.
Dataset Statistics
Total samples: 69
Total duration: 0h 15m 47s (0.26 h)
Per-split breakdown
Split
Samples
Total Duration
Avg Duration
train
69
0h 15m 47s (0.26 h)
13.73 s
Single split (no held-out test set) -- this is source/reference data, not a
benchmark split.
audio_source
