datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
mega-asr-conversational-overlap
Mega-ASR Conversational Overlap
Mega-ASR Conversational Overlap is a deterministic English ASR diagnostic set
derived from AirCaps/mega-asr-noise-a5sv2,
which in turn is sampled from the Mega-ASR training corpus
zhifeixie/Voices-in-the-Wild-2M.
The existing AirCaps dataset evaluates single-utterance acoustic robustness.
This companion dataset evaluates a different failure mode: two-turn conversational
continuity with slight overlap and unequal turn loudness. It does not replace… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-conversational-overlap.mega-asr-noise-a5sv2
Mega-ASR Noise A5SV2
Mega-ASR Noise A5SV2 is a deterministic, English-only robustness evaluation
subset derived from
zhifeixie/Voices-in-the-Wild-2M,
the training corpus released with Mega-ASR.
We sampled from Mega-ASR-Train, rather than the standard Mega-ASR test set,
because in our experiments the standard test set was not acoustically
challenging enough to clearly discriminate among robust ASR systems. This is a
derived evaluation set; it should not be mixed into training… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/mega-asr-noise-a5sv2.rir-mega-speech
RIR-Mega-Speech
Dataset Summary
RIR-Mega-Speech is a large-scale reverberant speech corpus created by convolving LibriSpeech utterances with simulated room impulse responses (RIRs sampled from the RIR-Mega collection). Each reverberant utterance includes per-file acoustic metadata computed from the source RIR, enabling controlled analysis of reverberation effects on speech processing systems.
This dataset emphasizes transparency and reproducibility: acoustic metrics are… See the full description on the dataset page: https://huggingface.co/datasets/mandipgoswami/rir-mega-speech.german-pronuncheck-mega-dataset
German PronunCheck Mega Dataset 🇩🇪
Dataset Summary
This is a highly curated, 123GB+ mega-dataset designed specifically for training and fine-tuning German Automatic Speech Recognition (ASR) and Computer-Assisted Pronunciation Training (CAPT) models, such as HuBERT and Wav2Vec2.
Composition
This dataset is a clean concatenation of three distinct open-source datasets:
Mozilla Common Voice 26.0 (German): Standard scripted crowdsourced speech.… See the full description on the dataset page: https://huggingface.co/datasets/q1805/german-pronuncheck-mega-dataset.
