datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
a5sv2-asr-benchmark-dataset
A5Sv2 ASR Benchmark Dataset
Public references, saved predictions, scores, and provenance for the
A5Sv2 ASR benchmark. The benchmark evaluates
streaming English ASR on four fixed public corpora with approximately equal normalized reference
word counts.
Corpus
Fixed selection
Reference words
Audio in this repository
Mega-ASR / Voices-in-the-Wild-2M
1,250 utterances, 250 per acoustic condition
32,928
Yes
AMI
7 scenario-only unseen-evaluation meetings
32,928
Yes
DiPCo… See the full description on the dataset page: https://huggingface.co/datasets/AirCaps/a5sv2-asr-benchmark-dataset.VoiceIsolation-Benchmark-Dataset
Voice Isolation Benchmark Dataset
265 real-world recordings for measuring how a second voice breaks speech-to-text, and how much Krisp Voice Isolation fixes it. Three scenarios, 47 speakers, real rooms, real headsets. No synthetic mixing.
265 recordings · 47 speakers · 3 scenarios · 65 scripts
Why this dataset exists
Modern STT engines handle noise well. They still fail when a second person talks near the microphone: they transcribe the wrong speaker, and voice… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset.VoiceIsolation-Benchmark-Dataset-Processed
Voice Isolation Benchmark – Processed Audios
Audio samples processed by four Krisp Voice Isolation models. Each scenario folder contains subfolders for every model, with one processed .wav file per original sample.
Voice Isolation Models
VI 2.5 Default (vi_2_5_default)
The main Voice Isolation model. Strongest at removing noise and other speakers. Goes fully silent when only a bystander is talking.
VI 2.5 Balanced (vi_2_5_balanced)… See the full description on the dataset page: https://huggingface.co/datasets/Krisp-AI/VoiceIsolation-Benchmark-Dataset-Processed.benchmark_datasetmusic-benchmark-datasetcall-center-samples-for-stt-benchmarkbenchmark_english_datasetbenchmark_dataset
