datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.nvidia-brain-noise-evaluation-dataset
Nvidia Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 64 samples organized across multiple splits and 32 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 2 samples
test: 2 samples
noisy-bg-snr-20: 2 samples
test: 2 samples
noisy-bg-snr-30: 2 samples
test: 2 samples
noisy-bg-snr-40: 2 samples
test: 2 samples
noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.denoised-evaluation-dataset
Denoised Subset Fixed
Dataset Description
This dataset contains 500 samples organized across multiple splits and 1 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
denoised: 500 samples
test: 500 samples
Usage
Load specific subset and split:
from datasets import load_dataset
# Load specific subset and split
dataset = load_dataset('sujalappa/denoised-evaluation-dataset'… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/denoised-evaluation-dataset.asr-evaluationsspeech-brain-noise-evaluation-dataset
Speech Brain Noise Evaluation Dataset
Dataset Description
This dataset contains 2,000 samples organized across multiple splits and 20 subsets.
The dataset includes audio data.
Dataset Structure
Subsets
This dataset includes the following subsets:
noisy-bg-snr-10: 100 samples
test: 100 samples
noisy-bg-snr-30: 100 samples
test: 100 samples
noisy-bg-snr-50: 100 samples
test: 100 samples
denoised-bg-snr-10: 100 samples
test: 100 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/speech-brain-noise-evaluation-dataset.polyvox-kpi-evaluation
PolyVox KPI Evaluation Dataset
This dataset supports PolyVox evaluation for Problem Statement 11: Real-Time Multi-User Smart Assistant for Dynamic and Noisy Smart Environments.
It contains compact evaluation mixtures for 2-speaker and 3-speaker speech separation under clean, noisy, overlap, and optional RIR-style conditions.
Contents
manifests/kpi_dataset_manifest.csv
dataset_summary.json
audio/ containing mixtures and clean source references
Source… See the full description on the dataset page: https://huggingface.co/datasets/isaacnetero/polyvox-kpi-evaluation.ASR_Evaluation_dataset
Dataset Card for asr-africa/ASR_Evaluation_dataset
Dataset Overview
Languages Covered: Afrikaans, Amharic, Bemba, Hausa, Igbo, Kinyarwanda, Lingala, Luganda, Oromo, Swahili, Wolof, Xhosa, Yoruba
Source: Transcriptions evaluated by native/advanced speakers of the languages.
Dataset Description
This dataset provides human evaluations of automatic speech recognition (ASR) outputs across 13 African languages. For each audio sample, it includes the model-generated… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/ASR_Evaluation_dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/priyamallojjala/eka-medical-asr-evaluation-dataset.bambara-asr-evaluation
bambara-asr-evaluation
A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference
transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-asr-evaluation", split="test")
print(ds[0]["text"], ds[0]["source_dataset"])
One config and one split, so no config argument is needed.
Config
Split
Rows
Audio
default
test
1,295
2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.
