datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.asr-evaluationsasr_evaluation_datasetsnepali_asr_evaluation_dataASR_Evaluation_dataset
Dataset Card for asr-africa/ASR_Evaluation_dataset
Dataset Overview
Languages Covered: Afrikaans, Amharic, Bemba, Hausa, Igbo, Kinyarwanda, Lingala, Luganda, Oromo, Swahili, Wolof, Xhosa, Yoruba
Source: Transcriptions evaluated by native/advanced speakers of the languages.
Dataset Description
This dataset provides human evaluations of automatic speech recognition (ASR) outputs across 13 African languages. For each audio sample, it includes the model-generated… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/ASR_Evaluation_dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/priyamallojjala/eka-medical-asr-evaluation-dataset.asr_evaluation
Not supported long audio (tested with 20min)
bambara-asr-evaluation
bambara-asr-evaluation
A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference
transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-asr-evaluation", split="test")
print(ds[0]["text"], ds[0]["source_dataset"])
One config and one split, so no config argument is needed.
Config
Split
Rows
Audio
default
test
1,295
2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.African-ASR-Domain-Adaptation-Evaluation
Dataset Card for Africa ASR Domain Adaptation Benchmark Dataset
This dataset forms the the Africa ASR domain adaptation benchmark. The goal of the dataset is to enable building of ASR models for African languages that can adapt to domians outside the training data.
The benchmark is made up of 2 languages, Wolof and Akan.
The training dataset for Akan is made up of general purpose data while the test dataset is financial data. The training dataset for Wolof is composed of general… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/African-ASR-Domain-Adaptation-Evaluation.asr-evaluation-telegramasr-evaluation-wmcASR-Evaluation-Audio
