datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.asr-evaluationseka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.asr-evaluationsnepali_asr_evaluation_datawelsh-asr-evaluation-setPodcast-ASR-Evaluationasr_evaluation_datasetseka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/priyamallojjala/eka-medical-asr-evaluation-dataset.Norm_Malayalam_Evaluation_samples
Malayalam ASR Reference Prediction dataset
This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset.
ASR Model Name: vrclc/Whisper_small_malayalam
Dataset: google/fleurs
Curated by: VRCLC
vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data.
The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model
The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Norm_Malayalam_Evaluation_samples.ASR_Evaluation_dataset
Dataset Card for asr-africa/ASR_Evaluation_dataset
Dataset Overview
Languages Covered: Afrikaans, Amharic, Bemba, Hausa, Igbo, Kinyarwanda, Lingala, Luganda, Oromo, Swahili, Wolof, Xhosa, Yoruba
Source: Transcriptions evaluated by native/advanced speakers of the languages.
Dataset Description
This dataset provides human evaluations of automatic speech recognition (ASR) outputs across 13 African languages. For each audio sample, it includes the model-generated… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/ASR_Evaluation_dataset.asr_evaluation
Not supported long audio (tested with 20min)
bambara-asr-evaluation
bambara-asr-evaluation
A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference
transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out.
Load
from datasets import load_dataset
ds = load_dataset("djelia/bambara-asr-evaluation", split="test")
print(ds[0]["text"], ds[0]["source_dataset"])
One config and one split, so no config argument is needed.
Config
Split
Rows
Audio
default
test
1,295
2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.African-ASR-Domain-Adaptation-Evaluation
Dataset Card for Africa ASR Domain Adaptation Benchmark Dataset
This dataset forms the the Africa ASR domain adaptation benchmark. The goal of the dataset is to enable building of ASR models for African languages that can adapt to domians outside the training data.
The benchmark is made up of 2 languages, Wolof and Akan.
The training dataset for Akan is made up of general purpose data while the test dataset is financial data. The training dataset for Wolof is composed of general… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/African-ASR-Domain-Adaptation-Evaluation.asr-evaluation-telegramasr-evaluation-wmcASR-Evaluation-Audio
