CoolFace
18 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ekacare /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K17 likes1.2k downloads1y agoHugging Face02KothapalliAnusha /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K0 likes72 downloads8mo agoHugging Face03speech-uk /asr-evaluationstabularautomatic-speech-recognition10K<n<100K0 likes40 downloads2y agoHugging Face04havahavai /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K1 likes40 downloads4mo agoHugging Face05prvInSpace /asr-evaluationstext10K<n<100K0 likes35 downloads1y agoHugging Face06wandererupak /nepali_asr_evaluation_dataaudio1K<n<10K0 likes21 downloads4mo agoHugging Face07prvInSpace /welsh-asr-evaluation-settext10K<n<100K0 likes20 downloads1y agoHugging Face08danielrosehill /Podcast-ASR-Evaluationtextn<1K0 likes16 downloads10mo agoHugging Face09djopin /asr_evaluation_datasetsaudion<1K0 likes14 downloads11mo agoHugging Face10priyamallojjala /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/priyamallojjala/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K1 likes13 downloads4mo agoHugging Face11asr-malayalam /Norm_Malayalam_Evaluation_samples Malayalam ASR Reference Prediction dataset This repository contains evaluation results from the Malayalam ASR model "vrclc/Whisper_small_malayalam" using the "google/fleurs" dataset. ASR Model Name: vrclc/Whisper_small_malayalam Dataset: google/fleurs Curated by: VRCLC vrclc/Whisper_small_malayalam was trained with 50 hours of Malayalam speech data. The test set of google/fleurs dataset which consists of Malayalam speech data was used to evaluate the model The evaluation of 500… See the full description on the dataset page: https://huggingface.co/datasets/asr-malayalam/Norm_Malayalam_Evaluation_samples.textsentence-similarityn<1K0 likes10 downloads2y agoHugging Face12asr-africa /ASR_Evaluation_datasetgated Dataset Card for asr-africa/ASR_Evaluation_dataset Dataset Overview Languages Covered: Afrikaans, Amharic, Bemba, Hausa, Igbo, Kinyarwanda, Lingala, Luganda, Oromo, Swahili, Wolof, Xhosa, Yoruba Source: Transcriptions evaluated by native/advanced speakers of the languages. Dataset Description This dataset provides human evaluations of automatic speech recognition (ASR) outputs across 13 African languages. For each audio sample, it includes the model-generated… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/ASR_Evaluation_dataset.audioautomatic-speech-recognition1K<n<10K1 likes10 downloads1y agoHugging Face13viethq5 /asr_evaluation Not supported long audio (tested with 20min) audio1K<n<10K0 likes7 downloads2y agoHugging Face14djelia /bambara-asr-evaluationgated bambara-asr-evaluation A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-asr-evaluation", split="test") print(ds[0]["text"], ds[0]["source_dataset"]) One config and one split, so no config argument is needed. Config Split Rows Audio default test 1,295 2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.audioautomatic-speech-recognition1K<n<10K0 likes5 downloads2mo agoHugging Face15asr-africa /African-ASR-Domain-Adaptation-Evaluationgated Dataset Card for Africa ASR Domain Adaptation Benchmark Dataset This dataset forms the the Africa ASR domain adaptation benchmark. The goal of the dataset is to enable building of ASR models for African languages that can adapt to domians outside the training data. The benchmark is made up of 2 languages, Wolof and Akan. The training dataset for Akan is made up of general purpose data while the test dataset is financial data. The training dataset for Wolof is composed of general… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/African-ASR-Domain-Adaptation-Evaluation.audio10K<n<100K0 likes2 downloads1y agoHugging Face16rinabuoy /asr-evaluation-telegramgatedaudion<1K0 likes1 downloads2y agoHugging Face17rinabuoy /asr-evaluation-wmcgatedaudion<1K0 likes1 downloads2y agoHugging Face18saanimustaf /ASR-Evaluation-Audioaudion<1K0 likes1 downloads6mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.