CoolFace
9 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ekacare /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K18 likes1.2k downloads1y agoHugging Face02sujalappa /nvidia-brain-noise-evaluation-dataset Nvidia Brain Noise Evaluation Dataset Dataset Description This dataset contains 64 samples organized across multiple splits and 32 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: noisy-bg-snr-10: 2 samples test: 2 samples noisy-bg-snr-20: 2 samples test: 2 samples noisy-bg-snr-30: 2 samples test: 2 samples noisy-bg-snr-40: 2 samples test: 2 samples noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.audioautomatic-speech-recognitionn<1K0 likes95 downloads1y agoHugging Face03KothapalliAnusha /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K0 likes71 downloads8mo agoHugging Face04havahavai /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K1 likes54 downloads4mo agoHugging Face05sujalappa /denoised-evaluation-dataset Denoised Subset Fixed Dataset Description This dataset contains 500 samples organized across multiple splits and 1 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: denoised: 500 samples test: 500 samples Usage Load specific subset and split: from datasets import load_dataset # Load specific subset and split dataset = load_dataset('sujalappa/denoised-evaluation-dataset'… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/denoised-evaluation-dataset.audioautomatic-speech-recognitionn<1K0 likes43 downloads1y agoHugging Face06sujalappa /speech-brain-noise-evaluation-dataset Speech Brain Noise Evaluation Dataset Dataset Description This dataset contains 2,000 samples organized across multiple splits and 20 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: noisy-bg-snr-10: 100 samples test: 100 samples noisy-bg-snr-30: 100 samples test: 100 samples noisy-bg-snr-50: 100 samples test: 100 samples denoised-bg-snr-10: 100 samples test: 100 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/speech-brain-noise-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K0 likes34 downloads1y agoHugging Face07priyamallojjala /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/priyamallojjala/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K1 likes17 downloads4mo agoHugging Face08asr-africa /ASR_Evaluation_datasetgated Dataset Card for asr-africa/ASR_Evaluation_dataset Dataset Overview Languages Covered: Afrikaans, Amharic, Bemba, Hausa, Igbo, Kinyarwanda, Lingala, Luganda, Oromo, Swahili, Wolof, Xhosa, Yoruba Source: Transcriptions evaluated by native/advanced speakers of the languages. Dataset Description This dataset provides human evaluations of automatic speech recognition (ASR) outputs across 13 African languages. For each audio sample, it includes the model-generated… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/ASR_Evaluation_dataset.audioautomatic-speech-recognition1K<n<10K1 likes11 downloads1y agoHugging Face09djelia /bambara-asr-evaluationgated bambara-asr-evaluation A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-asr-evaluation", split="test") print(ds[0]["text"], ds[0]["source_dataset"]) One config and one split, so no config argument is needed. Config Split Rows Audio default test 1,295 2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.audioautomatic-speech-recognition1K<n<10K0 likes7 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.