CoolFace
11 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ekacare /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K17 likes1.3k downloads1y agoHugging Face02sujalappa /nvidia-brain-noise-evaluation-dataset Nvidia Brain Noise Evaluation Dataset Dataset Description This dataset contains 64 samples organized across multiple splits and 32 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: noisy-bg-snr-10: 2 samples test: 2 samples noisy-bg-snr-20: 2 samples test: 2 samples noisy-bg-snr-30: 2 samples test: 2 samples noisy-bg-snr-40: 2 samples test: 2 samples noisy-bg-snr-50: 2 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/nvidia-brain-noise-evaluation-dataset.audioautomatic-speech-recognitionn<1K0 likes95 downloads1y agoHugging Face03KothapalliAnusha /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K0 likes71 downloads8mo agoHugging Face04havahavai /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K1 likes50 downloads4mo agoHugging Face05sujalappa /denoised-evaluation-dataset Denoised Subset Fixed Dataset Description This dataset contains 500 samples organized across multiple splits and 1 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: denoised: 500 samples test: 500 samples Usage Load specific subset and split: from datasets import load_dataset # Load specific subset and split dataset = load_dataset('sujalappa/denoised-evaluation-dataset'… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/denoised-evaluation-dataset.audioautomatic-speech-recognitionn<1K0 likes41 downloads1y agoHugging Face06speech-uk /asr-evaluationstabularautomatic-speech-recognition10K<n<100K0 likes38 downloads2y agoHugging Face07sujalappa /speech-brain-noise-evaluation-dataset Speech Brain Noise Evaluation Dataset Dataset Description This dataset contains 2,000 samples organized across multiple splits and 20 subsets. The dataset includes audio data. Dataset Structure Subsets This dataset includes the following subsets: noisy-bg-snr-10: 100 samples test: 100 samples noisy-bg-snr-30: 100 samples test: 100 samples noisy-bg-snr-50: 100 samples test: 100 samples denoised-bg-snr-10: 100 samples test: 100 samples… See the full description on the dataset page: https://huggingface.co/datasets/sujalappa/speech-brain-noise-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K0 likes35 downloads1y agoHugging Face08isaacnetero /polyvox-kpi-evaluation PolyVox KPI Evaluation Dataset This dataset supports PolyVox evaluation for Problem Statement 11: Real-Time Multi-User Smart Assistant for Dynamic and Noisy Smart Environments. It contains compact evaluation mixtures for 2-speaker and 3-speaker speech separation under clean, noisy, overlap, and optional RIR-style conditions. Contents manifests/kpi_dataset_manifest.csv dataset_summary.json audio/ containing mixtures and clean source references Source… See the full description on the dataset page: https://huggingface.co/datasets/isaacnetero/polyvox-kpi-evaluation.audioautomatic-speech-recognition1K<n<10K0 likes34 downloads3mo agoHugging Face09asr-africa /ASR_Evaluation_datasetgated Dataset Card for asr-africa/ASR_Evaluation_dataset Dataset Overview Languages Covered: Afrikaans, Amharic, Bemba, Hausa, Igbo, Kinyarwanda, Lingala, Luganda, Oromo, Swahili, Wolof, Xhosa, Yoruba Source: Transcriptions evaluated by native/advanced speakers of the languages. Dataset Description This dataset provides human evaluations of automatic speech recognition (ASR) outputs across 13 African languages. For each audio sample, it includes the model-generated… See the full description on the dataset page: https://huggingface.co/datasets/asr-africa/ASR_Evaluation_dataset.audioautomatic-speech-recognition1K<n<10K1 likes10 downloads1y agoHugging Face10priyamallojjala /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/priyamallojjala/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K1 likes10 downloads4mo agoHugging Face11djelia /bambara-asr-evaluationgated bambara-asr-evaluation A Bambara ASR benchmark: 1,295 utterances, 2.04 hours of 16 kHz audio with reference transcripts. Monolingual Bambara transcription — audio in, transcript out, WER out. Load from datasets import load_dataset ds = load_dataset("djelia/bambara-asr-evaluation", split="test") print(ds[0]["text"], ds[0]["source_dataset"]) One config and one split, so no config argument is needed. Config Split Rows Audio default test 1,295 2.043 h… See the full description on the dataset page: https://huggingface.co/datasets/djelia/bambara-asr-evaluation.audioautomatic-speech-recognition1K<n<10K0 likes5 downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.