CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01ekacare /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K18 likes1.2k downloads1y agoHugging Face02VSSA-SDSA /LT_Medical_S_corpusEnglish | Lietuvių English LT_Medical_S_corpus — Lithuanian Medical Speech Corpus A Lithuanian speech dataset of medical dictation audio (radiology and family medicine) with transcriptions, speaker metadata, and word-level timestamps. Columns Column Type Description audio Audio Audio sentence string Ground truth transcription duration_ms int Recording duration in milliseconds medical_area string RADIOLOGIJA or SEIMOS gender string MALE or… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_Medical_S_corpus.audioautomatic-speech-recognition10K<n<100K1 likes350 downloads5mo agoHugging Face03manhcuong2005 /medical_noise_data Medical Noise Dataset Dữ liệu âm thanh tiếng Việt đã được tăng cường nhiễu (noise augmentation + RIR convolution). Nguồn gốc Audio gốc: dolly-vn/dolly-audio-1000h-vietnamese Noise sources: YouTube extracted, WHAM!, Hospital ambient noise RIR: Real RIR far-field (RVB2014) Cách load from datasets import load_dataset ds = load_dataset("manhcuong2005/medical_noise_data") audio10K<n<100K1 likes259 downloads2mo agoHugging Face04Hani89 /Synthetic-Medical-Speech-Dataset Synthetic Medical Speech Dataset Overview Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.audioautomatic-speech-recognition10K<n<100K4 likes247 downloads1y agoHugging Face05jarvisx17 /Medical-ASR-ENaudio1K<n<10K9 likes198 downloads4y agoHugging Face06Yettiesoft /voice_medicalaudioautomatic-speech-recognitionn<1K0 likes159 downloads2y agoHugging Face07Yettiesoft /voice_medical_cut_smallaudio10K<n<100K0 likes143 downloads2y agoHugging Face08Yettiesoft /voice_medical_newaudio10K<n<100K0 likes141 downloads2y agoHugging Face09SilencioNetwork /medical-speech-dataset Medical Speech Dataset A protocol sample. 11 contributors, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment. Small by design — see What this is for below before downloading. Hours 0.69 Clips 33 Speakers 11 Origin varieties 11 Languages 2 Configs 2 Speaker metadata origin region / variety, mother tongue, gender, device, OS, recording environment… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/medical-speech-dataset.audioaudio-classificationn<1K0 likes114 downloads3d agoHugging Face10HumynLabs /medical-prescription-english-audio Medical Prescription English Audio Dataset Text spoken by all participants: "Doctor, my third visit, and I'm hopeful but not fully better. Joint pain eased slightly, yet mornings are tough, and I'm exhausted. The last prescription helped a bit. Can we adjust it? I want to feel like myself again." The dataset supports training and evaluation of models in: Automatic Speech Recognition (ASR) Emotional tone classification Voice synthesis and generation Emotion-aware conversational… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/medical-prescription-english-audio.audioaudio-classificationn<1K719 likes113 downloads1y agoHugging Face11shreyas1104 /medical-intent-audio-datasetaudio1K<n<10K1 likes110 downloads3y agoHugging Face12turkmedstt /medv3-turkish-medical-asr medv3 - Türkçe Sentetik Tıbbi Konuşma Korpusu Türkçe tıbbi konuşma tanıma araştırmaları için hazırlanmış sentetik konuşma korpusudur. Klinik cümleler Google Cloud Text-to-Speech Chirp 3 HD sesleriyle sentezlenmiştir. Önemli uyarılar Tüm kayıtlar sentetiktir (synthetic=true). Gerçek hasta veya klinisyen sesi ve kişisel sağlık verisi içermez. Tıbbi cihaz geliştirme onayı veya klinik doğrulama anlamına gelmez. Klinik karar için değil, araştırma ve ASR… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/medv3-turkish-medical-asr.audioautomatic-speech-recognition1K<n<10K4 likes105 downloads3mo agoHugging Face13ahmedelsayed /MedicalDatasetaudio1K<n<10K0 likes104 downloads1y agoHugging Face14yashtiwari /PaulMooney-Medical-ASR-Dataaudio1K<n<10K9 likes90 downloads3y agoHugging Face15adriana98 /medical_spanishaudion<1K5 likes88 downloads3y agoHugging Face16KothapalliAnusha /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K0 likes72 downloads8mo agoHugging Face17batyrme /medical-speech-datasetgatedaudio10K<n<100K2 likes63 downloads23d agoHugging Face18Shamus /Medical_Speech_Transcription_and_IntentThis dataset came from Kaggle and was contributed by Paul Mooney. https://www.kaggle.com/datasets/paultimothymooney/medical-speech-transcription-and-intent/data Context 8.5 hours of audio utterances paired with text for common medical symptoms. Content This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These… See the full description on the dataset page: https://huggingface.co/datasets/Shamus/Medical_Speech_Transcription_and_Intent.audio1K<n<10K3 likes61 downloads3y agoHugging Face19manhcuong2005 /vietnam_medical_noise_dataset Vietnam Medical Noise Dataset Bộ dữ liệu tiếng ồn môi trường và y tế tiếng Việt (~442 giờ audio, định dạng trực tiếp .wav kèm file metadata.parquet ở root). Thông tin dữ liệu: Format Audio: .WAV Sample Rate: 16,000 Hz Channels: 1 (Mono) Subtype: PCM_16 Metadata: Duy nhất file metadata.parquet ở root. Cấu trúc đường dẫn & tên file: <type_of_noise>/<index_folder>/<type_of_noise>-<folder>-<index_file>.wav (Ví dụ:… See the full description on the dataset page: https://huggingface.co/datasets/manhcuong2005/vietnam_medical_noise_dataset.audio10K<n<100K0 likes56 downloads2mo agoHugging Face20abar-uwc /medical-segmentation-dataset_v2audion<1K0 likes52 downloads1y agoHugging Face21havahavai /eka-medical-asr-evaluation-dataset Eka Medical ASR Evaluation Dataset Dataset Overview and Sourcing The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context. The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.audioautomatic-speech-recognition1K<n<10K1 likes50 downloads4mo agoHugging Face22zklehjruwehurfhqw /medical-speech-dataset Medical Speech Dataset A specialized speech dataset for healthcare AI applications featuring real medical terminology, clinical conversations, and domain-specific vocabulary. This dataset is curated from the complete-voiceai-speech-dataset and focuses specifically on medical domain speech data collected from real healthcare contexts. Dataset Overview Total audio files: 33 recordings Total duration: ~42 minutes Languages: English (native) + Global Medical… See the full description on the dataset page: https://huggingface.co/datasets/zklehjruwehurfhqw/medical-speech-dataset.audioautomatic-speech-recognitionn<1K0 likes48 downloads3mo agoHugging Face23juliasdata /medical-audio-sample-brazilian-portuguese Julia's Data: Brazilian Portuguese Medical Audio Sample Public sample of a Brazilian Portuguese medical audio dataset built for ASR, TTS, and conversational AI evaluation. This repository contains deidentified clinical source material transformed into five spoken content types and recorded by a human speaker. This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about 5.26 minutes of audio. Full dataset and commercial licensing: juliasdata.com Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.audioautomatic-speech-recognitionn<1K1 likes46 downloads6mo agoHugging Face24HieuNguyen203 /Vietnamese_Medical_Consultationaudioautomatic-speech-recognitionn<1K4 likes45 downloads1y agoHugging Face25HumynLabs /medical-opinion-english-audio Medical Opinion English Audio Dataset *This dataset contains intentionally low-quality (“B-grade”) data. It has been curated to include noisy, imperfect, or otherwise suboptimal samples for the purpose of testing model robustness and performance under degraded input conditions Text spoken by all participants: ""Doctor, another physician suggested my chest pain is stress-related, but I'm anxious. It feels like a heavy weight on my heart, and I struggle to breathe deeply. I'm scared.… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/medical-opinion-english-audio.audioaudio-classificationn<1K8 likes45 downloads1y agoHugging Face26MarieDeVox /english-vocal-medical-terminology-mini FREE PREVIEW: CLINICAL AI VOICE DATASET — MEDICAL TERMINOLOGY SERIES Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 48kHz Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, ethically sourced human voice data optimized specifically for training, benchmarking, and stress-testing clinical transcription models, medical speech-to-text (STT) pipelines, and health-tech conversational… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/english-vocal-medical-terminology-mini.audioautomatic-speech-recognitionn<1K2 likes45 downloads4mo agoHugging Face27shreyas1104 /medical-intent-audio-dataset-consolidatedaudio1K<n<10K0 likes42 downloads3y agoHugging Face28HumynLabs /medical-symptoms-english-audio Medical Symptoms English Audio Dataset *This dataset contains intentionally low-quality (“B-grade”) data. It has been curated to include noisy, imperfect, or otherwise suboptimal samples for the purpose of testing model robustness and performance under degraded input conditions Text spoken by all participants: "Doctor, I'm constantly tired, like a heavy fog I can't shake. Sharp headaches hit, worse at night, and sleep is tough. I get dizzy, and my stomach feels uneasy after meals.… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/medical-symptoms-english-audio.audioaudio-classificationn<1K5 likes41 downloads1y agoHugging Face29Dev372 /Medical_STT_Dataset_1.0audio1K<n<10K2 likes36 downloads2y agoHugging Face30yezarniko /medicinesaudio10K<n<100K0 likes32 downloads10mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.