datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.LT_Medical_S_corpusEnglish | Lietuvių
English
LT_Medical_S_corpus — Lithuanian Medical Speech Corpus
A Lithuanian speech dataset of medical dictation audio (radiology and family medicine) with transcriptions, speaker metadata, and word-level timestamps.
Columns
Column
Type
Description
audio
Audio
Audio
sentence
string
Ground truth transcription
duration_ms
int
Recording duration in milliseconds
medical_area
string
RADIOLOGIJA or SEIMOS
gender
string
MALE or… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_Medical_S_corpus.medical_noise_data
Medical Noise Dataset
Dữ liệu âm thanh tiếng Việt đã được tăng cường nhiễu (noise augmentation + RIR convolution).
Nguồn gốc
Audio gốc: dolly-vn/dolly-audio-1000h-vietnamese
Noise sources: YouTube extracted, WHAM!, Hospital ambient noise
RIR: Real RIR far-field (RVB2014)
Cách load
from datasets import load_dataset
ds = load_dataset("manhcuong2005/medical_noise_data")
Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.Medical-ASR-ENvoice_medicalvoice_medical_cut_smallvoice_medical_newmedical-speech-dataset
Medical Speech Dataset
A protocol sample. 11 contributors, recorded on their own devices in their own environments. Every clip carries origin region / variety, mother tongue, gender, device, OS, recording environment. Small by design — see What this is for below before downloading.
Hours
0.69
Clips
33
Speakers
11
Origin varieties
11
Languages
2
Configs
2
Speaker metadata
origin region / variety, mother tongue, gender, device, OS, recording environment… See the full description on the dataset page: https://huggingface.co/datasets/SilencioNetwork/medical-speech-dataset.medical-prescription-english-audio
Medical Prescription English Audio Dataset
Text spoken by all participants:
"Doctor, my third visit, and I'm hopeful but not fully better. Joint pain eased slightly, yet mornings are tough, and I'm exhausted. The last prescription helped a bit. Can we adjust it? I want to feel like myself again."
The dataset supports training and evaluation of models in:
Automatic Speech Recognition (ASR)
Emotional tone classification
Voice synthesis and generation
Emotion-aware conversational… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/medical-prescription-english-audio.medical-intent-audio-datasetmedv3-turkish-medical-asr
medv3 - Türkçe Sentetik Tıbbi Konuşma Korpusu
Türkçe tıbbi konuşma tanıma araştırmaları için hazırlanmış sentetik konuşma korpusudur.
Klinik cümleler Google Cloud Text-to-Speech Chirp 3 HD sesleriyle sentezlenmiştir.
Önemli uyarılar
Tüm kayıtlar sentetiktir (synthetic=true).
Gerçek hasta veya klinisyen sesi ve kişisel sağlık verisi içermez.
Tıbbi cihaz geliştirme onayı veya klinik doğrulama anlamına gelmez.
Klinik karar için değil, araştırma ve ASR… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/medv3-turkish-medical-asr.MedicalDatasetPaulMooney-Medical-ASR-Datamedical_spanisheka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.medical-speech-datasetMedical_Speech_Transcription_and_IntentThis dataset came from Kaggle and was contributed by Paul Mooney.
https://www.kaggle.com/datasets/paultimothymooney/medical-speech-transcription-and-intent/data
Context
8.5 hours of audio utterances paired with text for common medical symptoms.
Content
This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These… See the full description on the dataset page: https://huggingface.co/datasets/Shamus/Medical_Speech_Transcription_and_Intent.vietnam_medical_noise_dataset
Vietnam Medical Noise Dataset
Bộ dữ liệu tiếng ồn môi trường và y tế tiếng Việt (~442 giờ audio, định dạng trực tiếp .wav kèm file metadata.parquet ở root).
Thông tin dữ liệu:
Format Audio: .WAV
Sample Rate: 16,000 Hz
Channels: 1 (Mono)
Subtype: PCM_16
Metadata: Duy nhất file metadata.parquet ở root.
Cấu trúc đường dẫn & tên file:
<type_of_noise>/<index_folder>/<type_of_noise>-<folder>-<index_file>.wav
(Ví dụ:… See the full description on the dataset page: https://huggingface.co/datasets/manhcuong2005/vietnam_medical_noise_dataset.medical-segmentation-dataset_v2eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.medical-speech-dataset
Medical Speech Dataset
A specialized speech dataset for healthcare AI applications featuring real medical terminology, clinical conversations, and domain-specific vocabulary.
This dataset is curated from the complete-voiceai-speech-dataset and focuses specifically on medical domain speech data collected from real healthcare contexts.
Dataset Overview
Total audio files: 33 recordings
Total duration: ~42 minutes
Languages: English (native) + Global Medical… See the full description on the dataset page: https://huggingface.co/datasets/zklehjruwehurfhqw/medical-speech-dataset.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.Vietnamese_Medical_Consultationmedical-opinion-english-audio
Medical Opinion English Audio Dataset
*This dataset contains intentionally low-quality (“B-grade”) data. It has been curated to include noisy, imperfect, or otherwise suboptimal samples for the purpose of testing model robustness and performance under degraded input conditions
Text spoken by all participants:
""Doctor, another physician suggested my chest pain is stress-related, but I'm anxious. It feels like a heavy weight on my heart, and I struggle to breathe deeply. I'm scared.… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/medical-opinion-english-audio.english-vocal-medical-terminology-mini
FREE PREVIEW: CLINICAL AI VOICE DATASET — MEDICAL TERMINOLOGY SERIES
Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 48kHz
Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, ethically sourced human voice data optimized specifically for training, benchmarking, and stress-testing clinical transcription models, medical speech-to-text (STT) pipelines, and health-tech conversational… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/english-vocal-medical-terminology-mini.medical-intent-audio-dataset-consolidatedmedical-symptoms-english-audio
Medical Symptoms English Audio Dataset
*This dataset contains intentionally low-quality (“B-grade”) data. It has been curated to include noisy, imperfect, or otherwise suboptimal samples for the purpose of testing model robustness and performance under degraded input conditions
Text spoken by all participants:
"Doctor, I'm constantly tired, like a heavy fog I can't shake. Sharp headaches hit, worse at night, and sleep is tough. I get dizzy, and my stomach feels uneasy after meals.… See the full description on the dataset page: https://huggingface.co/datasets/HumynLabs/medical-symptoms-english-audio.Medical_STT_Dataset_1.0medicines
