datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.Medical_InterviewThe dataset was re-organized and used in the following paper. Please cite if you adopted the corpus in your work.
@inproceedings{liu2024post,
title={Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview},
author={Liu, Heyang and Wang, Yanfeng and Wang, Yu},
booktitle={Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)},
pages={12917--12926},
year={2024}
}
LT_Medical_S_corpusEnglish | Lietuvių
English
LT_Medical_S_corpus — Lithuanian Medical Speech Corpus
A Lithuanian speech dataset of medical dictation audio (radiology and family medicine) with transcriptions, speaker metadata, and word-level timestamps.
Columns
Column
Type
Description
audio
Audio
Audio
sentence
string
Ground truth transcription
duration_ms
int
Recording duration in milliseconds
medical_area
string
RADIOLOGIJA or SEIMOS
gender
string
MALE or… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_Medical_S_corpus.FR_VocalCommands_medical_dictation
FR_VocalCommands_medical_dictation
Commandes vocales de dictee medicale en francais (PraxyDictee) : edition,
selection, formatage, majuscules et dictee ponctuee, commentees avec des mots
medicaux (molecules et pathologies des jeux precedents).
element
valeur
textes
22500
fichiers audio
22500
moteurs
coqui XTTS v2 (300 voix clonees), edge-tts, gTTS
augmentations
low_rms (volume -22 dB), low_rms_speedup (atempo 1.2)
format
WAV mono 16 kHz
QC
inference NeMo… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/FR_VocalCommands_medical_dictation.Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.medical_asr_recording_datasetData Source
Kaggle Medical Speech, Transcription, and Intent
Context
8.5 hours of audio utterances paired with text for common medical symptoms.
Content
This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These audio snippets can be used to train conversational agents in the medical field.
This Figure Eight… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/medical_asr_recording_dataset.voice_medicalmedv3-turkish-medical-asr
medv3 - Türkçe Sentetik Tıbbi Konuşma Korpusu
Türkçe tıbbi konuşma tanıma araştırmaları için hazırlanmış sentetik konuşma korpusudur.
Klinik cümleler Google Cloud Text-to-Speech Chirp 3 HD sesleriyle sentezlenmiştir.
Önemli uyarılar
Tüm kayıtlar sentetiktir (synthetic=true).
Gerçek hasta veya klinisyen sesi ve kişisel sağlık verisi içermez.
Tıbbi cihaz geliştirme onayı veya klinik doğrulama anlamına gelmez.
Klinik karar için değil, araştırma ve ASR… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/medv3-turkish-medical-asr.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.medical-speech-dataset
Medical Speech Dataset
A specialized speech dataset for healthcare AI applications featuring real medical terminology, clinical conversations, and domain-specific vocabulary.
This dataset is curated from the complete-voiceai-speech-dataset and focuses specifically on medical domain speech data collected from real healthcare contexts.
Dataset Overview
Total audio files: 33 recordings
Total duration: ~42 minutes
Languages: English (native) + Global Medical… See the full description on the dataset page: https://huggingface.co/datasets/zklehjruwehurfhqw/medical-speech-dataset.Vietnamese_Medical_Consultationenglish-vocal-medical-terminology-mini
FREE PREVIEW: CLINICAL AI VOICE DATASET — MEDICAL TERMINOLOGY SERIES
Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 48kHz
Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, ethically sourced human voice data optimized specifically for training, benchmarking, and stress-testing clinical transcription models, medical speech-to-text (STT) pipelines, and health-tech conversational… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/english-vocal-medical-terminology-mini.MedicalLessons
Dataset Card for Medical Lessons Speech Corpus
Dataset Description
Dataset Summary
The Medical Lessons Speech Corpus is a specialized audio dataset designed to benchmark Automatic Speech Recognition (ASR) systems on highly technical, domain-specific language. It contains a collection of medical lesson audio files accompanied by their corresponding transcriptions.
A defining characteristic of this dataset is its high density of complex medical… See the full description on the dataset page: https://huggingface.co/datasets/litosway/MedicalLessons.medicineenglish-medical-speech-dataset
English Medical Speech Dataset
This dataset has moved.
This dataset is no longer actively hosted here. It is now published and maintained on Mozilla Data Collective by Proxima AI:
👉 https://mozilladatacollective.com/datasets/cmrxix9a7001rl407d629khq8
Dataset Summary
11,563 audio-text pairs of spoken English medical content, covering clinical abbreviations, diagnoses, and terminology across specialties including cardiology, gastroenterology, neurology, obstetrics… See the full description on the dataset page: https://huggingface.co/datasets/mahwizzzz/english-medical-speech-dataset.French-Medical-Transcription-Benchmark
🩺 French Medical Transcription Evaluation Dataset
Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped).
🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local
📊 Présentation du Dataset
L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.medical-terms-2025
Medical Terms 2025 — Medical ASR Benchmark
Entity-aware medical ASR benchmark — 50 hard rows with synthetic TTS audio of 2025 drug/condition terminology.
Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here.
Source
84 manually curated terms from 2025 FDA/EMA/WHO primary sources. Each term has source_url, source_date, and source_quality. Sentences generated by Gemini 2.5 Flash. Audio by Kokoro TTS via Trelis Studio… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/medical-terms-2025.Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/niamhtracey1/Synthetic-Medical-Speech-Dataset.general-medical-sentences-1
IntelMedica General Medical Sentences v1
Synthetic general medical terminology for broad clinical use sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
313,447
Train
219,412
Validation
47,017
Test
47,018
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
General
Category Distribution… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/general-medical-sentences-1.medical-conversation-deepgram-diarized-272
Medical Conversation Deepgram Diarized 272
This dataset package contains 272 simulated doctor-patient medical interview audio files, the original clean transcripts, and Deepgram-generated diarized timestamped transcripts.
The source audio/transcripts come from:
Fareez, F., Parikh, T., Wavell, C. et al. A dataset of simulated patient-physician medical interviews with a focus on respiratory cases. Scientific Data 9, 313 (2022). https://doi.org/10.1038/s41597-022-01423-1
Original… See the full description on the dataset page: https://huggingface.co/datasets/wmatbooth/medical-conversation-deepgram-diarized-272.speech-simulated-medical-exams
Speech Simulated Medical Exams
Simulated patient-physician medical exam conversations with rich speech metadata annotations. Built for training single-step ASR models that transcribe and annotate multiple concepts simultaneously, including speaker changes, emotions, intents, and roles.
Dataset Details
Property
Value
Examples
25,706
Language
English
Audio
16 kHz WAV
Source
Simulated medical interviews (respiratory focus)
Features… See the full description on the dataset page: https://huggingface.co/datasets/WhissleAI/speech-simulated-medical-exams.voxmind-medical-action-state-stress
VoxMind Medical Action-State Stress
This dataset contains a synthetic Chinese server-TTS benchmark for executable reliability in spoken medical front-desk tool agents.
The main split is Argument-Binding Stress: 140 cases targeting appointment IDs, dates, departments, patient relation, final-after-partial turns, and pending cancellation. It is designed to test whether a spoken tool agent selects the correct action, tool, arguments, and final state before side-effecting execution.… See the full description on the dataset page: https://huggingface.co/datasets/v1tavitavita/voxmind-medical-action-state-stress.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/priyamallojjala/eka-medical-asr-evaluation-dataset.icu_medicationsmedical-tts-parquet-2-16khz
IntelMedica Medical TTS Dataset v2 (16kHz)
Description
Synthetic medical speech dataset for fine-tuning Whisper-based ASR models on clinical and nursing terminology. Contains 101,475 audio-text pairs totaling 184.1 hours of speech at 16 kHz mono, generated using Kokoro-82M TTS with 19 voices across three English accent groups.
This is v2 -- a companion to the v1 dataset (125,500 samples, ~257 hours). v2 focuses on terms from additional data sources (RxNorm API, FDA… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-2-16khz.medical_asr_aligned_27_04
Medical ASR Aligned Dataset
Aligned Kazakh medical speech dataset from the «ТЕЛЕДӘРІГЕР» (TeleDoctor) TV program on Qazaqstan National Channel.
Dataset Description
Audio-transcript aligned segments of Kazakh-language medical TV broadcasts. Each segment contains the original audio chunk, ASR transcription, human reference transcription, and Character Error Rate (CER).
Only segments with CER < 25% are included.
Stats
Split
Segments
Avg CER
Avg Duration… See the full description on the dataset page: https://huggingface.co/datasets/RakhatM/medical_asr_aligned_27_04.hebrew_medical_audio
Hebrew Medical Audio Dataset
Overview
This dataset is published by Verbit.ai and contains over one thousand audio recordings of invented clinical summaries by 41 different speakers. Each recording is in Hebrew and represents a summary of a patient's visit, providing valuable insights into clinical interactions, diagnosis, treatment plans, and follow-up procedures. The recordings do not contain any personal or private information.
Dataset Structure
Audio… See the full description on the dataset page: https://huggingface.co/datasets/verbit/hebrew_medical_audio.medical-tts-parquet-1
IntelMedica Medical TTS Dataset v1 (24kHz) -- DEPRECATED
This dataset is deprecated. Please use intelmedica/medical-tts-parquet-1-16khz instead, which contains 125,500 samples (vs 10,000 here) at 16kHz sample rate optimized for ASR training.
Description
Synthetic medical speech dataset for training medical ASR models. This is the original 10K-sample version at 24kHz. It has been superseded by the 16kHz version with 12.5x more data.
Dataset Details
Samples:… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-1.
