datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/eka-medical-asr-evaluation-dataset.Medical_InterviewThe dataset was re-organized and used in the following paper. Please cite if you adopted the corpus in your work.
@inproceedings{liu2024post,
title={Post-decoder Biasing for End-to-End Speech Recognition of Multi-turn Medical Interview},
author={Liu, Heyang and Wang, Yanfeng and Wang, Yu},
booktitle={Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)},
pages={12917--12926},
year={2024}
}
LT_Medical_S_corpusEnglish | Lietuvių
English
LT_Medical_S_corpus — Lithuanian Medical Speech Corpus
A Lithuanian speech dataset of medical dictation audio (radiology and family medicine) with transcriptions, speaker metadata, and word-level timestamps.
Columns
Column
Type
Description
audio
Audio
Audio
sentence
string
Ground truth transcription
duration_ms
int
Recording duration in milliseconds
medical_area
string
RADIOLOGIJA or SEIMOS
gender
string
MALE or… See the full description on the dataset page: https://huggingface.co/datasets/VSSA-SDSA/LT_Medical_S_corpus.french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/Archime/french_tv_media_dataset_2026.FR_VocalCommands_medical_dictation
FR_VocalCommands_medical_dictation
Commandes vocales de dictee medicale en francais (PraxyDictee) : edition,
selection, formatage, majuscules et dictee ponctuee, commentees avec des mots
medicaux (molecules et pathologies des jeux precedents).
element
valeur
textes
22500
fichiers audio
22500
moteurs
coqui XTTS v2 (300 voix clonees), edge-tts, gTTS
augmentations
low_rms (volume -22 dB), low_rms_speedup (atempo 1.2)
format
WAV mono 16 kHz
QC
inference NeMo… See the full description on the dataset page: https://huggingface.co/datasets/PraxySante/FR_VocalCommands_medical_dictation.french_tv_media_dataset_2026
Dataset Card : A Multi-Domain Pseudo-Labeled ASR Corpus
Résumé (Abstract)
Ce corpus présente un jeu de données de reconnaissance automatique de la parole (ASR) en langue française, totalisant 97 heures d'audio annoté. Il est dérivé de flux de diffusion (broadcast) issus de France Télévisions, couvrant une diversité de domaines acoustiques et linguistiques (Information, Société, Divertissement, Documentaire, Sport). L'annotation a été réalisée via une méthodologie… See the full description on the dataset page: https://huggingface.co/datasets/madoss/french_tv_media_dataset_2026.Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/Synthetic-Medical-Speech-Dataset.medical_asr_recording_datasetData Source
Kaggle Medical Speech, Transcription, and Intent
Context
8.5 hours of audio utterances paired with text for common medical symptoms.
Content
This data contains thousands of audio utterances for common medical symptoms like “knee pain” or “headache,” totaling more than 8 hours in aggregate. Each utterance was created by individual human contributors based on a given symptom. These audio snippets can be used to train conversational agents in the medical field.
This Figure Eight… See the full description on the dataset page: https://huggingface.co/datasets/Hani89/medical_asr_recording_dataset.MediaSpeech
MediaSpeech
MediaSpeech is a dataset of Arabic, French, Spanish, and Turkish media speech built with the purpose of testing Automated Speech Recognition (ASR) systems performance. The dataset contains 10 hours of speech for each language provided.
The dataset consists of short speech segments automatically extracted from media videos available on YouTube and manually transcribed, with some pre-processing and post-processing.
Baseline models and WAV version of the dataset can be… See the full description on the dataset page: https://huggingface.co/datasets/ymoslem/MediaSpeech.voice_medicalai-researcher-roadmap-media
AI Researcher Roadmap Media
Optional video and subtitle assets for the
AI Researcher Roadmap
application.
Repository layout
manifest.json: file sizes and SHA-256 checksums used by the application.
videos/<stem>.mp4: lecture video.
subs/<stem>.<language>.vtt: subtitle tracks.
subs/<stem>.asr.<language>.vtt: ASR-generated subtitle tracks.
The application downloads only the selected lecture and its subtitle tracks.
Files are cached locally and can be played offline… See the full description on the dataset page: https://huggingface.co/datasets/kadirnar/ai-researcher-roadmap-media.MediBeng
Dataset Card for MediBeng
This dataset includes synthetic code-switched conversations in Bengali and English. It is designed to help train models for tasks like speech recognition (ASR), text-to-speech (TTS), and machine translation, focusing on bilingual code-switching in healthcare settings. The dataset is free to… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng.medv3-turkish-medical-asr
medv3 - Türkçe Sentetik Tıbbi Konuşma Korpusu
Türkçe tıbbi konuşma tanıma araştırmaları için hazırlanmış sentetik konuşma korpusudur.
Klinik cümleler Google Cloud Text-to-Speech Chirp 3 HD sesleriyle sentezlenmiştir.
Önemli uyarılar
Tüm kayıtlar sentetiktir (synthetic=true).
Gerçek hasta veya klinisyen sesi ve kişisel sağlık verisi içermez.
Tıbbi cihaz geliştirme onayı veya klinik doğrulama anlamına gelmez.
Klinik karar için değil, araştırma ve ASR… See the full description on the dataset page: https://huggingface.co/datasets/turkmedstt/medv3-turkish-medical-asr.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3,900+… See the full description on the dataset page: https://huggingface.co/datasets/KothapalliAnusha/eka-medical-asr-evaluation-dataset.media_queen_entertaiment_voices
Media Queen Entertainment Voices
Where the stars speak, and their stories come to life.
Media Queen Entertainment Voices is a massive, large-scale collection of 190,013 short audio segments (totaling approximately 125 hours of speech) derived from public videos by Media Queen Entertainment — a prominent digital media channel in Myanmar focused on celebrity news, lifestyle content, and in-depth interviews.
The source channel regularly features:
Interviews with artists, actors… See the full description on the dataset page: https://huggingface.co/datasets/freococo/media_queen_entertaiment_voices.eka-medical-asr-evaluation-dataset
Eka Medical ASR Evaluation Dataset
Dataset Overview and Sourcing
The Eka Medical ASR Evaluation Dataset enables comprehensive evaluation of automatic speech recognition systems designed to transcribe medical speech into accurate text—a fundamental component of any medical scribe system. This dataset captures the unique challenges of processing medical terminology, particularly branded drugs, which is specific to the Indian context.
The dataset comprises over 3… See the full description on the dataset page: https://huggingface.co/datasets/havahavai/eka-medical-asr-evaluation-dataset.medical-audio-sample-brazilian-portuguese
Julia's Data: Brazilian Portuguese Medical Audio Sample
Public sample of a Brazilian Portuguese medical audio dataset built for ASR,
TTS, and conversational AI evaluation. This repository contains deidentified
clinical source material transformed into five spoken content types and
recorded by a human speaker.
This sample includes 1 record, 20 aligned audio segments, 1 speaker, and about
5.26 minutes of audio.
Full dataset and commercial licensing: juliasdata.com
Commercial overview:… See the full description on the dataset page: https://huggingface.co/datasets/juliasdata/medical-audio-sample-brazilian-portuguese.medical-speech-dataset
Medical Speech Dataset
A specialized speech dataset for healthcare AI applications featuring real medical terminology, clinical conversations, and domain-specific vocabulary.
This dataset is curated from the complete-voiceai-speech-dataset and focuses specifically on medical domain speech data collected from real healthcare contexts.
Dataset Overview
Total audio files: 33 recordings
Total duration: ~42 minutes
Languages: English (native) + Global Medical… See the full description on the dataset page: https://huggingface.co/datasets/zklehjruwehurfhqw/medical-speech-dataset.Vietnamese_Medical_Consultationenglish-vocal-medical-terminology-mini
FREE PREVIEW: CLINICAL AI VOICE DATASET — MEDICAL TERMINOLOGY SERIES
Format: LJ Speech Standard Compliance | 24-bit Signed Linear PCM Mono WAV | 48kHz
Thank you for downloading this Developer Compatibility Sample Pack. This repository contains enterprise-grade, high-fidelity, ethically sourced human voice data optimized specifically for training, benchmarking, and stress-testing clinical transcription models, medical speech-to-text (STT) pipelines, and health-tech conversational… See the full description on the dataset page: https://huggingface.co/datasets/MarieDeVox/english-vocal-medical-terminology-mini.MEDISCO
Building MEDISCO: Indonesian Speech Corpus for Medical Domain
The dataset was published in the following paper:
Building MEDISCO: Indonesian Speech Corpus for Medical Domain (PDF | IEEEXplore)
Muhammad Reza Qorib and Mirna Adriani
2018 International Conference on Asian Language Processing (IALP)
Please look for the raw files (inside the "Files and versions" tab) as the dataset viewer parsed by Huggingface does not show the text transcript.
Please direct any questions to… See the full description on the dataset page: https://huggingface.co/datasets/mrqorib/MEDISCO.stt-mediaspeech-test
MediaSpeech — French test split
Split test de MediaSpeech (français) — extraits courts de médias (radio,
TV, podcasts) collectés par MTS AI. Empaqueté en Parquet shardé avec audio
FLAC embarqué.
Usage principal : benchmark ASR français (WER / CER) sur parole de
diffusion (broadcast / médias).
Contenu
2498 utterances (segments ~10 s)
Audio : FLAC 16 kHz mono PCM_16
Langue : français (fr)
Licence : CC-BY-4.0 (héritée de MediaSpeech / OpenSLR 108)
Durée totale : 10.00 h… See the full description on the dataset page: https://huggingface.co/datasets/ggfox00000/stt-mediaspeech-test.MedicalLessons
Dataset Card for Medical Lessons Speech Corpus
Dataset Description
Dataset Summary
The Medical Lessons Speech Corpus is a specialized audio dataset designed to benchmark Automatic Speech Recognition (ASR) systems on highly technical, domain-specific language. It contains a collection of medical lesson audio files accompanied by their corresponding transcriptions.
A defining characteristic of this dataset is its high density of complex medical… See the full description on the dataset page: https://huggingface.co/datasets/litosway/MedicalLessons.medicineenglish-medical-speech-dataset
English Medical Speech Dataset
This dataset has moved.
This dataset is no longer actively hosted here. It is now published and maintained on Mozilla Data Collective by Proxima AI:
👉 https://mozilladatacollective.com/datasets/cmrxix9a7001rl407d629khq8
Dataset Summary
11,563 audio-text pairs of spoken English medical content, covering clinical abbreviations, diagnoses, and terminology across specialties including cardiology, gastroenterology, neurology, obstetrics… See the full description on the dataset page: https://huggingface.co/datasets/mahwizzzz/english-medical-speech-dataset.MediBeng-FL
MediBeng-FL
A Federated Learning–Ready Extension of MediBeng
MediBeng-FL augments the original MediBeng dataset with rich,
realistic synthetic metadata designed to benchmark federated learning
(FL) algorithms on clinical Bengali-English code-switched speech.
Every sample retains the original six MediBeng columns — audio,
text, translation, speaker_name, utterance_pitch_mean,
utterance_pitch_std — and adds 9 new FL-metadata columns that
simulate the heterogeneity found in real… See the full description on the dataset page: https://huggingface.co/datasets/pr0mila-gh0sh/MediBeng-FL.French-Medical-Transcription-Benchmark
🩺 French Medical Transcription Evaluation Dataset
Ce dataset a été créé et ouvert à la communauté dans le cadre du développement R&D de LucioleScribe, la plateforme souveraine de transcription IA 100% locale, spécifiquement conçue pour les milieux médicaux et juridiques (compatibilité RGPD, HDS, et architectures Air-Gapped).
🔗 Découvrir LucioleScribe Édition Santé | ⚙️ Voir le Pipeline Technologique Local
📊 Présentation du Dataset
L'évaluation des modèles de… See the full description on the dataset page: https://huggingface.co/datasets/AWANNABY/French-Medical-Transcription-Benchmark.medical-terms-2025
Medical Terms 2025 — Medical ASR Benchmark
Entity-aware medical ASR benchmark — 50 hard rows with synthetic TTS audio of 2025 drug/condition terminology.
Prepared by Trelis Research. Watch more on Youtube or inquire about our custom voice AI (ASR/TTS) services here.
Source
84 manually curated terms from 2025 FDA/EMA/WHO primary sources. Each term has source_url, source_date, and source_quality. Sentences generated by Gemini 2.5 Flash. Audio by Kokoro TTS via Trelis Studio… See the full description on the dataset page: https://huggingface.co/datasets/Trelis/medical-terms-2025.Synthetic-Medical-Speech-Dataset
Synthetic Medical Speech Dataset
Overview
Synthetic Medical Speech Dataset is a synthetic dataset of audio–text pairs designed for developing and evaluating automatic speech recognition (ASR) models in the medical domain.The corpus contains thousands of short audio clips generated from medically relevant text using a text-to-speech (TTS) system.Each clip is paired with its corresponding transcript.Because all content is synthetically produced, the dataset does not contain… See the full description on the dataset page: https://huggingface.co/datasets/niamhtracey1/Synthetic-Medical-Speech-Dataset.general-medical-sentences-1
IntelMedica General Medical Sentences v1
Synthetic general medical terminology for broad clinical use sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
313,447
Train
219,412
Validation
47,017
Test
47,018
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
General
Category Distribution… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/general-medical-sentences-1.
