datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
nursing-sentences-1
IntelMedica Nursing Sentences v1
Synthetic nursing-specific clinical documentation sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
40,247
Train
28,173
Validation
6,037
Test
6,037
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
Nursing
Category Distribution
Category
Train
Val
Test… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/nursing-sentences-1.physician-sentences-1
IntelMedica Physician Sentences v1
Synthetic physician-specific clinical documentation sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
107,906
Train
75,534
Validation
16,186
Test
16,186
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
Physician
Category Distribution
Category
Train… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/physician-sentences-1.general-medical-sentences-1
IntelMedica General Medical Sentences v1
Synthetic general medical terminology for broad clinical use sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
313,447
Train
219,412
Validation
47,017
Test
47,018
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
General
Category Distribution… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/general-medical-sentences-1.medical-tts-parquet-2-16khz
IntelMedica Medical TTS Dataset v2 (16kHz)
Description
Synthetic medical speech dataset for fine-tuning Whisper-based ASR models on clinical and nursing terminology. Contains 101,475 audio-text pairs totaling 184.1 hours of speech at 16 kHz mono, generated using Kokoro-82M TTS with 19 voices across three English accent groups.
This is v2 -- a companion to the v1 dataset (125,500 samples, ~257 hours). v2 focuses on terms from additional data sources (RxNorm API, FDA… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-2-16khz.medical-tts-parquet-1
IntelMedica Medical TTS Dataset v1 (24kHz) -- DEPRECATED
This dataset is deprecated. Please use intelmedica/medical-tts-parquet-1-16khz instead, which contains 125,500 samples (vs 10,000 here) at 16kHz sample rate optimized for ASR training.
Description
Synthetic medical speech dataset for training medical ASR models. This is the original 10K-sample version at 24kHz. It has been superseded by the 16kHz version with 12.5x more data.
Dataset Details
Samples:… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/medical-tts-parquet-1.
