datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Taiwanese-Minnan-Example-Sentences
Taiwanese Minnan Example Sentences
The dataset consists of a collection of example sentences designed to aid in recognizing Taiwanese Minnan (Taiwanese Hokkien) for automatic speech recognition (ASR) tasks. This dataset is sourced from the Ministry of Education in Taiwan and aims to provide valuable linguistic resources for researchers and developers working on speech recognition systems.
Dataset Features
Source: Ministry of Education, Taiwan (Sutian Resource Center)
Text:… See the full description on the dataset page: https://huggingface.co/datasets/sarahwei/Taiwanese-Minnan-Example-Sentences.Tech-Sentences-For-ASR-Training
TechVoice Dataset
Work in Progress – This dataset is actively being expanded with new recordings.
Dataset Statistics
Metric
Current
Target
Progress
Duration
38m 43s
5h 0m 0s
██░░░░░░░░░░░░░░░░░░ 12.9%
Words
10,412
50,000
████░░░░░░░░░░░░░░░░ 20.8%
Total Recordings: 205 samples
Total Characters: 74,312
A specialized speech dataset for fine-tuning Automatic Speech Recognition (ASR) models on technical and developer vocabulary. Contains human-recorded… See the full description on the dataset page: https://huggingface.co/datasets/danielrosehill/Tech-Sentences-For-ASR-Training.swiss-german-city-sentences_v2
Swiss German City Sentences v2
Synthetic Swiss German speech dataset with city name sentences across multiple dialects.
ghana-sentences-synth
Ghana Sentences — Synthetic Speech
Synthetic speech for ghananlpcommunity/ghana-sentences, voiced with the VoxCPM2-Ghana TTS model, one config per language. Each language uses in-language reference voices from ghana-speech (Ga uses Dangme voices). Generated with ghana-speech-datagen.
Total so far: 582 clips · 0.8 hours
Language
Code
Clips
Hours
Fante
fat
582
0.78
Each <lang>/ folder holds wavs/ + metadata.jsonl ({"audio","text"}).
nursing-sentences-1
IntelMedica Nursing Sentences v1
Synthetic nursing-specific clinical documentation sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
40,247
Train
28,173
Validation
6,037
Test
6,037
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
Nursing
Category Distribution
Category
Train
Val
Test… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/nursing-sentences-1.physician-sentences-1
IntelMedica Physician Sentences v1
Synthetic physician-specific clinical documentation sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
107,906
Train
75,534
Validation
16,186
Test
16,186
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
Physician
Category Distribution
Category
Train… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/physician-sentences-1.general-medical-sentences-1
IntelMedica General Medical Sentences v1
Synthetic general medical terminology for broad clinical use sentences for training medical Automatic Speech Recognition (ASR) models. Part of the IntelMedica open-source medical AI initiative.
Overview
Stat
Value
Total rows
313,447
Train
219,412
Validation
47,017
Test
47,018
Split ratio
70 / 15 / 15 (stratified by category)
Language
English
Audience
General
Category Distribution… See the full description on the dataset page: https://huggingface.co/datasets/intelmedica/general-medical-sentences-1.
