datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Thai-Food-Ordering-Dataset
🍲 Thai Food Ordering Speech Dataset
A Specialized Speech Recognition Corpus for Thai Food Ordering and Restaurant Contexts
📌 Dataset Overview
The Thai Food Ordering Speech Dataset is a domain-specific audio dataset created to develop and enhance Automatic Speech Recognition (ASR) systems, specifically targeting Thai food ordering in food courts, street stalls, and dining environments.
In real-world food court operations, manual order… See the full description on the dataset page: https://huggingface.co/datasets/KittipatPaisanpudinun/Thai-Food-Ordering-Dataset.neuro-parakeet-food
neuro-whisper-v1
Dataset Description
This is a synthetic dataset for German medical speech recognition, specifically designed for fine-tuning ASR models on neuro-oncology and neurology terminology. The dataset provides a comprehensive coverage of German medical terminology in the neurology and neuro-oncology domains.
Data Generation
Voice Data: Synthetically generated using Resemble AI Chatterbox TTS
Text Data: Medical text generated with Qwen/Qwen3-30B-A3B… See the full description on the dataset page: https://huggingface.co/datasets/NeurologyAI/neuro-parakeet-food.
