datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.chatbot-arena-spoken-voicesTVSpeech
TVSpeech (Thai Video Speech)
TVSpeech is a Thai speech recognition benchmark dataset specifically designed as a Robustness Track for evaluating ASR models on real-world, in-the-wild Thai audio. The dataset consists of 570 utterances (3.75 hours) curated from diverse public media channels on YouTube under the Creative Commons Attribution (CC-BY) license, representing challenging acoustic and semantic complexity found in natural speech.
Dataset Overview
Language: Thai… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/TVSpeech.thaimos-tts-annotation
ThaiMOS (TTS MOS Evaluaution)
(Older) TTS synthesized speech with human evaluation
Mean Opinion Score (MOS)
Annotation was done by Datawow
Annonation aspect: sound quality, pronunciation, silence
This dataset was originally developed in 2024 based on older TTS models -- likely that patterns in this data may not be applicable to modern TTS systems.
Annotation Guideline
In directory pack, there are 12 directories each with 50 utterances.
Each subject carefully listens to… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thaimos-tts-annotation.tts_arena_resynthesized
