datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.thai-contextasr-bench
Thai Contextual-Biasing ASR Benchmark
TL;DR
Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant?
Each utterance comes with a bias list: entity strings (brands, person names,
places) that may or may not be spoken in the audio, written the way a real Thai user
would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.vaja-thai
Vaja-Thai (วาจา) — Combined Thai TTS Dataset
A unified, quality-filtered Thai speech dataset combining multiple sources for
Text-to-Speech (TTS) research. All audio is resampled to 24 kHz WAV format.
Dataset Summary
Metric
Value
Total samples
289,916
Total hours
554.6h
Sampling rate
24,000 Hz
Format
WAV 16-bit PCM
Language
Thai (ภาษาไทย)
Sources
Source
Samples
Hours
License
Description
tsync2
1,823
3.7h
CC-BY-NC-SA-3.0
NECTEC… See the full description on the dataset page: https://huggingface.co/datasets/dubbing-ai/vaja-thai.Thai-Food-Ordering-Dataset
🍲 Thai Food Ordering Speech Dataset
A Specialized Speech Recognition Corpus for Thai Food Ordering and Restaurant Contexts
📌 Dataset Overview
The Thai Food Ordering Speech Dataset is a domain-specific audio dataset created to develop and enhance Automatic Speech Recognition (ASR) systems, specifically targeting Thai food ordering in food courts, street stalls, and dining environments.
In real-world food court operations, manual order… See the full description on the dataset page: https://huggingface.co/datasets/KittipatPaisanpudinun/Thai-Food-Ordering-Dataset.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-dialect-isan-dataset.thai-speech-20k
🎤 Thai Speech 20K — Thai Speech Dataset for VibeVoice Fine-tuning
ไทย/English — ชุดข้อมูลเสียงพูดภาษาไทย 20,000 ประโยค สำหรับ fine-tune โมเดล TTSSource: Derived from Thanarit/Thai-Voice-Test7
🇹🇭 Dataset เสียงภาษาไทย 20,000 ตัวอย่าง ผู้พูด 1 คน (SPK_00001)🇬🇧 20,000 Thai speech utterances, single speaker (SPK_00001)
🏷️ Source
Field
Detail
Original Dataset
Thanarit/Thai-Voice-Test7
Original Creator
Thanarit
Upstream SourceGigaSpeech2 (filtered… See the full description on the dataset page: https://huggingface.co/datasets/hotdogs/thai-speech-20k.thai-tts-dataset
Thai TTS Dataset — Unlimited Mindset
เสียงพูดภาษาไทยคุณภาพสูงจาก Podcast สำหรับ fine-tune โมเดล Text-to-Speech
Dataset Summary
Metric
Value
Total samples
4,771
Train / Eval
4,533 / 238 (95% / 5%)
Total duration
14.92 hours (895.2 min)
Episodes
40
Speaker
Single speaker (male)
Language
Thai (th)
Audio format
WAV 22050Hz mono 16-bit
Source
Unlimited Mindset จิตไม่จำกัด
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/doyze/thai-tts-dataset.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/kritsanan/thai-dialect-isan-dataset.
