datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai-aligner-bench
Thai Aligner Bench
🚧 Development in progress.
How accurately can a forced aligner place Thai token and word boundaries in
speech? This is a self-contained benchmark: one Python file
(aligner_bench.py) plus 1,572 clips of Thai speech with frame-exact timing
ground truth. No Thai NLP stack or other code is needed — just
numpy soundfile torch torchaudio transformers.
The ground truth is what makes the dataset useful: the audio was rendered by a
TTS model whose duration predictor… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-aligner-bench.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.thai-contextasr-bench
Thai Contextual-Biasing ASR Benchmark
TL;DR
Does your Thai ASR system actually use the context you give it (e.g., a list of names, custom words from your own dictionary) — and does it hallucinate when the context is irrelevant?
Each utterance comes with a bias list: entity strings (brands, person names,
places) that may or may not be spoken in the audio, written the way a real Thai user
would write them — one list, mixed Thai and Latin script. A good system does… See the full description on the dataset page: https://huggingface.co/datasets/wayu-ai/thai-contextasr-bench.burmese-speech-refined-openslr-80
Burmese Speech Refined OpenSLR-80
Summary
This dataset is a speech dataset developed based on the original OpenSLR Dataset (SLR80), with the text and audio data carefully reviewed and further refined for Burmese language applications.
In the original OpenSLR Dataset, the Burmese text was transcribed based on how the words were pronounced in the corresponding audio recordings. In this dataset, the original audio and text data were used as a reference, and the text… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/burmese-speech-refined-openslr-80.vaja-thai
Vaja-Thai (วาจา) — Combined Thai TTS Dataset
A unified, quality-filtered Thai speech dataset combining multiple sources for
Text-to-Speech (TTS) research. All audio is resampled to 24 kHz WAV format.
Dataset Summary
Metric
Value
Total samples
289,916
Total hours
554.6h
Sampling rate
24,000 Hz
Format
WAV 16-bit PCM
Language
Thai (ภาษาไทย)
Sources
Source
Samples
Hours
License
Description
tsync2
1,823
3.7h
CC-BY-NC-SA-3.0
NECTEC… See the full description on the dataset page: https://huggingface.co/datasets/dubbing-ai/vaja-thai.myanmar-shopvoice
Myanmar Shopvoice Dataset (300 Sample)
This is a randomly sampled subset of 300 Myanmar shop voice audio chunks for Whisper fine-tuning and evaluation.
Dataset Details
Total Audio Chunks: 300
Audio Format: 16 kHz WAV, mono
Language: Myanmar (Burmese)
Features
audio: Audio feature (16 kHz WAV audio player)
transcription: Myanmar sentence transcription text
source_audio: Source continuous recording file name
source_line: Index line of transcript… See the full description on the dataset page: https://huggingface.co/datasets/thantzinphyo/myanmar-shopvoice.Thai-Food-Ordering-Dataset
🍲 Thai Food Ordering Speech Dataset
A Specialized Speech Recognition Corpus for Thai Food Ordering and Restaurant Contexts
📌 Dataset Overview
The Thai Food Ordering Speech Dataset is a domain-specific audio dataset created to develop and enhance Automatic Speech Recognition (ASR) systems, specifically targeting Thai food ordering in food courts, street stalls, and dining environments.
In real-world food court operations, manual order… See the full description on the dataset page: https://huggingface.co/datasets/KittipatPaisanpudinun/Thai-Food-Ordering-Dataset.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/thai-dialect-isan-dataset.thai-speech-20k
🎤 Thai Speech 20K — Thai Speech Dataset for VibeVoice Fine-tuning
ไทย/English — ชุดข้อมูลเสียงพูดภาษาไทย 20,000 ประโยค สำหรับ fine-tune โมเดล TTSSource: Derived from Thanarit/Thai-Voice-Test7
🇹🇭 Dataset เสียงภาษาไทย 20,000 ตัวอย่าง ผู้พูด 1 คน (SPK_00001)🇬🇧 20,000 Thai speech utterances, single speaker (SPK_00001)
🏷️ Source
Field
Detail
Original Dataset
Thanarit/Thai-Voice-Test7
Original Creator
Thanarit
Upstream SourceGigaSpeech2 (filtered… See the full description on the dataset page: https://huggingface.co/datasets/hotdogs/thai-speech-20k.thaha-research-data2-v2
Nepali Speech Dataset (YouTube-sourced)
441 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 441 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2.thai-tts-dataset
Thai TTS Dataset — Unlimited Mindset
เสียงพูดภาษาไทยคุณภาพสูงจาก Podcast สำหรับ fine-tune โมเดล Text-to-Speech
Dataset Summary
Metric
Value
Total samples
4,771
Train / Eval
4,533 / 238 (95% / 5%)
Total duration
14.92 hours (895.2 min)
Episodes
40
Speaker
Single speaker (male)
Language
Thai (th)
Audio format
WAV 22050Hz mono 16-bit
Source
Unlimited Mindset จิตไม่จำกัด
Usage
from datasets import load_dataset
ds =… See the full description on the dataset page: https://huggingface.co/datasets/doyze/thai-tts-dataset.thaha-research-data1
Nepali Speech Dataset (YouTube-sourced)
59 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 59 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data1.thaha-research-data2
Nepali Speech Dataset (YouTube-sourced)
63 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 63 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2.thaha-research-data2-v2-v2
Nepali Speech Dataset (YouTube-sourced)
76 labeled speech segments, split by channel (not by individual video) so the same speaker/recording can't appear in more than one split.
Splits
train: 76 segments
validation: 0 segments
test: 0 segments
Transcript columns — read this before training
Each segment carries three transcript variants. They are NOT interchangeable:
text_original — the YouTube caption text (if any) that overlapped this segment's… See the full description on the dataset page: https://huggingface.co/datasets/lilgoose777/thaha-research-data2-v2-v2.Thai-H2H-Call-center-audio-with-human-transcriptionThis dataset contains natural Thai-language conversations between human agents and human customers, designed to reflect realistic call center interactions across multiple domains. All conversations are conducted through unscripted role-playing, allowing for spontaneous and dynamic exchanges that closely mirror real-world scenarios.
🗣️ Speech Type: Human-to-human dialogues simulating customer-agent interactions.
🎭 Style: Non-scripted, spontaneous role-playing to capture authentic speech… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-H2H-Call-center-audio-with-human-transcription.thai-elderly-speech
Thai Elderly Speech Dataset (Combined Evaluation Set)
This dataset contains evaluation recordings for Thai elderly speech, combined from Healthcare and Smarthome domains.
Dataset Structure
After extracting Combined_Dataset.zip, the directory structure will look like this:
Combined_Dataset/
├── Train/ # 80% of the dataset (15,360 files)
│ ├── Accuracy_100/ # Files with 100% baseline accuracy
│ ├── Accuracy_50_99/ #… See the full description on the dataset page: https://huggingface.co/datasets/Phonsiri/thai-elderly-speech.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/kritsanan/thai-dialect-isan-dataset.synthetic-atc-speech
Synthetic ATC Speech
Synthetic English air-traffic-control speech created for research on robust
automatic speech recognition. The dataset contains 276,304 generated
utterances from 15,660 unique ATC transcripts.
The dataset accompanies:
Contrastive Regularization for Accent-Robust ASR
Robust ATC ASR code
UWB SupCon Hybrid model
UWB+ATCOSIM SupCon Hybrid model
Dataset Structure
The dataset provides one training split packaged as uncompressed WebDataset TAR… See the full description on the dataset page: https://huggingface.co/datasets/ThaiVanPhat95/synthetic-atc-speech.Thai-human-to-machine-call-center-audio-with-scriptThis dataset features natural Thai-language conversations between human speakers and machine agents, simulating real-world call center interactions across a variety of customer service domains. All dialogues are non-scripted and performed as role-play scenarios, capturing spontaneous, realistic exchanges.
🗣️ Speech Type: Human-to-machine conversations, simulating AI agents and human customers' dialogues.
🎭 Style: Spontaneous, unscripted role-playing, designed to reflect actual customer… See the full description on the dataset page: https://huggingface.co/datasets/AIxBlock/Thai-human-to-machine-call-center-audio-with-script.17-minute-world-languages_thai
[!NOTE]
Dataset origin: https://www.17-minute-world-languages.com/fr/thaï/
Site à scrapper
thaint_85k-mls-french-audio-datasetCe répertoire est vide, il a été créé pour améliorer le référencement du jeu de données thaint/85k-mls-french-audio-dataset.
