datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ThaiOCRBench
ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai
ThaiOCRBench is the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks.Inspired by OCRBench v2, it contains 2,808 human-annotated samples across 13 diverse tasks, including table parsing, chart understanding, full-page OCR, key information extraction, and visual question answering.
The benchmark enables standardized zero-shot… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiOCRBench.thai-dialect-isan-dataset
Dataset Card for Thai Dialect Isan Speech Corpus
Dataset Description
This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language.
The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.typhoon-s-instruct-post-training
Typhoon-S Instruct Post-Training
Dataset Summary
This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths.
The dataset follows a two-part mixture philosophy:
Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-instruct-post-training.chatbot-arena-spoken-voicestyphoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-sovereign-capability-dataset.ThaiSafetyBench
ThaiSafetyBench
⚠️ Warning: This dataset contains harmful and toxic language. It is intended for academic purposes only.
[ArXiv Paper] [Github] [Hugging Face Leaderboard 🤗]
The ThaiSafetyBench dataset comprises 1,889 malicious Thai-language prompts across various categories. In addition to translated malicious prompts, it includes prompts tailored to Thai culture, offering deeper insights into culturally specific attacks.
Note: The Monarchy type of harm has been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiSafetyBench.TVSpeech
TVSpeech (Thai Video Speech)
TVSpeech is a Thai speech recognition benchmark dataset specifically designed as a Robustness Track for evaluating ASR models on real-world, in-the-wild Thai audio. The dataset consists of 570 utterances (3.75 hours) curated from diverse public media channels on YouTube under the Creative Commons Attribution (CC-BY) license, representing challenging acoustic and semantic complexity found in natural speech.
Dataset Overview
Language: Thai… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/TVSpeech.gigaspeech2-typhoon
Gigaspeech2 Typhoon
Project page | Paper | GitHub
Gigaspeech2 Typhoon is a metadata-only reference dataset for Thai speech recognition benchmarking, specifically designed as an Accuracy Track for evaluating ASR models. The dataset contains 1,000 test samples with audio IDs and human transcriptions derived from the Gigaspeech2 corpus. Each audio_id directly links to the original Gigaspeech2 dataset, allowing users to download the corresponding audio.
Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/gigaspeech2-typhoon.typhoon-yolanda-tweets-fil-classification
TyphoonYolandaTweets_fil_Classification
Deduplicated copy of kornwtp/typhoon-yolanda-tweets-fil-classification.
Splits
split
rows
test
153
train
582
typhoon-audio-preview-data
Typhoon Audio Preview Data
Overview
This dataset is for aligning speech/audio representations with textual representations. It consists of {audio, instruction, response} examples in both Thai and English. This repository provides {instruction, response} pairs that we generated for Typhoon-Audio training. We do not own the original data sources (e.g., CommonVoice, LibriSpeech, etc), and you can download these datasets from the original sources, or contact {potsawee… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-audio-preview-data.typhoon-s-instruct-post-training
Typhoon-S Instruct Post-Training
Dataset Summary
This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths.
The dataset follows a two-part mixture philosophy:
Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-instruct-post-training.typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.typhoon-s-instruct-sft-single-turntyphoon-yolanda-tweets-fil-classificationsurvey_social_value_th2025
Social Attitudes and Values Survey Dataset
This dataset contains survey questions and responses designed to explore social attitudes and values among people in Thailand in 2025. It includes a comprehensive set of carefully crafted questions and collected responses aimed at facilitating research on social perspectives, values, cultural attitudes, as well as crowdsourcing algorithm research. This dataset was used to evaluate our proposed crowdsourcing algorithm [link to the paper to… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/survey_social_value_th2025.typhoon-yolanda-tweets-fil-classificationMath-reasoning-Opus4.6-typhoon-translated
Dataset Card for Math-reasoning-Opus4.6-typhoon-translated
Dataset Description
This dataset is a Thai-translated version of the Crownelius/Opus-4.6-Reasoning-3300x dataset. It is designed to train and evaluate mathematical reasoning capabilities in Thai language models.
The original English dataset was translated into Thai using the scb10x/typhoon-translate1.5-4b model, providing high-quality, localized mathematical problems, step-by-step thinking processes, and… See the full description on the dataset page: https://huggingface.co/datasets/Thiraput01/Math-reasoning-Opus4.6-typhoon-translated.asia-3w-for-typhoon-hagupit-ruby
3W (Who does What Where) for Typhoon Hagupit (Ruby)
Publisher: OCHA Philippines · Source: HDX · License: cc-by-igo · Updated: 2023-03-03
Abstract
Response assistance matrix for Typhoon Hagupit as of 15 Jan 2015
Each row in this dataset represents first-level administrative unit observations. Temporal coverage is indicated by the unnamed_9 column(s). Geographic scope: PHL.
Curated into ML-ready Parquet format by Electric Sheep Africa.
Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-3w-for-typhoon-hagupit-ruby.typhoon-t1-3b-sci-fm-iclr-2025-exp-dataset
Typhoon T1 3B ICLR 2025 SCI-FM Workshop Dataset
Paper Title: Typhoon T1: An Open Thai Reasoning ModelVenue: Open Science for Foundation Models (SCI-FM), ICLR 2025Paper Link: https://arxiv.org/abs/2502.09042Authors: Pittawat Taveekitworachai, Potsawee Manakul, Kasima Tharnpipitchai, and Kunat Pipatanakul
Dataset Details
This dataset is part of the experiments in the paper Typhoon T1: An Open Thai Reasoning Model, accepted at SCI-FM, ICLR 2025. Please refer to the paper… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-t1-3b-sci-fm-iclr-2025-exp-dataset.survey_questions_pew_claudeasia-impact-data-casualties-and-damage-typhoon-haiyan-yolanda
Impact data - casualties and damage - Typhoon Haiyan (Yolanda)
Publisher: Netherlands Red Cross - 510 · Source: HDX · License: cc-by · Updated: 2021-09-23
Abstract
Counts of damage and casualties from official data sets
Each row in this dataset represents first-level administrative unit observations. Data was last updated on HDX on 2021-09-23. Geographic scope: PHL.
Curated into ML-ready Parquet format by Electric Sheep Africa.
Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-impact-data-casualties-and-damage-typhoon-haiyan-yolanda.thai-tts-intelligiblity-eval
Thai-TTS-Intelligibility-Eval
Thai-TTS-Intelligibility-Eval is a curated evaluation set for measuring intelligibility of Thai Text-to-Speech (TTS) systems.All 290 items are short, challenging phrases that commonly trip up phoneme-to-grapheme converters, prosody models, or pronunciation lexicons.It is not intended for training; use it purely for benchmarking and regression tests.
Dataset Summary
Split
#Utterances
Description
easy
50
Everyday phrases that… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-tts-intelligiblity-eval.asia-demographics-typhoon-mangkhut-2018-twitter-data
Typhoon Mangkhut 2018 Twitter Data
Publisher: Qatar Computing Research Institute · Source: HDX · License: cc-by · Updated: 2024-09-13
Abstract
This is a Twitter dataset collected during the typhoon Mangkhut 2018 in the Philippines. The data was collected, processed, and analyzed by the AIDR (http://aidr.qcri.org) platform using state of the art machine learning techniques. The data includes the reports of number of injured and dead people, infrastructure damage reports… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-demographics-typhoon-mangkhut-2018-twitter-data.asia-logistics-philippines-typhoon-mangkhut-priority-in
Philippines - Typhoon Mangkhut - Priority Index (Estimated damage per municipality)
Publisher: Netherlands Red Cross - 510 · Source: HDX · License: cc-by · Updated: 2024-06-05
Abstract
Update 15/09 (POST-EVENT)
Now that the typhoon has passed the country, the model is not run with forecasted wind speeds and typhoon track any more, but with actual estimated wind speeds and typhoon track. They come from the same source (Tropical Storm Risk - UCL), and are of the… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-logistics-philippines-typhoon-mangkhut-priority-in.typhoon_QA_Setthaimos-tts-annotation
ThaiMOS (TTS MOS Evaluaution)
(Older) TTS synthesized speech with human evaluation
Mean Opinion Score (MOS)
Annotation was done by Datawow
Annonation aspect: sound quality, pronunciation, silence
This dataset was originally developed in 2024 based on older TTS models -- likely that patterns in this data may not be applicable to modern TTS systems.
Annotation Guideline
In directory pack, there are 12 directories each with 50 utterances.
Each subject carefully listens to… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thaimos-tts-annotation.survey_questions_pew_claude_1K2022-typhoons-cmatyphoon-t1-3b-research-preview-data
Typhoon T1 3B Research Preview Data
Overview
This is a dataset used to train our first open reasoning model, Typhoon T1 (Research Preview): llama-3.2-typhoon-t1-3b-research-preview. It's available in Alpaca format ({instruction, input, output}), although input for all records is null. We acknowledge the owners of the original data sources. Please visit our technical blog for more details on the original data sources.
Data Splits
This dataset consists of 55… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-t1-3b-research-preview-data.rl-code-math-v2
