CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01typhoon-ai /ThaiOCRBench ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai ThaiOCRBench is the first comprehensive benchmark for evaluating vision-language models (VLMs) on Thai text-rich visual understanding tasks.Inspired by OCRBench v2, it contains 2,808 human-annotated samples across 13 diverse tasks, including table parsing, chart understanding, full-page OCR, key information extraction, and visual question answering. The benchmark enables standardized zero-shot… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiOCRBench.imageimage-text-to-text1K<n<10K7 likes768 downloads10mo agoHugging Face02typhoon-ai /thai-dialect-isan-dataset Dataset Card for Thai Dialect Isan Speech Corpus Dataset Description This dataset contains audio recordings of Isan (Northeastern Thai) speech, paired with rich transcriptions and demographic metadata. It is designed to support Automatic Speech Recognition (ASR), dialect study, and text normalization tasks for the Isan language. The dataset features spontaneous responses to specific questions, covering two domains (General and Finance), recorded by speakers from different… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-dialect-isan-dataset.textautomatic-speech-recognition10K<n<100K5 likes471 downloads10mo agoHugging Face03typhoon-ai /typhoon-s-instruct-post-training Typhoon-S Instruct Post-Training Dataset Summary This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths. The dataset follows a two-part mixture philosophy: Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-instruct-post-training.texttext-generation100K<n<1M0 likes182 downloads8mo agoHugging Face04typhoon-ai /chatbot-arena-spoken-voicesaudio1K<n<10K0 likes142 downloads2y agoHugging Face05typhoon-ai /typhoon-s-sovereign-capability-dataset Typhoon-S Training Assets Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project. Datasets NitiBench (Legal Domain) nitibench_train_rl.parquet - RL training set (8,211 examples) nitibench_train_pretrain.parquet - Pretrain set (3,648 examples) nitibench_train_sft.parquet - SFT set (3,648 examples) nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-s-sovereign-capability-dataset.texttext-generation10K<n<100K0 likes124 downloads8mo agoHugging Face06typhoon-ai /ThaiSafetyBench ThaiSafetyBench ⚠️ Warning: This dataset contains harmful and toxic language. It is intended for academic purposes only. [ArXiv Paper] [Github] [Hugging Face Leaderboard 🤗] The ThaiSafetyBench dataset comprises 1,889 malicious Thai-language prompts across various categories. In addition to translated malicious prompts, it includes prompts tailored to Thai culture, offering deeper insights into culturally specific attacks. Note: The Monarchy type of harm has been removed from the… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ThaiSafetyBench.textquestion-answering1K<n<10K5 likes96 downloads7mo agoHugging Face07typhoon-ai /TVSpeech TVSpeech (Thai Video Speech) TVSpeech is a Thai speech recognition benchmark dataset specifically designed as a Robustness Track for evaluating ASR models on real-world, in-the-wild Thai audio. The dataset consists of 570 utterances (3.75 hours) curated from diverse public media channels on YouTube under the Creative Commons Attribution (CC-BY) license, representing challenging acoustic and semantic complexity found in natural speech. Dataset Overview Language: Thai… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/TVSpeech.audion<1K1 likes67 downloads8mo agoHugging Face08typhoon-ai /gigaspeech2-typhoon Gigaspeech2 Typhoon Project page | Paper | GitHub Gigaspeech2 Typhoon is a metadata-only reference dataset for Thai speech recognition benchmarking, specifically designed as an Accuracy Track for evaluating ASR models. The dataset contains 1,000 test samples with audio IDs and human transcriptions derived from the Gigaspeech2 corpus. Each audio_id directly links to the original Gigaspeech2 dataset, allowing users to download the corresponding audio. Dataset Overview… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/gigaspeech2-typhoon.textautomatic-speech-recognition1K<n<10K1 likes67 downloads4mo agoHugging Face09suwaimyo /typhoon-yolanda-tweets-fil-classification TyphoonYolandaTweets_fil_Classification Deduplicated copy of kornwtp/typhoon-yolanda-tweets-fil-classification. Splits split rows test 153 train 582 textn<1K0 likes65 downloads29d agoHugging Face10typhoon-ai /typhoon-audio-preview-data Typhoon Audio Preview Data Overview This dataset is for aligning speech/audio representations with textual representations. It consists of {audio, instruction, response} examples in both Thai and English. This repository provides {instruction, response} pairs that we generated for Typhoon-Audio training. We do not own the original data sources (e.g., CommonVoice, LibriSpeech, etc), and you can download these datasets from the original sources, or contact {potsawee… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-audio-preview-data.text1M<n<10M3 likes57 downloads2y agoHugging Face11wannaphong /typhoon-s-instruct-post-training Typhoon-S Instruct Post-Training Dataset Summary This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths. The dataset follows a two-part mixture philosophy: Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-instruct-post-training.texttext-generation100K<n<1M0 likes54 downloads6mo agoHugging Face12wannaphong /typhoon-s-sovereign-capability-dataset Typhoon-S Training Assets Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project. Datasets NitiBench (Legal Domain) nitibench_train_rl.parquet - RL training set (8,211 examples) nitibench_train_pretrain.parquet - Pretrain set (3,648 examples) nitibench_train_sft.parquet - SFT set (3,648 examples) nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.texttext-generation10K<n<100K0 likes53 downloads6mo agoHugging Face13PumeTu /typhoon-s-instruct-sft-single-turntext100K<n<1M0 likes44 downloads7mo agoHugging Face14puttatidam /typhoon-yolanda-tweets-fil-classificationtextn<1K0 likes32 downloads11d agoHugging Face15typhoon-ai /survey_social_value_th2025 Social Attitudes and Values Survey Dataset This dataset contains survey questions and responses designed to explore social attitudes and values among people in Thailand in 2025. It includes a comprehensive set of carefully crafted questions and collected responses aimed at facilitating research on social perspectives, values, cultural attitudes, as well as crowdsourcing algorithm research. This dataset was used to evaluate our proposed crowdsourcing algorithm [link to the paper to… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/survey_social_value_th2025.text1K<n<10K3 likes31 downloads2y agoHugging Face16kornwtp /typhoon-yolanda-tweets-fil-classificationtextn<1K0 likes29 downloads2y agoHugging Face17Thiraput01 /Math-reasoning-Opus4.6-typhoon-translated Dataset Card for Math-reasoning-Opus4.6-typhoon-translated Dataset Description This dataset is a Thai-translated version of the Crownelius/Opus-4.6-Reasoning-3300x dataset. It is designed to train and evaluate mathematical reasoning capabilities in Thai language models. The original English dataset was translated into Thai using the scb10x/typhoon-translate1.5-4b model, providing high-quality, localized mathematical problems, step-by-step thinking processes, and… See the full description on the dataset page: https://huggingface.co/datasets/Thiraput01/Math-reasoning-Opus4.6-typhoon-translated.texttext-generation1K<n<10K0 likes28 downloads6mo agoHugging Face18electricsheepasia /asia-3w-for-typhoon-hagupit-ruby 3W (Who does What Where) for Typhoon Hagupit (Ruby) Publisher: OCHA Philippines · Source: HDX · License: cc-by-igo · Updated: 2023-03-03 Abstract Response assistance matrix for Typhoon Hagupit as of 15 Jan 2015 Each row in this dataset represents first-level administrative unit observations. Temporal coverage is indicated by the unnamed_9 column(s). Geographic scope: PHL. Curated into ML-ready Parquet format by Electric Sheep Africa. Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-3w-for-typhoon-hagupit-ruby.texttabular-classificationn<1K0 likes26 downloads5mo agoHugging Face19typhoon-ai /typhoon-t1-3b-sci-fm-iclr-2025-exp-datasetgated Typhoon T1 3B ICLR 2025 SCI-FM Workshop Dataset Paper Title: Typhoon T1: An Open Thai Reasoning ModelVenue: Open Science for Foundation Models (SCI-FM), ICLR 2025Paper Link: https://arxiv.org/abs/2502.09042Authors: Pittawat Taveekitworachai, Potsawee Manakul, Kasima Tharnpipitchai, and Kunat Pipatanakul Dataset Details This dataset is part of the experiments in the paper Typhoon T1: An Open Thai Reasoning Model, accepted at SCI-FM, ICLR 2025. Please refer to the paper… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-t1-3b-sci-fm-iclr-2025-exp-dataset.text100K<n<1M0 likes25 downloads2y agoHugging Face20typhoon-ai /survey_questions_pew_claudetext1K<n<10K0 likes24 downloads2y agoHugging Face21electricsheepasia /asia-impact-data-casualties-and-damage-typhoon-haiyan-yolanda Impact data - casualties and damage - Typhoon Haiyan (Yolanda) Publisher: Netherlands Red Cross - 510 · Source: HDX · License: cc-by · Updated: 2021-09-23 Abstract Counts of damage and casualties from official data sets Each row in this dataset represents first-level administrative unit observations. Data was last updated on HDX on 2021-09-23. Geographic scope: PHL. Curated into ML-ready Parquet format by Electric Sheep Africa. Dataset Characteristics… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-impact-data-casualties-and-damage-typhoon-haiyan-yolanda.tabulartabular-classificationn<1K0 likes24 downloads5mo agoHugging Face22typhoon-ai /thai-tts-intelligiblity-eval Thai-TTS-Intelligibility-Eval Thai-TTS-Intelligibility-Eval is a curated evaluation set for measuring intelligibility of Thai Text-to-Speech (TTS) systems.All 290 items are short, challenging phrases that commonly trip up phoneme-to-grapheme converters, prosody models, or pronunciation lexicons.It is not intended for training; use it purely for benchmarking and regression tests. Dataset Summary Split #Utterances Description easy 50 Everyday phrases that… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai-tts-intelligiblity-eval.textn<1K0 likes23 downloads1y agoHugging Face23electricsheepasia /asia-demographics-typhoon-mangkhut-2018-twitter-data Typhoon Mangkhut 2018 Twitter Data Publisher: Qatar Computing Research Institute · Source: HDX · License: cc-by · Updated: 2024-09-13 Abstract This is a Twitter dataset collected during the typhoon Mangkhut 2018 in the Philippines. The data was collected, processed, and analyzed by the AIDR (http://aidr.qcri.org) platform using state of the art machine learning techniques. The data includes the reports of number of injured and dead people, infrastructure damage reports… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-demographics-typhoon-mangkhut-2018-twitter-data.tabulartabular-regressionn<1K0 likes19 downloads5mo agoHugging Face24electricsheepasia /asia-logistics-philippines-typhoon-mangkhut-priority-in Philippines - Typhoon Mangkhut - Priority Index (Estimated damage per municipality) Publisher: Netherlands Red Cross - 510 · Source: HDX · License: cc-by · Updated: 2024-06-05 Abstract Update 15/09 (POST-EVENT) Now that the typhoon has passed the country, the model is not run with forecasted wind speeds and typhoon track any more, but with actual estimated wind speeds and typhoon track. They come from the same source (Tropical Storm Risk - UCL), and are of the… See the full description on the dataset page: https://huggingface.co/datasets/electricsheepasia/asia-logistics-philippines-typhoon-mangkhut-priority-in.tabulartabular-regressionn<1K0 likes18 downloads5mo agoHugging Face25Dragon1218 /typhoon_QA_Settextn<1K0 likes15 downloads2y agoHugging Face26typhoon-ai /thaimos-tts-annotation ThaiMOS (TTS MOS Evaluaution) (Older) TTS synthesized speech with human evaluation Mean Opinion Score (MOS) Annotation was done by Datawow Annonation aspect: sound quality, pronunciation, silence This dataset was originally developed in 2024 based on older TTS models -- likely that patterns in this data may not be applicable to modern TTS systems. Annotation Guideline In directory pack, there are 12 directories each with 50 utterances. Each subject carefully listens to… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thaimos-tts-annotation.audio1K<n<10K0 likes13 downloads1y agoHugging Face27typhoon-ai /survey_questions_pew_claude_1Ktext1K<n<10K0 likes9 downloads2y agoHugging Face28sunhaozhepy /2022-typhoons-cmaimagen<1K0 likes8 downloads3y agoHugging Face29typhoon-ai /typhoon-t1-3b-research-preview-datagated Typhoon T1 3B Research Preview Data Overview This is a dataset used to train our first open reasoning model, Typhoon T1 (Research Preview): llama-3.2-typhoon-t1-3b-research-preview. It's available in Alpaca format ({instruction, input, output}), although input for all records is null. We acknowledge the owners of the original data sources. Please visit our technical blog for more details on the original data sources. Data Splits This dataset consists of 55… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-t1-3b-research-preview-data.text10K<n<100K1 likes8 downloads2y agoHugging Face30typhoon-ai /rl-code-math-v2gatedtext10K<n<100K0 likes8 downloads1y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.