datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
thai_exam
Dataset Card for Thai_Exam
ThaiExam is a Thai knowledge benchmarking dataset, consisting of multiple-choice questions from examinations in Thailand. The dataset was originally developed for evaluating Typhoon (Thai LLM). This dataset contains 5 splits corresponding to 5 examinations as follows:
ONET: The Ordinary National Educational Test (ONET) is an examination for students in Thailand. This dataset is based on the grade-12 ONET exam, comprising 4 subjects and each question has 5… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/thai_exam.ifeval-thDataset Card for IFEval-TH
IFEval-TH is a Thai version of IFEval. The original English instructions (https://huggingface.co/datasets/google/IFEval)
were translated into Thai using GPT-4, followed by a manual verification and correction process to ensure accuracy and content consistency.
Rows with poor translation quality or irrelevant context in Thai were removed from the dataset.
IFEval code modification
To use this dataset, you need to modify the IFEval code… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/ifeval-th.Pangpuriye-generated_by_typhoon
🤖 Super AI Engineer Development Program Season 4 - Pangpuriye House - Generated by Typhoon API
Pangpuriye's House Dataset - Generated Dataset from Typhoon API
This dataset is an output generated from the Typhoon API in the structure of SQL instruction for fine-tuning Pangpuriye's LLM. The dataset is set under cc-by-nc-2.0 license.
Content
The dataset consists of 16,125 rows of input, instruction, and output packed into a train set.
Each schema has its own CSV file… See the full description on the dataset page: https://huggingface.co/datasets/AIAT/Pangpuriye-generated_by_typhoon.translation_valPangpuriye-generated_by_typhoon
🤖 Super AI Engineer Development Program Season 4 - Pangpuriye House - Generated by Typhoon API
Pangpuriye's House Dataset - Generated Dataset from Typhoon API
This dataset is an output generated from the Typhoon API in the structure of SQL instruction for fine-tuning Pangpuriye's LLM. The dataset is set under cc-by-nc-2.0 license.
Content
The dataset consists of 16,125 rows of input, instruction, and output packed into a train set.
Each schema has its own CSV file… See the full description on the dataset page: https://huggingface.co/datasets/Highgroundbkk/Pangpuriye-generated_by_typhoon.Nexesenex__Llama_3.1_8b_Typhoon_v1.03-details
Dataset Card for Evaluation run of Nexesenex/Llama_3.1_8b_Typhoon_v1.03
Dataset automatically created during the evaluation run of model Nexesenex/Llama_3.1_8b_Typhoon_v1.03
The dataset is composed of 38 configuration(s), each one corresponding to one of the evaluated task.
The dataset has been created from 1 run(s). Each run can be found as a specific split in each configuration, the split being named using the timestamp of the run.The "train" split is always pointing to the… See the full description on the dataset page: https://huggingface.co/datasets/open-llm-leaderboard/Nexesenex__Llama_3.1_8b_Typhoon_v1.03-details.typhoon-r1-sft-data
Typhoon2 R1 Preview Data
Overview
This dataset is used to align Typhoon2-70B Instruct with DeepSeek-R1-70B Distill for the final merge. It is based on https://arxiv.org/abs/2502.09056 SFTv3 configuration.
Citation
@misc{pipatanakul2025adaptinglanguagespecificllmsreasoning,
title={Adapting Language-Specific LLMs to a Reasoning Model in One Day via Model Merging - An Open Recipe},
author={Kunat Pipatanakul and Pittawat Taveekitworachai and Potsawee… See the full description on the dataset page: https://huggingface.co/datasets/typhoon-ai/typhoon-r1-sft-data.preference-typhoonscb_mt_enth_2020_aqdf_1klivecodebench-th
