datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
gpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.SteeringSafety
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
🎯 Overview
SteeringSafety is a benchmark suite for evaluating representation steering methods across multiple safety perspectives.
SteeringSafety provides:
📊 A collection of 17 datasets including 7 perspectives for measuring safety behaviors.
🔧 A modular code framework implementing the taxonomy of training-free steering methods with standardized, interchangeable… See the full description on the dataset page: https://huggingface.co/datasets/WangResearchLab/SteeringSafety.WangchanLION-Web
Citation
@misc{phatthiyaphaibun2025mangosteenopenthaicorpus,
title={Mangosteen: An Open Thai Corpus for Language Model Pretraining},
author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong},
year={2025},
eprint={2507.14664},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2507.14664},
}
We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.WangchanThaiInstructWangchanThaiInstruct: An instruction-following Dataset for Culture-Aware, Multitask, and Multi-domain Evaluation in Thai (EMNLP'25)
WangchanThaiInstruct is a human-authored Thai dataset that improves instruction-following in low-resource settings, capturing cultural and domain-specific nuances across four domains and seven task types.
The evaluate code can be found at this github link
@inproceedings{limkonchotiwat2025thaiinstruct,
title = {WangchanThaiInstruct: An Instruction-Following… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanThaiInstruct.wikipedia-zh-mnbvc
zhwiki-mnbvc
分项目:爬取并处理中文维基百科语料
数据时间:202302-202305 (持续更新)
主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC
该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1
并且使用组员开发的去重工具进行数据格式化。
总行数(样本): 10,754,146
一个示例:
{
"文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt",
"是否待查文件": false,
"是否重复文件": false,
"文件大小": 558,
"simhash": 14363740497821204542,
"最长段落长度": 142,
"段落数": 6,
"去重段落数": 6,
"低质量段落数": 0,
"段落": [
{… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.wikipedia-monthly
🚀 Wikipedia Monthly
Last updated: March 14, 2026, 21:06 UTC
This repository provides monthly, multilingual dumps of Wikipedia, processed and prepared for easy use in NLP projects.
📊 Current Statistics
Metric
Current Export (March 2026)
All Exports (Total)
Languages
343
361
Articles
62.8M
62.8M
Usage
Load any language with a single line of code using 🤗 datasets.
latest always refers to the most recent dump, while dated configs refer to… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/wikipedia-monthly.OmniEAR
OmniEAR Expert Trajectory Dataset
Dataset Summary
The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.WangchanX-Legal-ThaiCCL-RAG
🏛️ WangchanX-Legal-ThaiCCL-RAG
[Technical Report]
The WangchanX-Legal-ThaiCCL-RAG dataset supports the development Retrieval-Augmented Generation (RAG) for Thai Legal question answering. This dataset is allows developers to finetune both retrieval model - to better retrieve relevant law section, and Large Language Model (LLM) - for instruction tuning. Our dataset supports Corporate and Commercial Law (thus ThaiCCL name). See legislation section for more details on supported… See the full description on the dataset page: https://huggingface.co/datasets/airesearch/WangchanX-Legal-ThaiCCL-RAG.KhanomTanLLM-pretrained-dataset
KhanomTanLLM pretrained dataset
This daataset collect all raw text for pretraining LLM.
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Tokens
53,376,211,711 Tokens
English: 31,629,984,243 Tokens
Thai: 12,785,565,497 Tokens
Code: 8,913,084,300 Toekns
Parallel data: 190,310,686 Tokens
Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer
All subset
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.WangchanLION-Curated
Citation
@misc{phatthiyaphaibun2025mangosteenopenthaicorpus,
title={Mangosteen: An Open Thai Corpus for Language Model Pretraining},
author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong},
year={2025},
eprint={2507.14664},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2507.14664},
}
We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Curated.PyCoder
PyCoder
This repository contains the dataset for the paper Syntax-Aware On-the-Fly Code Completion
The sample code to run the model can be found in directory: "assets/notebooks/inference.ipynb" in our GitHub: https://github.com/awsm-research/pycoder.
PyCoder is an auto code completion model which leverages a Multi-Task Training technique (MTT) to cooperatively
learn the code prediction task and the type prediction task. For the type prediction
task, we propose to leverage the… See the full description on the dataset page: https://huggingface.co/datasets/Wannita/PyCoder.WanJuan-Korean
💡 Introduction
WanJuan-Korean(万卷丝路-韩语) corpus, with a volume exceeding 280GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Korean.LumiRAG
LumiRAG: A Unified Multimodal RAG Large Model Bridging Text and Image Retrieval
Dataset Description
This is the dataset used in the paper "LumiRAG: A Unified Multimodal RAG Large Model Bridging Text and Image Retrieval".
Languages
English
Data Fields
{
"question"
"answer"
"imagepath"
"reward_method"
"data_source"
}
bridgeTLDR: This dataset is a real-world math tutoring dataset from the NAACL 2024 paper ``Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes''.
The dataset targets scenarios where the student makes a math mistake.
c_h is the conversation history
c_r is the original tutor's response
c_r_ is the experienced teacher's response
Optionally, there is other interesting metadata from our Bridge method:
e is the student error type that the experienced… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/bridge.step3.5-flash-sft
Step-3.5-Flash-SFT
Step-3.5-Flash-SFT is a general-domain supervised fine-tuning release for chat models.
This repository keeps the full training interface in one place:
json/: canonical raw training data
tokenizers/: tokenizer snapshots for Step-3.5-Flash and Qwen3, released to preserve chat-template alignment
compiled/: tokenizer-specific compiled shards for StepTronOSS training
Data Format
Each raw shard is a JSON file whose top level is a list of examples.… See the full description on the dataset page: https://huggingface.co/datasets/wangrongsheng/step3.5-flash-sft.mind2dialogue
Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States
This dataset contains synthetic conversations between a simulated user and an Oracle assistant. Both use the same evolving user state.
Paper · Code · Project page
from datasets import load_dataset
dataset = load_dataset(
"wannabeyourfriend-hf/mind2dialogue", "conversations", split="train"
)
Citation
@misc{wang2026mind2dialoguetraininghumanawarelanguage… See the full description on the dataset page: https://huggingface.co/datasets/wannabeyourfriend-hf/mind2dialogue.WanJuan-Thai
💡 Introduction
WanJuan-Thai (万卷丝路-泰语) corpus, with a volume exceeding 155GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Thai.PolyReal
PolyReal: A Benchmark for Real-World Polymer Science Workflows
Evaluating Multimodal Large Language Models Across the Full Lifecycle of Polymer Science
Contents
data/train.jsonl: Hugging Face viewer-friendly training split.
PolyReal.json: main dataset file.
ref/: referenced images and CSV files used by Path entries in the dataset.
🔥 Overview
PolyReal is a multimodal benchmark for real-world polymer science workflows.It is… See the full description on the dataset page: https://huggingface.co/datasets/wanhaoliu/PolyReal.classical-tcm-canon
Classical Chinese Medicine Canon — 中医经典文本数据集 (v1)
A curated Traditional Chinese Medicine (TCM) text dataset: clean, full-text
digitizations of the foundational Chinese medicine canon — the 内经 (Inner Canon),
难经, 伤寒论 (Treatise on Cold Damage), 金匮要略, and 温病 (warm-disease) classics —
assembled from public-domain source works. Useful for LLM training, RAG, and search
over classical Chinese medicine / 中医药 literature.
Summary
115 distinct works, 9,401,166 Chinese… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-tcm-canon.WanJuan-BaiHua
OpenDataLab近期拟发布万卷·百华大规模专业领域数据集,主要面向金融、能源、文化教育、政务、通信、交通运输、医疗健康、汽车、烟草、计算机等领域大模型训练,提供高质量、精细处理、领域分类的多模态专用语料。 在此我们将面向社区征集需求,若有相关行业数据集的需求,请填写以下调查问卷,我们会根据社区反馈来决定不同领域数据集的发布顺序,欢迎大家提供想法!
调查问卷链接:https://www.wjx.cn/vm/mxip3eE.aspx#
数据集样例详见:https://opendatalab.com/OpenDataLab/WanJuan-BaiHua
HealthCoreBench-preview
Kimi-K2.6-Reasoning-3300x-WandB
Kimi-K2.6-Reasoning-3300x-WandB
Kimi-K2.6-Reasoning-3300x-WandB is a W&B-only synthetic reasoning dataset generated with Kimi-K2.6 through Weights & Biases Inference.
This dataset is the pure W&B-generated subset from a larger planned 8,000-example Kimi reasoning distillation run. Generation stopped when the W&B quota limit was reached, and the completed accepted rows were audited, cleaned, and exported as a standalone dataset.
This release contains 3,303 accepted W&B-generated rows… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Reasoning-3300x-WandB.WanJuan-Arabic
💡 Introduction
WanJuan-Arabic(万卷丝路-阿拉伯语) corpus, with a volume exceeding 220GB, comprises 7 major categories and 34 subcategories. It covers a wide range of local-specific content, including history, politics, culture, real estate, shopping, weather, dining, encyclopedias, and professional knowledge. The rich thematic classification not only facilitates researchers in retrieving data according to specific needs but also ensures that the corpus can adapt to diverse research… See the full description on the dataset page: https://huggingface.co/datasets/opendatalab/WanJuan-Arabic.typhoon-s-instruct-post-training
Typhoon-S Instruct Post-Training
Dataset Summary
This dataset is a post-training corpus used in the Typhoon-S recipe for building Sovereign AI: high-performing, region- and domain-specific LLMs that remain localized, controllable, and resource-efficient. It is designed to help transform a sovereignty-adapted base model into a capable assistant while preserving target-language strengths.
The dataset follows a two-part mixture philosophy:
Target-language (Thai) alignment… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-instruct-post-training.LonsRex-Misinfomation-SFTMSMS-LIMO-v2-SFT
A Multi-Source Multi-Solution Long CoT SFT Dataset from LIMO-v2
typhoon-s-sovereign-capability-dataset
Typhoon-S Training Assets
Training and evaluation datasets for Section 3, Thai language models used in the Typhoon-S project.
Datasets
NitiBench (Legal Domain)
nitibench_train_rl.parquet - RL training set (8,211 examples)
nitibench_train_pretrain.parquet - Pretrain set (3,648 examples)
nitibench_train_sft.parquet - SFT set (3,648 examples)
nitibench_test.parquet - Test set (373 examples) (10% of https://huggingface.co/datasets/VISAI-AI/nitibench ccl split)… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/typhoon-s-sovereign-capability-dataset.KhanomTanLLM-pretrained-dataset-thai-subset
KhanomTanLLM pretrained dataset (Thai subset)
This daataset collect all raw text for pretraining LLM. (Thai subset)
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0
pythainlp/thai-tnhc2-books
pythainlp/thai-constitution-corpus
pythainlp/thai-it-books
pythainlp/prd_news_3011202
pythainlp/thailand-policy-statements
pythainlp/thai-cc-license
pythainlp/blognone_news
pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.PyCoder-Type
PyCoder
This repository contains the dataset for the paper Syntax-Aware On-the-Fly Code Completion
The sample code to run the model can be found in directory: "assets/notebooks/inference.ipynb" in our GitHub: https://github.com/awsm-research/pycoder.
PyCoder is an auto code completion model which leverages a Multi-Task Training technique (MTT) to cooperatively
learn the code prediction task and the type prediction task. For the type prediction
task, we propose to leverage the… See the full description on the dataset page: https://huggingface.co/datasets/Wannita/PyCoder-Type.
