datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.WangchanLION-Web
Citation
@misc{phatthiyaphaibun2025mangosteenopenthaicorpus,
title={Mangosteen: An Open Thai Corpus for Language Model Pretraining},
author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong},
year={2025},
eprint={2507.14664},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2507.14664},
}
We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.wikipedia-zh-mnbvc
zhwiki-mnbvc
分项目:爬取并处理中文维基百科语料
数据时间:202302-202305 (持续更新)
主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC
该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1
并且使用组员开发的去重工具进行数据格式化。
总行数(样本): 10,754,146
一个示例:
{
"文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt",
"是否待查文件": false,
"是否重复文件": false,
"文件大小": 558,
"simhash": 14363740497821204542,
"最长段落长度": 142,
"段落数": 6,
"去重段落数": 6,
"低质量段落数": 0,
"段落": [
{… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.OmniEAR
OmniEAR Expert Trajectory Dataset
Dataset Summary
The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.KhanomTanLLM-pretrained-dataset
KhanomTanLLM pretrained dataset
This daataset collect all raw text for pretraining LLM.
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Tokens
53,376,211,711 Tokens
English: 31,629,984,243 Tokens
Thai: 12,785,565,497 Tokens
Code: 8,913,084,300 Toekns
Parallel data: 190,310,686 Tokens
Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer
All subset
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.bridgeTLDR: This dataset is a real-world math tutoring dataset from the NAACL 2024 paper ``Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes''.
The dataset targets scenarios where the student makes a math mistake.
c_h is the conversation history
c_r is the original tutor's response
c_r_ is the experienced teacher's response
Optionally, there is other interesting metadata from our Bridge method:
e is the student error type that the experienced… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/bridge.PolyReal
PolyReal: A Benchmark for Real-World Polymer Science Workflows
Evaluating Multimodal Large Language Models Across the Full Lifecycle of Polymer Science
Contents
data/train.jsonl: Hugging Face viewer-friendly training split.
PolyReal.json: main dataset file.
ref/: referenced images and CSV files used by Path entries in the dataset.
🔥 Overview
PolyReal is a multimodal benchmark for real-world polymer science workflows.It is… See the full description on the dataset page: https://huggingface.co/datasets/wanhaoliu/PolyReal.Kimi-K2.6-Reasoning-3300x-WandB
Kimi-K2.6-Reasoning-3300x-WandB
Kimi-K2.6-Reasoning-3300x-WandB is a W&B-only synthetic reasoning dataset generated with Kimi-K2.6 through Weights & Biases Inference.
This dataset is the pure W&B-generated subset from a larger planned 8,000-example Kimi reasoning distillation run. Generation stopped when the W&B quota limit was reached, and the completed accepted rows were audited, cleaned, and exported as a standalone dataset.
This release contains 3,303 accepted W&B-generated rows… See the full description on the dataset page: https://huggingface.co/datasets/trjxter/Kimi-K2.6-Reasoning-3300x-WandB.MSMS-LIMO-v2-SFT
A Multi-Source Multi-Solution Long CoT SFT Dataset from LIMO-v2
KhanomTanLLM-pretrained-dataset-thai-subset
KhanomTanLLM pretrained dataset (Thai subset)
This daataset collect all raw text for pretraining LLM. (Thai subset)
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0
pythainlp/thai-tnhc2-books
pythainlp/thai-constitution-corpus
pythainlp/thai-it-books
pythainlp/prd_news_3011202
pythainlp/thailand-policy-statements
pythainlp/thai-cc-license
pythainlp/blognone_news
pythainlp/goethe-website… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset-thai-subset.LonsRex-Misinfomation-SFTHundredCV-Chat
百人对话数据集
HundredCV-Chat: A Dataset of Daily Chatting Developed on HundredCVs
简介
本项目提出一个全新的中文多轮对话数据集(HundredCV-Chat),该数据集由 100 位青年的简历数据集 HundredCVs 开发而来,共包含 24,750 组日常闲聊对话数据。
数据集具有如下特点:
自动化标注:HundredCV-Chat 中的对话均由 Deepseek-V3 大模型生成,不涉及任何人工标注,因此同时保证了大规模数据量和低成本优势。
多样性话题:HundredCV-Chat 中的对话话题涵盖了校园生活、工作经验、兴趣爱好、生活琐事等多个方面,与真实生活联系紧密,尤其适用于开发年轻化应用。
高质量对话:利用 Deepseek 强大的生成能力和全面的知识,HundredCV-Chat 的对话内容在流畅度、拟人性、多样性方面均显著优于现有的开源对话数据集。
数据样例… See the full description on the dataset page: https://huggingface.co/datasets/wangmingyuan/HundredCV-Chat.MSMS-AceReason-20K-SFT
A Multi-Source Multi-Solution Long CoT SFT Dataset from 20K AceReason Questions
Anna-CPsyCounD
Dataset Card for AnnaAgent Virtual Seeker Dataset
Dataset Overview
Repository: AnnaAgent GitHubPurpose: Supports the AI psychotherapy agent framework AnnaAgent by providing multidimensional configuration for virtual seekersLanguage: ChineseData Source: Synthetic data constructed using GPT-4o based on CPsyCounD datasetKey Features:
Contains long-term memory (historical counseling records) and short-term memory (current session state)
Supports dynamic emotion evolution… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/Anna-CPsyCounD.gsm8k_distilled
Additional Information
This dataset contains mathematical problem-solving traces generated using the CAMEL framework. Each entry includes:
A mathematical problem statement
A detailed step-by-step solution
tinker-rl-bench-wandb
TinkerRL-Bench W&B Run Archive
Full export of every Weights & Biases run under the arvindcr4-pes-university
entity, covering the experiments reported in our NeurIPS submission
"A Unified Benchmark for RL Post-Training of Language Models"
(repo).
Contents
File
Rows
Description
runs.jsonl
334
One record per run: project, run_id, run_name, state, config, summary, tags, url, runtime
history.jsonl
9,255
Per-step metric history (step, reward, loss, accuracy, etc.)… See the full description on the dataset page: https://huggingface.co/datasets/arvindcr4/tinker-rl-bench-wandb.claude-opus-4.6-4.7-reasoning-8.7k
Background
Ended up with some tokens to burn on a Claude Max plan. Assembly began during 4.6 and moved to 4.7. Model is tagged. The development evolved as it went along. The dataset has not been manually reviewed. It's entirely Claude developed.
Clarification on Reasoning
The reasoning is not Claude's actual chain-of-thought (cot) and is not summarized cot. It's a fully synthetic cot created as part of the Assistant response to mimic the type of "thinking"… See the full description on the dataset page: https://huggingface.co/datasets/wannipaa007/claude-opus-4.6-4.7-reasoning-8.7k.classical-chinese-punctuation
Classical Chinese Punctuation Restoration · 文言文断句·标点 💰 (Commercial Dataset)
This is a commercial dataset. A free 200-record sample is provided below
(sample.jsonl); the full 5.3M-pair corpus is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — built from public-domain classical works.
The task
Restore punctuation and sentence segmentation (句读) to unpunctuated Classical
Chinese — a… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-punctuation.classical-chinese-variant-collation
Classical Chinese Variant Collation · 校勘 💰 (Commercial Dataset)
This is a commercial dataset. A free 50-work preview sample is provided
below (sample.jsonl, texts truncated); the full set with complete aligned
texts is available upon request.
📧 To license / purchase or request a quote, email wangeksy@gmail.com.
✅ Cleared for commercial use — both members are public-domain classical works.
What this is
A textual-criticism dataset: classical works that survive… See the full description on the dataset page: https://huggingface.co/datasets/wangekxy/classical-chinese-variant-collation.wanli-dibao-corpus
Wanli Dibao Corpus / 萬曆邸鈔校訂語料庫
Dataset Description
Summary
English:
The Wanli Dibao Corpus is a structured, proofread digital corpus of the Wanli Dichao (萬曆邸鈔), a collection of manuscript copies of official gazettes (dibao 邸報) from the Wanli reign (1573–1620) of the Ming dynasty. The dibao system was the primary channel of official communication in imperial China, transmitting memorials, edicts, personnel appointments, and policy decisions from the capital to… See the full description on the dataset page: https://huggingface.co/datasets/dibao-research/wanli-dibao-corpus.WanJuanSiLu-sft
WanJuanSiLu sft subset
This dataset is built from opendatalab/WanJuanSiLu-Multimodal-5Languages in sft subset.
180,000 SFT data
Languages: Arabic, Russian, Korean, Vietnamese, and Thai
🤖Featured instructions for fine-tuning SFT data:
Cultural adversarial samples: Contains culturally relevant question-answer pairs designed by local residents to detect cultural bias in models
Hybrid quality inspection process: Rules + model scoring to filter translation data and… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/WanJuanSiLu-sft.wangweiFiqih_Wanita_Dataset
Dataset Fiqih Wanita 10K
Dataset fine-tuning chatbot ustadzah fiqih wanita berbahasa Indonesia.Berisi 10.000 pasangan tanya-jawab berdasarkan kitab Uyunul Masa-il Linnisa' madzhab Syafi'i.
Topik
Haidl, Nifas, Istihadloh, Darah Fasad, Thoharoh, Wudhu, Mandi Wajib,
Larangan saat haidl, Qodlo sholat & puasa, Haji & Umroh, Ramadhan,
Keputihan, KB, Kehamilan, Mitos, Relasi suami-istri, Dalil Al-Quran & Hadits.
Format
{
"messages": [
{"role": "system"… See the full description on the dataset page: https://huggingface.co/datasets/subhanghuf/Fiqih_Wanita_Dataset.dolma_web1_raw
Raw data for Thai Dolma
Commoncrawl
Fineweb2
