datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
WANLI
Dataset Card for WANLI
Dataset Summary
WANLI (Worker-AI Collaboration for NLI) is a collection of 108K English sentence pairs for the task of natural language inference (NLI).
Each example is created by first identifying a "pocket" of examples in MultiNLI (Williams et al., 2018) that share a challenging reasoning pattern, then instructing GPT-3 to write a new example with the same pattern.
The set of generated examples are automatically filtered to contain those most… See the full description on the dataset page: https://huggingface.co/datasets/alisawuffles/WANLI.temp-deduptemp-decoder-train-tokenizedWangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.wan2.2-LorasHealthCareMagic-100k-enWangchanLION-Web
Citation
@misc{phatthiyaphaibun2025mangosteenopenthaicorpus,
title={Mangosteen: An Open Thai Corpus for Language Model Pretraining},
author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong},
year={2025},
eprint={2507.14664},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2507.14664},
}
We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.C4-Eval
C4-Eval
C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation.
221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures.
1,105 evaluation instances: five task formulations for every base item.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.Wan2GPCrossPoint-Bench
CrossPoint-Bench
CrossPoint-Bench is a comprehensive benchmark for evaluating Vision-Language Models (VLMs) on cross-view point correspondence tasks. It assesses models' abilities to spatial understanding, and correspondence between different viewpoints.
Dataset Structure
CrossPoint-Bench/
├── CrossPoint-Bench.jsonl # Main benchmark data file
└── image/
├── origin_image/ # Original scene images organized by scene ID
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/WangYipu2002/CrossPoint-Bench.TravelUAV_data_jsonHANDALwikipedia-zh-mnbvc
zhwiki-mnbvc
分项目:爬取并处理中文维基百科语料
数据时间:202302-202305 (持续更新)
主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC
该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1
并且使用组员开发的去重工具进行数据格式化。
总行数(样本): 10,754,146
一个示例:
{
"文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt",
"是否待查文件": false,
"是否重复文件": false,
"文件大小": 558,
"simhash": 14363740497821204542,
"最长段落长度": 142,
"段落数": 6,
"去重段落数": 6,
"低质量段落数": 0,
"段落": [
{… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.WangchanThaiMedicalOmniEAR
OmniEAR Expert Trajectory Dataset
Dataset Summary
The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.comfyui-wan22-assetspii-bench-zh
PII Bench ZH
Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations.
This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets.
Disclaimer / 免责声明
This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.HawkEye-IT
Download Video
Please download the original videos from the provided links:
VideoChat: Based on InternVid, we created additional instruction data and used GPT-4 to condense the existing data.
VideoChatGPT: The original caption data was converted into conversation data based on the same VideoIDs.
Kinetics-710 & SthSthV2: Option candidates were generated from UMTtop-20 predictions.
NExTQA: Typos in the original sentences were corrected.
CLEVRER: For single-option multiple-choice QAs… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/HawkEye-IT.KhanomTanLLM-pretrained-dataset
KhanomTanLLM pretrained dataset
This daataset collect all raw text for pretraining LLM.
Codename: numfa v2
Repository: https://github.com/pythainlp/KhanomTanLLM
Tokens
53,376,211,711 Tokens
English: 31,629,984,243 Tokens
Thai: 12,785,565,497 Tokens
Code: 8,913,084,300 Toekns
Parallel data: 190,310,686 Tokens
Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer
All subset
Thai
pythainlp/thai_food_v1.0
pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.med-imagesFRED-SemBench
FRED-SemBench
FRED-SemBench is a 200-question benchmark candidate for evaluating whether LLM
agents retrieve macroeconomic answers with the intended concept, series,
transformation, unit, observation period, and data-vintage semantics.
This dataset accompanies the FinNLP 2026 paper
“FRED-SemBench: Evaluating Semantic Reliability in LLM Access to
Macroeconomic Data”
by Wilson Wang, Chandler Han, and Peter Zhang (Kairos-AI).
Status and scope
50 independently… See the full description on the dataset page: https://huggingface.co/datasets/wangjinh/FRED-SemBench.wandr
WANDR
Overview and provenance
WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, structured,
high-volume web research tasks. This dataset is a task-and-verification corpus,
not a question/answer collection: it contains no solver outputs or reference
answer sets. WANDR evaluation refetches cited pages and judges submitted records
against task-specific, reference-free specifications.
See the paper, the
blog,
and the evaluation repository.… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/wandr.icliniq-10k-enROM
ROM: Real-time Overthinking Mitigation
This dataset contains the Counterfactual Self-Correction (CSC) training data used in the paper ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention. It consists of 1,533 samples (740 efficient + 793 overthinking) derived from MATH500, used to train a lightweight hidden-state detector that identifies when a large reasoning model has reached the first correct solution and should stop generating.
Links
Paper: arXiv… See the full description on the dataset page: https://huggingface.co/datasets/xinyan-wang/ROM.IteraTeR_full_sentPaper: Understanding Iterative Revision from Human-Written Text
Authors: Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, Dongyeop Kang
Github repo: https://github.com/vipulraheja/IteraTeR
korean-vocabulary-5000
Koko Korean 5K — Multilingual Vocabulary Dataset
5,000 carefully curated Korean vocabulary entries with English translations,
romanization, contextual usage notes, and example sentences. Each entry is
also translated into 9 additional languages, giving researchers and
developers a high-quality parallel corpus of 50,000 aligned vocabulary
records anchored to Korean.
The dataset reflects real conversation patterns from K-dramas, K-pop, and
everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/korean-vocabulary-5000.bridgeTLDR: This dataset is a real-world math tutoring dataset from the NAACL 2024 paper ``Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes''.
The dataset targets scenarios where the student makes a math mistake.
c_h is the conversation history
c_r is the original tutor's response
c_r_ is the experienced teacher's response
Optionally, there is other interesting metadata from our Bridge method:
e is the student error type that the experienced… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/bridge.mini-carla-192x320-wan-2p2-vae
mini-carla-192x320-wan-2p2-vae
Wan2.2-VAE-encoded latents of a small CARLA driving dataset (192x320, native resolution,
no resize). Produced for training miniworld, a
minimal flow-matching world-model framework, by caching pixel clips through the frozen
pretrained Wan2.2 video VAE instead of a locally-trained one.
Data size
Source pixel dataset
mini_192x320_low: 320 episodes, 600 frames each (192x320, 20 fps) — ~33 GB
Clips in this cache
1,920 (6… See the full description on the dataset page: https://huggingface.co/datasets/kamwoh/mini-carla-192x320-wan-2p2-vae.IteraTeR_human_sentPaper: Understanding Iterative Revision from Human-Written Text
Authors: Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, Dongyeop Kang
Github repo: https://github.com/vipulraheja/IteraTeR
