CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01alisawuffles /WANLI Dataset Card for WANLI Dataset Summary WANLI (Worker-AI Collaboration for NLI) is a collection of 108K English sentence pairs for the task of natural language inference (NLI). Each example is created by first identifying a "pocket" of examples in MultiNLI (Williams et al., 2018) that share a challenging reasoning pattern, then instructing GPT-3 to write a new example with the same pattern. The set of generated examples are automatically filtered to contain those most… See the full description on the dataset page: https://huggingface.co/datasets/alisawuffles/WANLI.texttext-classification100K<n<1M12 likes9k downloads4y agoHugging Face02upup-ashton-wang /temp-deduptext1B<n<10B0 likes2.3k downloads5mo agoHugging Face03upup-ashton-wang /temp-decoder-train-tokenized1B<n<10B0 likes1.9k downloads5mo agoHugging Face04ThaiSyntheticQA /WangchanThaiInstruct_Multi-turn_Conversation_Dataset WangchanThaiInstruct Multi-turn Conversation Dataset We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language. Citation Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633 or BibTeX @dataset{thammaleelakul_2024_13132633, author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.texttext-generation1K<n<10K1 likes1.9k downloads2y agoHugging Face05simbahuang /wan22-animate-3k-opensource-data Wan2.2 Animate Open Dataset Pack This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment. The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards. Restore: cat datasets.tar.part-* | tar -xf - sha256sum -c SHA256SUMS After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.image10K<n<100K0 likes1.3k downloads3mo agoHugging Face06mikeo609 /wan2.2-Lorastabularn<1K12 likes1.2k downloads1y agoHugging Face07wangrongsheng /HealthCareMagic-100k-entext100K<n<1M22 likes789 downloads3y agoHugging Face08aisingapore /WangchanLION-Web Citation @misc{phatthiyaphaibun2025mangosteenopenthaicorpus, title={Mangosteen: An Open Thai Corpus for Language Model Pretraining}, author={Wannaphong Phatthiyaphaibun and Can Udomcharoenchaikit and Pakpoom Singkorapoom and Kunat Pipatanakul and Ekapol Chuangsuwanich and Peerat Limkonchotiwat and Sarana Nutanong}, year={2025}, eprint={2507.14664}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.14664}, } We… See the full description on the dataset page: https://huggingface.co/datasets/aisingapore/WangchanLION-Web.texttext-generation10M<n<100M3 likes736 downloads1y agoHugging Face09sci-m-wang /C4-Eval C4-Eval C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation. 221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures. 1,105 evaluation instances: five task formulations for every base item. Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.imageimage-text-to-text1K<n<10K0 likes693 downloads1mo agoHugging Face10luxxandfuxx /Wan2GPtabularn<1K1 likes656 downloads13d agoHugging Face11WangYipu2002 /CrossPoint-Bench CrossPoint-Bench CrossPoint-Bench is a comprehensive benchmark for evaluating Vision-Language Models (VLMs) on cross-view point correspondence tasks. It assesses models' abilities to spatial understanding, and correspondence between different viewpoints. Dataset Structure CrossPoint-Bench/ ├── CrossPoint-Bench.jsonl # Main benchmark data file └── image/ ├── origin_image/ # Original scene images organized by scene ID │ ├──… See the full description on the dataset page: https://huggingface.co/datasets/WangYipu2002/CrossPoint-Bench.imagevisual-question-answering1K<n<10K3 likes645 downloads9mo agoHugging Face12wangxiangyu0814 /TravelUAV_data_jsontext10K<n<100K0 likes546 downloads2y agoHugging Face13wangzeze /HANDALtext100K<n<1M0 likes492 downloads11mo agoHugging Face14wanng /wikipedia-zh-mnbvc zhwiki-mnbvc 分项目:爬取并处理中文维基百科语料 数据时间:202302-202305 (持续更新) 主项目:MNBVC(Massive Never-ending BT Vast Chinese corpus)超大规模中文语料集 https://github.com/esbatmop/MNBVC 该项目清洗流程主要参考:https://kexue.fm/archives/4176/comment-page-1 并且使用组员开发的去重工具进行数据格式化。 总行数(样本): 10,754,146 一个示例: { "文件名": "cleaned/zhwiki-20230420/folder_0/723712.txt", "是否待查文件": false, "是否重复文件": false, "文件大小": 558, "simhash": 14363740497821204542, "最长段落长度": 142, "段落数": 6, "去重段落数": 6, "低质量段落数": 0, "段落": [ {… See the full description on the dataset page: https://huggingface.co/datasets/wanng/wikipedia-zh-mnbvc.tabulartext-generation1M<n<10M6 likes381 downloads3y agoHugging Face15airesearch /WangchanThaiMedicaltext1M<n<10M0 likes339 downloads2y agoHugging Face16wangzx1210 /OmniEAR OmniEAR Expert Trajectory Dataset Dataset Summary The OmniEAR Expert Trajectory SFT Dataset is a comprehensive collection of high-quality expert demonstration trajectories specifically designed for supervised fine-tuning (SFT) of embodied reasoning models. This dataset contains 1,982 instruction-following examples across single-agent and multi-agent scenarios, focusing on physical interactions, tool usage, and collaborative reasoning in embodied environments.… See the full description on the dataset page: https://huggingface.co/datasets/wangzx1210/OmniEAR.texttext-generation10K<n<100K10 likes290 downloads1y agoHugging Face17KishoOoOo /comfyui-wan22-assetsaudion<1K1 likes276 downloads23d agoHugging Face18wan9yu /pii-bench-zh PII Bench ZH Chinese PII (Personally Identifiable Information) detection benchmark dataset. Two subsets covering formal and informal Chinese text, with character-level span annotations. This is the first open Chinese PII benchmark that covers locale-specific formats (phone, national ID, bank card, license plate, address) with precise offsets. Disclaimer / 免责声明 This dataset is 100% synthetic and intended solely for research and evaluation purposes. It does not contain any real… See the full description on the dataset page: https://huggingface.co/datasets/wan9yu/pii-bench-zh.texttoken-classification1K<n<10K0 likes255 downloads6mo agoHugging Face19wangyueqian /HawkEye-IT Download Video Please download the original videos from the provided links: VideoChat: Based on InternVid, we created additional instruction data and used GPT-4 to condense the existing data. VideoChatGPT: The original caption data was converted into conversation data based on the same VideoIDs. Kinetics-710 & SthSthV2: Option candidates were generated from UMTtop-20 predictions. NExTQA: Typos in the original sentences were corrected. CLEVRER: For single-option multiple-choice QAs… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/HawkEye-IT.textvisual-question-answering1M<n<10M0 likes250 downloads3y agoHugging Face20wannaphong /KhanomTanLLM-pretrained-dataset KhanomTanLLM pretrained dataset This daataset collect all raw text for pretraining LLM. Codename: numfa v2 Repository: https://github.com/pythainlp/KhanomTanLLM Tokens 53,376,211,711 Tokens English: 31,629,984,243 Tokens Thai: 12,785,565,497 Tokens Code: 8,913,084,300 Toekns Parallel data: 190,310,686 Tokens Based on Typhoon-7B (https://huggingface.co/scb10x/typhoon-7b) tokenizer All subset Thai pythainlp/thai_food_v1.0 pythainlp/thailaw-v1.0… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/KhanomTanLLM-pretrained-dataset.texttext-generation10M<n<100M1 likes235 downloads2y agoHugging Face21wanqi16 /med-imagestextn<1K0 likes228 downloads8mo agoHugging Face22wangjinh /FRED-SemBench FRED-SemBench FRED-SemBench is a 200-question benchmark candidate for evaluating whether LLM agents retrieve macroeconomic answers with the intended concept, series, transformation, unit, observation period, and data-vintage semantics. This dataset accompanies the FinNLP 2026 paper “FRED-SemBench: Evaluating Semantic Reliability in LLM Access to Macroeconomic Data” by Wilson Wang, Chandler Han, and Peter Zhang (Kairos-AI). Status and scope 50 independently… See the full description on the dataset page: https://huggingface.co/datasets/wangjinh/FRED-SemBench.tabularquestion-answeringn<1K1 likes188 downloads19d agoHugging Face23perplexity-ai /wandr WANDR Overview and provenance WANDR (Wide ANd Deep Research) is a benchmark of 500 realistic, structured, high-volume web research tasks. This dataset is a task-and-verification corpus, not a question/answer collection: it contains no solver outputs or reference answer sets. WANDR evaluation refetches cited pages and judges submitted records against task-specific, reference-free specifications. See the paper, the blog, and the evaluation repository.… See the full description on the dataset page: https://huggingface.co/datasets/perplexity-ai/wandr.textquestion-answeringn<1K2 likes167 downloads16d agoHugging Face24wangrongsheng /icliniq-10k-entext1K<n<10K4 likes161 downloads3y agoHugging Face25xinyan-wang /ROM ROM: Real-time Overthinking Mitigation This dataset contains the Counterfactual Self-Correction (CSC) training data used in the paper ROM: Real-time Overthinking Mitigation via Streaming Detection and Intervention. It consists of 1,533 samples (740 efficient + 793 overthinking) derived from MATH500, used to train a lightweight hidden-state detector that identifies when a large reasoning model has reached the first correct solution and should stop generating. Links Paper: arXiv… See the full description on the dataset page: https://huggingface.co/datasets/xinyan-wang/ROM.texttext-classification1K<n<10K1 likes146 downloads1mo agoHugging Face26wanyu /IteraTeR_full_sentPaper: Understanding Iterative Revision from Human-Written Text Authors: Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, Dongyeop Kang Github repo: https://github.com/vipulraheja/IteraTeR text100K<n<1M1 likes141 downloads4y agoHugging Face27wannaphong /korean-vocabulary-5000 Koko Korean 5K — Multilingual Vocabulary Dataset 5,000 carefully curated Korean vocabulary entries with English translations, romanization, contextual usage notes, and example sentences. Each entry is also translated into 9 additional languages, giving researchers and developers a high-quality parallel corpus of 50,000 aligned vocabulary records anchored to Korean. The dataset reflects real conversation patterns from K-dramas, K-pop, and everyday Korean — not textbook-only material… See the full description on the dataset page: https://huggingface.co/datasets/wannaphong/korean-vocabulary-5000.texttranslation10K<n<100K0 likes139 downloads5mo agoHugging Face28rose-e-wang /bridgeTLDR: This dataset is a real-world math tutoring dataset from the NAACL 2024 paper ``Bridging the Novice-Expert Gap via Models of Decision-Making: A Case Study on Remediating Math Mistakes''. The dataset targets scenarios where the student makes a math mistake. c_h is the conversation history c_r is the original tutor's response c_r_ is the experienced teacher's response Optionally, there is other interesting metadata from our Bridge method: e is the student error type that the experienced… See the full description on the dataset page: https://huggingface.co/datasets/rose-e-wang/bridge.texttext-generationn<1K5 likes128 downloads2y agoHugging Face29kamwoh /mini-carla-192x320-wan-2p2-vae mini-carla-192x320-wan-2p2-vae Wan2.2-VAE-encoded latents of a small CARLA driving dataset (192x320, native resolution, no resize). Produced for training miniworld, a minimal flow-matching world-model framework, by caching pixel clips through the frozen pretrained Wan2.2 video VAE instead of a locally-trained one. Data size Source pixel dataset mini_192x320_low: 320 episodes, 600 frames each (192x320, 20 fps) — ~33 GB Clips in this cache 1,920 (6… See the full description on the dataset page: https://huggingface.co/datasets/kamwoh/mini-carla-192x320-wan-2p2-vae.tabularothern<1K0 likes128 downloads1mo agoHugging Face30wanyu /IteraTeR_human_sentPaper: Understanding Iterative Revision from Human-Written Text Authors: Wanyu Du, Vipul Raheja, Dhruv Kumar, Zae Myung Kim, Melissa Lopez, Dongyeop Kang Github repo: https://github.com/vipulraheja/IteraTeR text1K<n<10K0 likes122 downloads4y agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.