datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Wan-Syn_77x448x832_600kX-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.Polymarket_data
Polymarket Data
Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze
A comprehensive dataset of 1.9 billion trading records from Polymarket, processed into multiple analysis-ready formats. Features cleaned data, unified token perspectives, and user-level transformations — ready for market research, behavioral studies, and quantitative analysis.
Zhengjie Wang1,2, Leiyu Chao1,3, Yu Bao1,4, Lian Cheng1,3, Jianhan Liao1,5, Yikang Li1,†
1Shanghai Innovation Institute… See the full description on the dataset page: https://huggingface.co/datasets/SII-WANGZJ/Polymarket_data.CT_DeepLesion-MedSAM2
CT_DeepLesion-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.Wan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset
FastVideo Team
Paper |
Github |
Project Page
Abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/Hahshshsshbs/Wan2.2-Syn-121x704x1280_32k.WANLI
Dataset Card for WANLI
Dataset Summary
WANLI (Worker-AI Collaboration for NLI) is a collection of 108K English sentence pairs for the task of natural language inference (NLI).
Each example is created by first identifying a "pocket" of examples in MultiNLI (Williams et al., 2018) that share a challenging reasoning pattern, then instructing GPT-3 to write a new example with the same pattern.
The set of generated examples are automatically filtered to contain those most… See the full description on the dataset page: https://huggingface.co/datasets/alisawuffles/WANLI.Marathi-WikipediaWan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset
FastVideo Team
Paper |
Github |
Project Page
Abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k.icrm-hitek-fulldbtruecallerdatatemp-dedup-krakenWan-Syn_77x448x832_600kgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.temp-dedupZhihu-KOL
Dataset Card for "Zhihu-KOL"
Zhihu data for training Open Assitant
More Information needed
temp-bert-train-tokenizedRAGTruth-processed
RAGTruth Dataset
Dataset Description
Dataset Summary
The RAGTruth dataset is designed for evaluating hallucinations in text generation models, particularly in retrieval-augmented generation (RAG) contexts. It contains examples of model outputs along with expert annotations indicating whether the outputs contain hallucinations.
Dataset Structure
Each example contains:
A query/question
Context passages
Model output
Hallucination labels (evident… See the full description on the dataset page: https://huggingface.co/datasets/wandb/RAGTruth-processed.WangchanThaiInstruct_Multi-turn_Conversation_Dataset
WangchanThaiInstruct Multi-turn Conversation Dataset
We create a Thai multi-turn conversation dataset from airesearch/WangchanThaiInstruct (Batch 1) by LLM. It was created from synthetic method using open source LLM in Thai language.
Citation
Thammaleelakul, S., & Phatthiyaphaibun, W. (2024). WangchanThaiInstruct Multi-turn Conversation Dataset [Data set]. Zenodo. https://doi.org/10.5281/zenodo.13132633
or BibTeX
@dataset{thammaleelakul_2024_13132633,
author =… See the full description on the dataset page: https://huggingface.co/datasets/ThaiSyntheticQA/WangchanThaiInstruct_Multi-turn_Conversation_Dataset.World2VLM
🌍 World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
📄 Paper •
💻 Code •
🤗 Dataset
✨ Overview
This repository provides a full dataset for the paper:
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial ReasoningWanyue Zhang et al., 2026
🔍 MotivationVision-Language Models (VLMs) excel at static visual understanding but struggle with dynamic spatial reasoning, such as predicting how a scene… See the full description on the dataset page: https://huggingface.co/datasets/WanyueZhang/World2VLM.UAV-FlowSteeringSafety
SteeringSafety: A Systematic Safety Evaluation Framework of Representation Steering in LLMs
🎯 Overview
SteeringSafety is a benchmark suite for evaluating representation steering methods across multiple safety perspectives.
SteeringSafety provides:
📊 A collection of 17 datasets including 7 perspectives for measuring safety behaviors.
🔧 A modular code framework implementing the taxonomy of training-free steering methods with standardized, interchangeable… See the full description on the dataset page: https://huggingface.co/datasets/WangResearchLab/SteeringSafety.wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.wan2.2-Lorastext-2-video-human-preferences-wan2.1
Rapidata Video Generation Alibaba Wan2.1 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.GameplayQA
GameplayQA: A Decision-Dense POV-Synced Multi-Video
Understanding Benchmark of 3D Virtual Agents
Yunzhe Wang
Runhui Xu
Kexin Zheng
Tianyi Zhang
Jayavibhav N. Kogundi
Soham Hans
Volkan Ustun
University of Southern California
ACL 2026
Corresponding Author: yunzhewa@usc.edu
Overview
GameplayQA is the first benchmark for POV-Synced Multi-Video Understanding and… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/GameplayQA.Wan-Syn_77x768x1280_250kUAV-Flow-Simfinqa-data-processed
FinQA Dataset (Processed)
Dataset Description
Dataset Summary
The FinQA dataset is designed for numerical reasoning over financial data, containing questions that require complex reasoning over tables and text from financial reports.
Dataset Statistics
Total examples: 8281
Training set size: 6624 examples
Test set size: 1657 examples
Dataset Structure
Each example contains:
Required columns:
query: The question to be answered (derived… See the full description on the dataset page: https://huggingface.co/datasets/wandb/finqa-data-processed.HealthCareMagic-100k-enWANLORA
