datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MulSeT
MulSeT: A Benchmark for Multi-view Spatial Understanding Tasks
Paper: Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
Code: https://github.com/WanyueZhang-ai/spatial-understanding
A high-level overview of the MulSeT benchmark. The dataset challenges models to integrate information from two distinct viewpoints of a 3D scene to answer spatial reasoning questions.
📝 Dataset Summary
MulSeT is a comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/WanyueZhang/MulSeT.Wan-Syn_77x448x832_600kX-EGO-CS
X-Ego-CS
Ten players. One match. Ten simultaneous first-person recordings, each paired
with a 64 Hz stream of that player's exact keyboard, mouse and view-angle
inputs — all on a common, measured clock.
Paper · Paper code · Collection pipeline
Cross-Ego Demo (Pistol Round)
Your browser cannot play this video —
download it instead.
All ten players' points of view, from the same pistol round, on one clock.
Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.Polymarket_data
Polymarket Data
Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze
A comprehensive dataset of 1.9 billion trading records from Polymarket, processed into multiple analysis-ready formats. Features cleaned data, unified token perspectives, and user-level transformations — ready for market research, behavioral studies, and quantitative analysis.
Zhengjie Wang1,2, Leiyu Chao1,3, Yu Bao1,4, Lian Cheng1,3, Jianhan Liao1,5, Yikang Li1,†
1Shanghai Innovation Institute… See the full description on the dataset page: https://huggingface.co/datasets/SII-WANGZJ/Polymarket_data.EgoLife_IMUTravelUAVCT_DeepLesion-MedSAM2
CT_DeepLesion-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.Wan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset
FastVideo Team
Paper |
Github |
Project Page
Abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/Hahshshsshbs/Wan2.2-Syn-121x704x1280_32k.WANLI
Dataset Card for WANLI
Dataset Summary
WANLI (Worker-AI Collaboration for NLI) is a collection of 108K English sentence pairs for the task of natural language inference (NLI).
Each example is created by first identifying a "pocket" of examples in MultiNLI (Williams et al., 2018) that share a challenging reasoning pattern, then instructing GPT-3 to write a new example with the same pattern.
The set of generated examples are automatically filtered to contain those most… See the full description on the dataset page: https://huggingface.co/datasets/alisawuffles/WANLI.Copernicus-Pretrain
Dataset Card for Copernicus-Pretrain
Copernicus-Pretrain is a large-scale EO pretraining dataset with 18.7M aligned images covering all major Sentinel missions (S1,2,3,5P).
Officially named Copernicus-Pretrain, also referred to as SSL4EO-S ("S" means Sentinel), as an extension of SSL4EO-S12 to the whole Sentinel series.
Dataset Details
Copernicus-Pretrain contains 18.7M aligned imagery from all major Sentinel missions in operation (Sentinel-1 SAR, Sentinel-2… See the full description on the dataset page: https://huggingface.co/datasets/wangyi111/Copernicus-Pretrain.Marathi-Wikipediawan2.2_loraEgoLife_EyeTracking_EyeGazeProactiveVideoQA
ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models
📄 arXiv Paper |
🖥️ Github Code |
📦 Data
Introduction
ProactiveVideoQA is the first comprehensive benchmark designed to evaluate a system's ability to engage in proactive interaction in multimodal dialogue settings.
Unlike traditional turn-by-turn dialogue systems, in proactive intraction model need to determine when to repsond during… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/ProactiveVideoQA.TravelUAV_envWan2.2-Syn-121x704x1280_32k
FastVideo Synthetic Wan2.2 720P dataset
FastVideo Team
Paper |
Github |
Project Page
Abstract
Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k.Wan_datasets
rCM: Score-Regularized Continuous-Time Consistency Model
Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models
Paper Paper2 | Website | Code
This repo holds Wan-synthesized datasets used for rCM training.
Citation
@article{zheng2025rcm,
title={Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency}… See the full description on the dataset page: https://huggingface.co/datasets/worstcoder/Wan_datasets.LLD-MMRI-MedSAM2
LLD-MMRI-MedSAM2 Dataset
Authors
Jun Ma* 1,2,
Zongxin Yang* 3,
Sumin Kim2,4,5,
Bihui Chen2,4,5,
Mohammed Baharoon2,3,5,
Adibvafa Fallahpour2,4,5,
Reza Asakereh4,7,
Hongwei Lyu4,
Bo Wang† 1,2,4,5,6
* Equal contribution † Corresponding author
1AI Collaborative Centre, University Health Network, Toronto, Canada
2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/LLD-MMRI-MedSAM2.icrm-hitek-fulldbtruecallerdatatemp-dedup-krakenwan22-processed-clipsgpqa
Dataset Card for GPQA
GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google.
We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.WeaveBench
WeaveBench
A long-horizon, real-world benchmark for computer-use agents with hybrid GUI + CLI + code interfaces.
🎊 Accepted to EMNLP 2026 Main Conference — see you in Budapest!
📄 Paper: arXiv:2606.09426
💻 Code: github.com/weavebench/WeaveBench
🌐 Website: weavebench.github.io
114 long-horizon, real-world tasks across 8 work domains, where every task requires the agent to interleave GUI clicks with shell/code in one trajectory. Scored by a trajectory-aware Agent-as-Judge… See the full description on the dataset page: https://huggingface.co/datasets/wanlilll/WeaveBench.imagenet_vae_mds_fp32fusionX_480p_wan21_latentsZhihu-KOL
Dataset Card for "Zhihu-KOL"
Zhihu data for training Open Assitant
More Information needed
Wan-Syn_77x448x832_600ktemp-dedupWanda
WANDA: Worlds in One Demo
A Synthetic Data Engine for Learning Open-World Mobile Manipulation
🌐 Project page: https://wanda.lecar-lab.org/ · 📄 Paper (PDF) · 💻 Code (coming soon) · 🕹️ Interactive 4D viewer
Lingxiao Guo*, Huanyu Li*, Guanya Shi — Carnegie Mellon University
*Equal contribution; order decided by a coin flip.
WANDA is a synthetic data engine that turns one human demonstration into diverse training data for
open-world mobile manipulation. From a… See the full description on the dataset page: https://huggingface.co/datasets/LeCAR-Lab/Wanda.
