CoolFace
30 shown

datasets

Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.

Clear all
01WanyueZhang /MulSeT MulSeT: A Benchmark for Multi-view Spatial Understanding Tasks Paper: Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture Code: https://github.com/WanyueZhang-ai/spatial-understanding A high-level overview of the MulSeT benchmark. The dataset challenges models to integrate information from two distinct viewpoints of a 3D scene to answer spatial reasoning questions. 📝 Dataset Summary MulSeT is a comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/WanyueZhang/MulSeT.image6 likes74k downloads11mo agoHugging Face02FastVideo /Wan-Syn_77x448x832_600ktext100K<n<1M8 likes50k downloads9mo agoHugging Face03wangyz1999 /X-EGO-CS X-Ego-CS Ten players. One match. Ten simultaneous first-person recordings, each paired with a 64 Hz stream of that player's exact keyboard, mouse and view-angle inputs — all on a common, measured clock. Paper · Paper code · Collection pipeline Cross-Ego Demo (Pistol Round) Your browser cannot play this video — download it instead. All ten players' points of view, from the same pistol round, on one clock. Note: this demo concatenates the ten streams… See the full description on the dataset page: https://huggingface.co/datasets/wangyz1999/X-EGO-CS.tabularvideo-classification10K<n<100K2 likes29k downloads4h agoHugging Face04SII-WANGZJ /Polymarket_data Polymarket Data Complete Data Infrastructure for Polymarket — Fetch, Process, Analyze A comprehensive dataset of 1.9 billion trading records from Polymarket, processed into multiple analysis-ready formats. Features cleaned data, unified token perspectives, and user-level transformations — ready for market research, behavioral studies, and quantitative analysis. Zhengjie Wang1,2, Leiyu Chao1,3, Yu Bao1,4, Lian Cheng1,3, Jianhan Liao1,5, Yikang Li1,† 1Shanghai Innovation Institute… See the full description on the dataset page: https://huggingface.co/datasets/SII-WANGZJ/Polymarket_data.tabular1B<n<10B83 likes21k downloads2mo agoHugging Face05Wangtwohappy /EgoLife_IMU0 likes21k downloads1y agoHugging Face06wangxiangyu0814 /TravelUAV4 likes16k downloads2y agoHugging Face07wanglab /CT_DeepLesion-MedSAM2 CT_DeepLesion-MedSAM2 Dataset Authors Jun Ma* 1,2, Zongxin Yang* 3, Sumin Kim2,4,5, Bihui Chen2,4,5, Mohammed Baharoon2,3,5, Adibvafa Fallahpour2,4,5, Reza Asakereh4,7, Hongwei Lyu4, Bo Wang† 1,2,4,5,6 * Equal contribution     † Corresponding author 1AI Collaborative Centre, University Health Network, Toronto, Canada 2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/CT_DeepLesion-MedSAM2.tabular10K<n<100K20 likes14k downloads1y agoHugging Face08Hahshshsshbs /Wan2.2-Syn-121x704x1280_32k FastVideo Synthetic Wan2.2 720P dataset FastVideo Team  Paper | Github | Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/Hahshshsshbs/Wan2.2-Syn-121x704x1280_32k.tabulartext-to-video10K<n<100K1 likes11k downloads3mo agoHugging Face09alisawuffles /WANLI Dataset Card for WANLI Dataset Summary WANLI (Worker-AI Collaboration for NLI) is a collection of 108K English sentence pairs for the task of natural language inference (NLI). Each example is created by first identifying a "pocket" of examples in MultiNLI (Williams et al., 2018) that share a challenging reasoning pattern, then instructing GPT-3 to write a new example with the same pattern. The set of generated examples are automatically filtered to contain those most… See the full description on the dataset page: https://huggingface.co/datasets/alisawuffles/WANLI.texttext-classification100K<n<1M12 likes8.8k downloads4y agoHugging Face10wangyi111 /Copernicus-Pretrain Dataset Card for Copernicus-Pretrain Copernicus-Pretrain is a large-scale EO pretraining dataset with 18.7M aligned images covering all major Sentinel missions (S1,2,3,5P). Officially named Copernicus-Pretrain, also referred to as SSL4EO-S ("S" means Sentinel), as an extension of SSL4EO-S12 to the whole Sentinel series. Dataset Details Copernicus-Pretrain contains 18.7M aligned imagery from all major Sentinel missions in operation (Sentinel-1 SAR, Sentinel-2… See the full description on the dataset page: https://huggingface.co/datasets/wangyi111/Copernicus-Pretrain.geospatialimage-classification10M<n<100M7 likes8.6k downloads1y agoHugging Face11Aarsh-Wankar /Marathi-Wikipediatext0 likes8.4k downloads2y agoHugging Face12mess2735 /wan2.2_lora0 likes7.5k downloads6mo agoHugging Face13Wangtwohappy /EgoLife_EyeTracking_EyeGazevideo0 likes7.5k downloads1y agoHugging Face14wangyueqian /ProactiveVideoQA ProactiveVideoQA: A Comprehensive Benchmark Evaluating Proactive Interactions in Video Large Language Models 📄 arXiv Paper | 🖥️ Github Code | 📦 Data Introduction ProactiveVideoQA is the first comprehensive benchmark designed to evaluate a system's ability to engage in proactive interaction in multimodal dialogue settings. Unlike traditional turn-by-turn dialogue systems, in proactive intraction model need to determine when to repsond during… See the full description on the dataset page: https://huggingface.co/datasets/wangyueqian/ProactiveVideoQA.video2 likes5.9k downloads1y agoHugging Face15wangxiangyu0814 /TravelUAV_env0 likes5.2k downloads2y agoHugging Face16FastVideo /Wan2.2-Syn-121x704x1280_32k FastVideo Synthetic Wan2.2 720P dataset FastVideo Team  Paper | Github | Project Page Abstract Scaling video diffusion transformers (DiTs) is limited by their quadratic 3D attention, even though most of the attention mass concentrates on a small subset of positions. We turn this observation into VSA, a trainable, hardware-efficient sparse attention that replaces full attention at \emph{both} training and inference. In VSA, a… See the full description on the dataset page: https://huggingface.co/datasets/FastVideo/Wan2.2-Syn-121x704x1280_32k.tabulartext-to-video10K<n<100K9 likes5k downloads11mo agoHugging Face17worstcoder /Wan_datasets rCM: Score-Regularized Continuous-Time Consistency Model Causal-rCM: A Unified Teacher-Forcing and Self-Forcing Open Recipe for Autoregressive Diffusion Distillation in Streaming Video Generation and Interactive World Models Paper Paper2 | Website | Code This repo holds Wan-synthesized datasets used for rCM training. Citation @article{zheng2025rcm, title={Large Scale Diffusion Distillation via Score-Regularized Continuous-Time Consistency}… See the full description on the dataset page: https://huggingface.co/datasets/worstcoder/Wan_datasets.n>1T9 likes4.4k downloads3mo agoHugging Face18wanglab /LLD-MMRI-MedSAM2 LLD-MMRI-MedSAM2 Dataset Authors Jun Ma* 1,2, Zongxin Yang* 3, Sumin Kim2,4,5, Bihui Chen2,4,5, Mohammed Baharoon2,3,5, Adibvafa Fallahpour2,4,5, Reza Asakereh4,7, Hongwei Lyu4, Bo Wang† 1,2,4,5,6 * Equal contribution     † Corresponding author 1AI Collaborative Centre, University Health Network, Toronto, Canada 2Vector… See the full description on the dataset page: https://huggingface.co/datasets/wanglab/LLD-MMRI-MedSAM2.image-segmentation1K<n<10K16 likes3.8k downloads1y agoHugging Face19wannabeyour /icrm-hitek-fulldbtext1B<n<10B0 likes3.2k downloads29d agoHugging Face20wannabeyour /truecallerdatatext100M<n<1B0 likes2.9k downloads29d agoHugging Face21upup-ashton-wang /temp-dedup-krakentext1B<n<10B0 likes2.7k downloads4mo agoHugging Face22NZC415 /wan22-processed-clips0 likes2.5k downloads14d agoHugging Face23Wanfq /gpqa Dataset Card for GPQA GPQA is a multiple-choice, Q&A dataset of very hard questions written and validated by experts in biology, physics, and chemistry. When attempting questions out of their own domain (e.g., a physicist answers a chemistry question), these experts get only 34% accuracy, despite spending >30m with full access to Google. We request that you do not reveal examples from this dataset in plain text or images online, to reduce the risk of leakage into foundation model… See the full description on the dataset page: https://huggingface.co/datasets/Wanfq/gpqa.tabularquestion-answering1K<n<10K0 likes2.5k downloads2y agoHugging Face24wanlilll /WeaveBench WeaveBench A long-horizon, real-world benchmark for computer-use agents with hybrid GUI + CLI + code interfaces. 🎊 Accepted to EMNLP 2026 Main Conference — see you in Budapest! 📄 Paper: arXiv:2606.09426 💻 Code: github.com/weavebench/WeaveBench 🌐 Website: weavebench.github.io 114 long-horizon, real-world tasks across 8 work domains, where every task requires the agent to interleave GUI clicks with shell/code in one trajectory. Scored by a trajectory-aware Agent-as-Judge… See the full description on the dataset page: https://huggingface.co/datasets/wanlilll/WeaveBench.othern<1K8 likes2.5k downloads1mo agoHugging Face25wangyanhui666 /imagenet_vae_mds_fp320 likes2.5k downloads2y agoHugging Face26guyuchao /fusionX_480p_wan21_latents0 likes2.4k downloads6mo agoHugging Face27wangrui6 /Zhihu-KOL Dataset Card for "Zhihu-KOL" Zhihu data for training Open Assitant More Information needed textquestion-answering1M<n<10M262 likes2.4k downloads3y agoHugging Face28rand0nmr /Wan-Syn_77x448x832_600ktext100K<n<1M0 likes2.4k downloads9mo agoHugging Face29upup-ashton-wang /temp-deduptext1B<n<10B0 likes2.3k downloads5mo agoHugging Face30LeCAR-Lab /Wanda WANDA: Worlds in One Demo A Synthetic Data Engine for Learning Open-World Mobile Manipulation 🌐 Project page: https://wanda.lecar-lab.org/ · 📄 Paper (PDF) · 💻 Code (coming soon) · 🕹️ Interactive 4D viewer Lingxiao Guo*, Huanyu Li*, Guanya Shi — Carnegie Mellon University *Equal contribution; order decided by a coin flip. WANDA is a synthetic data engine that turns one human demonstration into diverse training data for open-world mobile manipulation. From a… See the full description on the dataset page: https://huggingface.co/datasets/LeCAR-Lab/Wanda.videorobotics10K<n<100K3 likes2.1k downloads2mo agoHugging Face

Listings come live from the Hugging Face Hub API. CoolFace does not host these files.