datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
MulSeT
MulSeT: A Benchmark for Multi-view Spatial Understanding Tasks
Paper: Why Do MLLMs Struggle with Spatial Understanding? A Systematic Analysis from Data to Architecture
Code: https://github.com/WanyueZhang-ai/spatial-understanding
A high-level overview of the MulSeT benchmark. The dataset challenges models to integrate information from two distinct viewpoints of a 3D scene to answer spatial reasoning questions.
📝 Dataset Summary
MulSeT is a comprehensive benchmark… See the full description on the dataset page: https://huggingface.co/datasets/WanyueZhang/MulSeT.VLBreakBenchUAV-FlowWorld2VLM
🌍 World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
📄 Paper •
💻 Code •
🤗 Dataset
✨ Overview
This repository provides a full dataset for the paper:
World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial ReasoningWanyue Zhang et al., 2026
🔍 MotivationVision-Language Models (VLMs) excel at static visual understanding but struggle with dynamic spatial reasoning, such as predicting how a scene… See the full description on the dataset page: https://huggingface.co/datasets/WanyueZhang/World2VLM.AgenticOCR-SFT
AgenticOCR SFT Training Data
Supervised fine-tuning data for the AgenticOCR project.
The dataset contains 7,631 training records in sft_combined_0422.json. Image paths in each record are relative to the repository root and point into sft_images/.
fitvto-100k
FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On
The official preview dataset from the paper "FIT: A Large-Scale Dataset for Fit-Aware Virtual Try-On".
This dataset supports garment-centric virtual try-on and try-off research, containing 100,000 training and 5,000 evaluation triplets. Each sample pairs a person image with a layflat garment image and body/garment measurements.
Dataset Structure
Each split (train / eval) contains four aligned modalities — all… See the full description on the dataset page: https://huggingface.co/datasets/Yuanhao-Harry-Wang/fitvto-100k.wan22-animate-3k-opensource-data
Wan2.2 Animate Open Dataset Pack
This dataset repo stores the complete datasets/ directory used for the Wan2.2 TI2V 5B + One-to-All animate experiment.
The original tree contains more than 10,000 files in one directory, which Hugging Face git repositories reject as raw files. Therefore the dataset is stored as split tar shards.
Restore:
cat datasets.tar.part-* | tar -xf -
sha256sum -c SHA256SUMS
After extraction, the restored tree contains:… See the full description on the dataset page: https://huggingface.co/datasets/simbahuang/wan22-animate-3k-opensource-data.text-2-video-human-preferences-wan2.1
Rapidata Video Generation Alibaba Wan2.1 Human Preference
If you get value from this dataset and would like to see more in the future, please consider liking it.
This dataset was collected in ~1 hour total using the Rapidata Python API, accessible to anyone and ideal for large scale data annotation.
Overview
In this dataset, ~45'000 human annotations were collected to evaluate Alibaba Wan 2.1 video generation model on our benchmark. The up to date benchmark… See the full description on the dataset page: https://huggingface.co/datasets/Rapidata/text-2-video-human-preferences-wan2.1.UAV-Flow-Simdataset_all_new1C4-Eval
C4-Eval
C4-Eval is the evaluation set for C4 Bench, a Chengyu-based benchmark for measuring whether multimodal language models can understand cross-concept creativity. The release contains the original images, the corresponding idiom answers, and the complete task-specific questions used for evaluation.
221 base items: 37 human-designed seed figures and 184 bridge-controlled synthetic figures.
1,105 evaluation instances: five task formulations for every base item.
Language:… See the full description on the dataset page: https://huggingface.co/datasets/sci-m-wang/C4-Eval.UniMER_Dataset
UniMER Dataset
For detailed instructions on using the dataset, please refer to the project homepage: UniMERNet Homepage
Introduction
The UniMER dataset is a specialized collection curated to advance the field of Mathematical Expression Recognition (MER). It encompasses the comprehensive UniMER-1M training set, featuring over one million instances that represent a diverse and intricate range of mathematical expressions, coupled with the UniMER Test Set, meticulously… See the full description on the dataset page: https://huggingface.co/datasets/wanderkid/UniMER_Dataset.CrossPoint-Bench
CrossPoint-Bench
CrossPoint-Bench is a comprehensive benchmark for evaluating Vision-Language Models (VLMs) on cross-view point correspondence tasks. It assesses models' abilities to spatial understanding, and correspondence between different viewpoints.
Dataset Structure
CrossPoint-Bench/
├── CrossPoint-Bench.jsonl # Main benchmark data file
└── image/
├── origin_image/ # Original scene images organized by scene ID
│ ├──… See the full description on the dataset page: https://huggingface.co/datasets/WangYipu2002/CrossPoint-Bench.3DDFA-V3
Data Card for 3DDFA-V3
This repository provides the data used in CVPR2024 (Highlight) paper 3DDFA_V3. Please see our github repository for details.
Assets Summary
assets
├── face_model.npy # (face model)
├── large_base_net.pth
├── net_recon.pth # (backbone)
├── net_recon_mbnet.pth # (MobileNet-V3 backbone, optional)
├── retinaface_resnet50_2020-07-20_old_torch.pth
├── similarity_Lm3D_all.mat
├── indices_38365_35709.npy #… See the full description on the dataset page: https://huggingface.co/datasets/Zidu-Wang/3DDFA-V3.liberoThis dataset was created using LeRobot.
Dataset Structure
meta/info.json:
{
"codebase_version": "v2.0",
"robot_type": "panda",
"total_episodes": 379,
"total_frames": 101469,
"total_tasks": 10,
"total_videos": 0,
"total_chunks": 1,
"chunks_size": 1000,
"fps": 10,
"splits": {
"train": "0:379"
},
"data_path": "data/chunk-{episode_chunk:03d}/episode_{episode_index:06d}.parquet",
"video_path":… See the full description on the dataset page: https://huggingface.co/datasets/wangmingxuan/libero.AgenticOCR-RL
AgenticOCR RL Training Data
Reinforcement-learning (GRPO) data for the AgenticOCR project.
The dataset contains 20,269 training records in rl_combined_v2.json. Image paths in each record are relative to the repository root and point into rl_images/.
LMEE-Bench
Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied Exploration
CVPR 2026
[arXiv]
LMEE-Bench
lmee_bench_sub: Includes 58 tasks.
lmee_bench: Includes the full 166 tasks.
task_test: Trajectory data test set.
imageworldtree-datalaion-high-resolution-chinese
laion-high-resolution-chinese
简介 Brief Introduction
取自Laion5B-high-resolution多语言多模态数据集中的中文部分,一共2.66M个图文对。
A subset from Laion5B-high-resolution (a multimodal dataset), around 2.66M image-text pairs (only Chinese).
数据集信息 Dataset Information
大约一共2.66M个中文图文对。大约占用381MB空间(仅仅是url等文本信息,不包含图片)。
Homepage: laion-5b
Huggingface: laion/laion-high-resolution
下载 Download
mkdir release && cd release
for i in {00000..00015}; do wget… See the full description on the dataset page: https://huggingface.co/datasets/wanng/laion-high-resolution-chinese.wan_lorascoco2017_train_512x_image_caption_cannyhttps://github.com/wangherr/coco2017_for_huggingface
WanVideosMM-SafetyBenchcoco2017_train_1024x_image_caption_cannyhumanposerefCOCOg_9k_840_sam2_parquet
refCOCOg 9k 840 with SAM2 Masks
This dataset is derived from refCOCOg_9k_840 and adds a mask column generated offline with SAM2.
Each sample contains:
id: sample identifier
problem: referring expression / query
solution: original box and point annotations
image: RGB image stored as Hugging Face image bytes
img_height: original metadata height
img_width: original metadata width
mask: SAM2-generated binary mask stored as PNG bytes
The mask column is a pseudo-label generated from… See the full description on the dataset page: https://huggingface.co/datasets/wanwan1111/refCOCOg_9k_840_sam2_parquet.EuroSAT-SAR
EuroSAT-SAR: Land Use and Land Cover Classification with Sentinel-1
The EuroSAT-SAR dataset is a SAR version of the popular EuroSAT dataset. We matched each Sentinel-2 image in EuroSAT with one Sentinel-1 patch according to the geospatial coordinates, ending up with 27,000 dual-pol Sentinel-1 SAR images divided in 10 classes. The EuroSAT-SAR dataset was collected as one downstream task in the work FG-MAE to serve as a CIFAR-like, clean, balanced ML-ready dataset for remote sensing… See the full description on the dataset page: https://huggingface.co/datasets/wangyi111/EuroSAT-SAR.dataset_all_newcoco2017_train_512x_image_caption_depth
