datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VOST-TAS
[NeurIPS 2025] Tracking and Understanding Object Transformations
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Dataset Visualizations: GitHub
Paper: arXiv:2511.04678
Project Page: tubelet-graph.github.io
Project Repository: GitHub
Point of Contact: Yihong Sun
📊 Dataset Overview
VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.robomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.PhySciBench
PhySciBench
PhySciBench is a benchmark for evaluating deep-research capabilities in the physical sciences, introduced in "Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark" (arXiv:2606.18648).
📖 Paper: https://arxiv.org/abs/2606.18648
💻 Code & evaluation: https://github.com/yigengjiang/physci-deepresearch
Overview
PhySciBench comprises 200 expert-curated questions (a single test split), balanced between physics and… See the full description on the dataset page: https://huggingface.co/datasets/yigengx/PhySciBench.CHUBS
CHUBS: A Large-Scale Dataset of Chu Bamboo Slip Script
Code | Paper (upcoming)
Introduction
This is a large-scale dataset of Chu bamboo slip (CBS, Chinese: 楚简, chujian) script, an ancient Chinese script used during the Spring and Autumn period over 2,000 years ago. This dataset consists of two parts:
The main dataset where each example is an image and the corresponding text label. This part is contained in the glyphs.zip ZIP file.
A character detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/chen-yingfa/CHUBS.SMR
SliME SMR data card
Dataset details
Dataset type:
Generation of Source Data and Instruction Data. The creation of SMR involves a meticulous amalgamation of publicly available datasets, comprising Arxiv-QA, ScienceQA, MATH-Vision, TextBookQA, GeoQA3, Geometry3K, TabMWP, DVQA, AI2D, and ChartVQA.
The disparities between SMR and conventional instruction tuning datasets manifest in two key aspects:
Challenging Reasoning Tasks: Many of the tasks in Physical/Social… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/SMR.Resource2Skill
Resource2Skill — Skill Library
Executable skill libraries for the Resource2Skill
runtime: reusable, structured skills that a software agent browses, inspects,
and composes to operate real tools (Web, PowerPoint, Excel, Blender, and
REAPER-style audio) and produce artifacts.
This dataset is the skill data half of the project; the runnable runtime,
MCP servers, and CLI live in the code repository.
Contents
skills_wiki/ Structured wiki entries used for runtime… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/Resource2Skill.pointarena_dataset
Molmo2 PointArena SFT Data
26,596 supervised pointing examples used to fine-tune Molmo2-8B
(yiyangd/molmo2-8b-ft)
into a stronger PointArena solver (76.2% → up from 73.9% base, +2.3 pp).
Provenance
Each record is (image, query, answer):
image: A LAION-2B image sampled by reservoir sampling. Stored under
laion_images/<bucket>/<hash>.jpg (bucket is the first 2 hex chars of the
SHA-1 hash of the image URL, used to spread files across folders).
query: A natural-language… See the full description on the dataset page: https://huggingface.co/datasets/yiyangd/pointarena_dataset.HatefulIllusion_Dataset[Disclaimer] This dataset contains harmful content and can only be used for research or educational purposes!
Dataset Description
This dataset is generated and used in the paper:
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions (ICCV 2025)
It contains 2,160 (hateful) AI-generated optical illusions that hide three types of messages:
digits: 10 messages, 300 AI-generated illusions
hate slangs (hate speech): 23 messages, 690 AI-generated illusions
hate… See the full description on the dataset page: https://huggingface.co/datasets/yiting/HatefulIllusion_Dataset.PaperWritingBench
PaperWritingBench 🎻
PaperWritingBench is the first benchmark designed to evaluate how well autonomous AI research paper writing systems can synthesize raw research materials into submission-ready papers.
[Paper] [Project Page] [Code]
Dataset Structure
This repository contains:
datasets.zip: The full dataset containing cvpr2025 and iclr2025 folders with raw materials.
metadata.json: A JSON file listing metadata for all 200 papers, including venue, paper ID, number… See the full description on the dataset page: https://huggingface.co/datasets/yiwen-song/PaperWritingBench.MR2Bench
MR²Bench
MR²Bench (Multi-Response Multimodal Reward Bench) is a pair of benchmarks for evaluating multimodal reward models on N-way ranking tasks. Unlike existing benchmarks that only support pairwise comparisons, MR²Bench provides N-way human-annotated rankings over responses from multiple diverse models, enabling evaluation of both pairwise and listwise ranking capabilities.
This benchmark is described in our paper:
You Only Judge Once: Multi-response Reward Modeling in a… See the full description on the dataset page: https://huggingface.co/datasets/yinuoy/MR2Bench.RoMMathluxury-products-dataMedXpertQA
Dataset Card for MedXpertQA
MedXpertQA is a highly challenging and comprehensive benchmark designed to evaluate expert-level medical knowledge and advanced reasoning capabilities. It features both text-based and multimodal question-answering tasks, with the multimodal subset leveraging structured clinical information alongside images.
Dataset Description
MedXpertQA comprises 4,460 questions spanning diverse medical specialties, tasks, body systems, and image… See the full description on the dataset page: https://huggingface.co/datasets/yiyanhuang/MedXpertQA.SFTDPOdataset
