datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
VOST-TAS
[NeurIPS 2025] Tracking and Understanding Object Transformations
If you like our project, please give us a star ⭐ on GitHub for the latest update.
💡 Description
Dataset Visualizations: GitHub
Paper: arXiv:2511.04678
Project Page: tubelet-graph.github.io
Project Repository: GitHub
Point of Contact: Yihong Sun
📊 Dataset Overview
VOST-TAS (TrackAnyState) is an extended version of the VOST validation set with explicit transformation annotations for tracking and… See the full description on the dataset page: https://huggingface.co/datasets/yihongs/VOST-TAS.GITQA-Aug-LegacyMovie101
Movie101
[!NOTE]
Please carefully read the Movie101 license before using the data.Current dataset version: Movie101v2
Audio Description (AD) describes movie content in real time to help visually impaired individuals enjoy movies, where a narration speech briefly summarizes the ongoing plots during pauses in character dialogue, help its audience keep up with the movie.
The AD creation involves extensive work by human experts, which is costly and difficult to cover the vast array… See the full description on the dataset page: https://huggingface.co/datasets/yuezih/Movie101.robomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.V3Det_Backup
License:
V3Det Images: Around 90% images in V3Det were selected from the Bamboo Dataset, sourced from the Flickr website. The remaining 10% were directly crawled from the Flickr. We do not own the copyright of the images. Use of the images must abide by the Flickr Terms of Use. We only provide lists of image URLs without redistribution.
V3Det Annotations: The V3Det annotations, the category relationship tree, and related tools are licensed under a Creative Commons Attribution 4.0… See the full description on the dataset page: https://huggingface.co/datasets/yhcao/V3Det_Backup.androidlife-530
AndroidLife-530 — Android agent benchmark (real phone, real LLM)
AndroidLife runs Android agent tasks against a real phone (via ADB/MobileRun)
and a real LLM, and grades the agent on reaching a verifiable device end-state.
This repo ships the 530-task corpus plus everything needed to reproduce runs.
Benchmark, or template — your call. The 530 tasks are an extended version
of the benchmark, usable as a larger evaluation set for further benchmarking of
models beyond the 60-task… See the full description on the dataset page: https://huggingface.co/datasets/YuvrajSingh9886/androidlife-530.yoga_posesParseBench
ParseBench
Quick links: [🌐 Website] [📜 Paper] [💻 Code]
ParseBench is a benchmark for evaluating document parsing systems on real-world enterprise documents, with the following characteristics:
Multi-dimensional evaluation. The benchmark is stratified into five capability dimensions — tables, charts, content faithfulness, semantic formatting, and visual grounding — each with task-specific metrics designed to capture what agentic workflows depend on.
Real-world enterprise… See the full description on the dataset page: https://huggingface.co/datasets/Yoongls/ParseBench.astra-robodojo-rollouts
Astra RoboDojo Evaluation Records
Rollout records from the evaluations in GPT 6 Astra as an Embodied Policy,
by Jiayi Su, Yixin Zheng, Mi Yan, Li Yi, Zhizheng Zhang, and He Wang. This archive
includes action proposals, executed actions, observations, robot states,
model-provided explanations and reasoning summaries, and metadata for reproducing
the evaluation settings, together with a reader and documentation.
Report
Public controller source
Data schema and alignment… See the full description on the dataset page: https://huggingface.co/datasets/YuMoool/astra-robodojo-rollouts.MMMC
MMMC: Massive Multi-discipline Multimodal Coding Benchmark for Educational Video Generation
Dataset Summary
The MMMC (Massive Multi-discipline Multimodal Coding) benchmark is a curated dataset for Code2Video research, focusing on the automatic generation of professional, discipline-specific educational videos. Unlike pixel-only video datasets, MMMC provides structured metadata that links lecture content with executable code, visual references, and topic-level annotations… See the full description on the dataset page: https://huggingface.co/datasets/YanzheChen/MMMC.medical-prescription-datasetRiOSWorld
News
2025-05-31: We released our paper, environment and benchmark, and project page. Check it out!
Download and Setup
# DownLoad the RiOSWorld risk examples
dataset = load_dataset("JY-Young/RiOSWorld", split='test')
The environmental risk examples require specific configuration. For specific configuration processes, please refer to: https://github.com/yjyddq/RiOSWorld
Data Statistics
Topic Distribution of User Instruction… See the full description on the dataset page: https://huggingface.co/datasets/JY-Young/RiOSWorld.ALL-Bench-Leaderboard
🏆 ALL Bench Leaderboard 2026
The only AI benchmark dataset covering LLM · VLM · Agent · Image · Video · Music in a single unified file.
Dataset Summary
ALL Bench Leaderboard aggregates and cross-verifies benchmark scores for 90+ AI models across 6 modalities. Every numerical score is tagged with a confidence level (cross-verified, single-source, or self-reported) and its original source. The dataset is designed for researchers, developers, and… See the full description on the dataset page: https://huggingface.co/datasets/youssef3146/ALL-Bench-Leaderboard.Traffic-VQAGVLQA-AUGETMIFSLanguage: English | 简体中文
Towards Scalable RLVR: Multimodal Instruction Following Data Synthesis and Distillation
A high-quality multimodal dataset (141,559 samples, 19,692 images) for RLVR training and evaluation of Multi-modal Large Language Models (MLLMs) on instruction-following tasks. Every RL/Eval sample ships with executable Python code_checker functions that programmatically verify whether a model response satisfies all constraints in the instruction — enabling fully automated… See the full description on the dataset page: https://huggingface.co/datasets/yrzeng/MIFS.GVLQA-AUGNOMMInstruct-GPT4V
MMInstruct
The official implementation of the paper "MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity".
The data engine is available on GitHub at yuecao0119/MMInstruct.
Todo List
Data Engine.
Open Source Datasets.
Release the checkpoint.
Introduction
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
Instruction Annotation… See the full description on the dataset page: https://huggingface.co/datasets/yuecao0119/MMInstruct-GPT4V.glaucoma-expert-cot-raw-1077
Glaucoma Expert Chain-of-Thought
Ophthalmologist six-step reasoning reports for fundus photographs, each paired with
a binary glaucoma label. 1,074 cases from LAG and Papila.
Files
file
rows
split
expert_cot_trainval.jsonl
915
train (823) + val (92)
expert_cot_test.jsonl
159
test
images/
1,074
<source>_<id>.jpg
Record schema
{
"id": "1689",
"source": "LAG",
"image": "LAG_1689.jpg",
"split": "train"… See the full description on the dataset page: https://huggingface.co/datasets/yuzhench/glaucoma-expert-cot-raw-1077.SpinBench
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
🌐 Project page |
🤗 Dataset |
📑 Paper |
💻 Code
SpinBench is a cognitively grounded diagnostic benchmark for evaluating spatial reasoning in vision-language models (VLMs).
SpinBench is designed around the core challenge of spatial reasoning: perspective taking, the ability to reason about how scenes and object relations change under viewpoint transformation. Since perspective taking requires… See the full description on the dataset page: https://huggingface.co/datasets/YuyouZhang/SpinBench.SVRDD_YOLO
SVRDD Road Damage Detection
This repository is formatted for the Hugging Face Dataset Viewer. Each row
contains an image and its road-damage object annotations.
Load the dataset
from datasets import load_dataset
ds = load_dataset("YOUR_USERNAME/YOUR_DATASET_NAME")
example = ds["train"][0]
print(example["image"])
print(example["objects"])
Local unzip usage
If you want the images as regular local files, unzip the split archives:
unzip train.zip -d… See the full description on the dataset page: https://huggingface.co/datasets/ShuoZheLi/SVRDD_YOLO.tool-calling-mix
This is a dataset for fine-tuning a language model to use tools. I combined sources from various other tool calling datasets and added some non-tool calling examples to prevent catastrophic forgetting.
Dataset Overview
Motivation
This dataset was created to address the need for a diverse, high-quality dataset for training language models in tool usage. By combining multiple sources and including non-tool examples, it aims to produce models that can effectively use tools… See the full description on the dataset page: https://huggingface.co/datasets/younissk/tool-calling-mix.GVLQA-AUGLYGVLQA-AUGNSPhySciBench
PhySciBench
PhySciBench is a benchmark for evaluating deep-research capabilities in the physical sciences, introduced in "Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark" (arXiv:2606.18648).
📖 Paper: https://arxiv.org/abs/2606.18648
💻 Code & evaluation: https://github.com/yigengjiang/physci-deepresearch
Overview
PhySciBench comprises 200 expert-curated questions (a single test split), balanced between physics and… See the full description on the dataset page: https://huggingface.co/datasets/yigengx/PhySciBench.CHUBS
CHUBS: A Large-Scale Dataset of Chu Bamboo Slip Script
Code | Paper (upcoming)
Introduction
This is a large-scale dataset of Chu bamboo slip (CBS, Chinese: 楚简, chujian) script, an ancient Chinese script used during the Spring and Autumn period over 2,000 years ago. This dataset consists of two parts:
The main dataset where each example is an image and the corresponding text label. This part is contained in the glyphs.zip ZIP file.
A character detection dataset… See the full description on the dataset page: https://huggingface.co/datasets/chen-yingfa/CHUBS.Vietnamese-yfcc15m-OpenAICLIPSMR
SliME SMR data card
Dataset details
Dataset type:
Generation of Source Data and Instruction Data. The creation of SMR involves a meticulous amalgamation of publicly available datasets, comprising Arxiv-QA, ScienceQA, MATH-Vision, TextBookQA, GeoQA3, Geometry3K, TabMWP, DVQA, AI2D, and ChartVQA.
The disparities between SMR and conventional instruction tuning datasets manifest in two key aspects:
Challenging Reasoning Tasks: Many of the tasks in Physical/Social… See the full description on the dataset page: https://huggingface.co/datasets/yifanzhang114/SMR.Resource2Skill
Resource2Skill — Skill Library
Executable skill libraries for the Resource2Skill
runtime: reusable, structured skills that a software agent browses, inspects,
and composes to operate real tools (Web, PowerPoint, Excel, Blender, and
REAPER-style audio) and produce artifacts.
This dataset is the skill data half of the project; the runnable runtime,
MCP servers, and CLI live in the code repository.
Contents
skills_wiki/ Structured wiki entries used for runtime… See the full description on the dataset page: https://huggingface.co/datasets/YijiaFan/Resource2Skill.pointarena_dataset
Molmo2 PointArena SFT Data
26,596 supervised pointing examples used to fine-tune Molmo2-8B
(yiyangd/molmo2-8b-ft)
into a stronger PointArena solver (76.2% → up from 73.9% base, +2.3 pp).
Provenance
Each record is (image, query, answer):
image: A LAION-2B image sampled by reservoir sampling. Stored under
laion_images/<bucket>/<hash>.jpg (bucket is the first 2 hex chars of the
SHA-1 hash of the image URL, used to spread files across folders).
query: A natural-language… See the full description on the dataset page: https://huggingface.co/datasets/yiyangd/pointarena_dataset.
