datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Quick Start… See the full description on the dataset page: https://huggingface.co/datasets/UnipatAI/RoadmapBench.RoboPulsePlusPlus
RoboPulse++
RoboPulse++ is an interval-level benchmark introduced in PRM-as-a-Judge 1.5: A Toolkit for Robot Process Assessment for evaluating progress judge models throughout complete robot manipulation trajectories.
This Hugging Face release contains 700 episodes with natural-language task instructions, temporally ordered observations, and human-annotated progress intervals.
Overview
RoboPulse++ extends progress judging from isolated state pairs to temporal… See the full description on the dataset page: https://huggingface.co/datasets/lyy0715/RoboPulsePlusPlus.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.RoadmapBench
RoadmapBench
A benchmark for evaluating AI coding agents on multi-target, long-horizon software development tasks derived from open-source project version upgrades.
Overview
RoadmapBench contains 115 tasks spanning 17 open-source repositories across 5 programming languages (Python, TypeScript, Go, Rust, C++). Each task requires an agent to implement multiple interdependent features that correspond to a real version upgrade of the target project.
Task Structure… See the full description on the dataset page: https://huggingface.co/datasets/benchmark-anon-2026/RoadmapBench.regionreasoner_datarobomme_preprocessed_data
RoboMME Training Data (Pickle Format)
Arxiv Paper | HF Paper | Website | Benchmark Code | Policy Learning Code
This repo contains preprocessed pickle files for RoboMME training data and npy files for cached image tokens. We use this dataset in our MME-VLA experiments.
.
├── data # zipped pickle files
├── features # zipped precompute siglip embeddings
├── meta # statistics for robomme
├── memer # VLM subgoal training data for MemER (only used for symbolic… See the full description on the dataset page: https://huggingface.co/datasets/Yinpei/robomme_preprocessed_data.PhD
[CVPR2025 Highlight] PhD: A ChatGPT-Prompted Visual hallucination Evaluation Dataset
preprint
🔥 PhD-webdataset
To enhance usability and integration with evaluation frameworks like lmm-eval, we are pleased to offer a packaged version in webdataset format. This packaged version is designed to facilitate easier deployment and testing. For further details and access, please refer to our repository PhD-webdataset.
Please note that the data in both repositories is completely… See the full description on the dataset page: https://huggingface.co/datasets/AIMClab-RUC/PhD.bankertoolbench
BankerToolBench
BankerToolBench is a benchmark of 100 end-to-end investment banking tasks for
evaluating AI agents. Each task mirrors real junior-banker work — building
financial models, preparing pitch decks, writing memos — and produces multi-file
deliverables (Excel, PowerPoint, Word) that are scored against expert-authored
rubrics.
The benchmark was developed with 502 investment bankers from firms including
Goldman Sachs, JPMorgan, Evercore, and others. Human completion time… See the full description on the dataset page: https://huggingface.co/datasets/handshake-ai-research/bankertoolbench.lgg-mri-segmentation-research
LGG Brain MRI Segmentation with Genomic Clusters
This repository provides a Patient-Centric version of the Lower-Grade Glioma (LGG) Segmentation dataset. While other versions of this data exist, they often treat slices as independent images. This version preserves the 3D patient volume and integrates all genomic/clinical labels directly into a multimodal-ready format.
🌟 Why This Version?
Developed for Multimodal AI Research, this dataset addresses several limitations… See the full description on the dataset page: https://huggingface.co/datasets/Ehsan-rmz/lgg-mri-segmentation-research.VBVR-Bench-Data
VBVR: A Very Big Video Reasoning Suite
Overview
Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture,
enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality.
Systematically studying video reasoning and its scaling behavior suffers from a lack of… See the full description on the dataset page: https://huggingface.co/datasets/Video-Reason/VBVR-Bench-Data.Reason-RFT-CoT-Dataset
🤗 Reason-RFT CoT Dateset
The full dataset used in our project "Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning".
⭐️ Project │ 🌎 Github │ 🔥 Models │ 📑 ArXiv │ 💬 WeChat
🤖 RoboBrain: Aim to Explore ReasonRFT Paradigm to Enhance RoboBrain's Embodied Reasoning Capabilities.
♣️ Quick Start
Please refer to Dataset Preparation
🔥 Overview
Visual reasoning abilities play a crucial role in understanding complex multimodal… See the full description on the dataset page: https://huggingface.co/datasets/tanhuajie2001/Reason-RFT-CoT-Dataset.DGM4vsr_random
VSR: Visual Spatial Reasoning
This is the random set of VSR: Visual Spatial Reasoning (TACL 2023) [paper].
Usage
from datasets import load_dataset
data_files = {"train": "train.jsonl", "dev": "dev.jsonl", "test": "test.jsonl"}
dataset = load_dataset("cambridgeltl/vsr_random", data_files=data_files)
Note that the image files still need to be downloaded separately. See data/ for details.
Go to our github repo for more introductions.
Citation
If you find VSR… See the full description on the dataset page: https://huggingface.co/datasets/cambridgeltl/vsr_random.gdpval_preference_rubricsGMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.refute
Can AI read new science honestly?
Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next.
REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence.
Truth Score is the main result. It combines factual accuracy, flaw… See the full description on the dataset page: https://huggingface.co/datasets/BGPT-OFFICIAL/refute.warwick-second-life-dm-2025-raw
First-life and second-life battery degradation mode test data
BSEBench status: raw_mirror_pending_validation
This repository is a raw mirror of the Mendeley Data dataset Test_Data from Sadia Tasnim Mowri, associated with the University of Warwick. The source description states that the dataset was created to study the influence of first-life degradation mode on second-life performance and degradation, with first-life cells brought to around 80% SoH and then evaluated in second-life… See the full description on the dataset page: https://huggingface.co/datasets/bsebench-org/warwick-second-life-dm-2025-raw.sat-vl-sft-training-ready-v1
Dataset Summary
NuTonic/sat-bbox-metadata-sft-v1 is a metadata-first, procedural VLM SFT dataset built from an existing “sat-bbox” style dataset tree (Sentinel‑2 chips + per-tile JSON metadata sidecars, optionally paired Mapbox stills).
The goal is to create high-signal, production-shaped supervision for multimodal chat models:
Captioning for satellite chips
Grounding (bounding boxes in normalized coordinates) for land-cover regions
Class-focused captions and absence checks for… See the full description on the dataset page: https://huggingface.co/datasets/NuTonic/sat-vl-sft-training-ready-v1.vlm-info-loss-results
VLM Grounding Evaluation Results
Grounding evaluation results for vision-language models on robotics manipulation datasets.
Part of the vlm-info-loss project studying
how VLM connectors transform visual representations.
Background
Our embedding-level analysis shows VLM connectors perform a compress-then-expand transformation:
they sharpen dominant-object representations while compressing secondary-object category identity.
All tested models converge to ~83%… See the full description on the dataset page: https://huggingface.co/datasets/MicroAGI-Labs/vlm-info-loss-results.remote-jobs
Jobicy Remote Jobs
An automatically updated dataset of remote job opportunities published by Jobicy.
The dataset is designed for developers, researchers, analysts, search systems, AI agents, RAG applications, labor-market research, career tools, and other applications that need structured remote-job data.
Data source
The source data comes from the public Jobicy Remote Jobs API:
https://jobicy.com/api/v2/remote-jobs
Each record includes a canonical Jobicy job URL… See the full description on the dataset page: https://huggingface.co/datasets/jobicy/remote-jobs.RSCC-RSEdit-Test-Split
RSCC-RSEdit-Test-Split
This directory contains the test split for RSCC-RSEdit dataset.
Directory Structure
RSCC-RSEdit-Test-Split/
├── images/ # Original images (676 PNG files)
├── masks/ # Original grayscale masks (338 PNG files)
│ └── [mask files with pixel values 0,1,2,3,4]
├── masks_colorful/ # Colorful RGBA visualization masks (338 PNG files)
│ └── [same filenames as masks/, but in RGBA format with colors]
├──… See the full description on the dataset page: https://huggingface.co/datasets/BiliSakura/RSCC-RSEdit-Test-Split.astra-robodojo-rollouts
Astra RoboDojo Evaluation Records
Rollout records from the evaluations in GPT 6 Astra as an Embodied Policy,
by Jiayi Su, Yixin Zheng, Mi Yan, Li Yi, Zhizheng Zhang, and He Wang. This archive
includes action proposals, executed actions, observations, robot states,
model-provided explanations and reasoning summaries, and metadata for reproducing
the evaluation settings, together with a reader and documentation.
Report
Public controller source
Data schema and alignment… See the full description on the dataset page: https://huggingface.co/datasets/YuMoool/astra-robodojo-rollouts.Real-UI-Clickboxes
RUC: Real UI Clickboxes
Click carefully, even when the page is trying to trick you! 👀
Official Hugging Face release for RUC: Real UI Clickboxes, the dataset accompanying our ACL 2026 paper Don't Click That: Teaching Web Agents to Resist Deceptive Interfaces on deceptive UI understanding for web agents.
ACL Anthology: https://aclanthology.org/2026.acl-long.310/
PDF: https://aclanthology.org/2026.acl-long.310.pdf
DOI: https://doi.org/10.18653/v1/2026.acl-long.310… See the full description on the dataset page: https://huggingface.co/datasets/DUDE-Framework/Real-UI-Clickboxes.qwen36-kquant-offload-mtp-swebench-lite100-results
Qwen3.6 K-Quant Offload MTP SWE-bench Lite 100 Results
This dataset contains the complete 5-model x 100-prompt runtime benchmark artifacts plus a detailed statistical analysis layer.
Primary conclusion: hot30/cold30 was the best decode-throughput run, while Q4_K_M had the best total wall clock. The ATX hot30/cold30 quantization significantly outperformed both Q4_K_M and Q3_K_XL on paired decode throughput, but Q4_K_M remains the elapsed-time control.
The ATX/K3 hot10, hot20, and… See the full description on the dataset page: https://huggingface.co/datasets/jakeatx/qwen36-kquant-offload-mtp-swebench-lite100-results.SPADES-RGBRiOSWorld
News
2025-05-31: We released our paper, environment and benchmark, and project page. Check it out!
Download and Setup
# DownLoad the RiOSWorld risk examples
dataset = load_dataset("JY-Young/RiOSWorld", split='test')
The environmental risk examples require specific configuration. For specific configuration processes, please refer to: https://github.com/yjyddq/RiOSWorld
Data Statistics
Topic Distribution of User Instruction… See the full description on the dataset page: https://huggingface.co/datasets/JY-Young/RiOSWorld.Drivegpt4_raw_dataembodied-spatial-reasoning
Embodied Spatial Reasoning Tasks
Dataset Description
This dataset is part of the embodied-spatial-reasoning project, where the agent has to actively explore the environment to determine if certain spatial relationships hold true. The tasks involve spatial reasoning with various objects and scenes. Each task includes a query about the spatial relationships between objects within a scene, which the agent must verify through exploration.
Dataset Structure
The… See the full description on the dataset page: https://huggingface.co/datasets/thanhqt2002/embodied-spatial-reasoning.Penguin-Recap-I
Penguin-Recap-I
Penguin-Recap-I publishes recap metadata only. The repository does not contain
image binaries.
Included subsets
subset
collection
local source roots
expected records
datacomp_coyo_penguin
DataComp + COYO Penguin recap
datamultimodal/IMAGE/datacomp_1b, datamultimodal/IMAGE/coyo_700m
57,618,155
sa1b_penguin
SA-1B Penguin recap
datamultimodal/IMAGE/SA-1B
9,254,501
openimages_penguin
OpenImages Penguin recap… See the full description on the dataset page: https://huggingface.co/datasets/tencent/Penguin-Recap-I.RubricRM-Data
Link
GitHub: SKYLENAGE-AI/SKYLENAGE-JUDGER
Hugging Face Models:
skylenage-ai/SkyJM-Gen-4B
skylenage-ai/SkyJM-Gen-9B
skylenage-ai/SkyJM-Edit-4B
skylenage-ai/SkyJM-Edit-9B
Hugging Face Dataset: skylenage-ai/RubricRM-Data
ModelScope Models:
SKYLENAGE/SkyJM-Gen-4B
SKYLENAGE/SkyJM-Gen-9B
SKYLENAGE/SkyJM-Edit-4B
SKYLENAGE/SkyJM-Edit-9B
Citation
If you find this dataset useful, please cite our paper:
@misc{kan2026rubricrmgenerativerewardmodeling… See the full description on the dataset page: https://huggingface.co/datasets/skylenage-ai/RubricRM-Data.
