datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
GUIGuard-Bench
GUIGuard-Bench (Public Ladder)
GUIGuard-Bench is a cross-platform GUI agent benchmark for studying privacy risks and privacy-preserving execution in multimodal GUI agents.
This public-ladder release contains 121 GUI interaction trajectories (68 Android + 53 PC) for benchmark evaluation, with 26,407 region-level privacy annotations across 2,002 screenshots.
For the anonymous review version of the evaluation toolkit, see GUIGaurd-Bench-CA4F.
Dataset Summary
GUI agents… See the full description on the dataset page: https://huggingface.co/datasets/ShaofantuoshuzhengzhiSha/GUIGuard-Bench.Gastric-X
Gastric-X
Multi-phase abdominal CT cohort paired with structured laboratory panels
and free-text radiology reports, in proficient medical English with
the original Simplified Chinese preserved alongside.
Changelog
2026-06-26
Added per-phase organ masks (<phase>_organ_mask.nii.gz) — CADS
multi-organ segmentation on each phase's CT grid (e.g. label 6 = stomach);
all 4897 phases.
Added per-phase gastric tumor masks (<phase>_tumor_mask.nii.gz,
binary) — a patient's… See the full description on the dataset page: https://huggingface.co/datasets/HaoChen2/Gastric-X.t2-ragbench
Dataset Card for T2-RAGBench
Project Page | Paper | Code
IMPORTANT NOTICE:
We deleted VQAonBD from the dataset due to low quality of the question reformulations. If you still want to use it you will find the data in the previous commit history.
Dataset Description
Dataset Summary
T2-RAGBench is a benchmark dataset designed to evaluate Retrieval-Augmented Generation (RAG) on financial documents containing both text and tables. It consists of 23,088… See the full description on the dataset page: https://huggingface.co/datasets/G4KMU/t2-ragbench.MAPBench-V2For more details, please check our project page.
Paper: https://arxiv.org/abs/2601.05432
Repository: https://github.com/AMAP-ML/Thinking-with-Map
GITQA-Aug-LegacyGUI-Odyssey
Dataset Card for GUI Odyssey
News⭐️
A new and improved version of the GUIOdyssey dataset has been released! 🎉🎉
👉 Please use the latest version and refer to the updated README for the most up-to-date information.
We highly recommend using the new version for all training and evaluation!
Repository: https://github.com/OpenGVLab/GUI-Odyssey
Latest Version of Dataset: hflqf88888/GUIOdyssey
Paper: https://arxiv.org/pdf/2406.08451
Introduction
GUI Odyssey is… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/GUI-Odyssey.GMAI-VL-5.5M
GMAI-VL-5.5M Dataset
GMAI-VL-5.5M is a comprehensive, large-scale medical General Medical AI Vision-Language (GMAI-VL) dataset built specifically for training multimodal foundation models in the medical domain. It contains an extraordinary scale of high-quality instructions encompassing over 5.5 million multimodal question-answering pairs, carefully constructed based on hundreds of medical classification, segmentation, and detection datasets.
This repository… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-VL-5.5M.physical-ai-bench-generation
Physical AI Bench - Generation
Paper | Code
Dataset Description
The PAI-Bench is a benchmark to measure the progress of world models quantitatively.
The predict task contains a list of 1044 samples of text prompts, conditioning images, and qa pairs, covering Physical AI target domains including autonomous vehicle (AV) driving, robotics, industry (smart space), physics, human, and common sense. All the questions are binary questions, and the answer is either Yes or No. Our… See the full description on the dataset page: https://huggingface.co/datasets/shi-labs/physical-ai-bench-generation.gdpval_preference_rubricsGMAI-Reasoning10K
GMAI-Reasoning10K
Medical Reasoning dataset used in GMAI-VL-R1
Data description
GMAI-Reasoning10K is a high-quality medical image reasoning dataset containing 10,000 carefully selected samples. The data was collected from 95 medical datasets from reliable sources such as Kaggle, GrandChallenge, and Open-Release, covering 12 imaging modalities including X-ray, CT, and MRI.
Data preprocessing followed the standardization methods from SAMed-20M: 3D data (CT/MRI) had… See the full description on the dataset page: https://huggingface.co/datasets/General-Medical-AI/GMAI-Reasoning10K.gemma4-serving-bench-data
Gemma 4 12B (QAT-Q4_0) — Serving-Behavior Test Data
Test data, charts, and the running research log from an autonomous research
loop characterizing and tuning a Gemma 4 12B QAT-Q4_0 model served via
llama.cpp/llamafile on a single RTX 3080 Ti. Every ~30 min the loop
summarizes findings, proposes a goal, tests it end-to-end, documents success or
failure, and publishes here + to GitHub.
Model under test: gemma-4-12b-it-qat-q4_0.gguf (Google, June 2026), 128K
ctx, f16 KV, MTP… See the full description on the dataset page: https://huggingface.co/datasets/SEBK4C/gemma4-serving-bench-data.GC10-DET
GC10-DET: Metallic Surface Defect Detection Dataset (Object Detection)
Unofficial redistribution of the GC10-DET metallic surface defect detection dataset, reformatted into a standardized YOLO-compatible directory layout with a deterministic train/val/test split.
Disclaimer
This repository is not an official release of the GC10-DET dataset.
GC10-DET was created by Xiaoming Lv, Fajie Duan, Jia-jia Jiang, Xiao Fu, and Lin Gan (Tianjin University), who… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/GC10-DET.GWHD
GWHD 2021: Global Wheat Head Dataset (Object Detection)
Unofficial redistribution of the Global Wheat Head Dataset (GWHD) 2021 competition release, reformatted into a standardized YOLO-compatible directory layout.
Disclaimer
This repository is not an official release of the Global Wheat Head Dataset.
GWHD was created by the Global Wheat Head Detection consortium — a multi-institution, multi-country collaboration (David, Serouart, Madec, et al.) — who… See the full description on the dataset page: https://huggingface.co/datasets/dronefreak/GWHD.guertin-mcro-forensic-corpus-hearing-media
Guertin MCRO Forensic Corpus: Hearing Media
Contents: 215 hearing video clips and 213 WebVTT caption files for 4 hearings in State of Minnesota v. Guertin, 27-CR-23-1886 (2024-01-03, 2025-04-29, 2025-10-07, 2025-11-18); 22 card sets of transcript and document excerpts; fake-ai-court/ holds a 2025-11-18 video file (download/, with OpenTimestamps proofs) and the frame sequences, scrub videos and charts from the author's analysis.
Layout: <date>--video-clips/ (clips and captions);… See the full description on the dataset page: https://huggingface.co/datasets/Matt1up/guertin-mcro-forensic-corpus-hearing-media.STEM2Crystal-Bench
STEM2Crystal-Bench
STEM2Crystal-Bench is the benchmark for the paper "From Noisy STEM to Crystal Structure: Evidence-Structure CoDiffusion under Composition Constraints" (Chen & You, KDD 2026, Oral), which introduces STEM2Crystal CoDiffusion (SCCD). It evaluates methods that reconstruct a crystal structure from a noisy STEM image when the composition is known. The release has a large synthetic set with controlled noise and a small set of real STEM images, with ground-truth CIFs… See the full description on the dataset page: https://huggingface.co/datasets/gary23ai/STEM2Crystal-Bench.GUI-CC
GUI-CC
GUI-CC is a benchmark for evaluating the contextual consistency of GUI world models when
they are used as agent environments rather than as isolated next-screen predictors.
A GUI world model predicts the next interface given the current screenshot and an action.
When that prediction is fed back as the next state, the rollout must stay coherent: app identity,
navigation history, created entities, selected options, and action affordances all have to remain
mutually… See the full description on the dataset page: https://huggingface.co/datasets/minuzero/GUI-CC.FiVE-Fine-Grained-Video-Editing-Benchmark
FiVE-Bench
FiVE-Bench: A Fine-Grained Video Editing Benchmark for Evaluating Diffusion and Rectified Flow Models
Minghan Li1*, Chenxi Xie2*, Yichen Wu13, Lei Zhang2, Mengyu Wang1†
1Harvard University 2The Hong Kong Polytechnic University 3City University of Hong Kong
*Equal contribution †Corresponding Author
💜 Leaderboard (coming soon) |
💻 GitHub |
🤗 Hugging Face
📝 Project Page |
📰 Paper |
🎥 Video Demo
FiVE is a benchmark comprising 100 videos for… See the full description on the dataset page: https://huggingface.co/datasets/LIMinghan/FiVE-Fine-Grained-Video-Editing-Benchmark.omnidocbench-render-compare
OmniDocBench Render-and-Compare
This dataset contains the rendered HTML reconstructions and comparison images produced
by a render-and-compare pipeline — a reference-free visual similarity evaluation
framework for OCR systems.
Overview
The pipeline processes each page of OmniDocBench through
a Qwen3.5-122B-A10B OCR model, renders the structured output back to a PNG via HTML
(reconstructed.png), and compares it against the original page scan (masked_original.png)
using… See the full description on the dataset page: https://huggingface.co/datasets/gt-free-ocr-metrics/omnidocbench-render-compare.shelf-photos-batch1
shelf-photos-batch1
Working dataset for the shelf-monitoring pipeline: bootstrap labeling and
fine-tuning data management for Geraldine/rf-detr-nano-bookshelf.
Layout
photos/ — 25 of the library's own shelf photos, untouched (no crops, no upscaling)
original_images_library/ — 285 high-res images from
llabres/library-dataset (MIT),
used for continuous fine-tuning (domain shift: real library stacks)
dataset/Bookshelf-recognition-2.v1i.coco.zip — COCO export of… See the full description on the dataset page: https://huggingface.co/datasets/Geraldine/shelf-photos-batch1.omniact-gui-trajectories
OmniACT
OmniACT is a GUI trajectory dataset with single-step action traces grounded in screenshots.
Dataset Structure
.
├── README.md
├── .gitattributes
├── data/
│ └── train.jsonl
├── observations/
│ └── OmniACT_pilot_*/000/screenshot.jpg
└── env_meta/
└── OmniACT_pilot_*/000/metadata.json
Each row in data/train.jsonl is one trajectory. The main image path is stored in the top-level image field, and the same relative path is also used inside… See the full description on the dataset page: https://huggingface.co/datasets/Dhscl/omniact-gui-trajectories.noteflow-research-pilots
Keep the failed attempts. Check the artifact.
Versioned public development evidence from Robot Reel × Skills Anywhere × EvalArc, recorded 14 September 2026 on an NVIDIA L40S, with separate scripted Harbor controls on CPU and separate GPU context-control and agent-requested MCP handoff cohorts recorded 19 September 2026. This is an inspectable engineering casebook, not a held-out benchmark or training corpus with established efficacy.
Configuration
Actual experiment
What… See the full description on the dataset page: https://huggingface.co/datasets/glayguo/noteflow-research-pilots.KnowGen-Bench
KnowGen Benchmark
Project Page | Paper | Code
This repository contains the KnowGen benchmark data for Gen-Searcher: Reinforcing Agentic Search for Image Generation.
👀 Intro
We introduce Gen-Searcher, as the first attempt to train a multimodal deep research agent for image generation that requires complex real-world knowledge. Gen-Searcher can search the web, browse evidence, reason over multiple sources, and search visual referencesbefore generation, enabling… See the full description on the dataset page: https://huggingface.co/datasets/GenSearcher/KnowGen-Bench.gaokao-related-math
problems-scraper
爬取 出卷网 高考专区-数学试卷,转换为 JSONL 格式。
安装
pnpm install
用法
# 首次试跑 50 套
pnpm scrape:test
# 全量爬取 (3027 套,约 5-6 小时)
pnpm scrape:full
# 增量爬取 (每天最新)
pnpm scrape:resume
输出
data/
├── jsonl/problems.jsonl # 每行一个 JSON 题目
└── images/<id>/ # 试卷配图(散点图/几何图等本地副本)
state/
├── state.json # 爬取状态(maxSeenDate / fromDate)
└── seen.bin # 已抓试卷 ID 集合(断点续抓)
Cron(每日增量)
# 每天 3:00~6:00… See the full description on the dataset page: https://huggingface.co/datasets/cheese233/gaokao-related-math.GVLQA-AUGETFIFA_World_Cup_Games
FIFA World Cup Full-Match Video Dataset
This dataset provides YouTube full-match references and structured match annotations for classic FIFA World Cup games. Instead of redistributing video files, it stores YouTube video IDs and URLs, then links each retained match to structured football data including match metadata, team statistics, lineups, formations, text event timelines, and pre-match prediction polls.
The dataset is designed for long-form sports… See the full description on the dataset page: https://huggingface.co/datasets/choucsan/FIFA_World_Cup_Games.geometry3k
geometry3k
A slice of geometry problems with diagram images and structured annotations.
Example Preview
problem_text: In \odot K, M N = 16 and m \widehat M N = 98. Find the measure of L N. Round to the nearest hundredth.
answer: C (≙ 8.94)
code
import matplotlib.pyplot as plt
import numpy as np
# Define points
points = {"J": [160.0, 24.0], "K": [93.0, 97.0], "L": [33.0, 166.0], "M": [1.0, 96.0], "N": [98.0, 186.0], "P": [52.0, 144.0]}
# Define lines
lines… See the full description on the dataset page: https://huggingface.co/datasets/LibraTree/geometry3k.GSEval
🚀 GSEval - A Comprehensive Grounding Evaluation Benchmark
GSEval is a meticulously curated evaluation benchmark consisting of 3,800 images, designed to assess the performance of pixel-level and bounding box-level grounding models. It evaluates how well AI systems can understand and localize objects or regions in images based on natural language descriptions.
📊 Results
📊 Download GSEval
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/hustvl/GSEval.GVLQA-AUGNOExpertHTR-Dataset
ExpertHTR Dataset
Gated page-level handwritten text recognition data for the
ExpertHTR project.
This is a rights-filtered replacement export: all HWDB/CASIA records and
images have been removed. The repository remains gated because the remaining
upstream sources have different access conditions. It is a companion data
release for ExpertHTR, not the exact training snapshot for the published
seven-source checkpoint.
Included data
Split
Records
Purpose… See the full description on the dataset page: https://huggingface.co/datasets/DAIR-Group/ExpertHTR-Dataset.grounded-visual-spatial-reasoning
Grounded Visual Spatial Reasoning
Code for generating the annotations can be found here: github.com
Dataset Summary
This dataset extends the Visual Spatial Reasoning (VSR) dataset with visual grounding annotations: each caption is annotated with COCO-category object mentions, their positions , and corresponding bounding boxes in the image.
Data instance
Each sample instance has the following structure:
Field
Type
Description
image_file
string… See the full description on the dataset page: https://huggingface.co/datasets/tomhodemon/grounded-visual-spatial-reasoning.
