datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
XLRS-Bench_visual_grounding_en
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_en.HoliSpatial-3D-GroundingPR1-Datasets-GroundingXLRS-Bench_visual_grounding_zh
🐙GitHub
Information or evaluatation on this dataset can be found in this repo: https://github.com/AI9Stars/XLRS-Bench
📜Dataset License
Annotations of this dataset is released under a Creative Commons Attribution-NonCommercial 4.0 International License. For images from:
DOTARGB images from Google Earth and CycloMedia (for academic use only; commercial use is prohibited, and Google Earth terms of use apply).
ITCVDLicensed under CC-BY-NC-SA-4.0.
MiniFrance… See the full description on the dataset page: https://huggingface.co/datasets/initiacms/XLRS-Bench_visual_grounding_zh.FACTS-grounding-public
FACTS Grounding 1.0 Public Examples
860 public FACTS Grounding examples from Google DeepMind and Google Research
FACTS Grounding is a benchmark from Google DeepMind and Google Research designed to measure the performance of AI Models on factuality and grounding.
▶ FACTS Grounding Leaderboard on Kaggle▶ Technical Report▶ Evaluation Starter Code▶ Google DeepMind Blog Post
Usage
The FACTS Grounding benchmark evaluates the ability of Large Language Models (LLMs)… See the full description on the dataset page: https://huggingface.co/datasets/google/FACTS-grounding-public.easyr1-grounding-dataset-30k-not_grounded-SE-GUI-3B-2MPFiftyOne-GUI-Grounding-Train
Dataset Card for FiftyOne GUI Grounding Training Set
Dataset Details
Dataset Description
This dataset contains 739 annotated GUI screenshots designed for training computer vision models to understand and interact with graphical user interfaces. The dataset uses the specialized COCO4GUI format, which extends the standard COCO detection format to handle GUI-specific features, interaction sequences, and rich metadata.
The dataset captures real user… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/FiftyOne-GUI-Grounding-Train.FiftyOne-GUI-Grounding-Train-with-Synthetic
Dataset Card for FiftyOne GUI Grounding Training Set with Synthetic Augmentation
Dataset Details
Dataset Description
This dataset represents a significant expansion of the original FiftyOne GUI Grounding Training Set, growing from 739 real GUI screenshots to 4,036 total samples through systematic synthetic data generation. The dataset combines authentic GUI interactions with carefully crafted synthetic variants designed to improve model robustness… See the full description on the dataset page: https://huggingface.co/datasets/Voxel51/FiftyOne-GUI-Grounding-Train-with-Synthetic.GroundingME
(CVPR 2026) GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
🔍 Overview
Visual grounding—localizing objects from natural language descriptions—represents a critical bridge between language and vision understanding. While multimodal large language models (MLLMs) achieve impressive scores on existing benchmarks, a fundamental question remains: can MLLMs truly ground language in vision with human-like… See the full description on the dataset page: https://huggingface.co/datasets/lirang04/GroundingME.sequential-3d-grounding
Sequential 3D Visual Grounding Dataset
5 indoor datasets (ScanNet, HM3D, 3RScan, ARKitScenes, MultiScan) · 10301 scenes · 117885 sequences · 579872 steps
Overview
Dataset
Scenes
Train
Val
Test
ScanNet
1513
14501
1700
1808
HM3D
2302
33302
7663
4232
3RScan
1381
13826
4864
1937
ARKitScenes
4834
24349
1394
2653
MultiScan
271
4133
858
665
Total
10301
90111
16479
11295
Annotation
Each step has one target (the object to locate)… See the full description on the dataset page: https://huggingface.co/datasets/Ziyannn/sequential-3d-grounding.PanoCaps
PanoCaps: A Human-Annotated Benchmark for Panoptic Grounded Captioning
PanoCaps is a benchmark for panoptic grounded captioning: a model writes a full-scene caption and grounds every mentioned entity, things and stuff alike, to pixel-level masks.
It contains 3,470 images and 34K panoptic regions, averaging ~9 grounded entities per image, with >99% of regions grounded. Captions are human-written and verified, cover the entire visible scene, use open-vocabulary… See the full description on the dataset page: https://huggingface.co/datasets/Panorama-grounding/PanoCaps.grounding-data
grounding_data
The annotation tree behind a grounding/segmentation training stack, plus the images and
video frames that live alongside it. Unlike the companion
royguw/mm-olmo-images — which is
pixels only — this repo carries the annotations: the parquet caches, JSON label files
and vocabularies under each dataset's cache/ directory, which is where masks, boxes,
referring expressions, captions and category vocabularies actually live.
9,331,994 files / 1.20 TiB, packed as 255 tar… See the full description on the dataset page: https://huggingface.co/datasets/royguw/grounding-data.grounding_dataset
Grounding Dataset
A comprehensive, high-quality dataset for GUI element grounding tasks, curated from multiple authoritative sources to provide diverse, well-annotated interface interactions.
Overview
This dataset combines and standardizes annotations from five major GUI interaction datasets:
Aria-UI
OmniAct
Widget Caption
UI-Vision
OS-Atlas
Dataset Schema
Each sample contains the following fields:
Field
Type
Description
Example
dataset
string… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/grounding_dataset.highlevel_thinking_with_grounding_annotation_split1000_v3grounding_datasetFor desktop and web datasets in GUI grounding, the data is generally collected via screenshots alongside accessibility tools like A11y or HTML parsers to extract element structure and bounding boxes. However, these bounding boxes may sometimes be misaligned with the visual rendering due to UI animations or timing inconsistencies. In our work, we primarily rely on datasets curated from Aria-UI and OS-Atlas, which we found to be cleaner and better aligned than alternative data collections.
To… See the full description on the dataset page: https://huggingface.co/datasets/HelloKKMe/grounding_dataset.blip3-grounding-50m
BLIP3-GROUNDING-50M Dataset
Overview
The BLIP3-GROUNDING-50M dataset is designed to enhance the ability of Vision-Language Models (VLMs) to ground semantic concepts in visual features, which is crucial for tasks like object detection, semantic segmentation, and understanding referring expressions (e.g., "the object to the left of the dog"). Traditional datasets often lack the necessary granularity for such tasks, making it challenging for models to accurately localize and… See the full description on the dataset page: https://huggingface.co/datasets/Salesforce/blip3-grounding-50m.highlevel_thinking_with_grounding_annotation_split1000_v3_merged_promptsUI-Grounding-Benchmarks
UI-Grounding-Benchmarks
This is a collection of UI grounding benchmarks:
ScreenSpot
ScreenSpot-V2
ScreenSpot-Pro
OS-World-G
UI-Vision
Thanks for their great work!
This benchmark collection is used in the paper:
FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
🖼️ Project Page: https://showlab.github.io/FocusUI/
🏠 Github Repo: https://github.com/showlab/FocusUI
📝 Paper: https://arxiv.org/pdf/2601.03928
Model Zoo
Model
Backbone
🤗… See the full description on the dataset page: https://huggingface.co/datasets/yyyang/UI-Grounding-Benchmarks.ui-vision-grounding-4MPagent-grounding-dataecg-grounding-250-2500
Dataset Details
This is a randomly split five fold dataset of the ECG-Grouding dataset stratified based on patients (i.e., zero patient overlap between training and test).
The name of the dataset repository is ecg-grounding-250-2500 where 250 refers to the sampling frequency and 2500 denotes 10 seconds.
The code to do the splitting is here.
Any questions or issues, please do not hesitate to reach out to the maintainer of ECG-Bench.
synthetic-grounding-images
Synthetic grounding images
3,162 images generated to extend visual grounding to concrete words that no
photograph dataset covers, for Augustinian BabyLM
(paper, code).
How they were made
Starting from 1,986 concrete words with no image support, an LLM
(claude-sonnet-4-6, temperature 0.8) wrote short scene descriptions
placing as many target words as fit naturally into one scene. Each of the
1,054 resulting descriptions was rendered three times with SDXL-Turbo
(2… See the full description on the dataset page: https://huggingface.co/datasets/augustinian-babylm/synthetic-grounding-images.GroundingCapv2grounding-atlas
grounding-atlas: verifiable-signal pairs
Matched (representation, verifiable-property) pairs for measuring whether a
language model grounds the content of a scientific representation (a SMILES
string, a protein/DNA/RNA sequence, an expression vector, a spectrum, an image)
or merely its name. Each property is either an experimentally measured endpoint
or a closed-form function of the representation, so the representation is the
ground truth and grounding becomes directly… See the full description on the dataset page: https://huggingface.co/datasets/jang1563/grounding-atlas.minecraft-grounding-action-dataseteasyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b
easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b
This dataset was generated from filtered GTA shards with images streamed from ZIP archives.
Generated on: 2025-09-18 05:15:30 UTC
Script: push_easyr1_zip_shards_to_hf.py
Filters directory: /p/project1/synthlaion/awadalla1/gta-grounding-data-filters
JSONL glob: gta_shard_*zip.jsonl
Resize max: 4.0 MP
Prompt format: gta1 (output: coordinates)
Random seed: 42
Deduplicate: False
Debug images: True
System Prompt
You are… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-grounding-gta1-4MP-easy-qwen7b-hard-gta1-7b.GUI_groundingWe share a comprehensive GUI instruction grounding dataset. It contains 2646 desktop screenshot images along with their annotations.
The screenshot images can be found in the folder "images" and the annotation file can be found under "annotations".
The annotation JSON largely follows the coco annotation structure with a few additional keys.
We provide 7 screen element categories:
Button,
Tab,
Text field/area,
Link,
Checkbox,
Radio button,
and List.
fineweb_synth_dense_ocr_with_groundingeasyr1-21k-jedi-grounding-4MP
easyr1-21k-jedi-grounding-4MP
This dataset was generated using the EasyR1 grounding dataset pipeline.
Generation Details
Generated on: 2025-08-27 01:12:14 UTC
Script: push_easyr1_to_hf.py
Data directory: /lustre/fsw/portfolios/nvr/users/aawadalla/LLaMA-Factory/data
Parameters Used
Maximum samples: 21000
Image resize (max megapixels): 4.0 MP
Minimum native image resolution: 0.0 MP
Prompt format: gta1_with_resolution
Output format: coordinates
Random seed: 42… See the full description on the dataset page: https://huggingface.co/datasets/mlfoundations-cua-dev/easyr1-21k-jedi-grounding-4MP.fineweb_synth_dense_ocr_colors_with_grounding
