datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
RefCOCOg
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of RefCOCOg. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@inproceedings{kazemzadeh-etal-2014-referitgame,
title = "{R}efer{I}t{G}ame: Referring to Objects in… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/RefCOCOg.refcocog
Dataset Card for "refcocog"
More Information needed
refcocog_valrefCOCOg_9k_840_maskrefCOCOg_9k_840
Seg-Zero Dataset
This repository contains the training data for the Seg-Zero framework, as presented in the paper Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.
Seg-Zero is a novel framework that demonstrates remarkable generalizability and derives explicit chain-of-thought reasoning for image segmentation tasks through cognitive reinforcement. This dataset facilitates the training of such a system, where a reasoning model interprets user intentions and… See the full description on the dataset page: https://huggingface.co/datasets/Ricky06662/refCOCOg_9k_840.refcocogRefCOCOg_rej
(CVPR 2026) GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
Dataset Description
GroundingME is a benchmark for evaluating visual grounding capabilities in Multimodal Large Language Models (MLLMs), systematically challenging models across four critical dimensions: Discriminative, Spatial, Limited, and Rejection. Our evaluation of 25 state-of-the-art MLLMs reveals that most models score 0% on rejection tasks… See the full description on the dataset page: https://huggingface.co/datasets/lirang04/RefCOCOg_rej.refCOCOg_9k_840_sam2_parquet
refCOCOg 9k 840 with SAM2 Masks
This dataset is derived from refCOCOg_9k_840 and adds a mask column generated offline with SAM2.
Each sample contains:
id: sample identifier
problem: referring expression / query
solution: original box and point annotations
image: RGB image stored as Hugging Face image bytes
img_height: original metadata height
img_width: original metadata width
mask: SAM2-generated binary mask stored as PNG bytes
The mask column is a pseudo-label generated from… See the full description on the dataset page: https://huggingface.co/datasets/wanwan1111/refCOCOg_9k_840_sam2_parquet.refcocog_testRefCOCOg-RexThinker-20krefCOCOg_2k_840
Seg-Zero Dataset: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
This repository hosts the training dataset introduced in the paper Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.
Abstract
Traditional methods for reasoning segmentation rely on supervised fine-tuning with categorical labels and simple descriptions, limiting its out-of-domain generalization and lacking explicit reasoning processes. To address these limitations, we… See the full description on the dataset page: https://huggingface.co/datasets/Ricky06662/refCOCOg_2k_840.RefCOCOg_valRefCOCOg_recjxu124_refcocogrefCOCOg_in_tarRefCOCOg_testrefcocog_polygonsrefcocog_val_500refcocog-coco2017
refcocog with COCO 2017 Image Paths
This dataset is a version of the original refcocog dataset that uses COCO 2017 image paths instead of COCO 2014.
Changes from Original
Image paths updated from COCO 2014 format to COCO 2017 format
Images loaded from COCO 2017 directory structure
All other annotations remain unchanged
Usage
from datasets import load_dataset
dataset = load_dataset("jhkwak-bp/refcocog-coco2017")
Citation
Please… See the full description on the dataset page: https://huggingface.co/datasets/jhkwak-bp/refcocog-coco2017.RefCOCOgvlmevalkit-refcocogrefcocog_test_allrefCOCOg_9k_840refcocog_object_detectionrefcocog
Dataset Card for "refcocog"
More Information needed
Som_bench_refcocog_refseg
Som_bench_refcocog_refseg Dataset
This dataset is a processed version of the RefCOCOg dataset and is intended to be used as part of a benchmark, specifically mirroring the data splits and format used in the Set-of-Mark (SoM) benchmark. It is designed for evaluating visual grounding and related tasks.
Original Dataset:
This dataset is based on the RefCOCOg dataset. Please refer to the original RefCOCOg dataset for its terms of use and licensing.
Benchmark Reference:
This dataset… See the full description on the dataset page: https://huggingface.co/datasets/Rajarshi-Roy-research/Som_bench_refcocog_refseg.refcocog_val_allrefCOCOg_100refcocog-speechjxu124_refcocog_debug
