datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
CountBenchQAThis dataset was introduced in PaliGemma for evaluating counting in vision language models. This version only includes 491 images from the original CountBench dataset, since some of the original URLs can no longer be accessed.
Original Description
CountBench: We introduce a new object counting benchmark called CountBench,
automatically curated (and manually verified) from the publicly available
LAION-400M image-text dataset. CountBench contains a total of 540 images
containing… See the full description on the dataset page: https://huggingface.co/datasets/vikhyatk/CountBenchQA.CountQA
Dataset Summary
CountQA is the new benchmark designed to stress-test the Achilles' heel of even the most advanced Multimodal Large Language Models (MLLMs): object counting. While modern AI demonstrates stunning visual fluency, it often fails at this fundamental cognitive skill, a critical blind spot limiting its real-world reliability.
This dataset directly confronts that weakness with over 1,500 challenging question-answer pairs built on real-world images, hand-captured to feature… See the full description on the dataset page: https://huggingface.co/datasets/Jayant-Sravan/CountQA.pixmo-count
PixMo-Count
PixMo-Count is a dataset of images paired with objects and their point locations in the image.
It was built by running the Detic object detector on web images, and then filtering the data
to improve accuracy and diversity. The val and test sets are human-verified and only contain counts from 2 to 10.
PixMo-Count is a part of the PixMo dataset collection and was used to
augment the pointing capabilities of the Molmo family of models
Quick links:
📃 Paper
🎥 Blog with… See the full description on the dataset page: https://huggingface.co/datasets/allenai/pixmo-count.pixmo-point-count-concat_0-20Multi-Hop-Objects-Countingpixmo-point-count-gen-undcountry211
Dataset Card for Country211
The Country 211 Dataset from OpenAI.
This dataset was built by filtering the images from the YFCC100m dataset that have GPS coordinate corresponding to a ISO-3166 country code. The dataset is balanced by sampling 150 train images, 50 validation images, and 100 test images images for each country.
countbench
Dataset Card for "countbench"
This dataset was introduced in the paper Teaching CLIP to Count to Ten.
sd-prompt-image-in-the-wild-counterfeitworldcuisines_format_sea_country_only_with_metadatacounterfactual-physicspixmo-count-filtered-imgContainedclevr_count_70kThis dataset is borrowed from clevr_cogen_a_train
ai2thor-counting-largeVisual-Counterfact
Visual CounterFact: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfactuals
This dataset is part of the work "Pixels Versus Priors: Controlling Knowledge Priors in Vision-Language Models through Visual Counterfacts".📖 Read the Paper💾 GitHub Repository
Overview
Visual CounterFact is a novel dataset designed to investigate how Multimodal Large Language Models (MLLMs) balance memorized world knowledge priors (e.g., "strawberries are red")… See the full description on the dataset page: https://huggingface.co/datasets/mgolov/Visual-Counterfact.country211
Dataset Card for "country211"
More Information needed
shanghaitech-crowd-countingeuropean-countries-classifierscannet_countingcoco-counterfactual-conflict
COCO-Counterfactual Conflict
Image-text conflict dataset built from
Intel/COCO-Counterfactuals.
Each COCO-Counterfactuals example is a minimal pair of captions differing by a single noun
subject, with a matching image for each. We keep the truthful image (image_0) and its
caption as original_caption, and use the counterfactual caption as conflicting_caption.
The swapped noun is extracted automatically (image_bias = true noun, text_bias = altered
noun); the question and… See the full description on the dataset page: https://huggingface.co/datasets/multilingual-vlm-conflict/coco-counterfactual-conflict.pixmo-count-imagesCountHallu-Dataset-SimObject
CountHalluSet — SimObject
Rendered dataset from Counting Hallucinations in Diffusion Models
(arXiv:2510.13080). Part of CountHalluSet, a suite with well-defined counting
criteria used to measure counting hallucination — a diffusion model generating
the wrong number of instances, even for patterns absent from its training data.
What's inside
256×256 RGB rendered images of everyday objects, each labelled with the
per-class instance count over three object classes.… See the full description on the dataset page: https://huggingface.co/datasets/ShyFoo/CountHallu-Dataset-SimObject.countqa_lite
gsarch/countqa_lite
A deterministic lite evaluation subset of Jayant-Sravan/CountQA.
Source revision: f92cc6fe46542c61e2916e3d2ae9a911e2216b1a
Source split: test
Sampling seed: 43
Output rows: 500
Schema: unchanged from the upstream dataset
CountQA is sampled at the QA-pair level. Each output row retains the original schema and contains one-element questions and answers lists, so lmms-eval's existing countqa_process_docs produces exactly 500 prompts.
Generated by… See the full description on the dataset page: https://huggingface.co/datasets/gsarch/countqa_lite.single-plot-count-3kcounteranimalworldcuisines_format_sea_country_only_3ro_sft_pixmo_count
Dataset Description
PixmoCount is a dataset of images paired with number of objects in the image.
Here we provide the Romanian translation of the PixmoCount dataset, translated with Seed-X-PPO.
This dataset is part of the instruction finetune protocol for Romanian VLMs proposed in "Înțelegi românește?" A Recipe for Romanian Vision-Language Models (Masala et al., 2026).
Citation
@inproceedings{deitke2025molmo,
title={Molmo and pixmo: Open weights and open data… See the full description on the dataset page: https://huggingface.co/datasets/surogate/ro_sft_pixmo_count.Task2_Chicken_Counting_Train2sem-seg-country-safety-binscountry-flags-dataset
World Country Flags Dataset
This dataset contains flag images from sovereign states with their country names as labels.
Dataset Description
A comprehensive collection of flag images for all sovereign nations, organized for machine learning tasks.
Dataset Summary
Total Images: 195 country flags
Format: PNG (640x427 pixels)
Task: Image classification
Language: English country names
Dataset Structure
Each example contains:
image: The flag image in… See the full description on the dataset page: https://huggingface.co/datasets/Aniket96/country-flags-dataset.
