datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual_captions
Dataset Card for Conceptual Captions
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.google-streetview-images-by-country
Dataset Card for google streetview images by country
⚠️ There are still images that should be deleted, such as those with tags or those that didn't load correctly.
Dataset Structure
folder with the individual countries
images have the creation date and the map name in the file name.
Dataset Card Contact
use the community section
images per country
scin
SCIN Dataset
The SCIN (Skin Condition Image Network) open access dataset aims to supplement publicly available dermatology datasets from health system sources with representative images from internet users. To this end, the SCIN dataset was collected from Google Search users in the United States through a voluntary, consented image donation application. The SCIN dataset is intended for health education and research, and to increase the diversity of dermatology images available for… See the full description on the dataset page: https://huggingface.co/datasets/google/scin.dreambooth
Dataset Card for "dreambooth"
Dataset of the Google paper DreamBooth: Fine Tuning Text-to-Image Diffusion Models for Subject-Driven Generation
The dataset includes 30 subjects of 15 different classes. 9 out of these subjects are live subjects (dogs and cats) and 21 are objects. The dataset contains a variable number of images per subject (4-6). Images of the subjects are usually captured in different conditions, environments and under different angles.
We include a file… See the full description on the dataset page: https://huggingface.co/datasets/google/dreambooth.RSRCC
RSRCC (A Remote Sensing Regional Change Comprehension Benchmark Constructed via Retrieval-Augmented Best-of-N Ranking)
This repository hosts the RSRCC dataset introduced in RSRCC paper.
The dataset is designed for semantic change understanding in remote sensing, pairing multi-temporal image evidence with natural language questions and answers.
🛰️ This work was done by the RSFM (Remote Sensing Foundation Models) team from Google Research.
Official… See the full description on the dataset page: https://huggingface.co/datasets/google/RSRCC.polaris-bench
The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space
Polaris-Bench: Official Evaluation Dataset
Xia Hu1,
Zhenrui Yue1,
Brian Potetz1,
Howard Zhou1,
Leonidas Guibas1,2,
Chun-Ta Lu3,
Zhicheng Wang1
1Google DeepMind 2Stanford University 3Google Research
Overview
As current Multimodal Large Language Models (MLLMs) rapidly saturate canonical visual reasoning benchmarks, a key… See the full description on the dataset page: https://huggingface.co/datasets/google/polaris-bench.google-shopping-general-eval
Marqo Ecommerce Embedding Models
In this work, we introduce the GoogleShopping-1m dataset for evaluation. This dataset comes with the release of our state-of-the-art embedding models for ecommerce products: Marqo-Ecommerce-B and Marqo-Ecommerce-L.
Released Content:
Marqo-Ecommerce-B and Marqo-Ecommerce-L embedding models
GoogleShopping-1m and AmazonProducts-3m for evaluation
Evaluation Code
The benchmarking results show that the… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/google-shopping-general-eval.mathwriting-googlegoogle-landmarkstecci
TECCI: Tricky Edits of Collected and Curated Images
Subsets
Config
Description
Images
Instructions
ggis
TECCI-GGIS (generated automatically using Gemini 3 Pro)
1404
7020 (5 per image)
ircs
TECCI-IRCS (manually written edit instructions)
530
530 (1 per image)
Usage
from datasets import load_dataset
# Load a subset (only "test" split is available)
ds_ggis = load_dataset("google/tecci", "ggis", split="test")
ds_ircs =… See the full description on the dataset page: https://huggingface.co/datasets/google/tecci.crawl-google-image
Crawl Google Image
Crawl Google Image using Malay keywords, total 2046313 rows. Done by https://github.com/kurkurzz
Source code at https://github.com/mesolitica/malaysian-dataset/tree/master/crawl/google-image
OCR-Google_Books
Dataset Card for OCR-Google_Books
A line-to-text dataset for Tibetan OCR.
Dataset Details
Dataset Structure
Features:
id: Image file identifier
label: Text transcription
image: Image of a line of Tibetan text
Splits:
Train: 601,152 samples (37.3M characters)
Eval: 75,136 samples (4.7M characters)
Test: 75,168 samples (4.7M characters)
Uses
Direct Use
Training and evaluation of Tibetan OCR models
Multi-script OCR… See the full description on the dataset page: https://huggingface.co/datasets/openpecha/OCR-Google_Books.witWikipedia-based Image Text (WIT) Dataset is a large multimodal multilingual dataset.
WIT is composed of a curated set of 37.6 million entity rich image-text examples with 11.5 million unique images across 108 Wikipedia languages.
Its size enables WIT to be used as a pretraining dataset for multimodal machine learning models.receipts-google-ocrGooglePlayStoreApps
📱 Google Play Store 2020 — What Made an App Go Viral?
A data-driven EDA of 9,436 apps released in 2020, exploring the patterns behind viral success on the Play Store.
🎬 Video Walkthrough
Can't see the video? Click here to watch
🔬 Research Question
"What made an app released in 2020 go viral?"
Defined as reaching 1M+ installs by June 2021 — within 6–18 months of launch.
2020 was chosen deliberately: the COVID-19 pandemic drove unprecedented mobile app… See the full description on the dataset page: https://huggingface.co/datasets/yonilev/GooglePlayStoreApps.DOCCI-Critiquecrawl-google-image-malaysian-vehicle
Crawl Google Image Malaysian Car
Crawl Google Image using Malaysian car keywords.
Source code at https://github.com/mesolitica/malaysian-dataset/tree/master/crawl/google-image
invoices-google-ocrgoogle-illustrations-full
Google Illustrations: Complete 4K Multi-Layer Archive (2,052 Avatars)
A comprehensive, uncompressed 4096x4096 archival dataset of the complete Google Account Illustrations library, featuring all 59 artist collections, 2,052 unique scenes, decomposed layer assets, color presets, and semantic metadata.
Legal Disclaimer and Copyright Notice
PLEASE READ CAREFULLY:
This repository is an independent archival and educational compilation provided strictly for… See the full description on the dataset page: https://huggingface.co/datasets/Anson10124/google-illustrations-full.google-landmarks-v2-minigoogle_maps
Dataset Card for "google_maps"
More Information needed
google_landmarks_photos
Dataset Card for "google_landmarks_photos"
More Information needed
google_landmark_v2_qabooster_t1_Robot_of_Google_DeepMind_mujoco
Analysis and Visualization of Google DeepMind's Booster T1 MuJoCo Robot
This repository contains a detailed analysis, visualization, and educational breakdown of the Booster T1 humanoid robot model, originally provided by Google DeepMind's mujoco_menagerie. It serves as a practical dataset for anyone interested in robotics, physics simulation, and biomechanics using the MuJoCo physics engine.
This file is designed to provide clear, detailed documentation for users.
You can find the… See the full description on the dataset page: https://huggingface.co/datasets/MartialTerran/booster_t1_Robot_of_Google_DeepMind_mujoco.google_earth_fixedgoogle-open-images-hair-style-dataset
Google Open Images — Hair Style Dataset
🇺🇸 English | 🇰🇷 한국어
Overview
This dataset is a curated custom subset of the Google Open Images V7 dataset, specifically filtered to include images of humans with various hair styles.It is intended for use in computer vision research and applications such as hair style classification, person detection, and instance segmentation.
Split
Purpose
train
Model training
validation
Model evaluation / hyperparameter tuning… See the full description on the dataset page: https://huggingface.co/datasets/hwany79/google-open-images-hair-style-dataset.google-ads-benchmark-2026
Note on checksums. This README.md carries the YAML dataset-card header required by the Hugging Face hub, so its SHA-256 differs from the entry in checksums.txt; that entry refers to the canonical README in the GitHub mirror. All data files are byte-identical across mirrors. Load any table with load_dataset("ivitskiy/google-ads-benchmark-2026", "<config_name>").
Ivitskiy Ads Lab: Google Ads Panel & Benchmark Compilation 2026 (Open Research Dataset)
Two things in one package.… See the full description on the dataset page: https://huggingface.co/datasets/ivitskiy/google-ads-benchmark-2026.google-shopping-general-eval-100k
Marqo Ecommerce Embedding Models
In this work, we introduce the GoogleShopping-1m dataset for evaluation. This dataset comes with the release of our state-of-the-art embedding models for ecommerce products: Marqo-Ecommerce-B and Marqo-Ecommerce-L.
Released Content:
Marqo-Ecommerce-B and Marqo-Ecommerce-L embedding models
GoogleShopping-1m and AmazonProducts-3m for evaluation
Evaluation Code
The benchmarking results show that the… See the full description on the dataset page: https://huggingface.co/datasets/Marqo/google-shopping-general-eval-100k.crawl-google-image-malaysia-location
Crawl Google Image Malaysia Location
Crawl Google Image using Malaysia location keywords.
Source code at https://github.com/mesolitica/malaysian-dataset/tree/master/crawl/google-image
How to use using streaming
dataset = load_dataset('malaysia-ai/crawl-google-image-malaysia-location', streaming=True, split='train')
for row in dataset:
break
print(row['image'])
<PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=1200x1200 at 0x7F50536E7550>
wmt24pp-images
WMT24++ Source URLs & Images
This repository contains the source URLs and full-page document screenshots for each document in the data from
WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects.
These images preserve the original document structure of the translation segments with any embedded images, and may be used for multimodal translation or language understanding.
If you are interested in the human translations and post-edit data, please see here.If… See the full description on the dataset page: https://huggingface.co/datasets/google/wmt24pp-images.
