datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
amazon-berkeley-objects
Amazon Berkeley Objects (ABO)
A Hugging Face packaging of the Amazon Berkeley Objects (ABO) dataset. The
data content is the official CC BY 4.0 release from
https://amazon-berkeley-objects.s3.amazonaws.com/index.html. This mirror
changes only the packaging: files are grouped into typed Parquet shards, and
every original media file is preserved byte-for-byte and never transcoded.
Images use the datasets Image() feature, 3D product models use the native
Mesh() feature (original… See the full description on the dataset page: https://huggingface.co/datasets/suvadityamuk/amazon-berkeley-objects.omega-multimodal
OMEGA Labs Bittensor Subnet: Multimodal Dataset for AGI Research
Introduction
The OMEGA Labs Bittensor Subnet Dataset is a groundbreaking resource for accelerating Artificial General Intelligence (AGI) research and development. This dataset, powered by the Bittensor decentralized network, aims to be the world's largest multimodal dataset, capturing the vast landscape of human knowledge and creation.
With over 1 million hours of footage and 30 million+ 2-minute… See the full description on the dataset page: https://huggingface.co/datasets/omegalabsinc/omega-multimodal.openbrush
OpenBrush-75K
A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research.
Dataset Description
OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush.openbrush-75k
OpenBrush-75K
A curated dataset of 75,313 public domain artworks with rich, structured VLM-generated captions designed for training image generation models, fine-tuning vision-language models, and art analysis research.
Dataset Description
OpenBrush-75K contains high-quality reproductions of paintings from the Western art canon, spanning from the Renaissance to the early 20th century. Each image is paired with a detailed structured caption generated by a vision-language… See the full description on the dataset page: https://huggingface.co/datasets/Trever896/openbrush-75k.openbrush-landscapes
OpenBrush Landscapes
Every landscape painting from OpenBrush-75K — across all artists, movements, and centuries. Largest single-genre subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,612 you actually want.
Why this subset
Every landscape across the parent dataset's full range — Romantic wildernesses, Impressionist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-landscapes.ODELIA-Challenge-2025
ODELIA Challenge Dataset
This dataset is part of the ODELIA project, a European Horizon initiative focused on developing privacy-preserving, AI-driven diagnostic tools using swarm learning.
The dataset provided here represents a curated subset of data from the broader ODELIA consortium. It is designed to facilitate the development, benchmarking, and validation of AI algorithms that can operate effectively across a range of heterogeneous clinical settings.
The dataset contains… See the full description on the dataset page: https://huggingface.co/datasets/ODELIA-AI/ODELIA-Challenge-2025.openbrush-impressionism
OpenBrush Impressionism
Every Impressionist work from OpenBrush-75K — the largest movement subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 12,798 you actually want.
Why this subset
Broad-coverage subset for training on the Impressionist visual language: broken brushwork, light-on-color theory, plein-air staging, atmospheric… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionism.polyvore-outfits
Polyvore Outfits (Refactored Version)
This repository provides a refactored version of the Polyvore Outfits dataset, originally introduced in the paper "Learning Type-Aware Embeddings for Fashion Compatibility" by Mariya I. Vasileva et al.
📌 Overview
The goal of this refactoring is to improve usability and developer experience. While the core data remains identical to the original, the file structure and JSON schemas have been standardized to make it easier to load and… See the full description on the dataset page: https://huggingface.co/datasets/owj0421/polyvore-outfits.foodv2
Turkish Food & Nutrition Vision Dataset (Food v2)
High-resolution, human-POV Turkish food and packaged snack dataset with canonical nutrition macros and portion sizes.
📊 Dataset Summary
Repository: ozertuu/foodv2
Canonical Foods: 29,840
Images Total: 89,520
Average Images / Food: ~3.0
Image Sources: Yemeksepeti / DeliveryHero, GetirYemek, Migros Sanalmarket, Getir, Nefis Yemek Tarifleri, OpenFoodFacts.
🍽️ Features & Schema
image: PIL Image /… See the full description on the dataset page: https://huggingface.co/datasets/ozertuu/foodv2.openbrush-baroque
OpenBrush Baroque
Baroque works from OpenBrush-75K (~1600–1750) — chiaroscuro, religious painting, dramatic light.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,240 you actually want.
Why this subset
The canonical Baroque visual language — Caravaggio, Rembrandt, Vermeer, Velázquez, Rubens. Useful for models learning dramatic… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-baroque.Guardian-FailCoT-OOD-datasets
Guardian FailCoT — Out-of-Distribution Real-Robot Benchmarks
This repository bundles the three real-world failure-detection benchmarks used to evaluate the Guardian vision-language model in the paper Scaling Cross-Environment Failure Reasoning Data for Vision-Language Robotic Manipulation (Pacaud et al., 2026):
UR5-Fail — our newly collected three-view real-robot benchmark.
RoboFail — single-view real-robot manipulation failure benchmark from Liu et al. (CoRL 2023).
RoboVQA —… See the full description on the dataset page: https://huggingface.co/datasets/paulpacaud/Guardian-FailCoT-OOD-datasets.dental-implant-surgery-sample
Dental Implant Surgery — Multimodal Annotated Video (Sample Case)
A public sample from one complete All-on-4 full-arch mandibular dental implant
surgery: two synchronised camera angles, the operating surgeon narrating while
he works, and six layers of structured clinical annotation (L0–L5) tied frame by
frame to what he said.
This is a showcase slice, not the whole case. What is here is enough to judge
the structure, the annotation quality and the honesty of the documentation.… See the full description on the dataset page: https://huggingface.co/datasets/OralSurgery/dental-implant-surgery-sample.openbrush-impressionist-landscapes
OpenBrush Impressionist Landscapes
Cross-cut subset: Impressionist landscape paintings from OpenBrush-75K. The most-targeted style+genre combination for Impressionist landscape LoRA training.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 4,308 you actually want.
Why this subset
The intersection of the largest movement… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-impressionist-landscapes.Optimized_Video_Facial_Landmarks
Dataset Card for 478-Point Normalized 3D Facial Landmark Dataset
Dataset Description
This dataset provides pre-extracted, normalized 3D facial landmark features derived from the Video Emotion dataset. It is optimized for efficient training of emotion recognition and facial analysis models, bypassing the need to process large raw video files.
License: The extracted feature data in this Parquet file is licensed under Apache 2.0. Note that the original source video files may… See the full description on the dataset page: https://huggingface.co/datasets/PSewmuthu/Optimized_Video_Facial_Landmarks.AU-OPG
AU-OPG: Panoramic Dental Radiographs With Oriented Tooth-Level Annotations
AU-OPG (Ajman University Orthopantomography) is a dataset of 901 panoramic dental radiographs annotated for:
oriented tooth detection;
radiographic diagnosis; and
radiographic-evidence-based treatment planning.
The release contains 7,006 tooth-level annotations. Each annotated tooth has a tooth-aligned oriented bounding box, one diagnostic condition, and one corresponding treatment label. The predefined… See the full description on the dataset page: https://huggingface.co/datasets/YSFF/AU-OPG.openbrush-religious-art
OpenBrush Religious Art
Religious paintings from OpenBrush-75K — saints, biblical scenes, devotional works.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,119 you actually want.
Why this subset
A coherent visual genre: religious narrative painting from medieval through early modern. Heavy on Renaissance and Baroque eras. Common… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-religious-art.openart-items-artifacts
OpenArt — Items & Artifacts
openart-items-artifacts is the items artifacts subject collection of the OpenArt family
of open, public-domain art datasets: 25,750 works (11,317 paintings/illustrations · 14,216
photographed objects · 217 unclassified), each paired with a structured VLM caption plus
medium, attribution and inscription metadata.
Human-made objects and the decorative arts — vessels, tools, arms and armor, textiles, furniture and ornament — both as physical artifacts… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openart-items-artifacts.nationalmuseet-open-images
Nationalmuseet Open Images
This dataset is an independently harvested research dataset from Nationalmuseet Samlinger Online.
It contains metadata and optionally WebDataset image shards for Nationalmuseet asset records whose
rights.license is one of:
Public Domain
CC-BY
No known rights
Public Domain and CC-BY are the strict open-license subset. No known rights is kept as a
separate license bucket because Nationalmuseet says this label means that, to their best assessment,
the… See the full description on the dataset page: https://huggingface.co/datasets/V4ldeLund/nationalmuseet-open-images.openbrush-anonymous-masters
OpenBrush Anonymous Masters
Unattributed works from OpenBrush-75K — anonymous old masters across centuries and styles. Useful for broad-style training without artist-specific bias.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 41,914 you actually want.
Why this subset
37% of the parent dataset is unattributed — a substantial… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-anonymous-masters.openbrush-portraits
OpenBrush Portraits
Every portrait painting from OpenBrush-75K — across all artists, movements, and centuries.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 13,059 you actually want.
Why this subset
Portraits across the full historical range — Renaissance bust portraits, Baroque chiaroscuro, Rococo society, Romantic, Realist… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-portraits.openbrush-renaissance
OpenBrush Renaissance
Renaissance works from OpenBrush-75K, combining Northern, Early, High, and Mannerism Late Renaissance into one period subset.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 6,565 you actually want.
Why this subset
Combined Renaissance period subset spanning ~1300–1600. Heavy on religious painting, portraits… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-renaissance.openbrush-ukiyo-e
OpenBrush Ukiyo-e
Japanese Ukiyo-e woodblock prints from OpenBrush-75K — the only non-Western style in the parent dataset, separated here for trainers who want a dedicated Japanese woodblock corpus without downloading 75K Western paintings to filter.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 1,167 you actually want.
Why this… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-ukiyo-e.openm3chest-labels
OpenM3Chest Labels (OM3C)
JSON label files and Series UIDs from the OpenM3Chest dataset, prepared for fine-tuning medical vision-language models such as MedGemma.
Raw imaging data (DICOM) can be downloaded from IDC (Imaging Data Commons) using the Series Instance UIDs provided in unique_keys.txt.
Dataset Summary
OpenM3Chest is a medical multimodal multitask dataset for diagnosing chest abnormalities with a focus on lung cancer screening. The original raw data comes… See the full description on the dataset page: https://huggingface.co/datasets/UngLong/openm3chest-labels.sd2-staged-foreign-objects
SD2 staged-laboratory foreign-object frames (heydonto)
97 frames · 163 frame-level annotation rows · 8 staged laboratory takes · CC BY 4.0
This dataset discloses and carries the SD2 own-footage portion of the training data of the ORena SAVE FOCUS challenge entry's FRAME-track component: 163 of that component's 58,086 pooled training rows. The other sources of that corpus are not part of this dataset.
What is in it
frames/ — 97 JPEG frames (filename = first 16 hex… See the full description on the dataset page: https://huggingface.co/datasets/HeyDonto/sd2-staged-foreign-objects.kamari-safe-open-v0
Kámárí-Safe Open v0 (benchmark)
A frozen, leakage-free benchmark for African-tailored age verification. It holds manifests and
split tables, not raw images (paths, hashes, labels, skin band, quality). Use it to measure age
accuracy and, more importantly, child-safety.
Headline metric
Minor-Pass-Through Rate (MPTR) is the headline: the fraction of true minors a model passes as
adults, reported overall, at 21, and for dark + brown skin. Report MPTR alongside MAE; a… See the full description on the dataset page: https://huggingface.co/datasets/Shinzmann/kamari-safe-open-v0.OpenHotelsSample
OpenHotels Representative Sample
This repository contains a representative sample of OpenHotels for review and inspection. It mirrors the full OpenHotels release structure: image files are stored in tar shards under shards/, and metadata files describe the gallery, non-object query images, object-centric query images, and hotel classes.
The sample is intended for data-quality inspection, not benchmark reporting. Use the full OpenHotels dataset for final evaluation.… See the full description on the dataset page: https://huggingface.co/datasets/imagingforgood/OpenHotelsSample.osm-europe-1k
OSM-Europe-1k
A 1,000-image street-level geolocation benchmark for Europe, sampled from
the OpenStreetView-5M (OSV-5M)
test split. Intended as a contamination-free, openly-licensed reference set
for evaluating image→GPS models. Companion benchmark to a master's thesis at
FH JOANNEUM (Florian Leber, 2026).
What's in this repo:
The 1,000 OSM/Mapillary image bytes are bundled directly under
images/ (66 MB) — re-distribution is allowed by the upstream
CC-BY-SA 4.0 license, with… See the full description on the dataset page: https://huggingface.co/datasets/lebfla11/osm-europe-1k.openbrush-van-gogh
OpenBrush Van Gogh
All Vincent van Gogh works from OpenBrush-75K, with structured VLM captions.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 1,889 you actually want.
Why this subset
Van Gogh's catalog spans his Realism period (Dutch landscapes, peasant scenes) through his Post-Impressionist breakthroughs (Arles, Saint-Rémy… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-van-gogh.openbrush-rembrandt
OpenBrush Rembrandt
All Rembrandt works from OpenBrush-75K — paintings, etchings, and sketches.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 776 you actually want.
Why this subset
The defining body of work for Baroque chiaroscuro and dramatic light. Includes religious scenes, portraits, self-portraits, and biblical narratives.… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-rembrandt.openbrush-monet
OpenBrush Monet
All Claude Monet works from OpenBrush-75K, with structured VLM captions.
Curated subset of jaddai/openbrush. Same CC0 license, same caption schema, same VLM (Qwen3-VL-30B-A3B). This subset exists so you don't have to download 75,313 images to get to the 1,334 you actually want.
Why this subset
Monet's canonical Impressionist body of work — haystacks, water lilies, Rouen cathedrals, Giverny garden scenes — captioned with detail on light quality… See the full description on the dataset page: https://huggingface.co/datasets/jaddai/openbrush-monet.
