datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
fruit-vegetable-conceptsconcepts_v2_mergedbroden_conceptsmultimodal-peer-collaboration-samples
Multimodal Peer Collaboration Samples - Embodied Map Task with Two Camera Angles
Two non-experts collaborate to build working circuits under asymmetric information: the instructor has the manual, the student has the components, and synchronized audio and dual-camera video capture how shared understanding emerges.
▶ Watch the interactions · See Expert Instruction samples · Discuss the full collection
Sister collection: Expert Instruction, a teacher and a student in… See the full description on the dataset page: https://huggingface.co/datasets/fluid-concepts/multimodal-peer-collaboration-samples.Multi_Subject_Concepts
Dataset Card for Dataset Name
Dataset Summary
This dataset card aims to be a base template for new datasets. It has been generated using this raw template.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
[More Information Needed]
Dataset Structure
Data Instances
[More Information Needed]
Data Fields
[More Information Needed]
Data Splits
[More Information Needed]
Dataset Creation… See the full description on the dataset page: https://huggingface.co/datasets/ImagenHub/Multi_Subject_Concepts.abstract-concepts-diffusiondb-imagesioai2025-onsite-concepts-hint-descriptionsmultimodal-cultural-conceptsThis dataset encompasses a diverse range of cultural concepts from five different languages and cultural backgrounds: Indonesian, Swahili, Tamil, Turkish, and Chinese.
Specifically, it includes 236 concepts in Chinese, 128 in Indonesian, 202 in Swahili, 178 in Tamil, and 178 in Turkish, with each cultural concept represented by at least two images, totaling 2,235 high-quality images.
Each culture comprises ten primary categories, covering festivals, music, religion and beliefs, animals and… See the full description on the dataset page: https://huggingface.co/datasets/zhili312/multimodal-cultural-concepts.manumoi-conceptsphoto_concepts_dataset_smallconceptsbg_photo_concepts_bucketed_512
bg_photo_concepts_bucketed_512
Title: bg_photo_concepts_bucketed_512
Description: A recaptioned, self contained, bucketed and ready to train with version of https://huggingface.co/datasets/bghira/photo-concept-bucket, exported at 512^2 ish resolution buckets.
I lost the tracking data of which version of Gemini this was captioned with, likely 2.0 flash or 2.5 flash. The captions are on the long and datailed side and sometimes slightly redundant, but overall high quality.… See the full description on the dataset page: https://huggingface.co/datasets/data-archetype/bg_photo_concepts_bucketed_512.vqa-rad-sft-conceptsrobots-human-concepts
Robots — Human Concepts
Synthetic benchmark for evaluating Concept Bottleneck Models (CBMs) under finer-grained, human-annotated concepts. Same underlying robot images and labels as juliannski/robots-true-concepts, but the foot_shape ground-truth concept is replaced by 6 one-hot subtypes that a human annotator would actually see, modelling concept specification mismatch between annotators and the latent labeling rule.
Generated from
This dataset is the exact… See the full description on the dataset page: https://huggingface.co/datasets/juliannski/robots-human-concepts.abstract_concepts
Abstract Concepts Dataset
Contrastive image pairs for abstract concepts (e.g., networking vs isolation, hierarchy vs equality). Each row has two images: one expressing concept A and one expressing concept B, in the same setting. Generated with FLUX.1-dev or NVIDIA Sana, verified with Qwen2.5-VL-32B-Instruct.
How It Is Collected
The collect.py script:
Concept pairs: Each pair has concept_a, concept_b, a scene_template, and distinct prompts for each side. For example… See the full description on the dataset page: https://huggingface.co/datasets/nirmalendu01/abstract_concepts.DreamBooth_Concepts
Dataset Card for "dreambooth"
More Information needed
robots-true-concepts
Robots — True Concepts
Synthetic benchmark for evaluating Concept Bottleneck Models (CBMs). Robot images are generated deterministically with pycairo; binary labels follow a known disjunction-style rule over the 7 ground-truth concepts.
Generated from
This dataset is the exact output of the concept-benchmark Python package — seed and config are pinned for bit-identical reproduction.
# pip install concept-benchmark==0.3.1
from concept_benchmark.robots import… See the full description on the dataset page: https://huggingface.co/datasets/juliannski/robots-true-concepts.conceptsDreamBooth_Concepts_llavaLLaVa captioned datasets from https://huggingface.co/datasets/ImagenHub/DreamBooth_Concepts
ioai2025-onsite-concepts-hint-descriptionsvqa-rad-initial-conceptsvqa-rad-sft-concepts-v2vqa-rad-dino-conceptsioai2025-onsite-concepts-hint-descriptionshero-conceptsconceptsabstract-concepts-imagesvqa-initial-concepts-2abstract-concepts-images-editedvqa-initial-concepts-1
