datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
galaxy-descriptions
Galaxy Descriptions
Project Page | Code
This dataset provides galaxy cutout images, natural-language descriptions, text embeddings, and image embeddings for galaxies drawn from multiple imaging surveys (specifically Legacy DR10 and HSC PDR3 Wide).
Each row corresponds to a single galaxy and contains:
A preprocessed RGB galaxy image
A caption generated by gpt-4.1-mini
A single-sentence summary of the caption generated by gpt-4.1-nano
Text embeddings for the caption and summary… See the full description on the dataset page: https://huggingface.co/datasets/astronolan/galaxy-descriptions.amazon-product-descriptions-vlm
Amazon Multimodal Product dataset
This is a modfied and slim verison of bprateek/amazon_product_description helpful to get started training multimodal LLMs.
The description field was generated used Gemini Flash.
manipulation-init-frame-descriptions
Manipulation init-frame / description pairs
1000 (init frame image, task description) pairs randomly sampled
(seed=42) from the manipulation task family of
nvidia/PhysicalAI-WorldModel-Synthetic-Embodied-Robot-Scenes.
Each row is the first frame of a simulated manipulation clip paired with its
task description text, drawn from the following generators within the
manipulation task family: DreamZero, MimicGen (AgiBot G1 / Fourier GR-1 /
Galbot G1), and Simulario.
Columns:… See the full description on the dataset page: https://huggingface.co/datasets/khang123452/manipulation-init-frame-descriptions.blind-people-scene-descriptions
Merged Navigation-Focused Image Caption Dataset
This dataset is a combination and filtered version of two publicly available image captioning datasets, specifically curated to focus on images and captions relevant to navigation and scene understanding.
Source Datasets
This dataset is derived from the following two sources:
COCO Captions (jxie/coco_captions)
Original Hugging Face Hub ID: jxie/coco_captions
Link: https://huggingface.co/datasets/jxie/coco_captions
Original… See the full description on the dataset page: https://huggingface.co/datasets/mlevytskyi/blind-people-scene-descriptions.HumanML3D-500ms-FPP-descriptions-CoTs-1
HumanML3D 500ms First person perspective descriptions for CoTs
Introduction
This repository contains files of the Mr. Ri's and Ms. Tique's HumanML3D human motion dataset,
but also descriptions of the movements in first person perspective in 0.5 second time windows.
The descriptions were created synthetically with use of a multimodal LLM and are in json format. They can be found in comics_and_descriptions folder.
The dataset also contains motion capture data and… See the full description on the dataset page: https://huggingface.co/datasets/Wojtekb30/HumanML3D-500ms-FPP-descriptions-CoTs-1.mtg-scryfall-unique-artwork-20240809-with-card-art-descriptions-and-images-with-embeddingsmtg-scryfall-cropped-art-with-descriptionsConsumer-Product-Descriptions-50kmtg-scryfall-unique-artwork-20240809-with-card-art-descriptions-and-imagescharacters_descriptionsGUI-Dense-Descriptions
GUI Screenshots - Dense descrptions Dataset
RawDet-7-Object-Descriptions
RAWDet-7: object-description track
This is the 500-image object-description track from RAWDet-7. It is object-level description with set-of-marks, not ordinary whole-image captioning: each annotated object is identified by a numbered black square with a white outline and a colored number, and receives its own detailed caption.
The release contains the exact 500 held-out images used by the paper, their corresponding full-precision RAW files, cleaned detection JSON, all 17 marked… See the full description on the dataset page: https://huggingface.co/datasets/shashankskagnihotri/RawDet-7-Object-Descriptions.adaption-street-scene-descriptions
This dataset is a remastered version of
Reubencf/streetview-global
prepared using Adaption's Adaptive Data platform.
Street Scene Descriptions (Adaption)
10,117 globally-sampled street-level images paired with detailed textual
scene descriptions and rich structured metadata (setting, weather, time of
day, road type, geolocation). Each entry features a natural-language
caption describing the scene along with Adaption-sharpened
enhanced_prompt and enhanced_completion columns… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-street-scene-descriptions.ioai2025-onsite-concepts-hint-descriptionsbalenciaga_short_descriptions
Dataset Card for "balenciaga_short_descriptions"
More Information needed
ioai-chameleon-hint_descriptionshierarchical-descriptions-sample
Archaia — Hierarchical Artifact Descriptions
A multimodal dataset of excavated archaeological artifacts, each paired with:
Up to 3 photographs of the artifact
Original catalog description written by archaeologists at excavation time
5-level AI-generated hierarchical descriptions produced by GPT-4o-mini from the photographs and structured metadata
Full structured metadata (Munsell color, dimensions, material, trench coordinates, date range, etc.)
Artifacts come from multiple… See the full description on the dataset page: https://huggingface.co/datasets/archaia/hierarchical-descriptions-sample.chanel_short_descriptions
Dataset Card for "chanel_short_descriptions"
More Information needed
floorplan-descriptionsmtg-scryfall-unique-artwork-20240809-with-card-art-descriptionsprocess_descriptions_for_modeling
Dataset Card for Business Process Descriptions and Images
This dataset contains pairs of business process descriptions (both normal and enhanced versions) and corresponding image file paths. It is intended for tasks related to understanding and potentially visualizing business processes.
Dataset Details
Dataset Description
This dataset comprises textual descriptions of various business processes alongside paths to related images (presumably process models… See the full description on the dataset page: https://huggingface.co/datasets/bis-aifb-kit/process_descriptions_for_modeling.video_clip_descriptionsamazon-product-descriptions-vlmAnimeGirlz-with-descriptionspokemon-with-pokedex-descriptions
Dataset Card for "pokemon-with-pokedex-descriptions"
More Information needed
image_descriptions_cleanedchanel_long_descriptions
Dataset Card for "chanel_long_descriptions"
More Information needed
ioai2025-onsite-concepts-hint-descriptionsIT-Running-Shoes-Descriptions
Italian Running Shoes Dataset (Multimodal)
This dataset contains a collection of running shoe products specifically curated from the Italian market. It is designed for tasks such as Computer Vision (Image Classification), Natural Language Processing (NLP) in Italian, and E-commerce recommendation systems.
📂 Dataset Structure
The dataset consists of a central metadata file and a folder containing product images.
images/: Directory containing .jpg images of the shoes.… See the full description on the dataset page: https://huggingface.co/datasets/mennox/IT-Running-Shoes-Descriptions.Consumer-Product-Descriptions-50k-processed
