datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
blind-people-scene-descriptions
Merged Navigation-Focused Image Caption Dataset
This dataset is a combination and filtered version of two publicly available image captioning datasets, specifically curated to focus on images and captions relevant to navigation and scene understanding.
Source Datasets
This dataset is derived from the following two sources:
COCO Captions (jxie/coco_captions)
Original Hugging Face Hub ID: jxie/coco_captions
Link: https://huggingface.co/datasets/jxie/coco_captions
Original… See the full description on the dataset page: https://huggingface.co/datasets/mlevytskyi/blind-people-scene-descriptions.adaption-street-scene-descriptions
This dataset is a remastered version of
Reubencf/streetview-global
prepared using Adaption's Adaptive Data platform.
Street Scene Descriptions (Adaption)
10,117 globally-sampled street-level images paired with detailed textual
scene descriptions and rich structured metadata (setting, weather, time of
day, road type, geolocation). Each entry features a natural-language
caption describing the scene along with Adaption-sharpened
enhanced_prompt and enhanced_completion columns… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/adaption-street-scene-descriptions.street-scene-descriptions
This dataset is a remastered version of this dataset prepared using Adaption's Adaptive Data platform.
street_scene_descriptions
This dataset contains over 10,000 street-level image samples accompanied by detailed textual scene descriptions and rich metadata including geolocation, time, weather, and infrastructure details. Each entry features a natural language caption describing urban and residential environments captured from vehicle perspectives under varying lighting and… See the full description on the dataset page: https://huggingface.co/datasets/Reubencf/street-scene-descriptions.vip-scene-description
