datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
simple-image-captionstestImageCaptioning_CatalanThe dataset consists of 153,791 images, each accompanied by a description in Catalan. The images have been sourced from two repositories:
"yerevann/coco-karpathy" and "UCSC-VLAA/Recap-COCO-30K." This dataset is ideal for computer vision tasks, as it combines a wide variety of
images with detailed descriptions that can be useful for training machine learning models.
It is freely accessible to everyone, as long as proper credit is given to the original data sources. Thanks
image_caption_regularization
Regularization Image Caption Dataset
Number of Images: 1976
Source
This is a subset of tomg-group-umd/pixelprose, converted to .csv format.
Files
people.csv: 1976 images with captions that contain one of these terms: ['person', 'people', 'man', 'men', 'woman', 'women']
safety-image-captions-1Image_Captioning_and_Attribute_Tagsimage_caption_pairs_for_multimodal
