datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
conceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file.
We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs.
length_captchas
Length Captchas Dataset
Converted dataset from twitter_ai label tool.
dreamlip_long_captions
Dataset Card for DreamLIP-30M
Dataset Summary
DreamLIP-Long-Captions is a dataset consisting of ~30M image annotations, i.e. detailed long captions. In contrast with the curated style of other synthetic image caption annotations, DreamLIP-30M utilizes pre-trained Multi-modality Large Language Model to obtain detailed descriptions with an average length of 247. More precisely, the detailed descriptions are generated by asking the ShareGPT4V/InstructBLIP/LLava1.5 the… See the full description on the dataset page: https://huggingface.co/datasets/qidouxiong619/dreamlip_long_captions.character-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes.
Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by Claude Opus. Here is an example:
{… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.simple-image-captionstestGPT4V-captions-from-LVIS-typography
GPT4V-captions-from-LVIS-typography
by: Peter Bevan, 21 March 2023
This dataset is a typography subset of 220k-GPT4Vision-captions-from-LIVIS.
This dataset comprises a subset of 8,857 captioned images from the LVIS dataset. This subset was creating by selecting only image-caption pairs which contain typography that is accurately reflected in the caption.
The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/GPT4V-captions-from-LVIS-typography.JourneyBench_Captioningimage_caption_regularization
Regularization Image Caption Dataset
Number of Images: 1976
Source
This is a subset of tomg-group-umd/pixelprose, converted to .csv format.
Files
people.csv: 1976 images with captions that contain one of these terms: ['person', 'people', 'man', 'men', 'woman', 'women']
newyorker_caption_contest_testdefence-capabilitiesimage_caption_pairs_for_multimodal
