datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.conceptual-captions-12This file contains English captions from Conceptual 12M dataset by Google. Since we don't own the images, we have provided the link to images, name of downloaded file, and caption for that image in the TSV file.
We would like to thank Luke Melas for helping us get the cleaned CC-12M data on our TPU-VMs.
human-templated-captions-1bcsv delimiter is = ".,|,."
apparently python doesn't like multichar delimiters using the native csv so there's some issues with environments when loading.
This seemed like a good idea to avoid overlapping potential characters, but in practice it turned into additional overhead and bugs. I'll be manually converting the split to parquet and providing a proper file split soon.
Additionally with the parquet will introduce the large caption split; which are considerably longer captions for the… See the full description on the dataset page: https://huggingface.co/datasets/AbstractPhil/human-templated-captions-1b.DataComp_large_pool_BLIP2_captions
Dataset Card for DataComp_large_pool_BLIP2_captions
Dataset Summary
Supported Tasks and Leaderboards
We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp.
Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work.
Languages
Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_large_pool_BLIP2_captions.dreamlip_long_captions
Dataset Card for DreamLIP-30M
Dataset Summary
DreamLIP-Long-Captions is a dataset consisting of ~30M image annotations, i.e. detailed long captions. In contrast with the curated style of other synthetic image caption annotations, DreamLIP-30M utilizes pre-trained Multi-modality Large Language Model to obtain detailed descriptions with an average length of 247. More precisely, the detailed descriptions are generated by asking the ShareGPT4V/InstructBLIP/LLava1.5 the… See the full description on the dataset page: https://huggingface.co/datasets/qidouxiong619/dreamlip_long_captions.DataComp_medium_pool_BLIP2_captions
Dataset Card for DataComp_medium_pool_BLIP2_captions
Dataset Summary
Supported Tasks and Leaderboards
We have used this dataset for pre-training CLIP models and found that it rivals or outperforms models trained on raw web captions on average across the 38 evaluation tasks proposed by DataComp.
Refer to the DataComp leaderboard (https://www.datacomp.ai/leaderboard.html) for the top baselines uncovered in our work.
Languages
Primarily English.… See the full description on the dataset page: https://huggingface.co/datasets/thaottn/DataComp_medium_pool_BLIP2_captions.character-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes.
Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by Claude Opus. Here is an example:
{… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.florence2-ofa-captions-500
OFA Florence-2 Dataset (500 Samples)
This dataset was generated using microsoft/Florence-2-large on a subset of COCO 2017 Validation images.
It is pre-formatted for OFA Stage-1 fine-tuning (Headerless TSV, URL-safe base64, max 512x512 resolution).
simple-image-captionsnepali-meme-captions
NeMeme-CAP: Nepali Meme Captions
Dataset Summary
English-language captions generated by Google Gemini for the CHiPSAL 2026 SubtaskA Nepali Meme Datset.
The context-aware captions was generated accross the training, validation, and test splits.
Supported Tasks
Hateful Meme Classification: Predict whether the meme is non-hateful (label=0) and hateful (label=1).
Multimodal Meme Understanding: Useful as auxiliary text features or as ground-truth explanations for… See the full description on the dataset page: https://huggingface.co/datasets/Anish/nepali-meme-captions.PixLore-Rich-CaptionsRich image captioning dataset used for training PixLore model: https://arxiv.org/abs/2312.05349
"image_path" contains the path to the COCO dataset image (change the path accordingly),
"rich_caption" contains the rich caption created using the technique described in the paper.
The rest of the columns are used for debugging or improving the prompt.
imagenet1k_captions_minigpt4
ImageNet1k Captions Generated with MiniGPT-4
MiniGPT-4 captions generated for ImageNet1k images. Can be used for training/finetuning diffusion models for image generation.
ImageNet1k: link
MiniGPT-4: link
GPT4V-captions-from-LVIS-typography
GPT4V-captions-from-LVIS-typography
by: Peter Bevan, 21 March 2023
This dataset is a typography subset of 220k-GPT4Vision-captions-from-LIVIS.
This dataset comprises a subset of 8,857 captioned images from the LVIS dataset. This subset was creating by selecting only image-caption pairs which contain typography that is accurately reflected in the caption.
The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted… See the full description on the dataset page: https://huggingface.co/datasets/pbevan11/GPT4V-captions-from-LVIS-typography.t2i-diversity-gender-neutral-captionsThis dataset contains different synthetic captions for our image samples.
We have selected the best-performing caption set from our experiments, the random-length captions. Then, we have used Gemma-2-9b-it and instructed it to remove different genders from the captions. We obtained three sets from the original set, namly (i) all genders neutralized, (ii) only female gender neutralized, and (iii) only male gender neutralized. To this end, we have removed all gender indicative words such as… See the full description on the dataset page: https://huggingface.co/datasets/AIML-TUDA/t2i-diversity-gender-neutral-captions.3D-DST-captions
3D-DST-captions
As part of our data release in 3D-DST, we present MiniGPT4-generated captions for all 1000 classes in ImageNet-1k.
See wufeim/DST3D for synthetic data generation with 3D annotations using the captions here.
These captions can be used to produce other synthetic datasets for fair comparisons between different data generation procedures.
wikiart-captions-5k
WikiArt Captions — 5k subset
image_row — индекс в 5k-сабсете
caption — автоописание (nlpconnect/vit-gpt2-image-captioning)
Usage
from datasets import load_dataset
ds = load_dataset("USERNAME/wikiart-captions-5k", data_files="data/train/captions.csv")
ShapeNetCore_CaptionsThis ShapeNetCore_Captions dataset contains text captions of all the 52,472 models of the ShapeNetCore v2 dataset. As these were generated by LLMs, there is a chance that these captions will comprise slightly incorrect phrases or information. If something incorrect is discovered during manual inspection, kindly make an effort to share your findings with me.
unsafe-safe-captions-3.0
Safety Caption Pairs Deduplicated
Filtered using n-gram and word-diff rules.
Contains only pairs with 1–2 word changes and minimal n-gram overlap.
Sydney_captionswikiart-captions_5000wikiart-captions_81444safety-image-captions-1wikiart-captionswikiart_captionsPokemon_captionswikiart_captionswikiart_w_captionswikiart-blip-captionswikiart_captions
WikiArt Captions Subset — Multimodal Art Retrieval Dataset
This dataset is a curated subset of 6,000 paintings from the WikiArt collection.It was created as part of a project on multimodal art retrieval, combining visual, textual, and semantic information.
Each record represents one artwork and includes:
Field
Description
image_row
Row index in the source subset (integer)
caption
Automatically generated textual description (caption) using the BLIP model… See the full description on the dataset page: https://huggingface.co/datasets/Lizagrin/wikiart_captions.
