datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
newyorker_caption_contest
Dataset Card for New Yorker Caption Contest Benchmarks
Dataset Summary
See capcon.dev for more!
Data from:
Do Androids Laugh at Electric Sheep? Humor "Understanding" Benchmarks from The New Yorker Caption Contest
@inproceedings{hessel2023androids,
title={Do Androids Laugh at Electric Sheep? {Humor} ``Understanding''
Benchmarks from {The New Yorker Caption Contest}},
author={Hessel, Jack and Marasovi{\'c}, Ana and Hwang, Jena D. and Lee, Lillian
and… See the full description on the dataset page: https://huggingface.co/datasets/jmhessel/newyorker_caption_contest.danbooru-1024-eq-captioned
Danbooru 1024 e/q Captioned Dataset
59,495 high-resolution (1024px) anime-style images from Danbooru's explicit and questionable rated pools. Each image includes comprehensive JSON captions generated via MiniMax-M3 with structured per-character state-of-dress inventories, camera notes, mood palettes, and post-processing detections.
Directory Structure
danbooru-1024-eq-captioned.parquet <- consolidated metadata manifest
originals/ <-… See the full description on the dataset page: https://huggingface.co/datasets/quarterturn/danbooru-1024-eq-captioned.conceptual_captions
Dataset Card for Conceptual Captions
Dataset Summary
Conceptual Captions is a dataset consisting of ~3.3M images annotated with captions. In contrast with the curated style of other image caption annotations, Conceptual Caption images and their raw descriptions are harvested from the web, and therefore represent a wider variety of styles. More precisely, the raw descriptions are harvested from the Alt-text HTML attribute associated with web images. To arrive at the… See the full description on the dataset page: https://huggingface.co/datasets/google-research-datasets/conceptual_captions.conceptual-captions-12m-webdatasetBLIP3o-Pretrain-Long-Caption
BLIP3o Pretrain Long-Caption Dataset
This collection contains 27 million images, each paired with a long (~120 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct.
Download
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="BLIP3o/BLIP3o-Pretrain-Long-Caption",
repo_type="dataset"
)
Load Dataset without Extracting
You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Long-Caption.pokemon-blip-captions
Notice of DMCA Takedown Action
We have received a DMCA takedown notice from The Pokémon Company International, Inc.
In response to this action, we have taken down the dataset.
We appreciate your understanding.
BLIP3o-Pretrain-Short-Caption
BLIP3o Pretrain Short-Caption Dataset
This collection contains 5 million images, each paired with a short (~20 token) caption generated by Qwen/Qwen2.5-VL-7B-Instruct.
Download
from huggingface_hub import snapshot_download
snapshot_download(
repo_id="BLIP3o/BLIP3o-Pretrain-Short-Caption",
repo_type="dataset"
)
Load Dataset without Extracting
You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead:
from datasets import… See the full description on the dataset page: https://huggingface.co/datasets/BLIP3o/BLIP3o-Pretrain-Short-Caption.COCO-Caption
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of COCO-Caption-2014-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@misc{lin2015microsoft,
title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption.coco_captions
Dataset Card for "coco_captions"
More Information needed
COCO-Caption2017
Large-scale Multi-modality Models Evaluation Suite
Accelerating the development of large-scale multi-modality models (LMMs) with lmms-eval
🏠 Homepage | 📚 Documentation | 🤗 Huggingface Datasets
This Dataset
This is a formatted version of COCO-Caption-2017-version. It is used in our lmms-eval pipeline to allow for one-click evaluations of large multi-modality models.
@misc{lin2015microsoft,
title={Microsoft COCO: Common Objects in Context}… See the full description on the dataset page: https://huggingface.co/datasets/lmms-lab-encoder/COCO-Caption2017.ffhq512-captionlaion2b-en-a65_cogvlm2-4bit_captions
Abstract
This dataset contains image captions for the laion2B-en aesthetics>=6.5 image dataset using CogVLM2-4bit with the "laion-pop"-prompt to generate captions which were "likely" used in Stable Diffusion 3 training. From these image captions new synthetic images were generated using stable-diffusion-3-medium (batch-size=8).
The synthetic images are best viewed locally by cloning this repo with:
git lfs install
git clone… See the full description on the dataset page: https://huggingface.co/datasets/GeroldMeisinger/laion2b-en-a65_cogvlm2-4bit_captions.synthetic-dataset-1m-dalle3-high-quality-captions
Dataset Card for Dalle3 1 Million+ High Quality Captions
Alt name: Human Preference Synthetic Dataset
Example grids for landscapes, cats, creatures, and fantasy are also available.
Description:
This dataset comprises of AI-generated images sourced from various websites and individuals, primarily focusing on Dalle 3 content, along with contributions from other AI systems of sufficient quality like Stable Diffusion and Midjourney (MJ v5 and above). As users typically… See the full description on the dataset page: https://huggingface.co/datasets/ProGamerGov/synthetic-dataset-1m-dalle3-high-quality-captions.Gradients_Gradients_and_Text_Full_Logic_CaptionsCaptioned_COCOStuffCaptionQA
📌 CaptionQA Benchmark
A high-density, taxonomy-grounded benchmark for evaluating image caption quality and the alignment between image information and generated captions
📄 Paper: CaptionQA: Is Your Caption as Useful as the Image Itself? 📦 Evaluation Code: GitHub Repository
Sample Usage
You can load the dataset using the Hugging Face datasets library:
from datasets import load_dataset
# Load the entire dataset
dataset = load_dataset("Borise/CaptionQA")
# Load a… See the full description on the dataset page: https://huggingface.co/datasets/Borise/CaptionQA.newyorker_caption_ranking
New Yorker Caption Ranking Dataset
Dataset Descriptions
Homepage: https://nextml.github.io/caption-contest-data/
Repository: https://github.com/yguooo/cartoon-caption-generation
Paper: Humor in AI: Massive Scale Crowd-Sourced Preferences and Benchmarks for Cartoon Captioning
Point of Contact: yguo@cs.wisc.edu
Dataset Summary
We present a novel multimodal preference dataset for creative tasks, consisting of over 250 million human ratings on more than 2.2… See the full description on the dataset page: https://huggingface.co/datasets/yguooo/newyorker_caption_ranking.COCO_captions_train
Dataset Card for "COCO_captions_train"
More Information needed
InternVL-SA1B-Caption-WebDatasetThis repo contains the recaptioned SA1B images in webdataset format. The recaptioned prompts are from https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption
yoimiya-hd-txt-captioned
Yoimiya HD Captioned Dataset
A dataset of 877 high-definition anime illustrations of Yoimiya from Genshin Impact, with detailed text captions describing poses, clothing, expressions, and scene details.
Dataset Structure
train/ — Contains all images and metadata.jsonl
dataset_infos_manual.yaml — Configuration for the HuggingFace dataset viewer
pokemon-gpt4-captions
Dataset Card for "pokemon-gpt4-captions"
This dataset is just lambdalabs/pokemon-blip-captions but the captions come from GPT-4 (Turbo).
Code used to generate the captions:
import base64
from io import BytesIO
import requests
from PIL import Image
def encode_image(image):
buffered = BytesIO()
image.save(buffered, format="JPEG")
img_str = base64.b64encode(buffered.getvalue())
returnimg_str.decode("utf-8")
def create_payload(image_string):
payload = {… See the full description on the dataset page: https://huggingface.co/datasets/diffusers/pokemon-gpt4-captions.220k-GPT4Vision-captions-from-LIVIS
220k-GPT4Vision-captions-from-LVIS
by: Christoph Schuhmann, Peter Bevan, 21 Nov, 2023
This dataset comprises 220,000 captioned images from the LVIS dataset. The captions were generated by summarising the LVIS-Instruct4V dataset released by X2FD. The instructions are converted into captions using Mistral-7B-OpenOrca.
PROMPT
"""<<SYS>> You are a highly intelligent, empathic, helpful, respectful, and honest assistant with high emotional intelligence. Always… See the full description on the dataset page: https://huggingface.co/datasets/laion/220k-GPT4Vision-captions-from-LIVIS.GBC10M
Graph-based captioning (GBC) is a new image annotation paradigm that combines the strengths of long captions, region captions, and scene graphs
GBC interconnects region captions to create a unified description akin to a long caption, while also providing structural information similar to scene graphs.
** The associated data point can be found at demo/water_tower.json
Description and data format
The GBC10M dataset, derived from the original images in CC12M, is… See the full description on the dataset page: https://huggingface.co/datasets/graph-based-captions/GBC10M.anime-with-gpt4v-caption-for-lora
Anime style image - text by GPT4V small dataset
The text is as follows:
This is a charming anime-style illustration featuring a young girl as the main subject. The image predominantly uses a soft, pastel color palette, creating a gentle and whimsical ambiance. The main character has light blonde hair styled in two low twintails, secured with what could be interpreted as dark-colored hair ties or ribbons. She has large expressive blue eyes and a demure expression, with… See the full description on the dataset page: https://huggingface.co/datasets/alfredplpl/anime-with-gpt4v-caption-for-lora.midjourney-niji-1m-llavanext
Dataset Card for midjourney-niji-1m-llavanext
Dataset Summary
This is a dataset of 2,079,886 synthetic captions for 1,039,943 images from midjourney-v6-520k-raw and nijijourney-v6-520k-raw. The captions were produced using https://huggingface.co/lmms-lab/llama3-llava-next-8b inferenced in float16 after tags were generated with wd-swinv2-tagger-v3, followed by cleanup and shortening with Meta-Llama-3-8B.
All images with metadata are available as MozJPEG encoded JPEGs… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/midjourney-niji-1m-llavanext.anime-captionsSTAIR-Captions
Dataset Card for STAIR-Captions
Dataset Summary
STAIR Captions is a large-scale dataset containing 820,310 Japanese captions. This dataset can be used for caption generation, multimodal retrieval, and image generation.
Supported Tasks and Leaderboards
[More Information Needed]
Languages
The language data in JDocQA is in Japanese (BCP-47 ja-JP).
Dataset Structure
Data Instances
[More Information Needed]… See the full description on the dataset page: https://huggingface.co/datasets/shunk031/STAIR-Captions.COCO_captions_validation
Dataset Card for "COCO_captions_validation"
More Information needed
conceptual_captions_jsonscientific-figures-captions-context
Dataset Card for Scientific Figures, Captions, and Context
A novel vision-language dataset of scientific figures taken directly from research papers.
We scraped approximately ~150k papers, with about ~690k figures total. We extracted each figure's caption and label from the paper. In addition, we searched through each paper to find references of each figure and included the surrounding text as 'context' for this figure.
All figures were taken from arXiv research papers.… See the full description on the dataset page: https://huggingface.co/datasets/mawadalla/scientific-figures-captions-context.
