datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Llama3-SSL4EO-S12-v1.1-captions
Llama3-SSL4EO-S12-Captions
The captions are aligned with the SSL4EO-S12 v1.1 dataset and were automatically generated using the Llama3-LLaVA-Next-8B model.
Please find more information regarding the generation and evaluation in the Llama3-MS-CLIP paper.
Code: https://github.com/IBM/MS-CLIP
Data Structure
We provide the captions in two versions: As a single compressed Parquet file per split and as CSV files with 256 captions each that match the Zarr Zip files of the… See the full description on the dataset page: https://huggingface.co/datasets/ibm-esa-geospatial/Llama3-SSL4EO-S12-v1.1-captions.danbooru-multitier-captions-202606
Danbooru — multi-tier captions (202606)
Per-post Danbooru data for the 202606 crawl: native tags, the raw API metadata, model-generated
multi-tier natural-language captions (long / refined long / medium / short), and post flags.
One row per Danbooru post_id. Images are not included — each post is referenced by
post_id, danbooru_url, md5, and the Danbooru CDN URLs. (The two example previews below are
downscaled for illustration.)
Based on:… See the full description on the dataset page: https://huggingface.co/datasets/BootsofLagrangian/danbooru-multitier-captions-202606.conceptual-captions-cc12m-llavanext
Dataset Card for conceptual-captions-cc12m-llavanext
Dataset Summary
This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B.
Languages
The captions are in English.
Data Instances
An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.Tom_and_Jerry_captions3SRC46HSXPXFC63X34QGLDUCIXGZNOXR
icongenai-svg-captions
IconGenAI SVG Captions
Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models.
Part of the IconGenAI research project.
Files
Two files are provided at different stages of the processing pipeline:
File
Records
Purpose
icons_captioned_merged.jsonl
275,912
Full license-filtered corpus with VLM-generated captions and collection metadata
icons_training_captioned.jsonl227,821
Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.laion400m-person-captionsconceptual_captions_3m_zh_tiny_0
Dataset Card for "conceptual_captions_3m_zh_tiny_0"
More Information needed
conceptual_captions_3m_zh_tiny_5
Dataset Card for "conceptual_captions_3m_zh_tiny_5"
More Information needed
mpii-human-pose-captions
Dataset Card for MPII Human Pose Descriptions
Dataset Summary
The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations. These annotations are generated by various state-of-the-art language models (LLMs) and include detailed descriptions of the activities being performed, the count of people present, and their specific poses.
The dataset consists of the same image splits as provided in MMPose, with 14644… See the full description on the dataset page: https://huggingface.co/datasets/saifkhichi96/mpii-human-pose-captions.conceptual_captions_3m_zh_tiny_4
Dataset Card for "conceptual_captions_3m_zh_tiny_4"
More Information needed
conceptual_captions_3m_zh_tiny_2
Dataset Card for "conceptual_captions_3m_zh_tiny_2"
More Information needed
conceptual_captions_3m_zh_tiny_1
Dataset Card for "conceptual_captions_3m_zh_tiny_1"
More Information needed
howto100m_captions_with_verb_nounsconceptual_captions_3m_zh_tiny_3
Dataset Card for "conceptual_captions_3m_zh_tiny_3"
More Information needed
pexels-people-captions
Pexels people captions
37,412 photographs of people, each with a description written for it in
Chinese or English. The photographs themselves are not included: every row
carries a link to the original on Pexels instead.
What a row holds
column
meaning
image_id
stable identifier used across the corpus, pexels-<photo id>
photo_id
the Pexels photo id
page_url
the photo's page on pexels.com
image_url
the original image file as the API reports it… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/pexels-people-captions.ffhq_with_llava_shorter_captions_flux_latentsvintage-photography-450k-high-quality-captionsThis is a 450k image datastet focused on photography from the 20th century, and their analog aspect. Many of the images are in high resolution. This dataset currently has 20k images captioned with InternVL2 26B, and is a work in progress (I plan to caption the entire dataset and also have short captions for all of the images, compute is an issue for now).
vg-captions-graphs-processed-image-graphsActivityNet-Captions
ActivityNet-Captions
This dataset repo contains a curated subset of the ActivityNet-Captions dataset.
Filtering Logic {Video length ≤ 10 seconds:KeepVideo length > 10 seconds:Discard\text{Filtering Logic }\begin{cases}
\text{Video length }\leq \text{ 10 seconds:} & \text{Keep} \\
\text{Video length } > \text{ 10 seconds:} & \text{Discard}
\end{cases}Filtering Logic {Video length ≤ 10 seconds:Video length > 10 seconds:KeepDiscard
Analysis
There are 10,759 unique… See the full description on the dataset page: https://huggingface.co/datasets/qingy2024/ActivityNet-Captions.BToks-ActivityNet-Captionsleetcode_with_youtube_captionscharacter-captions-opusDeduplicated set of character portraits that have been described by Anthropic Claude Opus as characters with stories and visual attributes.
Images obtained from CivitAI by filtering for SD XL-derived models only. Original Stable Diffusion prompt and metadata is also included.
Each image is a portrait, meaning it's taller than it's wider, and has exactly one face in it. Face bounding boxes are provided.
Character-like description for each image is given by Claude Opus. Here is an example:
{… See the full description on the dataset page: https://huggingface.co/datasets/kubernetes-bad/character-captions-opus.florence2-ofa-captions-500
OFA Florence-2 Dataset (500 Samples)
This dataset was generated using microsoft/Florence-2-large on a subset of COCO 2017 Validation images.
It is pre-formatted for OFA Stage-1 fine-tuning (Headerless TSV, URL-safe base64, max 512x512 resolution).
vintage-photography-captions
Dataset Card for Vintage Photograph Captions Recaption
This dataset contains 445,271 recaptioned vintage photographs, derived from the vintage-photography-450k-high-quality-captions dataset. It provides high-quality bilingual (English and Chinese) captions, aesthetic scores, and other metadata generated using the Qwen2-VL model.
This dataset is a recaptioned version of SilentAntagonist/vintage-photography-450k-high-quality-captions. The original dataset contained 456,006… See the full description on the dataset page: https://huggingface.co/datasets/kaupane/vintage-photography-captions.datacomp_small_english_captions
Dataset Card for "datacomp_small_english_captions"
More Information needed
wit-captionsDataset derived from the original WIT dataset, with the following changes:
Removed all columns except for 'image_url', 'caption_reference_description', 'caption_attribution_description', 'mime_type', 'original_height', 'original_width'
All rows without either a 'caption_reference_description' or 'caption_attribution_description' have been removed
The image references were deduplicated on the image_url preserving the entry with the longest caption_reference_description
A new 'text' column was… See the full description on the dataset page: https://huggingface.co/datasets/fschieber/wit-captions.coco-transformed-captionslaion2b-aesthetic-squareish-captionsThis dataset contains image captions generated from LAION2B-en-aesthetic-square.
We started with ~300K images after size filtering (2.5k max w/h), a portion of the images were skipped due to inaccessible URLs.
The captions were generated over ~30 hours using Qwen3-VL-30B-A3B-Instruct on 1xH100 running SGLang with the prompt Describe the content of the provided image in detail, in plaintext. Do not make assumptions. Do not use special formatting. Avoid purple prose.
Total samples: 209,141… See the full description on the dataset page: https://huggingface.co/datasets/opendiffusionai/laion2b-aesthetic-squareish-captions.datacomp_small_english_captions_without_weird_characters
Dataset Card for "datacomp_small_english_captions_without_weird_characters"
More Information needed
datacomp_small_french_captions
Dataset Card for "datacomp_small_french_captions"
More Information needed
