datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
ActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/friedrichor/ActivityNet_Captions.training_setASID-1M
ASID-1M: Attribute-Structured and Quality-Verified Audiovisual Instructions
[🏠 Homepage] [📖 Arxiv Paper] [🤗 Models & Datasets] [💻 Code]
Introduction
We introduce ASID-1M, a large-scale audiovisual instruction dataset built to support universal video understanding with fine-grained, controllable supervision.
Most existing video-instruction data represents complex audiovisual content as a single, monolithic caption. This often leads to incomplete coverage (missing audio… See the full description on the dataset page: https://huggingface.co/datasets/AudioVisual-Caption/ASID-1M.conceptual_captions_jsonVideo-Detailed-Caption
Video Detailed Caption Benchmark
Resources
Website
arXiv: Paper
GitHub: Code
Huggingface: AuroraCap Model
Huggingface: VDC Benchmark
Huggingface: Trainset
Features
Benchmark Collection and Processing
We building VDC upon Panda-70M, Ego4D, Mixkit, Pixabay, and Pexels. Structured detailed captions construction pipeline. We develop a structured detailed captions construction pipeline to generate extra detailed descriptions from various… See the full description on the dataset page: https://huggingface.co/datasets/wchai/Video-Detailed-Caption.Minecraft-Skins-Captioned-1M
Dataset Card for Minecraft Skins
Dataset Summary
This dataset contains 981,079 unique Minecraft player skins collected from various sources. Each skin is stored as a base64-encoded image with a unique identifier.
Dataset Structure
Data Fields
This dataset includes the following fields:
hash: A data dependent hash. These hashes are generated from raw bytes and will be same if the skin is identical.
image: The skin image encoded in base64 format.… See the full description on the dataset page: https://huggingface.co/datasets/neurlang/Minecraft-Skins-Captioned-1M.pexels-568k-internvl2
Dataset Card for pexels-568k-internvl2
Dataset Summary
This is 567,573 synthetic captions for the images found in ptx0/photo-concept-bucket. The captions were produced using OpenGVLab/InternVL2-40B-AWQ. The dataset was grounded for captioning using the tags originally listed.
Languages
The text is in English, but occasionally text in images in other languages is transcribed.
Intended Usage
Training text-to-image models and other machine learning… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/pexels-568k-internvl2.conceptual-captions-cc12m-llavanext
Dataset Card for conceptual-captions-cc12m-llavanext
Dataset Summary
This is a data of 21,930,344 synthetic captions for 10,965,172 images from conceptual_12m. In the interest of reproducibility, an archive found here on Huggingface was used (cc12m-wds). The captions were produced using llama3-llava-next-8b inferenced in float16, followed by cleanup and shortening with Meta-Llama-3-8B.
Languages
The captions are in English.
Data Instances
An… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/conceptual-captions-cc12m-llavanext.icongenai-svg-captions
IconGenAI SVG Captions
Captioned SVG icons from the Iconify corpus, intended for fine-tuning text-to-SVG generation models.
Part of the IconGenAI research project.
Files
Two files are provided at different stages of the processing pipeline:
File
Records
Purpose
icons_captioned_merged.jsonl
275,912
Full license-filtered corpus with VLM-generated captions and collection metadata
icons_training_captioned.jsonl227,821
Quality-filtered, normalised subset… See the full description on the dataset page: https://huggingface.co/datasets/yauheniya-adesso/icongenai-svg-captions.Multi-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.emilia-captions-v3InternVL-SA-1B-Caption
Dataset Card for InternVL-SA-1B-Caption
Overview
The InternVL-SA-1B-Caption Dataset is a bilingual dataset created using the InternVL2-Llama3-76B model. The dataset contains 12 million image-caption pairs in both English and Chinese. All images are sourced from Meta’s SA-1B dataset, and captions were generated using specific prompts designed to minimize hallucinations and ensure accurate descriptions based on visible image content. The dataset is intended for use in tasks… See the full description on the dataset page: https://huggingface.co/datasets/OpenGVLab/InternVL-SA-1B-Caption.coco_captions_quintets
Dataset Card for "coco_captions"
Dataset Summary
COCO is a large-scale object detection, segmentation, and captioning dataset. This repo contains five captions per image; useful for sentence similarity tasks.
Disclaimer: The team releasing COCO did not upload the dataset to the Hub and did not write a dataset card.
These steps were done by the Hugging Face team.
Supported Tasks
Sentence Transformers training; useful for semantic search and sentence… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/coco_captions_quintets.anime-caption-danbooru-2021-sfw-5m-hq
Dataset Card for anime-caption-danbooru-2021-sfw-5m-hq
Dataset Summary
This is 5.71 M captions of 1.43 M images from a safe-for-work (SFW) filtered subset of the Danbooru 2021 dataset. There are 4 captions per image: 1 by CogVLM, 1 by llava-v1.6-34b, 1 llava-v1.6-34b cleaned, and 1 llava-v1.6-34b shortened. See the sections below for how they were generated.
Most captions are substantially larger than 77 tokens and are unsuitable for discrimination using current… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/anime-caption-danbooru-2021-sfw-5m-hq.LiDAR-LLM-Nu-Caption
Dataset Details
Dataset type:
This is the nu-Caption dataset, a QA dataset designed for training MLLM models on caption tasks in autonomous driving scenarios. It is built upon the NuScenes dataset.
Dataset keys:
"answer" is the output of the VLM models using image data. "answer_lidar" uses GPT4O-mini to filter information that cannot be obtained from the image data.
If you want to train the model like LiDAR-LLM, which only uses the LiDAR modality and does not use the vision modality… See the full description on the dataset page: https://huggingface.co/datasets/Senqiao/LiDAR-LLM-Nu-Caption.graph-captioning-train-onlympii-human-pose-captions
Dataset Card for MPII Human Pose Descriptions
Dataset Summary
The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations. These annotations are generated by various state-of-the-art language models (LLMs) and include detailed descriptions of the activities being performed, the count of people present, and their specific poses.
The dataset consists of the same image splits as provided in MMPose, with 14644… See the full description on the dataset page: https://huggingface.co/datasets/saifkhichi96/mpii-human-pose-captions.flickr30k_captions_quintets
Dataset Card for "flickr30k-captions"
Dataset Summary
We propose to use the visual denotations of linguistic expressions (i.e. the set of images they describe) to define novel denotational similarity metrics, which we show to be at least as beneficial as distributional similarities for two tasks that require semantic inference. To compute these denotational similarities, we construct a denotation graph, i.e. a subsumption hierarchy over constituents and their denotations… See the full description on the dataset page: https://huggingface.co/datasets/embedding-data/flickr30k_captions_quintets.ActivityNet_Captions
About
ActivityNet Captions contains 20K long-form videos (180s as average length) from YouTube and 100K captions. Most of the videos contain over 3 annotated events. We follow the existing works to concatenate multiple short temporal descriptions into long sentences and evaluate ‘paragraph-to-video’ retrieval on this benchmark.
We adopt the official split:
Train: 10,009 videos, 10,009 captions (concatenate from 37,421 short captions)
Test (Val1): 4,917 videos, 4,917 captions… See the full description on the dataset page: https://huggingface.co/datasets/Sreevardhan1729/ActivityNet_Captions.social-media-captions
Social Media Captions
Based on the Instagram Influencer Dataset from Seungbae Kim, Jyun-Yu Jiang, and Wei Wang
Extended with photo descriptions of ydshieh/vit-gpt2-coco-en model to create a dataset which can be used to finetune Llama-2.
20k smaller subset: Waterfront/social-media-captions-20k
10k smaller subset: Waterfront/social-media-captions-10k
CSFM-ImageNet1K-Caption
CSFM-ImageNet1K-Caption Dataset
Project Page | Paper | Code
This repository contains dataset associated with the paper "Better Source Better Flow: Learning Condition-Dependent Source Distribution for Flow Matching".
This dataset is used for training and evaluating Condition-dependent Source Flow Matching (CSFM), a framework that learns condition-dependent source distributions for flow matching. We recaptioned the ImageNet-1K dataset using Qwen3-VL-8B Instruct, resulting in detailed… See the full description on the dataset page: https://huggingface.co/datasets/junwann/CSFM-ImageNet1K-Caption.furry-e621-sfw-7m-hq
Dataset Card for furry-e621-sfw-7m-hq
Dataset Summary
This is 6.92 M captions of the images from the safe-for-work (SFW) split of e621 ("e926"). It extends to January 2023, before the widespread advent of machine learning images. It includes captions created by LLMs and a custom multilabel classifier along with CogVLM. There are 8 LLM (mistralai/Mistral-7B-v0.1) and 1 CogVLM (THUDM/CogVLM) captions per image.
Most captions are substantially larger than 77 tokens and are… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/furry-e621-sfw-7m-hq.Harbor-Photo-Captions
Harbor Photo Captions
This collection contains caption records for digitized waterfront photographs.
Rights register
Accession: HP-19
Rights statement: Public Domain Mark 1.0
Depositing archive: Tideglass Image Library
Caption normalization follows the catalog's controlled vocabulary.
SC-Captioner-data
Training and testing annotations for SC-Captioner.
All files are processed into llamafactory data format.
In train_coco6k.json, the "rejected" line means original gpt4 captions in RefinedCaps. They are not used in our self-correction training, but can be used for other purposes.
coco_2017_caption_trainsocial-media-captions-10k
Social Media Captions
Based on the Instagram Influencer Dataset from Seungbae Kim, Jyun-Yu Jiang, and Wei Wang
Extended with photo descriptions of ydshieh/vit-gpt2-coco-en model to create a dataset which can be used to finetune Llama-2.
60k complete dataset: Waterfront/social-media-captions
20k bigger subset: Waterfront/social-media-captions-20k
mistral_captionsramanv-image-captions-6TextOCR-GPT4o
Dataset Card for TextOCR-GPT4o
Dataset Summary
TextOCR-GPT4o is Meta's TextOCR dataset dataset captioned with emphasis on text OCR using GPT4o. To get the image, you will need to agree to their terms of service.
Supported Tasks
The TextOCR-GPT4o dataset is intended for generating benchmarks for comparison of an VLM to GPT4o.
Languages
The caption languages are in English, while various texts in images are in many languages such as Spanish, Japanese… See the full description on the dataset page: https://huggingface.co/datasets/CaptionEmporium/TextOCR-GPT4o.flickr30k-captions-translated-nlThis is a Dutch version of the Flickr30k captions dataset. Which we have auto-translated from English into Dutch using Meta's No Language Left Behind model, specifically the huggingface implementation. For more information about the use of this dataset please refer to the flicker terms of use
