datasets
Training and evaluation data, with the modality, task and licence stated up front. Listed live from the Hugging Face Hub.
Embodied-Captioning
Embodied Image Captioning – Manually Annotated Test Set
Paper: Embodied Image Captioning: Self-supervised Learning Agents for Spatially Coherent Image Descriptions (ICCV 2025)Authors: Tommaso Galliena, Tommaso Apicella, Stefano Rosa, Pietro Morerio, Alessio Del Bue, Lorenzo NataleAffiliations: Italian Institute of Technology (IIT), University of GenoaProject Website: https://hsp-iit.github.io/embodied-captioningCode: https://github.com/hsp-iit/embodied-captioning
📦… See the full description on the dataset page: https://huggingface.co/datasets/TommyBsk/Embodied-Captioning.nordjylland-news-image-captioning
Dataset Card for "nordjylland-news-image-captioning"
Dataset Summary
This dataset is a collection of image-caption pairs from the Danish newspaper TV2 Nord.
Supported Tasks and Leaderboards
Image captioning is the intended task for this dataset. No leaderboard is active at this point.
Languages
The dataset is available in Danish (da).
Dataset Structure
An example from the dataset looks as follows.
{
"file_name": "1.jpg",
"caption":… See the full description on the dataset page: https://huggingface.co/datasets/alexandrainst/nordjylland-news-image-captioning.Chinese_Children_Image_Captioning_Dataset_Split0
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split0.IntraOral_Gingivitis_Image_Captioning
A DENTAL INTRAORAL IMAGE DATASET OF GINGIVITIS FOR IMAGE CAPTIONING
Dataset Description
This dataset is a copy of A Dental IntraOral Image Dataset of Gingivitis for Image Captioning which is shared with the license CC BY 4.0.
This dataset contains 1,096 samples organized across multiple splits.
The dataset includes image data.
Splits
train: 732 samples
test: 182 samples
validation: 182 samples
Dataset Creation
This dataset was created using… See the full description on the dataset page: https://huggingface.co/datasets/ekacare/IntraOral_Gingivitis_Image_Captioning.image-captioning-turkish
Türkçe Image Captioning Veri Seti
Bu veri seti BLIP3o modelinin pretrain eğitiminde kullanılan BLIP3o-Pretrain-Long-Caption ve BLIP3o-Pretrain-Short-Caption veri setlerinin Türkçeye çevirilmiş bir alt parçasıdır. Orijinal veri setinin oluşturulması ile ilgili detaylı bilgiye BLIP-3o makalesi üzerinden ulaşabilirsiniz.
Veri seti Image-to-Text modellerinin eğitilmesinde veya ince ayar sürecinde kullanılabilir. Veri seti, orijinal veri setinin lisansı olan Apache 2.0 altında… See the full description on the dataset page: https://huggingface.co/datasets/ituperceptron/image-captioning-turkish.coco_captioning_complete_formatcoco2017-captioningChinese_Children_Image_Captioning_Dataset_Split1
CODP-1200:Children Oral Description of Picture(Chinese-Child-Captions)
CODP-1200: An AIGC based benchmark for assisting in child language acquisition
数据集介绍
目前已知最大的儿童图像描述数据集,children image captioning
共有1200张图片
每张图片对应五个中文描述,每两张图片为一组
描述文字600*5=3000
如果使用CODP-1200数据集,请引用以下文章
@article{LENG2024102627,
title = {CODP-1200: An AIGC based benchmark for assisting in child language acquisition},
journal = {Displays},
volume = {82},
pages = {102627},
year =… See the full description on the dataset page: https://huggingface.co/datasets/svjack/Chinese_Children_Image_Captioning_Dataset_Split1.Closed_Captioning_Lecture_DatasetMulti-Source-Video-Captioning
Multi-source Video Captioning (MSVC) Dataset Card
Dataset details
Dataset type:
MSVC is a set of collected video captioning data. It is constructed to ensure a robust and thorough evaluation of Video-LLMs' video-captioning capabilities.
Dataset detail:
MSVC is introduced to address limitations in existing video caption benchmarks, MSVC samples a total of 1,500 videos with human-annotated captions from MSVD, MSRVTT, and VATEX, ensuring diverse scenarios and domains.… See the full description on the dataset page: https://huggingface.co/datasets/DAMO-NLP-SG/Multi-Source-Video-Captioning.tacos-captioningPokemon-Captioning-Classification
2000+ download monthly. Really appreciate for all of you guys:
Buy me a coffee:
https://buymeacoffee.com/tridoan
Disclaimer: This model is provided "as-is" without any warranties. The authors are not responsible for any misuse or damages arising from its use.
COCO-Image-Captioningjoy-captioning-20250328b
Work In Progress
I'm still going back through my data to add in the URLs.
joy-captioning-20250408aThis is the dataset used to do the initial training for JoyCaption Beta One (https://huggingface.co/fancyfeast/llama-joycaption-beta-one-hf-llava), before post-training.
Contents
Most of the dataset focusses on descriptions and captions for images, with a smaller subset covering general VQA tasks.
Some of the questions and answers are human written, some are automated, some are machine written. The is_human column is True when the answer text is human written.
WARNING… See the full description on the dataset page: https://huggingface.co/datasets/fancyfeast/joy-captioning-20250408a.ru_image_captioningmy_image_captioning_datasetgraph-captioning-train-onlyMSVD-Video-Captioning-Vi
MSVD-Video-Captioning-Vi
📌 Overview
MSVD-Video-Captioning-Vi is a Vietnamese video captioning dataset derived from the MSVD dataset originally hosted by the user friedrichor on Hugging Face.
This dataset provides Vietnamese captions for short video clips and is intended for:
Video captioning research
Vision–Language model training
Multimodal instruction tuning
Video-to-text generation
🔁 Dataset Origin
This dataset is a translated and derived version… See the full description on the dataset page: https://huggingface.co/datasets/NTQAI/MSVD-Video-Captioning-Vi.Bean_Captioning_DatasetMega-Fast-KNN-Captioningucf101-captioned-mappednordjylland-news-image-captioningMSEarth_Captioningchart_captioning
Dataset Card for "chart_captioning"
More Information needed
coco_captioningfrom COCO val2014
image-captioning-idECG-CaptioningPersian-Image-Captioning
Dataset Card for "Persian-Image-Captioning"
More Information needed
X-ray_Image_captioning
